
We expected the difficult questions to be about the artificial intelligence (AI). When we brought together experts from policing, the AI industry, law, psychology and linguistics to pressure-test how AI evidence might be presented in court, we assumed the most searching questions would centre on the algorithms themselves — how they work, whether they can be trusted, how to make them transparent. Some of those issues did arise. But the insight that surfaced most often, was almost the opposite: the AI is rarely the real problem.

This emerged from a World Café (a simple, effective, and flexible format for hosting large group dialogue) held as part of PROBabLE Futures, a Responsible AI UK (RAi UK) Keystone project. We are designing two mock trials to explore AI evidence, including its explainability, in criminal trials, and we wanted expert scrutiny of our scenarios before running them. What our contributors told us has helped shape how we think about AI evidence, and it raises questions that reach beyond our own project.
The AI is rarely a standalone issue
AI evidence is rarely a standalone issue. A case can be weakened by the process surrounding a tool — how evidence was gathered, who oversaw the investigation, what was disclosed — rather than by the technology itself. These are not problems unique to AI; they can affect any piece of evidence.
That carries a risk at trial because attention naturally gravitates to the novel element. In this case the AI tool and how it reached its output is the novel element, whereas the actual weakness may lie in the human processes that acted on it. A jury focused on scrutinising an AI tool might overlook that the investigation itself was thin. The concern here is not that AI props up a weak case, but that arguing over the technology can draw attention away from other questions a trial may need to test.
Explainability could make or break a tool
When the AI is itself a pivotal piece of evidence, whether it is believed may rest on a single quality: can it be explained to a layperson. Several experts felt that whether a tool is judged credible will hinge less on its technical sophistication than on whether a jury, a judge, the lawyers and victim can understand what it did and how far to trust it.
We ask you to consider explaining how a large language model automated a witness statement or summarised a case file, and how reliable that is, to twelve people with no technical background and their own preconceptions about what “AI” means. That is a communicative challenge as much as a technical one. Explainability, the World Café discussion suggested, is where these tools will stand or fall when they reach court. It is not something that can be added as an afterthought.
Not all AI is treated equally
There is no single verdict on “AI in court.” Perceptions of credibility varied notably depending on the tool and how it was used. Contributors appeared more comfortable with something like AI audio transcription than with object recognition built from a witness description. Transcription works from a fixed recording that can be checked against the original, whereas an object identification that begins with a subjective witness account inherits that subjectivity which, in turn, makes it more open to challenge under cross-examination. Tools used to automate a routine task felt lower risk than tools whose outputs shape investigative decisions.
The implication is that each tool, and each use, will be judged on its own terms. Juries are likely to accept some far more readily than others, and the same technology may be reasonable in one role and highly contestable in another. Where in the process a tool sits, whether it quietly assists an investigation or stands as a piece of evidence, matters considerably.
Whose job is it to explain?
One of the most practical challenges came from policing. If a force uses facial recognition, it should already hold a clear policy and a verification-and-testing record that is ready to produce rather than hastily assembled in response to questions at trial. This points less to the expert’s role than to any given force’s own documentation. Ideally, how a tool works is a matter for an expert, while how a particular force used it should be recorded as routine due diligence, long before anyone reaches a courtroom.
The World Café also prompted useful discussion about who a suitable expert witness might be. The people who operate these systems often press a button and receive a result; explaining the mechanism behind it is a different skill entirely. If AI evidence is to be tested fairly, the courtroom will need people who can explain the technology in terms a jury can follow.
Juries will fill in the gaps
A final observation concerned how juries respond to what is left unsaid. Where there is a perceived gap such as a missing identification step, or no camera footage where you might expect it, jurors may construct their own explanation for it. No case ever contains every possible piece of evidence, so some gaps are inevitable. However, where AI evidence is concerned, it helps if the surrounding account is coherent and any obvious absences are explained, so the technology is not judged against gaps that have nothing to do with it. Gaps in an investigation then can strengthen or weaken how the evidence is received and are something to be mindful of rather than attempt to eliminate entirely.

Why it matters
AI is already entering UK policing and courts, largely ahead of the research and evaluation base needed to guide it. What this World Café made clear is that presenting AI evidence in a criminal trial raises not only technical challenges but procedural and communicative ones. Whether AI evidence can ultimately support fair outcomes is a question our mock trials are designed to test; what the café helped us see is that the questions likely to matter most may be less “is the algorithm accurate?” and more “was the process sound, is the documentation there, and can any of it be explained to the people who must weigh it?”
These insights are now feeding directly into the design of our mock trials, which run at the end of this year and into next. We will be testing how AI evidence holds up under genuine adversarial pressure, and whether explanations that sound convincing in a research lab survive cross-examination. You can follow the work through the PROBabLE Futures website, and we would welcome your thoughts, particularly from anyone working at a meeting point of AI, policing and the law.
By Katherine Mary Jones, Marion Oswald, Angela Paul, Lizzie Tiarks, Carole McCartney, Kyle Montague and Kyriakos N. Kotsoglou.

This blog was compiled and authored by members of the PROBabLE Futures project team. It reports issues and insights generated by human participants and project members during the World Café discussion; all issues and insights are human-generated. Claude (an AI assistant developed by Anthropic) was used to support drafting and expression, working from the team’s notes and initial drafts. The final text has been reviewed, edited and approved by the project team, who are responsible for its content and accuracy.
Join us as an official partner if your organisation is interested in Responsible AI research, innovation and skills. We are looking for partners including businesses, funders, charities, creative organisations, industry associations, think tanks and more. Find out more and sign up
RAi UK will store your data in accordance with the General Data Protection Regulation 2017 (GDPR). We will not share your data with any third parties and you are given opportunities to unsubscribe at any time within the electronic communications you receive, or by emailing info@rai.ac.uk