Mock Hot-Tub Cross Examination LIVE Q&A Session

Watch this symposium webinar on demand

Watch time:

min

In this expert Q&A session, we explore the mock hot-tub cross-examination technique and its impact on courtroom strategy. Leading litigation professionals share practical insights and answer common questions on concurrent evidence and expert testimony.
This webinar will cover
• Item
• Item
• Item
• Item
Live question and answers

Question

“Are those hallucinations just random errors, or are they more likely to occur in specific categories, such as lesser-known individuals, or perhaps niche allegations?”

Answer

They said it’s both: hallucinations can arise in patterned ways depending on the information landscape and how the model is constrained. LLMs behave as statistical next-token predictors, and hallucinations are more likely when there is too little information (long-tail/sparse topics, where the model fills gaps unless instructed not to) or too much / too common information (e.g., common names where identities can be conflated). They emphasised that users often can’t tell what drove the hallucination from the outside, so clear prompting and constraints are important because the model will otherwise tend to produce confident answers that “make you happy.”

Question

“Regarding hallucinations and filtering, when post-training filters are implemented. Do they reduce hallucinations, perhaps evenly across those categories, or perhaps certain content types like you mentioned, there may be reputational allegations. Do you think they would be more susceptible despite those layers?” 

Answer

They described filtering as imperfect and uneven: it can help in some cases but is “hit and miss,” especially where identity is ambiguous (e.g., filtering “bad actors” risks filtering out the exact person you’re asking about). They drew a distinction between post-generation filters—which are good at blocking obviously harmful/toxic or policy-violating content and enforcing style constraints—and factual reliability, where those filters struggle because they don’t independently verify truth and can miss fluent, benign-sounding misstatements (particularly in domains like law and medicine). For high-stakes corporate settings, they said stronger control comes from system design and prompting—e.g., instructing the agent not to make things up and, in some cases, limiting or avoiding automatic internet RAG and instead curating external content. 

Question

“In practical terms, how does RAG, alter responsibility? Does it retrieve and summarize those existing contents, or can it completely just generate entirely new associations?” 

Answer

They said RAG is intended to shift a model from free association toward grounded generation by retrieving relevant material, but it doesn’t eliminate responsibility or prevent the model from forming new (and sometimes wrong) associations, because the generation step can still synthesize beyond what is explicitly in the sources. Responsibility becomes “re-bundled” across the system: retrieval determines what evidence is surfaced and generation determines how it is woven into an answer, so governance must cover training/data, retrieval indexes, chunking and re-ranking. They also noted implementations differ by model—some pre-filter/pre-process retrieved results while others pass them through unfiltered—affecting hallucination risk. They added that RAG design choices (chunking strategy, retrieval structure) can materially affect cost and efficiency at scale, with poor approaches multiplying compute/token spend. 

Question

“Regarding the hot tub process side of things, are there any techniques or approaches you use as experts to address or mitigate the courts’ potential perception that you might be acting as a gun for hire, particularly in situations where very established experts in the same field, such as both of you in AI, appear to reach opposing conclusions.” 

Answer

They emphasised that in real hot tub practice, experts are officers of the court and their primary duty must clearly come through in their answers and courtroom demeanour; the hot tub typically follows an exchange of reports and unresolved differences before the court puts joint questions to experts. They also noted that courts can be wary of what appear to be “extreme” conclusions, even if an experienced expert can diagnose problems quickly, so credibility depends on showing a transparent, methodical basis—walking the judge through the frameworks, evidence and steps that support the opinion, rather than asserting conclusions. In short: stay anchored to the court duty, avoid performative advocacy, and justify conclusions in a way the judge can follow. 

Question

“How much does it actually help? If, say, for example, an AI-generated statement were challenged in court, what level of reconstruction is realistically possible to ground the expert opinion in objective system evidence?

Answer

They said that with good observability, engineers can often reconstruct key external facts—what prompt was used, what retrieval content was accessed, which model/version ran, and associated session/request IDs and token-level metadata—but AI introduces steps that make it hard to reconstruct the internal reasoning process. In public LLMs, hidden activations and specific training examples generally aren’t traceable (at best, post-hoc methods may approximate influences). They also warned not to rely on “reasoning” shown by some public LLM interfaces, arguing that it can be produced as output content rather than exposing true internal reasoning, reinforcing that public systems remain a black box behind the boundary. 

Question

“What about in outputs, It contains, such as citations and references, to index sources it relied on to generate that output. Would this help with reasoning?” 

Answer

They said citations can support a human-in-the-loop workflow and are common in corporate implementations, but they don’t automatically explain how the system combined materials—especially when many documents are involved—so the black-box problem can remain. They stressed that citations can be misleading or mismatched (a cited source may not actually support the claim), so users must verify them; one mitigation they suggested is prompting the system to include not only citations but also the specific excerpted text from the cited material that it relied on for the relevant point. 

Question

“LLMs are known to provide customers with an indemnity for copyright infringement. Do you see a world in which LLMs provide an indemnity for hallucinations, and should they be held accountable?” 

Answer

They were effectively skeptical: because LLMs are probabilistic rather than deterministic, providers can’t fully control hallucinations, so they’re unlikely to offer indemnities for them; you can reduce risk, but not drive it to zero consistently. They also suggested major providers are unlikely to accept broad liability and may position themselves as platforms while relying on disclaimers, with practical barriers (scale, resources, jurisdiction) making it hard for most individuals to hold large companies to account. 

Question

“When Dr. Ken refers to filters or prompting the LLM, is that via system prompts, or built-in AI guardrail tools, or both? And he’s asking if you could help understand the differences between the two.”

Answer

They explained that agents typically use two layers of prompting: a developer-set system/master prompt that should be hidden from end users (where you instruct the agent not to lie, not to make things up, whether to use RAG, how to cite sources, etc.), and the user prompt where users ask their questions. Guardrails/controls can exist as tools, but the core distinction they highlighted is that the master/system instructions define the agent’s baseline behaviour, while user prompts operate within that boundary. 

Featured Speakers

Learn from leading authorities

Dr Ken Tan
Howard Elliot

Trusted by 2,100+ legal firms worldwide

AUSTRALIA