World Congress 2026 Europe Jul 9, 2026 Session details

Truth, Lies, and Probabilities: Testing AI Hallucinations

Anastasia Simou

Why do conventional QA methods fail non-deterministic language models? AI hallucinations aren't bugs but inherent architectural features, demanding a radical shift toward statistical evaluation and adversarial stress testing.

Pause
Mute Enter Fullscreen
#1 about 4 min

The systemic risk of AI hallucinations in practice

Real-world examples of AI models generating fake citations demonstrate why factual fabrication is a systemic model flaw.

#2 about 4 min

Defining intrinsic and extrinsic AI hallucinations

Models present unique failure modes depending on whether they contradict valid prompt contexts or generate completely fabricated facts.

#3 about 5 min

How data compression distorts parametric knowledge

Dimensionality reduction within embedding spaces forces language models to merge semantic clusters and lose distinct factual accuracies.

#4 about 4 min

The probabilistic mechanics of token completion

Large language models construct their outputs based on statistical sequence probability rather than executing authentic database factual retrieval.

#5 about 4 min

Why traditional quality engineering fails for LLMs

Architectural non-determinism explicitly prevents modern quality engineers from applying standard reproducibility tests or isolated debugging methods against language models.

#6 about 3 min

The practical difficulty of establishing ground truth

Validating generative responses is notoriously challenging when testing source data remains incomplete, historically ambiguous, or constantly changing.

#7 about 2 min

A four-step process for ground truth evaluation

Testing administrators execute structural validation by curating verifiable records, altering prompt phrasing, and averaging large scale analytical runs.

#8 about 3 min

Stress testing language models with adversarial prompting

Engineers deliberately break model consistency by injecting false theoretical bounds, demanding highly suspicious specifics, and challenging unrepresented boundary constraints.

#9 about 2 min

Verifying model responses across contested knowledge limitations

Quality teams manage highly ambiguous subjects by leveraging automated temperature polling, counterfactual debate comparisons, and rigid source tracing logic.

#10 about 3 min

Mitigating hallucination behaviors through retrieval-augmented generation

Implementing external documentation pathways effectively restricts contextual failures despite adding maintenance complexity for chunk extraction parameters and embedding similarity thresholds.

#11 about 1 min

Establishing explicit model boundaries via prompt engineering

Programmers manipulate behavioral confidence boundaries by demanding models accurately report data gaps, append exact source citations, and self-assign numerical ratings.

#12 about 2 min

Calculating statistical variation and risk confidence intervals

Reconciling raw hallucination observations into standard error calculations accurately verifies whether generative application deployments survive strict organizational safety thresholds.

#13 about 3 min

Transitioning toward statistical evaluations for AI systems

Accepting non-deterministic design mechanics highlights the essential priority of adopting probability distributions and infrastructure maintenance into modern quality cycles.

Matching moments

4:08 min

Managing hallucination risks and bias in AI generation

Léonie Watson Léonie Watson +4 · A11y + AI

1:34 min

Mitigating the inherent challenges of generative AI tools

Mary Grygleski Mary Grygleski · LIVE

5:03 min

Evaluating the accuracy and correctness of generated content

Cheuk Ho · WWC 2023

10:17 min

Discussion on AI hallucinations and practical developer workflows

Akmal Chaudhri Akmal Chaudhri · LIVE

4:25 min

Overcoming AI hallucinations and restrictive content guardrails

Perf + AI

1:49 min

Addressing the risks of hallucination and chat model manipulation

Zachary Powell Zachary Powell · WWC 2024

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Who Tests the AI? Building Trustworthy AI Systems at Enterprise Scale

Him Raj Singh

PayPal, Manager, Software Engineer

Him Raj Singh
Open session

World Congress 2026 North America

Reinventing Testing Practices in the AI Era

Eric Deandrea

Java Champion & Senior Principal Software Engineer, IBM

Eric Deandrea
Open session

World Congress 2026 North America

AI That Argues With Itself: Building Self-Debating Systems That Catch Their Own Bugs

Shreya Singhal

AI Applied Scientist at Claritev

Shreya Singhal
Open session

World Congress 2026 North America

Engineering the Pivot: How Creative Strategy Solves the Hard Problems of AI Accuracy and Scale

Shruti Tiwari

AI/ML product manager, Dell

Shruti Tiwari
Open session

World Congress 2026 North America

Building Stuff with GenAI - The Open Minded Workshop beyond OpenAI

Andreas Erben

CTO for Applied AI and Metaverse at daenet

Andreas Erben
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan