World Congress 2026 Europe Jul 10, 2026 Session details

Testing AI Agents: Automated Evaluation for Chatbots & RAG Systems

Sebastian Messingfeld

Traditional string-matching tests are obsolete for non-deterministic AI agents. Discover how to build an AI testing pyramid using LLM-as-a-judge to catch silent chatbot regressions in your deployment pipelines.

Pause
Mute Enter Fullscreen
#1 about 9 min

Understanding AI chatbot vulnerabilities and stateful attacks

How real-world AI agents silently fail through prompt injections, outdated knowledge, and unintended system manipulation.

#2 about 3 min

The AI agent test pyramid and basic validations

Structuring efficient test suites with cheap schema checks, reference-based assertions, and deterministic unit tests.

#3 about 4 min

Introducing LLMs as judges for automated testing

Using independent language models to critically review chatbot outputs at scale for correctness and consistency.

#4 about 7 min

Evaluating AI outputs practically using the DeepEval framework

Using the DeepEval Python framework to score system responses against retrieved context via specialized metrics.

#5 about 2 min

Running automated AI testing continuously in CI environments

Scheduling periodic test runs in integration pipelines to guard against model drift and knowledge base changes.

#6 about 4 min

Empowering non-developers with BDD and Gherkin testing

Enabling product owners and content editors to define quality expectations through human-readable structured test scenarios.

#7 about 4 min

Key takeaways for reliable AI agent testing

A summarized blueprint for preventing regressions and maintaining cross-functional ownership of overall chatbot quality.

Matching moments

2:28 min

Introduction to building reliable AI agents in production

Max Tkacz Max Tkacz · World Congress 2025

4:07 min

Running automated evaluation tests for conversational tool selection

Rishabh Budhiraja Rishabh Budhiraja +1 · World Congress 2026 Europe

2:06 min

Rethinking team structures around AI agent capabilities

Mike Mike · World Congress 2025

5:27 min

Evaluating AI agents through unpredictable behavior and logic tests

Chris Heilmann +2 · LIVE

2:40 min

Reviewing live performance of self-correcting AI engineering agents

Ingo Eichhorst Ingo Eichhorst · World Congress 2026 Europe

3:19 min

Validating code generation with evals, datasets, and graders

Przemysław Ładyński Przemysław Ładyński · World Congress 2026 Europe

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 24, 2026 · 16:50–17:20

Stage 6

Who Tests the AI? Building Trustworthy AI Systems at Enterprise Scale

Him Raj Singh

Manager, Software Engineer at PayPal

Him Raj Singh
Open session

World Congress 2026 North America

September 24, 2026 · 11:20–11:25

Outdoor Stage

Finding the Edges: Testing, Evaluating, and Monitoring Voice AI Agents Before Your Users Do

Matt Wyman

CEO of Okareo

Matt Wyman
Open session

World Congress 2026 North America

September 25, 2026 · 11:40–12:10

Stage 2

Reinventing Testing Practices in the AI Era

Eric Deandrea

Java Champion & Senior Principal Software Engineer at IBM

Eric Deandrea
Open session

World Congress 2026 North America

September 24, 2026 · 11:40–12:10

Stage 4

Taming Rogue Agents: Observability-Driven Evaluation for Production Reliability

Anagha Rumade, Anjana Umapathy, Apoorva Jaiswal

Anagha Rumade
Anjana Umapathy
Apoorva Jaiswal
Open session

World Congress 2026 North America

September 24, 2026 · 11:00–11:30

Stage 5

The Missing Infrastructure for AI Agents

Chris Waterson

CTO and Co-Founder of Guild.ai

Chris Waterson
Open session

World Congress 2026 North America

September 24, 2026 · 15:10–15:20

Outdoor Stage

AI That Argues With Itself: Building Self-Debating Systems That Catch Their Own Bugs

Shreya Singhal

AI Applied Scientist at Claritev

Shreya Singhal