World Congress 2026 Europe • Jul 10, 2026 • Session details

Testing AI Agents: Automated Evaluation for Chatbots & RAG Systems

Sebastian Messingfeld

Traditional string-matching tests are obsolete for non-deterministic AI agents. Discover how to build an AI testing pyramid using LLM-as-a-judge to catch silent chatbot regressions in your deployment pipelines.

Pause
Mute Enter Fullscreen
#1 about 9 min

Understanding AI chatbot vulnerabilities and stateful attacks

How real-world AI agents silently fail through prompt injections, outdated knowledge, and unintended system manipulation.

#2 about 3 min

The AI agent test pyramid and basic validations

Structuring efficient test suites with cheap schema checks, reference-based assertions, and deterministic unit tests.

#3 about 4 min

Introducing LLMs as judges for automated testing

Using independent language models to critically review chatbot outputs at scale for correctness and consistency.

#4 about 7 min

Evaluating AI outputs practically using the DeepEval framework

Using the DeepEval Python framework to score system responses against retrieved context via specialized metrics.

#5 about 2 min

Running automated AI testing continuously in CI environments

Scheduling periodic test runs in integration pipelines to guard against model drift and knowledge base changes.

#6 about 4 min

Empowering non-developers with BDD and Gherkin testing

Enabling product owners and content editors to define quality expectations through human-readable structured test scenarios.

#7 about 4 min

Key takeaways for reliable AI agent testing

A summarized blueprint for preventing regressions and maintaining cross-functional ownership of overall chatbot quality.

Matching moments

2:28 min

Introduction to building reliable AI agents in production

Max Tkacz Max Tkacz · World Congress 2025

4:07 min

Running automated evaluation tests for conversational tool selection

Rishabh Budhiraja Rishabh Budhiraja +1 · World Congress 2026 Europe

2:06 min

Rethinking team structures around AI agent capabilities

Mike Mike · World Congress 2025

5:27 min

Evaluating AI agents through unpredictable behavior and logic tests

Chris Heilmann Chris Heilmann +2 · LIVE

2:40 min

Reviewing live performance of self-correcting AI engineering agents

Ingo Eichhorst Ingo Eichhorst · World Congress 2026 Europe

3:19 min

Validating code generation with evals, datasets, and graders

Przemysław Ładyński Przemysław Ładyński · World Congress 2026 Europe