World Congress 2026 Europe Jul 10, 2026 Session details

Testing AI Agents: Automated Evaluation for Chatbots & RAG Systems

Sebastian Messingfeld

Traditional string-matching tests are obsolete for non-deterministic AI agents. Discover how to build an AI testing pyramid using LLM-as-a-judge to catch silent chatbot regressions in your deployment pipelines.

Pause
Mute Enter Fullscreen
#1 about 9 min

Understanding AI chatbot vulnerabilities and stateful attacks

How real-world AI agents silently fail through prompt injections, outdated knowledge, and unintended system manipulation.

#2 about 3 min

The AI agent test pyramid and basic validations

Structuring efficient test suites with cheap schema checks, reference-based assertions, and deterministic unit tests.

#3 about 4 min

Introducing LLMs as judges for automated testing

Using independent language models to critically review chatbot outputs at scale for correctness and consistency.

#4 about 7 min

Evaluating AI outputs practically using the DeepEval framework

Using the DeepEval Python framework to score system responses against retrieved context via specialized metrics.

#5 about 2 min

Running automated AI testing continuously in CI environments

Scheduling periodic test runs in integration pipelines to guard against model drift and knowledge base changes.

#6 about 4 min

Empowering non-developers with BDD and Gherkin testing

Enabling product owners and content editors to define quality expectations through human-readable structured test scenarios.

#7 about 4 min

Key takeaways for reliable AI agent testing

A summarized blueprint for preventing regressions and maintaining cross-functional ownership of overall chatbot quality.

Matching moments

2:28 min

Introduction to building reliable AI agents in production

Max Tkacz Max Tkacz · WWC 2025

4:07 min

Running automated evaluation tests for conversational tool selection

Rishabh Budhiraja Rishabh Budhiraja +1 · WWC Europe 2026

2:06 min

Rethinking team structures around AI agent capabilities

Mike Mike · WWC 2025

5:27 min

Evaluating AI agents through unpredictable behavior and logic tests

Chris Heilmann +2 · LIVE

2:40 min

Reviewing live performance of self-correcting AI engineering agents

Ingo Eichhorst Ingo Eichhorst · WWC Europe 2026

3:19 min

Validating code generation with evals, datasets, and graders

Przemysław Ładyński Przemysław Ładyński · WWC Europe 2026

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Who Tests the AI? Building Trustworthy AI Systems at Enterprise Scale

Him Raj Singh

PayPal, Manager, Software Engineer

Him Raj Singh
Open session

World Congress 2026 North America

Reinventing Testing Practices in the AI Era

Eric Deandrea

Java Champion & Senior Principal Software Engineer, IBM

Eric Deandrea
Open session

World Congress 2026 North America

AI That Argues With Itself: Building Self-Debating Systems That Catch Their Own Bugs

Shreya Singhal

AI Applied Scientist at Claritev

Shreya Singhal
Open session

World Congress 2026 North America

Evals Are Infra: Building AI Systems Developers Can Actually Trust

Phoebe Wang

Member of Technical Staff at OpenAI

Phoebe Wang
Open session

World Congress 2026 North America

Your Evals Passed. Your Agent Just Emptied a Database.

Tejas Pravinbhai Patel

IEEE Award-Winning Researcher | Best Keynote Speaker | Sr. Software Engineer at Amazon | AI Systems & Agent Architect

Tejas Pravinbhai Patel
Open session

World Congress 2026 North America

Agents Can't Iterate Against Tests That Lie

Rocky Warren

Senior Staff Software Engineer at Clipboard

Rocky Warren