World Congress 2026 Europe
July 10, 2026 · 13:00–13:30
Stage 8 - powered by Red Hat
Beyond the Benchmark: How to Evaluate AI Agents in the Real World
Taylor Jordan Smith
Senior AI Developer Advocate at Red Hat
World Congress 2026 Europe
AI Agents, chatbots, and RAG systems are easy to prototype — but difficult to test reliably. Small changes to prompts, models, retrieval sources, or system instructions can silently change behavior, and classic assertions (string matching, snapshots) often fail to capture what actually matters: correctness, relevance, grounded answers, and consistent multi-turn dialogue.
In this talk, we’ll start with the common testing problems in real projects: “it worked yesterday”, hidden regressions, evaluation noise, and the challenge of aligning developers and stakeholders on what “good” means.
Then we’ll explore practical testing possibilities with evaluation frameworks like DeepEval: how to validate responses beyond keyword matching, how to structure test cases for both chatbots and retrieval-based assistants, how to define pragmatic quality gates, and how to run these checks continuously in suggests, then?
As a side topic, we’ll show how BDD/Gherkin can wrap these evaluations into human-readable scenarios (Given–When–Then), making expectations reviewable by non-developers while keeping the actual validation powered by automated evaluation metrics.
You’ll leave with a reusable blueprint for introducing automated AI evaluation into your development workflow — from local runs to CI pipelines with actionable reports.
World Congress 2026 Europe
July 10, 2026 · 13:00–13:30
Stage 8 - powered by Red Hat
Taylor Jordan Smith
Senior AI Developer Advocate at Red Hat
World Congress 2026 Europe
July 8, 2026 · 16:00–18:00
Room M2 (40 Seats)
Alfonso Graziano
AI Tech Lead at Nearform
World Congress 2026 Europe
July 10, 2026 · 11:00–11:30
Stage 1
Andrei Nutas
Technical Test Architect at Atos
World Congress 2026 Europe
July 9, 2026 · 10:50–11:20
Stage 9
Jakub Janczyk
Senior Engineer at Kit