World Congress 2026 Europe

The Art and Science behind evaluating AI Agents at scale

July 8, 2026 16:00 – 18:00 · 120 min Room M2 (40 Seats)

What this session covers

Evaluating AI Agent is a mix of science and art. Working with subject matter experts is more important than ever. New methods and best practices are emerging to evaluate these systems at scale.

In this talk we will discuss a case study of a production agent used across an entire company. We will discuss live evals, how to build a golden dataset, how to collaborate with SMEs, what worked and what didn’t over a 6 months project.

We will share some of the best practices we found are working well in production contexts after investing hundreds of hours analyzing evals, building reports and iterating.

This talk is not about just the theory, we will use a real case study and we will share all the info you need to really iterate fast and build evals that matter for your use case!

Related talks at this congress

Open session

World Congress 2026 Europe

July 10, 2026 · 13:00–13:30

Stage 8 - powered by Red Hat

Beyond the Benchmark: How to Evaluate AI Agents in the Real World

Taylor Jordan Smith

Senior AI Developer Advocate at Red Hat

Taylor Jordan Smith
Open session

World Congress 2026 Europe

July 10, 2026 · 14:20–14:50

Stage 6 - powered by Microsoft

Testing AI Agents: Automated Evaluation for Chatbots & RAG Systems

Sebastian Messingfeld

Staff Engineer at Eurowings Digital

Sebastian Messingfeld
Open session

World Congress 2026 Europe

July 9, 2026 · 10:50–11:20

Stage 2

What 500+ Production Environments Taught Us About Shipping AI Agents

Liran Hason

VP of AI at Coralogix

Liran Hason
Open session

World Congress 2026 Europe

July 9, 2026 · 16:10–16:40

Stage 3 - powered by AWS

5 things I wish I hadn’t done building my AI agent

Shachar Azriel

VP Product at Baz

Shachar Azriel
All sessions at this congress