World Congress 2026 Europe • Jul 10, 2026 • Session details

Beyond the Benchmark: How to Evaluate AI Agents in the Real World

Taylor Jordan Smith

Tyler Smith reveals that high benchmark scores mask fundamental AI execution failures. Ditch the happy-path demo and evaluate production agents using full-stack observability and continuous telemetry.

Pause
Mute Enter Fullscreen
#1 about 4 min

Moving beyond happy paths in AI agent development

The transition from simple demo applications to production-ready systems introduces complex interconnected components.

#2 about 2 min

Recognizing components in production-level agentic testing

Evaluating real-world readiness requires observing retrieval processes, tool execution, and failure recovery paths.

#3 about 3 min

Expanding technical scope from model to agent evaluation

Large language model accuracy metrics must evolve to capture multi-step autonomous behavior and tool usage.

#4 about 4 min

Demystifying the core technical layers of agent architectures

A complete agentic system comprises the underlying application, execution harnesses, and secure sandbox environments.

#5 about 2 min

Analyzing an academic research agent workflow

Practical examples demonstrate how an agent plans, retrieves data from varied sources, and presents autonomous results.

#6 about 2 min

Identifying critical performance areas for agent observation

Effective monitoring focuses on execution outcomes, resilience variations, operational performance, and token usage costs.

#7 about 4 min

Implementing a robust observability stack for agents

Structuring a deployment leverages metrics, logs, and traces through tools like Prometheus and OpenTelemetry.

#8 about 6 min

Demonstrating visual trace collection with Nvidia blueprints

Integrating open-source multi-agent workflows into OpenShift AI enables detailed visibility of model execution and specific tool spans.

#9 about 4 min

Designing customized evaluation pipelines from trace data

Meaningful systems transition from observed trace metrics to deterministic checks and language model judging.

#10 about 3 min

Creating a continuous iteration loop for agent applications

Feedback from automated evaluations and user interactions directly informs model fine-tuning and future capability deployments.

#11 about 2 min

Addressing human review scaling and tool integrations

Prioritizing high-risk components and leveraging automation tools helps manage the scaling challenges of human review.

Matching moments

2:28 min

Introduction to building reliable AI agents in production

Max Tkacz Max Tkacz · World Congress 2025

2:40 min

Reviewing live performance of self-correcting AI engineering agents

Ingo Eichhorst Ingo Eichhorst · World Congress 2026 Europe

2:43 min

Best practices for implementing reliable AI agent frameworks

Marcel Scherenberg Marcel Scherenberg · World Congress 2025

1:00 min

Limitations of current performance benchmarks for agentic AI

Manuel Weichselbaum Manuel Weichselbaum · World Congress 2026 Europe

5:27 min

Evaluating AI agents through unpredictable behavior and logic tests

Chris Heilmann +2 · LIVE

1:48 min

Adapting observability strategies for long-running enterprise AI agents

Christian Heilmann Christian Heilmann +3 · World Congress 2026 Europe

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 24, 2026 · 11:40–12:10

Stage 4

Taming Rogue Agents: Observability-Driven Evaluation for Production Reliability

Anagha Rumade, Anjana Umapathy, Apoorva Jaiswal

Anagha Rumade
Anjana Umapathy
Apoorva Jaiswal
Open session

World Congress 2026 North America

September 25, 2026 · 15:45–15:55

Outdoor Stage

Closing the Visibility Gap: Lessons from Safety Critical Agentic Systems

Vivek Pandit

Frontier AI Lead at Turing

Vivek Pandit
Open session

World Congress 2026 North America

September 24, 2026 · 11:20–11:25

Outdoor Stage

Finding the Edges: Testing, Evaluating, and Monitoring Voice AI Agents Before Your Users Do

Matt Wyman

CEO of Okareo

Matt Wyman
Open session

World Congress 2026 North America

September 24, 2026 · 11:00–11:30

Stage 5

The Missing Infrastructure for AI Agents

Chris Waterson

CTO and Co-Founder of Guild.ai

Chris Waterson
Open session

World Congress 2026 North America

September 24, 2026 · 11:40–12:10

Stage 5

Your Agents Need Observability Before They Need Better Models

Julia Furst Morgado

Principal Developer Relations Engineer at Dash0

Julia Furst Morgado
Open session

World Congress 2026 North America

September 25, 2026 · 15:30–16:00

Stage 4

Evals Are Infra: Building AI Systems Developers Can Actually Trust

Phoebe Wang

Member of Technical Staff at OpenAI

Phoebe Wang