World Congress 2026 Europe • Jul 10, 2026 • Session details

Beyond the Benchmark: How to Evaluate AI Agents in the Real World

Taylor Jordan Smith

Tyler Smith reveals that high benchmark scores mask fundamental AI execution failures. Ditch the happy-path demo and evaluate production agents using full-stack observability and continuous telemetry.

Pause
Mute Enter Fullscreen
#1 about 4 min

Moving beyond happy paths in AI agent development

The transition from simple demo applications to production-ready systems introduces complex interconnected components.

#2 about 2 min

Recognizing components in production-level agentic testing

Evaluating real-world readiness requires observing retrieval processes, tool execution, and failure recovery paths.

#3 about 3 min

Expanding technical scope from model to agent evaluation

Large language model accuracy metrics must evolve to capture multi-step autonomous behavior and tool usage.

#4 about 4 min

Demystifying the core technical layers of agent architectures

A complete agentic system comprises the underlying application, execution harnesses, and secure sandbox environments.

#5 about 2 min

Analyzing an academic research agent workflow

Practical examples demonstrate how an agent plans, retrieves data from varied sources, and presents autonomous results.

#6 about 2 min

Identifying critical performance areas for agent observation

Effective monitoring focuses on execution outcomes, resilience variations, operational performance, and token usage costs.

#7 about 4 min

Implementing a robust observability stack for agents

Structuring a deployment leverages metrics, logs, and traces through tools like Prometheus and OpenTelemetry.

#8 about 6 min

Demonstrating visual trace collection with Nvidia blueprints

Integrating open-source multi-agent workflows into OpenShift AI enables detailed visibility of model execution and specific tool spans.

#9 about 4 min

Designing customized evaluation pipelines from trace data

Meaningful systems transition from observed trace metrics to deterministic checks and language model judging.

#10 about 3 min

Creating a continuous iteration loop for agent applications

Feedback from automated evaluations and user interactions directly informs model fine-tuning and future capability deployments.

#11 about 2 min

Addressing human review scaling and tool integrations

Prioritizing high-risk components and leveraging automation tools helps manage the scaling challenges of human review.

Matching moments

2:28 min

Introduction to building reliable AI agents in production

Max Tkacz Max Tkacz · World Congress 2025

2:40 min

Reviewing live performance of self-correcting AI engineering agents

Ingo Eichhorst Ingo Eichhorst · World Congress 2026 Europe

2:43 min

Best practices for implementing reliable AI agent frameworks

Marcel Scherenberg Marcel Scherenberg · World Congress 2025

1:00 min

Limitations of current performance benchmarks for agentic AI

Manuel Weichselbaum Manuel Weichselbaum · World Congress 2026 Europe

5:27 min

Evaluating AI agents through unpredictable behavior and logic tests

Chris Heilmann Chris Heilmann +2 · LIVE

1:48 min

Adapting observability strategies for long-running enterprise AI agents

Christian Heilmann Christian Heilmann +3 · World Congress 2026 Europe