World Congress 2026 Europe Jul 10, 2026 Session details

Beyond the Benchmark: How to Evaluate AI Agents in the Real World

Taylor Jordan Smith

Tyler Smith reveals that high benchmark scores mask fundamental AI execution failures. Ditch the happy-path demo and evaluate production agents using full-stack observability and continuous telemetry.

Pause
Mute Enter Fullscreen
#1 about 4 min

Moving beyond happy paths in AI agent development

The transition from simple demo applications to production-ready systems introduces complex interconnected components.

#2 about 2 min

Recognizing components in production-level agentic testing

Evaluating real-world readiness requires observing retrieval processes, tool execution, and failure recovery paths.

#3 about 3 min

Expanding technical scope from model to agent evaluation

Large language model accuracy metrics must evolve to capture multi-step autonomous behavior and tool usage.

#4 about 4 min

Demystifying the core technical layers of agent architectures

A complete agentic system comprises the underlying application, execution harnesses, and secure sandbox environments.

#5 about 2 min

Analyzing an academic research agent workflow

Practical examples demonstrate how an agent plans, retrieves data from varied sources, and presents autonomous results.

#6 about 2 min

Identifying critical performance areas for agent observation

Effective monitoring focuses on execution outcomes, resilience variations, operational performance, and token usage costs.

#7 about 4 min

Implementing a robust observability stack for agents

Structuring a deployment leverages metrics, logs, and traces through tools like Prometheus and OpenTelemetry.

#8 about 6 min

Demonstrating visual trace collection with Nvidia blueprints

Integrating open-source multi-agent workflows into OpenShift AI enables detailed visibility of model execution and specific tool spans.

#9 about 4 min

Designing customized evaluation pipelines from trace data

Meaningful systems transition from observed trace metrics to deterministic checks and language model judging.

#10 about 3 min

Creating a continuous iteration loop for agent applications

Feedback from automated evaluations and user interactions directly informs model fine-tuning and future capability deployments.

#11 about 2 min

Addressing human review scaling and tool integrations

Prioritizing high-risk components and leveraging automation tools helps manage the scaling challenges of human review.

Matching moments

2:28 min

Introduction to building reliable AI agents in production

Max Tkacz Max Tkacz · WWC 2025

2:40 min

Reviewing live performance of self-correcting AI engineering agents

Ingo Eichhorst Ingo Eichhorst · WWC Europe 2026

2:43 min

Best practices for implementing reliable AI agent frameworks

Marcel Scherenberg Marcel Scherenberg · WWC 2025

1:00 min

Limitations of current performance benchmarks for agentic AI

Manuel Weichselbaum Manuel Weichselbaum · WWC Europe 2026

5:27 min

Evaluating AI agents through unpredictable behavior and logic tests

Chris Heilmann +2 · LIVE

1:48 min

Adapting observability strategies for long-running enterprise AI agents

Christian Heilmann Christian Heilmann +3 · WWC Europe 2026

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Closing the Visibility Gap: Lessons from Safety Critical Agentic Systems

Vivek Pandit

Principal Engineer at Cadence

Vivek Pandit
Open session

World Congress 2026 North America

Evals Are Infra: Building AI Systems Developers Can Actually Trust

Phoebe Wang

Member of Technical Staff at OpenAI

Phoebe Wang
Open session

World Congress 2026 North America

Your Evals Passed. Your Agent Just Emptied a Database.

Tejas Pravinbhai Patel

IEEE Award-Winning Researcher | Best Keynote Speaker | Sr. Software Engineer at Amazon | AI Systems & Agent Architect

Tejas Pravinbhai Patel
Open session

World Congress 2026 North America

Agents Can't Iterate Against Tests That Lie

Rocky Warren

Senior Staff Software Engineer at Clipboard

Rocky Warren
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

DeepAgents: Build Multi-Agent AI Systems That Actually Work

Anagha Rumade, Anjana Umapathy, Apoorva Jaiswal

Anagha Rumade
Anjana Umapathy
Apoorva Jaiswal