Skip to content

Session

Evals Are Infra: Building AI Systems Developers Can Actually Trust

with Phoebe Wang

About This Session

AI teams often treat evals as a gate: a benchmark score, a pass/fail dashboard, or one number that decides whether a system ships. That breaks down quickly in production. Users disagree about what "good" means, behavior shifts across workflows, and agent failures often hide inside traces, tools, permissions, and human handoffs rather than in the final answer. This talk reframes evals as production infrastructure. I will show how developers can move from mystery scores to systems that expose failure modes: plural rubrics for stakeholder disagreement, trace-level observability, human review loops, failure taxonomies, and semantic drift/anomaly detection. The goal is practical: help teams ship AI products developers can debug, operate, and trust.

Topics

  • Agentic AI
  • Generative AI (GenAI)
  • Large Language Models (LLMs)
  • LLMOps
  • Observability
  • Reliability