World Congress 2026 North America
September 25, 2026 · 15:30–16:00
Stage 4
Evals Are Infra: Building AI Systems Developers Can Actually Trust
Phoebe Wang
Member of Technical Staff at OpenAI
World Congress 2026 North America
AI evals behave like flaky tests. The model under test is probabilistic, and the LLM judging it often is too. The same eval can pass on one run and fail on the next. A single score is one sample, which makes it a shaky basis for a release gate.
Plumloom makes the eval score itself more reliable. It calibrates the judging standard, uses multiple independent judges, and for scenario evaluations runs repeated trials to report a confidence interval. That helps separate a real improvement from noise instead of relying on a single number.
Autoeval, our open-source CLI with Apache 2.0 licensing and MCP support, brings this into the release path. Bring the scenario, transcript, or OpenTelemetry/OpenInference trace your harness already produces. There is no SDK or observability pipeline to wire up, and CI can gate the release when a score misses your threshold.
This pitch covers what an eval really is, why single-run scores can mislead, and how to put reliable, gated evals into your terminal, agent harness, and CI.
World Congress 2026 North America
September 25, 2026 · 15:30–16:00
Stage 4
Phoebe Wang
Member of Technical Staff at OpenAI
World Congress 2026 North America
September 25, 2026 · 16:50–17:20
Stage 3
Tejas Pravinbhai Patel
IEEE Award-Winning Researcher | Best Keynote Speaker | Sr. Software Engineer at Amazon | AI Systems & Agent Architect
World Congress 2026 North America
September 24, 2026 · 11:40–12:10
Stage 4
Anagha Rumade, Anjana Umapathy, Apoorva Jaiswal
World Congress 2026 North America
September 24, 2026 · 17:30–18:00
Stage 6
Emmanuel Acheampong
Senior Manager Developer Relations at Crusoe AI