World Congress 2026 North America

21 Experiments in Six Weeks: A Playbook for Improving Your AI Agent

September 23–25, 2026

World Congress 2026 North America

September 23–25, 2026 · San José, CA

Attend in person

Get tickets

Watch remotely

Watch live with Pro

Pro

Can’t make it to San José? Watch this session live with Pro. You also get:

  • All full videos, bookmarks, and playlists
  • World Congress livestreams
See pricing

What this session covers

Writing the code was the easy part. But setting out to improve Sentry’s AI code review agent meant answering a much harder question: what does it mean to make your agent “better”, and how do you know if it’s working? Over the course of six weeks, our team ran 21 live experiments to find out. Along the way, we built an eval loop that moved from offline datasets and limited observability to A/B testing infrastructure, composite metrics, and detailed cost tracking in a single dashboard. And our most honest signal—whether the developer actually fixed the flagged bug—surfaced results that broke our intuition (like why a smarter model isn’t always the answer) and forced hard tradeoffs between code review quality and cost. In this talk, I’ll walk through the practical playbook for improving your AI agent: how to define “better”, how to measure it, and how to get quality gains without breaking the bank.

Related talks at this congress

Open session

World Congress 2026 North America

Agents Can't Iterate Against Tests That Lie

Rocky Warren

Senior Staff Software Engineer at Clipboard

Rocky Warren
Open session

World Congress 2026 North America

Taming Rogue Agents: Observability-Driven Evaluation for Production Reliability

Anagha Rumade, Anjana Umapathy, Apoorva Jaiswal

Anagha Rumade
Anjana Umapathy
Apoorva Jaiswal
Open session

World Congress 2026 North America

AI That Argues With Itself: Building Self-Debating Systems That Catch Their Own Bugs

Shreya Singhal

AI Applied Scientist at Claritev

Shreya Singhal
Open session

World Congress 2026 North America

Closing the Visibility Gap: Lessons from Safety Critical Agentic Systems

Vivek Pandit

Principal Engineer at Cadence

Vivek Pandit
All sessions at this congress