World Congress 2026 North America
Agents Can't Iterate Against Tests That Lie
Rocky Warren
Senior Staff Software Engineer at Clipboard
World Congress 2026 North America
World Congress 2026 North America
September 23–25, 2026 · San José, CA
Attend in person
Get ticketsWatch remotely
Pro
Can’t make it to San José? Watch this session live with Pro. You also get:
Writing the code was the easy part. But setting out to improve Sentry’s AI code review agent meant answering a much harder question: what does it mean to make your agent “better”, and how do you know if it’s working? Over the course of six weeks, our team ran 21 live experiments to find out. Along the way, we built an eval loop that moved from offline datasets and limited observability to A/B testing infrastructure, composite metrics, and detailed cost tracking in a single dashboard. And our most honest signal—whether the developer actually fixed the flagged bug—surfaced results that broke our intuition (like why a smarter model isn’t always the answer) and forced hard tradeoffs between code review quality and cost. In this talk, I’ll walk through the practical playbook for improving your AI agent: how to define “better”, how to measure it, and how to get quality gains without breaking the bank.
World Congress 2026 North America
Rocky Warren
Senior Staff Software Engineer at Clipboard
World Congress 2026 North America
Anagha Rumade, Anjana Umapathy, Apoorva Jaiswal
World Congress 2026 North America
Shreya Singhal
AI Applied Scientist at Claritev
World Congress 2026 North America
Vivek Pandit
Principal Engineer at Cadence