World Congress 2026 North America

Agents Can't Iterate Against Tests That Lie

September 25, 2026 15:30 – 16:00 · 30 min Stage 1

What this session covers

Coding agents write nearly all of your code. Now you need to prove it works.

Flaky end-to-end (E2E) tests in shared environments make that proof difficult and unreliable. With limited context, agents often reach for the wrong fix: increase the timeout or add a retry.

This talk is a practical case study in rebuilding trust in tests for AI-heavy engineering organizations. I’ll show the workflow that helped us reduce the share of PRs affected by E2E flakes from 100% to under 15% in six weeks, including the open-source libraries and agent skills we built to classify flaky tests, connect to observability signals, and decide which tests to keep, delete, or move down the pyramid.

Attendees will leave with a repeatable playbook for making coding agents safe to use at scale and for teaching them to root-cause failing tests rather than “fix” them with retries.

Related talks at this congress

Open session

World Congress 2026 North America

September 25, 2026 · 16:50–17:20

Stage 3

Your Evals Passed. Your Agent Just Emptied a Database.

Tejas Pravinbhai Patel

IEEE Award-Winning Researcher | Best Keynote Speaker | Sr. Software Engineer at Amazon | AI Systems & Agent Architect

Tejas Pravinbhai Patel
Open session

World Congress 2026 North America

September 24, 2026 · 09:45–10:15

Mainstage

Manufacturing trust: speed and safety in the age of agents

Mark Cavage, Michael Irwin, Hervé Bizira

Mark Cavage
Michael Irwin
Hervé Bizira
Open session

World Congress 2026 North America

September 25, 2026 · 12:20–12:50

Stage 4

Give the Agent a Budget, Not a Token

Sachin Malhotra

MTS @Anthropic

Sachin Malhotra
Open session

World Congress 2026 North America

September 24, 2026 · 13:30–14:00

Stage 1

How AI Agents Tripled Our Test Coverage on a 1.8M-Line iOS Codebase

Kush Agrawal

Staff Software Engineer, Platform

Kush Agrawal
All sessions at this congress