> Markdown version of [/events/world-congress-2026-north-america/sessions/1849-finding-the-edges](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1849-finding-the-edges). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Finding the Edges: Testing, Evaluating, and Monitoring Voice AI Agents Before Your Users Do - **Date:** Thursday, Sep 24, 2026 - **Time:** 11:20–11:25 (5 min) - **Room:** Outdoor Stage - **Event:** World Congress 2026 North America ## Description Every demo of a voice AI agent looks great — until real users start talking. They interrupt, mumble, switch languages mid-sentence, and ask the one question that sends the agent off the rails. Traditional scripted QA only verifies the behaviors you already thought of, which is exactly why so many voice agents fail in production in ways their teams never saw coming. This session walks through a practical, engineering-grade approach to shipping Voice AI agents you can trust, built on three pillars: simulation, evaluation, and monitoring. You'll see how synthetic "drivers" — AI-powered simulated users with distinct personalities, goals, and contexts — hold realistic multi-turn conversations with your agent across 30+ languages and real-world audio conditions (noise, crosstalk, clipping), actively exploring the edges scripted tests miss. We'll then look at how judge-based, symbolic, and audio evaluations turn those discoveries into CI/CD release gates, so a change that degrades conversation quality fails the build before it reaches users. Finally, we'll close the loop with production monitoring that captures real failures and automatically converts them into regression tests — so the same mistake never ships twice. You'll leave with a concrete blueprint for finding the edges of your Voice AI agent before your customers do. ## Speaker ### [Matt Wyman](https://www.wearedevelopers.com/@matt-wyman) CEO at Okareo ## Related talks at this congress - [Your Evals Passed. Your Agent Just Emptied a Database.](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1684-your-evals-passed) — Tejas Pravinbhai Patel - [Taming Rogue Agents: Observability-Driven Evaluation for Production Reliability](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1768-taming-rogue-agents) — Anagha Rumade, Anjana Umapathy, Apoorva Jaiswal - [Closing the Visibility Gap: Lessons from Safety Critical Agentic Systems](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1408-closing-the) — Vivek Pandit - [I Don't Trust AI Agents (And Neither Should You): Building Production-Ready Architectures](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1772-i-don-t-trust-ai) — Darko Mesaros ## Watch remotely Can’t make it to San José? Watch this session live with Pro. You also get: - All full videos, bookmarks, and playlists - World Congress livestreams [See pricing](https://www.wearedevelopers.com/pricing) ## Links - [Get tickets](https://www.wearedevelopers.com/world-congress-north-america/tickets)