> Markdown version of [/videos/100421-21-experiments-in-six-weeks-a-playbook-for-improving-your-ai-agent](https://www.wearedevelopers.com/videos/100421-21-experiments-in-six-weeks-a-playbook-for-improving-your-ai-agent). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # 21 Experiments in Six Weeks: A Playbook for Improving Your AI Agent Think upgrading to a smarter model automatically improves your AI agent? Sentry ran 21 A/B tests in six weeks and discovered that more expensive models can actually hurt performance. - **Speakers:** [Sofia Rest](https://www.wearedevelopers.com/@sofia-rest) - **Event:** World Congress 2026 North America - **Published:** September 26, 2026 - **Duration:** 16:30 - **URL:** https://www.wearedevelopers.com/videos/100421-21-experiments-in-six-weeks-a-playbook-for-improving-your-ai-agent ## Access Playback and chapters for this video are available with a Free account. ## Related Moments - [Reviewing live performance of self-correcting AI engineering agents](https://www.wearedevelopers.com/videos/100190-architecture-3-0-from-90-to-99-999-reliability-in-building-ai-systems) (from "Architecture 3.0: From 90% to 99.999% Reliability in Building AI Systems") - [Introduction to building reliable AI agents in production](https://www.wearedevelopers.com/videos/1523-the-ai-agent-path-to-prod-building-for-reliability) (from "The AI Agent Path to Prod: Building for Reliability") - [Shifting developer workloads and realistic AI productivity gains](https://www.wearedevelopers.com/videos/1830-wearedevelopers-live-speculaitions) (from "WeAreDevelopers LIVE - SpeculAItions") - [Implementing practical scrum adjustments to orchestrate AI usage correctly](https://www.wearedevelopers.com/videos/1950-the-scrum-master-as-an-orchestrator-guiding-human-ai-collaboration-in-modern-teams) (from "The Scrum Master as an Orchestrator: Guiding Human–AI Collaboration in Modern Teams") - [Key takeaways for reliable AI agent testing](https://www.wearedevelopers.com/videos/100300-testing-ai-agents-automated-evaluation-for-chatbots-rag-systems) (from "Testing AI Agents: Automated Evaluation for Chatbots & RAG Systems") - [Verifying the quality of rapidly generated AI code](https://www.wearedevelopers.com/videos/100069-building-the-next-generation-of-ai-developer-tools) (from "Building the next generation of AI developer tools") ## Related Articles - [Never delegate the understanding](https://www.wearedevelopers.com/magazine/749-never-delegate-the-understanding) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [The Overflow: AI and Agentic Coding](https://www.wearedevelopers.com/magazine/721-the-overflow-ai-and-agentic-coding) - [AI Eats the Verifiable First](https://www.wearedevelopers.com/magazine/765-ai-eats-the-verifiable-first) ## Related Jobs - [Senior AI/ML Engineer](https://www.wearedevelopers.com/jobs/48352-senior-ai-ml-engineer) at **PagerDuty** - [Senior AI Developer](https://www.wearedevelopers.com/jobs/ext/2836034-senior-ai-developer) at **PwC** - [Staff Software Engineer, GitHub Intelligence (Copilot Agents)](https://www.wearedevelopers.com/jobs/ext/2650582-staff-software-engineer-github-intelligence-copilot-agents) at **GitHub** - [Software Engineer - Video](https://www.wearedevelopers.com/jobs/ext/2600051-software-engineer-video) at **Twilio** - [Principal AI Forward Deployed Engineer](https://www.wearedevelopers.com/jobs/48531-principal-ai-forward-deployed-engineer) at **Expedia** - [Principal Product Manager, Agentic Evals](https://www.wearedevelopers.com/jobs/48537-principal-product-manager-agentic-evals) at **Expedia**