> Markdown version of [/videos/100548-running-ai-written-software-in-production?t=1616](https://www.wearedevelopers.com/videos/100548-running-ai-written-software-in-production?t=1616). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Running AI-Written Software in Production How do you debug production software that no human fully understands? Learn to build the rigorous evaluation frameworks and programmatic guardrails required to safely deploy AI-generated code. - **Speakers:** [Anurag Goel](https://www.wearedevelopers.com/@anurag-goel), [Milin Desai](https://www.wearedevelopers.com/@milin-desai), [Richard Rizk](https://www.wearedevelopers.com/@richard-rizk), [Sead Ahmetović](https://www.wearedevelopers.com/@sead-ahmetovic), [Jared Zoneraich](https://www.wearedevelopers.com/@jared-zoneraich) - **Event:** World Congress 2026 North America - **Published:** September 24, 2026 - **Duration:** 35:26 - **URL:** https://www.wearedevelopers.com/videos/100548-running-ai-written-software-in-production ## Summary The dramatic reduction in software creation costs driven by AI has led to an unprecedented surge in code generation, exponentially increasing deployment volumes and production incidents. With AI tools acting as autonomous engineering teammates—often authoring the vast majority of pull requests—organizations face the escalating challenge of running and debugging software that no human fully understands. This shift transforms traditional development pipelines into high-velocity systems where raw code generation is trivial, but producing verifiable, production-ready output remains the critical bottleneck. To manage this influx, engineering teams must pivot from prompt-based generation to rigorous AI observability and bespoke evaluation systems. Platforms are recording massive spikes in preview environments and error tracking as agents rapidly ship features. Ensuring code quality now requires specialized frameworks, such as frontier code evaluations, which codify human PR review standards to verify if generated code is genuinely mergeable. Furthermore, organizations must build custom evaluation layers that validate agent outputs against first principles, enabling agents to operate within sandboxed environments where they can iteratively test and verify their own work before deployment. The rise of autonomous coding dictates a transition toward application-defined compute, where agents dynamically provision infrastructure across cloud providers using machine-to-machine protocols rather than traditional graphical dashboards. This requires cloud vendors to enforce strict programmatic boundaries, automated rollbacks, and permission guardrails to safely contain agent actions. Simultaneously, as AI workloads handle sensitive enterprise data, securing the inference layer becomes paramount. True production safety demands end-to-end token-level encryption for open-source models, stripping away infrastructure-level vulnerabilities and ensuring that prompt data remains opaque to human administrators and third-party hosting environments. **Keywords:** ai code generation, production software debugging, pull request automation, ai observability, bespoke evaluation systems, frontier code evaluation, sandboxed execution environments, application-defined compute, dynamic cloud provisioning, token-level encryption, open-source model inference, automated code verification, infrastructure guardrails, machine-to-machine protocols, enterprise data privacy ## Chapters 1. **Tracking the exponential growth of software errors** (02:20) — An increase in AI-generated code directly leads to higher volumes of software errors and performance issues in production. 1. **Collaborating with AI agents for issue triage** (06:04) — Integrating AI agents into chat platforms enables automated triaging and pull request generation for incoming bug alerts. 1. **Deploying untested code and maintaining system understanding** (08:18) — The rapid generation of code via AI increases preview deployments while amplifying the risks of shipping unreviewed logic. 1. **Implementing end-to-end encryption for enterprise AI models** (11:24) — Securing tokens during inference prevents unauthorized infrastructure access and protects sensitive enterprise workloads from external breaches. 1. **Preparing cloud infrastructure for autonomous agent selection** (15:54) — Cloud providers must expose machine-friendly interfaces and robust guardrails to support agents that dynamically provision infrastructure. 1. **Implementing bespoke evaluation layers for AI observability** (20:25) — Establishing custom offline and online evaluation systems is necessary to track model hallucinations and validate code generation outcomes. 1. **Defining mergeability standards for AI-generated pull requests** (22:24) — High-quality software agents must test their own work in isolated sandboxes to ensure generated code meets human review standards. 1. **Moving toward application-defined compute for long-horizon agents** (26:56) — The rise of autonomous agents necessitates infrastructure that can dynamically provision compute resources and sustain long-running stateful tasks. 1. **Managing enterprise data privacy across global inference providers** (29:28) — Directing AI workloads to locally hosted and encrypted models mitigates data exposure and limits reliance on foreign infrastructure. ## Related Moments - [Managing the operational impact of AI-generated code volume](https://www.wearedevelopers.com/videos/100332-software-that-fixes-itself) (from "Software That Fixes Itself") - [High volume of AI-generated code entering production environments](https://www.wearedevelopers.com/videos/100512-the-broken-rung-how-ai-is-rebuilding-software-development-from-the-ground-up) (from "The Broken Rung: How AI is Rebuilding Software Development from the Ground Up") - [Shifting bottlenecks in the era of AI code generation](https://www.wearedevelopers.com/videos/100468-culture-doesn-t-scale-itself-leading-engineering-teams-through-hypergrowth-and-the-ai-transition) (from "Culture Doesn't Scale Itself: Leading Engineering Teams Through Hypergrowth and the AI Transition") - [Exploring modern AI SDKs and production deployment challenges](https://www.wearedevelopers.com/videos/1988-your-ai-agent-is-just-a-while-loop-with-an-api-call-let-me-prove-it) (from "Your AI Agent is just a while loop with an API call. Let me prove it") - [Shifting developer workloads and realistic AI productivity gains](https://www.wearedevelopers.com/videos/1830-wearedevelopers-live-speculaitions) (from "WeAreDevelopers LIVE - SpeculAItions") - [Moving beyond demos to build production-ready software](https://www.wearedevelopers.com/videos/100042-it-s-a-great-time-to-be-a-builder-leveraging-ai-for-good) (from "It's a Great Time to be a Builder: Leveraging AI for Good") ## Related Articles - [Transforming Software Development: The Role of AI and Developer Tools](https://www.wearedevelopers.com/magazine/527-transforming-software-development-the-role-of-ai-and-developer-tools) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) ## Related Jobs - [Senior AI Developer](https://www.wearedevelopers.com/jobs/ext/2836034-senior-ai-developer) at **PwC** - [Senior AI/ML Engineer](https://www.wearedevelopers.com/jobs/48352-senior-ai-ml-engineer) at **PagerDuty** - [Principal Software Engineer, AI Inference Cloud](https://www.wearedevelopers.com/jobs/ext/2854957-principal-software-engineer-ai-inference-cloud) at **ARM** - [Staff Software Engineer, AI Inference Cloud](https://www.wearedevelopers.com/jobs/ext/3347267-staff-software-engineer-ai-inference-cloud) at **ARM** - [Principal Software Engineer, AI Compute Platform](https://www.wearedevelopers.com/jobs/ext/2847710-principal-software-engineer-ai-compute-platform) at **ARM** - [AI Software Engineer (Germany)](https://www.wearedevelopers.com/jobs/48317-ai-software-engineer-germany) at **Sunhat**