> Markdown version of [/videos/484-i-broke-the-production](https://www.wearedevelopers.com/videos/484-i-broke-the-production). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # I broke the production A new hire brought down production with a rogue script. Discover why resilient engineering teams embrace blameless post-mortems to fortify systems instead of hunting for a scapegoat. - **Speakers:** Arto Liukkonen - **Event:** World Congress 2022 - **Published:** June 15, 2022 - **Duration:** 26:18 - **URL:** https://www.wearedevelopers.com/videos/484-i-broke-the-production ## Summary A developer shares a firsthand account of causing a massive production outage at an ad-tech company shortly after being hired. While running a data backfilling script against the Facebook API, unexpected cascading failures halted customer reporting and exhausted API limits. Despite the initial chaos, the focus of the story is not on the technical failure itself, but on how organizations should respond to and learn from catastrophic incidents. The core theme centers on establishing a "blameless culture," a concept borrowed from high-stakes industries like aviation and healthcare. The speaker challenges engineering teams to judge colleagues by their underlying intentions rather than their mistaken actions, emphasizing that no developer sets out to intentionally break production. If a single pull request or rogue script brings down the system, it represents a structural flaw, not a personal one. Teams are encouraged to adopt blameless post-mortems—following the incident response methodologies of companies like Google and Atlassian—using mistakes as valuable opportunities to fortify system resilience instead of hunting for a scapegoat. To cultivate this environment proactively, development teams can introduce pre-mortems to map out potential edge cases and anticipate failures before deployment. Furthermore, the speaker highlights the psychological importance of positive code reviews. Because human memory disproportionately retains negative feedback, developers should strive for a high ratio of positive comments to constructive criticism in pull requests. Ultimately, when production inevitably breaks, developers should feel safe enough to own their mistakes immediately, trusting that a supportive engineering culture will transform the outage into an organizational capability upgrade. **Keywords:** blameless culture, production outages, incident management, blameless post-mortems, psychological safety in engineering, positive code reviews, pre-mortem analysis, system resilience, chaos engineering concepts, API rate limiting, Grafana monitoring, pull request workflows, root cause analysis, human error in software development ## Chapters 1. **Recognizing the inevitability of breaking production environments** (00:05) — Sharing early career failures normalizes the reality that all developers will eventually cause widespread system chaos. 1. **Managing constant production updates at massive data scale** (02:25) — Processing petabytes of advertising data requires tightly managed continuous deployment pipelines supporting dozens of daily updates. 1. **Triggering unexpected outages during restrictive data backfills** (04:36) — Attempting massive data backfills directly in production due to compliance restrictions can easily trigger severe monitoring spikes. 1. **Reconciling individual actions with good technical intentions** (06:48) — Overcoming the human tendency to judge coworkers by actions rather than intentions fosters fairer incident assessments. 1. **Recognizing systemic flaws over pointing fingers at developers** (09:03) — Treating deployment outages as failures of the pull request process rather than individual mistakes prevents toxic blame loops. 1. **Fostering self awareness to prevent toxic team behaviors** (11:56) — Cultivating personal responsibility and empathy stops frustration from cascading downward into harmful workplace chain reactions. 1. **Implementing blameless postmortems to strengthen system resilience** (13:42) — Borrowing practices from the healthcare and aviation industries turns every technical mistake into an opportunity for infrastructure improvement. 1. **Applying positive feedback ratios to technical code reviews** (15:40) — Leverging a five-to-one positive feedback model during pull requests builds team trust and breaks the cycle of purely critical interactions. 1. **Anticipating systemic failures using proactive premortem exercises** (19:04) — Simulating disaster scenarios before code ships helps uncover hidden edge cases and vulnerabilities within complex architectures. 1. **Embracing organizational learning following stressful production incidents** (20:41) — Navigating extended recovery windows proves that eliminating blame allows teams to focus entirely on rapid system restoration. 1. **Handling customer impact and mitigating third party failures** (24:07) — Buffering incidents through customer support provides necessary space for engineers while past external routing failures offer valuable historical lessons. ## Related Moments - [Navigating and mitigating the impacts of broken engineering cultures](https://www.wearedevelopers.com/videos/1998-from-code-to-culture-why-leadership-determines-software-quality) (from "From Code to Culture: Why Leadership Determines Software Quality") - [The feedback loop of leadership choices and technical outcomes](https://www.wearedevelopers.com/videos/1998-from-code-to-culture-why-leadership-determines-software-quality) (from "From Code to Culture: Why Leadership Determines Software Quality") - [Building a corporate culture that actively celebrates failure](https://www.wearedevelopers.com/videos/1778-unlearning-is-the-new-learning-the-skills-shaping-the-future-of-work) (from "Unlearning Is the New Learning: The Skills Shaping the Future of Work") - [Writing blameless and detailed incident postmortems](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) (from "SRE Methods In an Agency Environment") - [Embracing failures and analyzing high-profile software engineering mistakes](https://www.wearedevelopers.com/videos/423-the-software-bug-all-stars-and-what-we-can-learn-from-them) (from "The Software Bug All Stars - and what we can learn from them") - [Establishing a constructive failure culture for continuous learning](https://www.wearedevelopers.com/videos/1852-a-players-what-it-really-takes-to-attract-them-and-keep-them) (from "A-Players: What It Really Takes to Attract Them and Keep Them") ## Related Articles - [Never delegate the understanding](https://www.wearedevelopers.com/magazine/749-never-delegate-the-understanding) - [Now is the time for industrialized software development](https://www.wearedevelopers.com/magazine/601-now-is-the-time-for-industrialized-software-development) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Exploring AI: Opportunities and Risks for Developers](https://www.wearedevelopers.com/magazine/522-exploring-ai-opportunities-and-risks-for-developers) ## Related Jobs - [Senior Engineer, Infrastructure Platform](https://www.wearedevelopers.com/jobs/ext/328836-senior-engineer-infrastructure-platform) at **Intercom, Inc.** - [Tribe Lead - ( Software) Engineering Centre of Excllence](https://www.wearedevelopers.com/jobs/ext/1475530-tribe-lead-software-engineering-centre-of-excllence) at **SD Worx** - [Principal Product Manager, Agent Platform](https://www.wearedevelopers.com/jobs/ext/277541-principal-product-manager-agent-platform) at **GitHub** - [Staff Developer Advocate, GitHub Security Lab](https://www.wearedevelopers.com/jobs/ext/1921051-staff-developer-advocate-github-security-lab) at **GitHub** - [Staff Software Engineer, Copilot Experiences](https://www.wearedevelopers.com/jobs/ext/164361-staff-software-engineer-copilot-experiences) at **GitHub** - [Senior Software Engineer, Enterprise Products](https://www.wearedevelopers.com/jobs/ext/1841248-senior-software-engineer-enterprise-products) at **GitHub**