World Congress 2022 Jun 15, 2022

I broke the production

Arto Liukkonen

A new hire brought down production with a rogue script. Discover why resilient engineering teams embrace blameless post-mortems to fortify systems instead of hunting for a scapegoat.

Pause
Mute Enter Fullscreen
#1 about 3 min

Recognizing the inevitability of breaking production environments

Sharing early career failures normalizes the reality that all developers will eventually cause widespread system chaos.

#2 about 3 min

Managing constant production updates at massive data scale

Processing petabytes of advertising data requires tightly managed continuous deployment pipelines supporting dozens of daily updates.

#3 about 3 min

Triggering unexpected outages during restrictive data backfills

Attempting massive data backfills directly in production due to compliance restrictions can easily trigger severe monitoring spikes.

#4 about 3 min

Reconciling individual actions with good technical intentions

Overcoming the human tendency to judge coworkers by actions rather than intentions fosters fairer incident assessments.

#5 about 3 min

Recognizing systemic flaws over pointing fingers at developers

Treating deployment outages as failures of the pull request process rather than individual mistakes prevents toxic blame loops.

#6 about 2 min

Fostering self awareness to prevent toxic team behaviors

Cultivating personal responsibility and empathy stops frustration from cascading downward into harmful workplace chain reactions.

#7 about 2 min

Implementing blameless postmortems to strengthen system resilience

Borrowing practices from the healthcare and aviation industries turns every technical mistake into an opportunity for infrastructure improvement.

#8 about 4 min

Applying positive feedback ratios to technical code reviews

Leverging a five-to-one positive feedback model during pull requests builds team trust and breaks the cycle of purely critical interactions.

#9 about 2 min

Anticipating systemic failures using proactive premortem exercises

Simulating disaster scenarios before code ships helps uncover hidden edge cases and vulnerabilities within complex architectures.

#10 about 4 min

Embracing organizational learning following stressful production incidents

Navigating extended recovery windows proves that eliminating blame allows teams to focus entirely on rapid system restoration.

#11 about 3 min

Handling customer impact and mitigating third party failures

Buffering incidents through customer support provides necessary space for engineers while past external routing failures offer valuable historical lessons.

Matching moments

2:08 min

Navigating and mitigating the impacts of broken engineering cultures

Oskar Kruschitz Oskar Kruschitz · Europe 2026 Virtual

1:49 min

The feedback loop of leadership choices and technical outcomes

Oskar Kruschitz Oskar Kruschitz · Europe 2026 Virtual

3:03 min

Building a corporate culture that actively celebrates failure

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

3:34 min

Writing blameless and detailed incident postmortems

Martin Beránek · LIVE

1:54 min

Embracing failures and analyzing high-profile software engineering mistakes

Christian Seifert · WWC 2022

1:02 min

Establishing a constructive failure culture for continuous learning

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

Upcoming sessions on this topic

Open session

World Congress 2026 North America

It passed auth, then production caught fire

Alex Olivier

Co-founder & CPO @ Cerbos | OpenID AuthZEN Co-chair

Alex Olivier
Open session

World Congress 2026 North America

The Geometry of Incidents: What User-Impact Shapes Reveal About Platform Architecture

Bala Subrahmanyam Kambala

Staff Platform Engineer at Oracle Cloud Infrastructure

Bala Subrahmanyam Kambala
Open session

World Congress 2026 North America

Culture Doesn't Scale Itself: Leading Engineering Teams Through Hypergrowth and the AI Transition

Thanos Baskous

VP of Engineering and Co-founder of Cogent Security

Thanos Baskous
Open session

World Congress 2026 North America

The Autonomous Performance Agent: A Netflix Production Story

Rajat Shah

Staff Software Engineer, AI Platform, Netflix

Rajat Shah
Open session

World Congress 2026 North America

Evals Are Infra: Building AI Systems Developers Can Actually Trust

Phoebe Wang

Member of Technical Staff at OpenAI

Phoebe Wang
Open session

World Congress 2026 North America

Boring Failover: Predictable Region Recovery Across 5,000 Microservices

Garvit Kataria, Sahil Sabharwal

Garvit Kataria
Sahil Sabharwal