World Congress 2022 Jun 15, 2022

What we Learned from Reading 100+ Kubernetes Post-Mortems

Noaa Barki

Subtle YAML errors and missing memory limits are silently crashing your Kubernetes clusters. Uncover the leading root causes from 100+ post-mortems and automate infrastructure validation directly within CI workflows.

Pause
Mute Enter Fullscreen
#1 about 7 min

Friction between developers and infrastructure teams

Developers prioritize feature delivery over infrastructure concepts like memory limits and resource configurations.

#2 about 2 min

Finding developers to champion infrastructure best practices

Identifying team members interested in both code and infrastructure helps bridge the knowledge gap.

#3 about 4 min

Common Kubernetes misconfigurations in production environments

Real-world examples show how missing concurrency policies or memory limits cause major performance outages.

#4 about 4 min

Rolling out policy enforcement to engineering teams slowly

Gradually introducing automated validations in local and CI environments prevents friction while protecting production.

#5 about 3 min

Three primary categories of Kubernetes infrastructure failures

Most production outages stem from YAML syntax errors, misunderstanding tool defaults, or misaligned internal practices.

#6 about 3 min

Validating resource syntax and schema without cluster access

Tools like yq and kubeconform ensure basic file format and schema correctness before deployment.

#7 about 4 min

Managing infrastructure policies and guiding developers effectively

Centralizing policy management and providing detailed remediation steps prevents developer frustration when builds fail.

#8 about 5 min

Automating customized policy enforcement with open source tools

Implementing tools like Datree or Conftest allows centralized policy control while enabling local developer testing.

Matching moments

4:40 min

Assessing common Kubernetes security incidents and misconfigurations

Rico Komenda Rico Komenda · WWC 2025

1:36 min

Automating infrastructure to eliminate traditional DevOps roles

Andrew Holway · LIVE

1:39 min

Introduction to Kubernetes security challenges and opportunities

Marc Nimmerrichter · WWC 2022

4:43 min

Managing the complexity of bare metal Kubernetes deployments

Josip Stuhli Josip Stuhli · WWC Europe 2026

3:54 min

Managing Kubernetes clusters using command line and YAML

Hannes Norbert Göring · LIVE

1:54 min

Managing deployments with Kubernetes and Terraform self-healing

Thomas Hartenstein · LIVE

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Boring Failover: Predictable Region Recovery Across 5,000 Microservices

Garvit Kataria, Sahil Sabharwal

Garvit Kataria
Sahil Sabharwal
Open session

World Congress 2026 North America

Run your agents in Kubernetes: Build once, deploy anywhere. But really?

Michal Salanci

Senior Systems Engineer at ESET Cybersecurity

Michal Salanci
Open session

World Congress 2026 North America

From Static Rules to Reasoning Platforms: Scaling Intelligent Canary Delivery in 2026

Daniel Oh

Senior Principal Developer Advocate

Daniel Oh
Open session

World Congress 2026 North America

Stop Running Mystery Meat in Production

Jeroen van Erp

Technical Advocate @ SUSE

Jeroen van Erp
Open session

World Congress 2026 North America

rm -rf: Horror Stories From Unsandboxed AI Agents (and How Docker Fixes This)

Rishab Kumar

Staff Developer Evangelist @ Twilio

Rishab Kumar
Open session

World Congress 2026 North America

The Geometry of Incidents: What User-Impact Shapes Reveal About Platform Architecture

Bala Subrahmanyam Kambala

Staff Platform Engineer at Oracle Cloud Infrastructure

Bala Subrahmanyam Kambala