World Congress 2022 Jun 15, 2022

What we Learned from Reading 100+ Kubernetes Post-Mortems

Noaa Barki

Subtle YAML errors and missing memory limits are silently crashing your Kubernetes clusters. Uncover the leading root causes from 100+ post-mortems and automate infrastructure validation directly within CI workflows.

Pause
Mute Enter Fullscreen
#1 about 7 min

Friction between developers and infrastructure teams

Developers prioritize feature delivery over infrastructure concepts like memory limits and resource configurations.

#2 about 2 min

Finding developers to champion infrastructure best practices

Identifying team members interested in both code and infrastructure helps bridge the knowledge gap.

#3 about 4 min

Common Kubernetes misconfigurations in production environments

Real-world examples show how missing concurrency policies or memory limits cause major performance outages.

#4 about 4 min

Rolling out policy enforcement to engineering teams slowly

Gradually introducing automated validations in local and CI environments prevents friction while protecting production.

#5 about 3 min

Three primary categories of Kubernetes infrastructure failures

Most production outages stem from YAML syntax errors, misunderstanding tool defaults, or misaligned internal practices.

#6 about 3 min

Validating resource syntax and schema without cluster access

Tools like yq and kubeconform ensure basic file format and schema correctness before deployment.

#7 about 4 min

Managing infrastructure policies and guiding developers effectively

Centralizing policy management and providing detailed remediation steps prevents developer frustration when builds fail.

#8 about 5 min

Automating customized policy enforcement with open source tools

Implementing tools like Datree or Conftest allows centralized policy control while enabling local developer testing.

Matching moments

4:40 min

Assessing common Kubernetes security incidents and misconfigurations

Rico Komenda Rico Komenda · World Congress 2025

1:36 min

Automating infrastructure to eliminate traditional DevOps roles

Andrew Holway · LIVE

1:39 min

Introduction to Kubernetes security challenges and opportunities

Marc Nimmerrichter · World Congress 2022

4:43 min

Managing the complexity of bare metal Kubernetes deployments

Josip Stuhli Josip Stuhli · World Congress 2026 Europe

3:54 min

Managing Kubernetes clusters using command line and YAML

Hannes Norbert Göring · LIVE

1:54 min

Managing deployments with Kubernetes and Terraform self-healing

Thomas Hartenstein · LIVE

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 24, 2026 · 17:30–18:00

Stage 4

Boring Failover: Predictable Region Recovery Across 5,000 Microservices

Garvit Kataria, Sahil Sabharwal

Garvit Kataria
Sahil Sabharwal
Open session

World Congress 2026 North America

September 24, 2026 · 15:30–16:00

Stage 9

Run your agents in Kubernetes: Build once, deploy anywhere. But really?

Michal Salanci

Senior Systems Engineer at ESET Cybersecurity

Michal Salanci
Open session

World Congress 2026 North America

September 24, 2026 · 10:20–10:50

Stage 2

From Static Rules to Reasoning Platforms: Scaling Intelligent Canary Delivery in 2026

Daniel Oh

Senior Principal Developer Advocate

Daniel Oh
Open session

World Congress 2026 North America

September 24, 2026 · 11:40–12:10

Stage 2

Stop Running Mystery Meat in Production

Jeroen van Erp

Technical Advocate @ SUSE

Jeroen van Erp
Open session

World Congress 2026 North America

September 25, 2026 · 10:20–10:50

Stage 4

rm -rf: Horror Stories From Unsandboxed AI Agents (and How Docker Fixes This)

Rishab Kumar

Staff Developer Evangelist @ Twilio

Rishab Kumar
Open session

World Congress 2026 North America

September 25, 2026 · 13:30–14:00

Outdoor Stage

The Geometry of Incidents: What User-Impact Shapes Reveal About Platform Architecture

Bala Subrahmanyam Kambala

Staff Platform Engineer at Oracle Cloud Infrastructure

Bala Subrahmanyam Kambala