WeAreDevelopers LIVE Dec 16, 2020

Applying Agile Principles to Incident Management

Tobias Dunn-Krahn

What if you treated incident response like a compressed agile sprint? Discover how to transform unstructured outages into automated workflows that build lasting system resilience.

Pause
Mute Enter Fullscreen
#1 about 7 min

Defining digital service incidents and key technical stakeholders

An overview of digital service interruptions and the various teams responsible for their remediation.

#2 about 4 min

Applying software development methodologies to incident response

Adapting core frameworks like agile iterations, devops culture, and sre automation to manage technical crises.

#3 about 3 min

Practicing incident management through failure Friday exercises

How regular, deliberate fault injection builds team confidence and refines remediation processes.

#4 about 3 min

Establishing a simulated technical environment for the workflow demo

Setting up a corporate architecture complete with standard devops toolchains to demonstrate platform integration.

#5 about 5 min

Automating critical diagnostics and major incident creation

Triggering preliminary system checks and multi-platform communication workflows immediately upon anomaly detection.

#6 about 6 min

Navigating the incident timeline and dynamic resolver engagement

Managing the remediation state while coordinating with necessary service-level specialists as new details emerge.

#7 about 3 min

Distilling preventative action items during post-incident review

Curating timeline logs and capturing tasks to ensure incident retrospectives yield concrete system enhancements.

#8 about 9 min

Customizing automation pipelines via the visual flow designer

Building custom toolchain integrations by parsing JSON payloads and mapping multi-step response execution graphs.

Matching moments

3:08 min

Shifting software delivery bottlenecks to operations and incident response

Milin Desai Milin Desai +1 · World Congress 2026 Europe

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

1:57 min

Streamlining incident response and root cause analysis automatically

Mike Mike · World Congress 2025

1:43 min

Integrating intelligent support into modern agile practices

Christopher May Christopher May · Europe 2026 Virtual

6:03 min

Engaging software developers deeply in secure engineering practices

Tanya Janca · World Congress 2021

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 25, 2026 · 09:00–09:30

Stage 5

Reinventing Incident Response with AI Agents and MCP

Jayant Tyagi

Lead Member of Technical Staff at Salesforce

Jayant Tyagi
Open session

World Congress 2026 North America

September 25, 2026 · 13:30–14:00

Outdoor Stage

The Geometry of Incidents: What User-Impact Shapes Reveal About Platform Architecture

Bala Subrahmanyam Kambala

Staff Platform Engineer at Oracle Cloud Infrastructure

Bala Subrahmanyam Kambala
Open session

World Congress 2026 North America

September 25, 2026 · 15:30–16:00

Stage 1

AI-Powered Incident Triage: How We Built GenAI Agents with MCPs to Automate On-Call Workflows

Prakshal Doshi

Site Reliability Engineer

Prakshal Doshi
Open session

World Congress 2026 North America

September 24, 2026 · 17:30–18:00

Stage 4

Boring Failover: Predictable Region Recovery Across 5,000 Microservices

Garvit Kataria, Sahil Sabharwal

Garvit Kataria
Sahil Sabharwal
Open session

World Congress 2026 North America

September 25, 2026 · 09:00–09:30

Stage 1

Test Before Release, Enforce at Runtime: Governance for Tool-Using AI Agents

Sachin Gupta

Member of Technical Staff 2 at eBay

Sachin Gupta
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 4

Agentic Drift: keeping pace with your agents

John Coghlan

Senior Director, Developer Advocacy at GitLab

John Coghlan