> Markdown version of [/videos/101-applying-agile-principles-to-incident-management?t=390](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management?t=390). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Applying Agile Principles to Incident Management What if you treated incident response like a compressed agile sprint? Discover how to transform unstructured outages into automated workflows that build lasting system resilience. - **Speakers:** Tobias Dunn-Krahn - **Event:** WeAreDevelopers LIVE - **Published:** December 16, 2020 - **Duration:** 37:33 - **URL:** https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management ## Summary Treating incident management natively as a compressed iteration of software development allows engineering teams to resolve outages and SLO breaches systematically. By applying Agile, DevOps, and Site Reliability Engineering (SRE) principles, organizations can transform unstructured emergencies into controlled, continuously improving workflows. Just as Agile prioritizes shipping working increments, effective incident response focuses on mitigating customer impact sequentially rather than withholding communication until a complete fix is deployed. Furthermore, incorporating Scrum's retrospective model ensures that every anomaly, performance degradation, or service disruption feeds directly into future system resilience. Bridging the cultural divide between operations, development, and customer support teams is critical during a major incident. To facilitate this communication without friction, DevOps philosophies emphasize meeting teams inside their native tooling. This means automatically generating Jira issues, ServiceNow tickets, Slack channels, and public status pages the moment a critical alert is triggered. To optimize this collaborative muscle, adopting proactive methodologies like "Failure Fridays"—intentionally sabotaging non-production environments that mirror live infrastructure—gives on-call resources a safe space to practice incident command. Deliberate chaos engineering effectively tests telemetry precision, identifies operational bottlenecks, and builds deep team confidence. When minutes count, executing advanced diagnostic automation preemptively shaves precious time off the resolution clock. Customized workflows can intercept telemetry alerts, query queuing infrastructure metrics, and correlate recent deployment commits before an on-call responder is even explicitly paged. By utilizing customized JavaScript scripts to parse JSON payloads and establish dynamic cross-linking between incident platforms, developers can easily eliminate manual coordination tasks. Ultimately, engineering teams can prioritize diagnosing complex microservice failures, executing rapid rollbacks, and leading impactful post-incident reviews instead of scrambling to manage internal communication channels. **Keywords:** agile incident management, SRE automation practices, SLO breach remediation, devops cultural bridging, incident diagnostic automation, post-incident retrospectives, chaos engineering exercises, cross-functional triage workflows, microservice degradation tracking, JSON alert parsing, servicenow ticket automation, on-call workflow integration, dynamic resolver escalation, telemetry anomaly detection, continuous process improvement ## Chapters 1. **Defining digital service incidents and key technical stakeholders** (00:17) — An overview of digital service interruptions and the various teams responsible for their remediation. 1. **Applying software development methodologies to incident response** (06:30) — Adapting core frameworks like agile iterations, devops culture, and sre automation to manage technical crises. 1. **Practicing incident management through failure Friday exercises** (10:23) — How regular, deliberate fault injection builds team confidence and refines remediation processes. 1. **Establishing a simulated technical environment for the workflow demo** (13:07) — Setting up a corporate architecture complete with standard devops toolchains to demonstrate platform integration. 1. **Automating critical diagnostics and major incident creation** (15:35) — Triggering preliminary system checks and multi-platform communication workflows immediately upon anomaly detection. 1. **Navigating the incident timeline and dynamic resolver engagement** (20:29) — Managing the remediation state while coordinating with necessary service-level specialists as new details emerge. 1. **Distilling preventative action items during post-incident review** (26:12) — Curating timeline logs and capturing tasks to ensure incident retrospectives yield concrete system enhancements. 1. **Customizing automation pipelines via the visual flow designer** (28:34) — Building custom toolchain integrations by parsing JSON payloads and mapping multi-step response execution graphs. ## Related Moments - [Shifting software delivery bottlenecks to operations and incident response](https://www.wearedevelopers.com/videos/100332-software-that-fixes-itself) (from "Software That Fixes Itself") - [Integrating service level objectives into incident management](https://www.wearedevelopers.com/videos/854-serverless-observability-where-slos-meet-transforms) (from "Serverless Observability: where SLOs meet transforms") - [Streamlining incident response and root cause analysis automatically](https://www.wearedevelopers.com/videos/1539-agentic-devops-how-ai-powered-automation-transforms-software-delivery-on-github-and-azure) (from "Agentic DevOps: How AI-Powered Automation Transforms Software Delivery on GitHub and Azure") - [Integrating intelligent support into modern agile practices](https://www.wearedevelopers.com/videos/2081-ai-and-agility-the-dynamic-duo-for-disruption) (from "AI and Agility: The Dynamic Duo for Disruption") - [Engaging software developers deeply in secure engineering practices](https://www.wearedevelopers.com/videos/193-building-security-champions) (from "Building Security Champions") - [Navigating specialized roles and toolsets across engineering teams](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) (from "Handling incidents collaboratively is like solving a rubix cube") ## Related Articles - [Now is the time for industrialized software development](https://www.wearedevelopers.com/magazine/601-now-is-the-time-for-industrialized-software-development) - [Events like RSAC Get You CISOs. Developers Decide What Actually Gets Deployed.](https://www.wearedevelopers.com/magazine/693-events-like-rsac-get-you-cisos-developers-decide-what-actually-gets-deployed) - [Walking Into The Era of Supply Chain Risks](https://www.wearedevelopers.com/magazine/106-walking-into-the-era-of-supply-chain-risks) - [Exploring AI: Opportunities and Risks for Developers](https://www.wearedevelopers.com/magazine/522-exploring-ai-opportunities-and-risks-for-developers) ## Related Jobs - [Senior Security Engineer, Incident Response](https://www.wearedevelopers.com/jobs/ext/1346434-senior-security-engineer-incident-response) at **Twilio** - [Senior Security Engineer, Incident Response](https://www.wearedevelopers.com/jobs/ext/1347114-senior-security-engineer-incident-response) at **Twilio** - [Senior Engineer, Infrastructure Platform](https://www.wearedevelopers.com/jobs/ext/328836-senior-engineer-infrastructure-platform) at **Intercom, Inc.** - [Security Engineer, Incident Response](https://www.wearedevelopers.com/jobs/ext/1249908-security-engineer-incident-response) at **Twilio** - [Engineer, Offensive Security Organization](https://www.wearedevelopers.com/jobs/ext/1992296-engineer-offensive-security-organization) at **Twilio** - [Senior Software Engineer, Enterprise Products](https://www.wearedevelopers.com/jobs/ext/1841248-senior-software-engineer-enterprise-products) at **GitHub**