> Markdown version of [/jobs/ext/3102041-sre-incident-response](https://www.wearedevelopers.com/jobs/ext/3102041-sre-incident-response). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE Incident Response - **Company:** Illumio - **Location:** Sunnyvale, CA, United States - **Experience:** Expert - **Salary:** $136,000.0 - $156,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Cloud Computing, DevOps, Distributed Systems, Monitoring of Systems, Reliability Engineering, Runbook, Google Cloud, Pagerduty - **Published:** September 27, 2026 - **Apply:** https://www.themuse.com/jobs/illumio/sr-sre-incident-response-bc6822 ## About the Role * 5+ years in incident management, SRE, operations, or related roles in fast-paced, production environments. * Strong understanding of infrastructure, distributed systems, and service monitoring. * Proven ability to lead under pressure and drive clarity during ambiguous, high-stakes situations. * Excellent verbal and written communication skills, including stakeholder and executive comms. * Experience with on-call practices, incident tooling (PagerDuty, Opsgenie, StatusPage, etc.), and observability platforms. * Familiarity with ITIL, SRE principles, and modern DevOps practices is a plus. Plus Factors: * Experience in a regulated or compliance-driven environment (e.g., SOC2, HIPAA, PCI). * Background in cloud infrastructure (Azure, AWS, GCP). * Hands-on experience with runbooks, automation, and reliability engineering. ## Description We're seeking an experienced Sr. SRE Incident Response Analyst to join our Site Reliability Engineering team. In this role, you'll be responsible for leading the response to high-impact incidents, ensuring clear communication, rapid resolution, and continuous improvement of our incident management practices. You will work closely with engineering, operations, product, and leadership teams to maintain service reliability and resilience. You'll be at the center of mission-critical operations, helping scale and mature our incident response function while working alongside a world-class SRE and engineering team. Your work will have a direct impact on customer trust, system resilience, and company success. * Own the end-to-end incident management lifecycle-from detection and triage to resolution and post-incident review. * Coordinate response efforts during critical incidents, including cross-functional teams and executive stakeholders. * Establish and enforce incident response protocols, including on-call processes and escalation procedures. * Lead post-incident reviews (PIRs) to identify root causes, document learnings, and drive follow-up action items. * Track incident trends and partner with SRE and Engineering to prioritize resilience improvements. * Ensure timely and accurate communication during incidents, both internally and externally (if needed). * Contribute to tooling and automation that improve detection, response, and visibility. * Provide guidance and training to engineers and teams on effective incident handling. ## Related Videos - [AI in Leadership: How Technology is Reshaping Executive Roles](https://www.wearedevelopers.com/videos/1705-ai-in-leadership-how-technology-is-reshaping-executive-roles) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Shifting Stress to Progress— Understanding DevOps to do DevOps Better](https://www.wearedevelopers.com/videos/268-shifting-stress-to-progress-understanding-devops-to-do-devops-better) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [3 Key Steps for Optimizing DevOps Workflows](https://www.wearedevelopers.com/videos/962-3-key-steps-for-optimizing-devops-workflows) - [Bridging AI and Nomad: a Go-based MCP Server for Cluster Control](https://www.wearedevelopers.com/videos/2063-bridging-ai-and-nomad-a-go-based-mcp-server-for-cluster-control) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [The Geometry of Incidents: Connecting User Impact to Architecture](https://www.wearedevelopers.com/magazine/764-the-geometry-of-incidents-connecting-user-impact-to-architecture) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)