SRE Incident Response

Illumio
Sunnyvale, CA, United States
9 days ago
Apply on www.themuse.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$136,000.0 - $156,000.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Cloud Computing DevOps Distributed Systems Monitoring of Systems Reliability Engineering Runbook Google Cloud Pagerduty

Job description

We’re seeking an experienced Sr. SRE Incident Response Analyst to join our Site Reliability Engineering team. In this role, you’ll be responsible for leading the response to high-impact incidents, ensuring clear communication, rapid resolution, and continuous improvement of our incident management practices. You will work closely with engineering, operations, product, and leadership teams to maintain service reliability and resilience.

You’ll be at the center of mission-critical operations, helping scale and mature our incident response function while working alongside a world-class SRE and engineering team. Your work will have a direct impact on customer trust, system resilience, and company success.

  • Own the end-to-end incident management lifecycle-from detection and triage to resolution and post-incident review.
  • Coordinate response efforts during critical incidents, including cross-functional teams and executive stakeholders.
  • Establish and enforce incident response protocols, including on-call processes and escalation procedures.
  • Lead post-incident reviews (PIRs) to identify root causes, document learnings, and drive follow-up action items.
  • Track incident trends and partner with SRE and Engineering to prioritize resilience improvements.
  • Ensure timely and accurate communication during incidents, both internally and externally (if needed).
  • Contribute to tooling and automation that improve detection, response, and visibility.
  • Provide guidance and training to engineers and teams on effective incident handling.

Requirements

  • 5+ years in incident management, SRE, operations, or related roles in fast-paced, production environments.
  • Strong understanding of infrastructure, distributed systems, and service monitoring.
  • Proven ability to lead under pressure and drive clarity during ambiguous, high-stakes situations.
  • Excellent verbal and written communication skills, including stakeholder and executive comms.
  • Experience with on-call practices, incident tooling (PagerDuty, Opsgenie, StatusPage, etc.), and observability platforms.
  • Familiarity with ITIL, SRE principles, and modern DevOps practices is a plus.

Plus Factors:

  • Experience in a regulated or compliance-driven environment (e.g., SOC2, HIPAA, PCI).
  • Background in cloud infrastructure (Azure, AWS, GCP).
  • Hands-on experience with runbooks, automation, and reliability engineering.

Benefits & conditions

$136,000 USD - $156,000 USD, At Illumio we offer a wide range of benefits to our eligible team members. Our benefit programs vary by location and can include Medical, Dental, Vision Coverage - Health and Dependent Savings Accounts - Life and Disability Programs - Paid Parental Leave - Voluntary Benefit Programs - Company Sponsored Wellness Program - Wellness Reimbursement Program - Retirement Savings - Equity Opportunities - Paid time off and Paid Holidays - Employee Incentive Program. #LI-KD1 #LI-ONSITE

About the company

Illumio is the leader in ransomware and breach containment, redefining how organizations contain cyberattacks and enable operational resilience. Powered by the Illumio AI Security Graph, our breach containment platform identifies and contains threats across hybrid multi-cloud environments - stopping the spread of attacks before they become disasters.

Recognized as a Leader in the Forrester Wave for Microsegmentation, Illumio enables Zero Trust, strengthening cyber resilience for the infrastructure, systems, and organizations that keep the world running.

Our Team’s Vision:

Our Engineering team is driven by a culture that thrives on visionary leadership, autonomy, and ownership, creating a dynamic synergy that drives us forward in the ever-evolving landscape of cybersecurity.

When you join our team, you become part of the leader in Zero Trust Segmentation. You’ll work with a cutting-edge technology stack that spans operating systems, distributed applications, and immersive UI/visualization tools.

We’re shaping the future of cybersecurity. And together, we will continue to build world-class products-led by people with different perspectives, backgrounds, and a commitment to innovation in a time when the world faces its greatest cybersecurity threats in history., Illumio believes that an environment of unique backgrounds, experiences, viewpoints, and individual contributions drives our success and makes us stronger together. We are dedicated to creating and maintaining a diverse culture and emphasizing inclusion and belonging.

All official job offers from our company are extended directly by our recruitment team and will be sent through an official DocuSign document for your review and signature. Please be aware that we do not ask for any personal information in the process of extending offers of employment, such as financial details or social security numbers. Upon acceptance of any offer, we will request such information as part of the onboarding process prior to or on your first day of employment, and only after completing a background check through an authorized third-party vendor. If you receive any communication asking for personal details outside of these processes, please contact us immediately to verify the authenticity of the request. Your security is important to us, and we are committed to a safe and transparent hiring experience.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.themuse.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:53 min

Applying software development methodologies to incident response

Tobias Dunn-Krahn · LIVE

1:57 min

Streamlining incident response and root cause analysis automatically

Mike Mike · World Congress 2025

2:50 min

Introduction and the value of runbooks

Hila Fish · World Congress 2023

1:22 min

Analyzing differences between mobile and traditional backend DevOps

Mete Baydar Mete Baydar · World Congress 2025

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

3:18 min

Automating service provisioning through developer self-service tools

Llywelyn Griffith-Swain · World Congress 2023

Videos

See all

Related articles

See all