Site Reliability Engineer (SRE)

Illumio
Sunnyvale, CA, United States
10 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Software as a Service Cloud Computing Computer Programming Continuous Delivery Continuous Integration DevOps Monitoring of Systems Internet Security Python (Programming Language)
+16 more
Network Security Performance Tuning Windows PowerShell Reliability Engineering Cloud Services Data Logging Scripting Malware HybridCloud Containerization Gitlab-ci Kubernetes Information Technology Docker Jenkins Microservices

Job description

We are looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background in AWS & Azure cloud platforms to play a key role in ensuring the reliability, scalability, and performance of our cloud-based systems and applications.

The ideal candidate will have hands-on experience in supporting, and managing AWS and Azure infrastructure, along with a passion for automation, continuous improvement, and collaboration with cross-functional teams.

If you are passionate about AWS and/or Azure cloud platform and have a track record of driving reliability, scalability, and performance in cloud-based environments, we’d love to hear from you. Apply now to be a part of our talented team!

  • Monitor system performance, application health, and infrastructure metrics using monitoring and logging services, and implement proactive measures to optimize performance and availability
  • Oncall duty for production uptime and support for customer escalations
  • Release upgrades and maintenance activities including hotfixes and infrastructure updates
  • Lead incident response and resolution efforts, conducting root cause analysis, implementing corrective actions, and documenting post-incident reviews
  • Implement security best practices and controls in the cloud environments to protect data, applications, and infrastructure, and ensure compliance with regulatory requirements
  • Drive continuous improvement initiatives to enhance reliability, scalability, and efficiency of infrastructure and services, leveraging automation and emerging technologies

Requirements

  • Bachelor’s degree in computer science, Engineering, or related field; or equivalent work experience
  • 5+ years of experience working as a Site Reliability Engineer (SRE) or similar role, with a focus on AWS and/or Azure cloud platform
  • Hands-on experience in designing, deploying, and managing AWS and/or Azure infrastructure, including compute, storage, networking, and security services
  • Proficiency in scripting and programming languages such as PowerShell, Python, or Go for automation and infrastructure management tasks
  • Strong understanding of CI/CD principles and experience with tools such as Azure DevOps, Jenkins, or GitLab CI/CD
  • Experience with containerization technologies (e.g., Docker, Kubernetes) and microservices architecture in AWS and Azure environments is a plus
  • Excellent analytical, problem-solving, and communication skills, with the ability to collaborate effectively with cross-functional teams
  • AWS or Azure certifications such as AWS/Azure Solutions Architect, Azure DevOps Engineer, or Azure Security Engineer are preferred, Amazon Web Services (AWS), Analysis Skills, Artificial Intelligence (AI), Automation, Cloud Computing, Communication Skills, Computer Science, Continuous Deployment/Delivery, Continuous Improvement, Continuous Integration, Corrective Action, Cross-Functional, Customer Escalations, Customer Support/Service, DevOps, Docker, Documentation, Emerging Technology, Hybrid Cloud, Incident Management, Incident Response, Internet Security, Jenkins, Leadership, Metrics, Microservices, Microsoft Windows Azure, Network Security, On Call, Performance Analysis, Performance Tuning/Optimization, Philosophy, Problem Solving Skills, Process Improvement, Production Support, Python Programming/Scripting Language, Ransomware, Reliability Engineering, Root Cause Analysis, Scripting (Scripting Languages), Security Attacks, Software as a Service (SaaS), Team Player, Windows PowerShell

About the company

Title: Sr. Site Reliability Engineer Organization: Illumio Location: Sunnyvale Description: Onwards Together!

Illumio is the leader in ransomware and breach containment, redefining how organizations contain cyberattacks and enable operational resilience. Powered by the Illumio AI Security Graph, our breach containment platform identifies and contains threats across hybrid multi-cloud environments - stopping the spread of attacks before they become disasters.Recognized as a Leader in the Forrester Wave for Microsegmentation, Illumio enables Zero Trust, strengthening cyber resilience for the infrastructure, systems, and organizations that keep the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters.Our Team’s Vision:

Our Engineering team is shaping the future of cybersecurity. We thrive on visionary leadership, autonomy, and ownership, fostering a culture of innovation that propels us forward in the ever-evolving cybersecurity landscape.

As a leader in Zero Trust Segmentation, we are redefining security for a world facing unprecedented cyber threats. You’ll work with a highly scalable SaaS service built using cloud-native technologies while simultaneously shipping the solution on-premises.

Our guiding philosophy in Engineering is to get things right through practicing disciplined engineering, focusing, not cutting corners, and of course having fun while we are at it. We believe in enabling ownership at all levels of the organization and empowering teams. If you thrive in this culture, come join us! Your Impact, Illumio believes that an environment of unique backgrounds, experiences, viewpoints, and individual contributions creates a culture of belonging, drives our future, and makes us stronger together in support of our customers and their success.

All official job offers from our company are extended directly by our recruitment team and will be sent through an official E-Signature document for your review and signature. Please be aware that we do not ask for any personal information in the process of extending offers of employment, such as financial details or social security numbers. Upon acceptance of any offer, we will request such information as part of the onboarding process prior to or on your first day of employment, and only after completing a background check through an authorized third-party vendor. If you receive any communication asking for personal details outside of these processes, please contact us immediately to verify the authenticity of the request. Your security is important to us, and we are committed to a safe and transparent hiring experience.

For roles in San Francisco and Los Angeles: Pursuant to the San Francisco Fair Chance Ordinance and the Los Angeles Fair Chance Initiative for Hiring, Illumio will consider for employment qualified applicants with arrest and conviction records.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · WWC 2023

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

Videos

See all

Related articles

See all