Sr Staff Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
Join a team of senior engineers operating in a large-scale, multi-cloud production environment supporting tens of thousands of enterprise customers worldwide. This is not a typical SRE role - you’ll work at the core of a complex, high-impact system alongside experienced DevOps professionals in a fast-paced, cybersecurity-focused organization., * Own and operate large-scale, global production environments across multiple cloud providers (GCP, AWS, Azure)
- Actively monitor, investigate, and resolve incidents triggered by automated alerting systems (PagerDuty / Incident Response)
- Drive end-to-end troubleshooting across complex, distributed systems with high context switching
- Design, deploy, and improve monitoring and observability systems (e.g., Prometheus, Grafana) - not just react to alerts
- Collaborate closely with internal teams (CX, CS, Engineering) to ensure system reliability and performance
- Work hands-on with modern DevOps and infrastructure tools including Kubernetes, Terraform, CI/CD pipelines, and GitOps workflows
- Develop and maintain automation and tooling (primarily in Python)
- Gain deep understanding of system architecture and interconnected services
- Contribute to a culture of operational excellence in a high-scale, high-availability environment
- Champion asynchronous communication, documentation, and tooling standards to ensure seamless collaboration across time zones and fully distributed teams.
- On call responsibilities:Daytime hours (12:00-20:00 CET/CEST, based on candidate location and team coverage needs)
Occasional weekends and holidays (rotation-based)
Requirements
- 5+ years of experience in SRE roles in production environments at scale
- Strong hands-on experience with Kubernetes and Terraform
- Strong hands-on experience with at least one major cloud platform (GCP or AWS required)
- Experience building and configuring monitoring systems (e.g., Prometheus, Grafana)
- Familiarity with CI/CD and GitOps tools (GitLab CI, GitHub Actions, Jenkins, Flux)
- Proficiency in Python for scripting and automation
- Proven success in a fully remote or distributed team environment, demonstrating strong self-management and time organization.
- Strong troubleshooting and problem-solving skills with a passion for incident handling
- Ability to work in fast-paced environments with high context switching
- Highly responsive, proactive, and ownership-driven
- Strong collaboration and communication skills
- Curious mindset and eagerness to learn
About the company
At Palo Alto Networks®, we’re united by a shared mission-to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.
Who We Are
In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us!
This role is remote, but distance is no barrier to impact. Our hybrid teams collaborate across geographies to solve big problems, stay close to our customers, and grow together. You will be part of a culture that values trust, accountability, and shared success where your work truly matters.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Is Software Engineering Over-Saturated?
Mastering Remote Work: Tips for Developers
Find a Developer Job: 12 Best Job Sites For Developers