> Markdown version of [/jobs/ext/2489947-manager-site-reliability-engineering](https://www.wearedevelopers.com/jobs/ext/2489947-manager-site-reliability-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Manager, Site Reliability Engineering - **Company:** Palo Alto Networks - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, ARM Architecture, Software as a Service, Cloud Computing, Cloud Engineering, DevOps, Distributed Systems, Python (Programming Language), Reliability Engineering, Ansible, Prometheus, Service-Oriented Architecture, Google Cloud, Cloud Platform System, Grafana, Containerization, Git Flow, Kubernetes, Terraform - **Published:** August 11, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=25b949fa97f591ea ## About the Role * 10+ years of experience in Site Reliability Engineering, DevOps, Production Engineering, Cloud Infrastructure, or related areas, including experience leading or managing engineering teams. * Strong technical knowledge of cloud platforms, preferably Google Cloud Platform (GCP). * Strong experience with Kubernetes, containerized environments, and large-scale distributed systems. * Experience with observability technologies such as Prometheus, Thanos, Grafana, OpenTelemetry, and cloud-native monitoring solutions. * Strong understanding of incident management, monitoring, alerting, SLOs, and reliability engineering practices. * Experience with automation and Infrastructure as Code using technologies such as Python, Terraform, Ansible, and GitOps. * Demonstrated ability to lead engineers, set priorities, drive execution, and manage multiple operational and technical initiatives. * Strong communication skills with the ability to collaborate and influence across engineering teams and global organizations., * Experience managing SRE, DevOps, Production Engineering, or infrastructure teams supporting large-scale SaaS platforms. * Experience operating highly available, multi-region cloud environments. * Experience driving automation, self-healing, or AI-assisted operational capabilities. * Strong operational leadership with experience managing high-severity production incidents. * Proven ability to develop engineers, improve team processes, and drive a culture of ownership and continuous improvement. ## Description * Lead, mentor, and develop a team of Site Reliability/Production Engineers, providing technical direction, coaching, and career development. * Own the reliability, availability, and operational health of critical Cortex services and infrastructure. * Drive improvements in monitoring, alerting, incident management, and observability to identify and resolve issues before they impact customers. * Lead major production incidents, ensure effective root cause analysis, and drive corrective and preventive actions. * Partner with Engineering teams to improve service architecture, production readiness, scalability, and operability. * Drive automation and self-healing solutions to reduce operational toil and improve engineering efficiency. * Establish clear priorities, operational processes, and measurable reliability goals for the team. * Collaborate with Production Engineering teams across regions to strengthen follow-the-sun operations and consistent operational practices. * Evaluate new technologies and drive adoption of solutions that improve reliability, scalability, and operational efficiency. ## Related Videos - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Eclipse Che for Infrastructure Automation](https://www.wearedevelopers.com/videos/1611-eclipse-che-for-infrastructure-automation) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)