> Markdown version of [/jobs/ext/1290184-senior-manager-site-reliability-engineering](https://www.wearedevelopers.com/jobs/ext/1290184-senior-manager-site-reliability-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Manager, Site Reliability Engineering - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $200,000.0 - $322,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Configuration Management, Computer Engineering, Information Technology Operations, Reliability Engineering, Information Technology, Data Analytics - **Published:** July 16, 2026 - **Apply:** https://www.juju.com/job/00000000gglxm0 ## About the Role + BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, other Engineering or related fields (or equivalent experience). + 5+ years of experience leading and managing global IT operations or service management teams, with growing scope and complexity. + 12+ overall years of experience in Site Reliability Engineering, IT Service Management, with a focus on Incident Management, Problem Management, and Configuration Management + Proven proficiency in Incident, Problem, and CM with a consistent record of delivering measurable gains in reliability and efficiency. + Demonstrated experience applying AI, automation, or advanced analytics to improve operational outcomes. + Solid understanding of observability, monitoring ecosystems, and modern reliability practices (SRE principles, SLOs, error budgets). + Demonstrated ability to move organizations from process-heavy to technology-focused operating models. + Strong leadership capability with experience building and scaling engineering-focused teams (SRE, SWE, or equivalent). + Ability to deliver executive-level communication and insights, translating operational signals into clear, actionable narratives for leadership. + Ability to build and lead a high-performing team of SREs and engineers, encouraging a culture of ownership, innovation, and continuous improvement. Ways to stand out from the crowd: + ITIL knowledge and/or certification + Experience building or scaling AI-powered operational platforms. + Ability to challenge traditional ITSM models and introduce innovative, scalable approaches. + A mentality passionate about automation first, prevention over reaction, and systems over process. ## Description + Manage the full lifecycle of Incident, Problem, and CM as a 24×7 operational function, ensuring high reliability and minimal business disruption. + Transform incident response by bringing to bear AI detection, correlation, and guided remediation, reducing time to detect, respond, and resolve. + Build and scale intelligent incident workflows that integrate monitoring, telemetry, and service context to enable faster and more consistent response. + Evolve Problem Management into a data-driven field, using AI and analytics to identify patterns, eliminate recurring issues, and drive systemic fixes. + Modernize CM by introducing risk-aware, data-driven decisioning, improving change success rates, and reducing blast radius. + Drive the adoption of observability as a foundation, ensuring service-level visibility, signal quality, and actionable insights across the IT ecosystem. + Lead the development of automation and orchestration platforms that reduce manual effort across the outage lifecycle, including detection, triage, communication, and RCA or equivalent experience. + Partner closely with engineering, infrastructure, and business teams to align operations with service reliability goals and SLOs. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Swapping a Data Warehouse at Runtime: Zero-Downtime Migration Without Changing a Single Client](https://www.wearedevelopers.com/videos/100311-swapping-a-data-warehouse-at-runtime-zero-downtime-migration-without-changing-a-single-client) - [Enabling intelligent logistics automation: home-grown Industrial IoT platform at Austrian Post](https://www.wearedevelopers.com/videos/2018-enabling-intelligent-logistics-automation-home-grown-industrial-iot-platform-at-austrian-post) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Agentic AI: building autonomous systems for developers](https://www.wearedevelopers.com/videos/100034-agentic-ai-building-autonomous-systems-for-developers) - [How Data is Shaping our Games](https://www.wearedevelopers.com/videos/176-how-data-is-shaping-our-games) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)