> Markdown version of [/jobs/ext/2867844-senior-staff-site-reliability-operations](https://www.wearedevelopers.com/jobs/ext/2867844-senior-staff-site-reliability-operations). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Staff Site Reliability Operations - **Company:** NVIDIA Corporation - **Location:** Seattle, WA, United States - **Experience:** Expert - **Salary:** $184,000.0 - $264,500.0 - **Contract:** Permanent contract - **Skills:** Microsoft Windows, Active Directory, Artificial Intelligence, Apple Mac Systems, Build Automation, Bash Shell, Cyber Security, Information Systems, Databases, Data Centers, Dynamic Host Configuration Protocol, Linux, Domain Name System (DNS), IT Management, Python (Programming Language), Kerberos (Protocol), Lightweight Directory Access Protocols (LDAP), Linux Servers, Simple Mail Transfer Protocols, System Center Configuration Manager, Networking Basics, Windows PowerShell, Reliability Engineering, Virtual Local Area Networks, Virtualization Technology, Software Vulnerability Management, Firewalls (Computer Science), Microsoft InTune, Information Technology, Deployment Automation, Patch Management, Casper Suite, Data Management, 3-tier Architectures, Servicenow - **Published:** September 12, 2026 - **Apply:** https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite/job/US-WA-Seattle/Senior-Staff-Site-Reliability-Operations_JR2025124 ## About the Role * 12+ years in enterprise support engineering, infrastructure, or end user services, including 5+ years in a senior, lead, or escalation-tier role in a multi-site environment. * Deep hands-on solving across Active Directory and hybrid Entra ID, Exchange hybrid, Windows and Linux server, virtualization, enterprise storage, and datacenter hardware. * Enterprise endpoint management (Intune, Autopilot, MECM/SCCM, Jamf, or equivalent), Windows 11, and the Microsoft 365 ecosystem, plus endpoint security and vulnerability remediation. * Database operations support and networking fundamentals - DNS, DHCP, VLAN, wireless, firewall policy, and switch-level troubleshooting. * Scripting and automation in Python, PowerShell, or Bash applied to real support problems, and ServiceNow or similar ITSM. * Demonstrated technical leadership without formal authority, excellent executive-level communication during incidents, and the rigor to pursue root cause over symptom clearing. * Willingness to work on-site and hands-on (including lifting and moving equipment), join an on-call rotation, and support after-hours maintenance windows and cutovers. * Bachelor's degree in Computer Science, Information Systems, or related field, or equivalent experience., * Experience supporting engineering, lab, R&D, or manufacturing environments with specialized equipment and non-standard availability requirements. * Local technical lead through a site buildout, relocation, or major migration; or experience influencing global standards and tooling roadmaps for site and regional needs. * Executive support programs, AV and hybrid conference room technologies, or build automation with measurable efficiency and experience benefits. ## Description NVIDIANs are inspired to excel and make a profound global impact. We are seeking a Site Reliability Operations Technical Lead to serve as the senior technical individual contributor for reliability and support at our Seattle, WA site. This role owns technical service delivery locally, acts as the support point of last resort before regional and global platform teams, and provides technical leadership to site support engineers. Be responsible for the hardest issues across Active Directory, Exchange, database platforms, and compute infrastructure, lead the site through major incidents, and drive out the recurring problems that consume the team's capacity. You will also lead site-level projects, represent local requirements in global initiatives, and set the technical standard the site support team works to. The successful candidate is equally comfortable running a root cause analysis, supporting an executive before an all-hands, and briefing IT leadership on site risk. What you'll be doing: * Own day-to-day site operations - incidents, requests, critical issues, and support coverage - with accountability for queue health, SLA attainment, backlog, and service quality, plus site asset and inventory management across lifecycle, refresh, procurement, and compliance. * Serve as Tier 3 escalation owner for the site and AMER across identity (AD, hybrid Entra ID, GPO, Kerberos/LDAP, SSO, MFA), messaging (Exchange hybrid mail flow, mailbox, SMTP relay), compute (Windows, Linux, macOS, virtualization, storage, and hands-on datacenter and lab hardware), and endpoint (M365, Teams, Intune, Autopilot, imaging through migrations) driving root cause and permanent fixes rather than repeat break-fix. * Own endpoint compliance, vulnerability remediation, patch management, and hardening; audit readiness and evidence; and partnership with InfoSec on incident response and privileged access. * Drive critical issues into global platform teams and vendors with reproduction cases and diagnostic evidence through to a committed fix. * Act as technical lead for site SRO engineers setting standards, reviewing work, directing blocking issues, building diagnostic rigor through mentorship, and owning the site knowledge base and runbook library. * Serve as the primary technical contact for site IT, partnering with employees, site and executive leadership, Facilities, Security, HR, and Procurement on incidents, planned changes, onboarding and moves, and office and lab expansions. * Build automation in PowerShell, Python, or Bash for diagnostics, remediation, health checks, and reporting; analyze ticket and reliability trends to eliminate top recurring drivers; and champion AI-driven and agentic solutions that advance SRO strategy. * Represent site and AMER priorities in regional and global IT initiatives, standards, and architecture forums, and lead operational decisions in the manager's absence. ## Related Videos - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Robots are coming into the wild! Full-Stack Robotics Engineers, be ready!](https://www.wearedevelopers.com/videos/479-robots-are-coming-into-the-wild-full-stack-robotics-engineers-be-ready) - [AI in Production: applied AI & enterprise use cases](https://www.wearedevelopers.com/videos/100130-ai-in-production-applied-ai-enterprise-use-cases) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)