> Markdown version of [/jobs/ext/1338668-site-reliability-engineer-hardware](https://www.wearedevelopers.com/jobs/ext/1338668-site-reliability-engineer-hardware). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - Hardware... - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $184,000.0 - $287,500.0 - **Contract:** Permanent contract - **Skills:** Systems Engineering, DevOps, Distributed Systems, Perl (Programming Language), Fault Tolerance, Python (Programming Language), Reliability Engineering, Prometheus, Ruby, System Availability, Large Language Models, Grafana, Generative AI, Infrastructure Automation Frameworks, Information Technology, Golang - **Published:** July 18, 2026 - **Apply:** https://www.juju.com/job/00000000ghgm4q ## About the Role + Degree in Computer Science or a related technical field involving coding, or equivalent experience. + 8+ years of experience in SRE, DevOps, or Production Engineering. + Strong understanding of SRE principles, including incident management, error budgets, SLOs, and SLAs. + Experience crafting and deploying systems that are fault-tolerant, performant, and supportable. + Background with infrastructure automation. + Experience running critical services in production. + Experience in one or more of the following: Python, Go, Perl, or Ruby. + Hands-on experience with observability platforms (e.g., Prometheus, Grafana). + Strong communication skills with the ability to convey technical concepts effectively to diverse audiences. + Flexibility and adaptability working in a fast-paced environment with evolving requirements. Ways to stand out from the crowd: + Expertise in establishing incident management and postmortem processes. + Experience driving adoption of common tools and processes across diverse groups. + Experience working with LLM/Generative AI/Agentic solutions to shorten mitigation time, lessen toil, and ensure Service Level Objectives are met. + Hands-on expertise operating and scaling distributed systems with tight SLAs, ensuring high availability and performance. ## Description At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation with consistent reliability and uptime. As an SRE here, you will be part of a welcoming team that values collaboration and creativity, empowering developers to make significant updates while sustaining efficient system function. What you'll be doing: + Develop and support guidelines for incident management, planned maintenance, and blameless postmortems. + Assist teams in responding to high severity incidents, driving root cause analysis, crafting high-quality postmortems, and developing post-incident corrective actions. + Define reliability and supportability metrics, Service Level Objectives, and error budgets. + Develop and drive the adoption of actionable, customer-centric monitoring and alerting. + Apply automation and Generative AI/Agentic solutions to minimize manual and tedious activities and boost customer support. ## Related Videos - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers)