> Markdown version of [/jobs/ext/230906-senior-site-reliability-engineer-hpc](https://www.wearedevelopers.com/jobs/ext/230906-senior-site-reliability-engineer-hpc). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - HPC - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $152,000.0 - $241,500.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Continuous Integration, Perl (Programming Language), Python (Programming Language), Machine Learning, Open Source Technology, Performance Tuning, Reliability Engineering, Ruby, Scripting, Google Cloud, Multi-Cloud, Infrastructure as Code (IaC), Containerization, Kubernetes, Information Technology, Data Analytics, Slurm, Oracle Cloud Infrastructure, Golang, Programming Languages - **Published:** May 25, 2026 - **Apply:** https://www.juju.com/job/00000000g2mw5j ## About the Role + B.S. degree in Computer Science or related technical field (or equivalent experience) with 5+ years professional experience building and supporting critical services. + Experience supporting large-scale HPC clusters using Slurm, LSF or Kubernetes clusters, including setup, tuning, and troubleshooting. + Proficiency in modern CI/CD techniques, and Infrastructure as Code (IaC) for managing services. + Strong experience crafting large-scale infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (AIOps/ML-driven signals) that materially reduce manual intervention. + Proficient in monitoring, metrics, container management, and log collection tools. + 5+ years of coding/scripting experience in at least two high-level programming languages such as Python, Go, Perl, or Ruby. + Mentored other engineers and influenced technical direction through design reviews, architecture documents, and strong partnership with product and leadership. + Creative problem solver with excellent debugging skills and strong communication and documentation abilities. Ways to stand out from the crowd: + Published technical write-ups or talks (conference presentations, meetups, engineering blogs) that deep-dive into real-world reliability, observability, or large-scale HPC/SRE problems and their solutions. + Maintainer or co-maintainer responsibilities for an open source component used in production (plugins, operators, exporters, controllers, or SDKs) at large scale. ## Description + Own SRE solutions end-to-end, from design and implementation to operation and continuous improvement, ensuring they integrate cleanly with HPC schedulers, storage, and network fabrics. + Use IaC(Infrastructure-as-Code) and config management to standardize and automate provisioning everywhere. + Deliver solutions in a globally distributed, multi-cloud hybrid environment - On-prem, AWS, GCP, and OCI. + Design for failure with redundancy, failure domains, progressive delivery, and strict change control. + Ensure the highest level of uptime and Quality of Service (QoS) for internal customers through operational excellence. + Conduct capacity management and planning to meet ongoing operational needs. + Detects performance issues and recommends solutions to maintain world-class service quality. + Collaborate with various teams in a fast-paced environment to ensure seamless project completion. + Participate in on-call, incident reviews, assist in root cause identification, and produce high-quality RCA reports. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Fireside Chat with Werner Vogels, VP & CTO, Amazon.com & Daniel Gebler, CTO at Picnic](https://www.wearedevelopers.com/videos/1405-fireside-chat-with-werner-vogels-vp-cto-amazon-com-daniel-gebler-cto-at-picnic) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)