> Markdown version of [/jobs/ext/2120092-senior-site-reliability-engineer-bcm-dgx-cloud](https://www.wearedevelopers.com/jobs/ext/2120092-senior-site-reliability-engineer-bcm-dgx-cloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer, BCM - DGX Cloud - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $180,000.0 - $281,250.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Cloud Computing, Computer Clusters, Data Centers, Linux, Monitoring of Systems, InfiniBand, Python (Programming Language), Reliability Engineering, Software Engineering, High Performance Computing, Sysadmin, Kubernetes, Information Technology, Slurm - **Published:** August 19, 2026 - **Apply:** https://www.jofdav.com/jobs/59311927-senior-site-reliability-engineer-bcm-dgx-cloud ## About the Role * Bachelor's Degree or equivalent experience in Computer Science or related field. * 8+ years of experience in site reliability engineering and/or software development roles. * Fluency in Python * In-depth knowledge of Linux and networking Ways to stand out from the crowd: * Experience with C++, high-performance computing, Kubernetes and/or system administration would be an asset * Previous experience as a system admin running BCM/Bright Cluster Manager/Base Command Manager clusters is a definite plus. * Proficiency with cluster networking including InfiniBand and Spectrum-X ## Description NVIDIANS immerse themselves in a diverse, supportive environment that encourages everyone to do their best work. Join the team and see how you can make a lasting impact on the world. NVIDIA Base Command Manager powers thousands of clusters worldwide, varying from a few to several thousands of nodes, and streamlines cluster provisioning, workload management, and infrastructure monitoring. It provides all the tools you need to deploy and run an AI data center. We take great pride in providing excellent, comprehensive support to our customers! Sr Site Reliability Engineer in this role will significantly impact and contribute to the overall success of both external customers running their clusters with NVIDIA solutions AND internal clusters used for research, operations, and next-generation projects. What you'll be doing: * Contributing to deployments and daily operations of large scale next-generation GPU platforms * Handling incidents in GPU clusters, bridging the gap between cluster operations and development * Designing and implementing small features in the Base Command Manager product to become intimately familiar with the workings of the product * Validating complex cluster configurations including Slurm and Kubernetes orchestrators for performance, scalability and resilience, ensuring they meet real-world customer scenarios. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Hacking MSSQL on Cloud. All of them. How I became sysadmin on Azure, AWS, GCP and Alibaba.](https://www.wearedevelopers.com/videos/100339-hacking-mssql-on-cloud-all-of-them-how-i-became-sysadmin-on-azure-aws-gcp-and-alibaba) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Fake or News: Self-Driving Cars on Subscription, Crypto Attacks Rising and Working While You Sleep - Théodore Lefèvre](https://www.wearedevelopers.com/videos/1793-fake-or-news-self-driving-cars-on-subscription-crypto-attacks-rising-and-working-while-you-sleep-theodore-lefevre) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)