> Markdown version of [/jobs/ext/2946931-senior-site-reliability-engineer-dgx-cloud](https://www.wearedevelopers.com/jobs/ext/2946931-senior-site-reliability-engineer-dgx-cloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer, DGX Cloud - **Company:** NVIDIA Corporation - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $49,920.0 - $56,160.0 - **Contract:** Temporary contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Microsoft Azure, Cloud Computing, Cloud Computing Security, Nvidia CUDA, Python (Programming Language), Networking Basics, Reliability Engineering, Ansible, Prometheus, Software Engineering, TCP/IP, Data Logging, Google Cloud, Pytorch, Large Language Models, Grafana, Containerization, Kubernetes, Infrastructure Automation Frameworks, Information Technology, TensorRT, Puppet, Terraform, Oracle Cloud Infrastructure, Splunk, Virtual Private Clouds, Elk Stack, Microservices - **Published:** September 16, 2026 - **Apply:** https://www.jofdav.com/jobs/59733639-senior-site-reliability-engineer-dgx-cloud ## About the Role * BS in Computer Science or related technical field, or equivalent experience. * 8+ years of experience operating production services. * Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture. * Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet). * Proficiency in at least one high-level programming language (e.g., Python, Go). * In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards. * Solid grasp of SRE principles, such as SLOs, SLIs, error budgets, and incident management. * Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc. Ways to stand out from the crowd: * Operating GPU-accelerated clusters with KubeVirt in production. * Applying generative-AI techniques to reduce operational toil. * Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions. * Experience operating and resolving problems in production AI inference workloads across the model-to-GPU stack, including vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, and GPU performance analysis. ## Description * Build, implement and support operational and reliability aspects of large-scale Kubernetes clusters with focus on performance at scale, real-time monitoring, logging, and alerting. * Define SLOs/SLIs, monitor error allowances, and streamline reporting. * Support services before they launch through system creation consulting, developing software tools, platforms and frameworks, capacity management, and launch reviews. * Maintain services once they are live by measuring and supervising availability, latency, and overall system health. * Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds. * Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity. * Lead triage and root-cause analysis of high-severity incidents. * Practice balanced incident response and blameless postmortems. * Participate in on-call rotation to support production services. ## Related Videos - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Automate everything via NodeJS and Puppeteer](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)