> Markdown version of [/jobs/ext/1322294-senior-platform-engineer-network-infrastructure-dgx-cloud](https://www.wearedevelopers.com/jobs/ext/1322294-senior-platform-engineer-network-infrastructure-dgx-cloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Platform Engineer, Network Infrastructure - DGX Cloud - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $176,000.0 - $276,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Applications Architecture, Common ISDN Application Programming Interface (CAPI), Cloud Computing, Continuous Integration, Data Centers, Distributed Systems, IP Routing, Python (Programming Language), Network Architecture, Network Service, Open Source Technology, Software Engineering, Computer Networking Systems, Cloud Platform System, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Bare Metal, Open Network Automation Platform - **Published:** July 17, 2026 - **Apply:** https://www.disabledperson.com/jobs/73716995-senior-platform-engineer-network-infrastructure-dgx-cloud ## About the Role * Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent experience. * 8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems. * Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery. * Proficiency in at least one general-purpose programming language, such as Go or Python. * Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery. * Experience deploying and supporting network automation or telemetry services on Kubernetes. * Experience with production on-call, incident response, root-cause analysis, and driving corrective actions to completion. Ways to Stand Out From the Crowd: * Strong knowledge of IP routing, data center fabrics, and cloud networking is a great plus. * Experience designing and operating large, multi-region Kubernetes fleets, including fleet-wide upgrades and recovery. * Hands-on experience with Cluster API (CAPI) and Metal3 for bare-metal provisioning, cluster lifecycle, machine remediation, and upgrades. * Experience building Kubernetes controllers or operators in Go using custom resources and reconciliation patterns. Experience designing or operating network automation and telemetry services on Kubernetes at global scale. * Contributions to Cluster API, Metal3, or other open-source Kubernetes infrastructure projects. ## Description Cloud Foundations Reliability (CFR) is part of NVIDIA's Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform and shared services used to provision, monitor, and operate NVIDIA's global network across data centers, colocation facilities, and cloud environments. The team owns the architecture and lifecycle of this platform, including cluster provisioning and upgrades, GitOps delivery, observability, capacity, and service enablement. We build software and automation to standardize how network platforms and services are deployed, scaled, and managed across environments., We are looking for a hands-on senior engineer to own the lifecycle and automation of the Kubernetes platform supporting GNI network systems. You will also provide production support for network services running on the platform, partnering with their engineering owners when issues or changes cross the platform boundary. You will take complex problems from design through production and remain accountable for the outcome. You will bring deep Kubernetes expertise and help establish consistent engineering practices across the US and Bangalore teams. This is a senior individual contributor role with end-to-end ownership and production responsibility. What You'll Be Doing: * Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and cloud environments. * Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery. * Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi-cluster delivery through GitOps. * Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features. * Diagnose complex Kubernetes platform and hosted-service failures involving control-plane health, cluster networking, storage, scheduling, workload placement, and multi-cluster dependencies. Drive issues from initial signal through verified resolution. * Define production-readiness and observability standards for the platform and hosted network services, including health signals, capacity, alerts, runbooks, and recovery. * Participate in CFR's production on-call rotation, including scheduled after-hours and weekend coverage. Lead incident response and recovery, then drive corrective actions to completion. ## Related Videos - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Bitcoin SV: The Massively Scaled Blockchain to Meet Developer Needs](https://www.wearedevelopers.com/videos/20-bitcoin-sv-the-massively-scaled-blockchain-to-meet-developer-needs) - [Single Server, Global Reach: Running a Worldwide Marketplace on Bare Metal in a Cloud-Dominated World](https://www.wearedevelopers.com/videos/1206-single-server-global-reach-running-a-worldwide-marketplace-on-bare-metal-in-a-cloud-dominated-world) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)