> Markdown version of [/jobs/ext/516847-senior-systems-software-engineer-kubernetes-node-lifecycle-dgx-cloud](https://www.wearedevelopers.com/jobs/ext/516847-senior-systems-software-engineer-kubernetes-node-lifecycle-dgx-cloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Systems Software Engineer, Kubernetes Node Lifecycle - DGX Cloud - **Company:** NVIDIA Ltd. - **Location:** Seattle, WA, United States - **Experience:** Expert - **Salary:** $184,000.0 - $356,500.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Common ISDN Application Programming Interface (CAPI), Cloud Computing, Computer Engineering, Software Debugging, Linux, Python (Programming Language), Node.Js, Software Deployment, Software Engineering, Google Cloud, Generative AI, Kubernetes, Information Technology, CIS Benchmarks, Oracle Cloud Infrastructure, Vulnerability Analysis, Golang - **Published:** June 13, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=3036a3d91be2a763 ## About the Role Do you have experience in Vulnerability scanning?, We are looking for a Senior Systems Software Engineer with strong experience in Kubernetes node engineering, OS image packaging, and cloud infrastructure. The ideal candidate will possess deep hyperscaler-level knowledge across the entire node lifecycle. This covers CAPI providers, bring-your-own-node onboarding, OS image build pipelines, packaging, and nodepool management. They must have the technical depth needed to maintain cluster reliability at frontier AI scale. In this vital role, you will manage the node layer within NVIDIA Kubernetes Engine (NKE). Your work will ensure it scales to fulfill DGX Cloud's two main goals: supporting internal researchers and enabling NCPs. Are you prepared to innovate?, * 8 years of experience with a background in systems software, cloud infrastructure, or Kubernetes node engineering. * Bachelor's or Master's degree in Engineering (Electrical, Computer Engineering, Computer Science) or equivalent experience. * Deep expertise in Cluster API (CAPI), including provider development and full machine lifecycle from provisioning to deletion. * Extensive experience with OS image build pipelines, node image packaging, and delivery systems for Kubernetes nodes (for example image-builder, containerd, cloud-init, packer). * Practical experience with bring-your-own-node models and integrating diverse hardware into live Kubernetes environments, including large-scale nodepool lifecycle management and upgrades. * Strong understanding of kubelet configuration, node bootstrap, and the Kubernetes node registration lifecycle. * Experience with node image security, including vulnerability scanning, patch automation, and compliance gating as part of image build pipelines. * Proficiency in Golang and/or Python, and hands-on experience with at least one major public cloud provider (GCP, AWS, Azure, OCI or equivalent). Ways to stand out from the crowd: * Direct experience building or maintaining node image pipelines for a hyperscaler Kubernetes distribution (GKE, EKS, AKS, OKE, or equivalent). * Experience with supply chain security and hardening for node images, including image signing, provenance attestation, SBOM generation, CIS benchmark consistency, and automated CVE remediation. * Experience with automated node provisioning and optimal sizing at scale (for example Karpenter, GKE NAP or similar) and how these interact with GPU workload scheduling. * Strong operational experience working with immutable OS image distributions (such as Flatcar, Bottlerocket, Azure Linux) and debugging node-layer failures in large Kubernetes clusters. * Proven background of upstream contributions to Cluster API, Kubernetes or related CNCF projects, combined with excellent communication and interpersonal abilities. ## Description * Direct the building and refinement of CAPI providers for NVIDIA Kubernetes Engine, maintaining steady, consistent, and scalable node provisioning across DGX Cloud and NCP environments. * Develop and maintain bring-your-own-node workflows that allow customers to integrate different NVIDIA hardware into NKE clusters while ensuring high operational consistency. * Coordinate OS image generation, packaging, deployment, and update processes for NKE nodes. Ensure images are fine-tuned for NVIDIA GPU workloads and satisfy enterprise- and cloud-grade security and compliance criteria. * Develop and sustain node image hardening pipelines, incorporating CIS benchmarks, automated CVE remediation, and promotion gates connected to security posture. * Develop and maintain automated test suites for node images. These tests verify accuracy across Kubernetes versions and NVIDIA hardware configurations. This process occurs prior to production deployment and facilitates continuous validation through modern CI/CD pipelines. * Handle nodepool lifecycle at scale, including provisioning, upgrades, drain and cordon workflows, and seamless node replacement across very large clusters with diverse NVIDIA hardware. * Examine, resolve, and determine underlying causes of node-layer faults in production NKE clusters, such as those involving image configuration, driver packaging, kubelet operation, and hardware activation, and review and optimize the node layer in real-world high-scale scenarios. * Partner with upstream communities including Cluster API, Kubernetes, and CNCF projects to establish node provisioning and lifecycle standards in accordance with NKE requirements. Communicate your progress and findings at internal and external gatherings such as KubeCon and GTC. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)