> Markdown version of [/jobs/ext/2620812-nvidia-ai-infrastructure-kubernetes-platform-engineer-dgx-systems-remote](https://www.wearedevelopers.com/jobs/ext/2620812-nvidia-ai-infrastructure-kubernetes-platform-engineer-dgx-systems-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # NVIDIA AI Infrastructure & Kubernetes Platform Engineer (DGX Systems) - REMOTE - **Company:** Catapult Solutions Group - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $125,000.0 - **Contract:** Permanent contract - **Skills:** Kubernetes Security, Application Programming Interfaces (APIs), Artificial Intelligence, Bash Shell, Computer Clusters, Computer Programming, Continuous Integration, DevOps, Firmware, InfiniBand, Python (Programming Language), Key Management, Network Segmentation, Role-Based Access Control, Remote Direct Memory Access, Ansible, Prometheus, Zero Trust Network Access, YAML, AI Infrastructure, Grafana, Git Flow, Kubernetes, Machine Learning Operations, Terraform, Grpc - **Published:** August 11, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=65e52d6dcac5359c ## About the Role * Certified Kubernetes Administrator (CKA) * Certified Kubernetes Application Developer (CKAD) * Certified Kubernetes Security Specialist (CKS) * NVIDIA Certified Associate: AI Infrastructure & Operations (NCA-AIIO) * NVIDIA Certified Professional: AI Infrastructure (NCP-AII) * NVIDIA Certified Professional: AI Operations (NCP-AIO) * NVIDIA Certified Professional: AI Networking (NCP-AIN) Expertise With: * DGX System, BasePOD, and SuperPOD Administration * BlueField DPU Configuration & Operations * InfiniBand Fabric and UFM Management * Base Command Manager for workload orchestration Technical Skills: * Kubernetes, Helm, GPU Operator, Kubeflow * DevOps tools: Ansible, Terraform, GitOps, CI/CD pipelines * Storage: NFS, BeeGFS, Lustre * Networking: RoCE, InfiniBand, DPU offload, gRPC, RDMA * Programming/scripting: Python, YAML, Bash ## Description We are seeking a highly skilled AI Infrastructure & Kubernetes Platform Engineer with a proven track record in deploying and managing NVIDIA DGX-based AI clusters, orchestrating containerized AI workloads using Kubernetes, and ensuring secure, high-throughput operations across InfiniBand-powered networks. The ideal candidate will hold a combination of Kubernetes certifications (CKA, CKAD, CKS) and NVIDIA certifications (NCA-AIIO, NCP-AIO, NCP-AII, NCP-AIN), coupled with hands-on training in DGX, BlueField, and high-speed network operations.This position plays a key role in supporting AI/ML infrastructure at scale, enabling efficient training and inference for complex models, and integrating NVIDIA's cutting-edge compute, storage, and fabric solutions with modern DevOps practices., * Oversee DGX system lifecycle operations including provisioning, monitoring, firmware upgrades, and capacity planning. * Operate Base Command Manager to manage GPU clusters, schedule workloads, and integrate with MLOps tools. * Perform DGX node health validation, NCCL interconnect testing, and NVLink topology verification following new deployments or hardware changes. Kubernetes Platform Engineering * Architect secure and scalable Kubernetes clusters optimized for GPU-accelerated workloads using NVIDIA GPU Operator. * Leverage expertise from CKA/CKAD/CKS to develop, deploy, and secure AI applications on Kubernetes. * Implement CI/CD pipelines and GitOps methodologies for deploying and managing ML workflows. High-Performance Networking & DPUs * Administer InfiniBand networks and BlueField DPUs using Unified Fabric Manager (UFM). * Enable NVLink/NVSwitch performance across GPU nodes and tune fabric configurations for minimal latency and maximum throughput. * Use BlueField for offloading storage, firewalling, and telemetry, enhancing AI workload security and performance. Security & Compliance * Apply best practices from the CKS certification to secure containerized AI environments. * Configure runtime security, secrets management, network segmentation, and auditing using DPU-enhanced Kubernetes deployments. * Support zero-trust architecture initiatives by enforcing workload identity, RBAC policies, and supply chain integrity across AI container images and model artifacts. Monitoring, Telemetry & Optimization * Monitor GPU, CPU, and I/O performance using NVIDIA DCGM, Prometheus, Grafana, and Base Command APIs. * Tune system performance and model training pipelines for cost-efficiency and throughput. * Build and maintain operational runbooks, incident response playbooks, and SLA reporting dashboards covering GPU utilization, thermal thresholds, and fabric health. ## Related Videos - [Exploring the Power of gRPC-Gateway for Writing RESTful Services](https://www.wearedevelopers.com/videos/2072-exploring-the-power-of-grpc-gateway-for-writing-restful-services) - [CI/CD with Github Actions](https://www.wearedevelopers.com/videos/856-ci-cd-with-github-actions) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Boosting OpenSearch Performance: gRPC Search in Action](https://www.wearedevelopers.com/videos/1964-boosting-opensearch-performance-grpc-search-in-action) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)