> Markdown version of [/jobs/ext/2198399-ai-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2198399-ai-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Infrastructure Engineer - **Company:** NVIDIA Ltd. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Computer Clusters, Continuous Integration, DevOps, Firmware, InfiniBand, Python (Programming Language), Key Management, Machine Learning, Network Segmentation, Role-Based Access Control, Remote Direct Memory Access, Ansible, Prometheus, YAML, AI Infrastructure, Grafana, Git Flow, Kubernetes, Machine Learning Operations, Terraform, Grpc - **Published:** August 23, 2026 - **Apply:** https://www.dice.com/job-detail/4b401d38-ff00-4213-8755-476ab3bfe421 ## About the Role The ideal candidate will have strong hands-on expertise across NVIDIA DGX systems, Kubernetes, NVIDIA GPU Operator, InfiniBand, BlueField DPUs, NVLink/NVSwitch, and AI infrastructure automation. This is a deeply technical role requiring experience managing high-performance GPU clusters and secure, scalable Kubernetes environments., * Strong hands-on experience with NVIDIA DGX, BasePOD, and SuperPOD environments. * Experience with Kubernetes, NVIDIA GPU Operator, Helm, and Kubeflow. * Proven experience administering InfiniBand and UFM. * Hands-on experience with NVIDIA BlueField DPUs. * Experience with Base Command Manager. * Strong knowledge of NVLink/NVSwitch and NCCL. * Strong scripting and automation skills using Python, YAML, and Bash. * Experience securing GPU-enabled Kubernetes environments. * CKA, CKAD, and CKS certifications. * NVIDIA certifications such as NCA-AIIO, NCP-AII, NCP-AIO, and NCP-AIN. Preferred Qualifications * Experience with Ansible, Terraform, GitOps, and CI/CD. * Experience with NFS, BeeGFS, or Lustre storage. * Knowledge of RoCE, InfiniBand, RDMA, gRPC, and DPU offload. * Experience supporting large-scale AI/ML infrastructure and MLOps environments. Technical Environment AI/GPU: NVIDIA DGX, BasePOD, SuperPOD, NVLink, NVSwitch, NCCL, NVIDIA DCGM Kubernetes: Kubernetes, GPU Operator, Helm, Kubeflow Networking: InfiniBand, UFM, BlueField DPU, RoCE, RDMA Automation: Python, Bash, YAML, Terraform, Ansible DevOps: GitOps, CI/CD Monitoring: Prometheus, Grafana Storage: NFS, BeeGFS, Lustre ## Description We are seeking a highly experienced Senior AI Infrastructure Engineer to deploy, operate, secure, and optimize NVIDIA DGX-based AI infrastructure supporting large-scale AI/ML training and inference workloads., * Manage DGX lifecycle operations including provisioning, monitoring, firmware upgrades, and capacity planning. * Use Base Command Manager for GPU cluster management and workload orchestration. * Perform DGX node health validation, NCCL interconnect testing, and NVLink topology verification., * Manage InfiniBand infrastructure using Unified Fabric Manager (UFM). * Configure and optimize NVLink/NVSwitch connectivity and performance. * Leverage BlueField DPUs for storage, firewalling, security, and telemetry offload. Security & Compliance * Apply CKS-level security practices to Kubernetes and containerized AI environments. * Implement RBAC, workload identity, secrets management, network segmentation, and auditing. * Support zero-trust security initiatives and container/model supply-chain security. Monitoring & Optimization * Monitor GPU, CPU, and I/O performance using NVIDIA DCGM, Prometheus, and Grafana. * Optimize GPU utilization, AI workload performance, and infrastructure efficiency. * Develop operational runbooks, incident response procedures, and SLA dashboards. ## Related Videos - [Exploring the Power of gRPC-Gateway for Writing RESTful Services](https://www.wearedevelopers.com/videos/2072-exploring-the-power-of-grpc-gateway-for-writing-restful-services) - [CI/CD with Github Actions](https://www.wearedevelopers.com/videos/856-ci-cd-with-github-actions) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Boosting OpenSearch Performance: gRPC Search in Action](https://www.wearedevelopers.com/videos/1964-boosting-opensearch-performance-grpc-search-in-action) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)