> Markdown version of [/jobs/ext/2193795-telecommute-principal-ai-infrastructure-architect](https://www.wearedevelopers.com/jobs/ext/2193795-telecommute-principal-ai-infrastructure-architect). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # TELECOMMUTE Principal AI Infrastructure Architect - **Company:** NVIDIA Ltd. - **Location:** United States (Remote available) - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, DevOps, InfiniBand, Python (Programming Language), Machine Learning, Ansible, Prometheus, YAML, AI Infrastructure, Grafana, Kubernetes Helm Charts, AI Platforms, Kubernetes, Infrastructure Automation Frameworks, Machine Learning Operations, Hardware Infrastructure, Terraform - **Published:** August 23, 2026 - **Apply:** https://www.dice.com/job-detail/360d7498-7232-470d-863e-4204a4699c6d ## About the Role * Hands-on experience running AI/ML workloads on NVIDIA DGX systems. * Strong expertise in Kubernetes administration, architecture, and security. * Deep experience with InfiniBand, UFM, and BlueField DPU administration. * Strong scripting and automation skills with Python, Bash, and YAML. * Experience designing scalable, secure, production-grade AI infrastructure. * CKA, CKAD, and CKS certifications. Preferred Qualifications * Experience with NVIDIA Base Command Manager. * Experience with NVIDIA GPU Operator. * Experience with Kubeflow/Kubeflow Pipelines. * Knowledge of MIG GPU partitioning. * Experience with Terraform and Ansible. * Experience with NVIDIA Triton Inference Server. * Strong collaboration skills with ML researchers, DevOps engineers, and infrastructure teams. Technical Environment GPU Infrastructure: NVIDIA DGX, BasePOD, SuperPOD, NVIDIA AI Enterprise Kubernetes: Kubernetes, GPU Operator, Helm, Custom Controllers, MIG Networking: InfiniBand, UFM, BlueField DPU MLOps: Kubeflow Pipelines, NVIDIA Triton Inference Server Automation: Terraform, Ansible, Python, Bash, YAML Monitoring: NVIDIA DCGM, Prometheus, Grafana ## Description We are seeking a Principal AI Infrastructure Architect to design, build, and operate secure, scalable GPU-accelerated AI platforms. The ideal candidate will have deep expertise in Kubernetes, NVIDIA DGX infrastructure, InfiniBand networking, BlueField DPUs, and MLOps platforms. This is a hands-on architecture role responsible for building production-grade AI infrastructure that supports large-scale ML training and inference workloads., * Integrate NVIDIA Base Command Manager with Kubernetes for GPU workload scheduling and resource optimization. * Design MIG-based GPU partitioning strategies for multi-tenant environments. * Develop and manage Helm charts, custom controllers, and GPU operators. DGX Infrastructure & Capacity Planning * Administer and optimize NVIDIA DGX BasePOD and SuperPOD environments. * Ensure optimal GPU, CPU, storage, and cluster performance. * Manage DGX system lifecycle, updates, and infrastructure operations. * Lead capacity planning for cluster expansion, including power, cooling, and storage requirements., * Automate infrastructure provisioning using Terraform and Ansible. * Develop automation using Python, Bash, and YAML. * Monitor AI infrastructure using NVIDIA DCGM, Prometheus, and Grafana. * Build MLOps workflows using Kubeflow Pipelines and NVIDIA Triton Inference Server. * Troubleshoot complex issues across hardware, networking, Kubernetes, and AI software layers. ## Related Videos - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) - [CI/CD with Github Actions](https://www.wearedevelopers.com/videos/856-ci-cd-with-github-actions) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)