TELECOMMUTE Principal AI Infrastructure Architect

NVIDIA Ltd.
United States
11 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Bash Shell DevOps InfiniBand Python (Programming Language) Machine Learning Ansible Prometheus YAML AI Infrastructure Grafana Kubernetes Helm Charts
+6 more
AI Platforms Kubernetes Infrastructure Automation Frameworks Machine Learning Operations Hardware Infrastructure Terraform

Job description

We are seeking a Principal AI Infrastructure Architect to design, build, and operate secure, scalable GPU-accelerated AI platforms. The ideal candidate will have deep expertise in Kubernetes, NVIDIA DGX infrastructure, InfiniBand networking, BlueField DPUs, and MLOps platforms.

This is a hands-on architecture role responsible for building production-grade AI infrastructure that supports large-scale ML training and inference workloads., * Integrate NVIDIA Base Command Manager with Kubernetes for GPU workload scheduling and resource optimization.

  • Design MIG-based GPU partitioning strategies for multi-tenant environments.
  • Develop and manage Helm charts, custom controllers, and GPU operators.

DGX Infrastructure & Capacity Planning

  • Administer and optimize NVIDIA DGX BasePOD and SuperPOD environments.
  • Ensure optimal GPU, CPU, storage, and cluster performance.
  • Manage DGX system lifecycle, updates, and infrastructure operations.
  • Lead capacity planning for cluster expansion, including power, cooling, and storage requirements., * Automate infrastructure provisioning using Terraform and Ansible.
  • Develop automation using Python, Bash, and YAML.
  • Monitor AI infrastructure using NVIDIA DCGM, Prometheus, and Grafana.
  • Build MLOps workflows using Kubeflow Pipelines and NVIDIA Triton Inference Server.
  • Troubleshoot complex issues across hardware, networking, Kubernetes, and AI software layers.

Requirements

  • Hands-on experience running AI/ML workloads on NVIDIA DGX systems.
  • Strong expertise in Kubernetes administration, architecture, and security.
  • Deep experience with InfiniBand, UFM, and BlueField DPU administration.
  • Strong scripting and automation skills with Python, Bash, and YAML.
  • Experience designing scalable, secure, production-grade AI infrastructure.
  • CKA, CKAD, and CKS certifications.

Preferred Qualifications

  • Experience with NVIDIA Base Command Manager.
  • Experience with NVIDIA GPU Operator.
  • Experience with Kubeflow/Kubeflow Pipelines.
  • Knowledge of MIG GPU partitioning.
  • Experience with Terraform and Ansible.
  • Experience with NVIDIA Triton Inference Server.
  • Strong collaboration skills with ML researchers, DevOps engineers, and infrastructure teams.

Technical Environment

GPU Infrastructure: NVIDIA DGX, BasePOD, SuperPOD, NVIDIA AI Enterprise Kubernetes: Kubernetes, GPU Operator, Helm, Custom Controllers, MIG Networking: InfiniBand, UFM, BlueField DPU MLOps: Kubeflow Pipelines, NVIDIA Triton Inference Server Automation: Terraform, Ansible, Python, Bash, YAML Monitoring: NVIDIA DCGM, Prometheus, Grafana

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · World Congress 2025

1:35 min

Centralizing configuration logic with native YAML block references

Matthieu Vincent Matthieu Vincent · Europe 2026 Virtual

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

2:00 min

Introduction to YAML syntax and basic formatting

Chris Ayers · LIVE

Videos

See all

Related articles

See all