TELECOMMUTE Principal AI Infrastructure Architect
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+6 more
Job description
We are seeking a Principal AI Infrastructure Architect to design, build, and operate secure, scalable GPU-accelerated AI platforms. The ideal candidate will have deep expertise in Kubernetes, NVIDIA DGX infrastructure, InfiniBand networking, BlueField DPUs, and MLOps platforms.
This is a hands-on architecture role responsible for building production-grade AI infrastructure that supports large-scale ML training and inference workloads., * Integrate NVIDIA Base Command Manager with Kubernetes for GPU workload scheduling and resource optimization.
- Design MIG-based GPU partitioning strategies for multi-tenant environments.
- Develop and manage Helm charts, custom controllers, and GPU operators.
DGX Infrastructure & Capacity Planning
- Administer and optimize NVIDIA DGX BasePOD and SuperPOD environments.
- Ensure optimal GPU, CPU, storage, and cluster performance.
- Manage DGX system lifecycle, updates, and infrastructure operations.
- Lead capacity planning for cluster expansion, including power, cooling, and storage requirements., * Automate infrastructure provisioning using Terraform and Ansible.
- Develop automation using Python, Bash, and YAML.
- Monitor AI infrastructure using NVIDIA DCGM, Prometheus, and Grafana.
- Build MLOps workflows using Kubeflow Pipelines and NVIDIA Triton Inference Server.
- Troubleshoot complex issues across hardware, networking, Kubernetes, and AI software layers.
Requirements
- Hands-on experience running AI/ML workloads on NVIDIA DGX systems.
- Strong expertise in Kubernetes administration, architecture, and security.
- Deep experience with InfiniBand, UFM, and BlueField DPU administration.
- Strong scripting and automation skills with Python, Bash, and YAML.
- Experience designing scalable, secure, production-grade AI infrastructure.
- CKA, CKAD, and CKS certifications.
Preferred Qualifications
- Experience with NVIDIA Base Command Manager.
- Experience with NVIDIA GPU Operator.
- Experience with Kubeflow/Kubeflow Pipelines.
- Knowledge of MIG GPU partitioning.
- Experience with Terraform and Ansible.
- Experience with NVIDIA Triton Inference Server.
- Strong collaboration skills with ML researchers, DevOps engineers, and infrastructure teams.
Technical Environment
GPU Infrastructure: NVIDIA DGX, BasePOD, SuperPOD, NVIDIA AI Enterprise Kubernetes: Kubernetes, GPU Operator, Helm, Custom Controllers, MIG Networking: InfiniBand, UFM, BlueField DPU MLOps: Kubeflow Pipelines, NVIDIA Triton Inference Server Automation: Terraform, Ansible, Python, Bash, YAML Monitoring: NVIDIA DCGM, Prometheus, Grafana
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
Stephan Gillich - Bringing AI Everywhere
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud