NVIDIA AI Infrastructure & Kubernetes Platform Engineer (DGX Systems) - REMOTE

Catapult Solutions Group
United States
24 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$125,000.0
Working hours
Regular working hours
Job source

Tech stack

Kubernetes Security Application Programming Interfaces (APIs) Artificial Intelligence Bash Shell Computer Clusters Computer Programming Continuous Integration DevOps Firmware InfiniBand Python (Programming Language) Key Management
+14 more
Network Segmentation Role-Based Access Control Remote Direct Memory Access Ansible Prometheus Zero Trust Network Access YAML AI Infrastructure Grafana Git Flow Kubernetes Machine Learning Operations Terraform Grpc

Job description

We are seeking a highly skilled AI Infrastructure & Kubernetes Platform Engineer with a proven track record in deploying and managing NVIDIA DGX-based AI clusters, orchestrating containerized AI workloads using Kubernetes, and ensuring secure, high-throughput operations across InfiniBand-powered networks. The ideal candidate will hold a combination of Kubernetes certifications (CKA, CKAD, CKS) and NVIDIA certifications (NCA-AIIO, NCP-AIO, NCP-AII, NCP-AIN), coupled with hands-on training in DGX, BlueField, and high-speed network operations.This position plays a key role in supporting AI/ML infrastructure at scale, enabling efficient training and inference for complex models, and integrating NVIDIA’s cutting-edge compute, storage, and fabric solutions with modern DevOps practices., * Oversee DGX system lifecycle operations including provisioning, monitoring, firmware upgrades, and capacity planning.

  • Operate Base Command Manager to manage GPU clusters, schedule workloads, and integrate with MLOps tools.
  • Perform DGX node health validation, NCCL interconnect testing, and NVLink topology verification following new deployments or hardware changes.

Kubernetes Platform Engineering

  • Architect secure and scalable Kubernetes clusters optimized for GPU-accelerated workloads using NVIDIA GPU Operator.
  • Leverage expertise from CKA/CKAD/CKS to develop, deploy, and secure AI applications on Kubernetes.
  • Implement CI/CD pipelines and GitOps methodologies for deploying and managing ML workflows.

High-Performance Networking & DPUs

  • Administer InfiniBand networks and BlueField DPUs using Unified Fabric Manager (UFM).
  • Enable NVLink/NVSwitch performance across GPU nodes and tune fabric configurations for minimal latency and maximum throughput.
  • Use BlueField for offloading storage, firewalling, and telemetry, enhancing AI workload security and performance.

Security & Compliance

  • Apply best practices from the CKS certification to secure containerized AI environments.
  • Configure runtime security, secrets management, network segmentation, and auditing using DPU-enhanced Kubernetes deployments.
  • Support zero-trust architecture initiatives by enforcing workload identity, RBAC policies, and supply chain integrity across AI container images and model artifacts.

Monitoring, Telemetry & Optimization

  • Monitor GPU, CPU, and I/O performance using NVIDIA DCGM, Prometheus, Grafana, and Base Command APIs.
  • Tune system performance and model training pipelines for cost-efficiency and throughput.
  • Build and maintain operational runbooks, incident response playbooks, and SLA reporting dashboards covering GPU utilization, thermal thresholds, and fabric health.

Requirements

  • Certified Kubernetes Administrator (CKA)
  • Certified Kubernetes Application Developer (CKAD)
  • Certified Kubernetes Security Specialist (CKS)
  • NVIDIA Certified Associate: AI Infrastructure & Operations (NCA-AIIO)
  • NVIDIA Certified Professional: AI Infrastructure (NCP-AII)
  • NVIDIA Certified Professional: AI Operations (NCP-AIO)
  • NVIDIA Certified Professional: AI Networking (NCP-AIN)

Expertise With:

  • DGX System, BasePOD, and SuperPOD Administration
  • BlueField DPU Configuration & Operations
  • InfiniBand Fabric and UFM Management
  • Base Command Manager for workload orchestration

Technical Skills:

  • Kubernetes, Helm, GPU Operator, Kubeflow
  • DevOps tools: Ansible, Terraform, GitOps, CI/CD pipelines
  • Storage: NFS, BeeGFS, Lustre
  • Networking: RoCE, InfiniBand, DPU offload, gRPC, RDMA
  • Programming/scripting: Python, YAML, Bash

Benefits & conditions

$125,000 an hour - Full-time, Contract

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:56 min

Establishing internal service communication with gRPC

Florian Bader Florian Bader · World Congress 2026 Europe

1:35 min

Centralizing configuration logic with native YAML block references

Matthieu Vincent Matthieu Vincent · Europe 2026 Virtual

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

33 sec

Supporting NVIDIA GB200 GPUs on Kubernetes

Kevin Klues Kevin Klues · World Congress 2025

1:11 min

Evaluating architectural trade-offs between REST and gRPC

Sakshi Nasha Sakshi Nasha · Europe 2026 Virtual

1:37 min

Unlocking direct GPU access within managed Kubernetes platforms

Kevin Klues Kevin Klues

Videos

See all

Related articles

See all