Principal Architect GPU Platform & Orchestration
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
We are building a GPU-as-a-Service and AI factory practice from the ground up, delivering multi tenant GPU platforms for enterprise and industrial customers. This is the senior technical seat on that platform. You will own the orchestration and multi-tenancy architecture that turns a GPU cluster into a consumable service, set the standards our global delivery team builds against, and serve as deputy to the practice lead in customer architecture engagements. This is an architecture role. You will design, review, and defend and lead an offshore engineering pod that executes. What you ll do Architect the GPUaaS control plane on Kubernetes and OpenShift: NVIDIA GPU Operator, Network Operator, device plugin, MIG manager, node feature discovery. Design multi-tenancy end to end MIG partitioning strategy, time-slicing tiers, namespace and RBAC model, network policy, quotas, priority classes, and tenant onboarding. Own GPU scheduling and allocation policy: gang scheduling (Kueue, Volcano), fair-share and preemption, topology-aware placement, and Slurm integration where customers run genuine batch HPC. Define the service catalog instance shapes, self-service request flow, and GPU metering for chargeback or showback from DCGM telemetry. Build and own the reusable platform blueprint: reference architecture, Terraform and Helm modules, GitOps patterns, and runbooks that every engagement starts from. Technically lead an offshore delivery pod set standards, run design reviews, gate deliverables before they reach a customer. Partner with the practice lead on customer discovery, solution design, and technical escalation; lead design sessions independently as the practice scales.
Requirements
Rate: $/hr on 1099 Experience: 7+ Years Platform Engineering; 4+ Years Kubernetes in Production Interview Mode: Virtual Practice: AI Infrastructure / GPU-as-a-Service, 7+ years platform engineering, with 4+ on Kubernetes in production; OpenShift experience valued. Demonstrated GPU workload orchestration GPU Operator, MIG, device plugin, GPU scheduling policy on real multi-node clusters. Real multi-tenancy design experience: isolation, quota, RBAC, network segmentation, and the failure modes each produces. Batch or HPC scheduling background (Slurm, LSF, PBS) or gang scheduling on Kubernetes. Strong IaC and GitOps: Terraform, Helm, Argo CD or Flux, Ansible. Experience leading distributed or offshore engineering teams through written standards rather than direct supervision. Customer-facing credibility you can whiteboard a design for a CTO and defend it under challenge. Nice to have NVIDIA AI Enterprise; Run:ai or equivalent GPU orchestration; consulting or professional-services background; internal developer platform / service catalog experience; CKA or CKS.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
I Gave a Video Editor More Autonomy Than a Trading Bot. On Purpose.
Highest Paying Tech Companies for Developers
Dev Digest 120 - Apple and peers