> Markdown version of [/jobs/ext/2984821-principal-architect-gpu-platform-orchestration](https://www.wearedevelopers.com/jobs/ext/2984821-principal-architect-gpu-platform-orchestration). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Architect GPU Platform & Orchestration - **Company:** VST Consulting, Inc - **Location:** Plano, TX, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computer Networks, Network Control, Network Segmentation, Octopus Deploy, OpenShift, Role-Based Access Control, Ansible, AI Infrastructure, Kubernetes, Slurm, Terraform - **Published:** September 18, 2026 - **Apply:** https://www.dice.com/job-detail/696cf8d9-90a0-4c14-bfa1-e62c388c65bd ## About the Role Rate: $/hr on 1099 Experience: 7+ Years Platform Engineering; 4+ Years Kubernetes in Production Interview Mode: Virtual Practice: AI Infrastructure / GPU-as-a-Service, 7+ years platform engineering, with 4+ on Kubernetes in production; OpenShift experience valued. Demonstrated GPU workload orchestration GPU Operator, MIG, device plugin, GPU scheduling policy on real multi-node clusters. Real multi-tenancy design experience: isolation, quota, RBAC, network segmentation, and the failure modes each produces. Batch or HPC scheduling background (Slurm, LSF, PBS) or gang scheduling on Kubernetes. Strong IaC and GitOps: Terraform, Helm, Argo CD or Flux, Ansible. Experience leading distributed or offshore engineering teams through written standards rather than direct supervision. Customer-facing credibility you can whiteboard a design for a CTO and defend it under challenge. Nice to have NVIDIA AI Enterprise; Run:ai or equivalent GPU orchestration; consulting or professional-services background; internal developer platform / service catalog experience; CKA or CKS. ## Description We are building a GPU-as-a-Service and AI factory practice from the ground up, delivering multi tenant GPU platforms for enterprise and industrial customers. This is the senior technical seat on that platform. You will own the orchestration and multi-tenancy architecture that turns a GPU cluster into a consumable service, set the standards our global delivery team builds against, and serve as deputy to the practice lead in customer architecture engagements. This is an architecture role. You will design, review, and defend and lead an offshore engineering pod that executes. What you ll do Architect the GPUaaS control plane on Kubernetes and OpenShift: NVIDIA GPU Operator, Network Operator, device plugin, MIG manager, node feature discovery. Design multi-tenancy end to end MIG partitioning strategy, time-slicing tiers, namespace and RBAC model, network policy, quotas, priority classes, and tenant onboarding. Own GPU scheduling and allocation policy: gang scheduling (Kueue, Volcano), fair-share and preemption, topology-aware placement, and Slurm integration where customers run genuine batch HPC. Define the service catalog instance shapes, self-service request flow, and GPU metering for chargeback or showback from DCGM telemetry. Build and own the reusable platform blueprint: reference architecture, Terraform and Helm modules, GitOps patterns, and runbooks that every engagement starts from. Technically lead an offshore delivery pod set standards, run design reviews, gate deliverables before they reach a customer. Partner with the practice lead on customer discovery, solution design, and technical escalation; lead design sessions independently as the practice scales. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [This Is Not Your Father's .NET](https://www.wearedevelopers.com/videos/967-this-is-not-your-father-s-net) - [Eclipse Che for Infrastructure Automation](https://www.wearedevelopers.com/videos/1611-eclipse-che-for-infrastructure-automation) - [Embracing the Hybrid Cloud: Unlocking Success with Ansible](https://www.wearedevelopers.com/videos/932-embracing-the-hybrid-cloud-unlocking-success-with-ansible) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [ I Gave a Video Editor More Autonomy Than a Trading Bot. On Purpose.](https://www.wearedevelopers.com/magazine/773-i-gave-a-video-editor-more-autonomy-than-a-trading-bot-on-purpose) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)