> Markdown version of [/jobs/ext/3071760-head-of-ai-infrastructure](https://www.wearedevelopers.com/jobs/ext/3071760-head-of-ai-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Head of AI Infrastructure - **Company:** Lhi Group Ltd - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $250,000.0 - $350,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computer Clusters, Linux, Distributed Data Store, InfiniBand, Python (Programming Language), Oracle (Applications), AI Platforms, Kubernetes, Bare Metal, Machine Learning Operations - **Published:** September 25, 2026 - **Apply:** https://www.dice.com/job-detail/e7a796b8-6f83-4365-b6c8-e139d50b6d9d ## About the Role * 12+ years in infrastructure, with hands-on ownership of physical production compute at scale (not only consuming public cloud services) * Experience at a hyperscaler, GPU cloud or neocloud: you know how large organizations keep large systems running * Practical AI inference knowledge, such as vLLM or other model serving frameworks * Depth in at least two of: inference serving, InfiniBand or RoCE fabrics, distributed storage, bare-metal Kubernetes * Experience running distributed or multi-site infrastructure * Strong Linux skills, plus Python or Go: you read the code and write the fix * Technical leadership experience, and the ambition to build and manage a team Nice to have * Edge or distributed-site infrastructure, such as CDN points of presence, cloud local zones or telecom edge * Experience with 1,000+ GPU clusters * Power-aware scheduling, demand response or GPU cluster power management * Background at a GPU, chip or infrastructure vendor, such as NVIDIA, AMD, Intel or Oracle ## Description This is the company's first dedicated infrastructure hire, reporting directly to the CTO. You'll start hands-on and own the AI platform end to end: compute, memory, network, storage, orchestration and field operations. As the fleet grows, you'll build and lead the infrastructure team. You'll also be the senior technical voice for customers and set the 12 to 18 month infrastructure roadmap. The work moves fast. New NVIDIA drivers ship weekly, new vLLM versions every two weeks and new models monthly, so continuous benchmarking and safe rollouts are at the heart of the job. What you'll do * Bring up, burn in and run GPU pods in production, and own the acceptance benchmarks every new pod and hardware generation must pass * Measure and improve inference performance across compute, memory, network and storage: tokens per second per kW, KV-cache offload, RoCEv2 and NCCL fabric performance, model cold-start * Test and roll out frequent driver, vLLM and model updates safely, with clear benchmarks and rollback plans * Run power-aware operations: GPU power caps, curtailment and workload drain coordinated with other on-site energy loads * Design for graceful failure, so work hands off cleanly between pods and sites * Operate a multi-tenant, bare-metal Kubernetes GPU platform against service level objectives, with 24/7 incident response * Write the runbooks field technicians follow at unmanned sites * Act as technical lead for customers, and set the 12 to 18 month infrastructure roadmap * Hire, build and lead the infrastructure team as the fleet scales ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)