Head of AI Infrastructure

Lhi Group Ltd
San Francisco, CA, United States
3 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$250,000.0 - $350,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Computer Clusters Linux Distributed Data Store InfiniBand Python (Programming Language) Oracle (Applications) AI Platforms Kubernetes Bare Metal Machine Learning Operations

Job description

This is the company’s first dedicated infrastructure hire, reporting directly to the CTO.

You’ll start hands-on and own the AI platform end to end: compute, memory, network, storage, orchestration and field operations. As the fleet grows, you’ll build and lead the infrastructure team. You’ll also be the senior technical voice for customers and set the 12 to 18 month infrastructure roadmap.

The work moves fast. New NVIDIA drivers ship weekly, new vLLM versions every two weeks and new models monthly, so continuous benchmarking and safe rollouts are at the heart of the job.

What you’ll do

  • Bring up, burn in and run GPU pods in production, and own the acceptance benchmarks every new pod and hardware generation must pass
  • Measure and improve inference performance across compute, memory, network and storage: tokens per second per kW, KV-cache offload, RoCEv2 and NCCL fabric performance, model cold-start
  • Test and roll out frequent driver, vLLM and model updates safely, with clear benchmarks and rollback plans
  • Run power-aware operations: GPU power caps, curtailment and workload drain coordinated with other on-site energy loads
  • Design for graceful failure, so work hands off cleanly between pods and sites
  • Operate a multi-tenant, bare-metal Kubernetes GPU platform against service level objectives, with 24/7 incident response
  • Write the runbooks field technicians follow at unmanned sites
  • Act as technical lead for customers, and set the 12 to 18 month infrastructure roadmap
  • Hire, build and lead the infrastructure team as the fleet scales

Requirements

  • 12+ years in infrastructure, with hands-on ownership of physical production compute at scale (not only consuming public cloud services)
  • Experience at a hyperscaler, GPU cloud or neocloud: you know how large organizations keep large systems running
  • Practical AI inference knowledge, such as vLLM or other model serving frameworks
  • Depth in at least two of: inference serving, InfiniBand or RoCE fabrics, distributed storage, bare-metal Kubernetes
  • Experience running distributed or multi-site infrastructure
  • Strong Linux skills, plus Python or Go: you read the code and write the fix
  • Technical leadership experience, and the ambition to build and manage a team

Nice to have

  • Edge or distributed-site infrastructure, such as CDN points of presence, cloud local zones or telecom edge
  • Experience with 1,000+ GPU clusters
  • Power-aware scheduling, demand response or GPU cluster power management
  • Background at a GPU, chip or infrastructure vendor, such as NVIDIA, AMD, Intel or Oracle

About the company

Our client is a venture-backed energy technology company building a new kind of AI compute network. Instead of one large data center, they deploy liquid-cooled NVIDIA GPU inference pods across a nationwide network of existing sites, running on power that is already permitted and in place.

The pods are designed to be modular and quick to deploy. They avoid new utility permits, water hookups, backup generators and major construction. Each pod shares its site’s electrical service with other energy infrastructure, so the platform balances power intelligently and fails over gracefully between sites.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · World Congress 2025

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:23 min

The AI workload technology stack and its components

Lerna Ekmekcioglu Lerna Ekmekcioglu · Europe 2026 Virtual

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · World Congress 2022

Videos

See all

Related articles

See all