VP of Engineering (AI)

Hyphen Hyphen LLC
San Francisco, CA, United States
2 days ago
Apply on www.juju.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Cloud Computing Computer Clusters Continuous Integration Software Debugging Linux Distributed Systems Reliability Engineering Cloud Platform System Kubernetes Storage Technologies

Job description

  • Lead the design and evolution of the AI cloud platform architecture - GPU orchestration, compute scheduling, networking, storage, and distributed systems
  • Build and scale large GPU clusters supporting customer workloads, including GPU provisioning, scheduling, utilization optimization, and capacity management
  • Personally participate in architecture reviews, system design, and key technical initiatives (expect 40%+ of time on technical contribution)
  • Act as the technical escalation point for complex infrastructure challenges - debug production issues, review proposals, and drive decisions
  • Establish best practices for Kubernetes, observability, CI/CD, security, and operational excellence
  • Build SRE and Platform Engineering functions from scratch - define SLOs, SLIs, incident response, and capacity planning
  • Recruit and develop world-class Infrastructure, Platform, and SRE teams
  • Partner with executive leadership on company strategy and infrastructure investments
  • Manage infrastructure budgets, vendor relationships, and capacity planning

Requirements

  • 12+ years building and operating large-scale infrastructure systems, with experience leading infrastructure organizations while remaining deeply hands-on technically.
  • Previous experience building or operating a cloud platform at scale - ideally GPU-native cloud infrastructure supporting AI training and inference workloads
  • Expert-level Kubernetes knowledge and experience designing multi-region cloud infrastructure
  • Deep expertise in Linux, networking, distributed systems, and storage architecture
  • Proven track record scaling infrastructure in high-growth startup environments - not just maintaining systems at large companies
  • Strong understanding of Infrastructure-as-Code, automation frameworks, observability, monitoring, and reliability engineering
  • Experience building highly available production systems with clear SLOs and incident response processes

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:15 min

Deploying local container pods to Kubernetes clusters

Stevan Le Meur Stevan Le Meur · World Congress 2024

2:27 min

Introduction to WebAssembly in a cloud computing context

Edo Edo · World Congress 2024

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all