VP of Engineering (AI)
Hyphen Hyphen LLC
San Francisco, CA, United States
2 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.juju.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source
Tech stack
Artificial Intelligence
Cloud Computing
Computer Clusters
Continuous Integration
Software Debugging
Linux
Distributed Systems
Reliability Engineering
Cloud Platform System
Kubernetes
Storage Technologies
Job description
- Lead the design and evolution of the AI cloud platform architecture - GPU orchestration, compute scheduling, networking, storage, and distributed systems
- Build and scale large GPU clusters supporting customer workloads, including GPU provisioning, scheduling, utilization optimization, and capacity management
- Personally participate in architecture reviews, system design, and key technical initiatives (expect 40%+ of time on technical contribution)
- Act as the technical escalation point for complex infrastructure challenges - debug production issues, review proposals, and drive decisions
- Establish best practices for Kubernetes, observability, CI/CD, security, and operational excellence
- Build SRE and Platform Engineering functions from scratch - define SLOs, SLIs, incident response, and capacity planning
- Recruit and develop world-class Infrastructure, Platform, and SRE teams
- Partner with executive leadership on company strategy and infrastructure investments
- Manage infrastructure budgets, vendor relationships, and capacity planning
Requirements
- 12+ years building and operating large-scale infrastructure systems, with experience leading infrastructure organizations while remaining deeply hands-on technically.
- Previous experience building or operating a cloud platform at scale - ideally GPU-native cloud infrastructure supporting AI training and inference workloads
- Expert-level Kubernetes knowledge and experience designing multi-region cloud infrastructure
- Deep expertise in Linux, networking, distributed systems, and storage architecture
- Proven track record scaling infrastructure in high-growth startup environments - not just maintaining systems at large companies
- Strong understanding of Infrastructure-as-Code, automation frameworks, observability, monitoring, and reliability engineering
- Experience building highly available production systems with clear SLOs and incident response processes
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.juju.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
DC
Daniel Cranney
almost 2 years ago
LM
Luis Minvielle
How to Become an AI Engineer
almost 3 years ago
BR
Benjamin Ruschin
Navigating the AI Shift
about 1 year ago
CS
Christina Schaireiter
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
4 months ago
DC
Daniel Cranney
What is Software Engineering in the Age of AI?
12 months ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
about 2 months ago