CPU/Storage/PoP-WAN Program Manager

OpenAI Inc.
San Francisco, CA, United States
15 days ago
Apply on diversityjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$226,000.0 - $285,000.0
Working hours
Regular working hours

Tech stack

Board Bringup Microsoft Azure Computer Clusters Data Centers Data Infrastructure Distributed Data Store Network Architecture Systems Architecture AI Infrastructure Network Routers Cloud Platform System Storage Technologies
+1 more
Network Server

Job description

We are seeking a highly technical Program Manager to lead execution across CPU, Storage, PoP, and WAN infrastructure programs that directly unlock OpenAI’s next generation compute capacity.

In this role, you will own complex cross-functional programs spanning compute cluster activation, storage deployment, PoP bring-up, and backbone expansion. You will coordinate hardware readiness, site readiness, network pathing, storage availability, vendor execution, and engineering dependencies required to turn contracted infrastructure into live training and inference capacity.

This role requires strong technical fluency across hardware systems, network infrastructure, storage architecture, and deployment execution. You should be comfortable operating from rack-level implementation details through executive-level capacity planning discussions.

This role is based in San Francisco, CA, with travel as needed., * Lead end-to-end execution of CPU / GPU cluster activation programs across OpenAI’s global infrastructure footprint

  • Drive readiness to convert contracted compute capacity into schedulable production clusters
  • Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives
  • Build integrated schedules spanning procurement, logistics, installation, storage readiness, network turn-up, testing, and production handoff
  • Coordinate BOM readiness, server delivery, racks, optics, cabling, storage hardware, and vendor milestones
  • Partner with engineering teams to align compute, storage, and networking dependencies before cluster activation
  • Manage deployment of storage systems supporting training and inference workloads, including readiness, validation, performance checks, and scaling plans
  • Coordinate backbone capacity expansion, cross-connects, inter-region pathing, and cloud interconnect readiness with Azure and third-party providers
  • Lead physical deployment execution including rack-and-stack, hardware bring-up, L1 validation, and site acceptance criteria
  • Build repeatable deployment playbooks, dashboards, governance cadences, and operating mechanisms for scale
  • Identify risks early across supply chain, site readiness, technical constraints, and vendor execution, then drive mitigation plans
  • Communicate milestones, escalations, and capacity forecasts to senior leadership

Requirements

  • 8+ years of experience in technical program management, infrastructure deployment, network deployment, or data center operations
  • Strong experience delivering programs involving compute, storage, networking, or large-scale infrastructure systems
  • Working knowledge of servers, clusters, storage arrays, routers, switches, optics, and structured cabling
  • Experience owning cross-functional programs across engineering, operations, supply chain, and external vendors
  • Strong understanding of deployment lifecycles from planning and procurement through production handoff
  • Ability to reason across physical infrastructure execution and logical systems architecture dependencies
  • Proven ability to build integrated schedules and drive accountability across multiple stakeholders
  • Strong executive communication skills with experience managing critical escalations and leadership updates
  • Comfortable operating in fast-moving environments with aggressive timelines and evolving priorities
  • Highly analytical with strong problem-solving and execution instincts

Preferred Skills

  • Experience at a hyperscaler, cloud provider, AI infrastructure company, or global network operator
  • Experience deploying GPU clusters, HPC systems, or large training environments
  • Familiarity with distributed storage systems and high-performance data infrastructure
  • Experience with PoP deployments, WAN backbone expansion, or global network buildouts
  • Experience working across first-party, colo, and cloud environments
  • Experience building repeatable infrastructure deployment systems in high-growth environments

About the company

The CPU / Storage / PoP / WAN team is responsible for the end-to-end infrastructure layers required to bring compute online: server and cluster activation, storage platforms, Points of Presence (PoPs), backbone connectivity, and global network expansion. We operate across first-party facilities, colocation environments, and strategic cloud partners to ensure OpenAI can scale reliably and quickly., OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity., At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · World Congress 2022

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:27 min

Binding executable logic into network servers

Saoussen Chaabnia Saoussen Chaabnia · Europe 2026 Virtual

4:03 min

Managing massive power consumption scaling in AI data centers

Stephan Gillich Stephan Gillich +3 · World Congress 2024

1:54 min

Speaker background and open source Kubernetes edge computing projects

Gaurav Gahlot Gaurav Gahlot · World Congress 2026 Europe

Videos

See all

Related articles

See all