> Markdown version of [/jobs/ext/3226852-principal-technical-program-manager-ai-cloud-infrastructure](https://www.wearedevelopers.com/jobs/ext/3226852-principal-technical-program-manager-ai-cloud-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Technical Program Manager, AI Cloud Infrastructure - **Company:** IREN LLC - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $200,000.0 - $260,000.0 - **Contract:** Permanent contract - **Skills:** Agile Methodology, Cloud Computing, Computer Clusters, Data Centers, Distributed Systems, Ethernet, Firmware, Issue Tracking Systems, InfiniBand, Networking Hardware, Scrum Methodology, Software Deployment, Software Engineering, Software Requirements Analysis, AI Infrastructure, Cloud Platform System, High Performance Computing, AI Platforms, Kubernetes, Infrastructure Automation Frameworks, Bare Metal, Enterprise Integration, Hardware Infrastructure, Server Operating Systems & Platforms - **Published:** September 13, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=e7b8b8a33cb566e2 ## About the Role * Bachelor's degree in a technical discipline or equivalent experience. * 10+ years of experience in technical program management, cloud infrastructure, data centers, hardware infrastructure, or software engineering programs. * Proven experience owning end-to-end delivery of complex infrastructure programs from planning and requirements through deployment, production readiness, customer launch, and operational handoff. * Strong understanding of bare metal GPU infrastructure, GPU server platforms, high-performance networking, storage, cloud platforms, distributed systems, and data center dependencies such as power, cooling, rack layout, and network readiness. * Experience leading GPU cluster bring-up, capacity enablement, hardware deployment, network integration, validation, and production launch across internal teams, vendors, and external partners. * Ability to manage critical paths, risks, dependencies, schedules, decision forums, executive communications, and customer-facing delivery commitments in high-ambiguity environments. Preferred Qualifications * Experience in hyperscale cloud, AI infrastructure, GPU cloud, bare metal cloud, or HPC environments. * Experience supporting NVIDIA-based AI platforms, high-density GPU clusters, InfiniBand or high-performance Ethernet fabrics, and cloud capacity bring-up with external partners. * Familiarity with Kubernetes, cloud-native platforms, provisioning, orchestration, automation, observability, fleet management, and infrastructure-as-code practices. * Experience partnering with solution architects, customer engineering, sales engineering, managed service providers, system integrators, and data center delivery teams. * PMP, PgMP, Agile, Scrum, or similar certifications. ## Description As a Principal Technical Program Manager, Bare Metal GPU & AI Cloud Delivery, you will own the end-to-end delivery of strategic AI Cloud programs, from infrastructure planning, capacity readiness, and bare metal GPU cluster bring-up through production launch, customer onboarding, and operational handoff. You will bring structure, execution discipline, and technical depth to programs spanning GPU systems, high-performance networking, storage, data center readiness, platform software, security, and customer-facing delivery. You will partner closely with infrastructure engineering, AI cloud software engineering, data center operations, networking, security, supply chain, finance, solution architecture, product, and customer-facing teams to deliver reliable, scalable, production-ready GPU cloud infrastructure for AI training and inference workloads at scale., This role requires a senior owner who can operate across hardware, data center, and cloud software domains; drive clarity in ambiguous environments; anticipate risks across vendors, facilities, and engineering teams; and ensure that every stage of AI Cloud delivery is planned, tracked, validated, and communicated with executive-level precision., * Own and drive end-to-end delivery processes for Bare Metal GPU and AI Cloud programs, from requirements definition, infrastructure design, capacity planning, procurement readiness, and deployment planning through commissioning, software deployment, customer handover, and operational steady state. * Define, standardize, and scale repeatable delivery processes across multiple data centers, ensuring each site follows clear playbooks, milestones, dependency tracking, readiness criteria, risk management, escalation paths, and handoff procedures. * Lead cross-functional execution across infrastructure design, data center operations, supply chain, logistics, infrastructure installation, cabling, networking, commissioning, cloud software deployment, security, customer engineering, and support teams. * Partner with infrastructure design teams to translate AI Cloud capacity, GPU cluster architecture, power, cooling, rack layout, network fabric, storage, and operational requirements into executable multi-site delivery plans. * Collaborate with supply chain and vendor partners to align GPU systems, network equipment, racks, optics, storage, firmware, spares, logistics, and site delivery schedules with program milestones and customer commitments. * Coordinate infrastructure installation and site readiness activities, including rack placement, power and cooling validation, network turn-up, cabling completion, hardware acceptance, and issue remediation across internal teams and external contractors. * Drive commissioning and production readiness reviews for each GPU cluster, ensuring hardware, firmware, networking, storage, automation, observability, security controls, support processes, and service acceptance criteria are fully validated before launch. * Lead software deployment readiness across provisioning, orchestration, monitoring, capacity management, customer onboarding workflows, and platform service enablement to ensure AI Cloud environments are production-ready. * Own customer handover planning and execution, including launch readiness, acceptance criteria, documentation, support transition, known-issue tracking, stakeholder communications, and post-launch stabilization. * Establish portfolio-level governance for concurrent data center and AI Cloud delivery programs, providing executive-level reporting on progress, risks, dependencies, escalations, site readiness, customer impact, and business outcomes. * Identify bottlenecks across process, tooling, vendor execution, installation workflows, commissioning, software deployment, and operational handoffs; drive durable improvements that increase deployment velocity, repeatability, quality, reliability, and cost efficiency. ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Single Server, Global Reach: Running a Worldwide Marketplace on Bare Metal in a Cloud-Dominated World](https://www.wearedevelopers.com/videos/1206-single-server-global-reach-running-a-worldwide-marketplace-on-bare-metal-in-a-cloud-dominated-world) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [How to Avoid Tech Hype Traps - Josip Stuhli](https://www.wearedevelopers.com/videos/1817-how-to-avoid-tech-hype-traps-josip-stuhli) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)