> Markdown version of [/jobs/ext/2708406-hardware-technical-program-manager-infrastructure-partner-operations](https://www.wearedevelopers.com/jobs/ext/2708406-hardware-technical-program-manager-infrastructure-partner-operations). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Hardware Technical Program Manager, Infrastructure Partner Operations - **Company:** OpenAI Inc. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Cloud Computing, Cloud Engineering, Information Systems, Computer Engineering, Data Centers, Distributed Systems, Executive Information Systems, Reliability Engineering, Cloud Services, AI Infrastructure, Google Cloud, High Performance Computing, Information Technology, Data Analytics, Performance Monitor, Hardware Infrastructure, Oracle Cloud Infrastructure - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/hardware-technical-program-manager-infrastructure-partner-operations-openai-8777805 ## About the Role * 7+ years of experience in Technical Program Management, Infrastructure Operations, Cloud Operations, Service Delivery, or Technical Account Management within large-scale infrastructure environments. * Experience managing operational relationships with external infrastructure providers, cloud service providers, hardware vendors, or strategic technology partners. * Strong understanding of hyperscale cloud infrastructure, data center operations, infrastructure delivery, or large-scale distributed systems. * Experience developing operational KPIs, SLAs, service health metrics, dashboards, and executive reporting for complex technical organizations. * Demonstrated success leading cross-functional operational programs involving both internal stakeholders and external partners. * Strong program management skills with the ability to drive accountability across organizations without direct authority. * Excellent written and verbal communication skills with experience presenting operational performance to senior technical and executive leadership. * Bachelor's degree in Engineering, Computer Science, Information Systems, Operations, or equivalent practical experience. Preferred Skills * Experience managing cloud infrastructure operations within organizations such as Microsoft Azure, Amazon Web Services (AWS), Google Cloud Platform (GCP), Oracle Cloud Infrastructure (OCI), or other hyperscale cloud providers. * Experience leading operational governance, service delivery, customer success engineering, technical account management, or infrastructure operations for enterprise cloud customers. * Strong understanding of service-level agreements (SLAs), operational KPIs, incident management, escalation processes, root cause analysis, and continuous service improvement methodologies. * Experience building executive dashboards, operational scorecards, business review frameworks, and data-driven performance reporting. * Familiarity with infrastructure operations supporting GPU infrastructure, AI infrastructure, high-performance computing (HPC), or hyperscale data center environments. * Experience managing complex cross-company technical relationships while balancing customer priorities, engineering constraints, and operational execution. * Proven ability to influence senior stakeholders across both internal teams and external partner organizations without direct authority. * Experience driving continuous operational improvements through metrics, process optimization, and structured governance. ## Description The Industrial Compute team is responsible for building the physical infrastructure that powers OpenAI's largest-scale AI systems. We design, deploy, and operate next-generation compute infrastructure across a rapidly expanding global footprint, combining OpenAI-owned infrastructure with strategic cloud and infrastructure partners to support frontier AI workloads. As our infrastructure footprint grows, operational excellence across third-party providers becomes increasingly critical. Our team ensures external infrastructure partners consistently deliver the reliability, performance, and operational maturity required to support OpenAI's rapidly expanding compute environment. About the Role We are seeking a Hardware Technical Program Manager, Infrastructure Partner Operations to lead operational delivery across OpenAI's third-party infrastructure partners, including major cloud service providers and strategic compute vendors. In this role, you will serve as the primary operational program manager for external infrastructure partners, driving accountability for service delivery, operational readiness, incident management, performance reporting, and continuous operational improvement. You will work closely with partner engineering and operations teams while coordinating internally across Hardware Engineering, Infrastructure Operations, Capacity Planning, Networking, Supply Chain, Deployment, Reliability Engineering, and executive leadership. Success in this role requires someone who understands how hyperscale infrastructure organizations operate, can establish strong operational governance with external partners, and is comfortable driving complex technical programs without direct ownership of the underlying infrastructure., * Own operational engagement with third-party infrastructure providers, ensuring consistent execution against operational commitments, service-level agreements (SLAs), and performance expectations. * Develop operational governance frameworks with strategic partners, including business reviews, operational scorecards, escalation processes, executive reporting, and performance improvement plans. * Define, track, and continuously improve key operational metrics related to infrastructure availability, deployment execution, incident response, operational health, service quality, and partner performance. * Build dashboards and reporting mechanisms that provide clear visibility into partner operational performance, risks, trends, and areas requiring executive attention. * Drive cross-functional coordination between OpenAI teams and external infrastructure providers to resolve operational issues, remove execution blockers, and improve delivery outcomes. * Lead operational escalations involving infrastructure availability, deployment execution, hardware operations, capacity delivery, or service performance, ensuring timely resolution and clear executive communication. * Establish repeatable operating rhythms with external partners, including weekly operational reviews, executive business reviews, service reviews, action tracking, and long-term improvement initiatives. * Partner with Capacity Planning, Hardware Operations, Networking, Deployment, Reliability Engineering, and Supply Chain teams to ensure external infrastructure providers remain aligned with OpenAI's operational priorities. * Identify systemic operational risks across partner organizations and proactively drive corrective actions that improve long-term operational effectiveness. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [The Cloud is Calling: Answer with In-Demand Skills](https://www.wearedevelopers.com/videos/945-the-cloud-is-calling-answer-with-in-demand-skills) - [The Sustainability Race: AI's Promises, Pitfalls and Potential](https://www.wearedevelopers.com/videos/100155-the-sustainability-race-ai-s-promises-pitfalls-and-potential) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [Reference Architecture of AI in the Cloud](https://www.wearedevelopers.com/videos/1613-reference-architecture-of-ai-in-the-cloud) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)