> Markdown version of [/jobs/ext/287231-principal-technical-program-manager-tpm-ai-infrastructure-operations](https://www.wearedevelopers.com/jobs/ext/287231-principal-technical-program-manager-tpm-ai-infrastructure-operations). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Technical Program Manager (TPM) - AI Infrastructure Operations - **Company:** NSCALE, LLC - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Agile Methodology, Artificial Intelligence, Computing Platforms, Computer Networks, Continuous Delivery, Continuous Integration, Data Centers, Linux, Distributed Systems, Network Interface Controllers, Firmware, InfiniBand, Networking Hardware, Uptime, Scrum Methodology, Remote Direct Memory Access, Software Engineering, AI Infrastructure, Graphics Processing Unit (GPU), Computer Network Operations, High Performance Computing, Mttr, Information Technology, Performance Monitor - **Published:** May 22, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=69945ce8465ba930 ## About the Role Do you have experience in Performance monitoring?, * Experience: 5+ years of experience in a Technical Program Management role, successfully driving large-scale, complex infrastructure or software engineering programs. * Technical Domain Knowledge: Strong foundational understanding of data center infrastructure, distributed systems, Linux, and networking concepts. * Program Management Rigor: Proven expertise in modern program management methodologies (Agile, Scrum, PMP certification preferred). Exceptional organizational, communication, and presentation skills. * Metrics-Driven Approach: Demonstrable experience in defining, tracking, and improving system performance based on operational metrics (e.g., Uptime, Availability, MTTR, SLOs/SLIs). * Execution in Ambiguity: Ability to thrive in a fast-paced, high-growth environment, managing multiple priorities and adapting to evolving technical requirements., * Direct experience managing programs related to data center infrastructure build-outs and hardware commissioning processes. * Specific domain knowledge of AI/HPC infrastructure, including NVIDIA GPUs, InfiniBand/RDMA networks, and the challenges of tightly-coupled systems. * Experience in a hyperscale or public cloud environment supporting 24/7 mission-critical services. * Familiarity with SRE principles, automation tooling, and continuous integration/continuous deployment (CI/CD) pipelines for infrastructure. * A Bachelor's or Master's degree in a technical field (Computer Science, Engineering, etc.) or equivalent practical experience. ## Description As a Technical Program Manager (TPM) for AI Infrastructure Operations, you will be the operational backbone of our high-scale, high-performance AI and High-Performance Computing (HPC) environment. You will be responsible for driving complex, cross-functional programs that ensure the stability, availability, and growth of our cutting-edge GPU fleet and Infiniband network fabrics. This role requires a blend of deep technical understanding, rigorous program management, and a relentless focus on delivering against key operational metrics (SLAs, Uptime, Availability). You will bridge the gap between engineering execution and strategic business goals, directly impacting our ability to serve customer workloads at scale., * Program Leadership: Own the planning, execution, and delivery of strategic operational programs, including new data center AI infrastructure build-outs, large-scale fleet software/firmware rollouts, and the implementation of new operational tooling (in partnership with SRE). * Metrics and Reporting: Establish, track, and drive accountability against critical infrastructure KPIs, specifically focusing on Availability (Target 97.5%) and Uptime (Target 99%). Develop clear dashboards and communication rhythms to provide leadership with real-time visibility into operational health, program status, and risk. * Process Engineering: Analyze and optimize operational workflows across Fleet Operations, Network Operations, and SRE teams. Drive the standardization of incident management, change management, and postmortem processes to reduce toil and improve Mean Time to Recovery (MTTR). * Cross-Functional Coordination: Serve as the primary liaison between engineering teams (Hardware, Compute Platform, Network), Data Center Operations, and external vendors (GPU, Network hardware). Proactively identify and resolve dependencies, risks, and roadblocks. * Capacity and Readiness: Partner with Data Science/Operation Programs to translate capacity planning models into actionable infrastructure delivery and readiness roadmaps. Ensure that new hardware (GPUs, NICs, switches) is successfully integrated into the operational control plane and meets go-live criteria. * Risk Management: Proactively identify technical, schedule, and resource risks related to AI infrastructure scaling and stability. Develop mitigation strategies and communicate impacts clearly to stakeholders. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)