> Markdown version of [/jobs/ext/2716127-technical-program-manager-cluster-orchestration-applied-training](https://www.wearedevelopers.com/jobs/ext/2716127-technical-program-manager-cluster-orchestration-applied-training). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Technical Program Manager - Cluster Orchestration & Applied Training - **Company:** Coreweave, Inc. - **Location:** Sunnyvale, United States - **Experience:** Expert - **Salary:** $237,000.0 - $261,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cloud Computing, Distributed Systems, Performance Tuning, Azure Machine Learning, Reinforcement Learning, AI Platforms, Kubernetes, Information Technology, Enterprise Integration, Slurm, Hardware Infrastructure - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/technical-program-manager-cluster-orchestration-applied-training-coreweave-2-7962184 ## About the Role This is a highly cross-functional role for a TPM who combines strong technical depth, excellent execution instincts, and the ability to bring structure and clarity to fast-moving infrastructure and AI platform initiatives., * Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience. * 8+ years of technical program management experience in cloud infrastructure, distributed systems, or AI/ML platforms. * Experience leading large-scale cross-functional programs involving scheduling systems, cluster infrastructure, or ML platform capabilities. * Strong technical fluency in Kubernetes, Slurm or comparable schedulers, distributed systems, and AI training workflows. * Demonstrated ability to define program metrics and deliver measurable outcomes in performance, reliability, scale, or operational maturity. * Excellent communication skills, with experience influencing engineering, product, and executive stakeholders. Preferred: * Experience with orchestration and scheduling technologies such as Kubernetes, Slurm, Kueue, Ray, or similar systems. * Familiarity with modern AI training and evaluation workflows, including pre-training, supervised fine-tuning, reinforcement learning, and experiment or sandbox environments. * Understanding of GPU infrastructure, cluster capacity planning, multi-tenant execution, and distributed training tradeoffs. * Experience building launch processes, release governance, dependency management, and operational review mechanisms in fast-scaling environments. * Familiarity with AI developer and research tooling such as W&B, SkyPilot, or adjacent ecosystem platforms. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren't a 100% skill or experience match. ## Description In this role, you will partner with engineering, product, infrastructure, and research-adjacent teams to improve both how workloads run on the cluster and how users interact with the training platform built on top of it. That includes driving programs across orchestration systems such as Slurm-on-Kubernetes (SUNK), Kueue, and workflow integrations, while also helping scale the environments, tooling, and operational mechanisms that make training and evaluation workflows easier to use., * Drive end-to-end program execution for cluster orchestration initiatives spanning workload scheduling, self-service provisioning, upgrade and migration flows, and platform integrations. * Lead cross-functional programs that improve how AI training, evaluation, RL, and mixed workloads run across CoreWeave clusters. * Partner with engineering and product leaders to define roadmap priorities and deliver measurable improvements in utilization, reliability, scalability, observability, and user experience. * Drive delivery for applied training initiatives across pre-training, fine-tuning, reinforcement learning, sandbox environments, and evaluation systems. * Coordinate dependencies across platform engineering, infrastructure, product, customer-facing teams, and ecosystem partners to ensure successful launches and clear operational ownership. * Build program mechanisms for release readiness, rollout planning, risk management, stakeholder communication, and post-launch review. * Establish success metrics, dashboards, and operating cadences to improve cluster efficiency, workload startup performance, time-to-research, and adoption of new platform capabilities. * Create clarity across ambiguous technical programs by aligning stakeholders, surfacing tradeoffs early, and driving decisions to resolution. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [This App Reached 10,000 Users in One Week. Here's How.](https://www.wearedevelopers.com/videos/100329-this-app-reached-10-000-users-in-one-week-here-s-how) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [Hiring AI Native Talents](https://www.wearedevelopers.com/videos/100268-hiring-ai-native-talents) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Should Tech Managers Be Developers First? Pros and Cons](https://www.wearedevelopers.com/magazine/327-should-tech-managers-be-developers-first-pros-and-cons) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)