Technical Program Manager
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
The AI/ML TPM team owns delivery and execution across CoreWeave’s AI/ML Platform Services organization. The team partners closely with Product, Engineering, Research, Infrastructure, and Go-to-Market teams to deliver scalable, reliable, and high-performance platforms that support the full AI lifecycle. AI/ML TPMs drive alignment and execution across highly technical, cross-functional teams to ensure the successful delivery of customer-facing infrastructure and platform capabilities used by researchers, engineers, and enterprise customers.
As a Technical Program Manager, you will lead complex, cross-functional programs across Performance & Benchmarking within our AI/ML Platform Services organization. This team is responsible for ensuring CoreWeave’s infrastructure is performant, stable, and validated for demanding AI workloads before and as it reaches customers. The work spans infrastructure verification, benchmarking, observability, and performance readiness across new hardware platforms, clusters, and model workloads. You will partner with engineering, infrastructure, product, capacity, and go-to-market teams to drive programs that improve workload performance, validate new environments, operationalize benchmarking frameworks, and create visibility into how CoreWeave systems perform across models, hardware generations, and deployment contexts., * Drive end-to-end program execution for performance and benchmarking initiatives spanning infrastructure validation, performance testing, benchmark execution, observability, and launch readiness
- Partner with engineering and infrastructure teams to deliver programs that verify new hardware platforms, clusters, and software environments meet CoreWeave standards for performance and stability
- Lead cross-functional efforts to operationalize benchmarking frameworks that measure model performance, runtime efficiency, GPU utilization, and workload reliability across environments
- Coordinate dependencies across platform engineering, infrastructure, capacity, product, and go-to-market teams to ensure performance findings are translated into roadmap priorities, customer readiness, and external proof points
- Build program mechanisms for release readiness, benchmark planning, risk management, issue escalation, and post-launch review for performance-sensitive infrastructure initiatives
- Establish dashboards, operating cadences, and success metrics to improve performance visibility, infrastructure validation coverage, benchmark repeatability, and time-to-readiness for new platforms
- Help drive prioritization across performance bottlenecks, test gaps, and benchmark requests by aligning stakeholders on goals, tradeoffs, and measurable outcomes
- Create clarity across ambiguous technical programs by aligning teams around performance goals, validation criteria, and execution milestones
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience
- 5+ years of technical program management experience in cloud infrastructure, distributed systems, high-performance computing, or AI/ML platforms
- Experience leading large-scale cross-functional programs involving performance engineering, benchmarking, validation systems, or infrastructure readiness
- Strong technical fluency in distributed systems, GPU or accelerator-based infrastructure, workload performance measurement, and large-scale infrastructure operations
- Demonstrated ability to define program metrics and drive measurable outcomes in performance, reliability, scale, or operational maturity
- Excellent communication skills, with experience influencing engineering, product, and infrastructure stakeholders
- Experience with AI/ML benchmarking, performance analysis, or infrastructure validation for training and inference workloads
- Familiarity with GPU cluster architecture, workload observability, hardware bring-up, and performance bottleneck analysis
- Understanding of benchmarking methodologies, reproducibility, test coverage, and the tradeoffs between performance, stability, utilization, and customer readiness
- Experience building launch processes, release governance, dependency management, and operational review mechanisms in fast-scaling environments
- Familiarity with translating technical performance data into actionable decisions for product, customer, or go-to-market audiences
Wondering if you’re a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams - even if you aren’t a 100% skill or experience match., This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. § 1157, or (iv) asylee under 8 U.S.C. § 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency. CoreWeave may, for legitimate business reasons, decline to pursue any export licensing process.
Benefits & conditions
The base salary range for this role is $177,000 to $237,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility)., In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include:
- Medical, dental, and vision insurance - 100% paid for by CoreWeave
- Company-paid Life Insurance
- Voluntary supplemental life insurance
- Short and long-term disability insurance
- Flexible Spending Account
- Health Savings Account
- Tuition Reimbursement
- Ability to Participate in Employee Stock Purchase Program (ESPP)
- Mental Wellness Benefits through Spring Health
- Family-Forming support provided by Carrot
- Paid Parental Leave
- Flexible, full-service childcare support with Kinside
- 401(k) with a generous employer match
- Flexible PTO
- Catered lunch each day in our office and data center locations
- A casual work environment
- A work culture focused on innovative disruption
About the company
CoreWeave is The Essential Cloud for AI . Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com., At CoreWeave, we work hard, have fun, and move fast! We’re in an exciting stage of hyper-growth that you will not want to miss out on. We’re not afraid of a little chaos, and we’re constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Makes WeAreDevelopers World Congress Different From Every Other Tech Event?
Dev Digest 121 - AI goes offline
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Dev Digest 120 - Apple and peers