Senior Technical Program Manager, GenAI and Models
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+6 more
Job description
How do bold AI research ideas become diligent training, evaluation, and production systems? NVIDIA’s Deep Learning Software team is looking for a Senior Technical Program Manager to lead programs across model pre-training, production RL runs, evaluation, and agentic AI infrastructure. We build the software foundations that help research and engineering teams train, evaluate, and deliver sophisticated AI models. We partner across research, platform engineering, distributed computing, evaluation, and open-source development to turn sophisticated technical goals into clear software plans. Come help us improve how sophisticated AI systems are built and validated!
What you’ll be doing:
- Lead multi-functional programs across training frameworks, evaluation environments, agent and model runtimes, datasets, verifiers, and distributed training infrastructure.
- Partner with AI researchers, engineering leaders, product, infrastructure, and QA teams to define roadmaps, achievements, release plans, and measurable success criteria.
- Coordinate large-scale RL training and evaluation experiments, including handling GPU resources, dependency tracking, run scheduling, results reporting, release readiness, technical decisions, integration plans, and program updates.
Requirements
- Bachelor’s degree in computer science, engineering, or a related technical field, or equivalent experience.
- 10+ years of technical program management, engineering program management, or related experience delivering sophisticated software platforms.
- Experience leading global, matrixed programs across research, software engineering, infrastructure, QA, release teams, and partner groups.
- Strong understanding of the AI model lifecycle, including training, post-training, evaluation, experimentation, production readiness, reinforcement learning concepts, and GPU-accelerated distributed systems.
- Experience running software releases across repositories, dependencies, test configurations, quality gates, collaborator approvals, open-source workflows, CI/CD systems, and tools such as GitHub, Git, Jira, Linear, Aha!, or Confluence., * Experience supporting reinforcement learning, post-training, agentic AI, or large-scale model evaluation programs.
- Familiarity with PPO, GRPO, asynchronous RL, distributed inference, rollout generation, policy optimization, evaluation harnesses, verifiers, benchmark development, or reproducible experimentation.
- Knowledge of GPU infrastructure, distributed training, Kubernetes, workload schedulers, cluster capacity management, performance analysis, open-source contributor workflows, release readiness, or operational metrics.
Benefits & conditions
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 258,750 USD for Level 4, and 200,000 USD - 322,000 USD for Level 5.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.disabledperson.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?
MLOps – What’s the deal behind it?
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence