Senior Machine Learning and Simulation Engineer...

NVIDIA Ltd.
Santa Clara, United States of America
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 224K

Job location

Santa Clara, United States of America

Tech stack

Big Data
C++
Computer Clusters
Computer Programming
Software Debugging
Job Scheduling
Python
Machine Learning
Software Engineering
AI Infrastructure
Reinforcement Learning
Large Language Models
Deep Learning
Kubernetes
Information Technology
Slurm
Data Pipelines

Job description

We are seeking exceptional Senior Machine Learning and Simulation Engineers to join NVIDIA's Autonomous Vehicles (AV) Simulation team! This role requires strong technical leadership and outstanding software engineering skills, coupled with deep expertise in both simulation and artificial intelligence, including deep learning, reinforcement learning, end-to-end driving and Physics AI models. The successful candidate will have a solid track record of productizing ML solutions for autonomous driving and simulation at scale.

This position centers on developing a Closed-Loop Simulation-based Reinforcement Learning (RL) framework in order to train advanced end-to-end AV models, such as Alpamayo R1 (https://research.nvidia.com/publication/2025-10_alpamayo-r1) . This position will design and improve the accuracy and performance of the RL framework and simulation, leveraging SOTA techs including NuRec (https://docs.nvidia.com/nurec/index.html) , Traffic Models (https://research.nvidia.com/labs/avg/publication/zhang.karkus.etal.cvpr2025/) , and Cosmos World Model (https://www.nvidia.com/en-us/ai/cosmos/) . Success in this role requires close collaboration with the AV Platform, AV Product, and Research teams.

What you will be doing:

  • Lead the design and development of large-scale RL training frameworks to accelerate the development of multi-modal AV foundation models.

  • Design, build, and optimize simulation and data processing pipelines to enable scalable training of driving policies.

  • Focus on measuring and enhancing simulation quality and refining the reward function for RL training.

  • Ensure the reliability and performance of training workflows on large GPU clusters through the development of robust monitoring and debugging tools.

  • Partner with researchers to integrate state-of-the-art model architectures into efficient and scalable training pipelines.

Requirements

  • Bachelor's degree in Computer Science, Robotics, Engineering, or a related field (or equivalent experience).

  • 12+ years of relevant professional experience encompassing large-scale ML training, AV systems, simulation, and AI infrastructure development.

  • Deep proficiency in RL algorithms, such as PPO and GRPO, including practical experience with hyperparameter tuning and reward function design.

  • Exceptional programming skills in C++ and Python, vital for developing efficient systems and data pipelines.

  • Extensive experience with large-scale GPU clusters, High-Performance Computing (HPC) environments, and job scheduling/orchestration tools (e.g., Kubernetes, SLURM).

Ways to stand out from the crowd:

  • Experience in RL infrastructure or general LLM training/fine-tuning infrastructure in industry.

  • Experience in simulation & closed-loop evaluation of autonomous driving end-to-end models.

  • Proven record on large-scale data pipeline development and algorithm optimization.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6.

Apply for this position