Senior Machine Learning and Simulation Engineer - Autonomous Vehicles

NVIDIA Corporation
Santa Clara, CA, United States
6 days ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$224,000.0
Working hours
Regular working hours

Tech stack

Big Data C++ (Programming Language) Computer Clusters Computer Programming Software Debugging Job Scheduling Python (Programming Language) Machine Learning Software Engineering AI Infrastructure Reinforcement Learning Large Language Models
+5 more
Deep Learning Kubernetes Information Technology Slurm Data Pipelines

Job description

This position centers on developing a Closed-Loop Simulation-based Reinforcement Learning (RL) framework in order to train advanced end-to-end AV models, such as Alpamayo R1. This position will design and improve the accuracy and performance of the RL framework and simulation, leveraging SOTA techs including NuRec, Traffic Models, and Cosmos World Model. Success in this role requires close collaboration with the AV Platform, AV Product, and Research teams.

What you will be doing:

  • Lead the design and development of large-scale RL training frameworks to accelerate the development of multi-modal AV foundation models.
  • Design, build, and optimize simulation and data processing pipelines to enable scalable training of driving policies.
  • Focus on measuring and enhancing simulation quality and refining the reward function for RL training.
  • Ensure the reliability and performance of training workflows on large GPU clusters through the development of robust monitoring and debugging tools.
  • Partner with researchers to integrate state-of-the-art model architectures into efficient and scalable training pipelines.

Requirements

We are seeking exceptional Senior Machine Learning and Simulation Engineers to join NVIDIA’s Autonomous Vehicles (AV) Simulation team! This role requires strong technical leadership and outstanding software engineering skills, coupled with deep expertise in both simulation and artificial intelligence, including deep learning, reinforcement learning, end-to-end driving and Physics AI models. The successful candidate will have a solid track record of productizing ML solutions for autonomous driving and simulation at scale., * Bachelor’s degree in Computer Science, Robotics, Engineering, or a related field (or equivalent experience).

  • 12+ years of relevant professional experience encompassing large-scale ML training, AV systems, simulation, and AI infrastructure development.
  • Deep proficiency in RL algorithms, such as PPO and GRPO, including practical experience with hyperparameter tuning and reward function design.
  • Exceptional programming skills in C++ and Python, vital for developing efficient systems and data pipelines.
  • Extensive experience with large-scale GPU clusters, High-Performance Computing (HPC) environments, and job scheduling/orchestration tools (e.g., Kubernetes, SLURM).

Ways to stand out from the crowd:

  • Experience in RL infrastructure or general LLM training/fine-tuning infrastructure in industry.
  • Experience in simulation & closed-loop evaluation of autonomous driving end-to-end models.
  • Proven record on large-scale data pipeline development and algorithm optimization.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:19 min

Advancing autonomous driving capabilities with specialized software talent

Katrin Lehmann Katrin Lehmann +1 · Coffee With Developers

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · World Congress 2022

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

Videos

See all

Related articles

See all