Senior Software Engineer, Evaluation Flywheel - Autonomous Vehicles

NVIDIA Ltd.
Santa Clara, CA, United States
13 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$224,000.0 - $356,500.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Information Engineering Python (Programming Language) Language Modeling Robotic Automation Software Software Construction Management of Software Versions Large Language Models Information Technology Machine Learning Operations Data Pipelines

Job description

NVIDIA is building the future of autonomous driving, and evaluation is how we know the drive is getting better. Our AV Evaluation team owns the metrics, golden datasets, and closed-loop evaluation workflows that decide what ships in NVIDIA’s self-driving stack - every model change runs through us. We are bringing frontier vision-language models and agentic techniques to evaluation at scale, and the early results are changing how our organization develops AI drivers!

We are looking for a senior engineer to own the engine that makes it all trustworthy: the eval flywheel. This is a hands-on technical leadership role - the flywheel’s architecture, quality, and adoption are owned end to end by this person, in code and in the room, not from a strategy document.

What you’ll be doing:

  • Owning the eval flywheel’s strategy and architecture: how road and simulation driving data becomes curated golden datasets, how metrics are measured against them (precision/recall), and how those results earn lasting trust with the teams that depend on them.
  • Setting the standard for evaluation quality: golden dataset curation, versioning, and health; metric performance measurement; and release processes that keep results dependable as the system evolves.
  • Building the tooling that helps our metric developers iterate quickly: self-serve dataset pipelines, metric performance measurement, and quality reporting used every day by the team and our partners.
  • Partnering with senior engineers and leaders across test engineering, behavior planning, and infrastructure - setting expectations, working through trade-offs, and being the voice of evaluation quality in cross-team decisions.
  • Working directly with AI model developers so evaluation iteration speed becomes an advantage for the whole program, including our push into learned, VLM-based evaluation.

Requirements

  • A track record of independent execution and technical leadership: finding the highest-leverage problem, driving it across team boundaries, and delivering without waiting to be asked.
  • Clear, proactive communication with engineers and senior leaders alike, at a fast pace.
  • BS or MS in Computer Science, Robotics, or a related field (or equivalent experience).
  • 12+ years building software, with significant time in autonomous vehicles, robotics, or large-scale ML systems.
  • Deep experience evaluating ML or robotic systems: metric design, ground-truth and golden dataset curation, precision/recall methodology, and the data pipelines behind them.
  • Strong Python and data engineering skills for production-scale pipelines.

Ways to stand out from the crowd:

  • Experience building an evaluation flywheel before - dataset curation, metric measurement, developer tooling - and the story of how it changed model development velocity.
  • Closed-loop simulation evaluation for autonomous driving, and the realism questions that come with it.
  • Applying LLMs or VLMs to evaluation, or productionization the infrastructure behind them.
  • A history of earning trust for metrics across skeptical partner teams.

Benefits & conditions

4.24.2 out of 5 stars 2788 San Tomas Expressway, Santa Clara, CA 95051 $224,000 - $356,500 a year - Full-time, Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff · WWC 2024

1:32 min

Predicting missing words using specialized language models

David vonThenen David vonThenen · WWC Europe 2026

2:59 min

Navigating uncertain execution and expanding technical capabilities

Paul Adams Paul Adams · WWC 2025

1:19 min

Advancing autonomous driving capabilities with specialized software talent

Katrin Lehmann Katrin Lehmann +1 · Coffee With Developers

47 sec

Building modern data pipelines for legacy exports

Dr. Alexander Wachtel Dr. Alexander Wachtel +1 · WWC 2025

10:40 min

Evaluating automotive software architectures and backend technologies

Georg Kühberger +1 · LIVE

Videos

See all

Related articles

See all