Senior Vision Language Model Engineer

NVIDIA Corporation
Santa Clara, CA, United States
1 day ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$184,000.0 - $287,500.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Computer Vision Computer Engineering Data Cleansing Data Discovery Data Files Software Debugging Distributed Computing Environment Python (Programming Language) Open Source Technology Robotic Automation Software Deep Learning
+1 more
Information Technology

Job description

  • Partner with our researchers to develop and evaluate prototypes of our latest models, such as VLMs and VLAs, for video search, video understanding, and more. Enable fundamental advances in autonomous driving, healthcare, and robotics.
  • Design and implement agentic data workflows that automate data discovery, labeling, evaluation, and retraining to maximize development velocity.
  • Build, curate, and maintain high-quality multimodal datasets (e.g., video, sensor, language/action traces) tailored for end-to-end physical AI problems, such as autonomous driving.
  • Explore and productize new data sources including simulation and synthetic data.
  • Use agentic AI workflows across the full applied research lifecycle.
  • Collaborate with research, model development, performance, and product teams.
  • Contribute to NVIDIA Cosmos Dataset Search and other core NVIDIA platforms and products.

Requirements

  • PhD with 4+ years, MS with 6+ years, or BS (or equivalent experience) with 8+ years of relevant experience in Computer Science, Computer Engineering, or a related technical field
  • Strong background in modern deep learning, including transformer-based architectures, video modeling, and multimodal VLM/VLA or foundation models.
  • Excellent experience training and deploying deep learning models on real-world datasets: data preprocessing, distributed training, evaluation, debugging, and iterative improvement.
  • Excellent experience with python and at least one deep learning framework.
  • Current with the latest research on image and video search in autonomous vehicles, healthcare, robotics, or related physical AI applications.
  • Fluent with agentic AI workflows across the full applied research lifecycle, including prototyping novel algorithms and search pipelines, benchmarking, and integrating prototypes in production codebases.
  • Clear and effective communication skills, with experience working well in a dynamic, product- and research-focused team.

Ways to Stand Out from the Crowd:

  • Strong track record publishing in top-tier conference such as CVPR, NeuRIPS, ICML, ECCV
  • Patents in video retrieval or related field
  • Strong coding architecture skills demonstrated through contributions to large internal or open-source projects.
  • Experience in robotic systems such as autonomous vehicles or humanoid robotics.

Come join us at NVIDIA and contribute to a team that is pushing the edges of what can be done in AI and computer vision. We’re looking for candidates who are innovative, ambitious, and ready to leave a lasting mark on the world!

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Balancing model training with data preparation realities

Lukas Kölbl · LIVE

1:25 min

Distinguishing artificial intelligence from deep learning

Sam Witteveen · Coffee With Developers

3:38 min

Reusing email software standards for HTTP file uploads

Imran Nazar · World Congress 2023

3:16 min

Composing real time video flow applications utilizing multiple AI models

Ankit Patel Ankit Patel · World Congress 2024

3:24 min

Harmonizing federal data records with automated software robots

Clemens Schwaiger · LIVE

2:17 min

Distinguishing between AI, machine learning, and deep learning

Mary Grygleski Mary Grygleski · LIVE

Videos

See all

Related articles

See all