> Markdown version of [/jobs/ext/2283330-senior-vision-language-model-engineer](https://www.wearedevelopers.com/jobs/ext/2283330-senior-vision-language-model-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Vision Language Model Engineer - **Company:** NVIDIA Corporation - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $184,000.0 - $287,500.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computer Vision, Computer Engineering, Data Cleansing, Data Discovery, Data Files, Software Debugging, Distributed Computing Environment, Python (Programming Language), Open Source Technology, Robotic Automation Software, Deep Learning, Information Technology - **Published:** August 28, 2026 - **Apply:** https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-Vision-Language-Model-Engineer_JR2020818 ## About the Role * PhD with 4+ years, MS with 6+ years, or BS (or equivalent experience) with 8+ years of relevant experience in Computer Science, Computer Engineering, or a related technical field * Strong background in modern deep learning, including transformer-based architectures, video modeling, and multimodal VLM/VLA or foundation models. * Excellent experience training and deploying deep learning models on real-world datasets: data preprocessing, distributed training, evaluation, debugging, and iterative improvement. * Excellent experience with python and at least one deep learning framework. * Current with the latest research on image and video search in autonomous vehicles, healthcare, robotics, or related physical AI applications. * Fluent with agentic AI workflows across the full applied research lifecycle, including prototyping novel algorithms and search pipelines, benchmarking, and integrating prototypes in production codebases. * Clear and effective communication skills, with experience working well in a dynamic, product- and research-focused team. Ways to Stand Out from the Crowd: * Strong track record publishing in top-tier conference such as CVPR, NeuRIPS, ICML, ECCV * Patents in video retrieval or related field * Strong coding architecture skills demonstrated through contributions to large internal or open-source projects. * Experience in robotic systems such as autonomous vehicles or humanoid robotics. Come join us at NVIDIA and contribute to a team that is pushing the edges of what can be done in AI and computer vision. We're looking for candidates who are innovative, ambitious, and ready to leave a lasting mark on the world! ## Description * Partner with our researchers to develop and evaluate prototypes of our latest models, such as VLMs and VLAs, for video search, video understanding, and more. Enable fundamental advances in autonomous driving, healthcare, and robotics. * Design and implement agentic data workflows that automate data discovery, labeling, evaluation, and retraining to maximize development velocity. * Build, curate, and maintain high-quality multimodal datasets (e.g., video, sensor, language/action traces) tailored for end-to-end physical AI problems, such as autonomous driving. * Explore and productize new data sources including simulation and synthetic data. * Use agentic AI workflows across the full applied research lifecycle. * Collaborate with research, model development, performance, and product teams. * Contribute to NVIDIA Cosmos Dataset Search and other core NVIDIA platforms and products. ## Related Videos - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [Nemotron: NVIDIA's open model strategy for developers](https://www.wearedevelopers.com/videos/100064-nemotron-nvidia-s-open-model-strategy-for-developers) - [RPA in the Public Sector](https://www.wearedevelopers.com/videos/86-rpa-in-the-public-sector) - [Introduction to TXT](https://www.wearedevelopers.com/videos/30-introduction-to-txt) - [30 Golden Rules of Deep Learning Performance](https://www.wearedevelopers.com/videos/11-30-golden-rules-of-deep-learning-performance) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)