Senior Research Engineer - Enterprise Products

NVIDIA Ltd.
Washington, DC, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$101,535.0 - $155,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Artificial Neural Networks Code Review Distributed Systems Machine Learning Natural Language Processing Open Source Technology Tensorflow System Software Speech Recognition Graphics Processing Unit (GPU) Pytorch
+4 more
Large Language Models Deep Learning Generative AI Information Technology

Job description

We are now looking for a Senior Research Engineer passionate about Generative AI inference. Are you excited to change the way people infuse AI into products and services? NVIDIA is at the forefront of generative AI models, from language to images. NVIDIA provides building blocks to democratize AI and make generative AI easy to develop, integrate, and deploy. Our team is dedicated to developing optimized inferencing technologies to support our growing generative AI needs. We contribute to all steps of the machine learning lifecycle: from conceptualization, to applied research, engineering for optimized inference, and deployment. Collaborate with research teams, engineers, and open-source community. What you will be doing: Design and evaluate routing policies for LLM traffic to best use mixture of model systems. Build and run agentic benchmarks (e.g., Terminal-Bench ) to measure algorithm quality, and turn results into calibration data and routing profiles Ship to an open-source repo: design docs, code review, docs, and community contributions Collaborating with engineering teams across all of NVIDIA to ensure our software integrates seamlessly up and down the NVIDIA accelerated serving stack.

Requirements

Bachelor’s of Master’s degree in Computer Science or equivalent experience. 8+ years of industry experience in Deep Learning frameworks (PyTorch or TensorFlow). Experience designing or running LLM evaluations/benchmarks - ideally agentic ones - and drawing statistically sound conclusions from them Understanding of modern techniques in Machine Learning, Deep Neural Networks, Natural Language Processing, or Speech Recognition. Empirical research mindset: forming hypotheses about new algorithms, running calibrations, iterating on results Strong communication and interpersonal skills, along with the ability to work in a dynamic and distributed team. A history of mentoring junior engineers and interns is a huge plus. A desire to constantly grow and learn new things. Strong computer science fundamentals - algorithms and data structures, computational complexity, parallel and distributed computing, system software. Ways to stand out from a crowd: Experience architecting or developing large-scale distributed systems for deep learning. Agentic benchmark creation and publications. Knowledge of CPU and/or GPU architecture.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 192,000 USD - 304,750 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and . Applications for this job will be accepted at least until July 14, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:36 min

Exploring high-level Python frameworks for accelerated enterprise artificial intelligence

Paul Graham Paul Graham · LIVE

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

1:39 min

Fundamentals of tensors and the TensorFlow library

Håkan Silfvernagel · LIVE

3:39 min

Addressing code review surrender and process exploitation

Laura Tacho Laura Tacho · WWC Europe 2026

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

4:41 min

Replacing PyTorch with ONNX runtime for AWS Lambda deployments

Marek Suppa · LIVE

Videos

See all

Related articles

See all