Senior Deep Learning Algorithm Engineer

NVIDIA Ltd.
Santa Clara, CA, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$62,400.0 - $112,320.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Application Frameworks C++ (Programming Language) Profiling Computer Programming Computer Engineering Software Debugging Distributed Systems Python (Programming Language) Open Source Technology Large Language Models Deep Learning
+6 more
AI Platforms Information Technology Low Latency Free and Open-Source Software Machine Learning Operations TensorRT

Job description

NVIDIA is seeking a Senior Deep Learning Algorithms Engineer to advance Dynamo, our open-source distributed inference platform for large-scale, low-latency AI services. You’ll lead architecture and performance work across Dynamo and open source frameworks. You’ll collaborate across research, software, systems, and hardware teams to make AI inference faster, more efficient, and easier to deploy. You’ll engage with the broader ecosystem, including vLLM, SGLang, and TensorRT-LLM as well as with external partners to build the best operating system for AI. If you’re excited by deep learning, performance engineering, and distributed systems, we’d love to hear from you.

What you’ll be doing:

  • Design, build, and maintain Dynamo integrations for open source frameworks vLLM, SGLang, TRTLLM.
  • Partner with open source communities to land measurable gains in latency, throughput, reliability, and efficiency.
  • Showcase NVIDIA token/watt leadership by pushing the pareto frontier on public/private benchmarks
  • Find and remove bottlenecks across runtimes, kernels, networking, routing, and orchestration.
  • Develop inference optimizations for scheduling, disaggregation, KV caching, and autoscaling.

Requirements

  • BS, MS, PhD in Computer Science, Electrical Engineering, Computer Engineering, or a related field (or equivalent experience).
  • 3+ years building, profiling, and debugging performance-critical distributed or ML systems.
  • Strong programming skills in Python and/or Rust, C++.
  • Understanding of modern ML architectures and inference techniques

Ways to stand out from the crowd:

  • High agency and a track record of leading ambiguous work end to end.
  • Experience with AI Accelerators
  • Open-source contributions / leadership
  • Research in ML inference or distributed systems.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

About the company

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you’re creative and autonomous, we want to hear from you!

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jofdav.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

5:48 min

Balancing delivery latency with stream reliability and scale

Phil Cluff · LIVE

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · WWC Europe 2026

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · WWC 2025

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

3:37 min

Accessing API documentation and testing remote driving latency

Alexandru Ciinaru Alexandru Ciinaru +3 · WWC 2025

Videos

See all

Related articles

See all