Senior AI Systems and Algorithms Engineer

NVIDIA Ltd.
Santa Clara, CA, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$62,400.0 - $112,320.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Big Data Software Debugging Distributed Computing Environment Distributed Systems Python (Programming Language) Machine Learning Language Modeling Open Source Technology Performance Tuning Software Engineering AI Infrastructure
+8 more
Reinforcement Learning Pytorch Large Language Models Deep Learning Information Technology Low Latency HuggingFace Stable Diffusion

Job description

NVIDIA is seeking a Senior GenAI Algorithms Engineer to advance the state of the art in foundation model development, training, and deployment. You will work at the intersection of large-scale distributed training, reinforcement learning for LLMs/VLMs, model efficiency, multimodal AI, and open-source AI infrastructure. This role spans the entire GenAI lifecycle from large-scale data preparation to training, post-training, inference optimization, and framework development. You will collaborate with research, product, and infrastructure teams to design new algorithms, optimize existing systems, and contribute to NVIDIA’s open-source AI stack, including Megatron-LM, Megatron Bridge, and NeMo-RL.

What You’ll Be Doing:

  • Data Curation & Readiness: Design scalable systems for preparing high-quality multimodal datasets for frontier foundation model training.
  • Training Efficiency: Develop algorithms and systems that improve the scalability, efficiency, and cost of large-scale pre-training and post-training.
  • Inference Efficiency: Advance techniques that improve inference performance, reduce deployment cost, and enable efficient serving across cloud and edge platforms.
  • Open-Source AI Infrastructure: Develop reusable infrastructure and contribute brand new model support to NVIDIA’s open-source GenAI training platform.

Requirements

  • MS or Ph.D in Computer Science, AI, Applied Mathematics, or a related field (or equivalent experience).
  • 5+ years of relevant industry experience.
  • Strong foundation in machine learning, deep learning, and optimization.
  • Excellent software engineering skills, including Python and PyTorch.
  • Experience building high-performance software for large-scale AI systems.
  • Strong analytical, debugging, and performance optimization skills.
  • Excellent communication and collaboration skills.

Ways to stand out from the crowd:

Experience in some of the following areas is highly desirable:

  • Large-Scale Training: Distributed training at scale, including Megatron-LM, Megatron Bridge, FSDP, TP/PP/CP/DP, heterogeneous or per-module parallelism, optimizer research, and efficient sparse or long-context attention.
  • LLM/VLM Post-Training: Supervised fine-tuning (SFT), reinforcement learning for LLMs (e.g., PPO, GRPO, asynchronous RL), and large-scale RL frameworks such as NeMo-RL.
  • Inference Efficiency: Model compression techniques including quantization (FP8, NVFP4, INT4), pruning, knowledge distillation, neural architecture search, and diffusion or non-autoregressive language models.
  • Open-Source AI Infrastructure: Contributing to open-source AI frameworks such as Megatron-LM, Megatron Bridge, NeMo-RL, or Hugging Face Transformers along with experience in GPU performance optimization, distributed systems, latency/throughput analysis, and profiling of large-scale AI workloads.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jofdav.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

4:41 min

Replacing PyTorch with ONNX runtime for AWS Lambda deployments

Marek Suppa · LIVE

5:01 min

Leveraging large language models for code optimization and development

Stephan Gillich Stephan Gillich +3 · WWC 2024

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

Videos

See all

Related articles

See all