Senior Software Engineer, NCCL

NVIDIA Corporation
Santa Clara, CA, United States
1 day ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$152,000.0 - $241,500.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) C++ (Programming Language) Computer Clusters Nvidia CUDA Computer Programming Computer Networks Software Debugging Linux InfiniBand Tensorflow Systems Architecture System Software
+5 more
Graphics Processing Unit (GPU) High Performance Computing Pytorch Deep Learning Parallel Computation

Job description

  • Design, implement and maintain highly-optimized communication runtimes for Deep Learning frameworks (e.g. NCCL for TensorFlow/Pytorch) and HPC programming interfaces (e.g. UCX for MPI/OpenSHMEM) on GPU clusters.
  • Participating in and contributing to parallel programming interface specifications like MPI/OpenSHMEM.
  • Design, implement and maintain system software that enables interactions among GPUs and interactions between GPUs and other system components.
  • Creating proof-of-concepts to evaluate and motivate extensions in programming models, new designs in runtimes and new features in hardware.

Requirements

We are looking for a highly motivated senior software engineer for an exciting role in our communication libraries and network software team. The position will be part of a fast-paced crew that develops and maintains software for complex heterogeneous computing systems that power disruptive products in High Performance Computing and Deep Learning., * M.S./Ph.D. degree in CS/CE or equivalent experience.

  • 5+ years of relevant experience.
  • Excellent C/C++ programming and debugging skills.
  • Strong experience with Linux.
  • Expert understanding of computer system architecture and operating systems.
  • Experience with parallel programming interfaces and communication runtimes.
  • Ability and flexibility to work and communicate effectively in a multi-national, multi-time-zone corporate environment.

Ways to stand out from the crowd:

  • Deep understanding of technology and passionate about what you do.
  • Experience with CUDA programming and NVIDIA GPUs.
  • Knowledge of high-performance networks like InfiniBand, iWARP etc.
  • Experience with HPC applications.
  • Experience with Deep Learning Frameworks such PyTorch, TensorFlow, etc.
  • Strong collaborative and interpersonal skills, specifically a proven ability to effectively guide and influence within a dynamic matrix environment.

Benefits & conditions

NVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most forward-thinking and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you’re a creative and autonomous engineer with real passion for technology, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

You will also be eligible for equity and benefits.

About the company

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues · World Congress 2025

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 · World Congress 2026 Europe

3:23 min

The AI workload technology stack and its components

Lerna Ekmekcioglu Lerna Ekmekcioglu · Europe 2026 Virtual

Videos

See all

Related articles

See all