Senior Deep Learning Sofware Infrastructure Engineer

NVIDIA Ltd.
United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Big Data Computer Clusters Computer Engineering File Systems Distributed Computing Environment Distributed Systems Python (Programming Language) Machine Learning Pytorch Deep Learning Kubernetes Information Technology
+1 more
Slurm

Job description

Experteer Overview In this role you will help advance NVIDIA’s Autonomous Vehicles initiative by building and scaling training libraries and infrastructure for multi-thousand GPU clusters. You will partner with research and platform teams to accelerate iteration, improve safety, and enable large-scale model training. This is a hands-on opportunity to shape the ML training stack and keep infrastructure robust and scalable as GPU capacity and data grow. You’ll work at the intersection of research, engineering, and platform development to deliver impactful, production-grade systems. Compensation / Benefits * Scale and harden deep learning infrastructure libraries for training on large GPU clusters * Improve efficiency of the training stack (data loaders, distributed training, scheduling, monitoring) * Build robust training pipelines for massive video datasets and rapid experimentation * Collaborate with researchers, model engineers, and platform teams to reduce stalls and improve training availability * Own core infrastructure components such as orchestration libraries, distributed training frameworks, and fault-resilient systems * Partner with leadership to ensure infrastructure scales with GPU capacity and data sizes while maintaining developer efficiency and stability Tasks * BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, or related field, or equivalent experience * 12+ years building and scaling high-performance distributed systems, ideally in ML, HPC, or large-scale data infrastructure * Extensive knowledge of deep learning frameworks (PyTorch preferred), large-scale training (DDP/FSDP, NCCL, tensor/pipeline parallelism), and performance profiling * Strong systems background: datacenter networking (RoCE, IB), parallel filesystems (Lustre), storage systems, schedulers (Slurm, Kubernetes, etc.) * Proficiency in Python with production-grade libraries, orchestration layers, and automation tools * Ability to work with multi-functional teams and translate requirements into robust systems Key requirements * equity * benefits * remote work * competitive base salary * career growth opportunities

Requirements

availability * Own core infrastructure components such as orchestration libraries, distributed training frameworks, and fault-resilient systems * Partner with leadership to ensure infrastructure scales with GPU capacity and data sizes while maintaining developer efficiency and stability Tasks * BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, or related field, or equivalent experience * 12+ years building and scaling high-performance distributed systems, ideally in ML, HPC, or large-scale data infrastructure * Extensive knowledge of deep learning frameworks (PyTorch preferred), large-scale training (DDP/FSDP, NCCL, tensor/pipeline parallelism), and performance profiling * Strong systems background: datacenter networking (RoCE, IB), parallel filesystems (Lustre), storage systems, schedulers (Slurm, Kubernetes, etc.) * Proficiency in Python with production-grade libraries, orchestration layers, and automation tools * Ability to work with multi-functional teams and aaaaaaaq _ requirements into robust systems Key requirements * equity * benefits * remote work * competitive base salary * career growth opportunities

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · WWC Europe 2026

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · WWC 2025

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

Videos

See all

Related articles

See all