Senior Deep Learning Sofware Infrastructure Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+1 more
Job description
Experteer Overview In this role you will help advance NVIDIA’s Autonomous Vehicles initiative by building and scaling training libraries and infrastructure for multi-thousand GPU clusters. You will partner with research and platform teams to accelerate iteration, improve safety, and enable large-scale model training. This is a hands-on opportunity to shape the ML training stack and keep infrastructure robust and scalable as GPU capacity and data grow. You’ll work at the intersection of research, engineering, and platform development to deliver impactful, production-grade systems. Compensation / Benefits * Scale and harden deep learning infrastructure libraries for training on large GPU clusters * Improve efficiency of the training stack (data loaders, distributed training, scheduling, monitoring) * Build robust training pipelines for massive video datasets and rapid experimentation * Collaborate with researchers, model engineers, and platform teams to reduce stalls and improve training availability * Own core infrastructure components such as orchestration libraries, distributed training frameworks, and fault-resilient systems * Partner with leadership to ensure infrastructure scales with GPU capacity and data sizes while maintaining developer efficiency and stability Tasks * BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, or related field, or equivalent experience * 12+ years building and scaling high-performance distributed systems, ideally in ML, HPC, or large-scale data infrastructure * Extensive knowledge of deep learning frameworks (PyTorch preferred), large-scale training (DDP/FSDP, NCCL, tensor/pipeline parallelism), and performance profiling * Strong systems background: datacenter networking (RoCE, IB), parallel filesystems (Lustre), storage systems, schedulers (Slurm, Kubernetes, etc.) * Proficiency in Python with production-grade libraries, orchestration layers, and automation tools * Ability to work with multi-functional teams and translate requirements into robust systems Key requirements * equity * benefits * remote work * competitive base salary * career growth opportunities
Requirements
availability * Own core infrastructure components such as orchestration libraries, distributed training frameworks, and fault-resilient systems * Partner with leadership to ensure infrastructure scales with GPU capacity and data sizes while maintaining developer efficiency and stability Tasks * BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, or related field, or equivalent experience * 12+ years building and scaling high-performance distributed systems, ideally in ML, HPC, or large-scale data infrastructure * Extensive knowledge of deep learning frameworks (PyTorch preferred), large-scale training (DDP/FSDP, NCCL, tensor/pipeline parallelism), and performance profiling * Strong systems background: datacenter networking (RoCE, IB), parallel filesystems (Lustre), storage systems, schedulers (Slurm, Kubernetes, etc.) * Proficiency in Python with production-grade libraries, orchestration layers, and automation tools * Ability to work with multi-functional teams and aaaaaaaq _ requirements into robust systems Key requirements * equity * benefits * remote work * competitive base salary * career growth opportunities
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
How to Become an AI Engineer
Fully Remote Software Engineer Jobs
Best Countries for Software Engineers