> Markdown version of [/jobs/ext/1971136-senior-deep-learning-sofware-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/1971136-senior-deep-learning-sofware-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Deep Learning Sofware Infrastructure Engineer - **Company:** NVIDIA Ltd. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Big Data, Computer Clusters, Computer Engineering, File Systems, Distributed Computing Environment, Distributed Systems, Python (Programming Language), Machine Learning, Pytorch, Deep Learning, Kubernetes, Information Technology, Slurm - **Published:** August 7, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/senior-deep-learning-sofware-infrastructure-engineer-usa-58831440 ## About the Role availability * Own core infrastructure components such as orchestration libraries, distributed training frameworks, and fault-resilient systems * Partner with leadership to ensure infrastructure scales with GPU capacity and data sizes while maintaining developer efficiency and stability Tasks * BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, or related field, or equivalent experience * 12+ years building and scaling high-performance distributed systems, ideally in ML, HPC, or large-scale data infrastructure * Extensive knowledge of deep learning frameworks (PyTorch preferred), large-scale training (DDP/FSDP, NCCL, tensor/pipeline parallelism), and performance profiling * Strong systems background: datacenter networking (RoCE, IB), parallel filesystems (Lustre), storage systems, schedulers (Slurm, Kubernetes, etc.) * Proficiency in Python with production-grade libraries, orchestration layers, and automation tools * Ability to work with multi-functional teams and aaaaaaaq _ requirements into robust systems Key requirements * equity * benefits * remote work * competitive base salary * career growth opportunities ## Description Experteer Overview In this role you will help advance NVIDIA's Autonomous Vehicles initiative by building and scaling training libraries and infrastructure for multi-thousand GPU clusters. You will partner with research and platform teams to accelerate iteration, improve safety, and enable large-scale model training. This is a hands-on opportunity to shape the ML training stack and keep infrastructure robust and scalable as GPU capacity and data grow. You'll work at the intersection of research, engineering, and platform development to deliver impactful, production-grade systems. Compensation / Benefits * Scale and harden deep learning infrastructure libraries for training on large GPU clusters * Improve efficiency of the training stack (data loaders, distributed training, scheduling, monitoring) * Build robust training pipelines for massive video datasets and rapid experimentation * Collaborate with researchers, model engineers, and platform teams to reduce stalls and improve training availability * Own core infrastructure components such as orchestration libraries, distributed training frameworks, and fault-resilient systems * Partner with leadership to ensure infrastructure scales with GPU capacity and data sizes while maintaining developer efficiency and stability Tasks * BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, or related field, or equivalent experience * 12+ years building and scaling high-performance distributed systems, ideally in ML, HPC, or large-scale data infrastructure * Extensive knowledge of deep learning frameworks (PyTorch preferred), large-scale training (DDP/FSDP, NCCL, tensor/pipeline parallelism), and performance profiling * Strong systems background: datacenter networking (RoCE, IB), parallel filesystems (Lustre), storage systems, schedulers (Slurm, Kubernetes, etc.) * Proficiency in Python with production-grade libraries, orchestration layers, and automation tools * Ability to work with multi-functional teams and translate requirements into robust systems Key requirements * equity * benefits * remote work * competitive base salary * career growth opportunities ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)