> Markdown version of [/jobs/ext/1505738-engineering-manager-deep-learning-inference](https://www.wearedevelopers.com/jobs/ext/1505738-engineering-manager-deep-learning-inference). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Engineering Manager, Deep Learning Inference - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Experienced - **Salary:** $224,000.0 - **Contract:** Permanent contract - **Skills:** Agile Methodology, Artificial Intelligence, Application Frameworks, C++ (Programming Language), Profiling, Collaborative Software, Nvidia CUDA, Computer Engineering, Python (Programming Language), Performance Tuning, Software Deployment, Software Engineering, Graphics Processing Unit (GPU), Pytorch, Large Language Models, Deep Learning, Gpu Programming, Information Technology, Machine Learning Operations, TensorRT - **Published:** July 30, 2026 - **Apply:** https://www.disabledperson.com/jobs/73923684-engineering-manager-deep-learning-inference ## About the Role * MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related field. * 6+ overall years of software development experience, including 3+ years in technical leadership or engineering management. * Strong background in C/C++ software design and development; proficiency in Python is a plus. * Hands-on experience with GPU programming (CUDA, Triton, CUTLASS) and performance optimization. * Proven record of deploying or optimizing deep learning models in production environments. * Experience leading teams using Agile or collaborative software development practices. Ways to Stand out from The Crowd: * Significant open-source contributions to deep learning or inference frameworks such as PyTorch, vLLM, SGLang, Triton, or TensorRT-LLM. * Deep understanding of multi-GPU communications (NIXL, NCCL, NVSHMEM) and distributed inference architectures. * Expertise in performance modeling, profiling, and system-level optimization across CPU and GPU platforms. * Proven ability to mentor engineers, guide architectural decisions, and deliver complex projects with measurable impact. * Publications, patents, or talks on LLM serving, model optimization, or GPU performance engineering. ## Description NVIDIA is seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI model deployment. You will shape the software powering today's most sophisticated AI systems - from large language models to multimodal generative AI - all accelerated on NVIDIA GPUs. The Deep Learning Inference team develops and optimizes open-source frameworks that make AI deployment scalable, efficient, and accessible - including SGLang, vLLM, and FlashInfer. Our work enables developers worldwide to harness NVIDIA accelerators for real-time inference at every scale, from datacenter clusters to edge devices. What you'll be doing: * Lead, mentor, and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software. * Guide the strategy, roadmap, and execution of NVIDIA's OSS inference frameworks engineering. * Partner with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines across NVIDIA accelerators. * Oversee performance tuning, profiling, and optimization of large-scale models for LLM, multimodal, and generative AI applications. * Guide engineers in adopting best practices for CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM). * Represent the team in roadmap and planning discussions, ensuring alignment with NVIDIA's broader AI and software strategies. * Foster a culture of technical excellence, open collaboration, and continuous innovation. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)