> Markdown version of [/jobs/ext/1958153-engineering-manager-deep-learning-inference](https://www.wearedevelopers.com/jobs/ext/1958153-engineering-manager-deep-learning-inference). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Engineering Manager, Deep Learning Inference - **Company:** NVIDIA Ltd. - **Location:** New York, NY, United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Agile Methodology, Artificial Intelligence, C++ (Programming Language), Profiling, Collaborative Software, Nvidia CUDA, Python (Programming Language), Performance Tuning, Software Engineering, Graphics Processing Unit (GPU), Large Language Models, Deep Learning, Gpu Programming, Machine Learning Operations - **Published:** August 6, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/engineering-manager-deep-learning-inference-new-york-usa-58815172 ## About the Role world-class for large-scale LLM, multimodal, and generative AI workloads. * Guide engineers on CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM). * Represent the team in roadmap and planning discussions to align with broader AI and software strategies. * Foster a culture of technical excellence, collaboration, and continuous innovation. Tasks * 6+ years of software development experience; 3+ years in technical leadership or engineering management. * Strong background in C/C++ software design; Python is a plus. * Hands-on GPU programming experience (CUDA, Triton, CUTLASS) and performance optimization. * Proven record of deploying or optimizing deep learning models in production environments. * Experience leading teams using Agile or collaborative software development practices. Key requirements * equity * comprehensive benefits package * base salary and variable compensation * hybrid/remote options * career advancement opportunities ## Description Experteer Overview As Manager, Deep Learning Inference Software, you lead a world-class team advancing AI model deployment on NVIDIA GPUs. You shape and execute the inference software strategy, partnering across compiler, libraries, and research groups to optimize end-to-end pipelines. You drive performance tuning for large-scale models and guide adoption of CUDA, Triton, CUTLASS, and multi-GPU techniques. This role blends technical leadership with strategic roadmap responsibilities, impacting real-time inference at datacenter and edge scales. Compensation / Benefits * Lead and mentor a high-performing engineering team focused on deep learning inference and GPU-accelerated software. * Set strategy, roadmap, and delivery for NVIDIA's inference frameworks engineering, with emphasis on Client AI. * Collaborate with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines across NVIDIA accelerators. * Oversee performance tuning, profiling, and optimization for large-scale LLM, multimodal, and generative AI workloads. * Guide engineers on CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM). * Represent the team in roadmap and planning discussions to align with broader AI and software strategies. * Foster a culture of technical excellence, collaboration, and continuous innovation. Tasks * 6+ years of software development experience; 3+ years in technical leadership or engineering management. * Strong background in C/C++ software design; Python is a plus. * Hands-on GPU programming experience (CUDA, Triton, CUTLASS) and performance optimization. * Proven record of deploying or optimizing deep learning models in production environments. * Experience leading teams using Agile or collaborative software development practices. Key requirements * equity * comprehensive benefits package * base salary and variable compensation * hybrid/remote options * career advancement opportunities ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [30 Golden Rules of Deep Learning Performance](https://www.wearedevelopers.com/videos/11-30-golden-rules-of-deep-learning-performance) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)