Engineering Manager, Deep Learning Inference
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+2 more
Job description
Experteer Overview As Manager, Deep Learning Inference Software, you lead a world-class team advancing AI model deployment on NVIDIA GPUs. You shape and execute the inference software strategy, partnering across compiler, libraries, and research groups to optimize end-to-end pipelines. You drive performance tuning for large-scale models and guide adoption of CUDA, Triton, CUTLASS, and multi-GPU techniques. This role blends technical leadership with strategic roadmap responsibilities, impacting real-time inference at datacenter and edge scales. Compensation / Benefits * Lead and mentor a high-performing engineering team focused on deep learning inference and GPU-accelerated software. * Set strategy, roadmap, and delivery for NVIDIA’s inference frameworks engineering, with emphasis on Client AI. * Collaborate with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines across NVIDIA accelerators. * Oversee performance tuning, profiling, and optimization for large-scale LLM, multimodal, and generative AI workloads. * Guide engineers on CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM). * Represent the team in roadmap and planning discussions to align with broader AI and software strategies. * Foster a culture of technical excellence, collaboration, and continuous innovation. Tasks * 6+ years of software development experience; 3+ years in technical leadership or engineering management. * Strong background in C/C++ software design; Python is a plus. * Hands-on GPU programming experience (CUDA, Triton, CUTLASS) and performance optimization. * Proven record of deploying or optimizing deep learning models in production environments. * Experience leading teams using Agile or collaborative software development practices. Key requirements * equity * comprehensive benefits package * base salary and variable compensation * hybrid/remote options * career advancement opportunities
Requirements
world-class for large-scale LLM, multimodal, and generative AI workloads. * Guide engineers on CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM). * Represent the team in roadmap and planning discussions to align with broader AI and software strategies. * Foster a culture of technical excellence, collaboration, and continuous innovation. Tasks * 6+ years of software development experience; 3+ years in technical leadership or engineering management. * Strong background in C/C++ software design; Python is a plus. * Hands-on GPU programming experience (CUDA, Triton, CUTLASS) and performance optimization. * Proven record of deploying or optimizing deep learning models in production environments. * Experience leading teams using Agile or collaborative software development practices. Key requirements * equity * comprehensive benefits package * base salary and variable compensation * hybrid/remote options * career advancement opportunities
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
MLOps And AI Driven Development
What Are Large Language Models?
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud