Engineering Manager, Deep Learning Inference

NVIDIA Ltd.
Aiken, TX, United States
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Agile Methodology Artificial Intelligence C++ (Programming Language) Profiling Collaborative Software Nvidia CUDA Python (Programming Language) Performance Tuning Software Engineering Graphics Processing Unit (GPU) Large Language Models Deep Learning
+2 more
Gpu Programming Machine Learning Operations

Job description

Experteer Overview As Manager, Deep Learning Inference Software, you lead a world-class team advancing AI model deployment on NVIDIA GPUs. You shape and execute the inference software strategy, partnering across compiler, libraries, and research groups to optimize end-to-end pipelines. You drive performance tuning for large-scale models and guide adoption of CUDA, Triton, CUTLASS, and multi-GPU techniques. This role blends technical leadership with strategic roadmap responsibilities, impacting real-time inference at datacenter and edge scales. Compensation / Benefits * Lead and mentor a high-performing engineering team focused on deep learning inference and GPU-accelerated software. * Set strategy, roadmap, and delivery for NVIDIA’s inference frameworks engineering, with emphasis on Client AI. * Collaborate with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines across NVIDIA accelerators. * Oversee performance tuning, profiling, and optimization for large-scale LLM, multimodal, and generative AI workloads. * Guide engineers on CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM). * Represent the team in roadmap and planning discussions to align with broader AI and software strategies. * Foster a culture of technical excellence, collaboration, and continuous innovation. Tasks * 6+ years of software development experience; 3+ years in technical leadership or engineering management. * Strong background in C/C++ software design; Python is a plus. * Hands-on GPU programming experience (CUDA, Triton, CUTLASS) and performance optimization. * Proven record of deploying or optimizing deep learning models in production environments. * Experience leading teams using Agile or collaborative software development practices. Key requirements * equity * comprehensive benefits package * base salary and variable compensation * hybrid/remote options * career advancement opportunities

Requirements

world-class for large-scale LLM, multimodal, and generative AI workloads. * Guide engineers on CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM). * Represent the team in roadmap and planning discussions to align with broader AI and software strategies. * Foster a culture of technical excellence, collaboration, and continuous innovation. Tasks * 6+ years of software development experience; 3+ years in technical leadership or engineering management. * Strong background in C/C++ software design; Python is a plus. * Hands-on GPU programming experience (CUDA, Triton, CUTLASS) and performance optimization. * Proven record of deploying or optimizing deep learning models in production environments. * Experience leading teams using Agile or collaborative software development practices. Key requirements * equity * comprehensive benefits package * base salary and variable compensation * hybrid/remote options * career advancement opportunities

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

1:25 min

Distinguishing artificial intelligence from deep learning

Sam Witteveen · Coffee With Developers

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann +2 · LIVE

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · WWC Europe 2026

3:30 min

Transitioning from CUDA software architect to user

Stephen Jones · Coffee With Developers

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 · WWC Europe 2026

Videos

See all

Related articles

See all