Senior Deep Learning Performance Architect - LPU

NVIDIA Corporation
United States
2 days ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$152,000.0 - $241,500.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence C++ (Programming Language) Compilers Profiling Nvidia CUDA Computer Programming General-Purpose Computing on Graphics Processing Units Python (Programming Language) Machine Learning OpenMP Software Architecture Systems Architecture
+4 more
Graphics Processing Unit (GPU) Application Specific Integrated Circuits Large Language Models Deep Learning

Job description

  • Design novel GPU and system architectures to advance the forefront of AI Inference performance and efficiency
  • Construct, investigate, and test popular deep learning algorithms and applications
  • Understand and analyze the relationship between hardware and software architectures as it influences future algorithms and applications
  • Build efficient power and performance models of AI inference stack, while capturing minimal but significant information to guide next-gen HW architecture
  • Collaborate across the company to guide the direction of AI, working with software, research, and product teams

Requirements

NVIDIA seeks a Senior DL Performance Architect to join our group of pioneers who enjoy pushing AI Inference performance boundaries. Our team focuses on ambitious hardware-software co-design to speed AI Inference workloads. This role gives an outstanding opportunity to develop world-class performance strategies, guide future GPU architecture decisions, and lead AI innovation. If you are passionate about AI efficiency Pareto curves, have a proven record of modeling LLM performance and architecting AI systems, and enjoy optimizing every cycle, this role may be perfect for you!, * A MS or PhD in a relevant field (CS, EE, Math) or equivalent experience, with 5+ years of relevant experience

  • Strong mathematical foundation in machine learning and deep learning
  • Expert programming skills in C, C++, and/or Python
  • Familiarity with GPU computing (CUDA or similar) and HPC (MPI, OpenMP) stack
  • Strong knowledge and coursework in computer architecture

Ways to stand out from the crowd:

  • Background with systems-level performance modeling, profiling, and analysis
  • Experience in characterizing and modeling system-level performance, accomplishing comparison studies, and documenting and publishing results
  • Background in improving AI Inference workloads by developing CUDA kernels or compilers for custom ASIC hardware

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · World Congress 2026 Europe

1:51 min

Evolution of custom compilers and virtual machines

Florian Rappl · LIVE

4:37 min

Simplifying parallel programming with the CUDA ecosystem

Paul Graham Paul Graham · LIVE

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

2:17 min

Comparing code profiling with surface level monitoring

Jérôme Vieilledent · LIVE

3:04 min

Clarifying compiled versus interpreted language implementations

Aleksandra Sikora · World Congress 2023

Videos

See all

Related articles

See all