Senior Deep Learning Performance Architect - LPU
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+4 more
Job description
- Design novel GPU and system architectures to advance the forefront of AI Inference performance and efficiency
- Construct, investigate, and test popular deep learning algorithms and applications
- Understand and analyze the relationship between hardware and software architectures as it influences future algorithms and applications
- Build efficient power and performance models of AI inference stack, while capturing minimal but significant information to guide next-gen HW architecture
- Collaborate across the company to guide the direction of AI, working with software, research, and product teams
Requirements
NVIDIA seeks a Senior DL Performance Architect to join our group of pioneers who enjoy pushing AI Inference performance boundaries. Our team focuses on ambitious hardware-software co-design to speed AI Inference workloads. This role gives an outstanding opportunity to develop world-class performance strategies, guide future GPU architecture decisions, and lead AI innovation. If you are passionate about AI efficiency Pareto curves, have a proven record of modeling LLM performance and architecting AI systems, and enjoy optimizing every cycle, this role may be perfect for you!, * A MS or PhD in a relevant field (CS, EE, Math) or equivalent experience, with 5+ years of relevant experience
- Strong mathematical foundation in machine learning and deep learning
- Expert programming skills in C, C++, and/or Python
- Familiarity with GPU computing (CUDA or similar) and HPC (MPI, OpenMP) stack
- Strong knowledge and coursework in computer architecture
Ways to stand out from the crowd:
- Background with systems-level performance modeling, profiling, and analysis
- Experience in characterizing and modeling system-level performance, accomplishing comparison studies, and documenting and publishing results
- Background in improving AI Inference workloads by developing CUDA kernels or compilers for custom ASIC hardware
Benefits & conditions
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
How to Become an AI Engineer
Stephan Gillich - Bringing AI Everywhere
MLOps – What’s the deal behind it?