Senior Performance Compiler Engineer - Triton

NVIDIA Ltd.
Santa Clara, United States of America
15 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 288K

Job location

Santa Clara, United States of America

Tech stack

Artificial Intelligence
Algorithm Design
Basic Linear Algebra Subprograms
C++
Nvidia CUDA
Computer Programming
Computer Engineering
Software Debugging
Python
Machine Learning
OpenMP
Open Source Technology
OpenCL
Software Engineering
High Performance Computing
Deep Learning
Parallel Computation
Gpu Programming
Information Technology

Job description

  • Investigating the latest and future NVIDIA GPU hardware architecture and programming models.
  • Working on the frontier of AI by understanding advanced algorithms (like attention sinks and MoEs) and numerics (like block-scaled floating point) to identify new opportunities for optimization.
  • Designing and implementing compiler technology using MLIR to optimize high-level kernel descriptions (written in Triton's Python DSL), with a focus on generating efficient, low-level GPU code. When vital, you'll also be able to use inline PTX to hand-tune critical code paths and extract peak performance from the hardware.
  • Engaging in a dynamic, iterative process of optimization-sometimes starting with the kernel, sometimes with the compiler-to find the most efficient path to peak performance.
  • Collaborating with teams across NVIDIA, including hardware architects and the CUDA compiler team, to influence future products and ensure we are always operating at maximum efficiency.

Requirements

Do you have experience in Software design?, Do you have a Master's degree?, * Bachelor, Masters or Ph.D. degree or equivalent experience in Computer Science, Computer Engineering, Applied Math, or a related field.

  • 8+ years of relevant industry experience in software development.
  • Demonstrated strong C++ programming and software design skills, with an emphasis on performance analysis and debugging.
  • Experienced in parallel programming, including CUDA/OpenCL GPU programming or other parallel models such as OpenMP.
  • Solid understanding of computer architecture and hands-on experience with assembly-level programming.

Ways to stand out from the crowd:

  • Experience in tuning BLAS or deep learning library kernels.
  • Background in numerics and linear algebra.
  • Experience with machine learning compilers like TVM or MLIR.
  • Contributions to open-source projects, especially in the AI/ML or compiler space.
  • Familiarity with the latest research in AI algorithms and numerics as well as a strong track record of contributions to open-source projects, particularly in the AI/ML, compiler, or high-performance computing domains.

Benefits & conditions

With competitive salaries and a generous benefits package, we are widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to unprecedented growth, our exclusive engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD.

You will also be eligible for equity and benefits.

About the company

NVIDIA's invention of the GPU 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI - the next era of computing - with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as "the AI computing company".

Apply for this position