Senior Software Engineer, CUTLASS Platform

NVIDIA Ltd.
Austin, TX, United States
3 months ago
Apply on indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$152,000.0 - $287,500.0
Working hours
Regular working hours
Job source

Tech stack

C++ (Programming Language) Code Generation Nvidia CUDA Computer Programming Computer Engineering Software Debugging Python (Programming Language) Software Engineering Graphics Processing Unit (GPU) Deep Learning Parallel Computation Backend
+2 more
Information Technology Software Coding

Job description

  • Develop core components of the CUTLASS platform including Tensor Core MMAs, copies, synchronization barriers, schedulers, and other GPU hardware features in CUDA C++ and CUTLASS Python DSL.
  • Contribute to the advancement of the MLIR-based backend compiler stack for the CUTLASS Python DSL by designing dialects and associated compiler passes.
  • Author example kernels utilizing CUTLASS abstractions to showcase the use of novel GPU hardware features that are crucial for achieving high performance.
  • Collaborate with GPU architecture, CUDA, and NVVM/PTX compiler teams to provide feedback on programming models and to assess the performance of future GPU hardware features.

Requirements

Do you have experience in Software coding?, Do you have a Master’s degree?, If you are passionate about designing abstractions for Tensor Core and related GPU hardware features in MLIR, Python, and C++ that enable writing high performance kernels, apply to join the CUTLASS team today!, * Masters or PhD degree in Computer Science, Computer Engineering, or related field (or equivalent experience).

  • 3+ years of relevant industry experience.
  • Strong proficiency in C++ programming and software design, including debugging, performance evaluation, and testing.
  • Experience working with high-performance code generation and knowledge of compiler transformations and optimizations.
  • A deep understanding of computer architecture and parallel computing programming models.

Ways to stand out from the crowd:

  • Experience writing high-performance kernels at low levels of abstractions like NVVM/ PTX for GPUs or other similar parallel processing architectures.
  • Hands-on compiler design experience, particularly in MLIR.
  • Understanding of deep learning models, algorithms, and frameworks.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

About the company

NVIDIA’s high-performance computing platforms are powering the AI revolution across many applications and industries. Within our software stack, CUTLASS stands out as a popular open-source ecosystem dedicated to high-performance linear algebra and Tensor Core primitives. Since 2017, it has provided the community with C++ and Python abstractions to implement custom matrix multiply (GEMM) and related math and deep learning computations on NVIDIA GPUs., NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hard working people in the world working for us. If you’re creative, autonomous, and love a challenge, consider joining our Deep Learning Library team and help us build the real-time, cost-effective computing platform driving our success in this exciting and quickly growing field.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

1:25 min

Distinguishing artificial intelligence from deep learning

Sam Witteveen · Coffee With Developers

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann +2 · LIVE

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

3:30 min

Transitioning from CUDA software architect to user

Stephen Jones · Coffee With Developers

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all