Sr. Software Engineer - AI Triton Kernels

Advanced Micro Devices, Inc.
San Jose, CA, United States
4 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Artificial Intelligence Computer Engineering Software Debugging Linux Kernel Open Source Technology Graphics Processing Unit (GPU) Pytorch Large Language Models Backend Information Technology

Job description

Experteer Overview As a Triton Kernel Engineer, you design and optimize high-performance GPU kernels for AI workloads on AMD Instinct hardware. You collaborate with research, compiler, and hardware teams to push Triton performance on AMD backends. You tackle bottlenecks across memory, scheduling, and ISA-level tuning to boost throughput. You will contribute to open-source Triton and ROCm ecosystems, shaping the AI software stack for scale and impact. Compensation / Benefits * Design, research, implement, and optimize high-performance matmul, attention, MoE, and fully fused transformer kernels using Triton for large-scale LLM and multimodal workloads * Own and productionize critical Triton/Gluon kernels within vLLM and SGL (e.g., paged attention, extend attention, MoE, quantized kernels) ensuring correctness, scalability, and peak throughput * Partner with compiler engineers to develop and maintain the Triton AMD backend across ROCm and the LLVM AMDGPU stack for CDNA and future architectures * Drive deep kernel-level optimizations across memory hierarchy (LDS, L2, HBM), wavefront execution, vectorization, MFMA utilization, occupancy, and instruction scheduling to maximize hardware efficiency * Perform profiling and microbenchmarking-led optimization on AMD Instinct GPUs using hardware counters and tracing tools; root-cause bottlenecks in memory bandwidth, latency hiding, synchronization, and register pressure * Debug and resolve performance and correctness issues end-to-end across PyTorch, vLLM/SGL runtimes, Triton IR/MLIR, ROCm runtime, and the LLVM AMDGPU backend * Contribute to open-source Triton, LLVM, and ROCm ecosystems Tasks * Deep experience in GPU kernel development, compiler backends, or performance engineering focused on AI/ML workloads * Strong hands-on expertise with Triton, including writing custom matmul, attention, and fused transformer kernels and understanding Triton IR lowering to GPU backends * Deep understanding of modern GPU architectures (wavefront execution, memory hierarchy, scheduling, occupancy) * Meaningful contributions to open-source projects such as Triton, Torch, vLLM, SGLang, MLIR, LLVM, or ROCm, with a collaborative and upstream-first engineering mindset * Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience Key requirements *

Requirements

push * Drive deep kernel-level optimizations across memory hierarchy (LDS, L2, HBM), wavefront execution, vectorization, MFMA utilization, occupancy, and instruction scheduling to maximize hardware efficiency * Perform profiling and microbenchmarking-led optimization on AMD Instinct GPUs using hardware counters and tracing tools; root-cause bottlenecks in memory bandwidth, latency hiding, synchronization, and register pressure * Debug and resolve performance and correctness issues end-to-end across PyTorch, vLLM/SGL runtimes, Triton IR/MLIR, ROCm runtime, and the LLVM AMDGPU backend * Contribute to open-source Triton, LLVM, and ROCm ecosystems Tasks * Deep experience in GPU kernel development, compiler backends, or performance engineering focused on AI/ML workloads * Strong hands-on expertise with Triton, including writing custom matmul, attention, and fused transformer kernels and understanding Triton IR lowering to GPU backends * Deep understanding of modern GPU architectures (wavefront aaa Triton, memory hierarchy, scheduling, occupancy) * Meaningful contributions to open-source projects such as Triton, Torch, vLLM, SGLang, MLIR, LLVM, or ROCm, with a collaborative and upstream-first engineering mindset * Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience Key requirements *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

1:15 min

Overcoming the challenges of modifying Linux kernel code

Ayesha Kaleem · WWC 2023

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · WWC Europe 2026

4:20 min

Utilizing AI and hardware acceleration for application code optimization

Stephan Gillich Stephan Gillich · WWC 2024

Videos

See all

Related articles

See all