Sr. Software Engineer - AI Triton Kernels
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Experteer Overview As a Triton Kernel Engineer, you design and optimize high-performance GPU kernels for AI workloads on AMD Instinct hardware. You collaborate with research, compiler, and hardware teams to push Triton performance on AMD backends. You tackle bottlenecks across memory, scheduling, and ISA-level tuning to boost throughput. You will contribute to open-source Triton and ROCm ecosystems, shaping the AI software stack for scale and impact. Compensation / Benefits * Design, research, implement, and optimize high-performance matmul, attention, MoE, and fully fused transformer kernels using Triton for large-scale LLM and multimodal workloads * Own and productionize critical Triton/Gluon kernels within vLLM and SGL (e.g., paged attention, extend attention, MoE, quantized kernels) ensuring correctness, scalability, and peak throughput * Partner with compiler engineers to develop and maintain the Triton AMD backend across ROCm and the LLVM AMDGPU stack for CDNA and future architectures * Drive deep kernel-level optimizations across memory hierarchy (LDS, L2, HBM), wavefront execution, vectorization, MFMA utilization, occupancy, and instruction scheduling to maximize hardware efficiency * Perform profiling and microbenchmarking-led optimization on AMD Instinct GPUs using hardware counters and tracing tools; root-cause bottlenecks in memory bandwidth, latency hiding, synchronization, and register pressure * Debug and resolve performance and correctness issues end-to-end across PyTorch, vLLM/SGL runtimes, Triton IR/MLIR, ROCm runtime, and the LLVM AMDGPU backend * Contribute to open-source Triton, LLVM, and ROCm ecosystems Tasks * Deep experience in GPU kernel development, compiler backends, or performance engineering focused on AI/ML workloads * Strong hands-on expertise with Triton, including writing custom matmul, attention, and fused transformer kernels and understanding Triton IR lowering to GPU backends * Deep understanding of modern GPU architectures (wavefront execution, memory hierarchy, scheduling, occupancy) * Meaningful contributions to open-source projects such as Triton, Torch, vLLM, SGLang, MLIR, LLVM, or ROCm, with a collaborative and upstream-first engineering mindset * Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience Key requirements *
Requirements
push * Drive deep kernel-level optimizations across memory hierarchy (LDS, L2, HBM), wavefront execution, vectorization, MFMA utilization, occupancy, and instruction scheduling to maximize hardware efficiency * Perform profiling and microbenchmarking-led optimization on AMD Instinct GPUs using hardware counters and tracing tools; root-cause bottlenecks in memory bandwidth, latency hiding, synchronization, and register pressure * Debug and resolve performance and correctness issues end-to-end across PyTorch, vLLM/SGL runtimes, Triton IR/MLIR, ROCm runtime, and the LLVM AMDGPU backend * Contribute to open-source Triton, LLVM, and ROCm ecosystems Tasks * Deep experience in GPU kernel development, compiler backends, or performance engineering focused on AI/ML workloads * Strong hands-on expertise with Triton, including writing custom matmul, attention, and fused transformer kernels and understanding Triton IR lowering to GPU backends * Deep understanding of modern GPU architectures (wavefront aaa Triton, memory hierarchy, scheduling, occupancy) * Meaningful contributions to open-source projects such as Triton, Torch, vLLM, SGLang, MLIR, LLVM, or ROCm, with a collaborative and upstream-first engineering mindset * Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience Key requirements *
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
Stephan Gillich - Bringing AI Everywhere
What Are Large Language Models?
Dev Digest 121 - AI goes offline