> Markdown version of [/jobs/ext/1990454-sr-software-engineer-ai-triton-kernels](https://www.wearedevelopers.com/jobs/ext/1990454-sr-software-engineer-ai-triton-kernels). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Software Engineer - AI Triton Kernels - **Company:** Advanced Micro Devices, Inc. - **Location:** San Jose, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computer Engineering, Software Debugging, Linux Kernel, Open Source Technology, Graphics Processing Unit (GPU), Pytorch, Large Language Models, Backend, Information Technology - **Published:** August 8, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/sr-software-engineer-ai-triton-kernels-san-jose-ca-usa-58853902 ## About the Role push * Drive deep kernel-level optimizations across memory hierarchy (LDS, L2, HBM), wavefront execution, vectorization, MFMA utilization, occupancy, and instruction scheduling to maximize hardware efficiency * Perform profiling and microbenchmarking-led optimization on AMD Instinct GPUs using hardware counters and tracing tools; root-cause bottlenecks in memory bandwidth, latency hiding, synchronization, and register pressure * Debug and resolve performance and correctness issues end-to-end across PyTorch, vLLM/SGL runtimes, Triton IR/MLIR, ROCm runtime, and the LLVM AMDGPU backend * Contribute to open-source Triton, LLVM, and ROCm ecosystems Tasks * Deep experience in GPU kernel development, compiler backends, or performance engineering focused on AI/ML workloads * Strong hands-on expertise with Triton, including writing custom matmul, attention, and fused transformer kernels and understanding Triton IR lowering to GPU backends * Deep understanding of modern GPU architectures (wavefront aaa Triton, memory hierarchy, scheduling, occupancy) * Meaningful contributions to open-source projects such as Triton, Torch, vLLM, SGLang, MLIR, LLVM, or ROCm, with a collaborative and upstream-first engineering mindset * Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience Key requirements * ## Description Experteer Overview As a Triton Kernel Engineer, you design and optimize high-performance GPU kernels for AI workloads on AMD Instinct hardware. You collaborate with research, compiler, and hardware teams to push Triton performance on AMD backends. You tackle bottlenecks across memory, scheduling, and ISA-level tuning to boost throughput. You will contribute to open-source Triton and ROCm ecosystems, shaping the AI software stack for scale and impact. Compensation / Benefits * Design, research, implement, and optimize high-performance matmul, attention, MoE, and fully fused transformer kernels using Triton for large-scale LLM and multimodal workloads * Own and productionize critical Triton/Gluon kernels within vLLM and SGL (e.g., paged attention, extend attention, MoE, quantized kernels) ensuring correctness, scalability, and peak throughput * Partner with compiler engineers to develop and maintain the Triton AMD backend across ROCm and the LLVM AMDGPU stack for CDNA and future architectures * Drive deep kernel-level optimizations across memory hierarchy (LDS, L2, HBM), wavefront execution, vectorization, MFMA utilization, occupancy, and instruction scheduling to maximize hardware efficiency * Perform profiling and microbenchmarking-led optimization on AMD Instinct GPUs using hardware counters and tracing tools; root-cause bottlenecks in memory bandwidth, latency hiding, synchronization, and register pressure * Debug and resolve performance and correctness issues end-to-end across PyTorch, vLLM/SGL runtimes, Triton IR/MLIR, ROCm runtime, and the LLVM AMDGPU backend * Contribute to open-source Triton, LLVM, and ROCm ecosystems Tasks * Deep experience in GPU kernel development, compiler backends, or performance engineering focused on AI/ML workloads * Strong hands-on expertise with Triton, including writing custom matmul, attention, and fused transformer kernels and understanding Triton IR lowering to GPU backends * Deep understanding of modern GPU architectures (wavefront execution, memory hierarchy, scheduling, occupancy) * Meaningful contributions to open-source projects such as Triton, Torch, vLLM, SGLang, MLIR, LLVM, or ROCm, with a collaborative and upstream-first engineering mindset * Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience Key requirements * ## Related Videos - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [From Model to Metal: An Open Source Stack for Accelerating Intelligence](https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data)