Triton Compiler and Kernel Software Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+11 more
Job description
We are seeking a Senior Triton Compiler and Kernel Engineer to advance Triton performance and capabilities on AMD GPUs.
Triton is an open-source language and compiler for developing high-performance GPU kernels in Python. It is a critical layer in the AI software stack, connecting frameworks and workloads to GPU hardware. Triton is strategic to AMD’s AI roadmap, and AMD is investing fully in making it a first-class platform for current and future AMD GPUs.
You will work across GPU architecture, compilers, kernels, multi-GPU communication, and AI frameworks while contributing to upstream Triton and AMD’s ROCm software stack., * Develop and optimize Triton compiler support for AMD GPUs.
- Improve compiler lowering, optimization, scheduling, and code generation.
- Create high-performance kernels for attention, GEMM, MoE, and other AI workloads.
- Develop and optimize multi-GPU kernels and communication primitives.
- Enable new AMD GPU and interconnect capabilities through effective Triton abstractions.
- Analyze compute, memory, communication, occupancy, register usage, and generated code.
- Resolve complex correctness and performance issues across single- and multi-GPU workloads.
- Collaborate with GPU architecture, ROCm, PyTorch, and AI framework teams.
- Contribute designs and implementations to upstream Triton and LLVM/MLIR.
- Provide technical leadership, code reviews, and mentoring., AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.
Requirements
The ideal candidate has strong experience in several of the following areas:
- GPU architecture and programming
- Compiler development and optimization
- High-performance GPU kernels
- Multi-GPU communication and collective operations
- Distributed AI training and inference
- AI workload and framework performance
- Low-level performance analysis
You can reason across the stack-from distributed AI algorithms and Triton programs to compiler transformations, communication libraries, generated instructions, and GPU hardware., * Experience with AMD GPU architecture and ROCm is highly desirable.
- Experience optimizing kernels with Triton, HIP, CUDA, or GPU assembly.
- Experience with Triton, LLVM, MLIR, or another optimizing compiler.
- Knowledge of GPU execution models, memory hierarchies, synchronization, and instruction pipelines.
- Experience with collective communication, distributed programming, and libraries such as RCCL or NCCL.
- Understanding of communication topologies, interconnects, synchronization, and communication-computation overlap.
- Familiarity with distributed training, inference, tensor parallelism, expert parallelism, or pipeline parallelism.
- Experience using AI coding tools and autonomous agents to accelerate software development, debugging, benchmarking, and kernel performance tuning.
- Familiarity with AI primitives, reduced-precision formats, and performance profiling.
- Contributions to complex or open-source software projects.
- Strong analytical, debugging, communication, and collaboration skills., * Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent
Benefits & conditions
$204,000.00/Yr.-$306,000.00/Yr.
About the company
At AMD, we believetechnology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.
Whetheryou’redesigning next-gen processors, enabling AI breakthroughs, orbringing leading edge products to market, every role at AMD contributes to something bigger- technologythat moves the world forward.Join us and, together, we’ll advance your career.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 157: CUDA in Python, Gemini Code Assist and Back-dooring LLMs
Stephan Gillich - Bringing AI Everywhere
Dev Digest 120 - Apple and peers
Dev Digest 121 - AI goes offline