Triton Compiler and Kernel Software Engineer

Advanced Micro Devices, Inc.
San Jose, CA, United States
4 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Compilers Code Generation Profiling Code Review Nvidia CUDA Computer Engineering Software Debugging Distributed Computing Environment Python (Programming Language) Open Source Technology Performance Tuning
+11 more
Software Engineering Graphics Processing Unit (GPU) Pytorch Multi-Agent Systems Distributed Programming Parallel Computation AI Coding Agents Information Technology Free and Open-Source Software ROCm MLIR (Multi-Level Intermediate Representation)

Job description

We are seeking a Senior Triton Compiler and Kernel Engineer to advance Triton performance and capabilities on AMD GPUs. Triton is an open-source language and compiler for developing high-performance GPU kernels in Python. It is a critical layer in the AI software stack, connecting frameworks and workloads to GPU hardware. Triton is strategic to AMD’s AI roadmap, and AMD is investing fully in making it a first-class platform for current and future AMD GPUs. You will work across GPU architecture, compilers, kernels, multi-GPU communication, and AI frameworks while contributing to upstream Triton and AMD’s ROCm software stack. THE PERSON: The ideal candidate has strong experience in several of the following areas:

  • GPU architecture and programming
  • Compiler development and optimization
  • High-performance GPU kernels
  • Multi-GPU communication and collective operations
  • Distributed AI training and inference
  • AI workload and framework performance
  • Low-level performance analysis

You can reason across the stack-from distributed AI algorithms and Triton programs to compiler transformations, communication libraries, generated instructions, and GPU hardware., * Develop and optimize Triton compiler support for AMD GPUs.

  • Improve compiler lowering, optimization, scheduling, and code generation.
  • Create high-performance kernels for attention, GEMM, MoE, and other AI workloads.
  • Develop and optimize multi-GPU kernels and communication primitives.
  • Enable new AMD GPU and interconnect capabilities through effective Triton abstractions.
  • Analyze compute, memory, communication, occupancy, register usage, and generated code.
  • Resolve complex correctness and performance issues across single- and multi-GPU workloads.
  • Collaborate with GPU architecture, ROCm, PyTorch, and AI framework teams.
  • Contribute designs and implementations to upstream Triton and LLVM/MLIR.
  • Provide technical leadership, code reviews, and mentoring.

Requirements

  • Experience with AMD GPU architecture and ROCm is highly desirable.
  • Experience optimizing kernels with Triton, HIP, CUDA, or GPU assembly.
  • Experience with Triton, LLVM, MLIR, or another optimizing compiler.
  • Knowledge of GPU execution models, memory hierarchies, synchronization, and instruction pipelines.
  • Experience with collective communication, distributed programming, and libraries such as RCCL or NCCL.
  • Understanding of communication topologies, interconnects, synchronization, and communication-computation overlap.
  • Familiarity with distributed training, inference, tensor parallelism, expert parallelism, or pipeline parallelism.
  • Experience using AI coding tools and autonomous agents to accelerate software development, debugging, benchmarking, and kernel performance tuning.
  • Familiarity with AI primitives, reduced-precision formats, and performance profiling.
  • Contributions to complex or open-source software projects.
  • Strong analytical, debugging, communication, and collaboration skills., * Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent

About the company

At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger - technology that moves the world forward. Join us and, together, we’ll advance your career., AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · World Congress 2024

1:51 min

Evolution of custom compilers and virtual machines

Florian Rappl · LIVE

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

3:00 min

Accelerating machine learning research with optimized compilers

Tanmay Bakshi · LIVE

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all