GPU Kernel Engineer - CUDA, Triton & Accelerator Performance
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking.
What You’ll Work On
You’ll work with GPU and accelerator kernel tasks involving:
- Kernel implementation and debugging
- CUDA and Triton optimization
- Translation between kernel frameworks
- Hardware migration
- Operator fusion
- Performance profiling and benchmarking
- Numerical correctness verification
- Compilation and runtime debugging
- Memory hierarchy optimization
- Kernel-level AI workload performance
You’ll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware., * Reviewing GPU and accelerator kernel implementations for correctness
- Comparing outputs against reference implementations
- Evaluating numerical tolerance thresholds
- Reviewing kernel benchmarks and determining whether comparisons are fair
- Identifying performance bottlenecks and optimization opportunities
- Assessing whether performance targets are realistic given hardware limits
- Reviewing kernel translations and hardware migrations
- Identifying compilation, driver, memory, shape, and runtime issues
- Determining whether technical tasks are genuinely difficult or incorrectly configured
- Providing clear, actionable technical feedback
Requirements
- 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels
- Strong experience with at least two of the following:
- CUDA
- Triton
- NKI / AWS Neuron
- Pallas / JAX
- Strong understanding of GPU performance optimization
- Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers
- Understanding of:
- Memory bandwidth
- Compute throughput
- GPU occupancy
- Shared memory
- Register pressure
- Memory coalescing
- Bank conflicts
- Strong understanding of floating-point numerical correctness and tolerance thresholds
- Experience debugging kernel compilation and runtime issues
- Ability to distinguish software defects, environment problems, and genuine optimization challenges
Relevant Experience
Candidates should have experience with several of the following types of work:
- Writing kernels from technical specifications
- Translating kernels between CUDA, Triton, or other frameworks
- Migrating kernels across hardware platforms
- Debugging incorrect kernel implementations
- Optimizing kernel performance
- Fusing multiple operations into optimized kernels
Nice to Have
- Experience across both NVIDIA GPU and custom accelerator ecosystems
- Experience with AWS Trainium, TPU, JAX, or other accelerators
- Compiler engineering experience
- Familiarity with MLIR, XLA, or intermediate representation lowering
- Contributions to GPU or ML kernel libraries
- Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls
- Experience with AI model evaluation, RLHF, or technical benchmark development
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 157: CUDA in Python, Gemini Code Assist and Back-dooring LLMs
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
MLOps And AI Driven Development
Dev Digest 121 - AI goes offline