GPU Kernel Engineer - CUDA, Triton & Accelerator Performance

Anyone Ai
Madrid, Spain
20 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Part-time (≤ 32 hours)
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Software Code Optimization Profiling Nvidia CUDA Software Debugging Performance Tuning

Job description

We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking.

What You’ll Work On

You’ll work with GPU and accelerator kernel tasks involving:

  • Kernel implementation and debugging
  • CUDA and Triton optimization
  • Translation between kernel frameworks
  • Hardware migration
  • Operator fusion
  • Performance profiling and benchmarking
  • Numerical correctness verification
  • Compilation and runtime debugging
  • Memory hierarchy optimization
  • Kernel-level AI workload performance

You’ll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware., * Reviewing GPU and accelerator kernel implementations for correctness

  • Comparing outputs against reference implementations
  • Evaluating numerical tolerance thresholds
  • Reviewing kernel benchmarks and determining whether comparisons are fair
  • Identifying performance bottlenecks and optimization opportunities
  • Assessing whether performance targets are realistic given hardware limits
  • Reviewing kernel translations and hardware migrations
  • Identifying compilation, driver, memory, shape, and runtime issues
  • Determining whether technical tasks are genuinely difficult or incorrectly configured
  • Providing clear, actionable technical feedback

Requirements

  • 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels
  • Strong experience with at least two of the following:
  • CUDA
  • Triton
  • NKI / AWS Neuron
  • Pallas / JAX
  • Strong understanding of GPU performance optimization
  • Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers
  • Understanding of:
  • Memory bandwidth
  • Compute throughput
  • GPU occupancy
  • Shared memory
  • Register pressure
  • Memory coalescing
  • Bank conflicts
  • Strong understanding of floating-point numerical correctness and tolerance thresholds
  • Experience debugging kernel compilation and runtime issues
  • Ability to distinguish software defects, environment problems, and genuine optimization challenges

Relevant Experience

Candidates should have experience with several of the following types of work:

  • Writing kernels from technical specifications
  • Translating kernels between CUDA, Triton, or other frameworks
  • Migrating kernels across hardware platforms
  • Debugging incorrect kernel implementations
  • Optimizing kernel performance
  • Fusing multiple operations into optimized kernels

Nice to Have

  • Experience across both NVIDIA GPU and custom accelerator ecosystems
  • Experience with AWS Trainium, TPU, JAX, or other accelerators
  • Compiler engineering experience
  • Familiarity with MLIR, XLA, or intermediate representation lowering
  • Contributions to GPU or ML kernel libraries
  • Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls
  • Experience with AI model evaluation, RLHF, or technical benchmark development

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

2:08 min

The vending machine trap in software debugging

Jen Callou Jen Callou · Europe 2026 Virtual

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann Chris Heilmann +2 · LIVE

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · World Congress 2026 Europe

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 · World Congress 2026 Europe

3:30 min

Transitioning from CUDA software architect to user

Stephen Jones · Coffee With Developers

Videos

See all

Related articles

See all