> Markdown version of [/jobs/ext/3244996-gpu-kernel-engineer-cuda-triton-accelerator-performance](https://www.wearedevelopers.com/jobs/ext/3244996-gpu-kernel-engineer-cuda-triton-accelerator-performance). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # GPU Kernel Engineer - CUDA, Triton & Accelerator Performance - **Company:** Anyone Ai - **Location:** Madrid, Spain (Remote available) - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Software Code Optimization, Profiling, Nvidia CUDA, Software Debugging, Performance Tuning - **Published:** September 15, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=c81767550c9d0484 ## About the Role * 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels * Strong experience with at least two of the following: + CUDA + Triton + NKI / AWS Neuron + Pallas / JAX * Strong understanding of GPU performance optimization * Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers * Understanding of: + Memory bandwidth + Compute throughput + GPU occupancy + Shared memory + Register pressure + Memory coalescing + Bank conflicts * Strong understanding of floating-point numerical correctness and tolerance thresholds * Experience debugging kernel compilation and runtime issues * Ability to distinguish software defects, environment problems, and genuine optimization challenges Relevant Experience Candidates should have experience with several of the following types of work: * Writing kernels from technical specifications * Translating kernels between CUDA, Triton, or other frameworks * Migrating kernels across hardware platforms * Debugging incorrect kernel implementations * Optimizing kernel performance * Fusing multiple operations into optimized kernels Nice to Have * Experience across both NVIDIA GPU and custom accelerator ecosystems * Experience with AWS Trainium, TPU, JAX, or other accelerators * Compiler engineering experience * Familiarity with MLIR, XLA, or intermediate representation lowering * Contributions to GPU or ML kernel libraries * Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls * Experience with AI model evaluation, RLHF, or technical benchmark development ## Description We're looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking. What You'll Work On You'll work with GPU and accelerator kernel tasks involving: * Kernel implementation and debugging * CUDA and Triton optimization * Translation between kernel frameworks * Hardware migration * Operator fusion * Performance profiling and benchmarking * Numerical correctness verification * Compilation and runtime debugging * Memory hierarchy optimization * Kernel-level AI workload performance You'll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware., * Reviewing GPU and accelerator kernel implementations for correctness * Comparing outputs against reference implementations * Evaluating numerical tolerance thresholds * Reviewing kernel benchmarks and determining whether comparisons are fair * Identifying performance bottlenecks and optimization opportunities * Assessing whether performance targets are realistic given hardware limits * Reviewing kernel translations and hardware migrations * Identifying compilation, driver, memory, shape, and runtime issues * Determining whether technical tasks are genuinely difficult or incorrectly configured * Providing clear, actionable technical feedback ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Using AI Without Losing Your Skills](https://www.wearedevelopers.com/videos/2045-using-ai-without-losing-your-skills) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/1112-accelerating-python-on-gpus) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/1521-accelerating-python-on-gpus) ## Related Articles - [What’s the latest in NVIDIA CUDA Python](https://www.wearedevelopers.com/magazine/568-what-s-the-latest-in-nvidia-cuda-python) - [Dev Digest 157: CUDA in Python, Gemini Code Assist and Back-dooring LLMs](https://www.wearedevelopers.com/magazine/557-dev-digest-157-cuda-in-python-gemini-code-assist-and-back-dooring-llms) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023)