> Markdown version of [/jobs/ext/2003778-kernel-engineer-compute-accelerator](https://www.wearedevelopers.com/jobs/ext/2003778-kernel-engineer-compute-accelerator). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Kernel Engineer (Compute / Accelerator) - **Company:** Density AI, Inc. - **Location:** Mountain View, CA, United States - **Salary:** $260,000.0 - $320,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Program Optimization, Profiling, Nvidia CUDA, Extract Transform Load (ETL), Memory Management, Field-Programmable Gate Array (FPGA), Python (Programming Language), Reduced Instruction Set Computing, Scientific Computating, SystemVerilog, Verilog - **Published:** August 9, 2026 - **Apply:** https://www.adzuna.com/details/5833062826 ## About the Role * C/C++ - production-grade systems code, not scripted glue. You'll write performance-critical kernels * CUDA or equivalent accelerator programming - deep experience writing GPU kernels, understanding warp/wavefront execution, memory coalescing, shared memory optimization. The mental model transfers directly * Computer architecture - you need to reason about pipelines, memory hierarchies, data movement costs, and how software maps to hardware * Performance profiling and optimization - you live in profilers. Identifying bottlenecks, measuring throughput, and iterating until kernels meet targets is the core loop * Tensor operations - practical understanding of GEMM, convolution, attention, reduction, and scatter/gather as they map to hardware * Python - for scripting, DSL integration, and profiling automation * (Optional) RISC-V, x86, or ARM64 ISA experience * (Optional) MLIR or LLVM compiler infrastructure * (Optional) HPC or scientific computing background (large-scale parallel compute intuition) * (Optional) FPGA or Verilog/SystemVerilog (ability to read RTL and reason about the hardware you're targeting) * (Optional) Familiarity with CUTLASS, Triton, or similar kernel libraries ## Description You will write, evaluate, and profile specialized compute kernels that run on a custom AI accelerator. This is the critical interface between high-level ML workloads and silicon - your code directly determines how effectively the hardware performs. You'll work closely with the architecture and compiler teams to define the kernel programming model, implement core tensor operations, and drive the performance profiling workflow that validates silicon design decisions. What you'll do * Write and optimize compute kernels for a custom AI accelerator - tensor operations, data movement patterns, memory hierarchy exploitation * Develop and maintain profiling infrastructure to measure kernel performance against architectural targets * Define and document shuffle patterns for ML kernel primitives across CPU-like control, tensor cores, and CUTLASS-style operations * Drive kernel DSL design decisions - thread spawn mechanisms, register passing conventions, and memory management strategies * Enable end-to-end kernel execution on the architectural simulator * Collaborate with the compiler team on the MLIR dialect - your kernels are the primary validation target * Create onboarding documentation and kernel writing guides for the broader team ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [CUDA Python: GPU programming for the modern developer](https://www.wearedevelopers.com/videos/100221-cuda-python-gpu-programming-for-the-modern-developer) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [The weekly developer show: Boosting Python with CUDA, CSS Updates & Navigating New Tech Stacks](https://www.wearedevelopers.com/videos/1293-the-weekly-developer-show-boosting-python-with-cuda-css-updates-navigating-new-tech-stacks) ## Related Articles - [What’s the latest in NVIDIA CUDA Python](https://www.wearedevelopers.com/magazine/568-what-s-the-latest-in-nvidia-cuda-python) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)