Kernel Engineer (Compute / Accelerator)

Density AI, Inc.
Mountain View, CA, United States
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$260,000.0 - $320,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence C++ (Programming Language) Program Optimization Profiling Nvidia CUDA Extract Transform Load (ETL) Memory Management Field-Programmable Gate Array (FPGA) Python (Programming Language) Reduced Instruction Set Computing Scientific Computating SystemVerilog
+1 more
Verilog

Job description

You will write, evaluate, and profile specialized compute kernels that run on a custom AI accelerator. This is the critical interface between high-level ML workloads and silicon - your code directly determines how effectively the hardware performs. You’ll work closely with the architecture and compiler teams to define the kernel programming model, implement core tensor operations, and drive the performance profiling workflow that validates silicon design decisions. What you’ll do

  • Write and optimize compute kernels for a custom AI accelerator - tensor operations, data movement patterns, memory hierarchy exploitation
  • Develop and maintain profiling infrastructure to measure kernel performance against architectural targets
  • Define and document shuffle patterns for ML kernel primitives across CPU-like control, tensor cores, and CUTLASS-style operations
  • Drive kernel DSL design decisions - thread spawn mechanisms, register passing conventions, and memory management strategies
  • Enable end-to-end kernel execution on the architectural simulator
  • Collaborate with the compiler team on the MLIR dialect - your kernels are the primary validation target
  • Create onboarding documentation and kernel writing guides for the broader team

Requirements

  • C/C++ - production-grade systems code, not scripted glue. You’ll write performance-critical kernels
  • CUDA or equivalent accelerator programming - deep experience writing GPU kernels, understanding warp/wavefront execution, memory coalescing, shared memory optimization. The mental model transfers directly
  • Computer architecture - you need to reason about pipelines, memory hierarchies, data movement costs, and how software maps to hardware
  • Performance profiling and optimization - you live in profilers. Identifying bottlenecks, measuring throughput, and iterating until kernels meet targets is the core loop
  • Tensor operations - practical understanding of GEMM, convolution, attention, reduction, and scatter/gather as they map to hardware
  • Python - for scripting, DSL integration, and profiling automation
  • (Optional) RISC-V, x86, or ARM64 ISA experience
  • (Optional) MLIR or LLVM compiler infrastructure
  • (Optional) HPC or scientific computing background (large-scale parallel compute intuition)
  • (Optional) FPGA or Verilog/SystemVerilog (ability to read RTL and reason about the hardware you’re targeting)
  • (Optional) Familiarity with CUTLASS, Triton, or similar kernel libraries

Benefits & conditions

Final offers depend on level, location, and skills relevant to the role. Additional compensation: equity grant per company guidelines; medical / dental / vision; 401(k); standard PTO. Visa Sponsorship

DensityAI sponsors qualified candidates for H-1B, O-1, TN, E-3, and other employment-based visas, and we welcome applicants on F-1 OPT and STEM-OPT. Work authorization is required at start; we provide immigration support to secure or transfer status. Equal Opportunity

DensityAI is an Equal Opportunity Employer. We do not discriminate on the basis of race, color, religious creed, national origin, ancestry, physical or mental disability, medical condition, genetic information, marital status, sex, gender, gender identity, gender expression, age (40+), sexual orientation, military or veteran status, pregnancy, or any other status protected by law. We comply with the California CROWN Act and provide reasonable accommodations on request.

Full compensation packages are based on candidate experience and relevant certifications.

California pay range

$260,000 - $320,000 USD

Compensation

Final offers depend on level, location, and skills relevant to the role. Additional compensation: equity grant per company guidelines; medical / dental / vision; 401(k); standard PTO. Visa Sponsorship

DensityAI sponsors qualified candidates for H-1B, O-1, TN, E-3, and other employment-based visas, and we welcome applicants on F-1 OPT and STEM-OPT. Work authorization is required at start; we provide immigration support to secure or transfer status. Export Controls

Aspects of this role may involve access to information subject to U.S. export controls (EAR/ITAR). We may discuss licensing or scope adjustments during the interview. Equal Opportunity

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · WWC 2024

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · WWC Europe 2026

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann +2 · LIVE

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 · WWC Europe 2026

3:30 min

Transitioning from CUDA software architect to user

Stephen Jones · Coffee With Developers

1:46 min

Navigating the CUDA ecosystem and abstraction layers

Paul Graham Paul Graham · WWC 2024

Videos

See all

Related articles

See all