performance engineer

VLLM LLC
San Francisco, CA, United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$200,000.0
Working hours
Regular working hours
Job source

Tech stack

C++ (Programming Language) Profiling Nvidia CUDA Python (Programming Language) Performance Tuning Graphics Processing Unit (GPU) Information Technology

Job description

We’re looking for a performance engineer to squeeze every FLOP out of modern accelerators. You’ll write the kernels and low-level optimizations that make vLLM the fastest inference engine in the world. Your code will run on hundreds of accelerator types, from NVIDIA GPUs to emerging silicon. When hardware vendors develop new chips, they integrate with vLLM. You’ll work directly with these teams to ensure we’re extracting maximum performance from every generation of hardware.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • Deep experience writing CUDA kernels or equivalent (CuTeDSL, Triton, TileLang, Pallas).
  • Strong understanding of GPU architecture: memory hierarchy, warp scheduling, tiling, tensor cores.
  • Proficiency in C++ and Python with demonstrated ability to write high-performance code.
  • Experience with profiling tools (Nsight, rocprof) and performance optimization methodologies.
  • Obsession with benchmarks and squeezing every percentage point of speedup.

Preferred qualifications:

  • Experience with ML-specific kernel optimization (FlashAttention, fused kernels).
  • Knowledge of quantization techniques (INT8, FP8, mixed-precision).
  • Familiarity with multiple accelerator platforms (NVIDIA, AMD, TPU, Intel).
  • Experience with compiler technologies (LLVM, MLIR, XLA).

Benefits & conditions

  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
  • Visa sponsorship: We sponsor visas on a case-by-case basis.
  • Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.

About the company

Inferact’s mission is to grow vLLM as the world’s AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware-a position that took years to build.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Profiling and debugging GPU code with Nsight developer tools

Paul Graham Paul Graham · World Congress 2025

8:32 min

Benchmarking GitOps engine constraints for extensive multi-cluster environments

Artem Lajko · Europe 2026 Virtual

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann +2 · LIVE

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · World Congress 2026 Europe

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 · World Congress 2026 Europe

1:31 min

Profiling computing workloads across environments using Performance Studio

Andrew Wafaa Andrew Wafaa · World Congress 2024

Videos

See all

Related articles

See all