> Markdown version of [/jobs/ext/2713180-performance-engineer](https://www.wearedevelopers.com/jobs/ext/2713180-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # performance engineer - **Company:** VLLM LLC - **Location:** San Francisco, CA, United States - **Salary:** $200,000.0 - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Profiling, Nvidia CUDA, Python (Programming Language), Performance Tuning, Graphics Processing Unit (GPU), Information Technology - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/member-of-technical-staff-kernel-engineering-inferact-8115126 ## About the Role * Bachelor's degree or equivalent experience in computer science, engineering, or similar. * Deep experience writing CUDA kernels or equivalent (CuTeDSL, Triton, TileLang, Pallas). * Strong understanding of GPU architecture: memory hierarchy, warp scheduling, tiling, tensor cores. * Proficiency in C++ and Python with demonstrated ability to write high-performance code. * Experience with profiling tools (Nsight, rocprof) and performance optimization methodologies. * Obsession with benchmarks and squeezing every percentage point of speedup. Preferred qualifications: * Experience with ML-specific kernel optimization (FlashAttention, fused kernels). * Knowledge of quantization techniques (INT8, FP8, mixed-precision). * Familiarity with multiple accelerator platforms (NVIDIA, AMD, TPU, Intel). * Experience with compiler technologies (LLVM, MLIR, XLA). ## Description We're looking for a performance engineer to squeeze every FLOP out of modern accelerators. You'll write the kernels and low-level optimizations that make vLLM the fastest inference engine in the world. Your code will run on hundreds of accelerator types, from NVIDIA GPUs to emerging silicon. When hardware vendors develop new chips, they integrate with vLLM. You'll work directly with these teams to ensure we're extracting maximum performance from every generation of hardware. ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Practical performance tuning for Serverless Java on AWS](https://www.wearedevelopers.com/videos/2075-practical-performance-tuning-for-serverless-java-on-aws) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [30 Golden Rules of Deep Learning Performance](https://www.wearedevelopers.com/videos/11-30-golden-rules-of-deep-learning-performance) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Dev Digest 157: CUDA in Python, Gemini Code Assist and Back-dooring LLMs](https://www.wearedevelopers.com/magazine/557-dev-digest-157-cuda-in-python-gemini-code-assist-and-back-dooring-llms)