> Markdown version of [/jobs/ext/177907-staff-software-engineer-genai-performance-and-kernel](https://www.wearedevelopers.com/jobs/ext/177907-staff-software-engineer-genai-performance-and-kernel). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer - GenAI Performance and Kernel - **Company:** Databricks - **Location:** San Francisco, CA, United States - **Salary:** $190,900.0 - $232,800.0 - **Contract:** Permanent contract - **Skills:** Systems Engineering, Profiling, Code Review, Nvidia CUDA, Software Debugging, Memory Management, Data Flow Control, Field-Programmable Gate Array (FPGA), Machine Learning, OpenCL, Performance Tuning, Systems Architecture, Backend, Perf (Linux), Information Technology, Optimization Algorithms, SAP Ariba, Machine Learning Operations - **Published:** May 31, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=41dbaf5648567ab3 ## About the Role Do you have experience in Tooling?, Do you have a Bachelor's degree?, * BS/MS/PhD in Computer Science, or a related field * Deep hands-on experience writing and tuning compute kernels (CUDA, Triton, OpenCL, LLVM IR, assembly or similar sort) for ML workloads * Strong knowledge of GPU/accelerator architecture: warp structure, memory hierarchy (global, shared, register, L1/L2 caches), tensor cores, scheduling, SM occupancy, etc. * Experience with advanced optimization techniques: tiling, blocking, software pipelining, vectorization, fusion, loop transformations, auto-tuning * Familiarity with ML-specific kernel libraries (cuBLAS, cuDNN, CUTLASS, oneDNN, etc.) or open kernels * Strong debugging and profiling skills (Nsight, NVProf, perf, vtune, custom instrumentation) * Experience reasoning about numerical stability, mixed precision, quantization, and error propagation * Experience in integrating optimized kernels into real-world ML inference systems; exposure to distributed inference pipelines, memory management, and runtime systems * Experience building high-performance products leveraging GPU acceleration * Excellent communication and leadership skills - able to drive design discussions, mentor colleagues, and make trade-offs visible * A track record of shipping performance-critical, high-quality production software * Bonus: published in systems/ML performance venues (e.g. MLSys, ASPLOS, ISCA, PPoPP), experience with custom accelerators or FPGA, experience with sparsity or model compression techniques Pay Range Transparency ## Description As a staff software engineer for GenAI Performance and Kernel, you will own the design, implementation, optimization, and correctness of the high-performance GPU kernels powering our GenAI inference stack. You will lead development of highly-tuned, low-level compute paths, manage trade-offs between hardware efficiency and generality, and mentor others in kernel-level performance engineering. You will work closely with ML researchers, systems engineers, and product teams to push the state-of-the-art in inference performance at scale., * Lead the design, implementation, benchmarking, and maintenance of core compute kernels (e.g. attention, MLP, softmax, layernorm, memory management) optimized for various hardware backends (GPU, accelerators) * Drive the performance roadmap for kernel-level improvements: vectorization, tensorization, tiling, fusion, mixed precision, sparsity, quantization, memory reuse, scheduling, auto-tuning, etc. * Integrate kernel optimizations with higher-level ML systems * Build and maintain profiling, instrumentation, and verification tooling to detect correctness, performance regressions, numerical issues, and hardware utilization gaps * Lead performance investigations and root-cause analysis on inference bottlenecks, e.g. memory bandwidth, cache contention, kernel launch overhead, tensor fragmentation * Establish coding patterns, abstractions, and frameworks to modularize kernels for reuse, cross-backend portability, and maintainability * Influence system architecture decisions to make kernel improvements more effective (e.g. memory layout, dataflow scheduling, kernel fusion boundaries) * Mentor and guide other engineers working on lower-level performance, provide code reviews, help set best practices * Collaborate with infrastructure, tooling, and ML teams to roll out kernel-level optimizations into production, and monitor their impact ## Related Videos - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Meet Your New BFF: Backend to Frontend without the Duct Tape](https://www.wearedevelopers.com/videos/682-meet-your-new-bff-backend-to-frontend-without-the-duct-tape) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Highest Paying Tech Companies in Europe](https://www.wearedevelopers.com/magazine/162-highest-paying-tech-companies-in-europe) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Software Developer Salary in Germany [2023]](https://www.wearedevelopers.com/magazine/194-software-developer-salary-in-germany-2023) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)