> Markdown version of [/jobs/ext/1239911-gpu-kernel-engineer](https://www.wearedevelopers.com/jobs/ext/1239911-gpu-kernel-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # GPU Kernel Engineer - **Company:** Typesafe AI Inc. - **Location:** San Francisco, CA, United States - **Salary:** $180,000.0 - $280,000.0 - **Contract:** Permanent contract - **Skills:** Nvidia CUDA, Large Language Models - **Published:** July 11, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a74c1dbffe57d090 ## About the Role * Have deep CUDA / GPU kernel expertise and a track record of real performance wins * Have built and optimized inference / training kernels * Have hands-on LLM training experience (real, not at a hobbyist level) * Reason from first principles about performance, memory, and parallelism * Are responsible, ownership-inclined team players who are mission aligned and excited to go all-in ## Description We're looking for a GPU kernel engineer with deep, low-level CUDA expertise to make our training and inference faster and more efficient. You'll write and optimize custom kernels, profile and eliminate bottlenecks, and work close to the metal across our model stack., * Write, optimize, and maintain high-performance GPU kernels (e.g., in CUDA / CuTe DSL) for training and inference * Profile end-to-end performance and eliminate bottlenecks across the stack * Partner with research and platform engineers to squeeze maximum throughput and minimum latency out of our hardware ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Creating Industry ready solutions with LLM Models](https://www.wearedevelopers.com/videos/899-creating-industry-ready-solutions-with-llm-models) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Bringing the power of AI to your application.](https://www.wearedevelopers.com/videos/1010-bringing-the-power-of-ai-to-your-application) - [Lies, Damned Lies and Large Language Models](https://www.wearedevelopers.com/videos/1231-lies-damned-lies-and-large-language-models) ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [What’s the latest in NVIDIA CUDA Python](https://www.wearedevelopers.com/magazine/568-what-s-the-latest-in-nvidia-cuda-python) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)