> Markdown version of [/jobs/ext/2649678-senior-ml-accelerator-engineer-gpu](https://www.wearedevelopers.com/jobs/ext/2649678-senior-ml-accelerator-engineer-gpu). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior ML Accelerator Engineer - GPU - **Company:** General Motors - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $170,100.0 - $258,300.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Compilers, Profiling, Code Review, Nvidia CUDA, Computer Programming, Software Debugging, Software Design Patterns, Hardware-In-The-Loop Simulation, Linux Kernel, Machine Learning, Software Architecture, Software Requirements Analysis, Systems Integration, Graphics Processing Unit (GPU), High Performance Computing, Real Time Systems, IT Architecture, Parallel Computation, Gpu Programming, Backend, Low Latency, GPT - **Published:** August 14, 2026 - **Apply:** https://dejobs.org/x/x/0F53D14ADC00469D93080C02D40B5A7B/job/ ## About the Role * Minimum 2+ years of relevant industry experience or equivalent experience * BS, MS or PhD in CS, or related technical field * Excellent GPU programming skills in CUDA, with a thorough understanding of parallel programming patterns and GPU architecture. * Hands-on experience benchmarking, profiling, debugging and optimizing accelerator libraries and kernels to extract optimal performance using the NSight suite of tools or similar. * Strong background in software architecture, library design, and design patterns. * Strong C++ programming skills with the ability to feel comfortable in large codebases. * Solid background in system performance, high performance computing and/or architecture-aware optimizations. * Strong communication skills and the ability to work collaboratively within a team * Excellent analytical and problem-solving skills What Will Give You A Competitive Edge (Preferred Qualifications) * 2+ years of relevant industry experience or equivalent experience * Experience with tensor core programming, CUTLASS and/or CuTe * Experience with ML model architectures, in particular transformer-based * Experience with low latency or real time systems * Experience with lower levels of an accelerator software stack (i.e. drivers, runtimes, and compilers) ## Description * Designing and implementing custom operators when vendor libraries hit their limits * Integrating those kernels deep into our ML runtime stack * Debugging and tuning GPU performance across the AV software stack, often on hardware-in-the-loop * We partner closely with AI Solutions, AI Compilers, AI Architecture, and AI Tooling to ensure models deploy efficiently to the car while consistently meeting strict latency, throughput, and reliability targets. If you enjoy pushing GPUs to their limits and seeing your work directly impact how autonomous vehicles perceive and act in the world, this is the team for you. What you'll be doing (Responsibilities) * Design, implement, benchmark, and iterate on CUDA-based kernels and custom operators to squeeze every last drop of performance out of on-vehicle inference workloads. * Build and improve tooling and infrastructure that make it easier to profile, debug, and validate CUDA kernels and accelerator-backend code across the AV stack. * Partner with AI Solutions, Compilers, and Architecture to translate model and system requirements into concrete kernel roadmaps, priorities, and project plans. * Collaborate with cross-functional teams (compiler, performance tooling, runtime, deployment solutions) to deliver reusable, reliable, high-performance libraries into production. * Maintain high technology standards, methodologies, processes, and guidelines for GPU kernel development and performance engineering through code review. * Manage relationships with internal customers to ensure our kernels and libraries meet real-world needs, This role is categorized as hybrid. This means the selected candidate is expected to report to a specific location at least 3 times a week {or other frequency dictated by their manager}. ## Related Videos - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Just-in-time Compilation in JVM](https://www.wearedevelopers.com/videos/240-just-in-time-compilation-in-jvm) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/1521-accelerating-python-on-gpus) - [Streaming AI Responses in Real-Time with SSE in Next.js & NestJS](https://www.wearedevelopers.com/videos/1630-streaming-ai-responses-in-real-time-with-sse-in-next-js-nestjs) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [The Prompt Engineer ✍️](https://www.wearedevelopers.com/magazine/216-the-prompt-engineer)