Senior ML Accelerator Engineer - GPU

General Motors
San Francisco, CA, United States
22 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Compensation
$170,100.0 - $258,300.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence C++ (Programming Language) Compilers Profiling Code Review Nvidia CUDA Computer Programming Software Debugging Software Design Patterns Hardware-In-The-Loop Simulation Linux Kernel Machine Learning
+12 more
Software Architecture Software Requirements Analysis Systems Integration Graphics Processing Unit (GPU) High Performance Computing Real Time Systems IT Architecture Parallel Computation Gpu Programming Backend Low Latency GPT

Job description

  • Designing and implementing custom operators when vendor libraries hit their limits
  • Integrating those kernels deep into our ML runtime stack
  • Debugging and tuning GPU performance across the AV software stack, often on hardware-in-the-loop
  • We partner closely with AI Solutions, AI Compilers, AI Architecture, and AI Tooling to ensure models deploy efficiently to the car while consistently meeting strict latency, throughput, and reliability targets. If you enjoy pushing GPUs to their limits and seeing your work directly impact how autonomous vehicles perceive and act in the world, this is the team for you.

What you’ll be doing (Responsibilities)

  • Design, implement, benchmark, and iterate on CUDA-based kernels and custom operators to squeeze every last drop of performance out of on-vehicle inference workloads.
  • Build and improve tooling and infrastructure that make it easier to profile, debug, and validate CUDA kernels and accelerator-backend code across the AV stack.
  • Partner with AI Solutions, Compilers, and Architecture to translate model and system requirements into concrete kernel roadmaps, priorities, and project plans.
  • Collaborate with cross-functional teams (compiler, performance tooling, runtime, deployment solutions) to deliver reusable, reliable, high-performance libraries into production.
  • Maintain high technology standards, methodologies, processes, and guidelines for GPU kernel development and performance engineering through code review.
  • Manage relationships with internal customers to ensure our kernels and libraries meet real-world needs, This role is categorized as hybrid. This means the selected candidate is expected to report to a specific location at least 3 times a week {or other frequency dictated by their manager}.

Requirements

  • Minimum 2+ years of relevant industry experience or equivalent experience
  • BS, MS or PhD in CS, or related technical field
  • Excellent GPU programming skills in CUDA, with a thorough understanding of parallel programming patterns and GPU architecture.
  • Hands-on experience benchmarking, profiling, debugging and optimizing accelerator libraries and kernels to extract optimal performance using the NSight suite of tools or similar.
  • Strong background in software architecture, library design, and design patterns.
  • Strong C++ programming skills with the ability to feel comfortable in large codebases.
  • Solid background in system performance, high performance computing and/or architecture-aware optimizations.
  • Strong communication skills and the ability to work collaboratively within a team
  • Excellent analytical and problem-solving skills

What Will Give You A Competitive Edge (Preferred Qualifications)

  • 2+ years of relevant industry experience or equivalent experience
  • Experience with tensor core programming, CUTLASS and/or CuTe
  • Experience with ML model architectures, in particular transformer-based
  • Experience with low latency or real time systems
  • Experience with lower levels of an accelerator software stack (i.e. drivers, runtimes, and compilers)

Benefits & conditions

Compensation: The compensation information is a good faith estimate only. It is based on what a successful applicant might be paid in accordance with applicable state laws. The compensation may not be representative for positions located outside of New York, Colorado, California, or Washington.

  • The salary range for this role: is $170,100 to $258,300 . The actual base salary a successful candidate will be offered within this range will vary based on factors relevant to the position.
  • Bonus Potential: An incentive pay program offers payouts based on company performance, job level, and individual performance.
  • Benefits: GM offers a variety of health and wellbeing benefit programs. Benefit options include medical, dental, vision, Health Savings Account, Flexible Spending Accounts, retirement savings plan, sickness and accident benefits, life insurance, paid vacation & holidays, tuition assistance programs, employee assistance program, GM vehicle discounts and more.

About the company

The AI Kernels team builds high-performance GPU kernels and custom libraries that sit at the heart of our on-vehicle ML inference for ADAS and autonomous driving . We own making core AI workloads faster, more reliable, and easier to maintain and deploy on real cars, under real-world constraints., We believe we all must make a choice every day - individually and collectively - to drive meaningful change through our words, our deeds and our culture. Every day, we want every employee to feel they belong to one General Motors team., General Motors is committed to being a workplace that is not only free of unlawful discrimination, but one that genuinely fosters inclusion and belonging. We strongly believe that providing an inclusive workplace creates an environment in which our employees can thrive and develop better products for our customers.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:20 min

Utilizing AI and hardware acceleration for application code optimization

Stephan Gillich Stephan Gillich · World Congress 2024

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · World Congress 2024

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

1:51 min

Evolution of custom compilers and virtual machines

Florian Rappl · LIVE

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 · World Congress 2026 Europe

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky · World Congress 2025

Videos

See all

Related articles

See all