Senior ML Inference Engineer - Platform

General Motors
Olympia, WA, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$128,700.0 - $261,300.0
Working hours
Regular working hours
Job source

Tech stack

Airflow Nvidia CUDA Cursor (Graphical User Interface Elements) Programming Tools Python (Programming Language) Machine Learning Workflow Management Systems Real Time Systems GitHub Copilot Pytorch Large Language Models Kubernetes
+6 more
Information Technology Low Latency ONNX (Open Neural Network Exchange) Format Free and Open-Source Software Machine Learning Operations TensorRT

Job description

About the Team

The Model Deployment & Inference Solutions team in GM AV deploys machine learning models from training frameworks (e.g. PyTorch) onto autonomous vehicle hardware. Our mission is two-fold: build the ML deployment platform that makes model rollouts fast and predictable, and optimize models so they meet the real-time latency and memory budgets required to run on-vehicle. Our work is on the critical path of GM’s publicly committed launch of eyes-off (hands-free, eyes-free) autonomous driving in 2028, debuting on the Cadillac Escalade IQ, building on Super Cruise’s billion-plus hands-free miles.

About the Role

This role sits in the team’s Platform pillar. We own the unified ML deployment platform that automates the path from a trained model to inference on the vehicle, along with the developer-experience and agentic-tooling layer that makes deployment self-serve for every ML model development team at GM.

What** you’ll **be doing (Responsibilities)

  • Design, build, andoperatethe ML deployment platform that automates the path from trained model to on-vehicle inference.

  • Drive cross-organization model deployments to the autonomous vehicle stack, partnering with model development teams to take high-value models from training to production on-vehicle.

  • Build agentic tools that diagnose and fix deployment-blocking issues, automating workflows currently performed manually by engineers.

  • Build the developer experience that ML model development teams use day to day: tooling, dashboards, automation, and observability.

  • Drive shift-left validation that surfaces deployment risk (compile, runtime, parity, latency) early in the model development cycle.

  • Build platform tools that integrate the work of our sister teams (kernels, compiler, reducedprecisionand parity) so their optimization wins land directly in the deployment workflow.

  • Partner with the team’s Performance pillar and model development teams across the AV organization., This role is based remotely, but if the selected candidate lives within a specific mile radius of a GM hub, they will be expected to report to the location three times a week {or other frequency dictated by your manager}.

Requirements

  • BS, MS, or PhD in Computer Science or a related technical field.

  • 3+ years of relevant industry experience.

  • Strong fundamentals and excellent coding ability in Python.

  • Experience building or operating production platform or infrastructure systems where reliability, observability, and extensibility matter.

  • Experience with ML model deployment, inference integration, model optimization workflows, or model serving infrastructure, with at least one prior context where you owned the path from a trained model to a running inference workload.

  • Experience using coding agents (Cursor, Claude Code, GitHub Copilot, or equivalent) as part of your engineering workflow.

  • Experience designing clean, well-tested software with clear interfaces and good abstractions.

  • Strong cross-team collaboration skills.

What Will Give You** A **Competitive Edge (Preferred Qualifications)

  • Experience building agentic or LLM-powered developer tooling.

  • Experience with ML or workflow orchestration frameworks (Airflow, Temporal, Flyte, Ray, Kubeflow, or equivalent).

  • Familiarity with the NVIDIA GPU stack at the integration level (CUDA-aware Python,TensorRT, Triton inference server,torch.compile, ONNX).

  • Experience with inference-serving frameworks (Triton,TorchServe, Ray Serve,vLLM) or edge-deployment toolchains.

  • Experience with low-latency or real-time systems.

  • Experience in autonomous vehicles, robotics, or other safety-critical ML deployment domains.

  • Open-source contributions toPyTorch, Ray, Airflow, Temporal,vLLM,TensorRT, or related projects.

  • 3+ years of relevant industry experience.

Compensation: The compensation information is a good faith estimate only. It is based on what a successful applicant might be paid in accordance with applicable state laws. The compensation may not be representative for positions located outside of New York, Colorado, California, or Washington.

Benefits & conditions

  • The salary range for this role: is $128,700 to $261,300. The actual base salary a successful candidate will be offered within this range will vary based on factors relevant to the position.

  • Bonus Potential: An incentive pay program offers payouts based on company performance, job level, and individual performance.

  • Benefits: GM offers a variety of health and wellbeing benefit programs. Benefit options include medical, dental, vision, Health Savings Account, Flexible Spending Accounts, retirement savings plan, sickness and accident benefits, life insurance, paid vacation & holidays, tuition assistance programs, employee assistance program, GM vehicle discounts and more.

About the company

We believe we all must make a choice every day - individually and collectively - to drive meaningful change through our words, our deeds and our culture. Every day, we want every employee to feel they belong to one General Motors team., General Motors is committed to being a workplace that is not only free of unlawful discrimination, but one that genuinely fosters inclusion and belonging. We strongly believe that providing an inclusive workplace creates an environment in which our employees can thrive and develop better products for our customers.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · WWC 2025

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · WWC Europe 2026

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · WWC 2024

Videos

See all

Related articles

See all