Senior Manager, AI Deployment

General Motors
Albany, NY, United States
2 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$296,300.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Computer Vision C++ (Programming Language) Program Optimization Profiling Nvidia CUDA Computer Engineering Extract Transform Load (ETL) Hardware-In-The-Loop Simulation Python (Programming Language) Machine Learning Software Deployment
+5 more
Graphics Processing Unit (GPU) Pytorch Information Technology Machine Learning Operations TensorRT

Job description

We are looking for a Senior Manager, AI Deployment to lead the strategy and execution of model performance and on-vehicle inference for autonomous driving. You will lead engineering managers and senior technical leaders working across model optimization, GPU systems, inference runtimes, and vehicle integration. You will set performance goals, guide optimization of complex autonomy models, and establish disciplined methods to measure latency, diagnose regressions, and validate improvements. Success requires strong technical judgment, people leadership, and the ability to make clear trade-offs among latency, memory, throughput, accuracy, power, and numerical parity.

What You’ll Do

  • Own the strategy, roadmap, and operating plan for AI model performance and inference quality.
  • Establish performance budgets for latency, throughput, memory, GPU utilization, power, and numerical parity.
  • Lead investigations into performance bottlenecks across model architecture, operators, kernels, memory movement, scheduling, runtime behavior, and hardware utilization.
  • Establish repeatable benchmarking and profiling practices across simulation, hardware-in-the-loop, bench, and vehicle environments.
  • Guide optimization through model architecture changes, operator and kernel improvements, memory optimization, scheduling, and hardware-aware execution.
  • Build performance dashboards, regression detection, benchmark automation, and root-cause diagnostics.
  • Partner with Embodied AI, model development, GPU kernel, runtime, system performance, vehicle integration, simulation, and safety teams.
  • Influence model design by translating profiling results into clear recommendations for model architects and researchers.
  • Represent AI Deployment in architecture reviews, program planning, and senior leadership discussions.

Leadership Responsibilities

  • Build and lead an inclusive, high-performing organization through hiring, coaching, feedback, and manager development.
  • Establish clear ownership, priorities, staffing plans, and operating rhythms across performance workstreams.
  • Define and manage KPIs for inference latency, latency variability, throughput, memory efficiency, GPU utilization, parity, and regression rate.
  • Balance near-term production needs with longer-term investments in profiling, optimization automation, reduced precision, and performance infrastructure.
  • Resolve cross-functional issues and align stakeholders when performance, quality, or implementation trade-offs are contested.
  • Develop technical leaders and succession plans in GPU performance, model optimization, inference systems, and numerical analysis., This role is categorized as remote. This means the selected candidate may be based anywhere in the country of work and is not expected to report to a GM worksite unless directed by their manager.

Requirements

  • Bachelor’s degree in Computer Science, Electrical or Computer Engineering, Robotics, Machine Learning, or a related field; advanced degree preferred, or equivalent experience.
  • 10+ years of experience in machine learning systems, model optimization, inference, GPU systems, robotics, autonomous driving, or a related field.
  • 5+ years of people-leadership experience, including experience leading managers or senior technical leaders.
  • Experience shipping production machine-learning inference systems on GPU, accelerator, robotics, automotive, or other edge hardware.
  • Strong understanding of the factors that determine model performance: architecture, tensor shapes, operators, kernels, memory movement, scheduling, runtime execution, and hardware utilization.
  • Hands-on experience with several of the following: PyTorch, CUDA, C++, Python, TensorRT, GPU profiling, benchmarking, performance analysis, or inference runtimes.
  • Experience with quantization, pruning, distillation, architecture optimization, kernel optimization, or memory optimization.
  • Experience building benchmark automation, performance regression detection, telemetry, dashboards, or profiling workflows.
  • Strong systems thinking, communication, decision-making, and cross-functional leadership skills.

What Will Give You a Competitive Edge

  • Experience optimizing real-time machine-learning systems for autonomous driving, robotics, embedded systems, or computer vision.
  • Deep experience with GPU performance, memory bandwidth, occupancy, synchronization, stream scheduling, or device-to-device data movement.
  • Experience with NVIDIA Nsight Systems, NVIDIA Nsight Compute, PyTorch Profiler, TensorRT profiling tools, or equivalent tools.
  • Experience deploying reduced-precision models and managing calibration, sensitivity, parity, and model-quality risks.
  • Experience optimizing transformer, vision, lidar, or multimodal workloads.
  • Experience measuring performance across simulation, hardware-in-the-loop, bench, and vehicle environments.
  • Experience with safety-critical or highly reliable systems.

Benefits & conditions

Compensation: The compensation information is a good faith estimate only. It is based on what a successful applicant might be paid in accordance with applicable state laws. The compensation may not be representative for positions located outside of New York, Colorado, California, or Washington

  • Compensation: The expected base compensation for this role is : $296,300 - $453,900 Actual base compensation within the identified range will vary based on factors relevant to the position.
  • Bonus Potential: An incentive pay program offers payouts based on company performance, job level, and individual performance.
  • Benefits: GM offers a variety of health and wellbeing benefit programs. Benefit options include medical, dental, vision, Health Savings Account, Flexible Spending Accounts, retirement savings plan, sickness and accident benefits, life insurance, paid vacation & holidays

About the company

General Motors is developing the software and artificial intelligence capabilities for the next generation of autonomous driving. Within AI Foundations, AI Acceleration makes machine learning models faster, more efficient, and more reliable on production vehicle hardware., We believe we all must make a choice every day - individually and collectively - to drive meaningful change through our words, our deeds and our culture. Every day, we want every employee to feel they belong to one General Motors team., General Motors is committed to being a workplace that is not only free of unlawful discrimination, but one that genuinely fosters inclusion and belonging. We strongly believe that providing an inclusive workplace creates an environment in which our employees can thrive and develop better products for our customers.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:28 min

Defining critical competencies for automotive AI engineering

Daniel Graff +1 · World Congress 2021

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · World Congress 2026 Europe

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

1:33 min

Summary of machine learning capabilities and engineering opportunities

Jan Zawadzki · LIVE

2:17 min

Comparing code profiling with surface level monitoring

Jérôme Vieilledent · LIVE

Videos

See all

Related articles

See all