> Markdown version of [/jobs/ext/2865322-senior-manager-ai-deployment](https://www.wearedevelopers.com/jobs/ext/2865322-senior-manager-ai-deployment). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Manager, AI Deployment - **Company:** General Motors - **Location:** Denver, CO, United States (Remote available) - **Experience:** Expert - **Salary:** $296,300.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computer Vision, C++ (Programming Language), Program Optimization, Profiling, Nvidia CUDA, Computer Engineering, Extract Transform Load (ETL), Hardware-In-The-Loop Simulation, Python (Programming Language), Machine Learning, Software Deployment, Graphics Processing Unit (GPU), Pytorch, Information Technology, Machine Learning Operations, TensorRT - **Published:** September 12, 2026 - **Apply:** https://dejobs.org/x/x/C78C1ABB1751441DAFC51585B6818B67/job/ ## About the Role * Bachelor's degree in Computer Science, Electrical or Computer Engineering, Robotics, Machine Learning, or a related field; advanced degree preferred, or equivalent experience. * 10+ years of experience in machine learning systems, model optimization, inference, GPU systems, robotics, autonomous driving, or a related field. * 5+ years of people-leadership experience, including experience leading managers or senior technical leaders. * Experience shipping production machine-learning inference systems on GPU, accelerator, robotics, automotive, or other edge hardware. * Strong understanding of the factors that determine model performance: architecture, tensor shapes, operators, kernels, memory movement, scheduling, runtime execution, and hardware utilization. * Hands-on experience with several of the following: PyTorch, CUDA, C++, Python, TensorRT, GPU profiling, benchmarking, performance analysis, or inference runtimes. * Experience with quantization, pruning, distillation, architecture optimization, kernel optimization, or memory optimization. * Experience building benchmark automation, performance regression detection, telemetry, dashboards, or profiling workflows. * Strong systems thinking, communication, decision-making, and cross-functional leadership skills. What Will Give You a Competitive Edge * Experience optimizing real-time machine-learning systems for autonomous driving, robotics, embedded systems, or computer vision. * Deep experience with GPU performance, memory bandwidth, occupancy, synchronization, stream scheduling, or device-to-device data movement. * Experience with NVIDIA Nsight Systems, NVIDIA Nsight Compute, PyTorch Profiler, TensorRT profiling tools, or equivalent tools. * Experience deploying reduced-precision models and managing calibration, sensitivity, parity, and model-quality risks. * Experience optimizing transformer, vision, lidar, or multimodal workloads. * Experience measuring performance across simulation, hardware-in-the-loop, bench, and vehicle environments. * Experience with safety-critical or highly reliable systems. ## Description We are looking for a Senior Manager, AI Deployment to lead the strategy and execution of model performance and on-vehicle inference for autonomous driving. You will lead engineering managers and senior technical leaders working across model optimization, GPU systems, inference runtimes, and vehicle integration. You will set performance goals, guide optimization of complex autonomy models, and establish disciplined methods to measure latency, diagnose regressions, and validate improvements. Success requires strong technical judgment, people leadership, and the ability to make clear trade-offs among latency, memory, throughput, accuracy, power, and numerical parity. What You'll Do * Own the strategy, roadmap, and operating plan for AI model performance and inference quality. * Establish performance budgets for latency, throughput, memory, GPU utilization, power, and numerical parity. * Lead investigations into performance bottlenecks across model architecture, operators, kernels, memory movement, scheduling, runtime behavior, and hardware utilization. * Establish repeatable benchmarking and profiling practices across simulation, hardware-in-the-loop, bench, and vehicle environments. * Guide optimization through model architecture changes, operator and kernel improvements, memory optimization, scheduling, and hardware-aware execution. * Build performance dashboards, regression detection, benchmark automation, and root-cause diagnostics. * Partner with Embodied AI, model development, GPU kernel, runtime, system performance, vehicle integration, simulation, and safety teams. * Influence model design by translating profiling results into clear recommendations for model architects and researchers. * Represent AI Deployment in architecture reviews, program planning, and senior leadership discussions. Leadership Responsibilities * Build and lead an inclusive, high-performing organization through hiring, coaching, feedback, and manager development. * Establish clear ownership, priorities, staffing plans, and operating rhythms across performance workstreams. * Define and manage KPIs for inference latency, latency variability, throughput, memory efficiency, GPU utilization, parity, and regression rate. * Balance near-term production needs with longer-term investments in profiling, optimization automation, reduced precision, and performance infrastructure. * Resolve cross-functional issues and align stakeholders when performance, quality, or implementation trade-offs are contested. * Develop technical leaders and succession plans in GPU performance, model optimization, inference systems, and numerical analysis., This role is categorized as remote. This means the selected candidate may be based anywhere in the country of work and is not expected to report to a GM worksite unless directed by their manager. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [How Machine Learning is turning the Automotive Industry upside down](https://www.wearedevelopers.com/videos/61-how-machine-learning-is-turning-the-automotive-industry-upside-down) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)