> Markdown version of [/jobs/ext/1792762-senior-ml-infrastructure-engineer-inference](https://www.wearedevelopers.com/jobs/ext/1792762-senior-ml-infrastructure-engineer-inference). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior ML Infrastructure Engineer, Inference... - **Company:** General Motors - **Location:** Warren, MI, United States (Remote available) - **Experience:** Expert - **Salary:** $155,420.0 - $205,900.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Big Data, C++ (Programming Language), Data Mining, Distributed Systems, Python (Programming Language), AI Infrastructure, Graphics Processing Unit (GPU), Backend, Hardware Acceleration, Machine Learning Operations, Software Version Control, Programming Languages - **Published:** July 9, 2026 - **Apply:** https://www.juju.com/job/00000000geslbv ## About the Role + 5+ years of industry experience, with focus on machine learning systems or high performance backend services. + Expertise in either Python, C++ or other relevant coding languages. + Expertise in ML inference, model serving frameworks (triton, rayserve, vLLM etc). + Strong communication skills and a proven ability to drive cross-functional initiatives. + Ability to thrive in a dynamic, multi-tasking environment with ever-evolving priorities. Preferred Qualifications + Deep expertise building zero-to-one ML infrastructure platforms. + Experience working with or designing interfaces, apis and clients for ML workflows. + Experience with Ray framework, and/or vLLM. + Experience with distributed systems, and handling large-scale data processing. + Familiarity with telemetry, and other feedback loops to inform product improvements. + Familiarity with hardware acceleration (GPUs) and optimizations for inference workloads. ## Description We are seeking a Senior ML Infrastructure engineer to help build and scale robust platforms for ML Inference workflows. In this role, you'll work closely with ML engineers and researchers to ensure efficient model serving and inference in production, for workflows such as data mining, labeling, model distillation, evaluations, simulations and more. This is a high-impact opportunity to influence the future of AI infrastructure at GM. You will play a key role in shaping the architecture, roadmap and user-experience of a robust ML inference service supporting real-time, batch, and experimental inference needs. The ideal candidate brings experience in designing distributed systems for ML, strong problem-solving skills, and a product mindset focused on platform usability and reliability. What you'll be doing: + Design and implement core platform backend software components. + Collaborate with ML engineers and researchers to understand critical workflows, parse them to platform requirements, and deliver incremental value. + Lead technical decision-making on model serving strategies, orchestration, caching, model versioning, and auto-scaling mechanisms for highly optimized use of accelerators. + Drive the development of monitoring, observability, and metrics to ensure reliability, performance, and resource optimization of inference services. + Proactively research and integrate state-of-the-art model serving frameworks, hardware accelerators, and distributed computing techniques. + Lead technical initiatives across GM's ML ecosystem. + Raise the engineering bar through technical leadership, establishing best practices. + Contribute to open source projects; represent GM in relevant communities., _Remote/Hybrid: This role is based remotely but if you live within a 50-mile radius of Mountain View, you are expected to report to that location three times a week, at minimum._ ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) - [How Machine Learning is turning the Automotive Industry upside down](https://www.wearedevelopers.com/videos/61-how-machine-learning-is-turning-the-automotive-industry-upside-down) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology)