> Markdown version of [/jobs/ext/1429463-staff-ml-engineer-ml-compute-platform](https://www.wearedevelopers.com/jobs/ext/1429463-staff-ml-engineer-ml-compute-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff ML Engineer, ML Compute Platform - **Company:** General Motors - **Location:** Warren, MI, United States - **Experience:** Expert - **Salary:** $195,000.0 - $298,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Computing Platforms, Microsoft Azure, C++ (Programming Language), Software Debugging, Distributed Systems, Python (Programming Language), Machine Learning, Open Source Technology, AI Infrastructure, Google Cloud, Pytorch, Backend, Kubernetes, Free and Open-Source Software, Machine Learning Operations, Golang, Programming Languages - **Published:** July 24, 2026 - **Apply:** https://dejobs.org/x/x/83D1A6E415954561BC714F2D5A8254E5/job/ ## About the Role * 8+ years of industry experience * Expertise in either Go, C++, Python or other relevant coding languages * Strong background with kubernetes at scale * Relevant experience building large-scale with distributed systems * Experience leading and driving large scale initiatives * Experience working with Google Cloud Platform, Microsoft Azure, or Amazon Web Services, * Hands-on experience building ML infrastructure platforms with strong developer/user experience * Experience working with or designing job orchestration interfaces, CLI tools, or web UIs for ML workflows * Familiarity with observability, telemetry, and user feedback loops to inform product improvements * Experience with GPU/TPU optimizations * Experience with training frameworks like PyTorch, TorchX * Experience with Ray framework * Leadership/active participation in the open source community * Experience infrastructure applications or similar experience ## Description The ML Compute Platform is part of the AI Compute Platform organization within Infrastructure Platforms. Our team owns the cloud-agnostic, reliable, and cost-efficient compute backend that powers GM AI. We're proud to serve as the AI infrastructure platform for teams developing autonomous vehicles (L3/L4/L5), as well as other groups building AI-driven products for GM and its customers. We enable rapid innovation and feature development by optimizing for high-priority, ML-centric use cases. Our platform supports the training and deployment of state-of-the-art (SOTA) machine learning models with a focus on performance, availability, concurrency, and scalability. We're committed to maximizing GPU utilization across platforms (B200, H100, A100, and more) while maintaining reliability and cost efficiency., We are seeking a Staff ML Engineer to help build and scale robust compute platforms for ML workflows. In this role, you'll work closely with ML engineers and researchers to ensure efficient model training and seamless deployment into production. This is a high-impact opportunity to influence the future of AI infrastructure at GM. You will play a key role in shaping the user-facing experience of the platform, ensuring that ML practitioners can discover, schedule, and debug jobs with ease. The ideal candidate brings experience in designing distributed systems for ML, strong problem-solving skills, and a product mindset focused on platform usability and reliability. What you'll be doing: * Design and implement core platform backend software components * Experience cloud platforms like GCP, Azure or on-prem * Collaborate with ML engineers and researchers to understand platform pain points and improve developer experience * Thrive in a dynamic, multi-tasking environment with ever-evolving priorities. Interface with other teams to incorporate their innovations and vice versa * Analyze and improve efficiency, scalability, and stability of various system resources * Lead large-scale technical initiatives across GM's ML ecosystem * Help raise the engineering bar through technical leadership and best practices * Contribute to and potentially lead open source projects; represent GM in relevant communities ## Related Videos - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)