> Markdown version of [/jobs/ext/1411840-senior-ml-infrastructure-engineer-embodied-ai-scaling-foundations](https://www.wearedevelopers.com/jobs/ext/1411840-senior-ml-infrastructure-engineer-embodied-ai-scaling-foundations). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior ML Infrastructure Engineer - Embodied AI Scaling Foundations - **Company:** General Motors - **Location:** Sunnyvale, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $153,200.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, C++ (Programming Language), CMake, Profiling, Software Quality, Microprocessors, Distributed Systems, Python (Programming Language), Machine Learning, Tensorflow, Graphics Processing Unit (GPU), Cloud Platform System, Pytorch, Deep Learning, Kubernetes, Information Technology, Optimization Algorithms, Build Tools, Machine Learning Operations, Software Coding, Docker - **Published:** July 23, 2026 - **Apply:** https://dejobs.org/x/x/5AAA799F7ACA42758333655279FBCC68/job/ ## About the Role * 3+ years of experience building large-scale distributed systems/applications or advanced ML Applications. * Proven track record of building robust frameworks with high-quality, long-lasting APIs. * Deep understanding and practical experience with machine learning algorithms. * Expertise in building reliable, highly performant, and cost-efficient systems leveraging modern cloud infrastructure. * Hands-on experience with the entire ML development lifecycle and MLOps practices. * Demonstrated ability to collaborate effectively across multiple teams and organizations. * Proficiency working with containerization and orchestration technologies (Docker, Kubernetes). * A strong passion for self-driving technology and its transformative potential. * Exceptional coding skills in Python or C++. * BS, MS, or PhD in Computer Science, Math, or equivalent practical experience. Exceptional candidates may also have: * Experience with distributed training methodologies. * A background in optimizing model training performance. * Experience scaling model training across large clusters of GPUs/CPUs or other accelerators. * Familiarity with deep learning frameworks such as PyTorch, TensorFlow, etc. * A strong grasp of performance profiling and state-of-the-art training optimization algorithms, including their performance characteristics and effect on model convergence. * Experience with advanced build systems (Bazel, Buck, Blaze, or Cmake). Remote/Hybrid: This role is based remotely but if you live within a 50-mile radius of an office, you are expected to report to that location three times a week, at minimum. ## Description * Lead the design, implementation, and deployment of scalable platforms and tools that drive machine learning model training and evaluation workflows across GM. * Own complex technical projects end-to-end, making key architectural decisions and technical trade-offs. You will be a core contributor to team planning, design reviews, and code quality. * Take a holistic view of projects, considering their impact across multiple teams, and proactively drive technical prioritization. Collaborate closely with partner teams to ensure maximum benefit from the systems we build. * Help shape our team through technical interviewing with high, well-calibrated standards, and play an essential role in recruiting. Mentor and onboard junior engineers and interns, helping them grow their careers. ## Related Videos - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Code to Road in < 12 hours](https://www.wearedevelopers.com/videos/1082-code-to-road-in-12-hours) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology)