Staff ML Engineer, ML Compute Platform

General Motors
Mountain View, CA, United States
19 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$195,000.0 - $298,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Computing Platforms Microsoft Azure C++ (Programming Language) Software Debugging Distributed Systems Python (Programming Language) Machine Learning Open Source Technology AI Infrastructure Google Cloud
+7 more
Pytorch Backend Kubernetes Free and Open-Source Software Machine Learning Operations Golang Programming Languages

Job description

The ML Compute Platform is part of the AI Compute Platform organization within Infrastructure Platforms. Our team owns the cloud-agnostic, reliable, and cost-efficient compute backend that powers GM AI. We’re proud to serve as the AI infrastructure platform for teams developing autonomous vehicles (L3/L4/L5), as well as other groups building AI-driven products for GM and its customers. We enable rapid innovation and feature development by optimizing for high-priority, ML-centric use cases. Our platform supports the training and deployment of state-of-the-art (SOTA) machine learning models with a focus on performance, availability, concurrency, and scalability. We’re committed to maximizing GPU utilization across platforms (B200, H100, A100, and more) while maintaining reliability and cost efficiency., We are seeking a Staff ML Engineer to help build and scale robust compute platforms for ML workflows. In this role, you’ll work closely with ML engineers and researchers to ensure efficient model training and seamless deployment into production. This is a high-impact opportunity to influence the future of AI infrastructure at GM.

You will play a key role in shaping the user-facing experience of the platform, ensuring that ML practitioners can discover, schedule, and debug jobs with ease. The ideal candidate brings experience in designing distributed systems for ML, strong problem-solving skills, and a product mindset focused on platform usability and reliability.

What you’ll be doing:

  • Design and implement core platform backend software components
  • Experience cloud platforms like GCP, Azure or on-prem
  • Collaborate with ML engineers and researchers to understand platform pain points and improve developer experience
  • Thrive in a dynamic, multi-tasking environment with ever-evolving priorities. Interface with other teams to incorporate their innovations and vice versa
  • Analyze and improve efficiency, scalability, and stability of various system resources
  • Lead large-scale technical initiatives across GM’s ML ecosystem
  • Help raise the engineering bar through technical leadership and best practices
  • Contribute to and potentially lead open source projects; represent GM in relevant communities

Requirements

  • 8+ years of industry experience
  • Expertise in either Go, C++, Python or other relevant coding languages
  • Strong background with kubernetes at scale
  • Relevant experience building large-scale with distributed systems
  • Experience leading and driving large scale initiatives
  • Experience working with Google Cloud Platform, Microsoft Azure, or Amazon Web Services, * Hands-on experience building ML infrastructure platforms with strong developer/user experience
  • Experience working with or designing job orchestration interfaces, CLI tools, or web UIs for ML workflows
  • Familiarity with observability, telemetry, and user feedback loops to inform product improvements
  • Experience with GPU/TPU optimizations
  • Experience with training frameworks like PyTorch, TorchX
  • Experience with Ray framework
  • Leadership/active participation in the open source community
  • Experience infrastructure applications or similar experience

Benefits & conditions

If you’re excited to tackle some of today’s most complex engineering challenges, see the impact of your work in real-world AV applications, and help shape the future of AI infrastructure at GM-this is the team for you.

Compensation: The compensation information is a good faith estimate only. It is based on what a successful applicant might be paid in accordance with applicable state laws. The compensation may not be representative for positions located outside of New York, Colorado, California, or Washington

  • Compensation: The expected base compensation for this role is : $195,000 - $298,000 Actual base compensation within the identified range will vary based on factors relevant to the position.
  • Bonus Potential: An incentive pay program offers payouts based on company performance, job level, and individual performance.
  • Benefits: GM offers a variety of health and wellbeing benefit programs. Benefit options include medical, dental, vision, Health Savings Account, Flexible Spending Accounts, retirement savings plan, sickness and accident benefits, life insurance, paid vacation & holidays.

About the company

We believe we all must make a choice every day - individually and collectively - to drive meaningful change through our words, our deeds and our culture. Every day, we want every employee to feel they belong to one General Motors team., General Motors is committed to being a workplace that is not only free of unlawful discrimination, but one that genuinely fosters inclusion and belonging. We strongly believe that providing an inclusive workplace creates an environment in which our employees can thrive and develop better products for our customers.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · WWC Europe 2026

1:33 min

Summary of machine learning capabilities and engineering opportunities

Jan Zawadzki · LIVE

Videos

See all

Related articles

See all