> Markdown version of [/jobs/ext/1893579-machine-learning-engineer-infra](https://www.wearedevelopers.com/jobs/ext/1893579-machine-learning-engineer-infra). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer (Infra) - **Company:** Waymo LLC - **Location:** Mountain View, CA, United States - **Experience:** Expert - **Salary:** $213,000.0 - $263,000.0 - **Contract:** Permanent contract - **Skills:** Training Data, Java (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, Information Engineering, Data Systems, Device Drivers, Distributed Systems, Machine Learning, Software Architecture, Systems Development Life Cycle, Tensorflow, Software Engineering, Systems Architecture, Extensible Markup Language (XML), Pytorch, Information Technology, Machine Learning Operations - **Published:** August 2, 2026 - **Apply:** https://www.careerbuilder.com/job-details/senior-machine-learning-engineer-infra-driver-understanding-and-evaluation-mountain-view-ca--1ed8bbca-8e48-47ab-b838-9f04be0f7c37 ## About the Role * M.S. or Ph.D. degree Computer Science, Machine Learning, Artificial Intelligence, or a related technical field, or equivalent practical experience. * 5+ years in machine learning infrastructure such as developing, designing, scaling, training, deploying, and optimizing large-scale machine learning systems from data to model. * A history of contributions to machine learning tooling and frameworks e.g. PyTorch, Jax, Tensorflow, Ray, or similar. The candidate should understand both the user facing API and the internal workings. * Strong expertise in distributed training techniques, including gradient sharding and optimization strategies for scaling large models across ML accelerator profiling tools to uncover performance bottlenecks. We prefer: * 7+ years in machine learning infrastructure such as developing, designing, scaling, training, deploying, and optimizing large-scale machine learning systems from data to model. * Experience in the autonomous vehicles domain, robotics, or complex simulation environments. * Familiarity with large-scale simulation platforms and their integration with ML training workflows., Application Programming Interface (API), Architectural Services, Artificial Intelligence (AI), Autonomous Driving Systems, Computer Science, Cross-Functional, Data Modeling, Data Partitioning, Develop Methodologies, Device Drivers, Distributed Computing, JAX (Java API for XML), Large-Scale Systems, Machine Learning, Machine Tool, Metrics, People Management, Performance Management, Process Improvement, Robotics, Scalable System Development, Simulation, Software Engineering, System Architecture, Technical Delivery, Training Data Sets, Training/Teaching, Use Cases, Vehicle Fleets ## Description The DUE Machine Learning team will build and operate scalable machine learning and data systems, simulation workflow and insight tools, improve and speed up the evaluation and onboard developer journeys. It will combine expert human judgements and advanced machine learning models to deliver training and evaluation data for hundreds of metrics and components that make up the Waymo driver. We are looking for researchers and software engineers who are passionate about developing machine learning techniques for the Evaluation systems on our autonomous vehicles, and have an incessant drive to improve the performance of our technology stack. You will: * Build scalable systems for training and fine-tuning large-scale models to evaluate interesting driving behaviors. * Work at the intersection of data engineering, model development, and simulation Provide guidance on architectural decisions and technical directions. Own large, complex systems, driving architectures that meet technical and business objectives. * Contribute to the production and optimization of machine learning models aiming to assess Waymo's expansive fleet of vehicles that cumulatively travel millions of miles. * Design and scale large distributed systems covering the ML lifecycle, supporting planet-scale dataset generation, model training, and evaluation. * Collaborate cross-functionally to derive performance and system-level requirements for large ML systems. Translate product/business goals into measurable technical deliverables, ensuring system component alignment. ## Related Videos - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [How Machine Learning is turning the Automotive Industry upside down](https://www.wearedevelopers.com/videos/61-how-machine-learning-is-turning-the-automotive-industry-upside-down) - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)