> Markdown version of [/jobs/ext/144639-machine-learning-systems-engineer](https://www.wearedevelopers.com/jobs/ext/144639-machine-learning-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Systems Engineer - **Company:** Motional LLC - **Location:** San Francisco, CA, United States (Remote available) - **Salary:** $144,000.0 - $192,000.0 - **Contract:** Permanent contract - **Skills:** Systems Engineering, Profiling, Nvidia CUDA, Computer Engineering, Extract Transform Load (ETL), Distributed Computing Environment, Python (Programming Language), Linux Kernel, Machine Learning, Tensorflow, Software Engineering, Pytorch, Information Technology, Data Analytics, Machine Learning Operations, Data Pipelines, Software Library - **Published:** May 13, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=5923701386985356 ## About the Role Do you have experience in Machine learning libraries?, Do you have a Master's degree?, * Education: Bachelor's, Master's degree, or PhD in Computer Science, Computer Engineering, or a related technical discipline. * Software Engineering: Strong proficiency in Python. * ML Frameworks: Extensive hands-on experience with PyTorch. * ML Knowledge: Experience optimizing machine learning model execution during training and inference, alongside a strong understanding of fundamental machine learning concepts, architectures, and processes. * Problem Solving: Exceptional analytical and problem-solving skills, with a bias for action and a data-driven approach to technical challenges. We encourage a hybrid schedule with in-office time at one of our locations in Boston, Pittsburgh, or Las Vegas to support collaboration, or this role can be fully remote. ## Description We are looking for a Machine Learning Systems Engineer to join our ML Acceleration team. In this role, you will be responsible for the core systems that enable our researchers to train frontier models at scale, focusing obsessively on speed, cost, reliability, and throughput. You will work at the intersection of machine learning research and high-performance systems engineering. Your work will directly impact our ability to scale large-scale distributed model training and reduce the time-to-convergence for our next generation of models., * Performance Profiling & Optimization: Utilize profiling tools (e.g., Nsight, PyTorch Profiler) to identify bottlenecks in data loading, gradient computation, and communication. Implement optimizations like kernel fusion, sharding, and tiling to improve step time. * Distributed Training: Optimize distributed training pipelines using frameworks such as PyTorch Distributed. * Kernel Development: Design and maintain high-performance GPU kernels in Triton or CUDA for state-of-the-art ML workloads. * Data Pipeline Engineering: Optimize robust data loading pipelines that maximize training throughput. ## Related Videos - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [Overview of Machine Learning in Python](https://www.wearedevelopers.com/videos/840-overview-of-machine-learning-in-python) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)