> Markdown version of [/jobs/ext/1889826-lead-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/1889826-lead-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Machine Learning Engineer - **Company:** Serve Robotics - **Location:** Los Angeles, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $177,000.0 - $215,000.0 - **Contract:** Permanent contract - **Skills:** Training Data, Computer-Aided Design, Artificial Neural Networks, Computer Vision, Cloud Computing, Computer Clusters, Computer Programming, Data Files, Extract Transform Load (ETL), Distributed Computing Environment, Memory Management, Fault Tolerance, Design of User Interfaces, Hardware Design, Human-Computer Interaction, Python (Programming Language), Machine Learning, Network Architecture, Data Processing, Scripting, Graphics Processing Unit (GPU), Delivery Pipeline, Model Validation, Usage Tracking, Information Technology, Optimization Algorithms, Data Management, Machine Learning Operations, Lidar, Data Pipelines - **Published:** July 31, 2026 - **Apply:** https://www.careerbuilder.com/job-details/lead-machine-learning-engineer-los-angeles-ca--6c8778ba-bb59-40d1-97e7-8687a61fb2e7 ## About the Role * Master's or PhD in Computer Science, Robotics, Electrical Engineering, Machine Learning, or a closely related technical discipline. * Minimum of 5 years of professional experience developing, training, and deploying machine learning models in production environments. * Hands-on experience training machine learning models across multiple GPUs or compute nodes, including familiarity with distributed training frameworks and large dataset handling. * Strong programming skills in Python for implementing machine learning models, data pipelines, and training workflows. * Solid knowledge of core concepts such as neural networks, optimization algorithms, loss functions, model evaluation, and training methodologies. What Makes You Stand out * Experience identifying and resolving training bottlenecks related to compute utilization, memory usage, and data throughput in machine learning systems. * Experience training machine learning models on robotics or autonomous driving datasets involving multimodal sensor inputs such as camera video, LiDAR point clouds, radar, or telemetry data. * Experience developing models that combine multiple data modalities (e.g., images, point clouds, and structured sensor data) into a unified learning system. * Peer-reviewed publications or significant research contributions in machine learning, robotics, or related areas. Please note: The listed base salary range applies to candidates based in the US. Compensation may vary depending on location, experience, and role alignment. We are open to qualified candidates working remotely in Canada * Canada - ALL: $177k - $215k CAD Skills: Analysis Skills, Autonomous Driving Systems, Cloud Computing, Computer Programming, Computer Science, Computer Vision, Data Management, Data Modeling, Data Processing, Data Sets, Electrical Engineering, GPU (Graphics Processing Unit), Hardware Design, Light Detection and Ranging (LiDAR)\Laser Detection and Ranging (LADAR), Machine Learning, Memory Hardware, Metrics, Network Architecture/Engineering, Neural Networks, Optimization Algorithm, Performance Analysis, Performance Management, Performance Modeling, Problem Solving Skills, Production Systems, Python Programming/Scripting Language, Robotics, Structured Data, Systems Maintenance, Team Player, Telemetry, Time Management, Training Data Sets, Ubiquity, User Interface/Experience (UI/UX), Vehicle Fleets ## Description * Design and maintain training systems that can process and learn from petabyte-scale multimodal datasets (e.g., video and point cloud data). This includes ensuring data is efficiently loaded, distributed, and processed across large GPU clusters. * Identify and resolve bottlenecks in the training pipeline, including data loading, preprocessing, model computation, and inter-node communication, to maximize GPU utilization and reduce training time. * Work with the ML team to develop and refine neural network architectures suitable for autonomy tasks, particularly those handling high-dimensional and sequential sensor data. * Create and adjust loss functions and training strategies that help the model learn effectively from complex multimodal inputs and improve autonomy performance. * Configure, monitor, and maintain large-scale distributed training jobs across multiple machines and GPUs, ensuring stability, fault tolerance, and efficient resource usage. * Implement scalable systems to preprocess, transform, and augment large robotics datasets so that they are suitable for model training. * Work closely with ML scientists and other engineers to integrate new models, experiments, and training approaches into the production training pipeline. * Analyze training metrics, model outputs, and experiment logs to assess model performance and guide improvements in architecture, data usage, or training strategies. * Develop tools and workflows that allow teams to run experiments, track results, and iterate quickly on new model ideas or training approaches. ## Related Videos - [How to develop an autonomous car end-to-end: Robotic Drive and the mobility revolution](https://www.wearedevelopers.com/videos/22-how-to-develop-an-autonomous-car-end-to-end-robotic-drive-and-the-mobility-revolution) - [How Machine Learning is turning the Automotive Industry upside down](https://www.wearedevelopers.com/videos/61-how-machine-learning-is-turning-the-automotive-industry-upside-down) - [Introduction to TXT](https://www.wearedevelopers.com/videos/30-introduction-to-txt) - [JavaScript? No. Java Scripts! - Scripting with Java](https://www.wearedevelopers.com/videos/2094-javascript-no-java-scripts-scripting-with-java) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Implementing continuous delivery in a data processing pipeline](https://www.wearedevelopers.com/videos/73-implementing-continuous-delivery-in-a-data-processing-pipeline) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)