> Markdown version of [/jobs/ext/511377-machine-learning-engineer-ii-autonomous-driving-inference-runtime](https://www.wearedevelopers.com/jobs/ext/511377-machine-learning-engineer-ii-autonomous-driving-inference-runtime). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer II - Autonomous Driving & Inference Runtime - **Company:** May Mobility - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $180,000.0 - $210,000.0 - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Profiling, Nvidia CUDA, Computer Engineering, Data Centers, Linux System Administration, Machine Learning, Robotic Automation Software, Software Engineering, Pytorch, Information Technology, ONNX (Open Neural Network Exchange) Format, Machine Learning Operations, TensorRT - **Published:** June 10, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=63fe36319cadcc5d ## About the Role Do you have experience in Software engineering?, Do you have a Master's degree?, We are seeking ML-Oriented Software Engineers with experience in robotics applications. As part of our Autonomous Driving ML team, you will use your knowledge of Software and Hardware concepts to deploy, optimize and scale State of the Art Machine Learning models for both Datacenter and Edge Vehicle devices., * Bachelor's or Master's degree in Robotics, Computer Science, Computer Engineering, or a related field with strong mathematical and engineering foundations. * A minimum of 2 years writing software to interface with GPU and ML systems. * Proficiency in C/C++/CUDA/PyTorch and experience in Linux environments. * Familiarity with basic Perception and Planning concepts in Autonomous Driving. Desirable * Familiarity with NVIDIA compute architectures (Ada, Hopper, Blackwell, etc). * Familiarity with common profiling tools such as Nsight, Pytorch Profiler, flamegraph. * Understanding of Quantization (INT8/FP8/FP16) and other compression techniques. * Familiarity with NVIDIA DRIVEOS architecture and SoCs (Orin/Thor). * Familiarity with techniques for scaling training throughput (batching, FSDP, streaming dataloaders). ## Description May Mobility is entering an exciting phase of growth as we expand our first-of-its-kind autonomous shuttle and mobility services across the nation. Launched in 2017 with a strong team of experienced roboticists and software engineers with decades of experience fielding robotic systems in the wild, May Mobility is looking to expand its team of robotics engineers with a background in robotics or autonomous vehicles., * Deploy and Optimize Machine Learning model architectures across May's Autonomous Driving training and inference stacks. * Own the model-compilation and deployment pipeline end-to-end. * Establish and defend latency/throughput budgets across the AV stack, including profiling, regression and integrity tests., * Architecting software for low-level GPU/CPU concurrency such as CUDA streams, pinned memory, kernel fusion and memory-layout optimization. * Use of compilation and runtime utilities such as TensorRT, ONNX and torch.compile for edge deployments. * Apply quantization, distillation, and pruning to fit models within onboard compute and memory budgets. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [How Machine Learning is turning the Automotive Industry upside down](https://www.wearedevelopers.com/videos/61-how-machine-learning-is-turning-the-automotive-industry-upside-down) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)