> Markdown version of [/jobs/ext/1897312-software-engineer-ml-performance-optimization-organization](https://www.wearedevelopers.com/jobs/ext/1897312-software-engineer-ml-performance-optimization-organization). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, ML Performance Optimization Organization - **Company:** Zoox - **Location:** Foster City, CA, United States - **Experience:** Experienced - **Salary:** $192,000.0 - $257,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Profiling, Computer Engineering, Python (Programming Language), Machine Learning, Performance Tuning, Software Architecture, Requirements Management, Software Engineering, Scripting, Graphics Processing Unit (GPU), Pytorch, Machine Learning Operations, TensorRT - **Published:** July 31, 2026 - **Apply:** https://www.careerbuilder.com/job-details/software-engineer-ml-performance-optimization-foster-city-ca--1c7d4356-721c-4483-b8bf-5ee6ab205e89 ## About the Role Note: You do not have to meet all the requirements below to be considered for this position: * 4+ years of total experience, including 2+ years of working on large-scale model training or inference platforms. * Experience with training frameworks like PyTorch, leveraging GPUs efficiently for distributed model training. * Experience with GPU-accelerated inference using TensorRT or similar frameworks. * Experience using profiling tools like NVIDIA's Nsight or PyTorch's Profiler for identifying model training and serving bottlenecks. * Proficient in Python or C++., Architectural Services, Artificial Intelligence (AI), Automotive Automation, Autonomous Driving Systems, C++ Programming Language, Computer Engineering, Cross-Functional, Ecosystems, GPU (Graphics Processing Unit), LinkedIn, Machine Learning, Performance Tuning/Optimization, Python Programming/Scripting Language, Requirements Management, Robotics, Simulation, Software Engineering, Team Building, Team Lead/Manager, Use Cases, Vehicle Fleets ## Description Are you excited to drive our ML Performance Optimization initiatives and make our ML models that enable autonomous driving as fast and efficient as possible? You will get to work with SOTA accelerators, cutting-edge techniques in distributed training, quantization, distillation, and pruning, among other things, working closely with all the Autonomy teams within Zoox - Perception, Prediction, Planner, Simulation, Collision Avoidance, and have the opportunity to significantly push the boundaries of how ML is practiced within Zoox. We build and operate the base layer of ML tools, model development, and serving systems that our applied research teams use for in- and off-vehicle ML use cases. You will work alongside a team of strong software engineers and act as a force multiplier for our internal customers. This team has many growth opportunities as we expand our robotaxi deployments and venture into new ML domains. If you want to learn more about our stack behind autonomous driving, please look here. If you want to learn more about our ML Infrastructure, here is one of our past talks at re:Invent. In this role, you will: * Design, implement, and operate cutting-edge ML Training OR Inference performance optimization techniques to scale our VLM, VLA, and Foundational models and deploy them efficiently in our robotaxi. * Collaborate closely with cross-functional teams, including ML researchers, software engineers, data engineers, and hardware engineers, to define requirements and align on architectural decisions. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)