> Markdown version of [/jobs/ext/2722730-machine-learning-engineer-perception-3d-segmentation](https://www.wearedevelopers.com/jobs/ext/2722730-machine-learning-engineer-perception-3d-segmentation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer - Perception 3D Segmentation - **Company:** Zoox - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Computer Vision, C++ (Programming Language), Nvidia CUDA, Python (Programming Language), Machine Learning, Sensor Fusion, Pytorch, Deep Learning, Gaussian, Information Technology, TensorRT, Lidar - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-machine-learning-engineer-perception-3d-segmentation-zoox-8768622 ## About the Role * MS or PhD in Computer Science, Robotics, Machine Learning, or related field with 6+ years of industry experience. * Deep expertise in 3D Computer Vision and Deep Learning, specifically with voxel-based or BEV (Bird's Eye View) architectures. * Strong proficiency in Python and deep learning frameworks (PyTorch) for model training and design as well as some experience in C++ for model integration. * Experience with multi-sensor fusion (Lidar, Camera, Radar) and handling temporal data sequences. * Experience with occupancy networks, implicit representations (NeRF/Gaussian Splats), or scene flow estimation., * Experience optimizing models for TensorRT/CUDA to achieve low-latency inference. * Familiarity with sparse convolutions or query-based architectures for efficient 3D processing. * Experience with Vision Language Model, or multi-modal 3D foundation model, or World Model, or VLA. ## Description The Perception team at Zoox is responsible for the robot's understanding of the world, fusing data from Lidar, Radar, and Cameras to create a unified representation of the environment. In this role, you will contribute to the development of our next-generation 3D occupancy and segmentation networks. You will architect and optimize high-performance deep learning models that generate dense, temporally consistent voxel representations of the driving environment. This work is critical for enabling our vehicle to navigate complex urban scenarios, handle rare obstacles, and drive safely in tight spaces by providing precise geometry and motion estimates to downstream planners., * Design and implement state-of-the-art multi-modal sensor fusion architectures (Lidar, Camera, Radar) to predict 3D occupancy, semantic segmentation, and flow . * Develop "vision-first" fusion strategies to enhance geometric understanding and reduce dependency on sparse sensor modalities . * Engineer temporal processing modules to improve the stability and consistency of predictions over time. * Optimize model architectures for real-time on-vehicle inference, balancing high-fidelity range extension with strict latency constraints . * Collaborate with downstream consumers (Tracking, Prediction, Planner) to refine geometric outputs, such as contours and free-space estimations, for complex maneuvering. ## Related Videos - [How to develop an autonomous car end-to-end: Robotic Drive and the mobility revolution](https://www.wearedevelopers.com/videos/22-how-to-develop-an-autonomous-car-end-to-end-robotic-drive-and-the-mobility-revolution) - [Finding the unknown unknowns: intelligent data collection for autonomous driving development](https://www.wearedevelopers.com/videos/519-finding-the-unknown-unknowns-intelligent-data-collection-for-autonomous-driving-development) - [WeAreDevelopers LIVE - Building The World’s Worst Image Editor™](https://www.wearedevelopers.com/videos/1836-wearedevelopers-live-building-the-world-s-worst-image-editor) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Confuse, Obfuscate, Disrupt: Using Adversarial Techniques for Better AI and True Anonymity](https://www.wearedevelopers.com/videos/1456-confuse-obfuscate-disrupt-using-adversarial-techniques-for-better-ai-and-true-anonymity) - [How Machine Learning is turning the Automotive Industry upside down](https://www.wearedevelopers.com/videos/61-how-machine-learning-is-turning-the-automotive-industry-upside-down) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)