Senior/Staff Machine Learning Engineer, Perception
Role details
Job location
Tech stack
Job description
We're looking for a skilled ML engineer to build the perception systems that give our autonomous machines human-like awareness in rugged, unstructured environments. You'll develop computer vision and machine learning systems that turns noisy camera and LiDAR data into robust 3D scene understanding - enabling heavy equipment to operate safely through dust, glare, occlusion, and whatever messy conditions a working site throws at it.
The field is moving past bounding-box detection and hand-tuned tracking toward learned, dense scene representations, foundation-model-driven data engines, and uncertainty-aware perception. You'll be at the center of that shift. This role is hands-on: you'll write production-grade software, distill and optimize models for embedded hardware, and validate your work on real machines at operating around the world., * Develop real-time perception models for open-world obstacle and terrain understanding.
- Build multi-modal fusion that combines camera and LiDAR into a unified 3D/BEV representation, robust to occlusions, sensor degradation, and GNSS outages.
- Optimize models for low-latency inference on resource-constrained hardware, balancing accuracy and performance.
- Design auto-labeling pipelines that leverage foundation models and teacher-student distillation to scale labeling and close the loop from real-world field interventions.
- Design data and evaluation pipelines that curate large multi-sensor datasets and surface failures fast, with strong visualization and debugging tooling.
- Analyze performance metrics and iterate on algorithms to improve accuracy and efficiency of various perception subsystems.
Requirements
- A MS/PhD in Computer Science, AI, or a related field, or 6+ years of industry experience building vision-based perception systems.
- Deep expertise developing and deploying modern perception models: detection, segmentation, mono/stereo/metric depth, BEV/occupancy, sensor fusion, and 3D scene understanding.
- Fluency adapting, fine-tuning, and distilling large pre-trained vision and vision-language models.
- Strong grounding in multi-sensor integration (camera, LiDAR, radar): calibration, spatiotemporal sync, and cross-modal fusion.
- Experience handling large datasets efficiently and organizing them for labeling, training and evaluation.
- Fluency in Python with PyTorch/TensorFlow/OpenCV and the ability to write efficient, production-ready code for real-time systems.
- Proven ability to design experiments, analyze metrics (mAP, IoU, latency/throughput, and calibration/ECE), and optimize to meet stringent real-world performance and safety requirements.
- An eagerness to get your hands dirty and agility in a fast-moving, collaborative, small team environment with lots of ownership.
What Makes You a Strong Fit
- Experience architecting multi-sensor ML systems from scratch.
- Experience building auto-labeling / data-engine flywheels at scale.
- Experience with compute-constrained pipelines including optimizing models to balance the accuracy vs. performance tradeoff, leveraging TensorRT, model quantization, etc.
- Familiarity with emerging predictive world models for anticipation, anomaly detection, or closed-loop simulation, and adjacent policy paradigms such as Vision-Language-Action (VLA) and World-Action (WAM) models.
- Experience with compute-constrained deployment: TensorRT, model quantization, and custom CUDA operations.
- Publications at top-tier perception/robotics venues (CVPR, ICRA, CoRL, RSS, etc.).
- Passion for how we feed, build, move, and maintain the world.
Benefits & conditions
Pulled from the full job description
- 401(k)
- Health insurance
- Vision insurance
- Health savings account
- Dental insurance
- Flexible spending account
- Stock options, The US base salary range for this full-time position is $200,000 to $280,000 + equity + benefits + unlimited PTO, * 100% covered medical, dental, and vision for the employee (partner, children, or family is additional)