> Markdown version of [/jobs/ext/3136280-technical-lead-autonomy-evaluation](https://www.wearedevelopers.com/jobs/ext/3136280-technical-lead-autonomy-evaluation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Technical Lead - Autonomy Evaluation - **Company:** ATOM, INC. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $185,000.0 - $242,000.0 - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Computer Clusters, Code Coverage, Code Review, Data Mining, Software Debugging, Middleware, Python (Programming Language), Machine Learning, Software Engineering, System Testing, Data Logging, Information Technology, Lidar - **Published:** September 29, 2026 - **Apply:** https://startup.jobs/technical-lead-autonomy-evaluation-atoms-10200585 ## About the Role * Bachelor's degree in Computer Science, Robotics, Electrical Engineering, Statistics, or a related field with 8+ years of relevant experience. * 2+ years as a technical lead or engineering manager for an evaluation, validation, or test team. * Ownership of an offline or system-level evaluation system for an autonomous vehicle or robotics program, or a comparable large-scale ML model evaluation system in production, with accountability for the metrics and for the release decisions made on them. * Experience designing safety or behavior metrics for an autonomous system and defending them to engineering and executive audiences. * Hands-on experience building log replay at scale, including scoring against ground truth in open-loop and closed-loop evaluation. * Experience running evaluation pipelines on GPU clusters or comparable large-scale compute on a recurring cadence, with responsibility for reproducibility, throughput, and compute cost. * Experience with evaluation dataset design, benchmarking, and regression detection across software or model versions. * Working command of applied statistics for system validation, including statistical uncertainty, sampling, rare-event analysis, and the data volume a given claim requires. * Working understanding of autonomous systems end to end, including sensing, perception, localization, planning, and controls. * Experience on an autonomous vehicle or robotics platform. * Working experience with a robotics middleware and its logging and replay tooling. * Experience handling autonomous driving or robotics sensor data across multiple modalities and timestamps, including camera, LiDAR, and radar. * Proficiency in Python and the ability to read, navigate, and debug existing C++ codebases. ## Description We are seeking a Technical Lead for Autonomy Evaluation to own the metrics our autonomy releases are judged on and the log replay and evaluation platform that computes them. In this role, you will take evaluation from recorded logs, to scored regression suites running on every candidate release, to the report a release passes before it reaches a vehicle, developing how replay, scoring, metrics, and test coverage come together into one platform engineers use every day. You will do this across on-road vehicles, and you will build the evaluation infrastructure that lets each new software release, sensor configuration, and platform be assessed seamlessly. What you'll do * Metrics. Own the safety and behavior metrics a release is judged on, including collision and near-miss measures, trajectory agreement against ground truth, and comfort, and the go/no-go release criteria built on them. * Log replay and scoring. Own the pipeline that converts recorded logs into scored test cases, covering open-loop and closed-loop evaluation of perception, localization, and planning against ground truth. * Regression and benchmarking. Own the system that benchmarks new software against old across the log corpus for every candidate release, including evaluation dataset design, regression detection, and triage of a regressed case to an owner. * Test coverage and data mining. Own how the corpus is mined for the long tail of rare and important events, including sampling strategy, selection of logs by expected value per replay-hour, and coverage of scenario classes the corpus does not yet contain. * Execution at scale. Own deterministic and reproducible execution of evaluation, run identity, traceability of every result to a software version, and cost per replay-hour as a tracked number. * Engineering practices. Set the practices for metric definitions, test design, code review, and the write-up of evaluation results. * Technical standard. Mentor the engineers who join the function, set the technical bar for their work, and participate in hiring., This role is based in our San Francisco office location. As a company driven by innovation and continuous change, close collaboration is essential. We're constantly reimagining our industry, creating new products, and refining our processes, and we do our best work together. That's why all of our office-based teams work onsite, five days a week. ## Related Videos - [How to develop an autonomous car end-to-end: Robotic Drive and the mobility revolution](https://www.wearedevelopers.com/videos/22-how-to-develop-an-autonomous-car-end-to-end-robotic-drive-and-the-mobility-revolution) - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [Developer’s Perspective: Overview of the Tezos Blockchain Ecosystem](https://www.wearedevelopers.com/videos/237-developer-s-perspective-overview-of-the-tezos-blockchain-ecosystem) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Automated Driving - Why is it so hard to introduce](https://www.wearedevelopers.com/videos/628-automated-driving-why-is-it-so-hard-to-introduce) - [Build a CI/CD pipeline to automate code reviews and ensure code quality](https://www.wearedevelopers.com/videos/349-build-a-ci-cd-pipeline-to-automate-code-reviews-and-ensure-code-quality) ## Related Articles - [ I Gave a Video Editor More Autonomy Than a Trading Bot. On Purpose.](https://www.wearedevelopers.com/magazine/773-i-gave-a-video-editor-more-autonomy-than-a-trading-bot-on-purpose) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [AI Eats the Verifiable First](https://www.wearedevelopers.com/magazine/765-ai-eats-the-verifiable-first) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Never delegate the understanding](https://www.wearedevelopers.com/magazine/749-never-delegate-the-understanding)