> Markdown version of [/jobs/ext/1264082-technical-lead-evaluation-infrastructure](https://www.wearedevelopers.com/jobs/ext/1264082-technical-lead-evaluation-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Technical Lead (Evaluation Infrastructure) - **Company:** Nuro Inc. - **Location:** Mountain View, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $193,900.0 - $291,100.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Business Analytics Applications, Application Integration Architecture, Big Data, C++ (Programming Language), Software Quality, Continuous Integration, Information Engineering, Data Infrastructure, Cursor (Graphical User Interface Elements), Distributed Systems, Python (Programming Language), Machine Learning, Software Deployment, Data Streaming, Software Organization, Large Language Models, Safety Critical Systems, AI Platforms, Kubernetes, Machine Learning Operations, Data Pipelines - **Published:** July 14, 2026 - **Apply:** https://nuro.ai/careersitem?gh_jid=7963886 ## About the Role * Engineering leadership: Experience setting technical vision, roadmap, and prioritization for a team operating at the intersection of autonomy, safety, and data infrastructure; a clear, concise communicator who partners effectively with PMs, engineers, and cross-functional stakeholders across Autonomy, Systems & Safety, and Simulation * You have a degree in B.Sc or M.Sc., plus 4 years of relevant work experience * Domain experience: Strong fluency in distributed systems, large-scale data and ML evaluation pipelines, metrics frameworks (heuristic and/or ML-based), and analytics platforms * Technical excellence: Ability and willingness to deep-dive into implementation; sets the technical bar for metric quality, pipeline rigor, and safety-critical engineering practice across the broader software organization; strong proficiency in Python, C++, or similar languages * AI-native mindset: Daily user of modern AI coding assistants and agentic tools (Claude Code, Cursor, and similar), with strong intuition for where they accelerate engineering work and where they don't; eager to apply LLMs and ML systems to evaluation problems, from automated triage and metric generation to natural-language analysis of fleet behavior; raises the team's productivity, code quality, and signal density through thoughtful AI integration * Knowledge of data engineering, and its tooling and best practices * Knowledge of batch and streaming data processing, warehousing, and analytics solutions * Experience with data workflow orchestration platforms * Prior experience building evaluation, validation, or analytics platforms, ideally in autonomy, robotics, or safety-critical systems ## Description * Evaluation Infrastructure plays a critical role at Nuro, directly enabling L4 driverless deployment. The team supports two demanding workloads: day-to-day Autonomy Evaluation that powers rapid software iteration, and large-scale Driverless Safety Validation that produces the rigorous evidence required to deploy autonomy on public roads * The Evaluation Infrastructure team builds the metrics framework, evaluation pipelines, introspection tooling, and analysis products that turn raw on-road and simulation logs into actionable insight * Our metrics stack spans both heuristic and ML-based approaches, covering everything from low-level component accuracy to end-to-end behavior quality. The platform empowers autonomy and Systems & Safety teams to run complex evaluations and validations across a wide range of configurations and scales, producing the high-fidelity metrics that drive both short-term iteration and long-term release confidence - in close partnership with Simulation and the broader AI Platform * As the Technical Lead, you will lead the team with deep technical guidance and rigor, setting the technical bar, shortening the time-to-signal for evaluation and the time-to-confidence for validation, so that both autonomy and Systems & Safety teams can iterate fast while deploying software safely * Build and own a unified metrics, evaluation, and validation platform - pipelines, introspection tooling, and analysis products that turn on-road and simulation logs into high-fidelity signals for autonomy iteration and driverless safety validation * Drive the technical bar for metric quality across both heuristic and ML-based approaches; invest in the scale, reliability, and CI/CD of the evaluation stack to shorten time-to-signal for evaluation and time-to-confidence for validation, and to meet high SLAs for downstream stakeholders * Mentor and grow the Evaluation Infrastructure team, and champion AI-native engineering practices that compound team velocity and code quality * Partner with Product, Autonomy, Systems & Safety, and Simulation teams to define and execute the vision and strategy for evaluation at Nuro ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [This App Reached 10,000 Users in One Week. Here's How.](https://www.wearedevelopers.com/videos/100329-this-app-reached-10-000-users-in-one-week-here-s-how) - [How to develop an autonomous car end-to-end: Robotic Drive and the mobility revolution](https://www.wearedevelopers.com/videos/22-how-to-develop-an-autonomous-car-end-to-end-robotic-drive-and-the-mobility-revolution) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)