> Markdown version of [/jobs/ext/3454229-ai-automation-engineer-real-world-test-lab](https://www.wearedevelopers.com/jobs/ext/3454229-ai-automation-engineer-real-world-test-lab). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Automation Engineer, Real-World Test Lab - **Company:** Niantic, Inc. - **Location:** San Francisco, CA, United States - **Salary:** $158,400.0 - $210,000.0 - **Contract:** Permanent contract - **Skills:** 3D Scanning, Application Programming Interfaces (APIs), Artificial Intelligence, Computer Vision, Cloud Storage, Python (Programming Language), Alwayson, Large Language Models, Backend - **Published:** September 30, 2026 - **Apply:** https://startup.jobs/ai-automation-engineer-real-world-test-lab-niantic-spatial-10227763 ## About the Role * Experience evaluating systems where correctness is graded rather than binary, meaning quality, accuracy, or latency thresholds rather than pass/fail assertions. * Strong Python, with fluency in orchestration, CI-style pipelines, cloud storage, and reproducible environments. * Worked with backend services and APIs you did not own, integrating against them without becoming a bottleneck for their team. * Written up a technical finding clearly enough that a non-author could act on it without a meeting. * A bachelor's degree in a relevant field, or equivalent experience. Nice to Have * Built agent-based or LLM-driven automation for a task that previously required human judgment. * Worked on evaluation or benchmarking for computer vision, 3D reconstruction, or spatial systems. * Instrumented dashboards or scorecards that leadership used for release decisions. * Operated data provisioning or registry infrastructure across multiple environments and access models. ## Description We're hiring an AI Automation Engineer, reporting to the Director of the Real-World Test Lab, to build the system that produces our evidence. Today our evaluations are manual, inconsistent, and slow. You'll design and own the automation that runs them end to end on every relevant release, with no human driving it, turning one-off experiments into an always-on service the whole company relies on. This role is about owning the evidence, not executing a test plan. You'll decide how each workflow gets exercised, and you'll be measured on whether the company can trust the results, not on how much automation exists. The work counts when it keeps running correctly months later, without you in the loop. You believe evaluation is engineering, not process. You've built systems that test other systems, and you know the difference between a script that works on your laptop and infrastructure a team can trust. When a result couldn't be reproduced, you fixed the tooling instead of arguing about the number. What You'll Do * Build the Evaluation Machine - Own the automation that executes Lab evaluations end to end: environment setup, run orchestration, artifact capture, and result collection. Make reruns free so we test constantly rather than occasionally. * Make Results Comparable - Instrument scorecards so a result can be compared across product versions, devices, and capture conditions. A number without its lineage is not evidence. * Automate the Agent-Driven Layer - Build the agent workflows that exercise priority customer use cases at realistic scale, across the graded difficulty spectrum from easy to frontier, and be honest about where agents cannot yet replace human judgment. * Kill Manual Work Permanently - Convert one-off experiments into standing protocols that run on every relevant release. Automate recurring inspection wherever it can be automated, and route what genuinely cannot to scalable human review. * Make Failures Actionable - Produce diagnostics precise enough that a finding reaches its owner with a reproducible case and data attached. Findings that need re-investigation before anyone can act on them are half-finished. * Keep Data From Being the Bottleneck - Work with the AI Data Manager so every run is reproducible from a known dataset state, without anyone downloading and re-uploading data by hand. What Success Looks Like * 30 days: One priority workflow evaluated end to end with no manual steps, with results in a comparable scorecard. * 60 days: Evaluations trigger automatically on relevant releases, with diagnostics that route failures to the right owners. * 90 days: Three priority workflows under standing automated evaluation, with version-over-version comparison for leadership., * Built and maintained production-grade automation or test infrastructure that other engineers relied on daily., * Intellectually honest. You design tests that produce answers people can trust, and you refuse to let a benchmark imply more than the evidence supports - including when the honest answer is inconvenient for your own work. * Pragmatic. You find the fastest path to a defensible result without going through one-way doors that undermine scaling. You have no patience for building infrastructure nobody has asked for yet. * AI-forward by instinct. You reach for agents and LLM-driven automation before you reach for headcount, and you're rigorous about verifying what they produce. * Builds for other people. You've maintained something after the initial excitement wore off. Your tooling is documented, your failures are legible, and colleagues use what you build without asking you to run it for them. ## Related Videos - [Agentic employees in world's most downloaded FinTech app](https://www.wearedevelopers.com/videos/100123-agentic-employees-in-world-s-most-downloaded-fintech-app) - [AI Space Factories, Hacking Self-Driving Cars & Detecting Deepfakes](https://www.wearedevelopers.com/videos/1812-ai-space-factories-hacking-self-driving-cars-detecting-deepfakes) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [AI in Production: applied AI & enterprise use cases](https://www.wearedevelopers.com/videos/100130-ai-in-production-applied-ai-enterprise-use-cases) - [IKEA Story: Transforming an Iconic Retail Brand](https://www.wearedevelopers.com/videos/444-ikea-story-transforming-an-iconic-retail-brand) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [AI Eats the Verifiable First](https://www.wearedevelopers.com/magazine/765-ai-eats-the-verifiable-first) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)