> Markdown version of [/jobs/ext/1902701-software-engineer-ai-evals](https://www.wearedevelopers.com/jobs/ext/1902701-software-engineer-ai-evals). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, AI Evals - **Company:** SENTRY - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $240,000.0 - $280,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Automation of Tests, Data Files, Data Infrastructure, Software Debugging, Programming Tools, Systems Analysis, Python (Programming Language), Machine Learning, Open Source Technology, Operational Databases, Regression Testing, Software Engineering, Test Case, TypeScript, Unstructured Data, Web Analytics, Scripting, Large Language Models, Build Management, Information Technology, Sentry, Production Code, Data Management, Machine Learning Operations - **Published:** July 31, 2026 - **Apply:** https://www.careerbuilder.com/job-details/senior-software-engineer-ai-evals-san-francisco-ca--9e15d6da-f67b-406f-911b-db7a5a1dfd5c ## About the Role * Minimum 5+ years of professional experience with a Bachelor's degree in computer science, machine learning, or a related field * Experience building testing, evaluation, or data infrastructure for complex systems (AI/ML experience strongly preferred) * Comfort writing production-quality code (we use Python and TypeScript) * Experience working with structured and unstructured datasets, labeling workflows, or data quality pipelines * Familiarity with modern ML systems and evaluation techniques (e.g., offline metrics, online evaluation, regression testing for models or prompts) * Bonus: experience evaluating LLMs, agentic systems, or AI-assisted developer tools, Artificial Intelligence (AI), Artificial Intelligence (AI) Agents, Benchmarking, Compensation and Benefits, Computer Science, Concrete, Cross-Functional, Data Management, Data Quality, Data Sets, Debugging Skills, Employee Benefits, Firefighting, Machine Learning, Open Source, Production Control, Programming Tools, Python Programming/Scripting Language, Quality Metrics, Regression Testing, Software Engineering, Systems Analysis, Test Automation, Test Case, Test Harness, Testing, Web Analytics ## Description As a Senior Software Engineer on Sentry's AI/ML team, you'll be responsible for building the evaluation infrastructure that measures the accuracy, reliability, and real-world performance of our AI systems. This role is critical to ensuring that our debugging agents and AI-powered features behave correctly, safely, and predictably as they scale. You'll design datasets, benchmarks, and test harnesses that turn ambiguous AI behavior into measurable signals, helping the team ship AI with confidence. In this role you will * Design and build robust evaluation frameworks to measure accuracy, reliability, regressions, and edge cases in AI systems * Create and curate high-quality datasets, golden test cases, and benchmarks grounded in real production data * Build automated test harnesses and metrics pipelines to continuously evaluate models, prompts, and agentic workflows * Partner closely with applied AI engineers and product leaders to define what "good" looks like and translate it into measurable criteria * Own the evaluation lifecycle for major AI initiatives, from early experimentation through production monitoring You'll love this job if you * Care deeply about correctness, rigor, and measurement in AI systems * Enjoy turning fuzzy product goals and model behavior into concrete tests and metrics * Like building foundational infrastructure that unlocks faster iteration and higher confidence for the entire AI team * Thrive in cross-functional environments and enjoy influencing model design through better evaluation ## Related Videos - [The AI-Ready Stack: Rethinking the Engineering Org of the Future](https://www.wearedevelopers.com/videos/1706-the-ai-ready-stack-rethinking-the-engineering-org-of-the-future) - [AIQSpecFlow: Improves and automates your agile process of specification and creation of testcases.](https://www.wearedevelopers.com/videos/100084-aiqspecflow-improves-and-automates-your-agile-process-of-specification-and-creation-of-testcases) - [JavaScript? No. Java Scripts! - Scripting with Java](https://www.wearedevelopers.com/videos/2094-javascript-no-java-scripts-scripting-with-java) - [From Doubt to Confidence: How Sentry Uses Verdaccio to Bulletproof SDK Releases](https://www.wearedevelopers.com/videos/739-from-doubt-to-confidence-how-sentry-uses-verdaccio-to-bulletproof-sdk-releases) - [AI as a Test Designer: Transforming Experience into Automated Testing](https://www.wearedevelopers.com/videos/1984-ai-as-a-test-designer-transforming-experience-into-automated-testing) - [Intermediate Bitcoin Script](https://www.wearedevelopers.com/videos/25-intermediate-bitcoin-script) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai)