> Markdown version of [/jobs/ext/1287607-software-engineering-manager-ai-platform](https://www.wearedevelopers.com/jobs/ext/1287607-software-engineering-manager-ai-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineering Manager, AI Platform - **Company:** Procore - **Location:** Austin, TX, United States - **Experience:** Expert - **Salary:** $168,560.0 - $231,770.0 - **Contract:** Permanent contract - **Skills:** A/B Testing, Artificial Intelligence, Code Coverage, Software Quality, Continuous Integration, Programming Tools, Distributed Systems, Information Retrieval, Python (Programming Language), Project Management Software, Software Engineering, Large Language Models, Reliability of Systems, Technical Debt, Backend, AI Platforms, Low Latency, Data Pipelines - **Published:** July 16, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/87290910/1 ## About the Role * 2+ years managing engineering teams or technical leads, with 7+ years total in software engineering. * Experience building evaluation, quality measurement, or observability platforms for LLM-based or agentic systems (RAG pipelines, multi-step agents, tool-use agents). * Strong understanding of evaluation methodologies: precision/recall, LLM-as-judge, human annotation, A/B testing, and statistical significance frameworks. * Proven ability to translate ambiguous problem spaces into clear technical strategies and executable roadmaps. * Hands-on technical depth in backend systems, data pipelines, or distributed infrastructure (Python, Go, or similar) * Familiarity with evaluation frameworks such as RAGAS, DeepEval, LangFuse, or custom eval harnesses. * Background in search relevance (NDCG, MRR) or information retrieval quality systems. * Experience with construction-tech, procurement, or enterprise B2B SaaS domains. ## Description We're looking for a Software Engineering Manager for our AI Evaluation Platform team to join Procore's Construction Intelligence organization. In this role, you'll build the infrastructure and tooling that enables users and internal teams to measure, benchmark, and improve the quality of AI agents - including Search Agent, RFI Create Agent, Invoice Agent, and future agentic products. You will own the end-to-end evaluation lifecycle: from defining quality metrics and building evaluation frameworks, to delivering intuitive interfaces that surface actionable insights about agent performance. This position reports into Sr Director of the Procore AI Engineering team and will be 2 days per week hybrid role in our Austin office. We're looking for someone to join us immediately. What you'll do: * Lead and grow a team of engineers focused on evaluation infrastructure, quality measurement, and developer tooling for AI agents. * Define the technical vision and roadmap for the Evaluation Platform - covering offline evaluations (batch benchmarks, regression suites) and online evaluations (live traffic quality monitoring, A/B testing). * Partner with AI/ML, Product, and Agent teams to define quality metrics for agents (relevance, accuracy, latency, safety, user satisfaction, token usage) and build automated pipelines to compute them at scale. * Design and deliver user-facing evaluation tools that allow customers and internal teams to assess agent output quality, compare model versions, and identify regressions. * Build frameworks for human-in-the-loop evaluation - annotation workflows, rating interfaces, and inter-rater reliability measurement. * Establish CI/CD quality gates so that new agent versions cannot ship without passing evaluation thresholds. * Drive engineering excellence: code quality, system reliability, test coverage, on-call health, and technical debt management. * Recruit, mentor, and develop engineers - fostering a culture of ownership, curiosity, and rigorous experimentation. ## Related Videos - [Bringing the power of AI to your application.](https://www.wearedevelopers.com/videos/1010-bringing-the-power-of-ai-to-your-application) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [AI Won't Fix Your Engineering Culture](https://www.wearedevelopers.com/videos/100266-ai-won-t-fix-your-engineering-culture) - [The AI-Ready Stack: Rethinking the Engineering Org of the Future](https://www.wearedevelopers.com/videos/1706-the-ai-ready-stack-rethinking-the-engineering-org-of-the-future) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) ## Related Articles - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)