AI Evaluation Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
We are seeking an AI Evaluation Engineer to establish and maintain the quality standards used to determine whether production AI and LLM systems are ready to launch. This role will build centralized evaluation frameworks, create golden datasets, establish measurable quality thresholds, audit AI evaluation results, and implement continuous production monitoring. The successful candidate will bring strong technical evaluation expertise along with the independence and analytical rigor required to objectively determine whether AI systems meet production quality and safety standards., * Design and build centralized LLM/AI evaluation frameworks and reusable evaluation templates.
- Establish evaluation standards that can be adopted consistently across multiple AI engineering teams.
- Create and maintain golden datasets in partnership with business, domain, and subject-matter experts.
- Develop evaluation suites covering functional quality, accuracy, safety, reliability, and domain-specific requirements.
- Implement and assess LLM-as-a-judge evaluation approaches while accounting for their limitations and failure modes.
- Define measurable pass/fail thresholds and production-readiness criteria for AI applications.
- Independently review and audit evaluation suites developed by individual engineering teams.
- Build regression testing approaches for LLM applications, agents, prompts, retrieval systems, and model changes.
- Establish production monitoring for AI quality degradation, drift, regressions, and incidents.
- Analyze and report evaluation pass rates, quality trends, regressions, and production incidents.
- Partner with AI engineers, platform engineers, product teams, and domain experts to continuously improve AI quality.
Requirements
- 4+ years of experience in ML/LLM evaluation, AI quality engineering, applied research engineering, or related areas.
- Hands-on experience designing and implementing AI/LLM evaluation frameworks.
- Experience developing golden/reference datasets and measurable evaluation criteria.
- Strong understanding of LLM-as-a-judge methodologies and associated failure modes.
- Experience evaluating LLM applications, agents, RAG systems, prompts, or other probabilistic AI systems.
- Strong statistical and analytical skills, including evaluation methodologies for relatively small sample sizes.
- Experience establishing quality thresholds and regression criteria.
- Strong engineering skills with the ability to build reusable evaluation tooling and automation.
- Ability to independently assess system quality and challenge release decisions when evaluation evidence does not meet established standards., * Healthcare, clinical, or other safety-critical AI evaluation experience.
- Experience designing medical-accuracy or domain-specific evaluation suites.
- AI red-teaming or adversarial testing experience.
- Experience with production AI monitoring, drift detection, and continuous evaluation.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Are Large Language Models?
MLOps And AI Driven Development
Trustworthy AI Starts at Deployment: 5 Checks Before You Ship
How to Use Generative AI to Accelerate Learning to Code