AI Evaluation Engineer

Acunor Infotech
United States
10 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Automated Storage and Retrieval Systems Regression Testing Large Language Models Build Management

Job description

We are seeking an AI Evaluation Engineer to establish and maintain the quality standards used to determine whether production AI and LLM systems are ready to launch. This role will build centralized evaluation frameworks, create golden datasets, establish measurable quality thresholds, audit AI evaluation results, and implement continuous production monitoring. The successful candidate will bring strong technical evaluation expertise along with the independence and analytical rigor required to objectively determine whether AI systems meet production quality and safety standards., * Design and build centralized LLM/AI evaluation frameworks and reusable evaluation templates.

  • Establish evaluation standards that can be adopted consistently across multiple AI engineering teams.
  • Create and maintain golden datasets in partnership with business, domain, and subject-matter experts.
  • Develop evaluation suites covering functional quality, accuracy, safety, reliability, and domain-specific requirements.
  • Implement and assess LLM-as-a-judge evaluation approaches while accounting for their limitations and failure modes.
  • Define measurable pass/fail thresholds and production-readiness criteria for AI applications.
  • Independently review and audit evaluation suites developed by individual engineering teams.
  • Build regression testing approaches for LLM applications, agents, prompts, retrieval systems, and model changes.
  • Establish production monitoring for AI quality degradation, drift, regressions, and incidents.
  • Analyze and report evaluation pass rates, quality trends, regressions, and production incidents.
  • Partner with AI engineers, platform engineers, product teams, and domain experts to continuously improve AI quality.

Requirements

  • 4+ years of experience in ML/LLM evaluation, AI quality engineering, applied research engineering, or related areas.
  • Hands-on experience designing and implementing AI/LLM evaluation frameworks.
  • Experience developing golden/reference datasets and measurable evaluation criteria.
  • Strong understanding of LLM-as-a-judge methodologies and associated failure modes.
  • Experience evaluating LLM applications, agents, RAG systems, prompts, or other probabilistic AI systems.
  • Strong statistical and analytical skills, including evaluation methodologies for relatively small sample sizes.
  • Experience establishing quality thresholds and regression criteria.
  • Strong engineering skills with the ability to build reusable evaluation tooling and automation.
  • Ability to independently assess system quality and challenge release decisions when evaluation evidence does not meet established standards., * Healthcare, clinical, or other safety-critical AI evaluation experience.
  • Experience designing medical-accuracy or domain-specific evaluation suites.
  • AI red-teaming or adversarial testing experience.
  • Experience with production AI monitoring, drift detection, and continuous evaluation.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:11 min

Integrating flat external dependencies without build management tools

Jens Knipper Jens Knipper · Europe 2026 Virtual

3:32 min

Building evaluation frameworks for automated regression testing

Andreas Erben Andreas Erben · World Congress 2025

3:32 min

Fundamentals and limitations of large language models

Krzystof Czieslak · LIVE

3:26 min

Introducing LLMs as judges for automated testing

Sebastian Messingfeld Sebastian Messingfeld · World Congress 2026 Europe

36 sec

Frustrations with the complexity of modern web development

Andrew Taylor · Coffee With Developers

2:49 min

Executing automated evaluation suites leveraging LLMs as judges

Csenge Szabo Csenge Szabo · Europe 2026 Virtual

Videos

See all

Related articles

See all