> Markdown version of [/jobs/ext/2271400-ai-evaluation-engineer](https://www.wearedevelopers.com/jobs/ext/2271400-ai-evaluation-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Evaluation Engineer - **Company:** Acunor Infotech - **Location:** United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Automated Storage and Retrieval Systems, Regression Testing, Large Language Models, Build Management - **Published:** August 27, 2026 - **Apply:** https://www.dice.com/job-detail/08e5b0e1-7c4a-425a-bc66-62aaf36f7b7d ## About the Role * 4+ years of experience in ML/LLM evaluation, AI quality engineering, applied research engineering, or related areas. * Hands-on experience designing and implementing AI/LLM evaluation frameworks. * Experience developing golden/reference datasets and measurable evaluation criteria. * Strong understanding of LLM-as-a-judge methodologies and associated failure modes. * Experience evaluating LLM applications, agents, RAG systems, prompts, or other probabilistic AI systems. * Strong statistical and analytical skills, including evaluation methodologies for relatively small sample sizes. * Experience establishing quality thresholds and regression criteria. * Strong engineering skills with the ability to build reusable evaluation tooling and automation. * Ability to independently assess system quality and challenge release decisions when evaluation evidence does not meet established standards., * Healthcare, clinical, or other safety-critical AI evaluation experience. * Experience designing medical-accuracy or domain-specific evaluation suites. * AI red-teaming or adversarial testing experience. * Experience with production AI monitoring, drift detection, and continuous evaluation. ## Description We are seeking an AI Evaluation Engineer to establish and maintain the quality standards used to determine whether production AI and LLM systems are ready to launch. This role will build centralized evaluation frameworks, create golden datasets, establish measurable quality thresholds, audit AI evaluation results, and implement continuous production monitoring. The successful candidate will bring strong technical evaluation expertise along with the independence and analytical rigor required to objectively determine whether AI systems meet production quality and safety standards., * Design and build centralized LLM/AI evaluation frameworks and reusable evaluation templates. * Establish evaluation standards that can be adopted consistently across multiple AI engineering teams. * Create and maintain golden datasets in partnership with business, domain, and subject-matter experts. * Develop evaluation suites covering functional quality, accuracy, safety, reliability, and domain-specific requirements. * Implement and assess LLM-as-a-judge evaluation approaches while accounting for their limitations and failure modes. * Define measurable pass/fail thresholds and production-readiness criteria for AI applications. * Independently review and audit evaluation suites developed by individual engineering teams. * Build regression testing approaches for LLM applications, agents, prompts, retrieval systems, and model changes. * Establish production monitoring for AI quality degradation, drift, regressions, and incidents. * Analyze and report evaluation pass rates, quality trends, regressions, and production incidents. * Partner with AI engineers, platform engineers, product teams, and domain experts to continuously improve AI quality. ## Related Videos - [Answering the Million Dollar Question: Why did I Break Production?](https://www.wearedevelopers.com/videos/1171-answering-the-million-dollar-question-why-did-i-break-production) - [Trunk-Based Development at Scale: Real-World Insights from a High-Traffic Luxury E-Commerce Platform](https://www.wearedevelopers.com/videos/1435-trunk-based-development-at-scale-real-world-insights-from-a-high-traffic-luxury-e-commerce-platform) - [Creating Industry ready solutions with LLM Models](https://www.wearedevelopers.com/videos/899-creating-industry-ready-solutions-with-llm-models) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [You are not my model anymore - understanding LLM model behavior](https://www.wearedevelopers.com/videos/1464-you-are-not-my-model-anymore-understanding-llm-model-behavior) - [Bringing the power of AI to your application.](https://www.wearedevelopers.com/videos/1010-bringing-the-power-of-ai-to-your-application) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [How to Use Generative AI to Accelerate Learning to Code](https://www.wearedevelopers.com/magazine/530-how-to-use-generative-ai-to-accelerate-learning-to-code) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)