> Markdown version of [/jobs/ext/2987730-ai-evaluation-infrastructure-production-readiness](https://www.wearedevelopers.com/jobs/ext/2987730-ai-evaluation-infrastructure-production-readiness). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Evaluation Infrastructure & Production Readiness - **Company:** Ibotix Us Inc. - **Location:** Charlotte, NC, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Software Quality, Customer Data Management, Statistical Hypothesis Testing, Machine Learning, Regression Testing, Management of Software Versions, Large Language Models, Model Validation, Machine Learning Operations - **Published:** September 18, 2026 - **Apply:** https://www.dice.com/job-detail/a8fdc916-c409-438d-8e35-37694a71f05c ## About the Role * 7+ years of relevant experience in software quality, data science, machine learning, or a closely related field, including experience leading a technical workstream. * Demonstrated experience designing evaluation methods for LLM or ML systems (accuracy, grounding, safety, and policy adherence). * Strong grasp of testing methodology: test-set design, regression testing, and metrics that hold up over time. * Statistical rigor: experimental design, significance testing, and sample sizing. * Experience with bias and fairness evaluation; familiarity with fair-lending concepts (ECOA / Reg B,disparate impact) is a strong plus. * Ability to define and enforce production-readiness gates and acceptance criteria. * Comfort partnering with risk, compliance, and legal to translate requirements into testable criteria., * Experience building evaluation harnesses, LLM-as-judge pipelines, or automated eval frameworks. * Familiarity with hallucination, faithfulness, and drift-detection techniques. * Experience setting vendor acceptance criteria and evaluating third-party AI solutions. * Model risk management or model validation background (e.g., SR 11-7-aligned practices). * Background in regulated financial services. ## Description AI quality cannot be proven once at launch it is an ongoing discipline. Because AI systems are probabilistic, quality can shift as models, prompts, content, vendors, retrieval, and user behavior change. This role builds the evaluation discipline and tooling that produces repeatable evidence for whether an AI system is accurate, safe, compliant, grounded, and ready to scale. It is the enterprise's most important AI risk-control role. What You'll Build : * Evaluation harnesses and tooling that can be run repeatedly across AI systems. * Golden test sets and scenario libraries covering expected behaviors and edge cases. * Regression testing to catch quality changes when models, prompts, or content change. * Hallucination testing and grounding/faithfulness testing. * Bias and fairness testing to detect discriminatory or disparate-impact outcomes across protected classes, in support of fair-lending obligations (e.g., ECOA / Reg B). * Adversarial and red-team testing, including jailbreak, prompt-injection, and harmful-output resistance. * Policy-adherence testing against enterprise, compliance, and regulatory requirements. * Drift detection and ongoing production (online) monitoring sampling live traffic, canary evaluations, and catching regressions after release. * Human and subject-matter-expert evaluation workflows, plus curation, labeling, and versioning of evaluation datasets (including governance of any customer data they contain). * Production-readiness gates that AI systems must pass before scaling. * Evaluation evidence and reporting that AI governance and model-risk committees rely on to make go/no-go decisions. * AI vendor acceptance criteria for third-party agents and solutions. Why This Role Matters : Without a rigorous evaluation function, AI scales on the strength of demos, pilots, anecdotes, or vendor claims rather than repeatable evidence. This role creates the evidence system leadership needs to decide, with confidence, whether an AI system is ready. Evaluations apply equally to agents built in-house and to third-party agents delivered by vendors. ## Related Videos - [Introduction to Responsible AI: Balancing Value and Risk](https://www.wearedevelopers.com/videos/1972-introduction-to-responsible-ai-balancing-value-and-risk) - [Let's Talk Quality! - Lilia Gargouri](https://www.wearedevelopers.com/videos/1815-let-s-talk-quality-lilia-gargouri) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Beyond the Benchmark: How to Evaluate AI Agents in the Real World](https://www.wearedevelopers.com/videos/100269-beyond-the-benchmark-how-to-evaluate-ai-agents-in-the-real-world) - [Let’s Talk Quality!](https://www.wearedevelopers.com/videos/100012-let-s-talk-quality) - [AI in Regulated Industry - Validating AI-Enabled Products with PLM and Digital Twins](https://www.wearedevelopers.com/videos/2065-ai-in-regulated-industry-validating-ai-enabled-products-with-plm-and-digital-twins) ## Related Articles - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [The State of WebDev AI 2025 Results: What Can We Learn?](https://www.wearedevelopers.com/magazine/581-the-state-of-webdev-ai-2025-results-what-can-we-learn) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [13 AI Tools You Have to Try](https://www.wearedevelopers.com/magazine/219-13-ai-tools-you-have-to-try)