AI Evaluation Infrastructure & Production Readiness

Ibotix Us Inc.
Charlotte, NC, United States
12 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Software Quality Customer Data Management Statistical Hypothesis Testing Machine Learning Regression Testing Management of Software Versions Large Language Models Model Validation Machine Learning Operations

Job description

AI quality cannot be proven once at launch it is an ongoing discipline. Because AI systems are probabilistic, quality can shift as models, prompts, content, vendors, retrieval, and user behavior change. This role builds the evaluation discipline and tooling that produces repeatable evidence for whether an AI system is accurate, safe, compliant, grounded, and ready to scale. It is the enterprise’s most important AI risk-control role.

What You’ll Build :

  • Evaluation harnesses and tooling that can be run repeatedly across AI systems.
  • Golden test sets and scenario libraries covering expected behaviors and edge cases.
  • Regression testing to catch quality changes when models, prompts, or content change.
  • Hallucination testing and grounding/faithfulness testing.
  • Bias and fairness testing to detect discriminatory or disparate-impact outcomes across protected classes, in support of fair-lending obligations (e.g., ECOA / Reg B).
  • Adversarial and red-team testing, including jailbreak, prompt-injection, and harmful-output resistance.
  • Policy-adherence testing against enterprise, compliance, and regulatory requirements.
  • Drift detection and ongoing production (online) monitoring sampling live traffic, canary evaluations, and catching regressions after release.
  • Human and subject-matter-expert evaluation workflows, plus curation, labeling, and versioning of evaluation datasets (including governance of any customer data they contain).
  • Production-readiness gates that AI systems must pass before scaling.
  • Evaluation evidence and reporting that AI governance and model-risk committees rely on to make go/no-go decisions.
  • AI vendor acceptance criteria for third-party agents and solutions.

Why This Role Matters : Without a rigorous evaluation function, AI scales on the strength of demos, pilots, anecdotes, or vendor claims rather than repeatable evidence. This role creates the evidence system leadership needs to decide, with confidence, whether an AI system is ready. Evaluations apply equally to agents built in-house and to third-party agents delivered by vendors.

Requirements

  • 7+ years of relevant experience in software quality, data science, machine learning, or a closely related field, including experience leading a technical workstream.
  • Demonstrated experience designing evaluation methods for LLM or ML systems (accuracy, grounding, safety, and policy adherence).
  • Strong grasp of testing methodology: test-set design, regression testing, and metrics that hold up over time.
  • Statistical rigor: experimental design, significance testing, and sample sizing.
  • Experience with bias and fairness evaluation; familiarity with fair-lending concepts (ECOA / Reg B,disparate impact) is a strong plus.
  • Ability to define and enforce production-readiness gates and acceptance criteria.
  • Comfort partnering with risk, compliance, and legal to translate requirements into testable criteria., * Experience building evaluation harnesses, LLM-as-judge pipelines, or automated eval frameworks.
  • Familiarity with hallucination, faithfulness, and drift-detection techniques.
  • Experience setting vendor acceptance criteria and evaluating third-party AI solutions.
  • Model risk management or model validation background (e.g., SR 11-7-aligned practices).
  • Background in regulated financial services.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:39 min

Evaluating current enterprise artificial intelligence production maturity

Christopher May Christopher May · Europe 2026 Virtual

1:48 min

Balancing code generation velocity with software quality standards

Lilia Gargouri Lilia Gargouri · Coffee With Developers

2:36 min

Applying supervised machine learning for practical rule extraction

Katja Träumner

5:08 min

Validating requests and responses using data transfer objects

Roman Alexis Anastasini · World Congress 2021

2:28 min

Introduction to building reliable AI agents in production

Max Tkacz Max Tkacz · World Congress 2025

1:57 min

Building a holistic system for software quality

Serge Baumberger Serge Baumberger · Europe 2026 Virtual

Videos

See all

Related articles

See all