AI/ML Test and Evaluation Engineer

Openteams, Inc.
Washington, DC, United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$145,000.0 - $250,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Intelligence Analysis Python (Programming Language) Machine Learning Open Source Technology Reference Data Tensorflow Software Engineering Pytorch Large Language Models Model Validation AI Platforms
+3 more
Information Technology HuggingFace Machine Learning Operations

Job description

We’re looking for a Senior AI/ML Test and Evaluation Engineer to build and operate the benchmarking and evaluation capability at the core of an AI platform. This is a role for someone who is more interested in what a model gets wrong than in what it gets right.

You build the evaluation harnesses - automated metrics paired with structured human expert judgment, applied to candidate models and to the agentic workflows built on top of them. You develop repeatable methodologies for comparing performance against current operational baselines, which means the comparison holds up when someone runs it again in six months with a different model. And you document the limitations and surface the failure modes that matter, including the ones nobody asked about.

Your reports go to senior stakeholders and inform decisions about which capabilities are ready to field. That’s the weight of the job: a benchmark that looks good and hides a failure mode is worse than no benchmark at all, and you’re the check against that., * Design, implement, and operate benchmark execution and evaluation harnesses for AI models and agentic workflows

  • Develop evaluation methodologies that combine automated metrics with structured human subject matter expert judgment
  • Curate and recommend candidate benchmarks based on mission needs and document the provenance of ground-truth and reference data
  • Produce defensible evaluation reports comparing candidate capabilities with current mission workflows, including documented limitations and failure modes
  • Define and contribute to common standards for benchmark expression, ingestion, and reporting
  • Support partner organizations and vendors as they integrate their capabilities with shared evaluation standards
  • Build lightweight expert-scoring workflows and measure inter-reviewer agreement for judgment-based evaluations
  • Participate in structured feedback sessions with mission end users and incorporate findings into the platform and evaluation methodology
  • Develop reference notebooks and example workflows that enable data-science-capable analysts to run, interpret, and extend evaluations
  • Document technical approaches, evaluation results, and key decisions for Government stakeholders and internal teams

Requirements

Clearance: An active U.S. security clearance is strongly preferred. Candidates without an active clearance may be considered for unclassified work but must be eligible to obtain and maintain a clearance., * U.S. citizenship and eligibility to obtain and maintain a U.S. security clearance

  • 6+ years of software engineering or machine learning engineering experience, including 3+ years evaluating, benchmarking, or deploying ML models in production or applied research environments
  • Strong Python proficiency in a machine learning or data science context
  • Hands-on experience with common ML frameworks and tooling, such as PyTorch and the Hugging Face ecosystem
  • Experience developing or using model evaluation harnesses, benchmark suites, or test and evaluation frameworks
  • Experience designing evaluation metrics and applying appropriate statistical rigor when interpreting and reporting results
  • Experience building repeatable and auditable evaluation pipelines with documented data provenance
  • Experience evaluating large language models or agentic workflows using task-based, metric-based, or judgment-based scoring
  • Strong written communication skills, including the ability to clearly explain evaluation methodologies, results, limitations, and failure modes to technical and nontechnical stakeholders
  • Ability to work effectively in an evolving environment and translate mission needs into practical evaluation approaches
  • Bachelor’s degree in computer science, mathematics, engineering, or a related field, or equivalent practical experience, * Prior AI/ML evaluation or test and evaluation experience supporting the Department of Defense, Intelligence Community, or another federal customer
  • Experience designing human-in-the-loop evaluations, measuring inter-rater reliability, or facilitating structured expert adjudication
  • Experience defining or implementing benchmark interchange formats or evaluation standards used across multiple organizations
  • Familiarity with intelligence analysis workflows or other high-stakes analytical domains
  • Experience working directly with Government stakeholders, mission users, or external technical partners
  • Contributions to open-source machine learning, benchmarking, or evaluation projects

Benefits & conditions

  • Medical, Dental & Vision - 100% paid for employees, 75% for dependents
  • 401(k) Match - Up to 5% with full vesting after 2 years
  • Unlimited PTO - With a required minimum of 15 days off annually
  • Fully Remote Setup - Includes up to $3,000 equipment reimbursement
  • Continuous Education - Includes up to $500 reimbursement
  • Disability & Life Insurance - 100% employer-paid
  • HSA & FSA Options - With monthly HSA contributions from OpenTeams

Grow With Us

At OpenTeams, growth isn’t just about the company-it’s about you.

We believe the best careers are built at the edge of your potential. That is where new tools, ideas, and technologies change the world. Here, you’ll work alongside pioneers of AI, solving problems that matter: making AI more transparent, more ethical, and more empowering. As your skills grow, our career framework provides a pathway and recognition of that increased impact.

Opportunities aren’t limited by geography. You’ll collaborate with global experts, contribute to open source projects that power the world’s technology, and stretch your skills daily. That global perspective and diversity makes our solution more universal and robust. We are committed to continuing to celebrate diversity on our team.

Supported people are successful people. We offer 100% employer paid medical premiums for employees and self-managed PTO with a minimum time off requirement, so that our teams are able to do their best work.

We invest in curiosity, creativity, and ownership. That means you’ll be trusted to boldly innovate, supported to learn fast, and celebrated for successful collaboration.

About the company

Every organization runs on intelligence: years of accumulated knowledge, decisions, and context. As AI takes on more of that work, companies face a choice: rent that intelligence from vendors who keep the data, the context, and the results, or own it.

OpenTeams exists to make ownership possible.

Founded by Travis Oliphant, creator of NumPy and SciPy, and built by people with deep roots across the open-source ecosystem, including NumPy, SciPy, PyTorch, and Jupyter, we help enterprises and governments build AI they control, govern, and evolve themselves.

If that sounds like your kind of work, we’d like to meet you.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:12 min

Navigating technical clarity as a global black belt

Chris Heilmann +2 · LIVE

1:39 min

Fundamentals of tensors and the TensorFlow library

Håkan Silfvernagel · LIVE

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

3:55 min

Evaluating central server APIs against edge deployment models

Hauke Brammer · World Congress 2021

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

1:16 min

Adjusting technical interviews for an AI-native industry

Ayotunde Obasa Ayotunde Obasa · Europe 2026 Virtual

Videos

See all

Related articles

See all