AI Evaluations Engineer

ConnexAI
Manchester, UK
10 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Data Analysis Automation of Tests Python (Programming Language) Regression Testing Scripting Large Language Models Machine Learning Operations

Job description

This role sits at the centre of how we measure and improve AI systems in production.

You’ll define what good performance means across LLMs, ASR, TTS, and full speech-to-speech pipelines, and build the datasets, metrics, and evaluation systems that make AI quality measurable and comparable in the real world.

You’ll work closely with engineering and product teams to ensure model changes lead to real improvements in user experience, not just better offline benchmarks.

What you’ll do

  • Design and run evaluations across LLM, ASR, TTS, and speech-to-speech systems
  • Build real-world datasets and test cases from production behaviour and edge cases
  • Define metrics and scorecards for model and system quality
  • Benchmark internal models against external and frontier systems
  • Build Python tools to automate evaluation workflows
  • Create internal leaderboards, red-teaming setups, and regression tests
  • Work with engineers and product teams to diagnose system failures
  • Turn vague product goals into measurable evaluation frameworks

What this role is about

  • Defining and measuring AI quality in production systems
  • Turning real user behaviour into structured evaluation signals
  • Ensuring model changes improve real-world performance
  • Understanding why AI systems fail, not just whether they do

What good looks like

  • You can translate improved quality into measurable metrics
  • You think in terms of system impact (before vs after), not just accuracy
  • You’re comfortable working across code, data, and production systems
  • You care about real-world behaviour, not just benchmarks

Core skills

  • Strong Python (scripting, data analysis, tooling)
  • Experience with ML systems, evaluation, or experimentation
  • Understanding of LLMs or speech systems (ASR / TTS)
  • Ability to design test cases and structured datasets
  • Comfortable working with engineers and product teams

Nice to have

  • Experience with LLM evaluation or benchmarking
  • Exposure to speech or multimodal systems
  • Familiarity with production APIs or ML systems
  • Experience with automated testing or CI-style workflows

Requirements

  • You think in terms of system impact (before vs after), not just accuracy
  • You’re comfortable working across code, data, and production systems
  • You care about real-world behaviour, not just benchmarks

Core skills

  • Strong Python (scripting, data analysis, tooling)
  • Experience with ML systems, evaluation, or experimentation
  • Understanding of LLMs or speech systems (ASR / TTS)
  • Ability to design test cases and structured datasets
  • Comfortable working with engineers and product teams, * Experience with LLM evaluation or benchmarking
  • Exposure to speech or multimodal systems
  • Familiarity with production APIs or ML systems
  • Experience with automated testing or CI-style workflows

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

1:48 min

Automating exploratory data analysis within training pipelines

Dora Petrella · World Congress 2023

3:32 min

Building evaluation frameworks for automated regression testing

Andreas Erben Andreas Erben · World Congress 2025

1:16 min

Adjusting technical interviews for an AI-native industry

Ayotunde Obasa Ayotunde Obasa · Europe 2026 Virtual

1:53 min

Evaluating traditional scripting languages for modern development tasks

Jens Knipper Jens Knipper · Europe 2026 Virtual

1:55 min

Essential resources for understanding agentic design and evaluation

Alfonso Graziano Alfonso Graziano · Coffee With Developers

Videos

See all

Related articles

See all