> Markdown version of [/jobs/ext/2041111-ai-evaluation-engineer](https://www.wearedevelopers.com/jobs/ext/2041111-ai-evaluation-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Evaluation Engineer - **Company:** Mindrift - **Location:** Birmingham, UK - **Experience:** Expert - **Salary:** £104,000.0 - **Contract:** Permanent contract - **Skills:** JavaScript (Programming Language), Artificial Intelligence, Python (Programming Language), PostgreSQL, Redis, Software Engineering, ReactJS, Prompt Engineering, Fastapi, Apache Kafka, Codebase, Virtual Agents, Software Coding, Docker - **Published:** August 13, 2026 - **Apply:** https://www.careerjet.co.uk/jobad/gb1e514621356fe9acef094d3d77b5dc27 ## About the Role * 5+ years in software development * Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis * Experience writing tests (functional, integration) * English proficiency - B2+ ## Description Please submit your CV in English and indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria within realistic simulated environments: * Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history * Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent * Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient * Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT * Not data labeling * Not prompt engineering * Not writing code from scratch - the agent writes most of the code; you guide and evaluate, Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds. How it works Apply * Pass qualification(s) * Join a project * Complete tasks * Get paid ## Related Videos - [Watch Tests Go Brrrr! : Getting Started with Cypress in ReactJS](https://www.wearedevelopers.com/videos/282-watch-tests-go-brrrr-getting-started-with-cypress-in-reactjs) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Postgres in the Age of AI (and Devin)](https://www.wearedevelopers.com/videos/1042-postgres-in-the-age-of-ai-and-devin) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) ## Related Articles - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [WeAreDevelopers Dev Digest Issue 116 - The new search wars…](https://www.wearedevelopers.com/magazine/445-wearedevelopers-dev-digest-issue-116-the-new-search-wars)