> Markdown version of [/jobs/ext/1947949-ai-evaluation-engineer-python-qa-or-security](https://www.wearedevelopers.com/jobs/ext/1947949-ai-evaluation-engineer-python-qa-or-security). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Evaluation Engineer (Python, QA or Security) - **Company:** Mindrift Alle - **Location:** München, Germany - **Experience:** Expert - **Salary:** €104,000.0 - **Contract:** Permanent contract - **Skills:** JavaScript (Programming Language), Artificial Intelligence, Python (Programming Language), PostgreSQL, Redis, Software Engineering, TypeScript, ReactJS, Prompt Engineering, Fastapi, Apache Kafka, Codebase, Virtual Agents, Software Coding, Docker - **Published:** August 6, 2026 - **Apply:** https://www.careerjet.de/jobad/dee3d87b23313d611f05079c4cd7b6f522 ## About the Role * 5+ years in software development * Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis * Experience writing tests (functional, integration) * English proficiency - B2+ ## Description Please submit your CV in English and indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria within realistic simulated environments: * Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history * Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent * Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient * Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT * Not data labeling * Not prompt engineering * Not writing code from scratch - the agent writes most of the code; you guide and evaluate, Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds. How it works Apply * Pass qualification(s) * Join a project * Complete tasks * Get paid ## Related Videos - [Watch Tests Go Brrrr! : Getting Started with Cypress in ReactJS](https://www.wearedevelopers.com/videos/282-watch-tests-go-brrrr-getting-started-with-cypress-in-reactjs) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Hiring AI Native Talents](https://www.wearedevelopers.com/videos/100268-hiring-ai-native-talents) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Testing AI Agents: Automated Evaluation for Chatbots & RAG Systems](https://www.wearedevelopers.com/videos/100300-testing-ai-agents-automated-evaluation-for-chatbots-rag-systems) ## Related Articles - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The Biggest German Tech Companies](https://www.wearedevelopers.com/magazine/424-the-biggest-german-tech-companies) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Best Coding Boot Camps in Germany](https://www.wearedevelopers.com/magazine/237-best-coding-boot-camps-in-germany)