> Markdown version of [/jobs/ext/2401390-ai-evaluation-engineer](https://www.wearedevelopers.com/jobs/ext/2401390-ai-evaluation-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Evaluation Engineer - **Company:** TXP - **Location:** London, UK - **Salary:** £182,000.0 - £195,000.0 - **Contract:** Temporary contract - **Skills:** Artificial Intelligence, Python (Programming Language), Software Engineering, Large Language Models, Prompt Engineering, Virtual Agents - **Published:** August 20, 2026 - **Apply:** https://www.totaljobs.com/job/evaluation-engineer/txp-job107875107 ## About the Role * Strong software engineering experience * Experience building and evaluating AI systems * Understanding of LLMs and agentic AI * Experience with AI evaluation frameworks * Ability to code independently * Strong problem-solving and communication skills * Comfortable working autonomously. Technical Experience AI/LLM Evaluation, RAG Evaluation, Ragas, Agentic AI Evaluation, Evaluation Harnesses, LLM Testing and Benchmarking, Prompt Engineering, Python Development, Model and Agent Integration. Ideal Background AI Evaluation Engineer, Applied AI Engineer, AI Engineer, LLM Engineer, Harness Engineer, Prompt Engineer or Software Engineer specialising in AI., Experience building evaluation tooling, evaluating LLMs, using Ragas, identifying AI failure modes, combining engineering with strategic thinking, and thriving in fast-moving environments. ## Description * Design evaluation frameworks and tooling * Develop evaluation harnesses * Evaluate model and agentic AI systems * Define metrics and methodologies * Build repeatable evaluation processes * Investigate AI failures * Work directly with government teams * Prototype and test solutions * Present findings and recommendations * Contribute to strategic direction. ## Related Videos - [Guiding Agentic AI with Vue](https://www.wearedevelopers.com/videos/2033-guiding-agentic-ai-with-vue) - [Same Words, Different Worlds: Who's in Control?](https://www.wearedevelopers.com/videos/2071-same-words-different-worlds-who-s-in-control) - [The Avengers Initiative (Practical Ethics for Software Engineers)](https://www.wearedevelopers.com/videos/2070-the-avengers-initiative-practical-ethics-for-software-engineers) - [Evals vs. Evil - AI and Package Security - Laurie Voss](https://www.wearedevelopers.com/videos/2131-evals-vs-evil-ai-and-package-security-laurie-voss) - [How to Stop Your Agents From Going Rogue - Arnav Gupta](https://www.wearedevelopers.com/videos/2152-how-to-stop-your-agents-from-going-rogue-arnav-gupta) - [Beyond the Benchmark: How to Evaluate AI Agents in the Real World](https://www.wearedevelopers.com/videos/100269-beyond-the-benchmark-how-to-evaluate-ai-agents-in-the-real-world) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Prompt Engineering is a Job of the Past](https://www.wearedevelopers.com/magazine/342-prompt-engineering-is-a-job-of-the-past) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)