AI Evaluation Engineer - SC Eligible - 4 Months - London, Bristol or Manchester
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
A major public sector organisation is seeking an experienced AI Evaluation Engineer to join a fast-paced and highly innovative technology team shaping the future direction of AI adoption across government. Working within a small, ambitious team operating with a start-up mindset, you will be responsible for designing, building and implementing robust AI evaluation frameworks, assessing the quality and performance of Large Language Models (LLMs), agentic AI solutions and emerging AI technologies.
This role requires a hands-on engineer who can build solutions from the ground up, develop evaluation harnesses, create automated testing frameworks, and confidently communicate outcomes to both technical and non-technical stakeholders. You will work across multiple government and cross-functional teams, helping identify opportunities, solve complex challenges, and establish best practices for AI evaluation and assurance., * Design, build and maintain AI evaluation frameworks for Large Language Models (LLMs) and agentic AI systems
- Develop automated evaluation pipelines, testing harnesses and quality assurance methodologies for AI solutions
- Measure model performance using quantitative and qualitative evaluation techniques
- Implement and utilise tools such as Ragas, DeepEval and other AI evaluation frameworks
- Build proof of concepts and production-ready Python-based solutions from the ground up
- Assess AI system outputs, identify weaknesses and recommend improvements to model performance
- Collaborate with multidisciplinary teams across government to understand business challenges and AI opportunities
- Present findings, recommendations and technical solutions to a range of stakeholders
- Provide technical leadership and contribute to strategic decision-making around AI adoption and evaluation
- Work independently while maintaining accountability and ownership of deliverables in a fast-moving environment
Requirements
The successful candidate will have strong software engineering experience, a deep understanding of AI model and agentic layers, and practical experience with AI evaluation tools such as Ragas, DeepEval, or similar frameworks. Strong communication, leadership and stakeholder engagement skills are essential, alongside the ability to thrive in environments with ambiguity and rapid change., * Strong commercial experience developing AI, Machine Learning or Generative AI solutions
- Proven experience designing and implementing AI evaluation frameworks and testing methodologies
- Hands-on software engineering expertise, particularly using Python
- Experience working with LLMs, AI model layers and agentic AI architectures
- Practical knowledge of tools such as Ragas, DeepEval, Promptfoo, LangSmith or similar evaluation platforms
- Demonstrable experience building AI harnesses, experimentation environments or benchmarking solutions
- Strong stakeholder management and communication skills, with the ability to present technical concepts clearly
- Comfortable operating in ambiguous environments and helping shape technical direction
- Experience working across multiple teams and influencing wider technical communities
- Public sector, government or highly regulated industry experience would be beneficial but is not essential, * Must be eligible for SC Clearance and meet Civil Service nationality requirements
- Active or transferable SC Clearance highly desirable
- Candidates can start on BPSS clearance while SC clearance progresses if required
- Hybrid working with attendance expected approximately 2 days per week in London, Bristol or Manchester
- Flexibility required for workshops, team days and stakeholder events across UK locations
- Must be comfortable working in a dynamic, evolving environment with changing priorities
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
Navigating the AI Shift
Data Engineer Salary UK
Dev Digest 121 - AI goes offline