AI Test Engineer

Governancereply
UK
6 days ago
Apply on www.apply4u.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Automation of Tests Encodings Python (Programming Language) NumPy Data Driven Tests Software Engineering Software Testing Automation Framework Test Data BLEU Score Feature Engineering Pytorch
+9 more
Evaluation Pipelines Large Language Models Reliability of Systems Generative AI Agentic-AI Pandas Pytest Playwright Data Pipelines

Job description

work across AI engineering, data pipelines, and quality automation, contributing both to architecture design and hands-on implementation of scalable testing and evaluation systems. This role is suited for someone who can operate independently, contribute to technical decisions, and help shape standards for AI system validation in production environments.Responsibilities:Design and develop Python-based frameworks for testing and evaluating AI agents and LLM-based systemsContribute to architecture design for evaluation pipelines, observability, and data-driven testing systemsBuild and maintain tools for test data generation (including synthetic and adversarial datasets)Define and implement evaluation strategies and quality KPIs for AI behavior (accuracy, robustness, bias, consistency, hallucination rate, etc.)Integrate LLM evaluation tools, CI/CD pipelines, and external AI platformsSupport continuous testing, benchmarking, and production monitoring of AI systemsCollaborate with AI engineers

Requirements

data engineers, and QA teams to improve system reliability and scalabilityAbout the candidate:2-5+ years of experience in Software Engineering, Test Automation, AI Engineering, or Data EngineeringStrong proficiency in Python and data/AI libraries (e.g., pandas, numpy, PyTorch or similar)Solid understanding of LLMs, NLP concepts, or generative AI systemsExperience with test automation frameworks (e.g., PyTest, Playwright) and CI/CD pipelinesFamiliarity with data pipelines, dataset design, or feature engineering is a strong plusExperience with evaluation metrics for NLP/AI systems (BLEU, ROUGE, embedding-based metrics, or custom scoring approaches)Ability to design scalable systems and work autonomously on complex technical problemsStrong analytical mindset and interest in AI system quality, reliability, and governanceReply is an Equal Opportunities Employer and committed to embracing diversity in the workplace. We provide equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type regardless of age, sexual orientation, gender, identity, pregnancy, religion, nationality, ethnic origin, disability, medical history, skin colour, marital status or parental status or any other characteristic protected by the Law. Reply is committed to making sure that our selection methods are fair to everyone. To help you during the recruitment process, please let us know of any Reasonable Adjustments you may need.Requisition ID11321- Posted10/02/#####-Technology-Job-Years of Experience (2)3-5 , 1-3Where (1)United Kingdom

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.apply4u.co.uk
Prepare application

Good distractions

Loading talks and stories from around this role…