> Markdown version of [/jobs/ext/2490227-ai-qa-tester](https://www.wearedevelopers.com/jobs/ext/2490227-ai-qa-tester). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI QA Tester - **Company:** Lorien - **Location:** London, UK (Remote available) - **Contract:** Temporary contract - **Skills:** JavaScript (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, Application Performance Management, Automation of Tests, Microsoft Azure, C Sharp (Programming Language), Continuous Integration, Document-Oriented Databases, Github, Apache JMeter, Python (Programming Language), Load Testing, Machine Learning, Open Web Application Security, Kusto Query Language, Strategies of Testing, TypeScript, Scripting, Performance Testing, Postman, Large Language Models, Concurrency, Pytest, Playwright, Data Management, Api Management - **Published:** August 21, 2026 - **Apply:** https://www.careerboard.com/pt/en/find-jobs-in-United-Kingdom/-346AC99B8CD9CE8C25/ ## About the Role * Strong QA engineering background with hands-on test automation, not manual Scripting alone. * Automation coding ability in Python, C# and/or JavaScript/TypeScript. * API testing experience, including contract testing, authentication, negative testing and mocked dependencies. * Demonstrable experience testing AI, machine learning or LLM-based systems, or a clear grasp of how to assure non-deterministic output using evaluation metrics and statistical thresholds. * Understanding of RAG architectures and where they fail - retrieval misses, stale indexes, chunk boundary loss, hallucination, ungrounded citations. * Awareness of LLM security risks (for example the OWASP Top 10 for LLM applications) and practical adversarial testing technique. * CI/CD integration of automated test suites (GitHub Actions or Azure DevOps). * Test data management, including handling of personal data safely in test environments. * Confidence challenging engineers and stakeholders on release readiness. * Experience with Azure AI Foundry evaluations, Azure AI Content Safety testing or PyRIT/other red-teaming tooling. * Performance and load testing tools (k6, JMeter, Azure Load Testing). * Azure Monitor/Application Insights and KQL for investigating failures from telemetry. * Accessibility and conversational UX testing experience. * Familiarity with UK GDPR, DPIA processes or regulated-sector assurance evidence. ## Description * Own and maintain the test strategy for the agentic platform, spanning the agent orchestration layer, the data APIs, the retrieval pipeline and the underlying Azure services. * Build and curate golden test datasets and expected-outcome sets with input from business subject-matter experts, covering happy paths, edge cases, ambiguous questions and out-of-scope requests. * Design and automate LLM evaluation: groundedness/faithfulness, answer relevance, retrieval precision and recall, citation correctness, task completion, tone and consistency - using Azure AI Foundry evaluation, LLM-as-judge and deterministic checks as appropriate. * Wire evaluations into CI/CD as release gates, so prompt, model, index or code changes cannot regress quality unnoticed; report trends over time. * Carry out adversarial and red-team testing: prompt injection, jailbreaks, indirect injection via knowledge-base content, data-exfiltration attempts, tool misuse and privilege escalation through agent tool calls. * Validate the safety and monitoring controls - Azure AI Content Safety categories and thresholds, refusal and escalation behaviour, and Azure AI Anomaly Detector signals - including deliberate false-negative and false-positive probing. * Test authorisation rigorously: confirm that agents and APIs never return member or document data outside the requesting user's entitlement, including via retrieval or summarisation side channels. * Build automated functional and contract test suites for the Agent Data API and Member Data API (for example pytest, Playwright, Postman/Newman, RestSharp) and for the end-to-end conversational flows. * Run non-functional testing: latency and response-time budgets, throughput and concurrency, token cost per interaction, rate-limit and failover behaviour, and resilience when BCUK or a dependency degrades. ## Related Videos - [Are Your APIs Ready for AI Agents](https://www.wearedevelopers.com/videos/2004-are-your-apis-ready-for-ai-agents) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [pytest: Simple, rapid and fun testing with Python](https://www.wearedevelopers.com/videos/213-pytest-simple-rapid-and-fun-testing-with-python) - [AI as a Test Designer: Transforming Experience into Automated Testing](https://www.wearedevelopers.com/videos/1984-ai-as-a-test-designer-transforming-experience-into-automated-testing) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [Testing AI Agents: Automated Evaluation for Chatbots & RAG Systems](https://www.wearedevelopers.com/videos/100300-testing-ai-agents-automated-evaluation-for-chatbots-rag-systems) ## Related Articles - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [13 AI Tools for Developers](https://www.wearedevelopers.com/magazine/302-13-ai-tools-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)