Quality Engineer (AI/LLM Test Automation)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+14 more
Job description
We are looking for a Senior Quality Engineer / SDET with strong hands-on experience in AI/LLM testing and test automation. This role focuses on validating AI-powered applications where traditional pass/fail testing is not enough. You will build automated testing and LLM evaluation frameworks to determine whether AI-generated responses are accurate, reliable, relevant, and production-ready. What You’ll Do Design and build automated tests for AI/LLM-powered applications using Playwright and TypeScript/Python. Create and maintain LLM evaluation datasets and regression test suites. Define ground truth, expected behavior, scoring criteria, and pass/fail thresholds for AI-generated outputs. Evaluate nondeterministic LLM responses using heuristic/code-based evaluators, AI/LLM judges, and comparison-based evaluation. Validate accuracy, semantic correctness, relevance, groundedness, hallucinations, and consistency of AI responses. Test AI workflows including RAG, agents, prompts, tool/API calls, and multi-step workflows. Use LangSmith for tracing, datasets, experiments, evaluation, and analyzing AI behavior. Work with LangGraph or similar agent frameworks to understand and test agent workflows and execution paths. Investigate failures and determine whether an issue is caused by LLM/model behavior, prompt changes, retrieval, or an actual product defect. Build automated regression checks and integrate AI evaluations into CI/CD pipelines. Identify and fix flaky automated tests and continuously improve test reliability. Use AI tools such as Claude, GitHub Copilot, Cursor, or similar tools to accelerate test creation, debugging, and failure analysis.
Requirements
5+ years of QA / SDET / Test Automation experience. 1 2+ years of hands-on experience testing AI/LLM applications. Hands-on experience with LLM evaluation / AI evaluation. Experience creating or working with evaluation datasets and ground truth. Understanding of LLM judges / AI-as-a-judge, evaluators, scoring, and quality thresholds. Experience testing nondeterministic AI-generated outputs. Hands-on LangSmith experience. Hands-on LangGraph or agent workflow experience. Strong Playwright experience, preferably with TypeScript. Strong Python and/or TypeScript skills. API, integration, regression, and end-to-end testing experience. Experience integrating automated tests into CI/CD. Strong understanding of Git and Agile development.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this role…