AI Test Automation Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+1 more
Job description
- Agent Test Design & Automation: Design and implement deterministic pytest suites to validate agent behavior, intent handling, tool-call correctness, strict JSON schemas, and failure paths .
- LLM Evaluation Engineering: Build and maintain LLM-as-a-judge evaluation scripts that score agent outputs for correctness, tone, and compliance, defining scoring thresholds and pass/fail release gates .
- Mocking & Trace Analysis: Build mocks and fixtures for LLM endpoints and core APIs so agent loops run repeatably in CI pipelines . Execute, mock, and capture traces from LangGraph or Microsoft Agent Framework agent loops within test suites .
- Guardrail & Security Testing: Continuously validate that agents refuse prompt injection attempts, do not hallucinate actions, and strictly follow compliance guardrails .
- CI/CD Integration & Reporting: Integrate test suites into GitLab CI with DevOps, leverage OPIK for trace analysis, and deliver automated quality evidence gating weekly drops .
Requirements
- Automation Mastery: 4+ years of test automation experience using Python and pytest (custom fixtures, mocking, and CI pipeline integration) .
- LLM & Agentic System Testing: 1+ years of hands-on experience evaluating non-deterministic AI applications .
- LLM Evaluation Engineering: Proven track record building LLM-as-a-judge scoring scripts and maintaining versioned evaluation datasets (golden paths, edge cases, regression suites) .
- Agent Deep-Dives & Traces: Ability to validate tool selection and parameters, inspect multi-step agent execution traces (e.g., LangGraph, Microsoft Agent Framework), and mock agent endpoints for CI workflows .
- Adversarial & Guardrail Testing: Experience designing negative-path test suites for prompt injections, hallucinated financial actions, and trust-level boundaries (Suggest / Approve / Act) .
- Observability & CI Integration: Hands-on integration of test stages into GitLab CI (or equivalent) using LLM observability tools like OPIK or similar trace analysis platforms ., * Test Automation (Essential): 4+ years of test automation with strong Python and pytest . Experience testing APIs, contract testing, and strict JSON schema validation .
- AI & LLM Testing (Essential): 1+ years of experience with behavioral acceptance criteria, evaluation datasets, and LLM-as-a-judge techniques . Practical understanding of managing non-determinism via statistical thresholds and automated judges .
- Frameworks (Essential): Hands-on familiarity with agentic frameworks (LangGraph, Microsoft Agent Framework, LangChain, or LlamaIndex) sufficient to execute and mock agent loops .
- Observability & Domain (Desirable): Experience with OPIK or equivalent trace analysis platforms . Prior experience testing in banking, fintech, or regulated financial environments .
- Knowledge of Arabic language is plus.
Benefits & conditions
Final compensation offered is based on multiple factors such as the specific role, hiring location, as well as individual skills, experience, and qualifications. In addition to competitive salaries, we offer a comprehensive benefits package. Learn more about life at Globant here: Globant Experience Guide .
About the company
At Globant, we are working to make the world a better place, one step at a time. We enhance business development and enterprise solutions to prepare them for a digital future. With a diverse and talented team present in more than 30 countries, we are strategic partners to leading global companies in their business process transformation.
Join Globant as an AI Test Automation Engineer and lead the quality transformation for our flagship agentic banking platform . We are looking for an automation expert skilled in Python, pytest, and LLM/Agentic system testing to solve one of tech’s newest challenges: systematically validating non-deterministic AI behavior . In this role, you will build automated evaluation datasets, inspect agent execution traces, and deploy LLM-as-a-judge frameworks to gate production releases with absolute confidence .
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this role…