Agentic QA Engineer - Generative AI & Multi-Agent Systems
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+8 more
Job description
Agentic AI is already running in production here, and it needs someone who can prove it works. You’ll set the testing approach for multi-agent systems - how resiliency, accuracy, latency, orchestration correctness, and behavior at scale actually get measured - and hold that standard from first commit to production. Frameworks, harnesses, and the QA function itself are yours to build, working directly with the Agentic Operations group. This is a build-it-yourself role, not an oversight one. Where You’ll Spend Your Time Setting the Bar
- The QA strategy for agentic and multi-agent systems is yours to define and defend across development, staging, and production
- Coach QA engineers and put real testing standards, harness coding guidelines, and review practice in place
- Work alongside Data Science, MLOps, and Platform teams to weave QA into the SDLC and into incident response
Testing the Agent Layer
- Write the tests that exercise agent orchestration, tool calling, planner-executor loops, and coordination between agents - how tasks get decomposed, whether handoffs stay intact, whether the system converges on its goal
- Check that state management, context windows, memory and knowledge stores, and prompt and graph correctness hold up as conditions shift
- Prove the orchestrator does what it claims under DAG execution, retries, branching, timeouts, and compensation paths
Proving Correctness
- Stand up ground-truth and reference pipelines that measure task accuracy through exact match, semantic similarity, and factuality checks
- Create macro validation frameworks that can judge outcomes across multi-step agent workflows, generation-plus-verification loops included
- Wire in guardrail checks for toxicity, PII, hallucination, and policy compliance
Breaking It on Purpose
- Fuzz the scenarios: adversarial inputs, prompt perturbations, tool latency spikes, degraded APIs
- Assemble resilience suites covering chaos experiments, failover, retry and backoff, circuit-breaking, and degraded-mode behavior
- Establish latency SLOs, then measure end-to-end response time across LLM calls, tool invocations, and queues
- Keep reliability honest with soak tests, canary verification, and automated rollback
Under Load
- Build load and stress tests that push multi-agent graphs on concurrency, throughput, queue depth, and backpressure
Shipping Safely
- Produce test artifacts other engineers can reuse - scenario configs, synthetic datasets, prompt libraries, agent graph fixtures, simulators
- Push testing into CI/CD as pre-merge gates, nightly runs, and canaries, and into production monitoring with alerting tied to KPIs
- Own release criteria and operational readiness across performance, security, compliance, and cost/latency budgets
Requirements
- 7+ years across software QA and testing, 2+ of them on AI/ML or LLM-based systems, with agentic and multi-agent architectures you have personally tested
- Python at a level where you’ve shipped your own test harnesses, simulators, and fixtures
- Real LLM evaluation experience - exact and soft match, BLEU/ROUGE, BERTScore, embedding-based semantic similarity - alongside guardrail and prompt testing
- Distributed systems testing chops: latency profiling, resiliency patterns including circuit breakers and retries, chaos engineering, message queues
- Hands-on with at least one orchestration framework: LangChain, LangGraph, LlamaIndex, DSPy, OpenAI Assistants/Actions, Azure OpenAI orchestration, or something comparable
- Comfortable inside CI/CD (GitHub Actions, Azure DevOps) and observability stacks (OpenTelemetry, Prometheus/Grafana, Datadog), plus feature flags and canaries
- A working grip on privacy, security, and compliance for AI systems - PII handling, content policy, model safety
- Communication and leadership range to hold your own across Operations, Data, and Engineering
About the company
- MLOps depth - model versioning, dataset management, evaluation pipelines - plus A/B experimentation for LLMs
- AWS, serverless, containerization, and event-driven architecture
- You’ve carried cost, latency, or SLA targets for production AI workloads before
Founded in 2014, is a global leader in delivering innovative IT solutions and services. Headquartered in the USA with a presence in the UK, Canada, and India, we specialize in offering industry-leading expertise and cutting-edge products that help our clients maximize their technological investments. Our focus on best-in-class solutions, a highly knowledgeable team, and proactive talent mapping ensure we remain at the forefront of the IT industry. ConglomerateIT is driven by our Center for Excellence and Innovation, an initiative dedicated to keeping us ahead in a rapidly evolving technology landscape. We understand that building strong relationships is key to our success, and this commitment has enabled us to partner with Fortune 500 companies and leading system integrators worldwide. Our ability to provide local talent on a global scale ensures that we can meet the contingent project requirements of our clients efficiently and effectively.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?
MLOps And AI Driven Development