Agentic QA Engineer - Generative AI & Multi-Agent Systems

ConglomerateIT LLC
Dallas, TX, United States
11 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Compensation
$104,000.0 - $116,480.0
Working hours
Regular working hours

Tech stack

Testing (Software) Application Programming Interfaces (APIs) Artificial Intelligence Microsoft Azure Encodings Continuous Integration Distributed Systems Github Python (Programming Language) Systems Development Life Cycle Queueing Systems Reliability Engineering
+8 more
Prometheus Datadog Large Language Models Grafana Multi-Agent Systems Generative AI Machine Learning Operations Virtual Agents

Job description

Agentic AI is already running in production here, and it needs someone who can prove it works. You’ll set the testing approach for multi-agent systems - how resiliency, accuracy, latency, orchestration correctness, and behavior at scale actually get measured - and hold that standard from first commit to production. Frameworks, harnesses, and the QA function itself are yours to build, working directly with the Agentic Operations group. This is a build-it-yourself role, not an oversight one. Where You’ll Spend Your Time Setting the Bar

  • The QA strategy for agentic and multi-agent systems is yours to define and defend across development, staging, and production
  • Coach QA engineers and put real testing standards, harness coding guidelines, and review practice in place
  • Work alongside Data Science, MLOps, and Platform teams to weave QA into the SDLC and into incident response

Testing the Agent Layer

  • Write the tests that exercise agent orchestration, tool calling, planner-executor loops, and coordination between agents - how tasks get decomposed, whether handoffs stay intact, whether the system converges on its goal
  • Check that state management, context windows, memory and knowledge stores, and prompt and graph correctness hold up as conditions shift
  • Prove the orchestrator does what it claims under DAG execution, retries, branching, timeouts, and compensation paths

Proving Correctness

  • Stand up ground-truth and reference pipelines that measure task accuracy through exact match, semantic similarity, and factuality checks
  • Create macro validation frameworks that can judge outcomes across multi-step agent workflows, generation-plus-verification loops included
  • Wire in guardrail checks for toxicity, PII, hallucination, and policy compliance

Breaking It on Purpose

  • Fuzz the scenarios: adversarial inputs, prompt perturbations, tool latency spikes, degraded APIs
  • Assemble resilience suites covering chaos experiments, failover, retry and backoff, circuit-breaking, and degraded-mode behavior
  • Establish latency SLOs, then measure end-to-end response time across LLM calls, tool invocations, and queues
  • Keep reliability honest with soak tests, canary verification, and automated rollback

Under Load

  • Build load and stress tests that push multi-agent graphs on concurrency, throughput, queue depth, and backpressure

Shipping Safely

  • Produce test artifacts other engineers can reuse - scenario configs, synthetic datasets, prompt libraries, agent graph fixtures, simulators
  • Push testing into CI/CD as pre-merge gates, nightly runs, and canaries, and into production monitoring with alerting tied to KPIs
  • Own release criteria and operational readiness across performance, security, compliance, and cost/latency budgets

Requirements

  • 7+ years across software QA and testing, 2+ of them on AI/ML or LLM-based systems, with agentic and multi-agent architectures you have personally tested
  • Python at a level where you’ve shipped your own test harnesses, simulators, and fixtures
  • Real LLM evaluation experience - exact and soft match, BLEU/ROUGE, BERTScore, embedding-based semantic similarity - alongside guardrail and prompt testing
  • Distributed systems testing chops: latency profiling, resiliency patterns including circuit breakers and retries, chaos engineering, message queues
  • Hands-on with at least one orchestration framework: LangChain, LangGraph, LlamaIndex, DSPy, OpenAI Assistants/Actions, Azure OpenAI orchestration, or something comparable
  • Comfortable inside CI/CD (GitHub Actions, Azure DevOps) and observability stacks (OpenTelemetry, Prometheus/Grafana, Datadog), plus feature flags and canaries
  • A working grip on privacy, security, and compliance for AI systems - PII handling, content policy, model safety
  • Communication and leadership range to hold your own across Operations, Data, and Engineering

About the company

  • MLOps depth - model versioning, dataset management, evaluation pipelines - plus A/B experimentation for LLMs
  • AWS, serverless, containerization, and event-driven architecture
  • You’ve carried cost, latency, or SLA targets for production AI workloads before

Founded in 2014, is a global leader in delivering innovative IT solutions and services. Headquartered in the USA with a presence in the UK, Canada, and India, we specialize in offering industry-leading expertise and cutting-edge products that help our clients maximize their technological investments. Our focus on best-in-class solutions, a highly knowledgeable team, and proactive talent mapping ensure we remain at the forefront of the IT industry. ConglomerateIT is driven by our Center for Excellence and Innovation, an initiative dedicated to keeping us ahead in a rapidly evolving technology landscape. We understand that building strong relationships is key to our success, and this commitment has enabled us to partner with Fortune 500 companies and leading system integrators worldwide. Our ability to provide local talent on a global scale ensures that we can meet the contingent project requirements of our clients efficiently and effectively.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

2:06 min

Elevating the QA engineering role for complex challenges

Ondřej Gróf Ondřej Gróf · World Congress 2026 Europe

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

3:45 min

Fusing developer experience and platform engineering for agentic SDLC

Julia Kordick Julia Kordick · World Congress 2026 Europe

Videos

See all

Related articles

See all