> Markdown version of [/jobs/ext/2847239-agentic-qa-engineer-generative-ai-multi-agent-systems](https://www.wearedevelopers.com/jobs/ext/2847239-agentic-qa-engineer-generative-ai-multi-agent-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Agentic QA Engineer - Generative AI & Multi-Agent Systems - **Company:** ConglomerateIT LLC - **Location:** Dallas, TX, United States - **Experience:** Expert - **Salary:** $104,000.0 - $116,480.0 - **Contract:** Temporary to permanent - **Skills:** Testing (Software), Application Programming Interfaces (APIs), Artificial Intelligence, Microsoft Azure, Encodings, Continuous Integration, Distributed Systems, Github, Python (Programming Language), Systems Development Life Cycle, Queueing Systems, Reliability Engineering, Prometheus, Datadog, Large Language Models, Grafana, Multi-Agent Systems, Generative AI, Machine Learning Operations, Virtual Agents - **Published:** September 11, 2026 - **Apply:** https://www.careerjet.com/jobad/us5996d4b003c5d6a2bb7d21b396893ec4 ## About the Role * 7+ years across software QA and testing, 2+ of them on AI/ML or LLM-based systems, with agentic and multi-agent architectures you have personally tested * Python at a level where you've shipped your own test harnesses, simulators, and fixtures * Real LLM evaluation experience - exact and soft match, BLEU/ROUGE, BERTScore, embedding-based semantic similarity - alongside guardrail and prompt testing * Distributed systems testing chops: latency profiling, resiliency patterns including circuit breakers and retries, chaos engineering, message queues * Hands-on with at least one orchestration framework: LangChain, LangGraph, LlamaIndex, DSPy, OpenAI Assistants/Actions, Azure OpenAI orchestration, or something comparable * Comfortable inside CI/CD (GitHub Actions, Azure DevOps) and observability stacks (OpenTelemetry, Prometheus/Grafana, Datadog), plus feature flags and canaries * A working grip on privacy, security, and compliance for AI systems - PII handling, content policy, model safety * Communication and leadership range to hold your own across Operations, Data, and Engineering ## Description Agentic AI is already running in production here, and it needs someone who can prove it works. You'll set the testing approach for multi-agent systems - how resiliency, accuracy, latency, orchestration correctness, and behavior at scale actually get measured - and hold that standard from first commit to production. Frameworks, harnesses, and the QA function itself are yours to build, working directly with the Agentic Operations group. This is a build-it-yourself role, not an oversight one. Where You'll Spend Your Time Setting the Bar * The QA strategy for agentic and multi-agent systems is yours to define and defend across development, staging, and production * Coach QA engineers and put real testing standards, harness coding guidelines, and review practice in place * Work alongside Data Science, MLOps, and Platform teams to weave QA into the SDLC and into incident response Testing the Agent Layer * Write the tests that exercise agent orchestration, tool calling, planner-executor loops, and coordination between agents - how tasks get decomposed, whether handoffs stay intact, whether the system converges on its goal * Check that state management, context windows, memory and knowledge stores, and prompt and graph correctness hold up as conditions shift * Prove the orchestrator does what it claims under DAG execution, retries, branching, timeouts, and compensation paths Proving Correctness * Stand up ground-truth and reference pipelines that measure task accuracy through exact match, semantic similarity, and factuality checks * Create macro validation frameworks that can judge outcomes across multi-step agent workflows, generation-plus-verification loops included * Wire in guardrail checks for toxicity, PII, hallucination, and policy compliance Breaking It on Purpose * Fuzz the scenarios: adversarial inputs, prompt perturbations, tool latency spikes, degraded APIs * Assemble resilience suites covering chaos experiments, failover, retry and backoff, circuit-breaking, and degraded-mode behavior * Establish latency SLOs, then measure end-to-end response time across LLM calls, tool invocations, and queues * Keep reliability honest with soak tests, canary verification, and automated rollback Under Load * Build load and stress tests that push multi-agent graphs on concurrency, throughput, queue depth, and backpressure Shipping Safely * Produce test artifacts other engineers can reuse - scenario configs, synthetic datasets, prompt libraries, agent graph fixtures, simulators * Push testing into CI/CD as pre-merge gates, nightly runs, and canaries, and into production monitoring with alerting tied to KPIs * Own release criteria and operational readiness across performance, security, compliance, and cost/latency budgets ## Related Videos - [Agentic AI - From Theory to Practice: Developing Multi-Agent AI Systems on Azure](https://www.wearedevelopers.com/videos/1532-agentic-ai-from-theory-to-practice-developing-multi-agent-ai-systems-on-azure) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Beyond Chatbots: How to build Agentic AI systems](https://www.wearedevelopers.com/videos/1629-beyond-chatbots-how-to-build-agentic-ai-systems) ## Related Articles - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [The Overflow: AI and Agentic Coding](https://www.wearedevelopers.com/magazine/721-the-overflow-ai-and-agentic-coding)