> Markdown version of [/jobs/ext/2580031-ai-qa-engineer-ai-test-architect](https://www.wearedevelopers.com/jobs/ext/2580031-ai-qa-engineer-ai-test-architect). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI QA Engineer / AI Test Architect - **Company:** TOGAL.AI INC - **Location:** United States (Remote available) - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Architectural Patterns, Automation of Tests, Code Coverage, Continuous Integration, Github, Monitoring of Systems, Apache JMeter, Python (Programming Language), Test Case, TypeScript, Management of Software Versions, Web Application Frameworks, Circleci, Postman, Large Language Models, Cypress (Programming Language), Gatling, Containerization, Kubernetes, Playwright, Machine Learning Operations, Grpc, Api Management, Docker, Programming Languages, Microservices - **Published:** August 3, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=f23e1d27ba11b6f8 ## About the Role * Test automation engineering: deep expertise in writing and maintaining automated test suites using modern frameworks (Playwright, Cypress, or similar) with at least one modern programming language, such as TypeScript (strongly preferred) or Python, specifically for building robust test libraries. * Non-functional test automation: hands-on experience designing and running performance, load, and stress tests with tools such as k6 or JMeter, including CI/CD integration and threshold-based quality gates. * AI/LLM testing literacy: practical understanding of what makes AI systems non-deterministic, and experience (or strong working knowledge) of testing LLM-based features for hallucination, consistency, safety, and latency. * Eval framework awareness: a working understanding of LLM evaluation concepts: scoring metrics (BLEU, ROUGE), LLM-as-judge patterns, and familiarity with at least one eval framework (DeepEval, RAGAS, etc.). * CI/CD and continuous testing: experience integrating test suites into pipelines (GitHub Actions, CircleCI, or equivalent) with a shift-left mindset that treats test failures as blocking signals, not background noise. * Quality ownership mentality: demonstrated ability to own quality outcomes, not just execute tasks; comfort setting standards, raising risk flags, and influencing cross-functional teams. * Solid understanding of architectural patterns, microservices, and API testing (REST, gRPC) using tools like Postman or custom frameworks. * Experience with containerization technologies (Docker) and orchestration (Kubernetes) as they relate to scalable testing environments. Nice-to-Haves * Experience with agentic systems testing: validating goal-completion rates, guardrail enforcement, and multi-step reasoning chains in LLM agent workflows. * Familiarity with AIOps/MLOps/LLMOps concepts: prompt versioning, model monitoring, canary deployments, and drift detection. * Experience with accessibility or security testing as part of a broader non-functional quality practice. * Background in red teaming, adversarial input testing, or prompt injection validation. * Experience working in a startup or scale-up environment where processes are built from scratch rather than inherited. ## Description * Own quality strategy end-to-end: define, implement, and continuously evolve testing standards across functional, non-functional, and AI-specific dimensions, ensuring quality is embedded from requirements through production. * Build and maintain non-functional test automation: design and run performance, load, and stress test suites (k6, JMeter, Gatling etc.) integrated directly into CI/CD pipelines, with quality gates that protect every release. * Design and operate (or contribute to) LLM/AI eval frameworks: establish evaluation pipelines (using tools such as DeepEval, Langfuse etc.) to assess AI feature quality across metrics including accuracy, hallucination rate, relevance, faithfulness, and safety. * Test AI features and agentic behaviours: validate non-deterministic outputs, prompt variability, model regression, guardrail enforcement, and multi-step agent task-completion rates as first-class quality concerns. * Champion shift-left and continuous testing: embed QA into planning, design review, and sprint ceremonies so defects are caught before they're coded, not after they ship. * Drive a quality engineering culture: act as a quality advocate across engineering, product, and AI teams; run blameless post-mortems, define quality metrics, and make test coverage and reliability visible to the whole organization. * Accelerate delivery through AI-assisted tooling: use AI coding assistants, self-healing automation, and intelligent test prioritization to increase the leverage of every hour spent on quality work. * Build observability into production: define and monitor post-release quality signals, model drift indicators, and SLO thresholds so the team can distinguish a regression from expected non-determinism., * Traditional QA foundations: solid understanding of deterministic testing: test planning, test case design, functional/regression/exploratory testing, defect lifecycle management, and quality metrics. ## Related Videos - [Exploring the Power of gRPC-Gateway for Writing RESTful Services](https://www.wearedevelopers.com/videos/2072-exploring-the-power-of-grpc-gateway-for-writing-restful-services) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [AI as a Test Designer: Transforming Experience into Automated Testing](https://www.wearedevelopers.com/videos/1984-ai-as-a-test-designer-transforming-experience-into-automated-testing) - [Boosting OpenSearch Performance: gRPC Search in Action](https://www.wearedevelopers.com/videos/1964-boosting-opensearch-performance-grpc-search-in-action) - [Let's Talk Quality! - Lilia Gargouri](https://www.wearedevelopers.com/videos/1815-let-s-talk-quality-lilia-gargouri) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [13 AI Tools for Developers](https://www.wearedevelopers.com/magazine/302-13-ai-tools-for-developers) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)