AI QA Engineer / AI Test Architect
TOGAL.AI INC
United States
about 1 month ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.indeed.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source
Tech stack
Artificial Intelligence
Architectural Patterns
Automation of Tests
Code Coverage
Continuous Integration
Github
Monitoring of Systems
Apache JMeter
Python (Programming Language)
Test Case
TypeScript
Management of Software Versions
+15 more
Web Application Frameworks
Circleci
Postman
Large Language Models
Cypress (Programming Language)
Gatling
Containerization
Kubernetes
Playwright
Machine Learning Operations
Grpc
Api Management
Docker
Programming Languages
Microservices
Job description
- Own quality strategy end-to-end: define, implement, and continuously evolve testing standards across functional, non-functional, and AI-specific dimensions, ensuring quality is embedded from requirements through production.
- Build and maintain non-functional test automation: design and run performance, load, and stress test suites (k6, JMeter, Gatling etc.) integrated directly into CI/CD pipelines, with quality gates that protect every release.
- Design and operate (or contribute to) LLM/AI eval frameworks: establish evaluation pipelines (using tools such as DeepEval, Langfuse etc.) to assess AI feature quality across metrics including accuracy, hallucination rate, relevance, faithfulness, and safety.
- Test AI features and agentic behaviours: validate non-deterministic outputs, prompt variability, model regression, guardrail enforcement, and multi-step agent task-completion rates as first-class quality concerns.
- Champion shift-left and continuous testing: embed QA into planning, design review, and sprint ceremonies so defects are caught before they’re coded, not after they ship.
- Drive a quality engineering culture: act as a quality advocate across engineering, product, and AI teams; run blameless post-mortems, define quality metrics, and make test coverage and reliability visible to the whole organization.
- Accelerate delivery through AI-assisted tooling: use AI coding assistants, self-healing automation, and intelligent test prioritization to increase the leverage of every hour spent on quality work.
- Build observability into production: define and monitor post-release quality signals, model drift indicators, and SLO thresholds so the team can distinguish a regression from expected non-determinism., * Traditional QA foundations: solid understanding of deterministic testing: test planning, test case design, functional/regression/exploratory testing, defect lifecycle management, and quality metrics.
Requirements
- Test automation engineering: deep expertise in writing and maintaining automated test suites using modern frameworks (Playwright, Cypress, or similar) with at least one modern programming language, such as TypeScript (strongly preferred) or Python, specifically for building robust test libraries.
- Non-functional test automation: hands-on experience designing and running performance, load, and stress tests with tools such as k6 or JMeter, including CI/CD integration and threshold-based quality gates.
- AI/LLM testing literacy: practical understanding of what makes AI systems non-deterministic, and experience (or strong working knowledge) of testing LLM-based features for hallucination, consistency, safety, and latency.
- Eval framework awareness: a working understanding of LLM evaluation concepts: scoring metrics (BLEU, ROUGE), LLM-as-judge patterns, and familiarity with at least one eval framework (DeepEval, RAGAS, etc.).
- CI/CD and continuous testing: experience integrating test suites into pipelines (GitHub Actions, CircleCI, or equivalent) with a shift-left mindset that treats test failures as blocking signals, not background noise.
- Quality ownership mentality: demonstrated ability to own quality outcomes, not just execute tasks; comfort setting standards, raising risk flags, and influencing cross-functional teams.
- Solid understanding of architectural patterns, microservices, and API testing (REST, gRPC) using tools like Postman or custom frameworks.
- Experience with containerization technologies (Docker) and orchestration (Kubernetes) as they relate to scalable testing environments.
Nice-to-Haves
- Experience with agentic systems testing: validating goal-completion rates, guardrail enforcement, and multi-step reasoning chains in LLM agent workflows.
- Familiarity with AIOps/MLOps/LLMOps concepts: prompt versioning, model monitoring, canary deployments, and drift detection.
- Experience with accessibility or security testing as part of a broader non-functional quality practice.
- Background in red teaming, adversarial input testing, or prompt injection validation.
- Experience working in a startup or scale-up environment where processes are built from scratch rather than inherited.
Benefits & conditions
- Join a dynamic team of AI-native engineering team.
- AI-native from day one - you’ll be building the quality discipline for a product that uses AI at its core, making every quality decision novel and impactful.
- Be the quality voice, not a quality follower - this role has direct influence over how Togal defines and measures product excellence.
- A culture of innovation, continuous learning, and high growth.
- Comprehensive benefits
- Competitive compensation package and flexible work arrangements.
About the company
Our collaborative platform enables real-time teamwork, instant drawing analysis, and features a revolutionary conversational AI interface that transforms how professionals interact with construction plans. Founded by construction industry veterans, our award-winning application automates the takeoff process, enabling estimators to analyze blueprints in seconds rather than hours or days.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
BB
Benedikt Bischof
over 4 years ago
ER
Erin Rifkin
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
about 1 year ago
LM
Luis Minvielle
What Are Large Language Models?
almost 3 years ago
BB
Benedikt Bischof
MLOps – What’s the deal behind it?
almost 4 years ago
LM
Luis Minvielle
13 AI Tools for Developers
almost 3 years ago
CH
Chris Heilmann
Dev Digest 137 - AI'm not sure about this
almost 2 years ago