> Markdown version of [/jobs/ext/2105165-software-engineer-in-test-ai-agentic-systems](https://www.wearedevelopers.com/jobs/ext/2105165-software-engineer-in-test-ai-agentic-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer in Test (AI Agentic Systems) - **Company:** CollectiveHealth, Inc. - **Location:** Lehi, UT, United States - **Experience:** Expert - **Salary:** $99,200.0 - $124,000.0 - **Contract:** Permanent contract - **Skills:** Testing (Software), Application Programming Interfaces (APIs), Artificial Intelligence, BigQuery, Cloud Computing, Continuous Integration, Information Engineering, Python (Programming Language), Mockito, SQL Databases, Data Logging, Large Language Models, Prompt Engineering, Pandas, Pytest, SDET, Web Api - **Published:** August 18, 2026 - **Apply:** https://www.dice.com/job-detail/4bc7eb1d-a001-420f-a9f7-2f8506d02022 ## About the Role + 5+ years of of experience in software testing and development + Python SDET Expertise: Expert in Python and pytest, specifically building custom mocking frameworks for external APIs (Vertex AI/ADK). + AI/LLM Observability: Hands-on experience with Vertex AI Experiments, Auto-SxS, and Cloud Logging for trace analysis. + Data Literacy: Expert-level SQL (BigQuery) and Pandas skills to "diff" massive datasets and identify adjudication discrepancies. + Prompt Engineering for QA: Ability to analyze "System Instructions" and refine prompts based on failed test cases to close logic gaps. + Architectural Testing: Experience testing multi-layer systems involving RAG (Vertex AI Search), state management (LangGraph), and function calling. * Preferred Skills (The "Nice-to-Haves") + Healthcare/Claims Domain: Familiarity with claims adjudication concepts (pend reason codes, COB, eligibility, stop-loss). + Compliance Knowledge: Understanding of HIPAA/PHI handling and writing test evidence for regulatory bodies (DOL/DOI). + Human-in-the-Loop Testing: Experience in "Shadow Mode" monitoring-comparing agent decisions against human expert (MCA) baselines. ## Description You will work at the intersection of Vertex AI, healthcare compliance, and high-scale data engineering. Your work directly determines whether claims are paid correctly and whether the company can withstand a Department of Labor (DOL) or state DOI audit. The stakes are real, the domain is hard, and the problems are genuinely novel. What you'll do: * Outcome Evaluation (The "What") + Golden Set Governance: Build and maintain a versioned library of "Grounding Data" results by working with senior claims examiners to define "Ground Truth." + Model-as-a-Judge Automation: Design automated "LLM-grading-LLM" workflows using custom rubrics to score factual grounding and policy compliance. + Semantic Assertion Framework: Develop testing libraries that move beyond string matching to validate semantic equivalence and numerical accuracy in agent outputs. * Trajectory Evaluation (The "How") + Function-Call Auditing: Use Vertex AI traces to programmatically verify that mandatory tools (via MCP) were invoked with correct arguments. + Orchestration Logic Validation: Assert that agents respect defined priorities across the four architectural layers: Data & Knowledge, Orchestration, Agentic Reasoning, and Tooling. + Reasoning Trace Auditing: Ensure every autonomous decision is traceable to a specific SOP sentence and a live API data point. * Continuous Automated Regression (The "Always") + CI/CD Integration: Every prompt or model update in Vertex AI Prompt Management must trigger an automated regression run. + Auto-SxS: Own the automated pairwise comparison process to detect logic drift between "New" and "Production" agent versions. + Mocking & Resilience: Build a Vertex AI/ADK mocking layer to simulate model responses, allowing for thousands of logic tests in seconds with zero API costs. ## Related Videos - [pytest: Simple, rapid and fun testing with Python](https://www.wearedevelopers.com/videos/213-pytest-simple-rapid-and-fun-testing-with-python) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [AI as a Test Designer: Transforming Experience into Automated Testing](https://www.wearedevelopers.com/videos/1984-ai-as-a-test-designer-transforming-experience-into-automated-testing) - [Automagic Configuration in Python](https://www.wearedevelopers.com/videos/363-automagic-configuration-in-python) - [Back to the Roots: Testing in the Age of AI](https://www.wearedevelopers.com/videos/100021-back-to-the-roots-testing-in-the-age-of-ai) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) ## Related Articles - [13 AI Tools for Developers](https://www.wearedevelopers.com/magazine/302-13-ai-tools-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [The State of WebDev AI 2025 Results: What Can We Learn?](https://www.wearedevelopers.com/magazine/581-the-state-of-webdev-ai-2025-results-what-can-we-learn)