> Markdown version of [/jobs/ext/1886880-senior-staff-software-ai-test-engineer-ai-engineering](https://www.wearedevelopers.com/jobs/ext/1886880-senior-staff-software-ai-test-engineer-ai-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior / Staff Software AI Test Engineer, AI Engineering - **Company:** TWG, INC. - **Location:** Santa Monica, CA, United States - **Experience:** Expert - **Salary:** $190,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Testing (Software), Clean Code Principles, Java (Programming Language), Test Suite, Application Programming Interfaces (APIs), Artificial Intelligence, Data Analysis, Apple IOS, Automation of Tests, Microsoft Azure, Cloud Computing, Code Coverage, Code Review, Continuous Delivery, Continuous Integration, Data Files, Software Debugging, Failover, Github, Home Automation, Web Servers, Java Web Services, Python (Programming Language), Load Testing, Object-Oriented Software Development, Systems Development Life Cycle, Next.js, Selenium, Software Engineering, SQL Databases, Systems Integration, Test Case, Test Data, Circleci, Data Processing, Scripting, Flexi (Photoshop Plugin), Performance Testing, Large Language Models, Model Validation, Cypress (Programming Language), Build Management, Pytest, Containerization, Gitlab-ci, Information Technology, Playwright, Production Code, Docker, SDET, Jenkins, Data Generation - **Published:** July 31, 2026 - **Apply:** https://www.careerbuilder.com/job-details/senior-staff-software-ai-test-engineer-ai-engineering-santa-monica-ca--2c116a12-dc70-4bd9-bcab-626f0b665922 ## About the Role You'll work shoulder-to-shoulder with AI engineers and data scientists, contributing production-quality code to shared repositories. The ideal candidate is a strong coder, fluent in Python and Java - who has shipped automated test infrastructure in a production environment and has hands-on experience evaluating LLM and agentic systems., * 3-7 years of software engineering experience, with a meaningful portion focused on test automation, SDET, or software engineering in test roles. * Expert-level Python. You write Python every day, design libraries other engineers use, and apply OOP and clean-code practices. * Hands-on Java experience, enough to read, write, and test Java services, not just touch them. * Working understanding of the LangGraph or Vercel frameworks: graph state, nodes, edges, tool calls, and how to write evals against agentic flows. * Demonstrated experience building eval sets for LLM models (this is critical to the role). * Experience testing across multiple client surfaces: iOS apps, plugins, and Chrome extensions. * Hands-on experience building automated test suites with frameworks such as pytest, Selenium, Playwright, Cypress, or similar. * Proven experience integrating test automation into CI/CD systems (GitHub Actions, Jenkins, CircleCI, GitLab CI, or similar). * Strong skills in data manipulation, test data preparation, and SQL. * Bachelor's degree or higher in Computer Science, Engineering, or a related field. Preferred Qualifications: * Experience with Azure (our primary cloud) and containerization (Docker). * Experience testing RAG pipelines, agentic workflows, or multi-step tool-calling systems., Alliance/Partner Management, Application Programming Interface (API), Artificial Intelligence (AI), Artificial Intelligence (AI) Agents, Artificial Intelligence (AI) Games, Automation, Benchmarking, Business Intelligence Software, Business Transformation, Cloud Computing, Code Coverage, Code Reviews, Computer Science, Continuous Deployment/Delivery, Continuous Integration, Data Analysis, Data Modeling, Data Science, Data Sets, Debugging Skills, Docker, Documentation, Engineering, Establish Priorities, Failover, Finance, Financial Services, GitHub, Home Automation, Injections, Insurance, Java, Java Testing, Jenkins, Load Testing, Machine Tool, Marketing, Metrics, Microsoft Windows Azure, Object Oriented Programming (OOP), Performance Testing, Product Development, Pytest, Python Programming/Scripting Language, Quality Engineering, Regulatory Compliance, Reporting Dashboards, SQL (Structured Query Language), Scalable System Development, Selenium, Software Design for Test (SDET), Software Engineering, Software Testing, Sports, Startup, Team Building, Test Automation, Test Case, Test Data, Test Harness, Test Plan/Schedule, Test Suite, Test Tools, Testing, Web Client Plug-ins, eCommerce, iOS ## Description TWG Global is seeking a Senior or Staff AI Software Engineer in Test to join our AI Engineering team building commercial-grade AI products. This is a software engineering role focused on test automation. You won't just write test cases, you'll design and build the frameworks, harnesses, evaluation infrastructure, and tooling that make testing AI agents and LLM-powered applications possible at scale. Our agents are written in LangGraph and run on Azure on the TWG side, with a parallel Vercel-based stack on the Palantir side. You'll write eval sets against both, and you'll validate the surfaces our users actually touch: iOS apps, plugins, and Chrome extensions, not just the model layer., Framework and harness engineering * Design and build scalable, reusable test automation frameworks for AI agents, LLM-powered applications, and underlying APIs. * Write clean, maintainable Python for test harnesses, eval pipelines, synthetic data generation utilities, and internal tooling. * Treat test code as production code: code review, type hints, documentation, library design. Evaluation infrastructure * Build evaluation infrastructure for benchmarking agent performance against SOTA LLMs, competitors, and internal baselines. * Own regression suites, golden datasets, rubric-based evals, and metric dashboards. * Build tooling for synthetic test data generation, edge-case discovery, and adversarial testing. Resilience and load * Design and run release, system, performance, and load tests against streaming, stateful, and async systems. * Build chaos and fault injection tooling for token expiry, connection pool exhaustion, provider failover, and cache pressure scenarios. * Drive contract testing across LLM providers (Bedrock, Anthropic, OpenAI) to catch parity drift. CI/CD and observability * Integrate automated tests into CI/CD so every model, prompt, and code change is validated before it ships. * Build trace-based assertions on LangGraph state, tool calls, and agent decisions - debugging an agent failure means replaying graph state, not re-running a prompt. * Make observability a first-class testing surface (LangSmith, audit logs). Human-in-the-loop and partnership * Implement HIL review workflows where automation alone cannot validate quality, then push the automation boundary outward. * Partner with AI engineers and data scientists on model evaluation, training and eval data prep, and root-cause debugging of complex end-to-end failures. * Champion quality engineering practices across the team: code review, coverage standards, observability, reproducibility. * Ensure user-centric validation so AI outputs are accurate, reliable, and meet real-world application needs. ## Related Videos - [Back to the Roots: Testing in the Age of AI](https://www.wearedevelopers.com/videos/100021-back-to-the-roots-testing-in-the-age-of-ai) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [AI as a Test Designer: Transforming Experience into Automated Testing](https://www.wearedevelopers.com/videos/1984-ai-as-a-test-designer-transforming-experience-into-automated-testing) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [The 8 Best Code Testing Tools](https://www.wearedevelopers.com/magazine/402-the-8-best-code-testing-tools) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [13 AI Tools for Developers](https://www.wearedevelopers.com/magazine/302-13-ai-tools-for-developers) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix)