> Markdown version of [/videos/1837-why-testing-matters-in-ai-luise-freese-and-elio-struyf](https://www.wearedevelopers.com/videos/1837-why-testing-matters-in-ai-luise-freese-and-elio-struyf). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Why Testing Matters in AI - Luise Freese and Elio Struyf Luise Freese and Elio Struyf reveal why letting AI test its own code destroys application stability. Learn how TDD and Playwright provide robust guardrails against unpredictable code generation. - **Speakers:** [Luise Freese](https://www.wearedevelopers.com/@luise-freese), Elio Struyf - **Event:** Coffee With Developers - **Published:** March 16, 2026 - **Duration:** 28:29 - **URL:** https://www.wearedevelopers.com/videos/1837-why-testing-matters-in-ai-luise-freese-and-elio-struyf ## Summary As AI tools and "vibe coding" accelerate software development, testing is increasingly viewed as a cumbersome afterthought. However, relying on AI agents to both write and evaluate code often results in tests that technically pass but fail to assert meaningful outcomes. The core challenge lies in maintaining application stability and accurate user experiences when automated code generation introduces unpredictable UI changes. By implementing robust guardrails, teams can prevent agents from merely satisfying their own generated tests without delivering actual business value. To combat this erratic output, adopting automated end-to-end (E2E) testing frameworks like Playwright combined with Gherkin becomes essential. Gherkin replaces vague "story bullshit" with a structured, shared language that aligns developers, QA engineers, and product teams before a single line of code is written. Establishing these upfront requirements ensures that AI-generated software meets specific user needs. This approach also sparks a renaissance for Test-Driven Development (TDD), where designing tests first sharpens feature understanding, prevents wasted effort, and anchors product stability against unreviewed modifications. Moving testing to the beginning of the development cycle preserves team energy and budget that are typically depleted by the end of a sprint. Furthermore, well-crafted E2E tests can double as living documentation, enabling technical writers and testers to build upon reliable base scenarios. When developers leverage these automated scripts, they not only guarantee consistent UI stability but can also repurpose them to guide AI agents through complex functional tasks, yielding a massive return on investment by safeguarding the ultimate end-user experience. **Keywords:** playwright automation testing, gherkin syntax requirements, end-to-end testing stability, AI-generated code quality, TDD methodology renaissance, vibe coding challenges, automated living documentation, QA engineering collaboration, software testing guardrails, cross-functional shared language, automated UI testing frameworks, agile user story limitations, AI testing agents, software testing ROI ## Chapters 1. **Adopting modern end-to-end testing for automated deployment pipelines** (00:00) — How quality engineering experience translates into adopting framework tools like Playwright for robust automated test pipelines. 1. **Defining actionable user stories with Gherkin and Playwright** (02:22) — How Gherkin provides structured syntax to describe actionable user stories for scalable automated testing environments. 1. **Writing meaningful test guardrails for AI generated code** (04:00) — Preventing AI coding agents from writing software that merely passes meaningless assertions without verifying actionable functional outcomes. 1. **Preventing unreviewed UI changes with continuous automated testing** (05:46) — Using stable end-to-end testing models to ensure automated coding tools do not silently alter existing user experiences. 1. **Adapting existing test automation scripts for AI agents** (07:23) — How developers can safely reuse existing happy-path test automation workflows to train external AI task agents. 1. **Handling complex API mocks and manual edge cases** (08:35) — Why human validation remains essential for breaking application inputs and correcting inaccurate AI generated API mocks. 1. **Creating shared testing language across diverse engineering disciplines** (11:00) — Reducing organizational friction by uniting developers, quality assurance engineers, and product teams around a common documentation syntax. 1. **Leveraging automated end-to-end tests as living technical documentation** (13:48) — Converting successful product workflows into automated tests to effortlessly generate reliable technical documentation for end users. 1. **Adopting upfront test design for agentic software workflows** (16:40) — Defining strict testing parameters upfront prevents organizations from wasting engineering hours maintaining incorrect AI code generations. 1. **Justifying the return on investment for testing education** (21:04) — Advocating for robust developer training budget by constantly highlighting the overarching business cost of delivering broken software. ## Related Moments - [The evolving role of software testing in AI development](https://www.wearedevelopers.com/videos/100021-back-to-the-roots-testing-in-the-age-of-ai) (from "Back to the Roots: Testing in the Age of AI") - [The future landscape of artificial intelligence in software testing](https://www.wearedevelopers.com/videos/1984-ai-as-a-test-designer-transforming-experience-into-automated-testing) (from "AI as a Test Designer: Transforming Experience into Automated Testing") - [Enforcing behavioral testing patterns in AI development tools](https://www.wearedevelopers.com/videos/100021-back-to-the-roots-testing-in-the-age-of-ai) (from "Back to the Roots: Testing in the Age of AI") - [Integrating artificial intelligence into software testing and requirement engineering](https://www.wearedevelopers.com/videos/100084-aiqspecflow-improves-and-automates-your-agile-process-of-specification-and-creation-of-testcases) (from "AIQSpecFlow: Improves and automates your agile process of specification and creation of testcases.") - [The evolution of software testing using artificial intelligence](https://www.wearedevelopers.com/videos/1984-ai-as-a-test-designer-transforming-experience-into-automated-testing) (from "AI as a Test Designer: Transforming Experience into Automated Testing") - [Exploring the expansion of end-to-end testing utility suites](https://www.wearedevelopers.com/videos/1678-the-2025-state-of-javascript-testing) (from "The 2025 State of JavaScript Testing") ## Related Articles - [Transforming Software Development: The Role of AI and Developer Tools](https://www.wearedevelopers.com/magazine/527-transforming-software-development-the-role-of-ai-and-developer-tools) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Exploring AI: Opportunities and Risks for Developers](https://www.wearedevelopers.com/magazine/522-exploring-ai-opportunities-and-risks-for-developers) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) ## Related Jobs - [Senior AI Frontend Engineer](https://www.wearedevelopers.com/jobs/ext/2552851-senior-ai-frontend-engineer) at **TeamViewer Germany GmbH,** - [AI Software Engineer (Germany)](https://www.wearedevelopers.com/jobs/48317-ai-software-engineer-germany) at **Sunhat** - [MLOps AI Engineer](https://www.wearedevelopers.com/jobs/ext/2565312-mlops-ai-engineer) at **TeamViewer Germany GmbH,** - [Senior Fullstack AI Engineer (React, Java)](https://www.wearedevelopers.com/jobs/ext/2618728-senior-fullstack-ai-engineer-react-java) at **TeamViewer Germany GmbH,** - [Senior AI Developer](https://www.wearedevelopers.com/jobs/ext/2836034-senior-ai-developer) at **PwC** - [Software Engineer - Video](https://www.wearedevelopers.com/jobs/ext/2600051-software-engineer-video) at **Twilio**