> Markdown version of [/videos/1622-data-governance-in-the-era-of-ai?t=371](https://www.wearedevelopers.com/videos/1622-data-governance-in-the-era-of-ai?t=371). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Governance in the Era of AI Feeding 900TB of raw data into AI does not yield insights. It creates expensive hallucinations. Kiwi.com implemented strict data mesh governance to transform this noise into actionable AI. - **Speakers:** [Kateřina Ščavnická](https://www.wearedevelopers.com/@katerina-scavnicka) - **Event:** World Congress 2025 - **Published:** August 20, 2025 - **Duration:** 20:04 - **URL:** https://www.wearedevelopers.com/videos/1622-data-governance-in-the-era-of-ai ## Summary Kiwi.com initially believed that feeding massive volumes of raw customer interaction data into AI would automatically yield deep business insights. However, managing a 900-terabyte BigQuery table containing 138 billion rows led to massive cloud costs, over 250 daily SQL queries, and a team of analysts dedicated solely to pipeline resuscitation. Early attempts to prompt GPT models failed entirely due to missing context, lack of documentation, and token limits, proving that without structure, AI simply hallucinated against the noise. The core revelation was that scalable AI implementation relies fundamentally on aggressively governed data, prompting a complete architectural rebuild from scratch. To reconstruct their data architecture, the engineering team adopted a data mesh philosophy where domain service producers act as explicit data owners. They defined strict criteria for a 'data product,' requiring clear ownership, robust documentation, business-logic data quality checks, unique connectability, and security. Realizing that less is more, they consolidated 3,000 interaction tracking events down to a highly relevant 200. A critical turning point was the introduction of a three-tier risk categorization system and formal data contracts for Tier 1 assets. These contracts bind producers and consumers to explicit SLAs and incident response protocols, transforming data reliability from an implicit assumption into an enforceable engineering standard. Rebuilding the data foundation required a painful, multi-year organizational mindset shift, but the operational gains were immense. The company collapsed hundreds of fragile dashboards down to under twenty critical views governed by a unified data model, freeing analysts to perform actual exploratory work. With semantic layers and business contexts finally documented and validated, AI adoption succeeded. They deployed semantic validation across GitLab repositories to catch schema drift, initiated ML-driven anomaly detection, and launched a custom Slack bot capable of accurately running complex business analysis over the governed dataset. Ultimately, unfiltered big data causes stakeholder confusion, while strict data governance translates raw volume into actionable AI capabilities. **Keywords:** data governance, data mesh architecture, bigquery cloud costs, data product strategy, data contract implementation, data risk categorization, data quality business checks, ai data hallucination, customer interaction tracking, distributed data ownership, service level agreements, schema semantic validation, conversational analytics bot, data model consolidation ## Chapters 1. **Struggling with ungoverned data lakes and massive storage costs** (01:22) — How hoarding raw customer interaction data led to massive storage costs and fragile analytics pipelines. 1. **Failing to analyze raw data with conversational AI** (05:21) — Why feeding poorly documented datasets to generative tools produces hallucinations and incorrect analysis. 1. **Adopting data mesh ownership and aggregated data models** (06:11) — Applying a data mesh philosophy ensures producers take responsibility for their domain datasets. 1. **Defining internal standards for high-quality data products** (09:02) — Evaluating internal data against six standard criteria including business checks and security. 1. **Categorizing data tiers and establishing data contracts** (11:29) — Prioritizing critical business pipelines by categorizing assets and enforcing service-level agreements between teams. 1. **Driving mindset shifts for effective data governance** (14:15) — Overcoming cultural resistance to fix broken pipelines creates a streamlined single source of truth. 1. **Enhancing data validation and self-service analytics with AI** (17:10) — Deploying artificial intelligence for semantic repository checks, anomaly detection, and conversational queries. ## Related Moments - [Root causes of underlying AI initiative failures](https://www.wearedevelopers.com/videos/100328-the-limits-of-llms-in-real-world-applications) (from "The Limits of LLMs in Real-World Applications") - [Overcoming artificial intelligence silos in the enterprise](https://www.wearedevelopers.com/videos/1525-beyond-gpt-building-unified-genai-platforms-for-the-enterprise-of-tomorrow) (from "Beyond GPT: Building Unified GenAI Platforms for the Enterprise of Tomorrow") - [Scaling generative AI use cases across large enterprises](https://www.wearedevelopers.com/videos/916-beyond-the-hype-real-world-ai-strategies-panel) (from "Beyond the Hype: Real-World AI Strategies Panel") - [Establishing data literacy and governance as prerequisites for adoption](https://www.wearedevelopers.com/videos/1248-ai-beyond-the-code-master-your-organisational-ai-implementation) (from "AI beyond the code: Master your organisational AI implementation.") - [Reinventing data governance as a collaborative business foundation](https://www.wearedevelopers.com/videos/1367-unlocking-value-from-data-the-key-to-smarter-business-decisions) (from "Unlocking Value from Data: The Key to Smarter Business Decisions-") - [Why agentic AI forces companies to fix data debt](https://www.wearedevelopers.com/videos/100286-the-missing-layer-between-enterprise-data-and-ai-agents) (from "The Missing Layer Between Enterprise Data and AI Agents") ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [WWC24 Talk - Scott Hanselman - AI: Superhero or Supervillain?](https://www.wearedevelopers.com/magazine/469-wwc24-talk-scott-hanselman-ai-superhero-or-supervillain) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) ## Related Jobs - [Data Scientist](https://www.wearedevelopers.com/jobs/ext/1351648-data-scientist) at **Almedia** - [AI Software Engineer (Germany)](https://www.wearedevelopers.com/jobs/48317-ai-software-engineer-germany) at **Sunhat** - [Staff Business Intelligence Engineer](https://www.wearedevelopers.com/jobs/ext/626164-staff-business-intelligence-engineer) at **Twilio** - [Senior Software Engineer, Data](https://www.wearedevelopers.com/jobs/48273-senior-software-engineer-data) at **Sportradar Media Services GmbH** - [Staff, Business Intelligence Engineer](https://www.wearedevelopers.com/jobs/ext/1401813-staff-business-intelligence-engineer) at **Twilio** - [Head of AI Applications](https://www.wearedevelopers.com/jobs/ext/1231536-head-of-ai-applications) at **ZEISS Group**