> Markdown version of [/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works?t=392](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works?t=392). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # From Messy Queries to Scalable Systems - How Data Engineering actually works A central data warehouse often transforms into a single point of failure. Learn how adopting a hybrid Data Mesh empowers domain teams and eliminates engineering bottlenecks. - **Speakers:** [Sandhya Menon](https://www.wearedevelopers.com/@sandhya-menon) - **Event:** World Congress 2026 Europe - **Published:** July 10, 2026 - **Duration:** 31:15 - **URL:** https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works ## Summary Before an organization can claim to be "AI-first," it must establish a deeply reliable, scalable data foundation. E.ON Energy's transition from siloed, undocumented SQL queries and manual Excel reports to a modern data ecosystem highlights the hidden costs of poor architecture. While establishing a central data warehouse initially provided structured reporting, it quickly became a severe bottleneck. As data requirements matured, the overarching lesson became perfectly clear: a single point of truth can easily transform into a single point of failure if a centralized team is required for every minor transformation, ultimately slowing time-to-market and compounding employee frustration. To decouple scaling limitations, the organization adopted a hybrid Data Mesh architecture, shifting technical ownership directly to domain business teams through an "Open Data Platform." Under this federated model, the central data engineering team provisions infrastructure rather than writing endless business logic. Domain experts are empowered with a modern tech stack featuring Snowflake as the cloud data platform, dbt Core for modular transformations, Apache Airflow for batch orchestration, and Atlan for curating a searchable internal data marketplace. This decentralization completely eliminated workflow bottlenecks, allowing distinct units to operate autonomously without compromising system-wide governance or structural standards. Thriving in a decentralized paradigm requires treating data engineering with the exact same rigor as complex software engineering. Relying on version control via Git, utilizing tagged releases, and establishing strict code formatting with automated linters like SQLFluff solidifies the discipline needed for trustworthy data products. At the same time, practitioners must avoid the trap of over-engineering; building hyper-complex, proactive systems without immediate stakeholder alignment wastes resources, whereas simple, iterative pipelines often hit business targets faster. Ultimately, "data engineering is not just SQL... it is system engineering of data," driving the stability, pipeline reliability, and cultural trust required to genuinely unlock the future potential of enterprise AI. **Keywords:** data engineering, data mesh architecture, snowflake cloud data platform, dbt core transformations, apache airflow orchestration, data vault modeling, decentralized domain ownership, federated data governance, open data platform, internal data marketplace, data cataloging with atlan, sqlfluff code linting, software engineering for data, overcoming data bottlenecks, ai data readiness ## Chapters 1. **Resolving the problem of messy ad-hoc queries** (03:36) — Unstructured SQL queries and scattered data files create silos and erode organizational trust. 1. **The true role and evolution of data engineering** (05:13) — Modern data engineering moves beyond query optimization to focus on reliable data movement, scalability, and automation. 1. **Transforming data architecture from on-premise to cloud** (06:32) — Transitioning from an on-premise database to Snowflake cloud storage utilizes Data Vault modeling to handle complex business acquisitions. 1. **The bottleneck of centralized data team ownership** (09:18) — Centralized data teams struggle to keep up with increasing business demands, leading to delayed innovations and employee frustration. 1. **Adopting a decentralized data mesh architecture model** (11:16) — Implementing data mesh principles distributes ownership to domain teams while maintaining federated governance and cross-domain data sharing. 1. **Empowering domain teams with an open data platform** (13:42) — Rethinking the central team's mandate to supply decentralized infrastructure and tools like Snowflake, dbt, and Airflow. 1. **Refactoring complex logic into scalable data products** (15:58) — Transforming a rigid top customer SQL query into a reusable, automated data pipeline establishes monitoring and organization-wide consensus. 1. **Core technical practices for robust data engineering** (18:26) — Essential disciplines include data modeling, pipeline orchestration, Git version control, standardization, and proactive monitoring. 1. **Common pitfalls in scaling modern data systems** (21:23) — Avoiding over-engineering, prioritizing documentation, securing stakeholder alignment, and planning for ongoing maintenance ensure long-term platform health. 1. **The future of data engineering and AI mesh** (22:57) — Preparing for metadata-driven platforms and AI mesh requires cementing a solid, scalable data foundation first. 1. **Strategies for transferring data ownership and adapting tools** (24:47) — Migrating responsibilities to domain teams demands continuous learning and adapting data tooling amidst rapid AI advancements. ## Related Moments - [Introduction to the speaker and data science background](https://www.wearedevelopers.com/videos/369-the-state-of-mlops-machine-learning-in-production-at-enterprise-scale) (from "The state of MLOps - machine learning in production at enterprise scale") - [Career evolution in data engineering and AI platforms](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) (from "Coffee with Developers - Maria Apazoglou") - [Introduction to analytical data formats for software developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) (from "Parquet, Delta, Iceberg & Ducklake - An introduction for developers") - [Embedding data engineering to solve database scalability](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) (from "Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure") - [Decentralizing data bottlenecks with data mesh principles](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) (from "Modern Data Architectures need Software Engineering") - [Solving complex platform architecture challenges at an enterprise scale](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) (from "Coffee with Developers - Maria Apazoglou") ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Why Event-Driven Architecture Isn’t About Speed (and When You Actually Need It)](https://www.wearedevelopers.com/magazine/745-why-event-driven-architecture-isn-t-about-speed-and-when-you-actually-need-it) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) ## Related Jobs - [Lead Software Engineer - Data Engineering](https://www.wearedevelopers.com/jobs/ext/2000968-lead-software-engineer-data-engineering) at **Dynatrace** - [Senior Data Engineer](https://www.wearedevelopers.com/jobs/ext/1589390-senior-data-engineer) at **Douglas GmbH** - [Data Scientist](https://www.wearedevelopers.com/jobs/ext/1351648-data-scientist) at **Almedia** - [Cloud Data Engineer (all genders)](https://www.wearedevelopers.com/jobs/48216-cloud-data-engineer-all-genders) at **msg** - [Principal Software Engineer, Enterprise AI Platform](https://www.wearedevelopers.com/jobs/ext/1467292-principal-software-engineer-enterprise-ai-platform) at **GitHub** - [Staff Business Intelligence Engineer](https://www.wearedevelopers.com/jobs/ext/626164-staff-business-intelligence-engineer) at **Twilio**