World Congress 2026 Europe

Parquet, Delta, Iceberg & Ducklake - An introduction for developers

July 9, 2026 13:30 – 14:00 · 30 min Stage 12

What this session covers

CSVs are inefficient. Everyone knows that. And yet they are probably the most widely used file formats, also for data scientists. At the same time, data engineers talk about Parquet, Iceberg, and Ducklake—and roll their eyes when someone actually still uses CSV or JSON.

As a software engineer, there’s often nothing left to do but close your eyes and go for it. Even if you don’t really understand it. You read CSV, you write Delta or Iceberg. The main thing is that the data guys don’t complain. But what are the differences? Why do we store data in files in the first place? What is all this metadata? And what do I really need and what not? And why can’t just look onto storage but also have to take the compute into account. It’s high time to dive in.

Related talks at this congress

Open session

World Congress 2026 Europe

July 10, 2026 · 14:20–14:50

Stage 10 - powered by TikTok

Swapping a Data Warehouse at Runtime: Zero-Downtime Migration Without Changing a Single Client

Max Fischer, Michael O'Toole

Max Fischer
Michael O'Toole
Open session

World Congress 2026 Europe

July 9, 2026 · 14:50–15:20

Stage 12

TiDB, One Layer at a Time: How Distributed SQL Became an Agentic AI Backbone

Daniël van Eeden, Mattias Jonsson

Daniël van Eeden
Mattias Jonsson
Open session

World Congress 2026 Europe

July 10, 2026 · 09:40–10:10

Stage 2

From Messy Queries to Scalable Systems - How Data Engineering actually works

Sandhya Menon

Lead Data Engineer at E.ON Energy Deutschland

Sandhya Menon
Open session

World Congress 2026 Europe

July 9, 2026 · 15:30–16:00

Stage 12

Beyond SQL Generation: How to Teach Agents What Your Database Actually Means

Celeste Horgan

Senior Developer Advocate at Snowflake

Celeste Horgan
All sessions at this congress