World Congress 2022 • Jun 15, 2022

Make Your Data FABulous

Philipp Krenn

Scaling Elasticsearch prematurely can silently destroy your scoring accuracy. Discover how to navigate the Fast, Accurate, and Big trade-offs to fix broken distributed searches.

Pause
Mute Enter Fullscreen
#1 about 5 min

Introduction and the CAP theorem for distributed systems

The CAP theorem dictates that distributed stateful systems can only guarantee two of three primary consistency attributes.

#2 about 1 min

Comparing database ACID consistency to CAP theorem timing

The concept of database transaction consistency differs fundamentally from the timing consistency defined in the CAP theorem.

#3 about 3 min

Explaining CAP theorem trade-offs with deserted island analogy

A simple thought experiment demonstrates how isolated network nodes must choose between remaining available or staying consistent.

#4 about 2 min

Understanding fast, accurate, and big data store trade-offs

Distributed data stores balance compromises between near real-time processing, exact calculation results, and multi-node scalability.

#5 about 2 min

Sharding principles and the history of distributed shards

Splitting indices into distinct shards allows data stores to parallelize workload distribution across multiple hardware nodes.

#6 about 5 min

Terms aggregation inaccuracies caused by skewed data routing

Combining local top-N document counts from separated cluster shards leads to missing metrics when data is unevenly distributed.

#7 about 3 min

Fixing distributed document aggregation counts with shard sizing

Adjusting the payload sizes fetched from remote nodes improves statistical accuracy but sequentially consumes more processing time.

#8 about 4 min

Document text search algorithms and BM25 relevance scoring

Search algorithms mathematically evaluate document relevance using term frequency, inverse document frequency, and analyzed field lengths.

#9 about 3 min

Distributed execution using search engine query-then-fetch mechanics

Coordinating nodes fetch preliminary relevance metrics first to parse combined rankings before retrieving large full-text documents.

#10 about 2 min

Correcting irregular relevance scores with global cluster statistics

Explicitly pre-fetching distributed data frequencies normalizes scoring variations that surface when shards possess disproportionate text token counts.

#11 about 5 min

Single shard index constraints and system performance trade-offs

Routing all records through a single isolated shard maximizes query accuracy at the strict expense of distributed scalability.

Matching moments

5:46 min

Maximizing data throughput via distributed stream sharding

David Leitner · World Congress 2022

1:56 min

Keeping systems straightforward to minimize performance bottlenecks at scale

Josip Stuhli Josip Stuhli · Coffee With Developers

4:30 min

Introducing data management and the shift to streaming

Mary Grygleski Mary Grygleski · LIVE

5:31 min

Applying distributed computing principles to agentic artificial intelligence

Mary Grygleski Mary Grygleski · Europe 2026 Virtual

4:33 min

Evaluating algorithmic trade-offs across distributed system implementations

Alexander Reelsen · LIVE

10:58 min

Exploring query boundaries, data storage, and architecture limits

Denis Washington +1 · World Congress 2021

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 24, 2026 · 12:50–13:20

Stage 3

Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases

Wei Hu

Senior Vice President of Research and Development

Wei Hu
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 3

Real-Time Data Platforms at Trillion-Event Scale

Diptamay Sanyal

Principal Engineer | Data, AI & Cybersecurity Platforms

Diptamay Sanyal
Open session

World Congress 2026 North America

September 24, 2026 · 11:00–11:30

Stage 3

Making Science Larger, not just Faster

Yuval Dvir

Commercial Executive, SandboxAQ

Yuval Dvir
Open session

World Congress 2026 North America

September 24, 2026 · 14:50–15:20

Stage 2

Databases in the Agent Era

Monica Sarbu

Founder and CEO of xata.io

Monica Sarbu
Open session

World Congress 2026 North America

September 24, 2026 · 17:30–18:00

Stage 3

Scaling Distributed Queues for AI workloads

Jasmit Kaur Saluja

Staff Software Engineer at Meta Platforms Inc

Jasmit Kaur Saluja
Open session

World Congress 2026 North America

September 25, 2026 · 16:50–17:20

Stage 7

You Can’t Rerun Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan