World Congress 2022 Jun 15, 2022

Make Your Data FABulous

Philipp Krenn

Scaling Elasticsearch prematurely can silently destroy your scoring accuracy. Discover how to navigate the Fast, Accurate, and Big trade-offs to fix broken distributed searches.

Pause
Mute Enter Fullscreen
#1 about 5 min

Introduction and the CAP theorem for distributed systems

The CAP theorem dictates that distributed stateful systems can only guarantee two of three primary consistency attributes.

#2 about 1 min

Comparing database ACID consistency to CAP theorem timing

The concept of database transaction consistency differs fundamentally from the timing consistency defined in the CAP theorem.

#3 about 3 min

Explaining CAP theorem trade-offs with deserted island analogy

A simple thought experiment demonstrates how isolated network nodes must choose between remaining available or staying consistent.

#4 about 2 min

Understanding fast, accurate, and big data store trade-offs

Distributed data stores balance compromises between near real-time processing, exact calculation results, and multi-node scalability.

#5 about 2 min

Sharding principles and the history of distributed shards

Splitting indices into distinct shards allows data stores to parallelize workload distribution across multiple hardware nodes.

#6 about 5 min

Terms aggregation inaccuracies caused by skewed data routing

Combining local top-N document counts from separated cluster shards leads to missing metrics when data is unevenly distributed.

#7 about 3 min

Fixing distributed document aggregation counts with shard sizing

Adjusting the payload sizes fetched from remote nodes improves statistical accuracy but sequentially consumes more processing time.

#8 about 4 min

Document text search algorithms and BM25 relevance scoring

Search algorithms mathematically evaluate document relevance using term frequency, inverse document frequency, and analyzed field lengths.

#9 about 3 min

Distributed execution using search engine query-then-fetch mechanics

Coordinating nodes fetch preliminary relevance metrics first to parse combined rankings before retrieving large full-text documents.

#10 about 2 min

Correcting irregular relevance scores with global cluster statistics

Explicitly pre-fetching distributed data frequencies normalizes scoring variations that surface when shards possess disproportionate text token counts.

#11 about 5 min

Single shard index constraints and system performance trade-offs

Routing all records through a single isolated shard maximizes query accuracy at the strict expense of distributed scalability.

Matching moments

5:46 min

Maximizing data throughput via distributed stream sharding

David Leitner · WWC 2022

1:56 min

Keeping systems straightforward to minimize performance bottlenecks at scale

Josip Stuhli Josip Stuhli · Coffee With Developers

4:30 min

Introducing data management and the shift to streaming

Mary Grygleski Mary Grygleski · LIVE

4:33 min

Evaluating algorithmic trade-offs across distributed system implementations

Alexander Reelsen · LIVE

10:58 min

Exploring query boundaries, data storage, and architecture limits

Denis Washington +1 · WWC 2021

3:55 min

Core architecture and federated execution of SearchOLAP

Andrey Abramov Andrey Abramov · WWC Europe 2026

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases

Wei Hu

Senior Vice President of Research and Development

Wei Hu
Open session

World Congress 2026 North America

Making Science Larger, not just Faster

Yuval Dvir

Commercial Executive, SandboxAQ

Yuval Dvir
Open session

World Congress 2026 North America

Databases in the Agent Era

Monica Sarbu

Founder and CEO of xata.io

Monica Sarbu
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

Boring Failover: Predictable Region Recovery Across 5,000 Microservices

Garvit Kataria, Sahil Sabharwal

Garvit Kataria
Sahil Sabharwal
Open session

World Congress 2026 North America

From Guesswork to Governance: Data Contracts Bring API Discipline to Apache Kafka

Sandon Jacobs

Senior Developer Advocate at IBM

Sandon Jacobs