WeAreDevelopers LIVE • May 12, 2021

Don't Change the Partition Count for Kafka Topics!

Dainius Jocas

A routine increase in Kafka partitions caused Vinted's Elasticsearch to silently serve stale data. Discover why you should never change partition counts on live, offset-dependent topics.

Pause
Mute Enter Fullscreen
#1 about 4 min

Introduction to the data indexing pipeline

Moving primary data from MySQL to Elasticsearch using Kafka creates a scalable indexing pipeline.

#2 about 2 min

Understanding Elasticsearch and optimistic concurrency control

Using optimistic concurrency control with document version numbers ensures only the newest updates are searchable.

#3 about 3 min

Kafka log compaction and tombstone messages

Configuring topics with infinite retention and log compaction prevents disk exhaustion while enabling safe re-indexing.

#4 about 2 min

Connecting systems safely using Kafka Connect

Using Kafka partition offsets as document versions allows Kafka Connect to safely parallelize indexing.

#5 about 3 min

Investigating reports of stale data in Elasticsearch

A production bug report reveals that tombstone messages are failing to delete obsolete Elasticsearch documents.

#6 about 4 min

Comparing Elasticsearch versions and Kafka offsets

Tracing the issue reveals older Elasticsearch documents possessing higher version numbers than newer Kafka offsets.

#7 about 2 min

Identifying the partition count root cause

Consulting documentation and metrics reveals that increasing the Kafka topic partition count triggered the inconsistency.

#8 about 4 min

How partition changes break message ordering

Altering the partition count changes message hashing, destroying the order guarantees for specific keys.

#9 about 3 min

Fixing the data inconsistency with full ingestion

Resolving the data inconsistency requires fully re-ingesting primary datastore records into newly created Kafka topics.

#10 about 2 min

Establishing best practices for Kafka partition sizing

Setting sensible default partition counts and avoiding modifications prevents data loss when relying on message ordering.

Matching moments

3:45 min

Reviewing core Apache Kafka architecture and distributed fundamentals

Kirill Kulikov · LIVE

3:06 min

Core concepts of Apache Kafka and topic topologies

Denis Washington +1 · World Congress 2021

2:33 min

Hidden costs of self-hosting and managed Kafka solutions

Bobur Umurzokov · LIVE

8:27 min

Solving race conditions and distributed state with repartitioning

Denis Washington +1 · World Congress 2021

2:00 min

Moving from traditional databases to decoupled event streaming

Gerard Klijs · LIVE

1:37 min

Introduction to Apache Kafka benchmarking and performance analysis

Kirill Kulikov · LIVE