> Markdown version of [/videos/138-don-t-change-the-partition-count-for-kafka-topics?t=1242](https://www.wearedevelopers.com/videos/138-don-t-change-the-partition-count-for-kafka-topics?t=1242). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Don't Change the Partition Count for Kafka Topics! A routine increase in Kafka partitions caused Vinted's Elasticsearch to silently serve stale data. Discover why you should never change partition counts on live, offset-dependent topics. - **Speakers:** Dainius Jocas - **Event:** WeAreDevelopers LIVE - **Published:** May 12, 2021 - **Duration:** 24:53 - **URL:** https://www.wearedevelopers.com/videos/138-don-t-change-the-partition-count-for-kafka-topics ## Summary Vinted’s search architecture relies on a data indexing pipeline moving records from MySQL through Kafka and Kafka Connect into Elasticsearch. To handle parallel indexing safely, the engineering team leveraged Elasticsearch's optimistic concurrency control by mapping the Kafka partition offset directly to the Elasticsearch document version number. By using log compaction and infinite retention, this setup theoretically guaranteed worry-free data consistency, ensuring Elasticsearch always stored the most recent listing updates. The architecture worked flawlessly until developers discovered Elasticsearch was silently serving stale data and ignoring tombstone deletion messages. Investigating the inconsistency revealed a bizarre scenario where newer Kafka messages possessed significantly lower offsets than older records stored in Elasticsearch. The root cause traced back to a seemingly routine infrastructure change: an SRE had increased the Kafka topic partition count from 6 to 24 to improve write throughput and node distribution. Because Kafka routes messages by hashing the key modulo the partition count, changing this variable forced existing keys into entirely new partitions with starting offsets near zero. Consequently, Elasticsearch rejected over 95% of subsequent document updates because their new offset numbers were lower than the previously stored versions. Resolving this bug required spinning up a fresh Kafka cluster and executing a full data reinjection from the primary MySQL datastore to restore chronological offset integrity. The critical lesson for distributed system design is clear: never increase partition counts on live Kafka topics if the downstream architecture relies on message ordering or offsets for concurrency control. Instead, provision a sufficiently high partition count from day one to accommodate future scaling. **Keywords:** kafka partition scaling, kafka connect integration, elasticsearch optimistic concurrency, distributed system data consistency, log compaction strategy, tombstone messages handling, kafka message ordering guarantees, offset-based versioning, data indexing pipeline, mysql to elasticsearch synchronization, kafka topic partitioning logic, stale data debugging, infrastructure scaling bugs, full data reingestion ## Chapters 1. **Introduction to the data indexing pipeline** (00:17) — Moving primary data from MySQL to Elasticsearch using Kafka creates a scalable indexing pipeline. 1. **Understanding Elasticsearch and optimistic concurrency control** (04:01) — Using optimistic concurrency control with document version numbers ensures only the newest updates are searchable. 1. **Kafka log compaction and tombstone messages** (05:37) — Configuring topics with infinite retention and log compaction prevents disk exhaustion while enabling safe re-indexing. 1. **Connecting systems safely using Kafka Connect** (07:57) — Using Kafka partition offsets as document versions allows Kafka Connect to safely parallelize indexing. 1. **Investigating reports of stale data in Elasticsearch** (09:35) — A production bug report reveals that tombstone messages are failing to delete obsolete Elasticsearch documents. 1. **Comparing Elasticsearch versions and Kafka offsets** (12:14) — Tracing the issue reveals older Elasticsearch documents possessing higher version numbers than newer Kafka offsets. 1. **Identifying the partition count root cause** (15:27) — Consulting documentation and metrics reveals that increasing the Kafka topic partition count triggered the inconsistency. 1. **How partition changes break message ordering** (16:47) — Altering the partition count changes message hashing, destroying the order guarantees for specific keys. 1. **Fixing the data inconsistency with full ingestion** (20:42) — Resolving the data inconsistency requires fully re-ingesting primary datastore records into newly created Kafka topics. 1. **Establishing best practices for Kafka partition sizing** (22:53) — Setting sensible default partition counts and avoiding modifications prevents data loss when relying on message ordering. ## Related Moments - [Reviewing core Apache Kafka architecture and distributed fundamentals](https://www.wearedevelopers.com/videos/76-how-to-benchmark-your-apache-kafka) (from "How to Benchmark Your Apache Kafka") - [Core concepts of Apache Kafka and topic topologies](https://www.wearedevelopers.com/videos/168-kafka-streams-microservices) (from "Kafka Streams Microservices") - [Hidden costs of self-hosting and managed Kafka solutions](https://www.wearedevelopers.com/videos/1233-python-based-data-streaming-pipelines-within-minutes) (from "Python-Based Data Streaming Pipelines Within Minutes") - [Solving race conditions and distributed state with repartitioning](https://www.wearedevelopers.com/videos/168-kafka-streams-microservices) (from "Kafka Streams Microservices") - [Moving from traditional databases to decoupled event streaming](https://www.wearedevelopers.com/videos/91-from-event-streaming-to-event-sourcing-101) (from "From event streaming to event sourcing 101") - [Introduction to Apache Kafka benchmarking and performance analysis](https://www.wearedevelopers.com/videos/76-how-to-benchmark-your-apache-kafka) (from "How to Benchmark Your Apache Kafka") ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Why Event-Driven Architecture Isn’t About Speed (and When You Actually Need It)](https://www.wearedevelopers.com/magazine/745-why-event-driven-architecture-isn-t-about-speed-and-when-you-actually-need-it) - [The Geometry of Incidents: Connecting User Impact to Architecture](https://www.wearedevelopers.com/magazine/764-the-geometry-of-incidents-connecting-user-impact-to-architecture) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) ## Related Jobs - [Sr. Open Source Software Engineer (Kafka)](https://www.wearedevelopers.com/jobs/48543-sr-open-source-software-engineer-kafka) at **NetApp** - [Staff Systems Engineer](https://www.wearedevelopers.com/jobs/48371-staff-systems-engineer) at **Bitdrift** - [Lead Telemetry Pipeline Specialist](https://www.wearedevelopers.com/jobs/48431-lead-telemetry-pipeline-specialist) at **Dynatrace** - [Principal Systems Engineer, Database Infrastructure](https://www.wearedevelopers.com/jobs/ext/2597685-principal-systems-engineer-database-infrastructure) at **GitHub** - [Staff Systems Engineer, Database Infrastructure](https://www.wearedevelopers.com/jobs/ext/2376353-staff-systems-engineer-database-infrastructure) at **GitHub** - [Principal Backend Engineer, Hub](https://www.wearedevelopers.com/jobs/48449-principal-backend-engineer-hub) at **Docker, Inc.**