> Markdown version of [/videos/1191-operating-etcd-for-managed-kubernetes](https://www.wearedevelopers.com/videos/1191-operating-etcd-for-managed-kubernetes). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Operating etcd for Managed Kubernetes Struggling to scale etcd for multi-tenant Kubernetes? Discover how Ionos rebuilt their architecture to eliminate network latency and execute zero-downtime control plane migrations. - **Speakers:** [Mario Valderrama](https://www.wearedevelopers.com/@mario-valderrama), [Tinashe Mundangepfupfu](https://www.wearedevelopers.com/@tinashe-mundangepfupfu) - **Event:** World Congress 2024 - **Published:** August 29, 2024 - **Duration:** 21:43 - **URL:** https://www.wearedevelopers.com/videos/1191-operating-etcd-for-managed-kubernetes ## Summary Scaling a managed Kubernetes service to tens of thousands of clusters requires radically rethinking etcd infrastructure. To minimize costs and operational overhead, Ionos adopted a multi-tenant etcd architecture using client-side name spacing via the API server's etcd prefix flag. Their infrastructure journey evolved from utilizing the CoreOS etcd operator and Bitnami helm charts to implementing a custom, straightforward Helm chart. This architectural shift resolved critical issues surrounding unstable rollouts and complex snapshot restorations that previously required ReadWriteMany storage.<br><br>Operating shared etcd clusters surfaces compounding performance challenges, as heavy traffic from Kubernetes events and leases in one tenant can impact the entire shared system. Continuous system health requires engineering teams to rigorously enable auto-compaction and defragment host memory to prevent key space exhaustion. Additionally, operating multi-tenant clusters demands scaling up parallel operations beyond default limits to accommodate high-write scenarios. Geographically dispersed etcd clusters present another major operational pitfall: network latency inevitably causes Raft revision drift between the leader and followers. Transitioning to a single data center with multiple availability zones, alongside aggressively tuning Raft heartbeat and election timeouts, ensures a much more stable consensus mechanism.<br><br>Executing seamless control plane migrations without global downtime requires an innovative approach to state transfer. By leveraging the `etcdctl make-mirror` command and artificially intercepting and bumping snapshot revision numbers, the team ensured that Kubernetes clients like Kubelet and Calico would not stall out due to revision mismatches upon reconnecting to the data store. Moving forward, the underlying infrastructure is evolving to support fully dedicated etcd clusters alongside the integration of modern control plane managers like Kamaji to further automate tenant isolation and resilience. **Keywords:** managed kubernetes scaling, etcd multi-tenancy, client-side name spacing, kubernetes control plane migration, etcd operator alternatives, custom helm charts, etcd auto-compaction, cluster defragmentation strategy, raft consensus revision drift, etcd network latency, etcdctl make-mirror, zero-downtime cluster migration, raft heartbeat tuning, boltdb storage management, kamaji control plane manager, multi-az etcd deployment ## Chapters 1. **Adopting and scaling managed Kubernetes environments** (00:01) — Tracking the operational metrics of scaling managed infrastructure reveals the centralized importance of the primary datastore. 1. **Overcoming cost constraints with shared etcd instances** (03:27) — Sharing control planes reduces infrastructure overhead and lowers operational hardware costs for multi-tenant environments. 1. **Evaluating etcd deployment operators and helm charts** (04:35) — Migrating from community operators to a simplified custom Helm chart avoids hanging deployments and scripting conflicts. 1. **Addressing performance impacts in multi-tenant etcd setups** (07:20) — High event churn across shared application environments creates compounding technical strain that requires active client bounds. 1. **Tuning database compaction and system defragmentation** (09:24) — Enabling automatic compaction and expanding parallel operation capacity maintains reliable database memory footprint limits. 1. **Mitigating geographical latency in redundant cluster layouts** (10:42) — Resolving internal raft consensus drift requires adjusting heartbeat intervals or reverting to localized network architectures. 1. **Executing zero-downtime cluster control plane migrations** (13:34) — Utilizing dedicated mirror implementations enables continuous data synchronization without disrupting live database client processes. 1. **Manipulating BoltDB snapshot revisions for seamless transitions** (16:18) — Forcing a higher revision key into underlying database snapshots guarantees uninterrupted connections for existing client watches. 1. **Exploring dedicated clusters and alternative deployment models** (19:22) — Transitioning away from manual cluster migrations prompts exploration into dedicated environments and specialized proxy controllers. 1. **Contrasting shared deployment strategies with external tools** (20:44) — Assessing alternative fleet management setups highlights crucial differences in how component statefulness and configurations are implemented. ## Related Moments - [Managing the complexity of bare metal Kubernetes deployments](https://www.wearedevelopers.com/videos/100135-the-new-shiny-syndrome-how-to-avoid-tech-hype-traps) (from "The New Shiny Syndrome: How to Avoid Tech Hype Traps") - [Running Kubernetes clusters efficiently on enterprise cloud platforms](https://www.wearedevelopers.com/videos/89-development-of-reactive-applications-with-quarkus) (from "Development of reactive applications with Quarkus") - [Adapting Kubernetes deployment patterns for heterogeneous edge device fleets](https://www.wearedevelopers.com/videos/100160-from-cloud-racks-to-control-cabinets-operating-kubernetes-on-edge-devices) (from "From Cloud Racks to Control Cabinets: Operating Kubernetes on Edge Devices") - [Customizing operator inputs and overriding default cluster configurations](https://www.wearedevelopers.com/videos/487-debug-a-kubernetes-operator) (from "Debug a Kubernetes Operator") - [Executing massive cloud network migrations while maintaining live systems](https://www.wearedevelopers.com/videos/100128-the-golden-age-of-email-owning-the-inbox-in-the-age-of-ai) (from "The Golden Age of Email: Owning the Inbox in the Age of AI") - [Managing heterogeneous deployment lifecycles via Kubernetes operators](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) (from "Kubernetes and Microservices with Multi-Model Databases") ## Related Articles - [Why Event-Driven Architecture Isn’t About Speed (and When You Actually Need It)](https://www.wearedevelopers.com/magazine/745-why-event-driven-architecture-isn-t-about-speed-and-when-you-actually-need-it) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) ## Related Jobs - [Platform Engineer (DevOps)](https://www.wearedevelopers.com/jobs/48264-platform-engineer-devops) at **WDW Consulting GmbH** - [DevOps Engineer (m/f/d)](https://www.wearedevelopers.com/jobs/48303-devops-engineer-m-f-d) at **basebox GmbH** - [Lead Cloud DevSecOps Engineer - Kubernetes](https://www.wearedevelopers.com/jobs/ext/1659167-lead-cloud-devsecops-engineer-kubernetes) at **BWI GmbH** - [Devops Engineer](https://www.wearedevelopers.com/jobs/ext/1940926-devops-engineer) at **Bitpanda** - [Senior Cloud Native Solution Architect (all genders welcome) - Kubernetes, CNCF, MlOps](https://www.wearedevelopers.com/jobs/ext/101488-senior-cloud-native-solution-architect-all-genders-welcome-kubernetes-cncf-mlops) at **Rosenxt Group** - [Senior Cloud Native Solution Architect (all genders welcome) - Kubernetes, CNCF, MlOps](https://www.wearedevelopers.com/jobs/ext/66342-senior-cloud-native-solution-architect-all-genders-welcome-kubernetes-cncf-mlops) at **Rosenxt Group**