> Markdown version of [/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure How did a five-person team scale infrastructure for 84 million users? Discover Yubo's lean GitOps strategy for conquering technical debt and slashing observability costs. - **Speakers:** [Mikael Robert](https://www.wearedevelopers.com/@mikael-robert) - **Event:** World Congress 2025 - **Published:** August 20, 2025 - **Duration:** 17:44 - **URL:** https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure ## Summary Yubo, a real-time social network with 84 million users, faced overwhelming technical debt, uncontrolled scaling costs, and a massive legacy Kubernetes infrastructure—all managed by a five-person team. Recognizing that constraint drives innovation, they shifted their operations strategy by internalizing data engineering within the core infrastructure team. This lean approach allowed them to optimize data architecture early in the feature-design phase, drastically reducing production database scaling alerts and accelerating development.\n\nBy standardizing their vast environment of databases and microservices, Yubo eliminated manual deployment in favor of customized Kubernetes custom resource definitions and Argo workflows. They established a robust configuration management database driven entirely by YAML, enabling complete developer autonomy without exposing core templates. Applications build, test, and deploy via standardized pipelines, while downstream, data replication is managed through change data capture pipelines streaming via Kafka. This decoupled model ensures that every microservice has fresh data without burdening a monolithic database, while their machine learning backend is managed with the exact same GitOps methodologies, differentiated only by GPU provisioning.\n\nHandling real-time metrics and logs for up to 300K database requests per second required a rigorous approach to observability to avoid crippling SaaS costs. The infrastructure team adopted dual shipping and metric compression as key cost-control strategies. By routing summarized, high-value daily metrics into Datadog and Amplitude, whilst pushing raw telemetry and high-cardinality traces into scalable open-source tools like VictoriaMetrics, Tempo, and Grafana, they struck a perfect balance between platform usability and tight budgets. This setup provides targeted dashboards for diverse stakeholders while maintaining full debugging capabilities based on lifecycle-aware data handling. **Keywords:** infrastructure scaling, gitops methodology, argo workflows, kubernetes operators, developer experience, self-service deployments, configuration management database, change data capture, event sourcing pattern, machine learning deployment, observability cost optimization, metric compression, high cardinality data, internal developer portal, database scalability, datadog integration, grafana correlation ## Chapters 1. **Operating global infrastructure with tiny engineering teams** (00:05) — Operating a massive social platform directly with just five individuals highlights how strict operational constraints force valuable automation patterns. 1. **Overcoming operational debt in unmanaged giant clusters** (01:48) — Moving beyond unmanaged single-cluster environments and unstructured hardware establishes the foundation for scalable declarative system management. 1. **Embedding data engineering to solve database scalability** (02:48) — Integrating data engineers directly alongside infrastructure early in feature development prevents production bottlenecks by anticipating specific indexing complexities. 1. **Automating delivery workflows via centralized gitops patterns** (04:43) — Adopting central repositories synchronized via automated deployment controllers prevents manual configuration drift across growing deployment states. 1. **Empowering developer autonomy with custom resource definitions** (05:37) — Masking intricate backend mechanisms behind custom orchestration boundaries securely empowers engineers to manage application delivery without bottlenecking infrastructure experts. 1. **Standardizing continuous integration pipelines across diverse languages** (09:03) — Orchestrating code compilation builds via unified declarative template configurations accelerates safe feature turnover for autonomous development squads. 1. **Distributing feature data asynchronously using change data capture** (10:00) — Allowing controlled datastore duplication via asynchronous streaming event feeds protects central data foundations from aggressive real-time querying bursts. 1. **Normalizing machine learning deployments via standard operational pipelines** (11:31) — Managing distinct inference algorithms through unified deployment lifecycles standardizes rollout methodologies alongside normal web framework release streams. 1. **Launching independent architectures rapidly utilizing global modular templates** (12:21) — Encoding comprehensive infrastructure layouts within unified configuration properties enables identical application verticals to physically materialize instantly. 1. **Combining saas and open-source observability for extreme scale** (13:12) — Blending highly visual premium interfaces with massive open-source data layers balances deep metric analysis against staggering scaling costs. 1. **Controlling observability costs through intelligent granular telemetry routing** (14:29) — Downsampling specific analytics for premium interfaces while retaining unabridged raw trace payloads natively solves financial forecasting realities without limiting resolution speed. ## Related Moments - [Managing high-traffic infrastructure without chasing technology hype](https://www.wearedevelopers.com/videos/1817-how-to-avoid-tech-hype-traps-josip-stuhli) (from "How to Avoid Tech Hype Traps - Josip Stuhli") - [Extreme engineering culture for massive web data operations](https://www.wearedevelopers.com/videos/100244-marketing-x-product-how-we-stopped-gaslighting-each-other-and-built-ai-products-that-actually-work) (from "Marketing x Product: How We Stopped Gaslighting Each Other and Built AI Products That Actually Work") - [Solving complex platform architecture challenges at an enterprise scale](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) (from "Coffee with Developers - Maria Apazoglou") - [Refactoring a complex legacy monolith into microservices](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) (from "Retooling and refactoring - an investment in people.") - [Modern application stacks and real-time data requirements](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) (from "Leveraging Real time data in FSIs") - [Boosting developer productivity via consolidated converged database architectures](https://www.wearedevelopers.com/videos/632-crypto-secure-data-management-with-in-database-blockchain) (from "Crypto-secure Data Management with In-Database Blockchain") ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [How Microsoft worked around a Git limitation to shrink a repository by 94%](https://www.wearedevelopers.com/magazine/499-how-microsoft-worked-around-a-git-limitation-to-shrink-a-repository-by-94) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) ## Related Jobs - [Lead Software Engineer - Data Engineering](https://www.wearedevelopers.com/jobs/ext/2000968-lead-software-engineer-data-engineering) at **Dynatrace** - [Staff Software Engineer, Database Infrastructure](https://www.wearedevelopers.com/jobs/ext/1470125-staff-software-engineer-database-infrastructure) at **GitHub** - [Principal Software Engineer, Database Infrastructure](https://www.wearedevelopers.com/jobs/ext/1465908-principal-software-engineer-database-infrastructure) at **GitHub** - [Staff Software Engineer](https://www.wearedevelopers.com/jobs/ext/1425755-staff-software-engineer) at **GitHub** - [Software Engineer, Platform Engineering (L2)](https://www.wearedevelopers.com/jobs/ext/1956829-software-engineer-platform-engineering-l2) at **Twilio** - [Senior Software Engineer, Client Apps Platform](https://www.wearedevelopers.com/jobs/ext/1773893-senior-software-engineer-client-apps-platform) at **GitHub**