How We Cut Our API's p99 Latency from Minutes to Under a Second
-
Deepak Agrawal
Atlassian
Principal Software Engineer
November 25–26, 2026
Bengaluru, India
Joining remotely?
Watch live with ProIn 2021, DigitalOcean’s internal observability hit a breaking point at 5,000 clusters. We were losing critical audit logs. Metrics scraping was failing under sheer volume. Our team faced a choice: stop growing or fundamentally re-architect.
We chose re-architecture. A year later, we’re seamlessly managing 20,000+ production clusters, processing 460+TB with zero log loss and complete metrics coverage. How? By rebuilding our observability stack with OTEL standards, building custom lightweight collectors, and leveraging Kubernetes-native patterns that scale automatically.
You’ll learn: 1. How we leveraged the upstream OTEL Operator to manage OTEL deployments across 20K+ clusters 2. What we got WRONG: First iteration of OTEL blasted our storage with 250M+ log files (we’ll show the mistakes so you don’t repeat them) 3. Operational efficiency: 4 engineers managing 20K+ clusters’ observability (previously 12+ engineers struggling at 5K)
Conference India 2026
Deepak Agrawal
Atlassian
Principal Software Engineer
Pritesh Kiri
Harness
Developer Relations Engineer
Hadar Geva
Myop
CTO & Co-founder
Antonio Mendoza Pérez
Temporal Technologies
Staff Developer Success Engineer
Gaurav Thadani
Temporal
Staff Developer Success Engineer