We Deleted Our Observability Stack and Rebuilt It with OTEL: 12 Engineers to 4 at 20K+ clusters
-
Yash Sharma
DigitalOcean
Developer Advocate
November 25–26, 2026
Bengaluru, India
Joining remotely?
Watch live with ProSix months ago, the most critical API in our platform had p99 latency that regularly spiked into minutes. Our primary database CPU was at capacity. Customers were timing out. We were three engineers staring at dashboards full of red.
This talk is the story of how we traced that meltdown to six compounding failures — and fixed all of them. Not one silver bullet, but six layers of the stack quietly making each other worse.
I’ll walk through each failure and the fix, with real metrics:
Queries running without statement_timeout — a single slow scan holding connections for 90+ seconds while the pool starved
CompletableFuture.supplyAsync() on Java’s shared ForkJoinPool — blocking JDBC calls saturating 11 threads and silently losing our @ReadOnly ThreadLocal, routing every read to the primary while the replica sat idle
Composite indexes on the wrong column order, dead JOINs the ORM added that the query never needed, and aggregations spilling to disk
A cache miss thundering-herd that turned a cold restart into a self-inflicted DDoS
For each one, I’ll show the before-and-after: the EXPLAIN ANALYZE output, the thread dump, the Grafana panel. I’ll demo a set of regex patterns we now run in CI to catch these anti-patterns before they reach production.
You’ll leave with a checklist of ten things to audit in your own Spring Boot + PostgreSQL services tonight — and the grep commands to do it.
Conference India 2026
Yash Sharma
DigitalOcean
Developer Advocate
Pritesh Kiri
Harness
Developer Relations Engineer
Hadar Geva
Myop
CTO & Co-founder
Andrei Tazetdinov
Dynatrace
React Native Developer