World Congress 2025 Aug 20, 2025 Session details

Planet-Scale Dashboards

Robert Lehmann

Stop manually copying dashboards for new microservices. Google solved this at a massive scale. Learn how dynamic service discovery enables automatic, context-aware observability.

Pause
Mute Enter Fullscreen
#1 about 3 min

Challenges of manual observability setups during service outages

The difficulties of configuring and utilizing manual monitoring tools while attempting to resolve active production incidents.

#2 about 3 min

Building global monitoring solutions within a monorepo architecture

How global data centers and single-repository structures shape the development of internal observability platforms.

#3 about 3 min

Tackling dashboard divergence across independent microservices

The operational impact and troubleshooting difficulties that arise when different service teams configure metrics independently.

#4 about 4 min

Standardizing reusable dashboards with template variables and dimensions

Leveraging dimension parameters to maintain consistency and establish deterministic links across different service instances.

#5 about 3 min

Filtering dashboard noise for context-specific service observability

The challenge of overwhelming metric views when reusing template files across distinct technological stacks.

#6 about 4 min

Automating dashboard relevance through service discovery and traits

Injecting scope contexts and dynamically discovering application settings to surface relevant metric displays automatically.

#7 about 2 min

Managing heterogeneous architecture entities using generic scope types

Organizing distinct infrastructural components like databases and machine learning models into unified monitoring namespaces.

#8 about 4 min

Limitations of infrastructure and dashboards as code approaches

Why declarative configuration files often fall out of sync with actual running application metrics.

#9 about 4 min

Handling massive global scale using pre-aggregation and graph variants

Swapping metric queries dynamically based on current filtering dimensions to maintain fast performance globally.

#10 about 2 min

Leveraging scope topology for systemic incident troubleshooting

Connecting related application tiers to intuitively guide operating teams through dependent services during an outage.

Matching moments

17:03 min

Introduction to metrics and observability challenges in monitoring

Liam Hurrell · LIVE

55 sec

Understanding observability through a practical dashboard analogy

Jemiah Sius Jemiah Sius · WWC 2023

1:59 min

Transitioning from reactive to proactive platform infrastructure governance

Neena Thomas Neena Thomas · WWC Europe 2026

41 sec

Operating critical stack observability and service monitoring

Michael Cade · LIVE

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

7:57 min

Visualizing test stability using service level agreement dashboards

Liam Hurrel · LIVE

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Microservice Cognitive Index for Deploy Diagnosis and Change Impact

Sachin Gupta

Member of Technical Staff 2 at eBay

Sachin Gupta
Open session

World Congress 2026 North America

The Geometry of Incidents: What User-Impact Shapes Reveal About Platform Architecture

Bala Subrahmanyam Kambala

Staff Platform Engineer at Oracle Cloud Infrastructure

Bala Subrahmanyam Kambala
Open session

World Congress 2026 North America

Boring Failover: Predictable Region Recovery Across 5,000 Microservices

Garvit Kataria, Sahil Sabharwal

Garvit Kataria
Sahil Sabharwal
Open session

World Congress 2026 North America

AI-Powered Incident Triage: How We Built GenAI Agents with MCPs to Automate On-Call Workflows

Prakshal Doshi

Site Reliability Engineer

Prakshal Doshi
Open session

World Congress 2026 North America

Closing the Visibility Gap: Lessons from Safety Critical Agentic Systems

Vivek Pandit

Principal Engineer at Cadence

Vivek Pandit
Open session

World Congress 2026 North America

From Static Rules to Reasoning Platforms: Scaling Intelligent Canary Delivery in 2026

Daniel Oh

Senior Principal Developer Advocate

Daniel Oh