World Congress 2023 Sep 21, 2023

Handling incidents collaboratively is like solving a rubix cube

Nele Uhlemann

When systems crash, isolated engineering teams scramble on different sides of the Rubik's cube. Discover how collaborative incident response untangles silos and speeds up recovery.

Pause
Mute Enter Fullscreen
#1 about 4 min

Navigating specialized roles and toolsets across engineering teams

Incident resolution requires cross-functional collaboration because backend developers, SREs, and DevOps professionals primarily operate within domain-specific workflows.

#2 about 3 min

Creating a shared understanding to resolve incidents quickly

Centralized platforms like shared notebooks establish data transparency to help diverse teams align on causality and fast fixes.

#3 about 3 min

Preventing future incidents through proper retry strategies

Establishing cross-team best practices determines whether network retries belong in the service mesh or the backend application codebase.

#4 about 2 min

Discovering incidents using logs, metrics, and traces

Instrumenting application code to expose continuous telemetry data allows operational teams to detect latency outliers and trace root causes.

#5 about 3 min

Standardizing telemetry extraction using the OpenTelemetry framework

Adopting a vendor-neutral standard simplifies application instrumentation while allowing infrastructure teams to switch backend storage layers effortlessly.

#6 about 2 min

Generating instant observability metrics with the Autometrics library

Function-level code decorators automatically produce latency, error, and request rate metrics without requiring developers to leave their IDE.

#7 about 4 min

Instrumenting a Python backend application with Prometheus metrics

Applying the Autometrics decorator to a FastAPI endpoint automatically exposes standard metrics for local querying and developer inspection.

#8 about 3 min

Defining programmatic service level objectives for specific endpoints

Attaching success rate objectives directly in application code translates logic-level expectations into actionable alerts on operational dashboards.

#9 about 4 min

Balancing tool consolidation and collaborative open source workflows

Embracing specific technologies for unique use cases while maintaining shared communication standards fosters effective cross-team incident management.

Matching moments

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

3:53 min

Applying software development methodologies to incident response

Tobias Dunn-Krahn · LIVE

1:57 min

Streamlining incident response and root cause analysis automatically

Mike Mike · WWC 2025

1:51 min

Leveraging scope topology for systemic incident troubleshooting

Robert Lehmann Robert Lehmann · WWC 2025

6:03 min

Engaging software developers deeply in secure engineering practices

Tanya Janca · WWC 2021

Upcoming sessions on this topic

Open session

World Congress 2026 North America

The Geometry of Incidents: What User-Impact Shapes Reveal About Platform Architecture

Bala Subrahmanyam Kambala

Staff Platform Engineer at Oracle Cloud Infrastructure

Bala Subrahmanyam Kambala
Open session

World Congress 2026 North America

Reinventing Incident Response with AI Agents and MCP

Jayant Tyagi

Lead Member of Technical Staff @ Salesforce

Jayant Tyagi
Open session

World Congress 2026 North America

AI-Powered Incident Triage: How We Built GenAI Agents with MCPs to Automate On-Call Workflows

Prakshal Doshi

Site Reliability Engineer

Prakshal Doshi
Open session

World Congress 2026 North America

Microservice Cognitive Index for Deploy Diagnosis and Change Impact

Sachin Gupta

Member of Technical Staff 2 at eBay

Sachin Gupta
Open session

World Congress 2026 North America

Closing the Visibility Gap: Lessons from Safety Critical Agentic Systems

Vivek Pandit

Principal Engineer at Cadence

Vivek Pandit
Open session

World Congress 2026 North America

Boring Failover: Predictable Region Recovery Across 5,000 Microservices

Garvit Kataria, Sahil Sabharwal

Garvit Kataria
Sahil Sabharwal