World Congress 2026 Europe - Virtual Stage Jul 2, 2026 Session details

Designing Reliable Distributed Systems: Failures, Retries & Idempotency

Violetta Pidvolotska

If your architectural diagram lacks failure paths, it depicts a demo, not a production system. Learn to prevent cascading outages using explicit timeouts, idempotency, and circuit breakers.

Pause
Mute Enter Fullscreen
#1 about 4 min

Identifying invisible reliability requirements in system design

Designing for happy paths without defining failure expectations sets teams up for inevitable production incidents.

#2 about 3 min

Diagnosing silent system degradation and resource exhaustion

Infinite network waits cause threads and memory to silently run out while monitoring dashboards appear healthy.

#3 about 4 min

Setting explicit network timeouts to create failure boundaries

Defining strict wait budgets based on real latency data prevents localized dependency slowness from spreading system-wide.

#4 about 3 min

Uncovering transient network errors by implementing fast failures

Adding strict network timeouts converts silent resource hanging into visible short-lived errors that can be explicitly handled.

#5 about 5 min

Preventing system retry storms using backoff and jitter

Implementing exponential backoff and randomized jitter prevents synchronized client retries from overwhelming struggling dependencies.

#6 about 3 min

Implementing fallback strategies for external dependency outages

Designing explicit secondary paths ensures acceptable service degradation when primary dependencies fail entirely.

#7 about 4 min

Preventing duplicate requests during unknown network states

Treating network timeouts as absolute failures leads to duplicate processing since the initial request may have succeeded.

#8 about 3 min

Building idempotent network operations to ensure safe retries

Requiring deterministic idempotency keys representing business intent ensures operations safely process exactly once regardless of retries.

#9 about 5 min

Solving concurrent race conditions with deterministic payload hashing

Using unique database constraints and deterministic hashing prevents double processing during concurrent requests and key reuse.

#10 about 4 min

Managing response caching to enable valid request retries

Storing only successful service responses ensures users can recover from temporary validation or insufficient funds errors.

#11 about 3 min

Protecting infrastructure using circuit breakers against persistent failures

Tripping a circuit breaker halts futile retry attempts but forces synchronous systems to reject all incoming traffic.

#12 about 3 min

Decoupling synchronous system dependencies using asynchronous message queues

Transitioning critical synchronous workflows to queue-based state machines preserves business operations during extended downstream provider outages.

#13 about 3 min

Adopting mental models for resilient distributed system architecture

Anticipating partial failures and understanding the true cost of mitigation strategies creates robust engineering design cultures.

Matching moments

2:20 min

Addressing the fallacies of distributed computing networks

Alexander Reelsen · LIVE

5:25 min

Implementing redundancy, failover, and architectural load balancing patterns

Mihaela-Roxana Ghidersa · LIVE

2:52 min

Mitigating latent system errors and designing for resilience

David Singleton David Singleton +1 · Coffee With Developers

3:26 min

Recognizing the danger of silent failures in resilient systems

Lyubomir Bozhinov Lyubomir Bozhinov · Europe 2026 Virtual

3:49 min

Enhancing system resilience and fault tolerance

Michael Eder +1 · LIVE

2:55 min

Configuring resilience pipelines and exponential backoff retry strategies

Sander ten Brinke Sander ten Brinke · WWC 2025

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases

Wei Hu

Senior Vice President of Research and Development

Wei Hu
Open session

World Congress 2026 North America

Boring Failover: Predictable Region Recovery Across 5,000 Microservices

Sahil Sabharwal, Garvit Kataria

Sahil Sabharwal
Garvit Kataria
Open session

World Congress 2026 North America

Designing APIs That Survive AI Agents at Scale

Phani Pendurthi

Mastercard, Principal Software Engineer

Phani Pendurthi
Open session

World Congress 2026 North America

How to generate business value through performance optimizations

Nikolai Sidiropulo

Software Engineer at Meta

Nikolai Sidiropulo
Open session

World Congress 2026 North America

It passed auth, then production caught fire

Alex Olivier

Co-founder & CPO @ Cerbos | OpenID AuthZEN Co-chair

Alex Olivier
Open session

World Congress 2026 North America

When Logging Becomes The Outage: Escaping the ECS Logging Trap

Rahul Tanniru

Senior Vice President Of Software Engineering, Jp Morgan Chase

Rahul Tanniru