World Congress 2026 Europe - Virtual Stage Jul 2, 2026 Session details

Designing Reliable Distributed Systems: Failures, Retries & Idempotency

Violetta Pidvolotska

If your architectural diagram lacks failure paths, it depicts a demo, not a production system. Learn to prevent cascading outages using explicit timeouts, idempotency, and circuit breakers.

Pause
Mute Enter Fullscreen
#1 about 4 min

Identifying invisible reliability requirements in system design

Designing for happy paths without defining failure expectations sets teams up for inevitable production incidents.

#2 about 3 min

Diagnosing silent system degradation and resource exhaustion

Infinite network waits cause threads and memory to silently run out while monitoring dashboards appear healthy.

#3 about 4 min

Setting explicit network timeouts to create failure boundaries

Defining strict wait budgets based on real latency data prevents localized dependency slowness from spreading system-wide.

#4 about 3 min

Uncovering transient network errors by implementing fast failures

Adding strict network timeouts converts silent resource hanging into visible short-lived errors that can be explicitly handled.

#5 about 5 min

Preventing system retry storms using backoff and jitter

Implementing exponential backoff and randomized jitter prevents synchronized client retries from overwhelming struggling dependencies.

#6 about 3 min

Implementing fallback strategies for external dependency outages

Designing explicit secondary paths ensures acceptable service degradation when primary dependencies fail entirely.

#7 about 4 min

Preventing duplicate requests during unknown network states

Treating network timeouts as absolute failures leads to duplicate processing since the initial request may have succeeded.

#8 about 3 min

Building idempotent network operations to ensure safe retries

Requiring deterministic idempotency keys representing business intent ensures operations safely process exactly once regardless of retries.

#9 about 5 min

Solving concurrent race conditions with deterministic payload hashing

Using unique database constraints and deterministic hashing prevents double processing during concurrent requests and key reuse.

#10 about 4 min

Managing response caching to enable valid request retries

Storing only successful service responses ensures users can recover from temporary validation or insufficient funds errors.

#11 about 3 min

Protecting infrastructure using circuit breakers against persistent failures

Tripping a circuit breaker halts futile retry attempts but forces synchronous systems to reject all incoming traffic.

#12 about 3 min

Decoupling synchronous system dependencies using asynchronous message queues

Transitioning critical synchronous workflows to queue-based state machines preserves business operations during extended downstream provider outages.

#13 about 3 min

Adopting mental models for resilient distributed system architecture

Anticipating partial failures and understanding the true cost of mitigation strategies creates robust engineering design cultures.

Matching moments

2:20 min

Addressing the fallacies of distributed computing networks

Alexander Reelsen · LIVE

5:25 min

Implementing redundancy, failover, and architectural load balancing patterns

Mihaela-Roxana Ghidersa · LIVE

2:52 min

Mitigating latent system errors and designing for resilience

David Singleton David Singleton +1 · Coffee With Developers

3:26 min

Recognizing the danger of silent failures in resilient systems

Lyubomir Bozhinov Lyubomir Bozhinov · Europe 2026 Virtual

3:49 min

Enhancing system resilience and fault tolerance

Michael Eder +1 · LIVE

2:55 min

Configuring resilience pipelines and exponential backoff retry strategies

Sander ten Brinke Sander ten Brinke · World Congress 2025

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 25, 2026 · 12:55–13:25

Stage 3

Replay-Safe Architecture: Building Event-Driven Systems That Can Recover With Confidence

Ishan Shah

PayPal, Software Engineer | Distributed Systems, AI, and Platform Engineering

Ishan Shah
Open session

World Congress 2026 North America

September 24, 2026 · 12:50–13:20

Stage 3

Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases

Wei Hu

Senior Vice President of Research and Development

Wei Hu
Open session

World Congress 2026 North America

September 24, 2026 · 17:30–18:00

Stage 4

Boring Failover: Predictable Region Recovery Across 5,000 Microservices

Garvit Kataria, Sahil Sabharwal

Garvit Kataria
Sahil Sabharwal
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 3

Real-Time Data Platforms at Trillion-Event Scale

Diptamay Sanyal

Principal Engineer | Data, AI & Cybersecurity Platforms

Diptamay Sanyal
Open session

World Congress 2026 North America

September 25, 2026 · 12:20–12:50

Stage 9

Designing APIs That Survive AI Agents at Scale

Phani Pendurthi

Mastercard, Principal Software Engineer

Phani Pendurthi
Open session

World Congress 2026 North America

September 24, 2026 · 14:50–15:20

Stage 9

How to generate business value through performance optimizations

Nikolai Sidiropulo

Software Engineer at Meta

Nikolai Sidiropulo