World Congress 2026 North America

Boring Failover: Predictable Region Recovery Across 5,000 Microservices

September 24, 2026 17:30 – 18:00 Β· 30 min Stage 4

What this session covers

At Intuit, 5,000+ services run at peak 1M+ TPS across TurboTax, QuickBooks, Credit Karma, and Mailchimp. Eighty percent are multi-region. Before EWOK β€” our Ecosystem Wide Orchestrator Kit β€” disaster recovery meant thousands of non-standard DR scripts, per-team runbooks, and 8,000+ engineers each solving the same problems differently. Game days were feared. MTTR was unpredictable. DR was treated as β€œthe database snapshot” β€” not the full stack.

This talk is the story of how we made regional failover boring β€” predictable, repeatable, automated, no heroics. I will walk through the architecture and the hard lessons:

β€’ A declarative YAML DSL for DR plans β€” stages, parallel blocks, per-stage agent versioning (armador/v1, database/v1, route53/v1), inline IAM role assumption.

β€’ AWS Step Functions as the DAG orchestrator β€” durable state for long-running promotions (Aurora global cluster failover, Redis replication switch), and a visual audit log that doubles as the incident timeline.

β€’ A Golang control plane on Kubernetes running goroutine-parallel mutations across thousands of namespaces.

β€’ A Python Agent Framework with an ABC contract β€” PreCheck, Failover, PostCheck β€” encapsulating IAM, logging, metrics. Product teams shipped a Redis agent in days without platform bottleneck.

β€’ Specialized agents per layer: armador (compute/capacity), Route53 (DNS cutover), Database (Aurora global failover), Redis (flush + replication).

β€’ The multi-workload problem everyone skips β€” cron jobs, async consumers, stateful tiers, caches. Parallel suspend, scale, resume, dial.

β€’ Progressive Dial β€” incremental traffic shift with error-gated automatic rollback.

β€’ Auto Failover via our Alert2Incident framework.

β€’ Same machinery for migrations β€” same-region failover as an upgrade feature.

You leave with: concrete patterns for declarative DR, a replicable agent contract for inner-sourcing reliability, and a checklist of the workload types your DR plan probably does not cover yet.

Related talks at this congress

Open session

World Congress 2026 North America

September 25, 2026 Β· 12:55–13:25

Stage 3

Replay-Safe Architecture: Building Event-Driven Systems That Can Recover With Confidence

Ishan Shah

PayPal, Software Engineer | Distributed Systems, AI, and Platform Engineering

Ishan Shah
Open session

World Congress 2026 North America

September 25, 2026 Β· 16:50–17:20

Stage 3

Your Evals Passed. Your Agent Just Emptied a Database.

Tejas Pravinbhai Patel

IEEE Award-Winning Researcher | Best Keynote Speaker | Sr. Software Engineer at Amazon | AI Systems & Agent Architect

Tejas Pravinbhai Patel
Open session

World Congress 2026 North America

September 25, 2026 Β· 15:30–16:00

Stage 1

Agents Can't Iterate Against Tests That Lie

Rocky Warren

Senior Staff Software Engineer at Clipboard

Rocky Warren
Open session

World Congress 2026 North America

September 24, 2026 Β· 16:50–17:20

Stage 2

It passed auth, then production caught fire

Alex Olivier

Co-founder & CPO @ Cerbos | OpenID AuthZEN Co-chair

Alex Olivier
All sessions at this congress