> Markdown version of [/events/world-congress-2026-north-america/sessions/1722-the-geometry-of](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1722-the-geometry-of). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # The Geometry of Incidents: What User-Impact Shapes Reveal About Platform Architecture - **Event:** World Congress 2026 North America ## Description Incidents are usually reviewed as timelines: what failed, who owned it, and how we restored service. That works well for understanding a single outage. But when you operate platforms used by many services, the root causes change while the user-impact patterns start to look familiar. This talk introduces incident shapes: a way to look at failures by the pattern they draw across impact, time, and blast radius. I’ll use a few public incidents as reference points, including CrowdStrike’s outage, AWS’s DynamoDB outage where retry amplification played a role, Cloudflare’s WAF incident, and GitHub’s 2018 database failover incident. The incident shapes are useful because they change the questions we ask. A sudden spike makes us look at rollout containment and rollback paths. A slow burn pushes us to examine saturation, queues, and retries. A repeating sawtooth suggests the system may be recovering temporarily without becoming stable. Fan-out patterns expose the risk of shared platform layers. Boundary shifts are often the hardest to catch: one layer reports success, while users are still having a bad experience. The main idea is simple: the shape of user impact can tell us what the what the architecture failed to protect against. Attendees will learn how to quantify impact using breadth, depth, and duration, and how to connect those shapes to engineering responses such as staged rollouts, rollback automation, retry budgets, semantic canaries, cell isolation, contract checks, and end-to-end verification. The goal is to make postmortems more useful: not just to explain what happened, but to help design platforms that are harder to break in the same way twice. ## Speaker ### [Bala Subrahmanyam Kambala](https://www.wearedevelopers.com/@bala-subrahmanyam-kambala) Staff Platform Engineer at Oracle Cloud Infrastructure ## Related talks at this congress - [When Logging Becomes The Outage: Escaping the ECS Logging Trap](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1670-when-logging-becomes) — Rahul Tanniru - [Microservice Cognitive Index for Deploy Diagnosis and Change Impact](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1423-microservice) — Sachin Gupta - [Reinventing Incident Response with AI Agents and MCP](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1429-reinventing-incident) — Jayant Tyagi - [Boring Failover: Predictable Region Recovery Across 5,000 Microservices](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1698-boring-failover) — Garvit Kataria, Sahil Sabharwal ## Watch remotely Can’t make it to San José? Watch this session live with Pro. You also get: - All full videos, bookmarks, and playlists - World Congress livestreams [See pricing](https://www.wearedevelopers.com/pricing) ## Links - [Get tickets](https://www.wearedevelopers.com/world-congress-north-america/tickets)