WeAreDevelopers LIVE Mar 22, 2024

System Resilience: Surviving the Software Storm

Mihaela-Roxana Ghidersa

Stop relying on reactive firefighting to handle unexpected system outages. Discover how to build fault-tolerant architectures and resilient engineering teams that prevent catastrophic user-facing failures.

Pause
Mute Enter Fullscreen
#1 about 4 min

Understanding system resilience and the costs of failure

Unexpected traffic spikes and critical component failures lead to significant financial and brand damage if systems are not properly hardened.

#2 about 5 min

Differentiating between hidden software faults and complete failures

Unnoticed code issues hiding in complex architectures escalate into user-facing failures under rare conditions unless mitigated by fault tolerance.

#3 about 5 min

Navigating complexity and anti-patterns in modern software architecture

Over-engineering or incorrectly applying distributed network patterns creates error-prone infrastructures and difficult state management rather than actual resilience.

#4 about 4 min

Building robust system resilience across all stack layers

Coordinating infrastructure backups, graceful application degradation, database replication, and cohesive team communication forms a truly robust overall system.

#5 about 6 min

Implementing redundancy, failover, and architectural load balancing patterns

Applying application decoupling, availability zones, and traffic distribution protects systems from targeted failure clustering while emphasizing contextual architectural fitness.

#6 about 3 min

Crafting an effective disaster recovery and communication plan

Conducting organizational risk assessments and prioritizing critical asset recovery limits the damage of unexpected disruptions like cyber attacks or hardware failure.

#7 about 3 min

Applying secure coding practices and proactive system monitoring

Training developers to mitigate common application vulnerabilities and implementing continuous system scanning prevents minor disruptions from escalating into major operational outages.

#8 about 5 min

Maintaining continuous testing and learning from system failures

Automating quality checks and analyzing past incidents establishes actionable technical feedback loops that constantly harden systems against future defects.

#9 about 3 min

Focusing on core software quality attributes for foundational resilience

Focusing on baseline performance, security, and maintainability metrics empowers individual engineers to improve structural architectural strength regardless of their overarching organizational influence.

#10 about 3 min

Embracing machine learning and predictive system behavior analytics

Fostering a continuous learning mindset prepares engineering teams to leverage predictive data models to anticipate and prevent application downtime under production stress.

#11 about 4 min

Balancing strict security practices with application performance requirements

Aligning technical trade-offs with specific business priorities guarantees that robust structural data protection correctly supports the intended overall system business goals without workflow friction.

#12 about 5 min

Prioritizing critical software components for targeted resiliency infrastructure

Applying the shearing layers component concept aids technical leaders in designing flexible application software configurations supported by highly stable architectural foundational investments.

#13 about 2 min

Pitching technical resiliency initiatives to business decision makers

Documenting specific user service availability risks alongside projected financial impact estimates builds compelling organizational arguments for prioritizing deep structural platform improvements over rapid end-user feature delivery.

#14 about 4 min

Cultivating passion for code quality and continuous learning

Experiencing resilient team engineering cultures that deeply value intentional database schema design and scalable coding components translates daily programming routines into sustained professional motivation.

Matching moments

4:40 min

Defining software resilience and layers of system architecture

Sander ten Brinke Sander ten Brinke · WWC 2025

3:00 min

Identifying examples of system resilience and fragility in technology

Jan de Vries Jan de Vries · WWC Europe 2026

2:52 min

Mitigating latent system errors and designing for resilience

David Singleton David Singleton +1 · Coffee With Developers

1:05 min

Differentiating fragile, robust, resilient, and antifragile systems

Jan de Vries Jan de Vries · WWC Europe 2026

3:26 min

Adopting practical engineering standard changes for resilient system architectures

Aaditya Binod Yadav Aaditya Binod Yadav · Europe 2026 Virtual

3:39 min

The growing cost of system outages and microservice failures

Sander ten Brinke Sander ten Brinke · WWC 2025

Upcoming sessions on this topic

Open session

World Congress 2026 North America

The Geometry of Incidents: What User-Impact Shapes Reveal About Platform Architecture

Bala Subrahmanyam Kambala

Staff Platform Engineer at Oracle Cloud Infrastructure

Bala Subrahmanyam Kambala
Open session

World Congress 2026 North America

From Legacy to Longevity – How a 17-Year-Old System stayed Modern

Alexander Arians

Empowering Engineers to drive innovation

Alexander Arians
Open session

World Congress 2026 North America

Boring Failover: Predictable Region Recovery Across 5,000 Microservices

Garvit Kataria, Sahil Sabharwal

Garvit Kataria
Sahil Sabharwal
Open session

World Congress 2026 North America

Reinventing Testing Practices in the AI Era

Eric Deandrea

Java Champion & Senior Principal Software Engineer, IBM

Eric Deandrea
Open session

World Congress 2026 North America

It passed auth, then production caught fire

Alex Olivier

Co-founder & CPO @ Cerbos | OpenID AuthZEN Co-chair

Alex Olivier
Open session

World Congress 2026 North America

The Broken Rung: How AI is Rebuilding Software Development from the Ground Up

Tomislav Tipurić

Chief Technology Officer, Nephos

Tomislav Tipurić