> Markdown version of [/videos/482-building-a-culture-from-chaos?t=111](https://www.wearedevelopers.com/videos/482-building-a-culture-from-chaos?t=111). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Building a culture from chaos Stop assuming architectural resilience. Intentionally break your systems to uncover hidden failures. Discover how chaos engineering builds both robust microservices and a psychologically safe organizational culture. - **Speakers:** [Steve Upton](https://www.wearedevelopers.com/@steve-upton) - **Event:** World Congress 2022 - **Published:** June 15, 2022 - **Duration:** 44:44 - **URL:** https://www.wearedevelopers.com/videos/482-building-a-culture-from-chaos ## Summary Tracing its roots to Netflix's cloud migration and microservices transformation, chaos engineering emerged as a disciplined approach to uncovering single points of failure. Rather than assuming architectural resilience, engineers utilized frameworks like Chaos Monkey to continuously validate their systems by intentionally inducing turbulence. This practice established a core engineering loop: trigger an experiment, learn from the resulting failure, and iteratively improve system hardiness. Beyond infrastructural hardening, modern technical environments act as complex adaptive systems characterized by unpredictable interdependence and non-linear relationships. In these shifting ecosystems, traditional "big design up-front" methodologies inevitably fail because the system constantly alters in response to technical and human inputs. Navigating this reality requires substituting absolute predictability with agile responses to change, recognizing that unexpected consequences are a natural, unavoidable byproduct of complex environments. Building a resilient engineering culture requires more than corporate value statements—it demands making safe-to-fail experimentation an organizational habit. By designing targeted probes, carefully limiting the blast radius, ensuring deep observability, and defining clear rollback strategies, engineering teams can sustainably test their boundaries. While the first-order effect of chaos engineering is demonstrably improved backend resilience, its profound second-order effect is transforming how a company culturally addresses complexity, psychological safety, and continuous adaptation. **Keywords:** chaos engineering methodologies, complex adaptive systems, cloud infrastructure migration, microservices architecture resilience, single points of failure, continuous resilience testing, chaos monkey framework, safe-to-fail experiments, limiting blast radius, system observability and monitoring, organizational culture transformation, learning from technical failures, agile change management, big design up-front anti-pattern, rollback and dampening strategies ## Chapters 1. **Transitioning architecture to microservices at Netflix** (01:51) — Splitting monolithic applications into microservices exponentially increases points of failure and demands resilient architecture. 1. **Testing system resilience with the Simian Army** (04:59) — Simulating random server shutdowns and latency changes validates the structural hypotheses of resilient architecture. 1. **The core experimental loop of chaos engineering** (08:09) — Inducing controlled failures creates crucial opportunities to learn from incidents and systematically improve production reliability. 1. **Defining the core properties of complex systems** (11:47) — Software systems scale in complexity through high component multiplicity and non-linear interdependent connections. 1. **How complex adaptive systems evolve and adapt** (16:08) — Human observers and past interactions constantly alter complex environments, guaranteeing unintended cascading consequences. 1. **Applying agile principles to complex system uncertainty** (23:24) — Rigid upfront planning fails in adaptive environments, forcing teams toward iterative methodologies that continuously respond to changing variables. 1. **Embedding failure acceptance into daily engineering culture** (27:34) — True cultural change requires engineering organizations to repeatedly practice learning from mistakes rather than merely declaring support for psychological safety. 1. **Designing and safely executing controlled chaos experiments** (34:15) — Effective failure testing requires thoughtful target selection accompanied by a limited blast radius and automated rollback mechanisms. 1. **Safely probing and intervening in complex ecosystems** (37:24) — Managing unpredictability involves testing coherent hypotheses through safe-to-fail diagnostic probes equipped with clear dampening plans. 1. **Second-order cultural effects of practicing chaos engineering** (41:20) — Beyond technical resilience, routine failure testing fundamentally reshapes how engineering teams analyze and approach unknown systemic risks. ## Related Moments - [Identifying examples of system resilience and fragility in technology](https://www.wearedevelopers.com/videos/100037-beyond-resilience-architecting-antifragile-systems) (from "Beyond Resilience: Architecting Antifragile Systems") - [Implementing proactive learning through scheduled chaos engineering](https://www.wearedevelopers.com/videos/465-empathy-the-secret-sauce-of-resilience) (from "Empathy: The secret sauce of Resilience") - [Navigating and mitigating the impacts of broken engineering cultures](https://www.wearedevelopers.com/videos/1998-from-code-to-culture-why-leadership-determines-software-quality) (from "From Code to Culture: Why Leadership Determines Software Quality") - [Establishing psychological safety to encourage continuous system experimentation](https://www.wearedevelopers.com/videos/100037-beyond-resilience-architecting-antifragile-systems) (from "Beyond Resilience: Architecting Antifragile Systems") - [Validating resilience assumptions through chaos engineering experiments](https://www.wearedevelopers.com/videos/2106-resilient-by-design-building-robust-architectures-in-high-stakes-financial-systems) (from "Resilient by Design: Building Robust Architectures in High-Stakes Financial Systems") - [Balancing failure culture with continuous operational quality standards](https://www.wearedevelopers.com/videos/1717-leading-through-stagility-human-capital-trends-that-redefine-work) (from "Leading Through Stagility: Human Capital Trends That Redefine Work") ## Related Articles - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Never delegate the understanding](https://www.wearedevelopers.com/magazine/749-never-delegate-the-understanding) - [Now is the time for industrialized software development](https://www.wearedevelopers.com/magazine/601-now-is-the-time-for-industrialized-software-development) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) ## Related Jobs - [Tribe Lead - ( Software) Engineering Centre of Excllence](https://www.wearedevelopers.com/jobs/ext/1475530-tribe-lead-software-engineering-centre-of-excllence) at **SD Worx** - [Senior Engineer, Infrastructure Platform](https://www.wearedevelopers.com/jobs/ext/328836-senior-engineer-infrastructure-platform) at **Intercom, Inc.** - [Software Engineer, Platform Engineering (L2)](https://www.wearedevelopers.com/jobs/ext/1956829-software-engineer-platform-engineering-l2) at **Twilio** - [Software Engineer L2 - Cloud Infrastructure](https://www.wearedevelopers.com/jobs/ext/1282024-software-engineer-l2-cloud-infrastructure) at **Twilio** - [Cloud Foundations Team](https://www.wearedevelopers.com/jobs/ext/1483289-cloud-foundations-team) at **GitHub** - [Software Engineer L2 - Cloud Infrastructure](https://www.wearedevelopers.com/jobs/ext/1293339-software-engineer-l2-cloud-infrastructure) at **Twilio**