> Markdown version of [/jobs/ext/2103773-manager-software-engineering-resilience-engineering](https://www.wearedevelopers.com/jobs/ext/2103773-manager-software-engineering-resilience-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Manager, Software Engineering (Resilience Engineering) - **Company:** Affirm - **Location:** Los Angeles, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $181,000.0 - $241,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon Web Services, Audit Trail, Cloud Computing, Computer Programming, Distributed Systems, Fault Tolerance, Python (Programming Language), Load Testing, Reliability Engineering, Software Engineering, System Testing, Datadog, Rate Limiting, Kotlin, Kubernetes, Low Latency, Build Tools - **Published:** August 18, 2026 - **Apply:** https://job-boards.greenhouse.io/affirm/jobs/7701911003 ## About the Role * Proven experience leading engineering teams in reliability, infrastructure, or distributed systems. * Hands-on experience with production load testing, chaos engineering, or large-scale system validation. * Experience with leveraging a chaos engineering vendor such as Gremlin, Harness, or something similar. * Strong understanding of failure modes in distributed systems, including latency, partial failure, and cascading outages. * Experience building or operating systems with strong safety guarantees (isolation, rate limiting, guardrails, auditability). * Familiarity with cloud-native environments (AWS, Kubernetes) and observability tooling. * Strong programming background (e.g., Python, Kotlin, Java, or similar). * Excellent problem-solving skills and the ability to balance long-term resilience investments with immediate business needs. * Strong communication and leadership skills, with a track record of influencing engineering practices across teams. ## Description We are seeking a seasoned Engineering Manager to lead our Resilience Engineering team. This role is critical in ensuring the safety and reliability of our production systems through proactive validation techniques, including production load testing and chaos engineering. You will lead the development of systems and practices that allow engineers to safely test system behavior under stress and failure conditions in production, ensuring issues are discovered and mitigated before they impact real users. What you'll do Leadership & Strategy * Define and drive the vision for resilience engineering at Affirm, with a focus on production load testing and chaos engineering as first-class engineering practices. * Lead and mentor a team of engineers building platforms and tooling for safe production experimentation. * Partner with infrastructure, product, and security leadership to embed resilience validation into the software development lifecycle. * Establish best practices for safely testing system limits and failure scenarios in production. Systems & Operations * Own the design and evolution of platforms that enable safe, controlled production load testing and fault injection. * Ensure strong safeguards are in place, including isolation boundaries, approval workflows, and automated rollback mechanisms to protect real users. * Build systems that provide end-to-end observability, traceability, and auditability for all resilience experiments. * Drive reliability improvements by systematically identifying weaknesses through load testing and chaos experiments. * Establish monitoring, alerting, and incident response practices tailored to proactive resilience validation. Collaboration & Enablement * Work closely with engineering teams to design and execute production load tests and chaos experiments safely. * Partner with infrastructure teams to build guardrails around tests and experimentations. * Enable teams to adopt resilience practices by providing reusable tooling, frameworks, and standardized workflows. * Identify systemic weaknesses and lead cross-functional efforts to improve reliability and fault tolerance. * Evangelize a culture of "test failure before failure tests you" across the organization. ## Related Videos - [Kotlin Multiplatform - True power of native code reuse](https://www.wearedevelopers.com/videos/4-kotlin-multiplatform-true-power-of-native-code-reuse) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Empathy: The secret sauce of Resilience](https://www.wearedevelopers.com/videos/465-empathy-the-secret-sauce-of-resilience) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)