> Markdown version of [/jobs/ext/2587107-aws-infrastructure-cloud-resilience-engineer](https://www.wearedevelopers.com/jobs/ext/2587107-aws-infrastructure-cloud-resilience-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AWS Infrastructure & Cloud Resilience Engineer - **Company:** Cognizant Technology Solutions Corporation - **Location:** Dallas, TX, United States (Remote available) - **Salary:** $90,000.0 - $115,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Cloud Computing, Software Design Documents, DevOps, Failure Mode Effects Analysis, Reliability Engineering, Site Reliability Engineering Practices, Runbook, TypeScript, Cloudwatch - **Published:** August 14, 2026 - **Apply:** https://www.juju.com/job/00000000gn4bwt ## About the Role * Resiliency testing (Level 300) - AWS Fault Injection Service (FIS) and Resilience Hub; leading FMEA and resiliency game days; failover/recovery validation. * Multi-region resilience patterns (Level 300) - Active-active (3-region), warm standby, and pilot light patterns across AZ/region failures. * Infrastructure as code (Level 300) - CDK TypeScript constructs for resilience and observability configuration. * Observability (Level 200) - CloudWatch, AppSignals, ADOT/X-Ray to measure resilience test outcomes. * SRE practices (Level 200) - Incident response, change management, and continuous DR testing embedded in SRE ways of working. General Requirements * AWS Professional certification preferred (e.g. Solutions Architect - Professional, DevOps Engineer - Professional, or Security - Specialty as relevant to the role). * Prior delivery experience in a large, regulated enterprise environment (financial services strongly preferred); comfortable operating under change-control and audit scrutiny. * Able to produce clear written documentation (architecture decision records, runbooks, design docs) suitable for client Tech Risk review. * Strong stakeholder communication; can work directly with client engineering, security, and SRE teams as an embedded SME. * Works to the stated engagement location / time-zone overlap and follows AWS ProServe and Customer onboarding, security, and vetting requirements ## Description This role develops resilience testing scenarios and runbooks, runs game days, and validates failover across AZ and region failure scenarios for target applications. Mandatory skill: CDK (medium), AWS FIS(AWS Resilience Hub),Chaos Engineering, Gremlin (Expert), DR Testing (Multi Region) (medium) Role Scope * Deploy and integrate AWS Resilience Hub and Fault Injection Service (FIS) within the customer environment and bundle into CFT/RefStack. * Lead Failure Mode & Effects Analysis (FMEA) and resiliency game days; remediate issues identified during testing. * Validate multi-region resiliency patterns - active-active (3-region), warm standby, pilot light - across AZ and region failure scenarios. * Develop resilience testing scenarios and runbooks for customer application archetypes and produce resilience posture assessments. ## Related Videos - [Azure-Well Architected Framework - designing mission critical workloads in practice](https://www.wearedevelopers.com/videos/1529-azure-well-architected-framework-designing-mission-critical-workloads-in-practice) - [Do TypeScript without TypeScript](https://www.wearedevelopers.com/videos/327-do-typescript-without-typescript) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [3 Key Steps for Optimizing DevOps Workflows](https://www.wearedevelopers.com/videos/962-3-key-steps-for-optimizing-devops-workflows) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Top Must-Visit Developer Conferences in the US in 2026](https://www.wearedevelopers.com/magazine/679-top-must-visit-developer-conferences-in-the-us-in-2026) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)