AWS Infrastructure & Cloud Resilience Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
This role develops resilience testing scenarios and runbooks, runs game days, and validates failover across AZ and region failure scenarios for target applications.
Mandatory skill: CDK (medium), AWS FIS(AWS Resilience Hub),Chaos Engineering, Gremlin (Expert), DR Testing (Multi Region) (medium)
Role Scope
-
Deploy and integrate AWS Resilience Hub and Fault Injection Service (FIS) within the customer environment and bundle into CFT/RefStack.
-
Lead Failure Mode & Effects Analysis (FMEA) and resiliency game days; remediate issues identified during testing.
-
Validate multi-region resiliency patterns - active-active (3-region), warm standby, pilot light - across AZ and region failure scenarios.
-
Develop resilience testing scenarios and runbooks for customer application archetypes and produce resilience posture assessments.
Requirements
-
Resiliency testing (Level 300) - AWS Fault Injection Service (FIS) and Resilience Hub; leading FMEA and resiliency game days; failover/recovery validation.
-
Multi-region resilience patterns (Level 300) - Active-active (3-region), warm standby, and pilot light patterns across AZ/region failures.
-
Infrastructure as code (Level 300) - CDK TypeScript constructs for resilience and observability configuration.
-
Observability (Level 200) - CloudWatch, AppSignals, ADOT/X-Ray to measure resilience test outcomes.
-
SRE practices (Level 200) - Incident response, change management, and continuous DR testing embedded in SRE ways of working.
General Requirements
-
AWS Professional certification preferred (e.g. Solutions Architect - Professional, DevOps Engineer - Professional, or Security - Specialty as relevant to the role).
-
Prior delivery experience in a large, regulated enterprise environment (financial services strongly preferred); comfortable operating under change-control and audit scrutiny.
-
Able to produce clear written documentation (architecture decision records, runbooks, design docs) suitable for client Tech Risk review.
-
Strong stakeholder communication; can work directly with client engineering, security, and SRE teams as an embedded SME.
-
Works to the stated engagement location / time-zone overlap and follows AWS ProServe and Customer onboarding, security, and vetting requirements
Benefits & conditions
Compensation: We are offering between $90,000 - $115,000. Applications will be accepted until Aug 21, 2026.Cognizant will only consider applicants for this position who are legally authorized to work in Canada without requiring employer sponsorship, now or at any time in the future.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Is Software Engineering Over-Saturated?
Top Must-Visit Developer Conferences in the US in 2026
Trustworthy AI Starts at Deployment: 5 Checks Before You Ship