Reliability Engineer

Compu-Vision - IT
Centreville, VA, United States
7 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Amazon Web Services Automation of Tests Cloud Computing Computer Programming Disaster Recovery Python (Programming Language) YAML Cloud Platform System System Availability AWS Lambda Infrastructure as Code (IaC) Cloudformation
+2 more
Terraform Serverless Computing

Job description

We are seeking an experienced Reliability Engineer with strong expertise in AWS cloud environments, multi-region workloads, resiliency, and failover optimization. The ideal candidate will have hands-on experience designing, supporting, and optimizing highly available AWS infrastructure across multiple regions, with a strong understanding of Infrastructure as Code and automation., * Design, implement, and optimize reliability and resiliency strategies for multi-region AWS workloads.

  • Analyze and improve AWS infrastructure failover mechanisms and recovery processes.
  • Support high availability, disaster recovery, and business continuity initiatives.
  • Develop and maintain Infrastructure as Code using AWS CloudFormation and Terraform.
  • Build, maintain, and enhance AWS automation for resiliency and failover.
  • Work with serverless AWS services including Lambda and Step Functions.
  • Support and troubleshoot AWS messaging and storage services, including AWS MQ and AWS EFS.
  • Identify potential reliability risks, bottlenecks, and single points of failure.
  • Test and validate failover and recovery procedures.
  • Collaborate with engineering and cloud teams to improve system availability and operational resilience.
  • Modify existing automation code to address reliability and resiliency requirements.

Requirements

Required AWS Skills

  • Strong experience with AWS cloud environments.
  • Experience supporting multi-region AWS workloads.
  • Strong understanding of AWS infrastructure resiliency and failover.
  • Hands-on experience with Infrastructure as Code (IaC).
  • Strong experience with:

  • AWS CloudFormation
  • CloudFormation YAML
  • Terraform
  • AWS Lambda
  • AWS Step Functions
  • AWS MQ
  • AWS EFS

Programming Requirements

  • Strong Python experience.
  • Ability to understand existing Python-based resiliency and automation code.
  • Ability to modify, troubleshoot, and enhance existing automation scripts/code.
  • Understanding of automation and scripting practices for cloud infrastructure.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:35 min

Centralizing configuration logic with native YAML block references

Matthieu Vincent Matthieu Vincent · Europe 2026 Virtual

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

2:45 min

Executing event-driven code using AWS Lambda functions

Sebastien Stormacq Sebastien Stormacq · LIVE

2:33 min

Evaluating complementary and alternative cloud deployment frameworks

Talia Nassi · LIVE

2:00 min

Introduction to YAML syntax and basic formatting

Chris Ayers · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all