Reliability Engineer
Compu-Vision - IT
Centreville, VA, United States
7 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.careerjet.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source
Tech stack
Amazon Web Services
Automation of Tests
Cloud Computing
Computer Programming
Disaster Recovery
Python (Programming Language)
YAML
Cloud Platform System
System Availability
AWS Lambda
Infrastructure as Code (IaC)
Cloudformation
+2 more
Terraform
Serverless Computing
Job description
We are seeking an experienced Reliability Engineer with strong expertise in AWS cloud environments, multi-region workloads, resiliency, and failover optimization. The ideal candidate will have hands-on experience designing, supporting, and optimizing highly available AWS infrastructure across multiple regions, with a strong understanding of Infrastructure as Code and automation., * Design, implement, and optimize reliability and resiliency strategies for multi-region AWS workloads.
- Analyze and improve AWS infrastructure failover mechanisms and recovery processes.
- Support high availability, disaster recovery, and business continuity initiatives.
- Develop and maintain Infrastructure as Code using AWS CloudFormation and Terraform.
- Build, maintain, and enhance AWS automation for resiliency and failover.
- Work with serverless AWS services including Lambda and Step Functions.
- Support and troubleshoot AWS messaging and storage services, including AWS MQ and AWS EFS.
- Identify potential reliability risks, bottlenecks, and single points of failure.
- Test and validate failover and recovery procedures.
- Collaborate with engineering and cloud teams to improve system availability and operational resilience.
- Modify existing automation code to address reliability and resiliency requirements.
Requirements
Required AWS Skills
- Strong experience with AWS cloud environments.
- Experience supporting multi-region AWS workloads.
- Strong understanding of AWS infrastructure resiliency and failover.
- Hands-on experience with Infrastructure as Code (IaC).
-
Strong experience with:
- AWS CloudFormation
- CloudFormation YAML
- Terraform
- AWS Lambda
- AWS Step Functions
- AWS MQ
- AWS EFS
Programming Requirements
- Strong Python experience.
- Ability to understand existing Python-based resiliency and automation code.
- Ability to modify, troubleshoot, and enhance existing automation scripts/code.
- Understanding of automation and scripting practices for cloud infrastructure.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.careerjet.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
AJ
Austin Joy
over 4 years ago
LM
Luis Minvielle
7 Cloud Computing Trends Coming in 2025 for Developers
over 2 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
23 days ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
LM
Luis Minvielle
Fully Remote Software Engineer Jobs
over 2 years ago