Cloud Site Reliability Engineer (SRE)

Akaasa Technologies
Charlotte, NC, United States
about 2 months ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Backup Devices Cloud Computing Cloud Engineering Databases Data Centers Data Recovery Disaster Recovery Reliability Engineering Runbook Backup and Restore
+5 more
Google Cloud Software Troubleshooting Multi-Cloud HybridCloud Infrastructure Automation Frameworks

Job description

seeking an experienced Cloud Site Reliability Engineer (Cloud SRE) to support a critical infrastructure transition initiative involving disaster recovery (DR) modernization, backup environment migration, and cloud infrastructure resiliency. This role will focus on ensuring the continuity, availability, recoverability, and operational stability of business-critical applications and services throughout a large-scale infrastructure transformation. This is a hands-on contract position responsible for planning, executing, validating, and operationalizing disaster recovery and backup solutions while collaborating closely with infrastructure, cloud, application, and security teams., Disaster Recovery & Infrastructure Migration

  • Plan and execute migrations of disaster recovery and backup environments associated with data center consolidations, infrastructure modernization efforts, or cloud transformation initiatives.
  • Responsible for supporting and improving DR processes and testing initiatives.
  • Support the relocation, re-platforming, or modernization of DR environments for applications, databases, and business-critical services.
  • Ensure recovery architectures meet established recovery objectives, resiliency requirements, and operational standards.

Disaster Recovery Testing & Validation

  • Plan, coordinate, and execute disaster recovery exercises and failover testing activities.
  • Validate backup recovery procedures, restoration processes, recovery sequencing, and operational readiness.
  • Identify recovery gaps, operational risks, and remediation opportunities to improve resiliency and recoverability.
  • Document test results, lessons learned, and recommended improvements.

Cloud & Hybrid Infrastructure Operations

  • Support cloud and hybrid infrastructure platforms related to backup, disaster recovery, and business continuity.
  • Assist in the implementation and operationalization of cloud-based backup and recovery solutions.
  • Execute operational runbooks, standard operating procedures, and recovery processes aligned with approved designs and governance standards.
  • Contribute to infrastructure reliability, automation, monitoring, and operational excellence initiatives.

Operational Coordination & Documentation

  • Collaborate with application owners, infrastructure engineers, cloud teams, security teams, and third-party vendors during migration and testing activities.
  • Track project milestones, dependencies, risks, and execution progress.
  • Maintain accurate documentation of environments, configurations, recovery procedures, and operational outcomes.
  • Provide status updates and technical recommendations to project stakeholders., Title: Senior Site Reliability Engineer (SRE) - Warehouse / Infrastructure Location: Charlotte, NC Duration: 12 month with extension Work Requirements: US Citizen, GC Holders or…
  • 2 days ago, Title: Site Reliability Engineer Duration: Contract Location: Charlotte, NC About the Role The Client Document Generation team is seeking a Senior Software Engineer ( IT Onsho…
  • 16 days ago +

Requirements

  • 5+ years of experience in Site Reliability Engineering (SRE), Cloud Operations, Infrastructure Engineering, or related disciplines.
  • MUST HAVE Disaster Recovery (DR) testing experience.
  • MUST HAVE experience with automation and reliability engineering
  • Hands-on experience supporting enterprise disaster recovery environments, backup systems, and business continuity initiatives.
  • MUST HAVE strong understanding of disaster recovery concepts, resiliency strategies, recovery testing, backup technologies, and infrastructure migration methodologies.
  • Experience operating within cloud, hybrid cloud, or multi-cloud environments.
  • Proven ability to execute complex infrastructure projects within defined timelines and operational constraints.
  • Strong troubleshooting, analytical, and problem-solving skills.
  • Excellent communication, collaboration, and documentation abilities., * Experience supporting data center migrations, colocation exits, or infrastructure modernization projects.
  • Knowledge of cloud-native backup and disaster recovery platforms.
  • Experience with AWS, Microsoft Azure, Google Cloud Platform (GCP), or hybrid cloud environments.
  • Familiarity with infrastructure automation and operational tooling.
  • Experience working within regulated industries such as financial services, healthcare, insurance, or government.
  • Strong operational discipline and process-oriented mindset.

Technical Skills

  • Disaster Recovery Planning & Execution
  • Backup & Recovery Solutions
  • Cloud Infrastructure (AWS, Azure, GCP)
  • Hybrid Infrastructure Operations
  • Infrastructure Reliability & Resiliency
  • Recovery Testing & Validation
  • Operational Runbooks & Documentation
  • Infrastructure Migration & Transformation
  • Monitoring & Operational Support
  • Incident Response & Problem Resolution

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

2:50 min

Introduction and the value of runbooks

Hila Fish · World Congress 2023

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:32 min

Structuring automated incident workflows between runbooks and raw models

Aram Hakobyan Aram Hakobyan +1 · World Congress 2026 Europe

4:01 min

Managing application isolation via pluggable database models

Wei Hu Wei Hu · World Congress 2022

Videos

See all

Related articles

See all