Information Technology - DevOps Engineer (Cloud Engineer)

Xcelo Group Inc
Chicago, IL, United States
2 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Active Directory Amazon Web Services Microsoft Azure Backup Devices Bash Shell Software as a Service Cloud Computing Cloud Engineering Cyber Security Databases Data Recovery Linux
+40 more
Disaster Recovery Distributed Systems Multi-Factor Authentication Github Monitoring of Systems Hyper-V Identity and Access Management Python (Programming Language) Key Management Windows Servers NetBackup OpenShift Windows PowerShell Reliability Engineering Ansible Prometheus Zero Trust Network Access Virtual Machines Google Cloud Enterprise Software Applications System Availability Grafana HybridCloud Infrastructure as Code (IaC) Kubernetes Storage Technologies Information Technology Bare Metal Performance Monitor Data Management CIS Benchmarks Veeam Terraform Splunk Network Server Dynatrace Commvault Elk Stack Servicenow Vmware

Job description

  • Design, engineer, and maintain highly available enterprise backup and recovery platforms using SRE principles.
  • Define and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for backup and recovery services.
  • Develop automation to reduce operational toil and improve platform reliability.
  • Perform Root Cause Analysis (RCA) and implement permanent corrective actions.
  • Improve platform reliability, scalability, performance, availability, and recoverability.
  • Build proactive monitoring, alerting, and observability for backup and cyber recovery platforms.
  • Participate in incident response, major incident management, and recovery operations.
  • Design and administer enterprise backup solutions across on-premises, cloud, SaaS, virtual machines, physical servers, databases, Kubernetes/OpenShift, NAS, object storage, and enterprise applications.
  • Engineer immutable backup architectures supporting ransomware resilience.
  • Optimize backup performance, retention, replication, encryption, and recovery objectives.
  • Implement policy-based backup automation and lifecycle management.
  • Ensure backup and recovery environments meet defined RPO and RTO requirements.
  • Design and implement air-gapped recovery vaults, Clean Rooms, IRE environments, and immutable storage architectures.
  • Develop secure recovery workflows for cyberattack and ransomware scenarios.
  • Automate malware scanning, recovery-point validation, and recovery readiness checks.
  • Design and test recovery orchestration for severe cyber disruption scenarios.
  • Work closely with Cyber Security teams to develop ransomware resilience strategies.
  • Develop Infrastructure as Code and Recovery as Code solutions.
  • Create automated recovery runbooks using Ansible, Terraform, Python, PowerShell, and GitHub Actions.
  • Automate recovery validation, compliance reporting, and evidence generation.
  • Implement monitoring for backup success rates, replication health, cyber vault health, recovery readiness, storage utilization, and infrastructure dependencies.
  • Build dashboards for operational teams and executive leadership.
  • Integrate backup and recovery platforms with Dynatrace, Grafana, Prometheus, Splunk, and other enterprise monitoring tools.
  • Plan and execute cyber recovery exercises, Clean Room validation, air-gap recovery testing, isolated recovery exercises, BMR testing, and Disaster Recovery testing.
  • Validate application recoverability against defined business RTO/RPO requirements.
  • Prepare executive-level reporting on cyber recovery readiness, resilience testing, risks, and remediation activities.

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent professional experience., * 7+ years of experience in Backup Engineering, Infrastructure Engineering, Site Reliability Engineering, or related infrastructure roles.
  • 5+ years designing and supporting enterprise backup and recovery solutions.
  • 3+ years supporting cyber recovery, cyber resiliency, or ransomware recovery architectures.
  • Hands-on experience implementing SRE principles, automation, reliability engineering, monitoring, and operational resilience.
  • Strong knowledge of distributed systems, high availability, disaster recovery, RPO/RTO, and enterprise infrastructure architecture.

Backup & Recovery Technologies

Hands-on experience with one or more of the following:

  • Cohesity
  • Dell PowerProtect Data Manager
  • Dell Data Domain
  • Dell Cyber Recovery
  • Rubrik
  • Commvault
  • Veritas NetBackup
  • Veeam

Cyber Recovery & Resiliency

Strong experience with:

  • Air-gapped cyber recovery vaults
  • Immutable backups and immutable storage
  • Clean Rooms
  • Isolated Recovery Environments (IRE)
  • Recovery orchestration
  • Cyber resilience testing
  • Ransomware recovery
  • Recovery validation
  • Bare Metal Recovery (BMR)
  • Secure recovery workflows
  • Disaster Recovery testing
  • RPO/RTO validation

Cloud & Infrastructure

Experience supporting:

  • Microsoft Azure
  • Amazon Web Services (AWS)
  • Google Cloud Platform (Google Cloud Platform)
  • Cloud-native backup and recovery
  • Cross-region recovery
  • Hybrid-cloud resiliency
  • VMware
  • Hyper-V
  • Kubernetes
  • OpenShift
  • Linux
  • Windows Server
  • Active Directory
  • Enterprise storage platforms

Automation / Infrastructure as Code

Strong hands-on experience with:

  • Ansible
  • Terraform
  • Python
  • PowerShell
  • Bash
  • GitHub
  • GitHub Actions
  • CI/CD pipelines
  • Infrastructure as Code (IaC)
  • Recovery as Code

Observability & Monitoring

Experience with:

  • Dynatrace
  • Grafana
  • Prometheus
  • Splunk
  • ELK Stack
  • ServiceNow

Security & Compliance

Strong understanding of:

  • Zero Trust architecture
  • NIST Cybersecurity Framework
  • CIS Controls
  • Encryption and key management
  • Identity and Access Management (IAM)
  • Multi-Factor Authentication (MFA)
  • Secure recovery processes, * Experience working within financial services, banking, insurance, or another highly regulated industry.
  • Experience supporting GSIB cyber resiliency programs.
  • Understanding of regulatory expectations from organizations such as the Federal Reserve, OCC, and FFIEC.
  • Experience with chaos engineering and resilience testing.
  • Strong understanding of SRE reliability metrics and operational excellence practices.
  • Experience implementing AIOps, intelligent monitoring, or predictive analytics.
  • Excellent troubleshooting and root cause analysis skills.
  • Ability to lead cross-functional technical recovery initiatives.
  • Strong communication and executive presentation skills.
  • Proven ability to influence engineering standards and improve operational reliability.
  • Strong commitment to automation, continuous improvement, and resilience engineering.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:40 min

Managing containerized infrastructure with Podman Desktop

Cedric Clyburn Cedric Clyburn +1 · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:46 min

Navigating a career in cloud transformation consulting

Piet Van Dongen · LIVE

1:41 min

Parallels between cloud and legacy infrastructure lock-ins

Björn Stahl Björn Stahl · World Congress 2024

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all