SRE/Backup Engineer

Spectraforce
Chicago, IL, United States
1 day ago
Apply on leoforce.us
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Active Directory Artificial Intelligence Amazon Web Services Microsoft Azure Backup Devices Bash Shell Software as a Service Cloud Computing Cloud Engineering Cyber Security Databases Linux
+43 more
Disaster Recovery Distributed Systems Multi-Factor Authentication Github Monitoring of Systems Hyper-V Identity and Access Management Python (Programming Language) Key Management Windows Servers NetBackup OpenShift Windows PowerShell Reliability Engineering Ansible Prometheus Zero Trust Network Access Software Engineering Backup and Restore Automatic Programming Google Cloud Enterprise Software Applications System Availability Grafana Virtual Environment HybridCloud Infrastructure as Code (IaC) Kubernetes Storage Technologies Information Technology Cybercrime Bare Metal Performance Monitor Data Management CIS Benchmarks Veeam Terraform Splunk Dynatrace Commvault Elk Stack Servicenow Vmware

Job description

We are seeking a highly technical Senior Site Reliability Engineer (SRE) with deep expertise in enterprise backup engineering, cyber recovery, and platform resiliency. This role will be responsible for engineering highly available, secure, and automated recovery capabilities that protect the organization against operational failures, ransomware, and other cyber threats. The ideal candidate combines traditional SRE principles-automation, observability, reliability engineering, and resilience-with extensive experience designing and operating enterprise backup platforms, immutable storage, air-gapped cyber vaults, isolated recovery environments (IREs), and recovery orchestration. This individual will partner closely with Infrastructure, Cyber Security, Cloud Engineering, Application Development, and Disaster Recovery teams to ensure critical services remain recoverable, resilient, and continuously validated., Site Reliability Engineering

  • Engineer and maintain highly available, resilient enterprise platforms using SRE principles.
  • Define and measure Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for backup and recovery services.
  • Develop automation to reduce operational toil and improve reliability.
  • Perform root cause analysis (RCA) and implement permanent corrective actions.
  • Continuously improve platform reliability, scalability, performance, and recoverability.
  • Establish proactive monitoring, alerting, and observability for backup and cyber recovery platforms.
  • Participate in incident response and major incident recovery activities.

Backup Engineering

  • Design, implement, and administer enterprise backup and recovery solutions across on-premises, cloud, and SaaS platforms.
  • Engineer immutable backup architectures that support ransomware resilience.
  • Design backup strategies for: o Virtual environments o Physical servers o Databases o Kubernetes/OpenShift o Cloud-native workloads o NAS/Object Storage o Enterprise applications

  • Optimize backup performance, retention, replication, encryption, and recovery objectives.
  • Implement policy-based backup automation and lifecycle management.
  • Ensure compliance with enterprise RPO and RTO requirements.

Cyber Recovery Engineering

  • Design and implement enterprise cyber recovery solutions including: o Air-gapped recovery vaults o Clean Rooms o Isolated Recovery Environments (IRE) o Immutable storage architectures

  • Develop secure recovery workflows following cyberattack scenarios.
  • Engineer automated malware scanning and recovery validation processes.
  • Design and test recovery orchestration for severe-but-plausible cyber events.
  • Support recovery point validation and promotion into production recovery environments.
  • Collaborate with Cyber Security teams on ransomware resilience strategies.

Recovery Automation

  • Develop Infrastructure as Code (IaC) and Recovery as Code automation.
  • Build automated recovery runbooks using tools such as: o Ansible o Terraform o PowerShell o Python o GitHub Actions

  • Automate recovery validation, reporting, and compliance evidence generation.
  • Eliminate manual recovery processes wherever possible.

Observability & Monitoring

  • Implement monitoring for: o Backup success rates o Replication health o Recovery readiness o Storage utilization o Cyber vault health o Infrastructure dependencies

  • Build dashboards for executive and operational visibility.
  • Integrate with enterprise observability platforms (e.g., Dynatrace, Grafana, Splunk, Prometheus).

Cyber Resiliency Testing

  • Plan and execute: o Cyber recovery exercises o Clean room validation o Air-gap recovery testing o Full isolated recovery environment exercises o Bare Metal Recovery (BMR) testing o Disaster Recovery testing

  • Validate application recoverability against defined RTO/RPO objectives.
  • Produce executive reporting on recovery readiness and testing outcomes.

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent experience.
  • 7+ years in Backup Engineering, Infrastructure Engineering, or Site Reliability Engineering.
  • 5+ years designing enterprise backup solutions.
  • 3+ years supporting cyber recovery architectures.
  • Experience implementing SRE principles within enterprise infrastructure environments.
  • Strong understanding of distributed systems and high availability architectures.

Required Technical Skills Backup Technologies Experience with one or more:

  • Cohesity
  • Dell PowerProtect Data Manager
  • Dell Data Domain
  • Dell Cyber Recovery
  • Rubrik
  • Commvault
  • Veritas NetBackup
  • Veeam

Cyber Recovery Experience designing and operating:

  • Air-gapped vaults
  • Immutable backups
  • Clean Rooms
  • Isolated Recovery Environments (IRE)
  • Recovery orchestration
  • Cyber resilience testing
  • Ransomware recovery
  • Recovery validation

Cloud Platforms Experience with:

  • Microsoft Azure
  • AWS
  • Google Cloud Platform Including:

  • Cloud-native backup
  • Cross-region recovery
  • Hybrid cloud resiliency

Infrastructure

  • VMware
  • Hyper-V
  • Kubernetes
  • OpenShift
  • Linux
  • Windows Server
  • Active Directory
  • Enterprise storage platforms

Automation Experience with:

  • Ansible
  • Terraform
  • Python
  • PowerShell
  • Bash
  • GitHub
  • GitHub Actions
  • CI/CD pipelines

Observability Experience with:

  • Dynatrace
  • Grafana
  • Prometheus
  • Splunk
  • ELK Stack
  • ServiceNow

Security Knowledge of:

  • Zero Trust architecture
  • NIST Cybersecurity Framework
  • CIS Controls
  • Encryption and key management
  • Identity and Access Management (IAM)
  • Multi-factor authentication (MFA)
  • Secure recovery processes

Preferred Qualifications

  • Experience in financial services or another highly regulated industry.
  • Experience supporting GSIB cyber resiliency programs.
  • Knowledge of regulatory expectations from agencies such as the Federal Reserve, OCC, or FFIEC.
  • Experience with chaos engineering and resilience testing.
  • Familiarity with SRE tooling and reliability metrics.
  • Experience implementing AI-assisted operations (AIOps) and predictive analytics.

Leadership Competencies

  • Strong systems thinking and engineering mindset.
  • Excellent troubleshooting and root cause analysis skills.
  • Ability to lead cross-functional technical recovery efforts.
  • Strong communication and executive presentation skills.
  • Proven ability to influence engineering standards and drive operational excellence.
  • Commitment to continuous improvement through automation and reliability engineering.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on leoforce.us
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

1:40 min

Managing containerized infrastructure with Podman Desktop

Cedric Clyburn Cedric Clyburn +1 · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

1:41 min

Parallels between cloud and legacy infrastructure lock-ins

Björn Stahl Björn Stahl · World Congress 2024

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all