Site Reliability Engineer

Everforth Apex
Chicago, IL, United States
1 day ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Active Directory Artificial Intelligence Amazon Web Services Microsoft Azure Backup Devices Bash Shell Software as a Service Cloud Computing Cloud Engineering Cyber Security Continuous Integration Linux
+34 more
Distributed Systems Hyper-V Identity and Access Management Python (Programming Language) Windows Servers NetBackup OpenShift Windows PowerShell Reliability Engineering Ansible Prometheus Zero Trust Network Access Software Engineering Backup and Restore Automatic Programming Google Cloud System Availability Grafana HybridCloud Infrastructure as Code (IaC) Kubernetes Information Technology Cybercrime Performance Monitor Data Management CIS Benchmarks Veeam Terraform Splunk Dynatrace Commvault Elk Stack Servicenow Vmware

Job description

  • Engineer and maintain highly available, resilient enterprise platforms using SRE principles.
  • Design, implement, and administer enterprise backup and recovery solutions across on-premises, cloud, and SaaS platforms.
  • Design and implement enterprise cyber recovery solutions including air-gapped recovery vaults and isolated recovery environments.
  • Develop Infrastructure as Code (IaC) and Recovery as Code automation using tools like Ansible, Terraform, and Python.
  • Implement proactive monitoring, alerting, and observability for backup and cyber recovery platforms.
  • Plan and execute cyber recovery exercises, including clean room validation and full isolated recovery testing.
  • Perform root cause analysis (RCA) and implement permanent corrective actions for incidents.
  • Collaborate with Cyber Security, Infrastructure, and Application Development teams on resilience strategies.

Requirements

We are seeking a highly technical Senior Site Reliability Engineer (SRE) with deep expertise in enterprise backup engineering, cyber recovery, and platform resiliency. This role will be responsible for engineering highly available, secure, and automated recovery capabilities that protect the organization against operational failures, ransomware, and other cyber threats. The ideal candidate combines SRE principles with extensive experience in designing and operating enterprise backup platforms, immutable storage, and recovery orchestration., Education: Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent experience., * 7+ years in Backup Engineering, Infrastructure Engineering, or Site Reliability Engineering.

  • 5+ years designing enterprise backup solutions.
  • 3+ years supporting cyber recovery architectures.
  • Experience implementing SRE principles within enterprise infrastructure environments.
  • A strong understanding of distributed systems and high availability architectures.

Technical Skills:

  • Backup Technologies: Experience with one or more: Cohesity, Dell PowerProtect Data Manager, Dell Data Domain, Dell Cyber Recovery, Rubrik, Commvault, Veritas NetBackup, or Veeam.
  • Cyber Recovery: Experience designing and operating air-gapped vaults, immutable backups, isolated recovery environments (IRE), and recovery orchestration.
  • Cloud Platforms: Experience with Microsoft Azure, AWS, or Google Cloud Platform, including cloud-native backup and hybrid cloud resiliency.
  • Infrastructure: VMware, Hyper-V, Kubernetes, OpenShift, Linux, Windows Server, Active Directory, and enterprise storage platforms.
  • Automation: Ansible, Terraform, Python, PowerShell, Bash, and CI/CD pipelines.
  • Observability: Dynatrace, Grafana, Prometheus, Splunk, ELK Stack, or ServiceNow.
  • Security Knowledge: Zero Trust architecture, NIST Cybersecurity Framework, CIS Controls, encryption, and Identity and Access Management (IAM).

Preferred Qualifications

  • Experience in financial services or another highly regulated industry.
  • Experience supporting GSIB cyber resiliency programs.
  • Knowledge of regulatory expectations from agencies such as the Federal Reserve, OCC, or FFIEC.
  • Experience with chaos engineering and resilience testing.
  • Familiarity with SRE tooling and reliability metrics.
  • Experience implementing AI-assisted operations (AIOps) and predictive analytics.

About the company

Everforth Apex is a world-class IT services company that serves thousands of clients across the globe. When you join Everforth Apex, you become part of a team that values innovation, collaboration, and continuous learning. We offer quality career resources, training, certifications, development opportunities, and a comprehensive benefits package. Our commitment to excellence is reflected in many awards, including ClearlyRateds Best of Staffing in Talent Satisfaction in the United States and Great Place to Work in the United Kingdom and Mexico.

Everforth Apex uses a virtual recruiter as part of the application process. Click for more details. By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails from Everforth Apex and its affiliates, and contracted partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or undelivered messages. You can reply STOP to cancel and HELP for help. You can access our privacy policy at

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

1:40 min

Managing containerized infrastructure with Podman Desktop

Cedric Clyburn Cedric Clyburn +1 · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

1:41 min

Parallels between cloud and legacy infrastructure lock-ins

Björn Stahl Björn Stahl · World Congress 2024

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all