Site Reliability Engineer

Everforth Apex
Minneapolis, MN, United States
6 days ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$143,520.0 - $153,920.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) .NET Framework Artificial Intelligence Application Performance Management Application Release Automation Databases Continuous Integration Machine Learning Microsoft SQL Server Oracle (Applications) Reliability Engineering Ansible
+13 more
Software Engineering Unix Commands Enterprise Software Applications Large Language Models Grafana HybridCloud IBM UrbanCode Deploy Virtual Agents Terraform Splunk Appdynamics Jenkins Artifactory

Job description

In this contingent resource assignment, you will consult on complex initiatives with broad impact and large-scale planning for Software Engineering. This role functions as a senior Site Reliability Engineer supporting enterprise production environments, platform reliability, observability, automation, incident management, and operational excellence initiatives. The successful candidate will join a platform management organization responsible for L2/L3 production support, focusing on improving reliability, implementing automation, and maturing the team’s SRE capabilities., * Provide L2/L3 support for critical enterprise applications and drive reliability improvements across production environments.

  • Lead incident, problem, and change management processes, including root cause analysis and corrective action planning.
  • Build and maintain observability solutions using platforms like Splunk, AppDynamics, Grafana, and BigPanda.
  • Support CI/CD pipelines and enhance release automation using tools such as Jenkins, Artifactory, UDeploy, and Terraform.
  • Develop operational automation and self-healing capabilities, primarily using Ansible, to reduce manual effort.
  • Support AI/ML, LLM, and Agentic AI-based application environments, partnering with development teams to manage reliability.
  • Manage operational risk and performance across on-prem, hybrid, and public cloud infrastructures.
  • Support large-scale Java/.NET applications and assist with troubleshooting for Oracle/MSSQL databases.

Requirements

Experience: 8+ years of experience in Software Engineering, Site Reliability Engineering (SRE), Production Support, or Platform Engineering.

Technical Skills:

  • Expertise with observability platforms (e.g., AppDynamics, Splunk, Grafana, BigPanda, Application Insights).
  • Experience with CI/CD tools (e.g., Jenkins, Artifactory, UDeploy, Terraform).
  • Automation experience using Ansible.
  • Experience supporting AI/ML, LLM, and Agentic AI platforms in production.
  • Experience with L2 level troubleshooting using Unix commands.
  • Familiarity with leading large-scale production support using ITIL practices.

Preferred Qualifications

  • Experience in enterprise banking or another highly regulated industry.
  • Knowledge of public, hybrid, and on-prem cloud platforms.
  • Experience with resilient system design.
  • Background in supporting large-scale Java and .NET applications.
  • Experience with Oracle and MSSQL database support.

About the company

Everforth Apex is a world-class IT services company that serves thousands of clients across the globe. When you join Everforth Apex, you become part of a team that values innovation, collaboration, and continuous learning. We offer quality career resources, training, certifications, development opportunities, and a comprehensive benefits package. Our commitment to excellence is reflected in many awards, including ClearlyRateds Best of Staffing in Talent Satisfaction in the United States and Great Place to Work in the United Kingdom and Mexico.

Everforth Apex uses a virtual recruiter as part of the application process. Click for more details. By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails from Everforth Apex and its affiliates, and contracted partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or undelivered messages. You can reply STOP to cancel and HELP for help. You can access our privacy policy at

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · WWC 2024

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · WWC 2023

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · WWC 2025

Videos

See all

Related articles

See all