Reliability Engineer

Robert Half
Avon, MN, United States
3 days ago
Apply on www.juju.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Application Performance Management Distributed Systems Monitoring of Systems Supervisory Control and Data Acquisition (SCADA) Information Technology Operations Reliability Engineering Cloud Services Ansible Scripting Cloud Platform System Reliability of Systems Infrastructure Automation Frameworks
+3 more
Terraform Splunk Dynatrace

Job description

  • Build and enhance highly available infrastructure that supports office locations, remote field environments, networking needs, cloud services, and edge-based systems.

  • Direct incident response efforts for service disruptions, coordinate restoration activities, and lead root cause investigations to prevent repeat issues.

  • Create and maintain monitoring, alerting, and observability capabilities that improve visibility into system health, uptime, and application performance.

  • Work closely with construction, engineering, and field personnel to ensure technology reliability aligns with project schedules, operational demands, and safety expectations.

  • Implement automated infrastructure deployment and recovery processes using Infrastructure as Code and configuration management tools such as Terraform and Ansible.

  • Establish service reliability targets, manage service level objectives, and use error budgets to guide operational decisions and continuous improvement.

  • Strengthen the security posture of remote and field-deployed systems by improving hardening practices and secure access methods.

  • Provide guidance to less experienced engineers and help foster a culture centered on reliability, accountability, and operational excellence.

  • Identify process improvements that reduce inefficiencies, simplify support efforts, and improve overall service delivery.

Requirements

We are looking for a Senior Reliability Engineer to strengthen the stability, scalability, and performance of technology systems that support construction and field operations. This position plays a key role in building dependable infrastructure across remote job sites, field offices, and cloud environments so teams can work efficiently with minimal disruption. The ideal candidate brings a strong SRE mindset, combines technical depth with practical problem-solving, and partners effectively with operational teams to maintain business continuity and system resilience., * Relevant experience in site reliability, infrastructure engineering, or a closely related IT operations role.

  • Strong understanding of cloud platforms, enterprise networking, and edge computing concepts in distributed environments.
  • Hands-on experience with Infrastructure as Code, including tools such as Terraform and Ansible.
  • Proficiency with monitoring and observability platforms such as Splunk and Dynatrace, along with scripting or automation capabilities.
  • Familiarity with construction, industrial, or field-based technology environments is strongly preferred.
  • Knowledge of systems used in operational technology settings, including SCADA, IoT-connected devices, or similar enterprise platforms.
  • Excellent written and verbal communication skills with the ability to collaborate effectively across technical and operational teams.
  • Ability to handle sensitive information with discretion while working independently and contributing positively in a team setting.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:53 min

Evaluating traditional scripting languages for modern development tasks

Jens Knipper Jens Knipper · Europe 2026 Virtual

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · World Congress 2025

Videos

See all

Related articles

See all