Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+3 more
Job description
-
Build and enhance highly available infrastructure that supports office locations, remote field environments, networking needs, cloud services, and edge-based systems.
-
Direct incident response efforts for service disruptions, coordinate restoration activities, and lead root cause investigations to prevent repeat issues.
-
Create and maintain monitoring, alerting, and observability capabilities that improve visibility into system health, uptime, and application performance.
-
Work closely with construction, engineering, and field personnel to ensure technology reliability aligns with project schedules, operational demands, and safety expectations.
-
Implement automated infrastructure deployment and recovery processes using Infrastructure as Code and configuration management tools such as Terraform and Ansible.
-
Establish service reliability targets, manage service level objectives, and use error budgets to guide operational decisions and continuous improvement.
-
Strengthen the security posture of remote and field-deployed systems by improving hardening practices and secure access methods.
-
Provide guidance to less experienced engineers and help foster a culture centered on reliability, accountability, and operational excellence.
-
Identify process improvements that reduce inefficiencies, simplify support efforts, and improve overall service delivery.
Requirements
We are looking for a Senior Reliability Engineer to strengthen the stability, scalability, and performance of technology systems that support construction and field operations. This position plays a key role in building dependable infrastructure across remote job sites, field offices, and cloud environments so teams can work efficiently with minimal disruption. The ideal candidate brings a strong SRE mindset, combines technical depth with practical problem-solving, and partners effectively with operational teams to maintain business continuity and system resilience., * Relevant experience in site reliability, infrastructure engineering, or a closely related IT operations role.
- Strong understanding of cloud platforms, enterprise networking, and edge computing concepts in distributed environments.
- Hands-on experience with Infrastructure as Code, including tools such as Terraform and Ansible.
- Proficiency with monitoring and observability platforms such as Splunk and Dynatrace, along with scripting or automation capabilities.
- Familiarity with construction, industrial, or field-based technology environments is strongly preferred.
- Knowledge of systems used in operational technology settings, including SCADA, IoT-connected devices, or similar enterprise platforms.
- Excellent written and verbal communication skills with the ability to collaborate effectively across technical and operational teams.
- Ability to handle sensitive information with discretion while working independently and contributing positively in a team setting.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Where To Find Software Engineering Jobs
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Find a Developer Job: 12 Best Job Sites For Developers