Site Reliability Engineer

Insight Global
Chandler, AZ, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Confluence JIRA Microsoft Azure Cloud Computing Continuous Integration DevOps Github Python (Programming Language) Windows PowerShell Scrum Methodology Reliability Engineering
+15 more
Ansible Prometheus YAML Cloud Monitoring Grafana Backend Hashicorp Bitbucket Terraform Splunk Ansible Tower Dynatrace Jenkins Servicenow Artifactory

Job description

Insight Global is seeking a Senior Terraform Production Support Engineer for a top enterprise client. This role focuses on supporting and administering Terraform Enterprise and enterprise automation platforms in a high-availability production environment. The ideal candidate brings strong Terraform and IaC expertise, deep production support experience, and a passion for root cause analysis, monitoring, and reliability engineering. You’ll work closely with cloud, DevOps, and product engineering teams to improve automation, streamline deployments, support cutovers, and ensure platform stability across AWS and hybrid environments. This is a senior-level opportunity with leadership visibility, especially for candidates stepping into a Lead role in Arizona.

Requirements

  • 5-8 years of Production Support / SRE / DevOps experience

  • Strong Terraform experience

o Terraform Enterprise development and administration (backend/platform)

  • Cutover experience (HUGE plus, near must-have)

  • Change Management experience

o Incident * Problem Management * RCA / PKE

o Hands-on with ITSM tools (Remedy and/or ServiceNow)

  • Ansible Platform / ATaaS

o Administration experience (not just a user)

o Onboarding applications to Ansible Tower

  • Infrastructure as Code (IaC) using Terraform, HCL, YAML

  • Cloud Monitoring & APM

o Prometheus, Grafana, Dynatrace, Splunk (any combination)

o Building dashboards, consoles, and alerts

  • Cloud experience (AWS required; Azure/GCP a plus)

  • Python and/or PowerShell scripting

  • Strong root cause analysis and troubleshooting skills

  • Experience supporting production environments * HashiCorp Terraform Enterprise certifications

  • CI/CD tools: GitHub, Jenkins, Artifactory

  • Horizon toolset (Ansible, Jira, Confluence, Bitbucket)

  • Automation tools:

o Cutover

o TrueSight Orchestration (TSO)

o TrueSight Server Automation (TSSA)

  • Experience with Remedy * ServiceNow migration

  • Agile/Scrum experience

  • Cloud certifications (AWS/Azure/GCP)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:38 min

Managing and versioning system prompts as YAML files

Kevin Lewis Kevin Lewis +1 · WWC 2025

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

1:31 min

Orchestrating generative configurations using standardized YAML files

Han Xiao · WWC 2022

Videos

See all

Related articles

See all