Site Reliability Engineer

Insight Global
Plano, TX, United States
1 day ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$122,720.0 - $153,920.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Confluence JIRA Microsoft Azure Cloud Computing Continuous Integration DevOps Github Python (Programming Language) Windows PowerShell Scrum Methodology Reliability Engineering
+15 more
Ansible Prometheus YAML Cloud Monitoring Grafana Backend Hashicorp Bitbucket Terraform Splunk Ansible Tower Dynatrace Jenkins Servicenow Artifactory

Job description

Insight Global is seeking a Site Reliability Engineer for a top enterprise client. This role focuses on supporting and administering Terraform Enterprise and enterprise automation platforms in a high-availability production environment. The ideal candidate brings strong Terraform and IaC expertise, deep production support experience, and a passion for root cause analysis, monitoring, and reliability engineering. You’ll work closely with cloud, DevOps, and product engineering teams to improve automation, streamline deployments, support cutovers, and ensure platform stability across AWS and hybrid environments. This is a senior-level opportunity with leadership visibility, especially for candidates stepping into a Lead role in Arizona.

Requirements

  • 5-8 years of SRE / DevOps experience Azure cloud experience

  • Strong Terraform experience o Terraform Enterprise development and administration (backend/platform)

  • Cutover experience (HUGE plus, near must-have)
  • Change Management experience o Incident * Problem Management * RCA / PKE o Hands-on with ITSM tools (Remedy and/or ServiceNow)

  • Ansible Platform / ATaaS o Administration experience (not just a user) o Onboarding applications to Ansible Tower

  • Infrastructure as Code (IaC) using Terraform, HCL, YAML
  • Cloud Monitoring & APM o Prometheus, Grafana, Dynatrace, Splunk (any combination) o Building dashboards, consoles, and alerts

  • Python and/or PowerShell scripting
  • Strong root cause analysis and troubleshooting skills
  • Experience supporting production environments

Nice to Have Skills & Experience

  • HashiCorp Terraform Enterprise certifications
  • CI/CD tools: GitHub, Jenkins, Artifactory
  • Horizon toolset (Ansible, Jira, Confluence, Bitbucket)
  • Automation tools: o Cutover o TrueSight Orchestration (TSO) o TrueSight Server Automation (TSSA)

  • Experience with Remedy * ServiceNow migration
  • Agile/Scrum experience
  • Cloud certifications (AWS/Azure/GCP)

Benefits & conditions

Benefit packages for this role will start on the 1st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.insightglobal.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:38 min

Managing and versioning system prompts as YAML files

Kevin Lewis Kevin Lewis +1 · WWC 2025

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

1:31 min

Orchestrating generative configurations using standardized YAML files

Han Xiao · WWC 2022

Videos

See all

Related articles

See all