Site Reliability Engineer (SRE)

Ascii Group, LLC
Dallas, TX, United States
1 day ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$106,080.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Application Performance Management Automation of Tests Microsoft Azure Backup Devices Bash Shell Cloud Computing Continuous Integration DevOps Disaster Recovery Github Python (Programming Language)
+18 more
Linux System Administration Log Analysis Windows PowerShell Reliability Engineering Ansible Prometheus Data Logging Google Cloud Cloud Monitoring System Availability Grafana Mttr Kubernetes Infrastructure Automation Frameworks Terraform Docker Elk Stack Jenkins

Requirements

· 10+ years of experience in Site Reliability Engineering, DevOps, cloud infrastructure, Linux administration, production operations, or platform engineering.

· Strong hands-on experience designing, operating, and supporting cloud infrastructure across AWS, Azure, and Google Cloud Platform.

· Deep experience with Kubernetes platforms such as EKS, AKS, GKE, and containerization using Docker.

· Experience building and maintaining CI/CD pipelines using Jenkins, GitHub Actions.

· Strong observability experience with Prometheus, Grafana, ELK Stack, OpenSearch, Log Analytics, Application Insights, and Google Cloud Platform Cloud Monitoring.

· Experience with disaster recovery, high availability, backup automation, multi-region failover, and recovery validation.

· Hands-on scripting and automation experience using Python, Bash, PowerShell, and Ansible.

· Linux systems administration experience across enterprise production environments.

Key Responsibilities:

· Site Reliability Engineering and Reliability Governance (SLIs, SLOs, error budgets, reliability reviews, and blameless postmortem practices, MTTD, MTTR, recurring incidents)

· Kubernetes, Containers, and Platform Engineering (Kubernetes, EKS, AKS, GKE, ECS, and Docker-based platforms)

· Infrastructure as Code and Automation (Terraform)

· CI/CD and Release Reliability (GitHub Actions, blue-green, canary, rolling deployments, automated rollback, deployment validation, and automated testing)

· Observability, Monitoring, and Logging (Prometheus, Grafana)

· Disaster Recovery, High Availability, and Resilience

· Security, Compliance, and Cloud Governance

· Linux Systems Administration and Production Support

Have Skills

· Google Cloud Platform (Google Cloud Platform)

· SRE

· Kubernetes & Docker

· Terraform & Infrastructure Automation

· CI/CD & Release Engineering

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all