Site Reliability Engineer (SRE)

The Smart
Southlake, TX, United States
12 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Compensation
$83,200.0 - $166,400.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Applications Architecture Microsoft Azure Cloud Computing Computer Networks Linux Disaster Recovery Monitoring of Systems Python (Programming Language) Reliability Engineering Ansible Prometheus
+8 more
Datadog System Availability Grafana Kubernetes Infrastructure Automation Frameworks Information Technology Terraform Splunk

Job description

As a Site Reliability Engineer (SRE), you will be responsible for improving the reliability, scalability, and operational efficiency of production systems through automation, observability, and incident management. The ideal candidate will have strong Python development expertise, hands-on experience supporting large-scale production environments, and a proven track record of reducing operational toil through automation., Develop Python-based automation solutions to eliminate manual operational tasks and improve efficiency. Support and maintain large-scale production systems, ensuring high availability and reliability. Participate in incident response, troubleshooting, root cause analysis, and problem remediation activities. Automate infrastructure management across cloud, Kubernetes, Linux, and Windows environments. Implement and support infrastructure automation using tools such as Terraform and Ansible. Build and maintain observability solutions including dashboards, alerts, metrics, logs, and monitoring frameworks. Drive operational improvements to reduce recurring incidents and improve system stability. Perform performance analysis, capacity planning, and system health assessments. Support disaster recovery, failover testing, and operational readiness initiatives. Evaluate and implement emerging observability, automation, and AIOps capabilities.

Requirements

Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience. 3-5 years of hands-on Site Reliability Engineering or Production Engineering experience supporting large-scale production systems. Strong Python programming expertise with demonstrated experience building automation tools, frameworks, and operational solutions. Proven experience in production operations, incident response, root cause analysis, and reliability engineering. Experience automating operational processes and reducing manual toil through engineering solutions. Hands-on experience with Kubernetes and cloud platforms such as GCP, AWS, or Azure. Experience with monitoring and observability tools including Splunk, Grafana, Prometheus, Datadog, or similar platforms. Strong understanding of Linux systems, networking concepts, and distributed application architectures. Experience with Infrastructure as Code and configuration management tools such as Terraform and Ansible. Excellent analytical, troubleshooting, and problem-solving skills with the ability to perform effectively in mission-critical environments. The hourly range for roles of this nature are $40.00 to $80.00/hr. Rates are heavily dependent on skills, experience, location, and industry. cyberThink is an Equal Opportunity Employer.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all