Senior Site Reliability Engineer (SRE) - Hybrid

The Smart
Austin, TX, United States
10 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$124,800.0 - $166,400.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) .NET Framework Artificial Intelligence Automation of Tests Bash Shell Cloud Computing Computer Networks Databases Continuous Integration Dynamic Host Configuration Protocol Linux Distributed Systems
+26 more
Domain Name System (DNS) IBM WebSphere MQ Python (Programming Language) Machine Learning Enterprise Messaging Systems MongoDB Routing Oracle (Applications) Performance Tuning Windows PowerShell RabbitMQ Reliability Engineering Software Deployment Software Engineering SQL Databases System Availability Grafana Reliability of Systems Firewalls (Computer Science) Git Flow Information Technology Deployment Automation Apache Kafka Splunk Appdynamics Programming Languages

Job description

As a Senior Site Reliability Engineer (SRE), you will lead reliability, observability, automation, and operational excellence initiatives for enterprise-scale applications and platforms. This role combines software engineering, infrastructure operations, and AI/ML-driven observability to enhance system reliability, reduce operational toil, and drive automation across mission-critical environments., Design and implement automation solutions that improve operational efficiency and platform reliability. Lead AI/ML-driven observability, anomaly detection, predictive alerting, and AIOps initiatives. Develop scripts, tools, and frameworks to automate infrastructure and operational processes. Expand automation coverage across deployment, monitoring, alerting, and self-healing workflows. Collaborate with engineering, operations, and product teams to improve system availability and performance. Troubleshoot critical production issues and drive root cause analysis and remediation efforts. Build and maintain monitoring dashboards, alerting frameworks, and operational intelligence solutions. Support CI/CD, GitOps, and deployment automation strategies to accelerate software delivery. Conduct capacity planning, performance analysis, and operational readiness assessments. Participate in on-call support and ensure reliable operation of mission-critical systems.

Requirements

Bachelor’s degree in Computer Science, Engineering, or a related technical field. 6-8 years of experience supporting and administering enterprise-scale production environments. 6-8 years of experience developing automation scripts, monitoring solutions, dashboards, and alerting frameworks. Strong experience with Linux and Windows system administration, troubleshooting, tuning, and deployments. Proficiency in one or more programming languages including Python, PowerShell, Java, .NET, or Bash. Experience with cloud platforms, application deployments, migrations, and high-availability architectures. Knowledge of networking concepts including DNS, DHCP, firewalls, routing, and distributed systems. Experience with observability tools such as Splunk, AppDynamics, or similar monitoring platforms. Strong understanding of databases and messaging technologies including SQL, Oracle, MongoDB, Kafka, RabbitMQ, Solace, or IBM MQ. Proven experience applying AI/ML, AIOps, predictive alerting, or ML-driven observability solutions in production environments. The hourly range for roles of this nature are $60.00 to $80.00/hr. Rates are heavily dependent on skills, experience, location, and industry. cyberThink is an Equal Opportunity Employer.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Reevaluating engineering careers at major technology corporations

Chris Heilmann +2 · LIVE

2:04 min

Enhancing network privacy with routing fees and onion routing

Andreas M Antonopoulos · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:51 min

Overview of the three Google Maps routing applications

Germán Álvarez · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all