Site Reliability Engineer (SRE)

Randstad
Plano, TX, United States
10 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Bash Shell Databases Continuous Integration DevOps Monitoring of Systems Python (Programming Language) Windows PowerShell Site Reliability Engineering Practices Ansible Prometheus
+8 more
Google Cloud Delivery Pipeline Grafana Reliability of Systems Apache Kafka Terraform Docker Jenkins

Job description

  • Design and maintain highly available scalable and faulttolerant systems
  • Implement and manage monitoring ing and observability tools Grafana Prometheus etc
  • Participate in incident management RCA and postmortems
  • Automate operational tasks using scripting and tools Python Bash etc
  • Collaborate with development infrastructure and support teams to improve system reliability
  • Drive adoption of SRE practices like SLIs SLOs and error budgets
  • Ensure performance optimisation capacity planning and system stability
  • Build and maintain automation pipelines and infrastructure as code

Requirements

Strong experience in SRE Production Support DevOps environment

Handson with LinuxUnix systems

Experience with cloud platforms AWS Azure Google Cloud Platform

Expertise in Docker Kubernetes

Knowledge of monitoring tools Grafana Prometheus OpenSearch Instana

Strong troubleshooting incident management skills

Experience with Infrastructure as Code Terraform Ansible

Programmingscripting using Python Bash PowerShell

skills:

cloud platforms AWS,Ansible,Kafka,Bash,CICD,databases,automation pipelines,DevOps environment,Docker Kubernetes,Grafana,Jenkins,Azure Google Cloud Platform,monitoring tools,Prometheus,Python,system reliability,SRE practices,Terraform,PowerShell,Automate,Automation & Scripting,budgets,capacity planning,Strategy Design,incident management,ITIL processes,infrastructure,Production Support

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · World Congress 2023

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all