Site Reliability Engineer

e-Solutions Inc
Leeds, UK
3 days ago
Apply on www.reed.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
£72,800.0 - £78,000.0
Working hours
Regular working hours
Job source

Tech stack

Cloud Computing Computer Programming Continuous Integration DevOps Python (Programming Language) Reliability Engineering Prometheus Scripting Cloud Platform System System Availability Grafana Reliability of Systems
+7 more
Kubernetes Infrastructure Automation Frameworks Deployment Automation Azure AKS Terraform AWS EKS Jenkins

Job description

We are looking for a passionate and experienced Site Reliability Engineer (SRE) to join our Cloud Platform team. The ideal candidate will have hands-on experience managing large-scale Kubernetes clusters on public cloud environments (AKS, EKS, or GKE) and a strong understanding of modern SRE and DevOps practices. You will be responsible for ensuring high availability, reliability, scalability, and performance of our cloud-native infrastructure and CI/CD systems., * Manage, monitor, and optimize large-scale Kubernetes clusters hosted on public cloud platforms (Azure AKS, AWS EKS, or Google GKE).

  • Implement and maintain infrastructure as code using tools such as Terraform.

  • Collaborate with development and operations teams to improve system reliability and deployment automation.

  • Build and maintain CI/CD pipelines using Jenkins or similar tools.

  • Troubleshoot production issues, conduct root cause analysis, and implement preventive measures.

  • Automate operational tasks using Python or other scripting languages.

  • Contribute to observability and monitoring improvements using modern tools and best practices.

  • Participate in on-call rotations and incident response processes.

Requirements

  • 5-9 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles.

  • Strong hands-on experience managing Kubernetes clusters in production (AKS/EKS/GKE).

  • Proficiency with Terraform and cloud infrastructure automation.

  • Practical experience with Jenkins and CI/CD pipeline management.

  • Sound understanding of SRE principles (incident management, blameless postmortems, capacity planning, error budgets, etc.).

  • Good programming or scripting skills in Python (preferred) or similar languages.

  • Strong analytical, troubleshooting, and problem-solving abilities.

  • Excellent written and verbal communication skills.

  • Experience with Prometheus, Grafana, or OpenTelemetry for observability.

  • Exposure to GitOps practices and tools (e.g. Flux). Skills

  • Jenkins CI/CD
  • Kubernetes
  • Python
  • Site Reliability Engineer (Job Titles)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.reed.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:49 min

Container hosting options available on Amazon Web Services

Federico Fregosi · World Congress 2022

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · World Congress 2023

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

5:02 min

Migrating stateful applications between Kubernetes clusters

Michael Cade · World Congress 2022

57 sec

Extracting API schemas automatically during continuous integration builds

Axel Barbier · World Congress 2023

Videos

See all

Related articles

See all