Site Reliability Engineer

REVYBE IT RECRUITMENT LIMITED
London, UK
7 days ago
Apply on www.totaljobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
£85,000.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Software as a Service Cloud Computing DevOps Disaster Recovery Github Monitoring of Systems Performance Tuning Reliability Engineering Software Engineering Data Logging Scripting
+3 more
Reliability of Systems Kubernetes Terraform

Job description

This is a hands-on role focused on AWS, Kubernetes, Terraform, observability, monitoring, and automation, working closely with software engineering teams to improve platform reliability and developer experience.

You’ll have genuine ownership and the opportunity to influence how the platform evolves as the business continues to scale.

What You’ll Be Doing

  • Design, build, and maintain highly available and scalable AWS infrastructure
  • Manage and optimise Kubernetes environments and containerised workloads
  • Build and maintain infrastructure using Terraform and Infrastructure as Code principles
  • Develop and optimise CI/CD pipelines using GitHub Actions
  • Build and improve comprehensive monitoring and observability across the platform
  • Implement and maintain effective logging, metrics, tracing, alerting, and dashboards
  • Define and improve SLIs, SLOs, and reliability metrics
  • Proactively identify and resolve performance, availability, and reliability issues
  • Lead and contribute to incident response, troubleshooting, and root cause analysis
  • Automate operational processes and eliminate repetitive manual tasks
  • Work closely with software engineers to improve deployment processes, system reliability, and developer experience
  • Help improve platform resilience, scalability, and disaster recovery capabilities
  • Contribute to capacity planning and performance optimisation as the platform scales
  • Establish and champion SRE best practices across the wider engineering function

Requirements

  • Proven commercial experience working as an SRE, DevOps Engineer, Platform Engineer, or similar
  • Strong hands-on experience with AWS
  • Strong experience working with Kubernetes
  • Excellent experience with Terraform and Infrastructure as Code
  • Strong experience building and managing GitHub Actions CI/CD pipelines
  • Solid experience with monitoring and observability tooling
  • Strong understanding of metrics, logging, tracing, alerting, and system health
  • Experience troubleshooting complex production environments
  • Understanding of SLIs, SLOs, SLAs, and error budgets
  • Experience with incident management and root cause analysis
  • Good understanding of cloud networking, security, and infrastructure fundamentals
  • Strong scripting/automation skills
  • A strong understanding of reliability, scalability, performance, and availability
  • Excellent communication skills and the ability to work closely with software engineering teams
  • A proactive mindset and genuine passion for automation and continuous improvement, If you’re an experienced SRE, Platform Engineer or DevOps Engineer who enjoys solving complex reliability challenges and wants to have a real impact within a rapidly growing SaaS business, we’d love to hear from you.

About the company

The company is open to speaking with engineers who may not have experience across every technology listed above.

If you have strong foundations in AWS, Kubernetes, Terraform, and cloud infrastructure, along with a genuine interest in reliability and observability, we’d still love to hear from you.

Why Join?

  • Join a fast-growing SaaS company at an exciting stage of its journey
  • Work with a modern AWS and Kubernetes environment
  • Take ownership of reliability, automation, and platform performance
  • Work with modern observability and monitoring technologies
  • Have genuine influence over engineering and platform decisions
  • Work closely with talented software engineering teams
  • Clear opportunities to progress as the business continues to scale
  • Help shape and mature the company’s SRE practices

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.totaljobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

2:33 min

Advocating for SRE practices within agency environments

Martin Beránek · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

Videos

See all

Related articles

See all