Site Reliability Engineer

Connells Group
Milton Keynes, UK
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
£40,000.0 - £55,000.0
Working hours
Regular working hours

Tech stack

Application Performance Management Microsoft Azure C Sharp (Programming Language) Cloud Computing DevOps Domain Name System (DNS) Github Monitoring of Systems Windows PowerShell Reliability Engineering Software Reliability Testing Site Reliability Engineering Practices
+11 more
Next.js Scripting Performance Testing ReactJS Reliability of Systems Azure Powershell Firewalls (Computer Science) Kubernetes ISO/IEC 27002 Terraform Docker

Job description

We are seeking an experienced Site Reliability Engineer (SRE) to join our Group Technology Team in Milton Keynes.

ConnellsX is Connells Group Technology’s internal developer platform, built on Microsoft Azure. It simplifies cloud hosting, embeds security and compliance by default, and enables a frictionless developer experience. As part of the team building and operating this platform, you will play a hands-on role in ensuring it is reliable, scalable, and observable.

You will help establish and mature SRE practices, focusing on:

  • Monitoring and observability
  • Incident response
  • Post-incident review
  • Reliability testing and capacity planning
  • Toil reduction
  • Enabling development velocity, * Support teams using ConnellsX and respond to incidents in a structured, blameless way
  • Investigate root causes and drive post-incident actions to completion
  • Define SLIs, contribute to SLOs, and monitor error budgets
  • Build dashboards, alerts, and runbooks to improve visibility
  • Automate repetitive tasks to reduce operational toil
  • Collaborate with cross-functional teams to enhance reliability and observability
  • Support performance testing and capacity planning
  • Proactively identify and prioritise reliability improvements

Requirements

  • Hands-on experience with Azure Monitoring (Application Insights, Alerts, Action Groups)
  • Strong knowledge of OpenTelemetry (including Kubernetes)
  • Scripting/automation using PowerShell and/or Azure CLI
  • Experience with Terraform and GitHub Actions
  • Ability to define SLIs/SLOs and manage error budgets
  • Incident response and post-incident review experience
  • Familiarity with Docker and Kubernetes
  • Strong communication and documentation skills

Desirable:

  • Working knowledge of .NET/C# and React/NextJS
  • Experience with cloud cost optimisation
  • Knowledge of Azure networking (DNS, VNets, Firewalls)
  • Understanding of security frameworks (e.g. ISO 27002, NIST CSF)
  • Azure certifications

About You:

You may come from SRE, DevOps, platform engineering, or operations backgrounds. What matters is hands-on experience running production systems, managing incidents, creating runbooks and automating repetitive work. The focus is on identifying root causes and systemic issues, reducing manual toil through automation, and maintaining reliability by applying SRE principles and using data-driven metrics (SLIs/SLOs).

You understand reliability is about balance, not perfection, and can make data-driven trade-offs between stability and delivery. You are curious, collaborative, and take shared responsibility for system reliability.

Please note that we are unable to provide visa sponsorship. Applicants must have the right to work in the UK.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.co.uk

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

43 sec

Software engineering journey and local Manchester roots

Jonathan Tang · Coffee With Developers

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

6:58 min

Building engineering communities and finding technical inspiration

Videos

See all

Related articles

See all