Site Reliability Engineer

Biometric Talent
Manchester, UK
1 day ago
Apply on itjobpro.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
£40,000.0 - £65,000.0
Working hours
Regular working hours
Job source

Tech stack

JavaScript (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Business Analytics Applications Information Technology Operations Python (Programming Language) Reliability Engineering Ansible Shell Script Software Engineering Large Language Models Grafana
+9 more
Reliability of Systems Kubernetes Performance Monitor Software Coding Terraform Splunk New Relic (SaaS) Pagerduty Golang

Job description

We’re supporting our client with the appointment of an experienced Site Reliability Engineer (SRE) to join their Platform Engineering team.

This is a software-focused SRE position, where your software engineering skills will be used to solve operational problems and improve the reliability, observability and performance of a large-scale production environment.

Working closely with Development, Platform Delivery and IT Operations teams, you’ll help build the tooling, automation, monitoring and practices that keep critical systems reliable and resilient.

If you’re a software engineer who enjoys solving complex operational problems, writing code and using engineering to improve system reliability, this could be an interesting next step.

How you’ll spend your day

You’ll work across software engineering, observability, automation and incident management, with a focus on proactively improving the health and performance of critical services.

Key responsibilities include:

  • Writing and contributing to code that improves service reliability and observability

  • Developing tools, operational APIs and automation to improve system management

  • Establishing proactive monitoring and alerting across complex platforms

  • Implementing service instrumentation using OpenTelemetry

  • Building sophisticated dashboards using Grafana, Splunk and New Relic

  • Automating manual processes and reducing operational toil

  • Working with Infrastructure as Code and orchestration technologies

  • Supporting live incident resolution and contributing to post-mortem analysis

  • Carrying out root cause analysis and implementing effective remediation

  • Driving initiatives to improve system reliability, performance and observability

  • Maintaining and administering existing monitoring and analytics platforms

  • Working with IT Operations to provide critical tooling and capabilities

  • Sharing knowledge and mentoring colleagues on new technologies and practices

The role also has a strong focus on AI-enabled engineering. You’ll use AI tools, LLM platforms and coding assistants in your day-to-day work to improve productivity, reduce toil and explore new approaches to autonomous operations, telemetry and system health.

Technology environment

The successful candidate will work across a modern engineering environment, with technologies including:

Python, Golang and JavaScript

OpenTelemetry

Grafana

Splunk

New Relic

PagerDuty

Ansible

Terraform

Infrastructure as Code

Shell scripting

Automation and orchestration platforms

AI tools, LLM platforms and coding assistants, Should we both wish to proceed, we will submit your details to the client and be in touch regarding the outcome and any further steps.

Requirements

Our client is looking for a strong software engineer with SRE experience, or someone who has a strong software engineering background and has moved into Site Reliability Engineering.

You’ll ideally have:

  • Strong software engineering experience, particularly with Python or Golang

  • Experience with monitoring, alerting and observability

  • Knowledge of OpenTelemetry and modern observability practices

  • Experience establishing proactive monitoring and alerting for complex platforms

  • Strong understanding of SRE principles, including SLIs and SLOs

  • Experience with modern software development practices and lifecycles

  • Proficiency in shell scripting

  • Experience with Infrastructure as Code, automation and orchestration, ideally using Terraform and Ansible

  • Experience with tools such as Grafana, Splunk, New Relic and PagerDuty

  • Experience working within large-scale, 24/7 enterprise environments where availability and stability are critical

  • Strong incident management, troubleshooting and root cause analysis experience

They’re also looking for an AI-native approach to engineering, with hands-on experience using LLM platforms and coding assistants to improve productivity and quality. Experience or interest in using AI for telemetry, predictive insights and root-cause analysis would be particularly relevant.

Benefits & conditions

  • Pension Scheme

  • Hybrid Working

  • Flexible Working Hours - 40-hour workweek with flexibility in how hours are structured.

  • Generous Annual Leave - 25 days holiday + your birthday off, plus bank holidays. Option to buy or sell up to 5 additional days.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on itjobpro.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · World Congress 2025

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all