Site Reliability Engineer - Observability & Automation

Manchester Digital
Manchester, UK
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Reliability Engineering Ansible Grafana Reliability of Systems Terraform Splunk

Requirements

Manchester Digital is seeking a Site Reliability Engineer to enhance system reliability and observability. This role focuses on monitoring critical systems and improving operational efficiency through engineering solutions.The ideal candidate will be proficient in tools like Splunk and Grafana, with experience in automation using Ansible and Terraform. You will collaborate across teams to establish reliability practices and participate in live incident resolution, ensuring our systems meet user demands. #J-18808-Ljbffr

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on apply4u.co.uk

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · WWC 2024

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

Videos

See all

Related articles

See all