> Markdown version of [/jobs/ext/3059098-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3059098-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Biometric Talent - **Location:** Bolton, UK - **Salary:** £40,000.0 - £65,000.0 - **Contract:** Permanent contract - **Skills:** JavaScript (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, Business Analytics Applications, Information Technology Operations, Python (Programming Language), Reliability Engineering, Ansible, Shell Script, Software Engineering, Large Language Models, Grafana, Reliability of Systems, Performance Monitor, Terraform, Splunk, New Relic (SaaS), Pagerduty, Golang - **Published:** September 25, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5896695435 ## About the Role * Strong software engineering experience, particularly with Python or Golang * Experience with monitoring, alerting and observability * Knowledge of OpenTelemetry and modern observability practices * Experience establishing proactive monitoring and alerting for complex platforms * Strong understanding of SRE principles, including SLIs and SLOs * Experience with modern software development practices and lifecycles * Proficiency in shell scripting * Experience with Infrastructure as Code, automation and orchestration, ideally using Terraform and Ansible * Experience with tools such as Grafana, Splunk, New Relic and PagerDuty * Experience working within large-scale, 24/7 enterprise environments where availability and stability are critical * Strong incident management, troubleshooting and root cause analysis experience * Hands-on experience using LLM platforms and coding assistants to improve productivity and quality * Experience or interest in using AI for telemetry, predictive insights and root-cause analysis ## Description * Write and contribute to code that improves service reliability and observability * Develop tools, operational APIs and automation to improve system management * Establish proactive monitoring and alerting across complex platforms * Implement service instrumentation using OpenTelemetry * Build sophisticated dashboards using Grafana, Splunk and New Relic * Automate manual processes and reduce operational toil * Work with Infrastructure as Code and orchestration technologies * Support live incident resolution and contribute to post-mortem analysis * Carry out root cause analysis and implement effective remediation * Drive initiatives to improve system reliability, performance and observability * Maintain and administer existing monitoring and analytics platforms * Work with IT Operations to provide critical tooling and capabilities * Share knowledge and mentor colleagues on new technologies and practices * Use AI tools, LLM platforms and coding assistants in day-to-day work to improve productivity, reduce toil and explore new approaches to autonomous operations, telemetry and system health Technologies: * AI * Ansible * Golang * Grafana * Incident Management * Support * LLM * OpenTelemetry * PagerDuty * Python * Splunk * Terraform * JavaScript ## Related Videos - [Monitoring as Code - Managing your dashboards at scale](https://www.wearedevelopers.com/videos/753-monitoring-as-code-managing-your-dashboards-at-scale) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)