Site Reliability Engineer

Biometric Talent
Bolton, UK
7 days ago
Apply on www.adzuna.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
£40,000.0 - £65,000.0
Working hours
Regular working hours

Tech stack

JavaScript (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Business Analytics Applications Information Technology Operations Python (Programming Language) Reliability Engineering Ansible Shell Script Software Engineering Large Language Models Grafana
+7 more
Reliability of Systems Performance Monitor Terraform Splunk New Relic (SaaS) Pagerduty Golang

Job description

  • Write and contribute to code that improves service reliability and observability
  • Develop tools, operational APIs and automation to improve system management
  • Establish proactive monitoring and alerting across complex platforms
  • Implement service instrumentation using OpenTelemetry
  • Build sophisticated dashboards using Grafana, Splunk and New Relic
  • Automate manual processes and reduce operational toil
  • Work with Infrastructure as Code and orchestration technologies
  • Support live incident resolution and contribute to post-mortem analysis
  • Carry out root cause analysis and implement effective remediation
  • Drive initiatives to improve system reliability, performance and observability
  • Maintain and administer existing monitoring and analytics platforms
  • Work with IT Operations to provide critical tooling and capabilities
  • Share knowledge and mentor colleagues on new technologies and practices
  • Use AI tools, LLM platforms and coding assistants in day-to-day work to improve productivity, reduce toil and explore new approaches to autonomous operations, telemetry and system health

Technologies:

  • AI
  • Ansible
  • Golang
  • Grafana
  • Incident Management
  • Support
  • LLM
  • OpenTelemetry
  • PagerDuty
  • Python
  • Splunk
  • Terraform
  • JavaScript

Requirements

  • Strong software engineering experience, particularly with Python or Golang
  • Experience with monitoring, alerting and observability
  • Knowledge of OpenTelemetry and modern observability practices
  • Experience establishing proactive monitoring and alerting for complex platforms
  • Strong understanding of SRE principles, including SLIs and SLOs
  • Experience with modern software development practices and lifecycles
  • Proficiency in shell scripting
  • Experience with Infrastructure as Code, automation and orchestration, ideally using Terraform and Ansible
  • Experience with tools such as Grafana, Splunk, New Relic and PagerDuty
  • Experience working within large-scale, 24/7 enterprise environments where availability and stability are critical
  • Strong incident management, troubleshooting and root cause analysis experience
  • Hands-on experience using LLM platforms and coding assistants to improve productivity and quality
  • Experience or interest in using AI for telemetry, predictive insights and root-cause analysis

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:07 min

Architecting the availability stack with Prometheus and Grafana

Gabriel Labachelerie · World Congress 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all