Site Reliability Engineer
Biometric Talent
Bolton, UK
7 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.adzuna.co.uk
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
£40,000.0 - £65,000.0
Working hours
Regular working hours
Job source
Tech stack
JavaScript (Programming Language)
Application Programming Interfaces (APIs)
Artificial Intelligence
Business Analytics Applications
Information Technology Operations
Python (Programming Language)
Reliability Engineering
Ansible
Shell Script
Software Engineering
Large Language Models
Grafana
+7 more
Reliability of Systems
Performance Monitor
Terraform
Splunk
New Relic (SaaS)
Pagerduty
Golang
Job description
- Write and contribute to code that improves service reliability and observability
- Develop tools, operational APIs and automation to improve system management
- Establish proactive monitoring and alerting across complex platforms
- Implement service instrumentation using OpenTelemetry
- Build sophisticated dashboards using Grafana, Splunk and New Relic
- Automate manual processes and reduce operational toil
- Work with Infrastructure as Code and orchestration technologies
- Support live incident resolution and contribute to post-mortem analysis
- Carry out root cause analysis and implement effective remediation
- Drive initiatives to improve system reliability, performance and observability
- Maintain and administer existing monitoring and analytics platforms
- Work with IT Operations to provide critical tooling and capabilities
- Share knowledge and mentor colleagues on new technologies and practices
- Use AI tools, LLM platforms and coding assistants in day-to-day work to improve productivity, reduce toil and explore new approaches to autonomous operations, telemetry and system health
Technologies:
- AI
- Ansible
- Golang
- Grafana
- Incident Management
- Support
- LLM
- OpenTelemetry
- PagerDuty
- Python
- Splunk
- Terraform
- JavaScript
Requirements
- Strong software engineering experience, particularly with Python or Golang
- Experience with monitoring, alerting and observability
- Knowledge of OpenTelemetry and modern observability practices
- Experience establishing proactive monitoring and alerting for complex platforms
- Strong understanding of SRE principles, including SLIs and SLOs
- Experience with modern software development practices and lifecycles
- Proficiency in shell scripting
- Experience with Infrastructure as Code, automation and orchestration, ideally using Terraform and Ansible
- Experience with tools such as Grafana, Splunk, New Relic and PagerDuty
- Experience working within large-scale, 24/7 enterprise environments where availability and stability are critical
- Strong incident management, troubleshooting and root cause analysis experience
- Hands-on experience using LLM platforms and coding assistants to improve productivity and quality
- Experience or interest in using AI for telemetry, predictive insights and root-cause analysis
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.adzuna.co.uk
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
over 2 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
about 2 months ago
LM
Luis Minvielle
Why Upskilling And Reskilling is Important For Developers
over 2 years ago
EM
Eli McGarvie
Data Engineer Salary UK
over 3 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
over 2 years ago
LM
Luis Minvielle
Fully Remote Software Engineer Jobs
over 2 years ago