Site Reliability Engineer

Incite Insight
Bristol, UK
15 days ago
Apply on www.adzuna.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
ÂŁ44,088.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence JIRA Data Centers Data Deduplication Data Center Infrastructure Management (CIM) DevOps Monitoring of Systems Python (Programming Language) Reliability Engineering Prometheus Software Engineering
+7 more
Systems Integration Large Language Models Grafana Slack Hardware Infrastructure Dynatrace Servicenow

Job description

We are recruiting for a growing technology infrastructure business that is building a new operational capability to support large-scale, high-performance computing environments. This is an excellent opportunity for an experienced Site Reliability Engineer or Platform Engineer who enjoys automating things rather than repeatedly fixing them manually. The role sits at the intersection of infrastructure, operations and software engineering. You will use Python and modern automation techniques to improve reliability, reduce manual workload and make incident response faster and more effective. What you’ll be doing

You’ll build Python-based automation around incident management, operational runbooks and routine infrastructure tasks. You’ll integrate monitoring, infrastructure and ITSM platforms through APIs, helping improve the quality of alerts through better correlation, enrichment, suppression and deduplication. You’ll also develop internal tools, command-line utilities, dashboards and potentially ChatOps capabilities that allow operational teams to resolve issues more quickly. A major part of the role will be taking existing operational processes and asking: “Why are we still doing this manually?” You’ll then design a safe, controlled and auditable way of automating it., ServiceNow, Halo, Jira Service Management, OpenTelemetry, distributed tracing, Slack/Teams automation, datacentre or colocation environments, GPU infrastructure, DCIM, IPAM, virtualisation platforms or LLM-assisted operational automation. This is not an AI/ML development position. We’re looking for someone who understands production infrastructure and can use software engineering and automation to make that infrastructure more reliable. You’ll be joining a growing organisation where you’ll have considerable autonomy and the opportunity to help shape the SRE and operational automation capability rather than simply inherit an established environment.

Requirements

You should have good commercial experience in Site Reliability Engineering, Platform Engineering, DevOps or production infrastructure operations, together with strong hands-on Python automation skills. You’ll also need experience with:

  • Monitoring and observability tools such as Prometheus, Grafana or similar
  • Production incident management and/or on-call environments
  • Automating operational runbooks and repetitive infrastructure processes
  • APIs and systems integration
  • Version-controlled automation and operational tooling

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein ¡ LIVE

1:14 min

Automating user bug triage and resolutions using Slack agents

Brian Lovin Brian Lovin ¡ World Congress 2026 Europe

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum ¡ World Congress 2026 Europe

8:02 min

Integrating service level objectives into incident management

Diana Todea ¡ LIVE

5:47 min

Integrating user stories and test automation via Jira tools

Christoph Ruggenthaler ¡ LIVE

6:58 min

Building engineering communities and finding technical inspiration

Videos

See all

Related articles

See all