Site Reliability Engineer

Era4
UK
7 days ago
Apply on www.adzuna.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
£63,583.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence JIRA Data Centers Data Deduplication Data Center Infrastructure Management (CIM) Python (Programming Language) Reliability Engineering Prometheus Software Engineering Systems Integration AI Infrastructure
+5 more
Data Logging Large Language Models Grafana Dynatrace Servicenow

Job description

We’re hiring SRE/Platform engineers with an automation bias to help build Era4’s operations capability from the ground up. You’ll turn runbooks, alerts and operational workflows into safe, auditable automation and internal tooling that improves reliability across our AI infrastructure and datacentre platform.

This is a Platform / SRE role with software engineering, not an AI model-building role. You’ll work closely with operations, platform and engineering teams to reduce manual toil, improve alert quality, and speed up incident response., * Build Python-based automation for incident triage, runbook execution, and routine operational tasks.

  • Integrate observability, ITSM and infrastructure APIs to enrich alerts and automate workflows.
  • Improve monitoring signal quality through correlation, enrichment, suppression and deduplication.
  • Build internal tools and self-service capabilities such as CLI utilities, ChatOps integrations and dashboards.
  • Maintain version-controlled runbook-as-code and automation libraries.
  • Translate post-incident learnings into better tooling, automation and operational standards.
  • Support safe, auditable automation for higher-risk actions with appropriate approval controls.

Requirements

  • Experience in SRE, Platform Engineering, or production infrastructure operations.
  • Hands-on experience with observability/monitoring tooling (for example Prometheus, Grafana or similar).
  • Exposure to incident management / on-call and converting manual runbooks into automation.
  • Experience with Python for automation, APIs and integrations.

Nice To Have:

  • GPU, datacentre or colocation infrastructure experience.
  • ITSM integrations (ServiceNow, Halo, Jira Service Management or similar).
  • ChatOps tooling (Slack or Microsoft Teams bots).
  • OpenTelemetry, logging or distributed tracing experience.
  • DCIM, IPAM or hypervisor-control-plane integrations.
  • Experience with LLM-assisted or agent-based operational automation.

About the company

Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · World Congress 2025

1:07 min

Architecting the availability stack with Prometheus and Grafana

Gabriel Labachelerie · World Congress 2023

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

12:08 min

Comparing Keptn orchestration capabilities against alternative software operators

Thomas Schütz · LIVE

Videos

See all

Related articles

See all