Site Reliability Engineer

Lorien
Midlothian, UK
about 1 month ago

Role details

Contract type
Contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Amazon Web Services Cloud Computing Distributed Systems Reliability Engineering VMware Infrastructure Cloud Platform System Grafana Apigee Splunk Vmware
+1 more
Microservices

Job description

Site Reliability Engineer

Hybrid Working - Edinburgh - 2 days a month on site.

Financial Services

Lorien’s leading banking client is looking for an experienced Site Reliability Engineer (SRE) to join an established engineering team responsible for supporting and improving the reliability, availability and performance of critical production services. This is an operationally focused role where you’ll work closely with engineering teams to ensure systems remain resilient, observable and highly available.

The ideal candidate will have a strong production engineering or SRE background, be comfortable working within live environments, with good skills of working with Grafana, Open Telemetry, Splunk, and knowledge of APIs.

This role is based in Edinburgh.

This role will be Via Umbrella.

Working in a Hybrid Model of 2 days a month on site.

Key Experience

  • Support and maintain production systems, ensuring high levels of availability and reliability.
  • Respond to, troubleshoot and resolve production incidents, taking ownership through to resolution.
  • Focus on incident response, service restoration and operational excellence (approximately 70% of the role).
  • Improve system observability, monitoring and alerting capabilities.
  • Work closely with development teams to enhance the reliability and operability of applications.
  • Analyse production issues and identify opportunities for automation and continuous improvement.
  • Participate in root cause analysis and implement preventative measures to reduce recurring incidents.
  • Support cloud-based infrastructure and operational tooling as the organisation transitions from VMware to AWS..

Key Skills

  • Approximately 6+ years’ experience within Site Reliability Engineering, Production Engineering or a similar operational engineering role.
  • Strong hands-on experience supporting live production environments.
  • Excellent troubleshooting and incident management skills.
  • Experience with observability and monitoring platforms, including:

  • Grafana
  • Open Telemetry
  • Splunk

Good understanding of cloud platforms (AWS experience preferred).

Strong knowledge of APIs and API troubleshooting.

Experience working within modern distributed systems and production environments.

Ability to quickly become productive within an existing engineering team.

Highly Desirable

  • Experience with Apigee API Management.
  • Experience supporting Java-based microservices.
  • Knowledge of VMware infrastructure.
  • Experience supporting organisations migrating from VMware to AWS.

IND_PC3

Guidant, Carbon60, Lorien & SRG - The Impellam Group Portfolio are acting as an Employment Business in relation to this vacancy.

Requirements

  • Approximately 6+ years’ experience within Site Reliability Engineering, Production Engineering or a similar operational engineering role.
  • Strong hands-on experience supporting live production environments.
  • Excellent troubleshooting and incident management skills.
  • Experience with observability and monitoring platforms, including:

  • Grafana
  • Open Telemetry
  • Splunk

Good understanding of cloud platforms (AWS experience preferred).

Strong knowledge of APIs and API troubleshooting.

Experience working within modern distributed systems and production environments.

Ability to quickly become productive within an existing engineering team.

Highly Desirable

  • Experience with Apigee API Management.
  • Experience supporting Java-based microservices.
  • Knowledge of VMware infrastructure.
  • Experience supporting organisations migrating from VMware to AWS.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on computerjobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:40 min

Managing containerized infrastructure with Podman Desktop

Cedric Clyburn Cedric Clyburn +1 · WWC 2025

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

1:41 min

Parallels between cloud and legacy infrastructure lock-ins

Björn Stahl Björn Stahl · WWC 2024

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all