Senior Site Reliability Engineer

GCS Ltd
Glasgow, UK
11 days ago
Apply on find.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
£75,000.0 - £95,000.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Amazon Web Services Microsoft Azure Bash Shell C Sharp (Programming Language) Cloud Computing Computer Programming Distributed Systems Python (Programming Language) Reliability Engineering Prometheus Software Engineering
+4 more
Scripting Grafana Kubernetes Golang

Job description

Build and maintain reliable, scalable and secure infrastructure platforms and solutions. Apply SRE and software engineering practices to improve reliability, availability and performance. Monitor systems, manage incidents and lead complex troubleshooting and root cause analysis. Develop automation using programming and scripting to reduce manual intervention and improve efficiency. Develop and improve observability, monitoring, instrumentation and performance capabilities. Use data and reliability metrics to drive continuous improvement and optimisation. Lead technical discussions, blameless retrospectives and problem-solving activities. Work with architects, engineers and stakeholders to define requirements and deliver effective solutions. Provide technical leadership, mentoring and guidance while helping to drive SRE maturity across teams.

Requirements

We’re recruiting for an experienced Senior Site Reliability Engineer to drive reliability, scalability and performance across critical banking systems. This role combines hands-on SRE engineering with technical leadership, with a strong focus on observability, automation, continuous improvement and optimisation., 5+ years’ experience in SRE, Production Engineering, Platform Engineering or a related discipline. Strong practical experience with SRE principles and production environments. Strong programming/scripting skills, such as Python, Go, Java, C# or Bash. Proven experience in incident management, troubleshooting and root cause analysis. Strong observability and monitoring experience. Experience with AWS, Azure or GCP. Good understanding of operating systems, networking, cloud infrastructure and automation. Experience with Infrastructure-as-Code. Strong communication and technical leadership skills.

Desirable Skills: SLOs, SLIs, SLAs and error budgets. Prometheus, Grafana, Elastic/ELK or OpenTelemetry. Kubernetes, containers and distributed systems. Performance and resilience engineering. Experience driving SRE maturity across engineering teams. Financial services or banking experience.

This is an excellent opportunity for a Senior SRE to combine hands-on engineering with technical leadership and help improve reliability, observability and performance across critical banking systems.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on find.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

Videos

See all

Related articles

See all