Senior Site Reliability Engineer

Selby Jennings
Greater London, UK
14 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Cloud Foundry Linux Python (Programming Language) Reliability Engineering Containerization Kubernetes Docker

Job description

Our client, a leading systematic hedge fund, is seeking a Senior Site Reliability Engineer to join their London-based platform team. In this role, you will focus on enhancing the reliability, resilience, and day-to-day operability of a rapidly scaling engineering platform. You will work closely with software engineers and platform owners to strengthen observability, improve incident response processes, and drive measurable reliability outcomes., * Own the effectiveness of the observability platform, ensuring high-quality signals, alert fidelity, and ongoing suitability as the platform scales.

  • Build and maintain actionable, low-noise dashboards and alerting across metrics and logs.
  • Define and apply SLIs and SLOs where they support operational decision-making.
  • Apply IaaC across observability and supporting systems.
  • Improve the reliability, scalability, and operability of core services through hands-on engineering changes.

Requirements

To be successful, you will bring hands-on experience applying SRE principles in production environments, alongside strong expertise in Linux systems. You must be capable of building and operating containerized workloads using tools such as Docker or Podman, and hold strong experience in Go and/or Python. We are looking for a highly technical individual with strong Infrastructure-as-Code proficiency, and the ability to effectively query, interpret, and reason about metrics using PromQL. A key part of this role will involve owning and improving the overall effectiveness of the platform's observability., * Strong practical experience applying SRE principles in production environments.

  • Strong development experience in Go and/or Python.
  • OpenTelemetry experience (metrics, logs, traces).
  • Kubernetes and cloud-native platform experience.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all