Site Reliability Engineer

Selby Jennings
London, UK
1 day ago
Apply on www.adzuna.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
£73,358.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) C++ (Programming Language) Cloud Computing Linux Distributed Systems Monitoring of Systems High-Frequency Trading Python (Programming Language) Reliability Engineering Prometheus Software Engineering Grafana
+4 more
Kubernetes Infrastructure Automation Frameworks Splunk Golang

Job description

Our client, a world renowned hedge fund, is seeking a Site Reliability Engineer to join its world-class engineering organization in London. This role sits at the intersection of software engineering and infrastructure, focusing on the reliability, scalability, and performance of the technology platforms that power global trading and investment operations.

Working closely with software engineers, quantitative researchers, traders, and infrastructure teams, you will be responsible for building automation, improving observability, and ensuring critical production systems operate at the highest levels of availability and efficiency., * Design, build, and maintain highly reliable, scalable, and automated infrastructure platforms.

  • Drive improvements in system performance, monitoring, observability, and operational efficiency.
  • Troubleshoot and resolve complex production incidents across distributed systems.
  • Develop tools and automation to reduce operational overhead and improve platform resilience.
  • Partner with engineering teams to improve system design, deployment processes, and operational readiness.
  • Participate in incident management and root cause analysis, ensuring lessons learned are incorporated into future improvements.
  • Support mission-critical trading and research environments in a fast-paced, high-performance setting.

Requirements

  • Strong software engineering skills in Python, Go, C++, Java, or a similar language.
  • Deep Linux systems knowledge and experience operating large-scale production environments.
  • Experience with Kubernetes, containerisation technologies, and cloud infrastructure.
  • Strong understanding of networking, distributed systems, and infrastructure automation.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Splunk, or similar.
  • Proven track record of solving complex reliability, scalability, or performance challenges.
  • Excellent problem-solving skills and ability to operate effectively in high-pressure environments.

Preferred Backgrounds

  • Technology companies operating large-scale distributed systems.
  • High-frequency trading firms, hedge funds, or electronic trading environments.
  • Cloud infrastructure, platform engineering, or production engineering teams.
  • Software engineers with a strong interest in reliability and infrastructure.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all