Senior Site Reliability Engineer

Understanding Recruitment
Greater London, UK
4 days ago
Apply on www.understandingrecruitment.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
£150,000.0 - £200,000.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Continuous Integration Linux DevOps Programming Tools Key Management Networking Basics Reliability Engineering Ansible Data Logging Cloud Platform System Delivery Pipeline
+4 more
Software Troubleshooting Infrastructure Automation Frameworks Low Latency Terraform

Job description

📍 London

💰 £150,000 - £200,000+ Base + Bonus + Equity

We’re partnered with a technology company building high-performance infrastructure for decentralised financial markets.

They’re looking for a Senior Site Reliability Engineer to improve the reliability, observability and operational tooling behind a latency-sensitive production platform.

The role covers production infrastructure, monitoring and alerting, incident diagnosis, deployment workflows, infrastructure automation and developer tooling. There is also a strong Linux and systems element, particularly around networking, host performance and running high-performance services in production.

The platform is still relatively early, so there is plenty of scope to improve how things are operated, introduce better automation and help set the standards the wider engineering team works to.

Responsibilities

  • Improve the reliability and operability of production systems.
  • Build and improve monitoring, logging, tracing, dashboards and alerting.
  • Improve incident diagnosis, root cause analysis and operational workflows.
  • Build safer and more repeatable deployment and rollback processes.
  • Automate repetitive operational and infrastructure work.
  • Improve CI/CD pipelines and release processes.
  • Develop internal tooling that helps engineers operate production systems more effectively.
  • Improve the developer experience from local development through to production.
  • Work with Linux systems, networking, host configuration and resource contention.
  • Contribute to infrastructure security, access controls, secrets management and system hardening.

The systems are latency-sensitive, so the role can extend into areas such as host-level tuning, kernel settings, CPU isolation and networking behaviour.

Skills & Experience

  • Strong experience in Site Reliability Engineering, Platform Engineering, DevOps or Infrastructure Engineering.
  • Experience operating production infrastructure in cloud environments.
  • Strong Linux systems knowledge and understanding of networking fundamentals.
  • Experience with monitoring, observability and alerting.
  • Strong troubleshooting and root cause analysis skills.
  • Experience with CI/CD and infrastructure automation.
  • AWS, Terraform or Ansible experience would be advantageous.
  • Experience with high-performance, high-throughput or latency-sensitive systems would be particularly valuable.
  • Comfortable taking ownership of problems and driving improvements independently.

Benefits

  • £150,000 - £200,000+ base salary.
  • Significant performance-based bonus + Equity
  • Private healthcare.
  • UK visa sponsorship available.
  • Engineering-led organisation - built prioritising engineering culture
  • Direct influence over reliability, tooling and engineering practices.
  • Opportunity to work alongside a small, elite team.
  • Exposure to complex, latency-sensitive production systems.

Interested?

Contact Chris Williams with any questions.

Requirements

  • Strong experience in Site Reliability Engineering, Platform Engineering, DevOps or Infrastructure Engineering.
  • Experience operating production infrastructure in cloud environments.
  • Strong Linux systems knowledge and understanding of networking fundamentals.
  • Experience with monitoring, observability and alerting.
  • Strong troubleshooting and root cause analysis skills.
  • Experience with CI/CD and infrastructure automation.
  • AWS, Terraform or Ansible experience would be advantageous.
  • Experience with high-performance, high-throughput or latency-sensitive systems would be particularly valuable.
  • Comfortable taking ownership of problems and driving improvements independently.

Benefits & conditions

  • £150,000 - £200,000+ base salary.
  • Significant performance-based bonus + Equity
  • Private healthcare.
  • UK visa sponsorship available.
  • Engineering-led organisation - built prioritising engineering culture
  • Direct influence over reliability, tooling and engineering practices.
  • Opportunity to work alongside a small, elite team.
  • Exposure to complex, latency-sensitive production systems.

Interested?

Contact Chris Williams with any questions.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.understandingrecruitment.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

6:58 min

Building engineering communities and finding technical inspiration

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

Videos

See all

Related articles

See all