Platform SRE Engineer

FT Select
UK
2 days ago
Apply on ftselect.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Microsoft Azure Disaster Recovery Prometheus Datadog Grafana Splunk New Relic (SaaS) Dynatrace Service Stack

Job description

  • Validate backup and restore processes for Tier 0/Tier 1 platforms.
  • Improve and test recovery runbooks.
  • Identify and close observability and alerting gaps.
  • Support service dependency mapping.
  • Carry out resilience and failover testing.
  • Implement or improve rate limiting and circuit breakers.
  • Contribute to disaster recovery evidence.
  • Support cyber recovery exercises with Engineering Leads and SOC.

Requirements

  • Platform engineering or SRE background with a resilience focus.
  • Experience with observability tooling (e.g. Azure App Insights, Datadog, Dynatrace, Prometheus/Grafana, Splunk, New Relic).
  • Hands-on with backup, restore, and disaster recovery processes in cloud environments.
  • Familiarity with operational resilience frameworks (DORA, PRA/FCA requirements desirable).
  • Experience writing and testing recovery runbooks.
  • Comfortable working across multiple platform teams with varying technology stacks.
  • Financial services or other regulated industry experience desirable.

About the company

FT Select is a niche technology search and selection company, specialising in recruiting elite Sales Executives and high-end Technology specialists for high-growth tech businesses.______________ About the Company A well-known leading global financial services organisation.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on ftselect.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:07 min

Architecting the availability stack with Prometheus and Grafana

Gabriel Labachelerie · World Congress 2023

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel · World Congress 2025

Videos

See all

Related articles

See all