Site Reliability Engineer- London

FDM Group
London, UK
19 days ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Data Analysis Microsoft Azure Backup Devices Bash Shell DevOps Elasticsearch Monitoring of Systems Python (Programming Language) Performance Tuning Windows PowerShell Reliability Engineering
+7 more
Data Logging Scripting Grafana Database Optimization Containerization Kubernetes Docker

Job description

FDM is a global business and technology consultancy seeking a Site Reliability Engineer to work for our client within the Finance sector. This is initially a 6 month contract with very good prospects to extend and will be a hybrid role that will be based in London. Our client is seeking an experienced Site Reliability Engineer (SRE) with a strong focus on Observability and Monitoring Platforms. The successful candidate will play a key role in enhancing the organisation’s monitoring, alerting, and operational visibility capabilities across critical engineering systems. This role requires hands-on expertise in the deployment, administration, and optimisation of OpenSearch, alongside experience with Grafana, Geneos, and automation/scripting technologies. Particular emphasis will be placed on the candidate’s ability to design, deploy, and support enterprise-grade OpenSearch environments. Responsibilities:

  • Lead the design, deployment, configuration, and ongoing management of OpenSearch clusters and associated observability tooling.
  • Develop and maintain scalable monitoring, logging, and alerting solutions for business-critical applications and infrastructure.
  • Build and enhance observability dashboards using Grafana.
  • Support and optimise existing Geneos monitoring implementations.
  • Create and maintain automation scripts to streamline operational processes and improve reliability.
  • Collaborate with engineering, infrastructure, and support teams to improve system resilience and operational performance.
  • Define and implement SRE best practices, including monitoring standards, alert management, incident response, and operational readiness.
  • Perform troubleshooting and root cause analysis of platform and application issues.
  • Support capacity planning, performance tuning, and platform optimisation initiatives.
  • Contribute to documentation, operational procedures, and knowledge sharing within the engineering team.

Requirements

  • Extensive hands-on experience deploying and managing OpenSearch in production environments.
  • Deep understanding of OpenSearch architecture, cluster design, indexing strategies, shard management, and performance tuning.
  • Experience implementing log aggregation, search, analytics, and observability use cases using OpenSearch.
  • Knowledge of OpenSearch security, access controls, backups, upgrades, and operational best practices.

Monitoring & Observability

  • Strong experience with Grafana, including dashboard development, alerting, and data source integration.
  • Experience with enterprise monitoring platforms, specifically Geneos.
  • Understanding of modern observability principles, including metrics, logs, traces, alerting, and service health monitoring.

Scripting & Automation

  • Strong scripting skills in one or more of:
  • Python
  • Shell/Bash
  • PowerShell
  • Experience automating operational tasks and monitoring workflows.

SRE / Platform Engineering

  • Proven experience in an SRE, Platform Engineering, DevOps, or Infrastructure Engineering role.
  • Strong troubleshooting and problem-solving capabilities.
  • Experience supporting highly available and business-critical systems.
  • Understanding of incident management, resilience engineering, and operational excellence practices.

Desirable Skills

  • Experience with cloud platforms (Azure, AWS, or GCP).
  • Knowledge of containerisation technologies (Docker, Kubernetes).
  • Experience with CI/CD pipelines and Infrastructure as Code.
  • Experience working within financial services or regulated environments.
  • Familiarity with Elasticsearch ecosystems and migration strategies to OpenSearch., The ideal candidate will be a hands-on engineer who combines deep technical expertise with a pragmatic operational mindset. They will be comfortable working independently, driving observability improvements, and collaborating across engineering teams to deliver reliable and scalable monitoring solutions. Key attributes:

  • Strong ownership mentality.
  • Excellent analytical and troubleshooting skills.
  • Effective stakeholder communication.
  • Ability to operate in fast-paced production environments.
  • Focus on reliability, automation, and continuous improvements

Benefits & conditions

Pulled from the full job description

  • Annual leave
  • Company pension

About the company

FDM is an award-winning global leader in tech and business talent solutions, backed by more than 35 years of industry experience. We have centres across Europe, North America, and Asia-Pacific, and a global workforce of over 2500 employees. FDM has shown exponential growth throughout the years, firmly establishing itself as an award-winning employer, currently listed on the FTSE4Good Index and as a 2026 Financial Times UK ‘Best Employer’.

Diversity and Inclusion

FDM Group is an equal opportunity employer, and all qualified applicants will receive consideration for employment without regard to race, colour, religion, sex, sexual orientation, national origin, age, disability, veteran status or any other status protected by federal, provincial or local laws.

Why join us

  • Career coaching, mentoring and access to upskilling throughout your entire FDM career
  • Assignments with global companies and opportunities to work abroad
  • Opportunity to re-skill and up-skill into new areas, develop non-linear career paths and build a skillset within your field
  • Annual leave and work-place pension

You must create an Indeed account before continuing to the company website to apply

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on uk.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all