Lead Site Reliability Engineer

Motive Partners
Greater London, UK
1 day ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Amazon Web Services Application Integration Architecture Microsoft Azure Bash Shell Cloud Computing Computer Networks Networking Basics Reliability Engineering Newrelic Software Deployment Systems Integration Load Balancing
+6 more
System Availability Delivery Pipeline Grafana Kubernetes Terraform Splunk

Job description

Role Purpose

The Site Reliability Engineer will work closely with Application, Infrastructure, and Network Engineering teams to ensure the reliability, scalability, and performance of FNZ platforms. This role focuses on deploying, integrating, and providing ongoing operational support for mission-critical systems, leveraging modern automation and cloud-native practices.

Key Responsibilities

  • Maintain high availability and performance of FNZ platforms.
  • Implement monitoring, alerting, and observability solutions to proactively detect and resolve issues.
  • Collaborate with engineering teams to design and implement robust deployment pipelines.
  • Ensure smooth integration of applications with infrastructure and network components.
  • Use Terraform for provisioning and managing infrastructure across environments.
  • Operate and optimise workloads on-prem and public cloud.
  • Manage and troubleshoot application delivery networks, load balancing, and traffic routing.
  • Configure and support F5 Distributed Cloud or similar CDN/ADC technologies.
  • Participate in on-call rotations, perform root cause analysis, and implement preventive measures.
  • Work cross-functionally with Application, Infrastructure, and Network Engineering teams to deliver reliable services.

Required Skills & Experience

  • Kubernetes (K8s): Deep understanding of container orchestration and cluster management.
  • Terraform: Strong experience in Infrastructure as Code for cloud and on-prem environments.
  • Public Cloud: Hands-on experience with AWS, Azure, or GCP.
  • F5 Distributed Cloud or Similar: Knowledge of CDN/ADC platforms and their integration.
  • Networking Fundamentals: Expertise in application delivery networks, load balancing, traffic routing, and troubleshooting.
  • Observability Tools: Familiarity with Splunk, NewRelic, or similar.
  • Scripting & Automation: Proficiency in Terraform, Bash, or similar languages.

Desirable Skills

  • Experience with CI/CD pipelines and GitOps workflows.
  • Knowledge of SRE principles.
  • Familiarity with security best practices.

Key Attributes

  • Strong problem-solving and troubleshooting skills.
  • Ability to work collaboratively across multiple teams.
  • Passion for automation and reducing operational toil.

Reporting Line

Reports to: Head of Platform Operations/Application Engineering.

Works closely with Application Engineering, Infrastructure Engineering, Network Engineering teams.

Requirements

  • Kubernetes (K8s): Deep understanding of container orchestration and cluster management.
  • Terraform: Strong experience in Infrastructure as Code for cloud and on-prem environments.
  • Public Cloud: Hands-on experience with AWS, Azure, or GCP.
  • F5 Distributed Cloud or Similar: Knowledge of CDN/ADC platforms and their integration.
  • Networking Fundamentals: Expertise in application delivery networks, load balancing, traffic routing, and troubleshooting.
  • Observability Tools: Familiarity with Splunk, NewRelic, or similar.
  • Scripting & Automation: Proficiency in Terraform, Bash, or similar languages.

Desirable Skills

  • Experience with CI/CD pipelines and GitOps workflows.
  • Knowledge of SRE principles.
  • Familiarity with security best practices., * Strong problem-solving and troubleshooting skills.
  • Ability to work collaboratively across multiple teams.
  • Passion for automation and reducing operational toil.

Reporting Line

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

Videos

See all

Related articles

See all