Senior Site Reliability Engineer - 12 months - London - £425/day - Inside IR35

Hamilton Barnes
yesterday

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
£ 111K

Job location

Tech stack

Amazon Web Services (AWS)
Amazon Web Services (AWS)
Amazon Web Services (AWS)
Bash
Cloud Computing
Configuration Management
Continuous Integration
Linux
DevOps
Distributed Systems
Identity and Access Management
Python
Reliability Engineering
Cloud Services
Datadog
Multi-Cloud
HybridCloud
Infrastructure as Code (IaC)
Amazon Web Services (AWS)
Gitlab
Amazon Web Services (AWS)
Containerization
Gitlab-ci
Deployment Automation
Functional Programming
Cloudwatch
Terraform
Docker
Jenkins
Microservices

Job description

  • Architect, implement, and maintain observability platforms using Datadog and Geneos to ensure comprehensive monitoring and alerting
  • Design and manage scalable infrastructure using Terraform and Infrastructure as Code (IaC) principles
  • Champion GitOps methodologies using GitLab for CI/CD, configuration management, and deployment automation
  • Optimise alerting strategies to reduce noise and improve actionable insights
  • Oversee and continuously optimise cloud cost management strategies for observability infrastructure in line with Cloud FinOps principles
  • Lead incident response, perform root cause analysis, and drive continuous improvement through blameless post-mortems
  • Collaborate with software developers across multiple geographies and cross-functional teams to deliver systems within agreed timelines
  • Collaborate with executive leadership and cross-functional stakeholders to align infrastructure strategy with long-term business objectives
  • Mentor junior engineers and contribute to the evolution of SRE best practices
  • Stay updated on industry best practices and emerging technologies in observability, DevOps, and cloud

Requirements

We are seeking a highly experienced and hands-on Senior Site Reliability Engineer to join a global technology services organisation on a 12-month hybrid contract based in London City. There are 2 positions available. The successful candidate will bring a strong background in observability, automation, cloud infrastructure, and DevOps practices on AWS, driving reliability and performance across complex distributed systems whilst designing and optimising a best-in-class monitoring infrastructure using Datadog and Geneos. Experience in financial services, capital markets, or fintech environments is highly advantageous., * 7+ years of experience in Site Reliability Engineering, DevOps, or related roles (essential)

  • Strong hands-on experience with AWS cloud services including EC2, S3, RDS, Lambda, VPC, IAM, CloudWatch, EKS, and ECS (essential)
  • Strong hands-on expertise in Infrastructure as Code using Terraform (essential)
  • Deep expertise in Datadog for metrics, logs, traces, and dashboards (essential)
  • Proficiency with Geneos for Real Time monitoring and alerting (essential)
  • Deep knowledge of CI/CD tools including GitLab CI and Jenkins for Java and Python-based microservices architectures
  • Solid understanding of GitLab and GitOps workflows
  • Strong Scripting and automation skills in Python and Bash
  • Experience with containerisation and orchestration using Docker and Kubernetes
  • Solid understanding of Linux systems, networking, and distributed systems
  • Desirable: AWS certifications, familiarity with SLOs/SLIs and error budgets, experience in multi-cloud or hybrid cloud environments, knowledge of ITIL or SRE frameworks, and experience working in regulated or high-availability environments such as fintech or capital markets
  • Bonus: experience in Equity or Fixed Income and working knowledge of Benchmarks and Indices

Apply for this position