Site Reliability Engineer (AWS, Terraform, Ansible, Python, GitOps, SRE, Grafana/Datadog) - Permanen

Scope AT
Charing Cross, United Kingdom
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English

Job location

Charing Cross, United Kingdom

Tech stack

Amazon Web Services (AWS)
Cloud Computing
Python
Powershell
Reliability Engineering
Ansible
Software Engineering
Datadog
Scripting (Bash/Python/Go/Ruby)
Grafana
Git Flow
Terraform
Dynatrace

Job description

Site Reliability Engineer (AWS, Terraform, Ansible, Python, GitOps, SRE, Grafana/Datadog) - Permanent - London, As a Site Reliability Engineer, you will be responsible for implementing and evolving SRE methodologies across cloud infrastructure while helping to automate operational processes and reduce manual intervention. You'll work closely with infrastructure and engineering teams to improve platform reliability and operational excellence., * Drive the adoption and implementation of Site Reliability Engineering (SRE) principles across cloud-hosted platforms.

  • Improve platform reliability through automation, observability, and continuous improvement initiatives.
  • Define and implement SLAs, SLOs and SLIs to improve service performance and availability.
  • Identify operational toil and automate repetitive tasks using Infrastructure as Code and automation tools.
  • Develop, review and troubleshoot production automation and infrastructure code.
  • Enhance GitOps capabilities using Terraform and Ansible Automation Platform.
  • Support multi-region, cloud-based environments.
  • Participate in an on-call rota, managing production incidents and ensuring platform stability.
  • Perform root cause analysis and drive preventative improvements following incidents.
  • Collaborate with infrastructure, engineering and operational teams to improve deployment processes and cloud operations.

Requirements

  • Previous experience in a Site Reliability Engineering (SRE) or Infrastructure Operations role.
  • Strong operational support experience, including incident management, root cause analysis and on-call support.
  • Experience implementing SRE methodologies within enterprise environments.
  • Strong Scripting skills using Python, Ansible or PowerShell.
  • Hands-on experience with AWS and/or GCP cloud platforms.
  • Experience with Terraform, Infrastructure as Code and GitOps practices.
  • Experience with observability and monitoring platforms such as Grafana, Datadog or Dynatrace.
  • Strong troubleshooting and analytical skills.
  • Excellent communication skills with both technical and business stakeholders.
  • Experience working within regulated financial services or banking environments is highly desirable.

Desirable Experience

  • Software development background.
  • Experience with Ansible Automation Platform.
  • Knowledge of ITIL.
  • AWS or Terraform certifications.

Benefits & conditions

  • Permanent opportunity.
  • London based with 2 days per week onsite.
  • Opportunity to work on highly available, enterprise-scale cloud platforms.
  • Exposure to modern cloud technologies, automation and SRE best practices.
  • Collaborative engineering culture focused on innovation, reliability and continuous improvement.

Apply for this position