Site Reliability Engineer - Terraform, Ansible, Python, AWS, GCP, GitOps, Infrastructure as Code

Scope AT
Charing Cross, United Kingdom
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English

Job location

Charing Cross, United Kingdom

Tech stack

Amazon Web Services (AWS)
Cloud Computing
Python
Powershell
Reliability Engineering
Cloud Services
Ansible
Datadog
Scripting (Bash/Python/Go/Ruby)
Google Cloud Platform
Grafana
Git Flow
Infrastructure Automation Frameworks
Terraform
Dynatrace

Job description

Join our Platform Engineering team to improve the reliability, automation, and performance of cloud-based services. You'll apply Site Reliability Engineering (SRE) practices to enhance observability, automate operations, and support highly available production platforms., * Drive SRE best practices across cloud infrastructure and platform operations.

  • Improve monitoring, alerting, SLAs/SLOs/SLIs, and platform observability.
  • Automate operational tasks using Infrastructure as Code and GitOps.
  • Develop and maintain automation using Terraform, Ansible, and Scripting.
  • Support production systems, participate in an on-call rota, and lead incident resolution and continuous improvement.

Requirements

  • Experience supporting cloud infrastructure in a production environment.
  • Knowledge of SRE principles, incident management, and root cause analysis.
  • Strong Scripting skills (Python, Ansible, or PowerShell).
  • Experience with AWS, GCP, or similar cloud platforms, Infrastructure as Code, and GitOps.
  • Familiarity with observability tools such as Grafana, Datadog, or Dynatrace.
  • Strong problem-solving, communication, and stakeholder management skills.

Apply for this position