Site Reliability Engineer - Terraform, Ansible, Python, AWS, GCP, GitOps, Infrastructure as Code
Scope AT
Charing Cross, United Kingdom
yesterday
Role details
Contract type
Permanent contract Employment type
Full-time (> 32 hours) Working hours
Regular working hours Languages
EnglishJob location
Charing Cross, United Kingdom
Tech stack
Amazon Web Services (AWS)
Cloud Computing
Python
Powershell
Reliability Engineering
Cloud Services
Ansible
Datadog
Scripting (Bash/Python/Go/Ruby)
Google Cloud Platform
Grafana
Git Flow
Infrastructure Automation Frameworks
Terraform
Dynatrace
Job description
Join our Platform Engineering team to improve the reliability, automation, and performance of cloud-based services. You'll apply Site Reliability Engineering (SRE) practices to enhance observability, automate operations, and support highly available production platforms., * Drive SRE best practices across cloud infrastructure and platform operations.
- Improve monitoring, alerting, SLAs/SLOs/SLIs, and platform observability.
- Automate operational tasks using Infrastructure as Code and GitOps.
- Develop and maintain automation using Terraform, Ansible, and Scripting.
- Support production systems, participate in an on-call rota, and lead incident resolution and continuous improvement.
Requirements
- Experience supporting cloud infrastructure in a production environment.
- Knowledge of SRE principles, incident management, and root cause analysis.
- Strong Scripting skills (Python, Ansible, or PowerShell).
- Experience with AWS, GCP, or similar cloud platforms, Infrastructure as Code, and GitOps.
- Familiarity with observability tools such as Grafana, Datadog, or Dynatrace.
- Strong problem-solving, communication, and stakeholder management skills.