AVP Site Reliability Engineer - SRE/Infrastructure/Python/Powershell/AWS/Observability/ITIL - PERM
Scope AT
Manor Park, United Kingdom
2 days ago
Role details
Contract type
Permanent contract Employment type
Full-time (> 32 hours) Working hours
Regular working hours Languages
English Experience level
IntermediateJob location
Manor Park, United Kingdom
Tech stack
Amazon Web Services (AWS)
Cloud Computing
Python
Powershell
Reliability Engineering
Ansible
Datadog
Scripting (Bash/Python/Go/Ruby)
Grafana
Production Code
Terraform
Dynatrace
Job description
The role is primarily responsible for developing SRE methodologies and ensuring they are applied to the Cloud hosted environment. In addition, the role will act as a central point of expertise for SRE automation across the Platform Operations team., * Responsible for driving the implementation of SRE methodologies, collaborating closely with other infrastructure teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence.
- Drives continuous improvement in system observability, alerting, and capacity planning through the definition and implementation of SLA, SLOs & SLIs
- Define and enhance frameworks for Toil identification, analysis & remediation to identify opportunities to eliminate or automate remediation of recurring tasks and issues
- Develops secure high-quality production code, and reviews and debugs code written by others.
- Build out and enhance GitOps capabilities for use in the Cloud hosted environments using tools such as Terraform and Ansible Automation Platform
- Provide on-call support and escalation for Cloud & Automation related issues ensuring that Production stability is the primary requirement.
- Ensure risks and stability issues in the cloud hosted environment are understood and addressed where possible through SRE best practices as part of any incident postmortems.
Requirements
Minimum Job-Related Experience Required:
- Must have strong technical operational support experience within an infrastructure services team performing on-call duties such as handling tickets, owning incidents & investigating their root cause
- Minimum of 2 years experience applying SRE methodologies within a support team and an understanding of Service Level metrics associated with this.
- Strong knowledge of at least 1 Scripting language, preferably either Python or Ansible. PowerShell would also be a positive
- Experience with supporting and building multi environment, multi region platforms with cloud providers such as AWS/GCP and managing them through Infrastructure as Code and GitOps methodologies
- Experience of Observability/APM tools (eg Grafana/Datadog/Dynatrace).