AVP Site Reliability Engineer - SRE/Infrastructure/Python/Powershell/AWS/Observability/ITIL - PERM

Scope AT
Charing Cross, United Kingdom
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Intermediate

Job location

Charing Cross, United Kingdom

Tech stack

Amazon Web Services (AWS)
Cloud Computing
Python
Powershell
Reliability Engineering
Ansible
Datadog
Scripting (Bash/Python/Go/Ruby)
Grafana
Production Code
Terraform
Dynatrace

Job description

AVP Site Reliability Engineer - SRE/Infrastructure/Python/Powershell/AWS/Observability/ITIL - PERM - Financial Services

Job Purpose: The role is primarily responsible for developing SRE methodologies and ensuring they are applied to the Cloud hosted environment. In addition, the role will act as a central point of expertise for SRE automation across the Platform Operations team.

Essential Job Functions:

  • Responsible for driving the implementation of SRE methodologies, collaborating closely with other infrastructure teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence.
  • Drives continuous improvement in system observability, alerting, and capacity planning through the definition and implementation of SLA, SLOs & SLIs
  • Define and enhance frameworks for Toil identification, analysis & remediation to identify opportunities to eliminate or automate remediation of recurring tasks and issues
  • Develops secure high-quality production code, and reviews and debugs code written by others.
  • Build out and enhance GitOps capabilities for use in the Cloud hosted environments using tools such as Terraform and Ansible Automation Platform
  • Provide on-call support and escalation for Cloud & Automation related issues ensuring that Production stability is the primary requirement.
  • Ensure risks and stability issues in the cloud hosted environment are understood and addressed where possible through SRE best practices as part of any incident postmortems.

Minimum Job-Related Experience Required:

  • Must have strong technical operational support experience within an infrastructure services team performing on-call duties such as handling tickets, owning incidents & investigating their root cause
  • Minimum of 2 years experience applying SRE methodologies within a support team and an understanding of Service Level metrics associated with this.
  • Strong knowledge of at least 1 Scripting language, preferably either Python or Ansible. PowerShell would also be a positive
  • Experience with supporting and building multi environment, multi region platforms with cloud providers such as AWS/GCP and managing them through Infrastructure as Code and GitOps methodologies
  • Experience of Observability/APM tools (eg Grafana/Datadog/Dynatrace).

Permanent Role based in Canary Wharf - Hybrid Working

By applying to this job you are sending us your CV, which may contain personal information. Please refer to our Privacy Notice to understand how we process this information. In short, in order to supply you with work finding services, we will hold and process your personal data, and only with your express permission we will share this personal data with a client (or a third party working on behalf of the client) by email or by upload to the Client/third parties vendor management system. By giving us permission to send your CV to a client, this constitutes permission to share the personal data that would be necessary to consider your application, interview you (Phone/video/face to face) and if successful hire you.

Scope AT acts as an employment agency for Permanent Recruitment and an employment business for the supply of temporary workers. By applying for this job you accept the Terms and Conditions, Data Protection Policy, Privacy Notice and Disclaimers which can be found at our website.

Requirements

  • Must have strong technical operational support experience within an infrastructure services team performing on-call duties such as handling tickets, owning incidents & investigating their root cause
  • Minimum of 2 years experience applying SRE methodologies within a support team and an understanding of Service Level metrics associated with this.
  • Strong knowledge of at least 1 Scripting language, preferably either Python or Ansible. PowerShell would also be a positive
  • Experience with supporting and building multi environment, multi region platforms with cloud providers such as AWS/GCP and managing them through Infrastructure as Code and GitOps methodologies
  • Experience of Observability/APM tools (eg Grafana/Datadog/Dynatrace).

Apply for this position