Site Reliability Engineer

Paritas Recruitment
Glasgow, UK
1 day ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Bash Shell Continuous Integration Data as a Services Information Engineering Disaster Recovery Identity and Access Management Python (Programming Language) Reliability Engineering Scripting
+10 more
Snowflake Infrastructure as Code (IaC) Amazon Virtual Private Cloud (VPC) Cloudformation Infrastructure Automation Frameworks Data Management Cloudwatch Terraform Data Pipelines Databricks

Job description

We are recruiting an AWS Site Reliability Engineer (SRE) to support a cloud-native data platform for a major international financial services organisation. The platform is built on AWS, with core components including Snowflake and Databricks, and underpins critical analytics and data services used across the business.

This role focuses on reliability engineering, automation, observability, and resilience. You will work closely with data engineering and platform teams to ensure the platform is scalable, highly available, and operationally robust in a regulated, high-availability environment., * Design, build, and maintain automation for infrastructure provisioning, platform operations, and incident response using Infrastructure as Code (IaC) and CI/CD

  • Lead resiliency and disaster recovery (DR) planning, including DR testing, failure scenarios, and recovery validation across AWS and data platform services
  • Define and manage SLIs, SLOs, and SLAs for critical data pipelines and platform services, using error budgets to drive reliability improvements
  • Build and operate comprehensive observability solutions (metrics, logs, traces, alerting) across AWS, Snowflake, and Databricks workloads
  • Partner with data engineering and platform teams to embed reliability-by-design into architecture and delivery
  • Perform root cause analysis (RCA) on incidents and drive continuous improvement to reduce operational toil
  • Own and drive resolution of incidents and service requests raised by platform consumers, identifying recurring issues and automating fixes to improve reliability and user experience

Requirements

  • Strong practical experience applying Site Reliability Engineering (SRE) principles, including SLO/SLI/SLA design and error budgets
  • Proven production experience with AWS (e.g. EC2, S3, IAM, VPC, CloudWatch)
  • Hands-on experience with automation and Infrastructure as Code (Terraform, CloudFormation, or CDK)
  • Experience building and operating observability and monitoring solutions
  • Scripting experience in Python and/or Bash
  • Exposure to data platforms such as Snowflake and/or Databricks

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

1:33 min

Integrating internal APIs and maintaining data sovereignty

Mahran Meißner Mahran Meißner · World Congress 2026 Europe

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:32 min

Overview of Terraform and Terraform Cloud features

Devlin Duldulao · LIVE

Videos

See all

Related articles

See all