AWS Site Reliability Engineer

Marks Sattin Ltd
Glasgow, UK
1 day ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Bash Shell Big Data Continuous Integration Data as a Services Disaster Recovery Identity and Access Management Python (Programming Language) Reliability Engineering Scripting
+9 more
Snowflake Grafana Amazon Virtual Private Cloud (VPC) Cloudformation Infrastructure Automation Frameworks Data Management Cloudwatch Terraform Databricks

Job description

An established technology-driven organisation is seeking an experienced Site Reliability Engineer (SRE) in Glasgow to strengthen and scale their cloud-native data platform, utilising AWS, Snowflake, and Databricks. This position offers the opportunity to drive automation, resilience, and operational excellence across critical data services., * Automate infrastructure provisioning and platform operations using Infrastructure as Code and CI/CD tools.

  • Lead and execute reliability initiatives including disaster recovery planning, failure testing, and resilience validation.
  • Define and manage service health metrics (SLIs/SLOs/SLAs) to drive measurable improvements in reliability.
  • Build observability solutions to monitor AWS, Snowflake, and Databricks workloads.
  • Collaborate with engineering teams to embed reliability best practices throughout platform development.
  • Analyse incidents and proactively address root causes to improve availability and performance.
  • Provide operational support, drive incident resolution, and implement automated fixes for recurring issues.

Requirements

  • Strong knowledge of SRE principles and practical experience defining SLAs, SLOs, and error budgets.
  • Demonstrated AWS expertise (e.g., EC2, S3, IAM, VPC, CloudWatch) in production environments.
  • Experience with observability tools, monitoring, and alerting practices.
  • Proficient in automation, Infrastructure as Code (Terraform, CloudFormation, or CDK), and scripting (Python/Bash).
  • Exposure to Snowflake and/or Databricks data platforms.
  • Background in DR/chaos engineering, CI/CD pipelines, GitOps, or supporting large-scale data environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

4:23 min

Reviewing AWS infrastructure deployment configuration and planning

Devlin Duldulao · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

1:24 min

Evaluating formal AWS certifications versus raw practical engineering experience

Jan Giacomelli · LIVE

Videos

See all

Related articles

See all