Site Reliability Engineer

Insight
Bournemouth, UK
1 day ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Java (Programming Language) Amazon Web Services Data Analysis Microsoft Azure Cloud Computing Information Engineering DevOps Github Reliability Engineering Prometheus Data Logging Transaction Processing (Computing)
+10 more
Google Cloud Cloud Platform System Grafana Spring-boot Kubernetes Data Analytics Cloud Migration Terraform Splunk Jenkins

Job description

  • Reliability & Scalability: Design, implement, and maintain systems that are robust, scalable, and highly available, supporting millions of daily transactions.
  • Cloud Migration: Lead and support migration of applications and infrastructure to public cloud platforms, ensuring best practices in security, reliability, and cost management.
  • Automation & Infrastructure as Code: Develop and maintain automation scripts and infrastructure using Kubernetes and Terraform.
  • Monitoring & Incident Response: Build and enhance monitoring, alerting, and observability solutions. Respond to incidents, perform root cause analysis, and drive continuous improvement.
  • Collaboration: Partner with software engineers, product managers, and business stakeholders to deliver solutions that meet business needs and operational requirements.
  • Analytics & Data Insights: Leverage cloud-based analytics tools to monitor system health, optimize performance, and extract actionable insights.
  • Continuous Improvement: Identify and implement opportunities to improve reliability, efficiency, and scalability of the platform.

Requirements

  • Proven experience as a Site Reliability Engineer, DevOps Engineer, or similar role supporting large-scale, mission-critical systems.
  • Strong hands-on experience with Kubernetes and Terraform.
  • Experience deploying and operating applications in public cloud environments (AWS, Azure, GCP).
  • Solid understanding of Java and Spring Boot applications.
  • Experience with monitoring, logging, and observability tools (Prometheus, Grafana, ELK, Splunk).
  • Strong troubleshooting and problem-solving skills.
  • Excellent communication and collaboration skills., * Experience in financial services or payments/transaction processing environments.
  • Familiarity with cloud-based analytics platforms and data engineering concepts.
  • Experience with CI/CD pipelines and automation tools (Jenkins, GitHub Actions).
  • Knowledge of security best practices in cloud environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

Videos

See all

Related articles

See all