Senior Site Reliability Engineer

Tenth Revolution Group
Knutsford, UK
13 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Systems Engineering Cloud Computing Computer Programming Python (Programming Language) Performance Tuning Windows PowerShell Reliability Engineering Site Reliability Engineering Practices Software Engineering Scripting Grafana Reliability of Systems
+1 more
Performance Monitor

Job description

Senior Site Reliability Engineer - Knutsford (hybrid, 2 days per week in office). A leading Financial Services firm is recruiting for a Senior Site Reliability Engineer to become part of a newly formed Core SRE Team that will establish a Centre of Excellence to enhance and promote SRE best practices., As a key hire, you will raise awareness and drive adoption of SRE methodologies within various teams. This is a hands-on engineering role where you will design, build, and optimise automation frameworks, observability tools, and incident response mechanisms. You will act as a trusted advisor, providing strategic guidance and consultative support to help teams improve reliability, scalability, and efficiency., * Availability, performance, and scalability of systems and services through proactive monitoring, maintenance, and capacity planning.

  • Resolution, analysis and response to system outages and disruptions, and implementation of measures to prevent similar incidents from recurring.
  • Development of tools and scripts to automate operational processes, reducing manual workload, increasing efficiency, and improving system resilience.
  • Monitoring and optimisation of system performance and resource usage, identifying bottlenecks, and implementing best practices for performance tuning.
  • Collaboration with development teams to integrate best practices for reliability, scalability, and performance into the software development lifecycle, and work closely with other teams to ensure smooth and efficient operations.

Requirements

  • Proficiency in Programming and Scripting - languages such as Python, Powershell, or Go for automating routine tasks and system deployments.
  • Incident Management and Troubleshooting - ability to manage incidents effectively, troubleshoot issues swiftly, and perform root cause analysis to prevent future incidents.
  • Systems Engineering and Automation - understanding of operating systems, networking, and cloud infrastructure; proficiency in automation tools for maintaining system reliability at scale.
  • Influential Communication Skills - ability to communicate effectively with team members and stakeholders to drive alignment and foster a collaborative environment for SRE practices.
  • Knowledge of Cloud Computing - familiarity with cloud platforms and services as infrastructure moves to the cloud.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

2:27 min

Introduction to WebAssembly in a cloud computing context

Edo Edo · World Congress 2024

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

Videos

See all

Related articles

See all