SRE

Technopride Ltd
Brighton and Hove, UK
11 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Distributed Systems Information Technology Operations Python (Programming Language) Reliability Engineering Ansible Workflow Management Systems Datadog Scripting Mttr
+5 more
Reliability of Systems Containerization Dynatrace Docker Microservices

Job description

Role: SRE

Location: Hove, UK

Is it Permanent / Contract: Open for both Perm/Contract

Is it Onsite/Remote/Hybrid: 2 days per week from office

No. of Positions: 1

We are seeking an experienced Site Reliability Engineer (SRE) to drive the modernization of IT operations through the implementation of observability practices, automation, and reliability engineering principles. The role requires a strategic thinker with strong hands-on expertise who can enhance system reliability, scalability, and operational efficiency while reducing manual operational tasks.

The successful candidate will work closely with engineering, architecture, and product teams to implement modern reliability practices, automate operational workflows, and establish robust monitoring and incident management frameworks.

Key Responsibilities

  • Collaborate with engineering teams to modernize IT operations by improving observability, automation, and operational efficiency.
  • Design and implement observability platforms to effectively monitor system health, performance, and reliability.
  • Develop strategies for AI-driven alerting and proactive anomaly detection to reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).
  • Establish and enforce SRE best practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
  • Define and implement an AIOps roadmap to enhance operational intelligence and automation.
  • Automate repetitive operational tasks (toil reduction) using scripting, orchestration tools, and automation frameworks.
  • Implement self-healing systems and automated incident response mechanisms to support autonomous operations.
  • Collaborate with cross-functional teams to ensure systems are scalable, resilient, and maintainable.
  • Lead incident management, root cause analysis, and post-incident improvement initiatives.
  • Promote shift-left reliability practices across engineering and product teams.
  • Mentor team members and advocate for a culture of reliability, automation, and continuous improvement.

Required Skills & Experience

  • Strong expertise in Site Reliability Engineering (SRE) principles and practices.
  • Hands-on experience implementing observability solutions, particularly with Dynatrace and Datadog.
  • Strong scripting and automation experience using Python and Ansible.
  • Experience working with cloud platforms such as AWS and Azure.
  • Solid understanding of containerization and orchestration technologies, including Docker and Kubernetes.
  • Experience working with cloud-native distributed systems and microservices architectures.

Requirements

  • Strong expertise in Site Reliability Engineering (SRE) principles and practices.
  • Hands-on experience implementing observability solutions, particularly with Dynatrace and Datadog.
  • Strong scripting and automation experience using Python and Ansible.
  • Experience working with cloud platforms such as AWS and Azure.
  • Solid understanding of containerization and orchestration technologies, including Docker and Kubernetes.
  • Experience working with cloud-native distributed systems and microservices architectures.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Advocating for SRE practices within agency environments

Martin Beránek · LIVE

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · World Congress 2025

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all