Senior Site Reliability Engineer

VIQU Ltd
Wavendon, UK
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
£65,000.0 - £75,000.0
Working hours
Regular working hours
Job source

Tech stack

Microsoft Windows Artificial Intelligence Application Performance Management Microsoft Azure Software as a Service Cloud Computing Linux DevOps Linux System Administration Reliability Engineering Prometheus Virtual Machines
+5 more
Datadog Grafana Mttr Kubernetes Terraform

Job description

VIQU have partnered with a well-established B2B SaaS company who are going through a significant platform transformation. and so are hiring for a Senior Site Reliability Engineer to build stability, respond to live incidents, and assist with system upkeep. The role will also play a key part in on implementing and adopting new tooling and processes surrounding the wider transformation., * Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles.

  • Regularly use Datadog and other observability tools for application performance monitoring.
  • Implement new ways of working, helping to shape how the organisation responds and recovers to incidents.
  • Take ownership of incident resolutions.

  • Actively drive down key reliability metrics (MTTR, incident frequency, on-call toil) by evaluating key incidents.
  • Work on an a on call rota, ensuring you are available to respond to incidents during this time.

  • Identify areas for automation and help implement changes that raise the bar for reliability.

Requirements

  • Previous experience as a Site Reliability Engineer or similar (cloud, infrastructure, DevOps or platform engineering) within a customer facing environment e.g SaaS or MSP.
  • Strong hands-on experience with both Azure, and on-premise virtual machines.

  • Experience withInfrastructure as Code / Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor).
  • Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways of working.
  • Ability to communicate across internal teams and external customers.
  • Skilled in networking across both cloud (Azure) and on premise environments.
  • Either Windows or Linux systems administration skills (Linux preferred).
  • Previous use of AI tools to enhance efficiency.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on find.jobs

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all