Site Reliability Engineer

Darktrace Ltd
Cambridge, UK
1 day ago
Apply on www.totaljobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Cloud Computing Computer Programming Databases Data Infrastructure DevOps Monitoring of Systems Python (Programming Language) Key Management Open Source Technology Reliability Engineering
+8 more
Data Streaming Systems Integration Data Logging Google Cloud Reliability of Systems Kubernetes Dynatrace Devsecops

Job description

We’re looking for a Site Reliability Engineer (SRE) to bring deep expertise in a key reliability domain and help shape the future of our platform reliability strategy.

SRE sits at the heart of our operational trifecta alongside Platform Engineering and DevSecOps. In this role, you’ll act as the go-to authority in your area of specialism, working across teams to embed best practices, solve complex reliability challenges, and improve system resilience at scale., Domain Expertise & Strategy

  • Act as the subject matter expert in your chosen reliability domain
  • Define and implement standards, frameworks, and best practices across SRE, Platform Engineering, and DevSecOps
  • Stay current with industry trends and bring innovative ideas into the organisation

Engineering & Delivery

  • Design and implement solutions to complex, cross-cutting reliability challenges
  • Build tooling, automation, and frameworks to improve system resilience and scalability
  • Lead deep-dive investigations into systemic issues and drive long-term fixes

Collaboration & Platform Integration

  • Partner with Platform Engineering to ensure your domain is embedded within the internal developer platform
  • Collaborate with DevSecOps to integrate security, compliance, and resilience practices
  • Contribute to cross-team initiatives that improve reliability across the stack

Incident & Operational Excellence

  • Play a key role in incident response, particularly within your specialism
  • Contribute to on-call rotations and continuous improvement of operational processes
  • Develop runbooks, documentation, and training materials to support teams

Requirements

  • Proven experience in Site Reliability Engineering, DevOps, or infrastructure engineering
  • Deep expertise in at least one of the following areas:
  • Observability & monitoring (metrics, logging, distributed tracing)
  • Performance engineering & capacity planning
  • Data infrastructure reliability (databases, streaming, pipelines)
  • Security-focused SRE (hardening, compliance automation, secrets management)
  • Network reliability & traffic management
  • Strong programming skills (e.g. Go, Python, or similar)
  • Experience with cloud platforms (AWS, GCP, Azure) and Kubernetes
  • Strong communication skills, with the ability to explain complex technical concepts clearly
  • Self-driven with the ability to identify and prioritise high-impact work independently

Desirable

  • Experience building internal developer platforms or tooling
  • Contributions to open-source, technical blogs, or public speaking
  • Experience working in regulated environments
  • Familiarity with SLO frameworks and error budget management
  • Relevant certifications in your specialist domain

Success Measures

  • Improved reliability and performance within your domain of specialism
  • Adoption of best practices across SRE, Platform Engineering, and DevSecOps
  • Reduction in incidents and faster resolution times
  • Scalable, well-integrated solutions within the internal platform
  • Strong collaboration across teams and measurable improvements in operational maturity

Benefits & conditions

  • 23 days’ holiday + all public holidays, rising to 25 days after 2 years of service,
  • Additional day off for your birthday,
  • Private medical insurance which covers you, your cohabiting partner and children,
  • Life insurance of 4 times your base salary,
  • Salary sacrifice pension scheme,
  • Enhanced family leave,
  • Confidential Employee Assistance Program,
  • Cycle to work scheme.

About the company

Darktrace is a global leader in AI for cybersecurity that keeps organizations ahead of the changing threat landscape every day. Founded in 2013, Darktrace provides the essential cybersecurity platform protecting nearly 10,000 organizations from unknown threats using its proprietary AI., The Darktrace Active AI Security Platform delivers a proactive approach to cyber resilience to secure the business across the entire digital estate - from network to cloud to email. Breakthrough innovations from our R&D teams have resulted in over 200 patent applications filed. Darktrace’s platform and services are supported by over 2,400 employees around the world. To learn more, visit http://www.darktrace.com.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.totaljobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · World Congress 2025

1:15 min

Key lessons learned from implementing automated mobile DevSecOps

Moataz Nabil Moataz Nabil · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

12:08 min

Comparing Keptn orchestration capabilities against alternative software operators

Thomas Schütz · LIVE

Videos

See all

Related articles

See all