> Markdown version of [/jobs/ext/2066792-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2066792-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Darktrace Ltd - **Location:** Cambridge, UK - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Cloud Computing, Computer Programming, Databases, Data Infrastructure, DevOps, Monitoring of Systems, Python (Programming Language), Key Management, Open Source Technology, Reliability Engineering, Data Streaming, Systems Integration, Data Logging, Google Cloud, Reliability of Systems, Kubernetes, Dynatrace, Devsecops - **Published:** August 15, 2026 - **Apply:** https://www.totaljobs.com/job/site-reliability-engineer/darktrace-job107848199 ## About the Role * Proven experience in Site Reliability Engineering, DevOps, or infrastructure engineering * Deep expertise in at least one of the following areas: + Observability & monitoring (metrics, logging, distributed tracing) + Performance engineering & capacity planning + Data infrastructure reliability (databases, streaming, pipelines) + Security-focused SRE (hardening, compliance automation, secrets management) + Network reliability & traffic management * Strong programming skills (e.g. Go, Python, or similar) * Experience with cloud platforms (AWS, GCP, Azure) and Kubernetes * Strong communication skills, with the ability to explain complex technical concepts clearly * Self-driven with the ability to identify and prioritise high-impact work independently Desirable * Experience building internal developer platforms or tooling * Contributions to open-source, technical blogs, or public speaking * Experience working in regulated environments * Familiarity with SLO frameworks and error budget management * Relevant certifications in your specialist domain Success Measures * Improved reliability and performance within your domain of specialism * Adoption of best practices across SRE, Platform Engineering, and DevSecOps * Reduction in incidents and faster resolution times * Scalable, well-integrated solutions within the internal platform * Strong collaboration across teams and measurable improvements in operational maturity ## Description We're looking for a Site Reliability Engineer (SRE) to bring deep expertise in a key reliability domain and help shape the future of our platform reliability strategy. SRE sits at the heart of our operational trifecta alongside Platform Engineering and DevSecOps. In this role, you'll act as the go-to authority in your area of specialism, working across teams to embed best practices, solve complex reliability challenges, and improve system resilience at scale., Domain Expertise & Strategy * Act as the subject matter expert in your chosen reliability domain * Define and implement standards, frameworks, and best practices across SRE, Platform Engineering, and DevSecOps * Stay current with industry trends and bring innovative ideas into the organisation Engineering & Delivery * Design and implement solutions to complex, cross-cutting reliability challenges * Build tooling, automation, and frameworks to improve system resilience and scalability * Lead deep-dive investigations into systemic issues and drive long-term fixes Collaboration & Platform Integration * Partner with Platform Engineering to ensure your domain is embedded within the internal developer platform * Collaborate with DevSecOps to integrate security, compliance, and resilience practices * Contribute to cross-team initiatives that improve reliability across the stack Incident & Operational Excellence * Play a key role in incident response, particularly within your specialism * Contribute to on-call rotations and continuous improvement of operational processes * Develop runbooks, documentation, and training materials to support teams ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [DevSecOps: Injecting Security into Mobile CI/CD Pipelines](https://www.wearedevelopers.com/videos/273-devsecops-injecting-security-into-mobile-ci-cd-pipelines) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [DevSecOps culture](https://www.wearedevelopers.com/videos/783-devsecops-culture) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [7 Most Popular Web Developer Jobs in Europe](https://www.wearedevelopers.com/magazine/163-7-most-popular-web-developer-jobs-in-europe)