> Markdown version of [/jobs/ext/2389038-engineering-manager-site-reliability-eu-uk-remote](https://www.wearedevelopers.com/jobs/ext/2389038-engineering-manager-site-reliability-eu-uk-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Engineering Manager - Site Reliability *EU/UK remote* - **Company:** Pliant - **Location:** Germany (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Cursor (Graphical User Interface Elements), PCI Data Security Standards, Reliability Engineering, Software Engineering, Datadog, Kubernetes, Terraform - **Published:** August 3, 2026 - **Apply:** https://de.indeed.com/viewjob?jk=bf1384c1ebde24f0 ## About the Role * 7-10 years of engineering experience, including at least 3 years directly managing engineers * A track record of hiring and developing engineers, with specific people you've levelled or promoted * Hands-on production or reliability engineering background. This is not a first management role, and you've carried a pager yourself * Strong AWS and Terraform experience, comfortable working inside a managed infrastructure-as-code pipeline * Experience building or running an on-call rotation and incident management process, not just participating in one * Strong platform observability experience. You know the difference between a good dashboard and a useless one * Clear communication for a technical, cross-team audience * A track record of pushing reliability practices upstream into product engineering teams, not just reacting to incidents after the fact * Proficiency with AI-assisted development (Claude Code, Cursor). Comfortable reviewing AI-written PRs as rigorously as any other, Terraform, Spacelift, AWS (including a dedicated PCI-scoped account), Datadog. We're subject to PCI DSS, SOC 2, and ISO 27001, and Platform Core's migration toward Kubernetes will increasingly shape what reliability looks like here too. ## Description You're the first hire for a Site Reliability function that doesn't exist yet at Pliant. Today, reliability is a responsibility scattered across many different teams, each with their own priorities: on-call rotation is only just being introduced, there's no framework for SLOs, and nobody's job is reliability instead of firefighting it on the side of something else. You'll standardize the practice at Pliant while hiring the engineers to run it, coaching engineering teams to shift reliability from a burden into an integral part of the software development lifecycle., * Define the framework other teams use to set their own SLOs and error budgets, educating and supporting product teams along the way * Own blameless post-mortems and root-cause fixes; repeat incidents are treated as a process gap, not a signal about whoever was paged * Implement production readiness reviews so nothing new ships without one * Close gaps in Datadog observability coverage, including missing alerts, dashboard blind spots, and noisy pages that erode trust in on-call * Hire and build the team from the ground up, setting the technical and cultural bar for every engineer who joins after you, * The first few months are about building the on-call rotation and incident process from scratch, since neither exists yet, and hiring the first engineers onto the team * By mid-year, you have SLOs defined for the services that matter most, a real incident review process, and at least one engineer on the team besides you * By year one, Site Reliability is a function other teams actually route to, not something they route around, and repeat incidents are trending down because the root-cause fixes stuck ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)