> Markdown version of [/jobs/ext/2305233-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2305233-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Helsing - **Location:** London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Cloud Engineering, Computer Networks, Software Debugging, Distributed Systems, Python (Programming Language), Machine Learning, Reliability Engineering, Ansible, Prometheus, Zero Trust Network Access, Software Engineering, Policy as Code, Scripting, Istio, Grafana, Reliability of Systems, Build Management, Containerization, Templating, Kubernetes, Infrastructure Automation Frameworks, Linkerd (Service Mesh), Machine Learning Operations, Terraform, Dynatrace - **Published:** August 30, 2026 - **Apply:** https://www.collegerecruiter.com/job/2815109980-site-reliability-engineer ## About the Role * Scripting: experience in either Python, Go, Rust or Bash/ Shell for automation and tooling. * Experience with GitOps workflows and CI/CD automation. * Kubernetes Expertise: deep experience operating production Kubernetes clusters, writing custom controllers/operators, and implementing service mesh architectures (Istio/Linkerd). * Cloud-Native Technologies: hands-on experience with CNCF ecosystem, e.g. including Helm, ArgoCD, Flux and container runtime security tools like Falco. * Observability Stack: expert-level knowledge of Grafana, Prometheus, Loki, Tempo, and OpenTelemetry. Experience building custom dashboards, alerts, and SLI/SLO frameworks. * Networking: expert understanding of networking concepts, protocols and security. * MLOps Platforms: experience with Kubeflow, MLflow, or similar platforms. * Infrastructure as Code: proficiency with Terraform, Ansible, and Kubernetes manifest templating. Experience with policy-as-code tools like OPA/Gatekeeper. * System Administration: deep understanding of Linux/Unix system administration and highly available, distributed systems. * Comfortable building out data and telemetry pipelines for debugging and future-proofing solutions., * Have a high level of personal integrity, reliability, and attention to detail. * Have a software engineering mindset with a passion for building platforms and tools that multiply developer productivity. * Have experience running cloud-native workloads in on-premises or air-gapped environments. * Are willing to relocate to Munich, London, or Paris. ## Description Much of our work takes place in high-security on-premise environments, and we are looking for a Site Reliability Engineer to support our high security environments. Your role as a Site Reliability Engineer will be to design, implement, and manage our on-premise Kubernetes infrastructure. We are looking for engineers with a strong work ethic and prioritisation skills. We value team players who communicate clearly, share knowledge generously, and collaborate effectively to move their team - and our mission-forward. Day-to-Day * Design and build cloud-native infrastructure platforms on-premises, focusing on Kubernetes-based solutions that enable our development teams to operate services at scale. * Create robust observability frameworks using Grafana, Prometheus, and distributed tracing to ensure system reliability and performance. * Architect and implement secure, multi-tenant Kubernetes clusters with strong access controls, policy-as-code governance, and zero-trust networking between red and black network domains. Develop operators and controllers to automate infrastructure provisioning and compliance. * Build and maintain MLOps platforms enabling AI researchers to deploy, monitor, and scale machine learning models in production. * Collaborate closely with our Security teams to implement supply chain security, container scanning, and runtime protection across our cloud-native stack. ## Related Videos - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [Get started with securing your cloud-native Java microservices applications](https://www.wearedevelopers.com/videos/123-get-started-with-securing-your-cloud-native-java-microservices-applications) ## Related Articles - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)