> Markdown version of [/jobs/ext/66943-site-reliability-engineer-kubernetes-multi-cloud-uk-based](https://www.wearedevelopers.com/jobs/ext/66943-site-reliability-engineer-kubernetes-multi-cloud-uk-based). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (Kubernetes / Multi-Cloud) UK Based - **Company:** Synalogik - **Location:** Hereford, UK (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Continuous Integration, Monitoring of Systems, Identity and Access Management, Key Management, Network Security, Role-Based Access Control, Reliability Engineering, Prometheus, Data Logging, Cloud Monitoring, Autoscaling, Istio, System Availability, Grafana, Kubernetes Helm Charts, Multi-Cloud, Containerization, Kubernetes, Cloudwatch, Terraform, Docker - **Published:** May 31, 2026 - **Apply:** https://www.apply4u.co.uk/jobs/x/37792239/ ## About the Role Experience with Azure and/or AWS Familiarity with networking, IAM, and core services Infrastructure as Code Experience with Terraform Observability Familiarity with monitoring/logging tools (Prometheus, Grafana, loki) Other Technical Skills Helm Charts / Kustomize creation and maintenance Containers (Docker) Exposure to both Azure and AWS GitOps tools (ArgoCD / Flux) Autoscaling tools (KEDA, Karpenter) Service mesh exposure Personal Attributes Problem-solving mindset Willingness to learn Proactive and dependable Qualifications 2-4 years of experience in cloud/SRE/platform roles Location - Hybrd (Hereford based) or Remote Employment Type - Full Time Residency - You must have been Resident in the UK for 5 years to meet security Clearance for this role ## Description We are looking for a Site Reliability Engineer (SRE) to join an established and growing SRE team supporting Kubernetes-based platforms running across Azure and AWS This role focuses on maintaining reliable, scalable, and observable systems, working closely with engineering teams to ensure services run smoothly in production. You will contribute to the operation of managed Kubernetes platforms (AKS/EKS), supporting best practices in monitoring, automation, and incident response, while continuing to develop your expertise in cloud-native technologies., Site Reliability Engineering Participate in incident response, troubleshooting, and post-incident reviews Help reduce operational toil through automation and process improvements Contribute to improving system availability, performance, and scalability Maintain and improve runbooks and operational documentation Participate in 24/7 On-Call Rota Support the deployment and operation of AKS and EKS clusters Assist with cluster upgrades, scaling, and maintenance Work with autoscaling tools (Cluster Autoscaler, KEDA, Karpenter) Help improve workload reliability and performance Support networking, identity, compute, and storage services Assist with maintaining secure and scalable environments Observability & Monitoring Work with Prometheus, Grafana, OpenTelemetry, Azure Monitor, and CloudWatch Build dashboards, alerts, and logging/tracing pipelines Support monitoring aligned to SLIs/SLOs Security & Compliance Implement RBAC/IAM, secrets management, and network security controls Support compliance and security requirements CI/CD & Automation Contribute to GitOps workflows (ArgoCD / Flux) Assist in automating deployments and operations Work closely with software engineers and product owners Contribute to team discussions, reviews, and planning Hands-on Kubernetes experience (AKS, EKS, or similar) Understanding of cluster architecture, networking, and scaling Cloud ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Monoskope: Developer Self-Service Across Clusters](https://www.wearedevelopers.com/videos/329-monoskope-developer-self-service-across-clusters) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fullstack Developer Salary UK](https://www.wearedevelopers.com/magazine/251-fullstack-developer-salary-uk) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Best Job Boards for Remote Work for Developers](https://www.wearedevelopers.com/magazine/290-best-job-boards-for-remote-work-for-developers)