> Markdown version of [/jobs/ext/2727306-senior-site-reliability-engineer-aws-eks](https://www.wearedevelopers.com/jobs/ext/2727306-senior-site-reliability-engineer-aws-eks). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer (AWS / EKS) - **Company:** Salve.Inno Consulting - **Location:** Cambridge, UK (Remote available) - **Experience:** Expert - **Salary:** £78,260.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Cloud Engineering, Disaster Recovery, Identity and Access Management, Octopus Deploy, Reliability Engineering, Prometheus, Runbook, Software Deployment, Datadog, Autoscaling, Grafana, Event Driven Architecture, Kubernetes, Infrastructure Automation Frameworks, Apache Kafka, Terraform - **Published:** September 5, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5869808074 ## About the Role We are looking for a Senior Site Reliability Engineer with deep, hands-on experience operating highly available production environments on AWS and Amazon EKS. This is a true SRE position, not a cloud architecture, infrastructure design, or monitoring-focused role. You will take direct ownership of production reliability, participate in the on-call rotation, respond to critical incidents, troubleshoot complex Kubernetes and distributed-system failures, and drive permanent improvements following incidents. The role also requires strong technical communication. You will interact directly with customers during technical discussions and production escalations, clearly explaining issues, making sound technical decisions, and driving problems through to resolution., * Significant professional experience as a hands-on Site Reliability Engineer, Production Engineer or senior Platform Engineer with direct production ownership. * Several years of recent, hands-on experience operating production environments on AWS. * Strong, demonstrable experience operating Amazon EKS in production. * Deep Kubernetes operational knowledge beyond application deployment, including cluster administration, upgrades, nodes, autoscaling, networking, troubleshooting and production failure scenarios. * Proven participation in a production on-call/pager rotation. * Demonstrable ownership of significant production incidents, including troubleshooting, mitigation, recovery, RCA and post-incident improvements. * Practical experience with SLIs, SLOs, error budgets, alerting and runbooks. * Strong Infrastructure as Code experience with Terraform and/or Terragrunt. * Production experience with Kubernetes delivery and GitOps practices; Argo CD or FluxCD strongly preferred. * Strong production observability experience with technologies such as Prometheus, Grafana, OpenTelemetry, Datadog or ELK. * Experience operating highly available, distributed production systems. * Strong understanding of AWS networking, IAM, security, availability and resilience. * Experience implementing and testing disaster recovery strategies with measurable RTO/RPO objectives. * Proven external customer-facing technical experience, including technical discussions, production escalations, architecture/reliability conversations or incident communication. * Ability to explain complex technical problems clearly and make sound decisions during high-pressure production incidents. * Strong troubleshooting mindset and ability to work independently during complex production failures. * Strong professional English communication skills (min. C1) for regular interaction with clients, * A coherent track record demonstrating sustained hands-on production engineering ownership. ## Description * Own the reliability, availability and operational health of production services running on AWS and Amazon EKS. * Operate and troubleshoot Kubernetes clusters in production, including cluster lifecycle, upgrades, node management, networking, scaling, capacity and workload reliability. * Participate actively in on-call and pager rotations and take ownership of production incidents. * Lead or play a key technical role during P1/P2 and Sev1/Sev2 incidents, including diagnosis, mitigation, recovery and communication. * Coordinate technical incident bridges and communicate directly with customers during production escalations when required. * Perform root cause analysis and lead blameless postmortems, ensuring incidents result in concrete engineering improvements. * Define, monitor and improve SLIs, SLOs and error budgets for production services. * Develop and maintain actionable alerts, operational runbooks and automated remediation. * Build and improve infrastructure using Terraform/Terragrunt and Infrastructure as Code practices. * Operate GitOps-based delivery environments using tools such as Argo CD or FluxCD. * Improve Kubernetes scaling and efficiency using technologies such as Karpenter, KEDA and native Kubernetes autoscaling capabilities. * Build and improve observability using technologies such as Prometheus, Grafana, OpenTelemetry, Datadog and/or ELK. * Support highly available distributed and event-driven systems, including environments using technologies such as Kafka/MSK. * Design, implement and validate disaster recovery and business continuity mechanisms against measurable RTO and RPO objectives. * Identify recurring operational problems and eliminate toil through automation and engineering. * Improve AWS performance, scalability, security and cost efficiency across production environments. * Work closely with software, platform and engineering teams to build reliability into systems throughout the development lifecycle. * Contribute to continuous improvement of incident management, operational readiness and SRE engineering practices. ## Related Videos - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Bridging AI and Nomad: a Go-based MCP Server for Cluster Control](https://www.wearedevelopers.com/videos/2063-bridging-ai-and-nomad-a-go-based-mcp-server-for-cluster-control) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025)