> Markdown version of [/jobs/ext/1413845-site-reliability-engineer-kubernetes](https://www.wearedevelopers.com/jobs/ext/1413845-site-reliability-engineer-kubernetes). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - Kubernetes - **Company:** Huxley Associates - **Location:** Amsterdam, Netherlands - **Contract:** Temporary contract - **Skills:** Software Debugging, Reliability Engineering, Prometheus, Grafana, Technical Debt, Kubernetes, Puppet, Terraform - **Published:** July 24, 2026 - **Apply:** https://www.huxley.com/en-gb/job/site-reliability-engineer---kubernetes/4064889 ## About the Role * StrongKubernetes operational experience, including: + Node management + Kubernetes upgrades + Workload debugging + Cluster health management * Experience with managed Kubernetes platforms (EKS or equivalent) and/or on-premises Kubernetes environments. * Experience withobservability tooling such as: + Prometheus + VictoriaMetrics + Grafana + Alerting pipelines * Strong written and verbal communication skills. * Ability to work independently, prioritise effectively, and meet commitments., * Terraform for infrastructure provisioning. * Puppet or similar configuration management tools. * AWS experience. * Experience supporting internal developer platforms or infrastructure teams., Be comfortable tackling unfamiliar technologies and challenges by asking the right questions and continuously learning. ## Description This is a hands-on operational Site Reliability Engineering role focused on keeping the Kubernetes platform healthy, up-to-date, and well-supported. You will spend most of your time on cluster maintenance, component upgrades, and helping engineering teams successfully run their workloads on the platform. In addition, you may act as a consultant to product teams on Kubernetes best practices and reliability topics., * Plan and execute Kubernetes version upgrades across EKS and on-premises clusters, coordinating with internal teams to minimise disruption. * Perform routine maintenance, including add-on upgrades, storage and networking configuration, and upgrades for monitoring, security, and other platform tooling. * Monitor cluster health across the fleet and proactively address degradation signals before they become incidents. Internal Customer Support * Act as the first point of contact for engineering teams running workloads on the platform. * Triage issues, diagnose failures, and guide teams towards resolution. * Help teams understand platform capabilities, quota management, and best practices for running reliable workloads. * Evaluate quota requests and usage requirements against platform capacity. * Contribute to runbooks and FAQs to reduce recurring support requests. Toil Reduction & Automation * Identify repetitive manual tasks and reduce them through scripting and automation. * Flag and address technical debt that increases operational risk or slows down delivery. * Partner with the wider platform team on tooling improvements that reduce operational burden across the fleet., * A pure software development role. The focus is operational excellence, platform reliability, and customer support rather than feature development. * A solo contributor role. Collaboration, escalation, pairing, and knowledge sharing are essential. * A reactive-only role. Proactive maintenance, automation, and continuous improvement are equally important as incident response. ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Automate everything via NodeJS and Puppeteer](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) - [One Platform Could Not Fit Them All](https://www.wearedevelopers.com/videos/1919-one-platform-could-not-fit-them-all) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Terraform for Developers](https://www.wearedevelopers.com/videos/3-terraform-for-developers) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) ## Related Articles - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)