Site Reliability Engineer - Kubernetes

Huxley Associates
Amsterdam, Netherlands
18 days ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Software Debugging Reliability Engineering Prometheus Grafana Technical Debt Kubernetes Puppet Terraform

Job description

This is a hands-on operational Site Reliability Engineering role focused on keeping the Kubernetes platform healthy, up-to-date, and well-supported. You will spend most of your time on cluster maintenance, component upgrades, and helping engineering teams successfully run their workloads on the platform. In addition, you may act as a consultant to product teams on Kubernetes best practices and reliability topics., * Plan and execute Kubernetes version upgrades across EKS and on-premises clusters, coordinating with internal teams to minimise disruption.

  • Perform routine maintenance, including add-on upgrades, storage and networking configuration, and upgrades for monitoring, security, and other platform tooling.
  • Monitor cluster health across the fleet and proactively address degradation signals before they become incidents.

Internal Customer Support

  • Act as the first point of contact for engineering teams running workloads on the platform.
  • Triage issues, diagnose failures, and guide teams towards resolution.
  • Help teams understand platform capabilities, quota management, and best practices for running reliable workloads.
  • Evaluate quota requests and usage requirements against platform capacity.
  • Contribute to runbooks and FAQs to reduce recurring support requests.

Toil Reduction & Automation

  • Identify repetitive manual tasks and reduce them through scripting and automation.
  • Flag and address technical debt that increases operational risk or slows down delivery.
  • Partner with the wider platform team on tooling improvements that reduce operational burden across the fleet., * A pure software development role. The focus is operational excellence, platform reliability, and customer support rather than feature development.
  • A solo contributor role. Collaboration, escalation, pairing, and knowledge sharing are essential.
  • A reactive-only role. Proactive maintenance, automation, and continuous improvement are equally important as incident response.

Requirements

  • StrongKubernetes operational experience, including:
  • Node management
  • Kubernetes upgrades
  • Workload debugging
  • Cluster health management
  • Experience with managed Kubernetes platforms (EKS or equivalent) and/or on-premises Kubernetes environments.
  • Experience withobservability tooling such as:
  • Prometheus
  • VictoriaMetrics
  • Grafana
  • Alerting pipelines
  • Strong written and verbal communication skills.
  • Ability to work independently, prioritise effectively, and meet commitments., * Terraform for infrastructure provisioning.
  • Puppet or similar configuration management tools.
  • AWS experience.
  • Experience supporting internal developer platforms or infrastructure teams., Be comfortable tackling unfamiliar technologies and challenges by asking the right questions and continuously learning.

About the company

The Kubernetes Platform team builds and operates an internal Kubernetes platform underpinning hundreds of engineering teams. The team manages a large fleet of EKS and on-premises clusters across multiple regions and focuses on treating operations as a software problem. When something is repetitive, it is automated. When something breaks, lessons are learned and improvements are made.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.huxley.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:26 min

Understanding Puppeteer and its underlying architectural design

Miki Lombardi · JS Congress

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

4:16 min

Managing large deployments and staging environments

Alexander Bubeck · WWC 2023

4:01 min

Comparing Terraform to popular configuration management tools

Devlin Duldulao · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all