Site Reliability Engineer (Europe)

Factorial
Madrid, Spain
6 days ago
Apply on www.adzuna.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
5 years minimum
Compensation
€65,000.0 - €80,000.0
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Bash Shell Continuous Integration Data as a Services Linux Github Python (Programming Language) MySQL Octopus Deploy Ruby on Rails Redis Prometheus Delivery Pipeline
+6 more
Caching Kubernetes Bare Metal Vertica Terraform Docker

Job description

Every engineer here waits on CI several times a day. We want someone who finds that wait personally annoying. Why this role exists Hundreds of developers push to a large Ruby on Rails monorepo, and every push lands on a runner fleet that the Developer Experience team builds, runs and keeps fast. That fleet is self-hosted Kubernetes on bare metal in European datacentres, in the low hundreds of machines, running ephemeral GitHub Actions runners that peak above a thousand concurrent pods. We own all of it: the hardware, the clusters, the runner images, the caches and the delivery pipelines on top. We add nodes by hand rather than letting an autoscaler do it, so capacity planning is genuinely part of the job. The leverage is unusual. Shave a minute off the average build and every engineer in the company gets that minute back, several times a day. The mission

  • Own the CI and CD Kubernetes clusters end to end: capacity, reliability, upgrades, security and cost.
  • Run the self-hosted GitHub Actions platform at scale with Actions Runner Controller. Runner scale sets, ephemeral pods, docker-in-docker, and the runner images themselves.
  • Keep the data services our tests depend on fast and healthy. MySQL, Redis and ClickHouse come up per job, alongside Rails application containers.
  • Design and tune the caching that makes builds quick: node-local overlay, image and artifact caches we build ourselves, plus pull-through registry mirrors.
  • Bring queue wait and build duration down, with measurement behind it. Instrument the platform, set targets, and show the improvement.
  • Manage it all as code. Terraform planned and applied from pull requests, GitOps delivery with Flux and Argo CD, Kustomize and Helm for manifests.
  • Provision and operate bare metal. Linux, networking across datacentre segments, and troubleshooting that sometimes ends up at the disk or the kernel.
  • Lead incident response for the platform, run blameless postmortems, and turn each one into an alert, a guardrail or a runbook.
  • Treat the platform as a product whose users are engineers. Talk to them, watch where they get stuck, and build paths that make the right thing easy. Your day to day

  • Rightsizing runner tiers against a week of CPU and memory data, then opening the pull request that changes them.
  • Tracking a flaky job from a red check down to the container runtime, the kernel, or a clock that drifted.
  • Adding machines to the fleet: install, enrol, network, verify, document.
  • Cutting p95 queue wait by working out which stage actually blocks.
  • Upgrading a cluster or a controller without anyone noticing.
  • Pairing with a product engineer on a workflow that has been slow for a reason nobody has looked at yet.

Requirements

  • 5+ years running production infrastructure, with Kubernetes at the centre of it.
  • Real depth in Kubernetes: scheduling, requests and limits, evictions, node pressure, DaemonSets, controllers and operators. You have debugged a cluster that was lying to you.
  • Strong Linux and container fundamentals. containerd or Docker internals, cgroups, namespaces, storage drivers, networking.
  • Terraform and infrastructure as code, with a feel for module design and state hygiene.
  • CI/CD at the platform level: shared workflows, reusable templates, runner architecture, build caching.
  • Comfortable with GitHub Actions, including self-hosted runners.
  • You script to automate. Bash, plus at least one of Python or Go.
  • Observability as a working method. Metrics, logs, traces, dashboards and alerts that people trust, with OpenTelemetry and Prometheus style tooling.
  • Operational maturity. On-call, incident command, postmortems, and the discipline to chase causes.
  • Capacity and cost awareness. You can size a fleet and explain the bill.
  • Clear written English. You document what you build, because the next person on call will not be you.

About the company

  • Large Ruby on Rails test suites, and what it takes to parallelise them honestly.
  • Operating MySQL, Redis or ClickHouse yourself.
  • Bare metal: provisioning, hardware failure, and the different rhythm it has compared to cloud.
  • Overlay and WireGuard style mesh networking, VLANs, and Kubernetes CNIs such as Cilium.
  • Building internal developer tooling, portals or platform APIs.
  • Supply chain and build security. Runner isolation, secret scoping, image provenance.
  • Working across more than one cloud. We use AWS, Azure and Cloudflare alongside our own hardware.
  • Contributions to open source infrastructure projects. How we work At Factorial, we believe the best products are built when people come together, in person, to collaborate, challenge ideas, and move fast. That’s why our Engineering Team follows an office-first, flexible approach. We work on-site several days a week (80%), using that time to connect, align, and innovate as a team. However, we also support remote work when it makes sense (20%) for deep focus or personal needs. Perks and benefits

  • High growth, multicultural and friendly environment
  • Continuous training and learning based on your needs
  • Alan private health insurance
  • Healthy life with Wellhub (gyms, pools, outdoor classes)
  • Save expenses with Cobee
  • Language classes with Preply
  • Get the most out of your salary with Payflow And when at the office:

  • Breakfast in the office and organic fruit
  • Nora and Apeteat discounts
  • Pet friendly About Factorial Factorial is an innovative Business Management Software solution designed to streamline company processes for small and medium-sized enterprises. Founded in 2016, our mission is to help companies automate workflows, centralize people data, and make better business decisions. With customers across over 60 countries, we’ve built a diverse team that’s driving change in the tech space. Our values

  • We own it: We take responsibility for every project.
  • We learn and teach: We share knowledge every day.
  • We partner: Every decision is a team decision.
  • We grow fast: We act fast and learn from our mistakes. Diversity Diversity is part of our culture. We have more than 43 nationalities and provide an inclusive environment for all. Please feel free to apply however suits you best (blind resume, identity pronouns, cover letter, etc.). We do not discriminate; we encourage everyone to join us! Inscribirse en esta oferta

Estadísticas para este empleo

Comparación salarial

Este empleo

Media Nacional

Media - Otros trabajos Comunidad de Madrid

  • Media

Salarios Número de empleos en el rango salarial: Recibir ofertas similares por correo electrónico Al crear una alerta, aceptas nuestros Términos y condiciones y Política de privacidad, y el uso de cookies., Aircall.io Madrid Staff Software Engineer EUR 65000-80000 per year Green Eagle Solutions Madrid, Community of Madrid, 28008 AI & Automation Engineer 35000-70000 GRUPO VERICAT Madrid Site Reliability Engineer (Europe) Arango Madrid Volver a la última búsqueda Trabajos ) Comunidad de Madrid ) Madrid ) Otros trabajos ) Staff Engineer - Performance, Reliability & AI Automation ( volver a la última búsqueda Recibir ofertas similares por correo electrónico No gracias, llévame a la oferta de empleo Al crear una alerta, aceptas nuestros Términos y condiciones y Política de privacidad, y el uso de cookies. Inscribirse en esta oferta

Profesiones

  • Técnico
  • Recepciónista
  • Administrador
  • Ventas
  • Enfermero

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

8:22 min

Simulating a Linux terminal and running Spring Boot

Jakov Semenski · LIVE

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · World Congress 2022

Videos

See all

Related articles

See all