> Markdown version of [/jobs/ext/2977861-site-reliability-engineer-europe](https://www.wearedevelopers.com/jobs/ext/2977861-site-reliability-engineer-europe). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (Europe) - **Company:** Factorial - **Location:** Madrid, Spain (Remote available) - **Salary:** €65,000.0 - €80,000.0 - **Contract:** Permanent contract - **Skills:** Bash Shell, Continuous Integration, Data as a Services, Linux, Github, Python (Programming Language), MySQL, Octopus Deploy, Ruby on Rails, Redis, Prometheus, Delivery Pipeline, Caching, Kubernetes, Bare Metal, Vertica, Terraform, Docker - **Published:** September 18, 2026 - **Apply:** https://www.adzuna.es/contact-us.html ## About the Role + 5+ years running production infrastructure, with Kubernetes at the centre of it. + Real depth in Kubernetes: scheduling, requests and limits, evictions, node pressure, DaemonSets, controllers and operators. You have debugged a cluster that was lying to you. + Strong Linux and container fundamentals. containerd or Docker internals, cgroups, namespaces, storage drivers, networking. + Terraform and infrastructure as code, with a feel for module design and state hygiene. + CI/CD at the platform level: shared workflows, reusable templates, runner architecture, build caching. + Comfortable with GitHub Actions, including self-hosted runners. + You script to automate. Bash, plus at least one of Python or Go. + Observability as a working method. Metrics, logs, traces, dashboards and alerts that people trust, with OpenTelemetry and Prometheus style tooling. + Operational maturity. On-call, incident command, postmortems, and the discipline to chase causes. + Capacity and cost awareness. You can size a fleet and explain the bill. + Clear written English. You document what you build, because the next person on call will not be you. ## Description Every engineer here waits on CI several times a day. We want someone who finds that wait personally annoying. Why this role exists Hundreds of developers push to a large Ruby on Rails monorepo, and every push lands on a runner fleet that the Developer Experience team builds, runs and keeps fast. That fleet is self-hosted Kubernetes on bare metal in European datacentres, in the low hundreds of machines, running ephemeral GitHub Actions runners that peak above a thousand concurrent pods. We own all of it: the hardware, the clusters, the runner images, the caches and the delivery pipelines on top. We add nodes by hand rather than letting an autoscaler do it, so capacity planning is genuinely part of the job. The leverage is unusual. Shave a minute off the average build and every engineer in the company gets that minute back, several times a day. The mission + Own the CI and CD Kubernetes clusters end to end: capacity, reliability, upgrades, security and cost. + Run the self-hosted GitHub Actions platform at scale with Actions Runner Controller. Runner scale sets, ephemeral pods, docker-in-docker, and the runner images themselves. + Keep the data services our tests depend on fast and healthy. MySQL, Redis and ClickHouse come up per job, alongside Rails application containers. + Design and tune the caching that makes builds quick: node-local overlay, image and artifact caches we build ourselves, plus pull-through registry mirrors. + Bring queue wait and build duration down, with measurement behind it. Instrument the platform, set targets, and show the improvement. + Manage it all as code. Terraform planned and applied from pull requests, GitOps delivery with Flux and Argo CD, Kustomize and Helm for manifests. + Provision and operate bare metal. Linux, networking across datacentre segments, and troubleshooting that sometimes ends up at the disk or the kernel. + Lead incident response for the platform, run blameless postmortems, and turn each one into an alert, a guardrail or a runbook. + Treat the platform as a product whose users are engineers. Talk to them, watch where they get stuck, and build paths that make the right thing easy. Your day to day + Rightsizing runner tiers against a week of CPU and memory data, then opening the pull request that changes them. + Tracking a flaky job from a red check down to the container runtime, the kernel, or a clock that drifted. + Adding machines to the fleet: install, enrol, network, verify, document. + Cutting p95 queue wait by working out which stage actually blocks. + Upgrading a cluster or a controller without anyone noticing. + Pairing with a product engineer on a workflow that has been slow for a reason nobody has looked at yet. ## Related Videos - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Coding for Good: Achieving social change with an app](https://www.wearedevelopers.com/videos/1645-coding-for-good-achieving-social-change-with-an-app) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)