> Markdown version of [/jobs/ext/2729349-senior-site-reliability-engineer-sre-kubernetes-hybrid-cloud](https://www.wearedevelopers.com/jobs/ext/2729349-senior-site-reliability-engineer-sre-kubernetes-hybrid-cloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer / SRE - Kubernetes & Hybrid Cloud - **Company:** FactFinder - **Location:** München, Germany - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, VMware ESX Servers, Octopus Deploy, OpenStack, Reliability Engineering, Prometheus, Virtual Local Area Networks, VMware VSphere, Ceph (Software), Load Balancing, Autoscaling, Grafana, HybridCloud, Git Flow, Kubernetes, Deployment Automation - **Published:** September 5, 2026 - **Apply:** https://www.adzuna.de/details/5870690671 ## About the Role * Kubernetes in production - built, not just used: you've set up and maintained clusters on your own servers (e.g.kubeadm, RKE2, k3s) and know cluster lifecycle and upgrades - managed-only experience isn't enough for this role * Lived SRE practice: SLOs, error budgets, incident management, on-call * Hands-on experience with GitOpsor comparable infrastructure/deployment automation - experience with Argo CD or Flux is a strong plus * Solid observability skills - metrics, logs, traces, alerting that people trust * A strong automation instinct - you'drather fix a problem's cause than repeat its workaround * A collaborative, enabling mindset - you see SRE as a service to our developers: you ask what they need, discuss trade-offs openly, anddon'tfall in love with your own solution Nice-to-haves (genuinely optional - we'll teach you the rest): * Harvester, KubeVirt, vSphere/ESXi, OpenStack or similar virtualization/HCI platforms * Container storage (Longhorn, Ceph) and datacenter networking (load balancing, ingress, VLAN) * Auto-scaling (HPA, VPA, KEDA, clusterautoscaler) and capacity/cost planning * Experience building Kubernetes operators/CRDs * German language skills Certifications (CKA, CKS) are welcome but no substitute for hands-on experience - in the tech interview we'll ask about what you've actually built and operated. You don't tick every box - or your title was never "SRE"? Apply anyway. If you've owned production systems, handled incidents and worked deeply with Kubernetes, we want to hear from you - production experience and engineering mindset matter more to us than titles or buzzwords. ## Description * Tech stack: Kubernetes on our own servers, Harvester (KubeVirt), Argo CD/Flux, Prometheus/Grafana, Longhorn/Ceph * Team: A growing SRE team - you report to our CTPO for now and to the Team Lead SRE we're hiring next; two system administrators in Pforzheim run the physical hardware * Process: Intro call · take-home task (~2h) · 90-min tech interview with our developers · leadership conversation · meet the team * Languages: Fluent English required; German is a plus, not a must Why this role is special Most SRE jobs today mean clicking around a managed cloud console. This one doesn't. We run our own hardware in Frankfurt and are building a modern private cloud platform on Kubernetes and Harvester - on-prem by default, with elastic burst into the public cloud and the option to go cloud-only later. You won't inherit a finished SRE practice: you'll help define it, side by side with our Berlin development teams - and you won't do it alone, a Team Lead SRE hire is coming next. SRE here is an enabling discipline: you build what our developers need to ship reliably, while two system administrators in Pforzheim run the physical hardware. And the impact is direct - our product discovery technology powers more than 2,000 European online shops (Intersport, SPAR, Douglas and more), handling billions of shopper queries a year. When product discovery is slow or down, our customers lose revenue in real time. Your first 90 days You get to know both products, join the on-call rotation with a buddy, and own your first reliability topic - SLOs for one product, alerting that actually helps at 3 a.m., or automating away a piece of toil. By day 90 you've shipped visible improvements and know where you want to take the platform next. Your mission * Define and own SLOs, SLIs and error budgets; drive data-informed reliability decisions * Lead incident response end-to-end: fast detection, clear communication, blameless postmortems - and reduce whole classes of incidents structurally, not case by case * Eliminate toil through automation and GitOps; evolve our observability (metrics, logs, traces, alerting, runbooks) across two different stacks * Help build our custom Kubernetes operator (CRDs) that makes stateful search clusters declarative, self-healing and safely upgradable - and roll out the auto-scaling (HPA/VPA, KEDA, cluster auto scaler) today's architecture makes hard * Plan capacity, performance and cost across on-premises and cloud - including the large-catalogue and peak-season loads our merchants care about - and use AI tools wherever they measurably speed up diagnosis and operations, * Impact from day one: Your work directly influences the revenue of leading eCommerce brands across Europe. * Modern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path toward the cloud - with room to build things right. * AI-first mindset: We use AI as a real part of our daily work, not as a buzzword. * Ownership & growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role. * Flexible work: Hybrid work model three office days per week with a focus on outcomes. * Strong team: Experienced engineers, an open feedback culture, and an environment where reliability is treated as a real engineering discipline. ## Related Videos - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Why Git Still Matters](https://www.wearedevelopers.com/videos/100288-why-git-still-matters) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Finding Jobs in Germany](https://www.wearedevelopers.com/magazine/375-finding-jobs-in-germany) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where to Find Entry-Level Software Engineering Jobs](https://www.wearedevelopers.com/magazine/397-where-to-find-entry-level-software-engineering-jobs)