Senior Site Reliability Engineer (all genders)

FactFinder
Pforzheim, Germany
25 days ago
Apply on www.adzuna.de
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Artificial Intelligence Software as a Service VMware ESX Servers Octopus Deploy OpenStack Reliability Engineering Software Tools Virtual Local Area Networks Virtualization Technology VMware VSphere Ceph (Software) Load Balancing
+3 more
Autoscaling Git Flow Kubernetes

Job description

FACT-Finder builds product discovery technology for eCommerce and is trusted by leading online shops across Europe with its two products Next Generation and Infinity. Both products are moving toward a modern, hybrid platform based on Kubernetes and Harvester - with the option to scale fully into the cloud in the mid-term. As a Senior Site Reliability Engineer (SRE), you make sure our systems stay fast, available, and scalable throughout this transformation. You work closely with the Hosting team and experienced engineers, and actively shape our journey toward a modern SaaS company., * You define and own SLOs, SLIs, and error budgets across both products and make data-driven decisions on reliability and performance.

  • You drive incident response: fast detection, clear communication, blameless postmortems, and meaningful follow-through.
  • You consistently reduce manual work through automation and GitOps (e.g. Argo CD / Flux) and build out self-healing and self-service capabilities.
  • You support the development of an NG Search Operator (custom Kubernetes operator / CRDs) and the rollout of auto-scaling (HPA, VPA, KEDA, cluster autoscaler).
  • You evolve our observability - metrics, logs, traces, alerting, and runbooks that actually help on call.
  • You plan capacity and cost across on-premise (Frankfurt, Stockholm) and cloud - including burst scenarios into the public cloud.
  • You leverage AI tools to noticeably accelerate diagnosis, alerting, and operational workflows., * Impact from day one: Your work directly influences the revenue of leading eCommerce brands across Europe.
  • Modern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path toward the cloud - with room to build things right.
  • AI-first mindset: We use AI as a real part of our daily work, not as a buzzword.
  • Ownership & growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role.
  • Flexible work: Hybrid work model with a focus on outcomes.
  • Strong team: Experienced engineers, an open feedback culture, and an environment where reliability is treated as a real engineering discipline.

Requirements

  • Experience as an SRE, infrastructure, or production engineer in a SaaS or platform environment - or a strong software/operations background with a clear drive to grow into an SRE role.
  • Solid understanding of SLOs, error budgets, incident management, and observability.
  • Hands-on experience with Kubernetes and interest in cluster lifecycle, upgrades, and operator patterns.
  • Experience with or strong interest in Harvester or comparable HCI/virtualization platforms (KubeVirt, vSphere/ESXi, OpenStack).
  • Familiarity with GitOps (Argo CD / Flux), container storage (Longhorn, Ceph), and Kubernetes networking (load balancing, ingress).
  • Knowledge of auto-scaling primitives (HPA, VPA, cluster autoscaler, KEDA) and capacity planning on-prem and in the cloud.
  • Understanding of networking in production-grade datacenters (incl. VLAN).
  • A strong automation instinct and a mindset to structurally eliminate toil.
  • Practical experience using AI tools in day-to-day operations.
  • Fluent English; German is a plus.

About the company

We are one of the leading Product Discovery Platforms for European eCommerce. Our search, navigation and recommendations handle billions of shopper queries every year - driving a combined GMV of more than €160 billion across our customers. These include Intersport, Spar, Douglas and more than 2,000 further B2B and B2C eCommerce companies across Europe.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.de
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

2:10 min

Transitioning domain specific analytics from OpenStack to Kubernetes

Ingvord Ingvord · World Congress 2024

5:02 min

Mapping Git flow branches to application tester segments

Majid Hajian · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

1:29 min

Transitioning from content management to cloud infrastructure

Matt Butcher · World Congress 2023

3:53 min

Introduction to git flow and clean feature branches

Johannes Haux · World Congress 2022

Videos

See all

Related articles

See all