Site Reliability Engineer

Raw Group España
Barcelona, Spain
25 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
4 years minimum
Working hours
Shift work
Languages
German
Job source

Tech stack

Linux DevOps PostgreSQL Load Testing Performance Tuning Reliability Engineering Site Reliability Engineering Practices Prometheus Data Logging Grafana Kubernetes Helm Charts Software Troubleshooting
+1 more
Kubernetes

Job description

We are looking for an experienced Site Reliability Engineer to ensure the stability, scalability, and operational excellence of a Kubernetes-based platform running in a hybrid environment. The project is entering a pivotal phase, with a major go-live planned for mid-February and a target audience of 75,000 users. User onboarding is already underway, with over 5,000 users connected and 15,000-20,000 expected to be active by year-end. While the system is stable, we anticipate increased activity and new challenges in January, February, and after the go-live-making this an exciting opportunity to make a real impact. The role focuses on performance optimization, scaling strategies, observability, and reliability engineering., Operate and optimize hybrid infrastructure (on-prem & STACKIT) Manage and scale Kubernetes clusters Optimize Helm charts, resource usage, and autoscaling Conduct performance, load, and stress testing Ensure reliability, availability, and monitoring of production systems Tune and operate PostgreSQL Operate and optimize vector databases (e.g. Qdrant) Implement monitoring, logging, and alerting Support incident response and capacity planning

Requirements

4+ years of experience as SRE / DevOps Engineer Strong hands-on experience with Kubernetes in production Experience working with hybrid infrastructure (on-prem + cloud) Solid knowledge of PostgreSQL performance tuning and scaling Experience with Qdrant or other vector databases Experience with Helm, Kubernetes autoscaling, and resource optimization Familiarity with observability stacks (Prometheus, Grafana, ELK/Loki) Understanding of performance engineering and load testing Experience with Linux systems and networking Strong troubleshooting and incident-management skills Nice to Have: Experience with STACKIT or other sovereign clouds Experience with PgBouncer Knowledge of SRE practices (SLO/SLI) Experience in regulated or public-sector environments German language skills

Benefits & conditions

Flexible working format - remote, office-based or flexible A competitive salary and good compensation package Personalized career growth Professional development tools (mentorship program, tech talks and trainings, centers of excellence, and more) Active tech communities with regular knowledge sharing Education reimbursement Memorable anniversary presents Corporate events and team buildings Other location-specific benefits *not applicable for freelancers Inscribirse en esta oferta Recibir ofertas similares por correo electrónico Al crear una alerta, aceptas nuestros Términos y condiciones y Política de privacidad, y el uso de cookies.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

1:06 min

Developer experience and project variety at scale

Alexandra Petri · WWC 2023

Videos

See all

Related articles

See all