Senior Site Reliability Engineer

Lodgify
Madrid, Spain
7 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
7 years minimum
Working hours
Regular working hours
Languages
Spanish

Tech stack

Application Programming Interfaces (APIs) Cloud Computing Databases Disaster Recovery Python (Programming Language) Reliability Engineering Site Reliability Engineering Practices Prometheus Data Streaming Datadog Grafana Caching
+1 more
Kubernetes

Job description

Experteer Overview In this Senior Site Reliability Engineer role, you will strengthen the reliability and scalability of Lodgify’s shared infrastructure and product services within the Platform team.You’ll define and drive practical SRE standards, improve observability, and reduce toil to empower engineering teams in production.You’ll work on critical infrastructure, deployment safety, incident response, and cost-efficient resilience as the company scales globally.You’ll collaborate across Engineering, Platform, Security, and Product to turn reliability into a measurable business advantage.This is a chance to shape a high-velocity, international tech stack with a strong focus on production excellence and a Compensaciones / Beneficios * Define meaningful SLIs, SLOs, and reliability targets for the platform * Collaborate with software teams to establish observability best practices, SLIs/SLOs, and reliability * Strengthen production readiness through ownership improvements, alerting, runbooks, scaling assumptions, rollback paths, and failure-mode planning * Improve reliability, scalability, and performance of cloud, Kubernetes, and shared infrastructure during growth and traffic spikes * Build actionable observability using metrics, logs, traces, and Golden Signals with Datadog, Prometheus, and Grafana * Implement operational and security best practices via guidelines, policies, and automation * Reduce alert noise and improve signal quality for faster issue resolution * Automate repetitive operational tasks using Python or other languages * Provide self-service Internal Developer Platform features via APIs and Kubernetes operators * Improve deployment safety, rollbackability, and release observability * Enhance reliability of stateful systems (databases, caches, queues, streaming platforms) * Participate in on-call, troubleshoot, incident response, and blameless post-incident reviews * Execute disaster recovery drills and analyze cloud usage for cost/resource efficiency gains without compromising reliability Responsabilidades * 7+ years of production experience with Kubernetes-based platforms and cloud infrastructure * Strong grasp of SRE practices including SLIs, SLOs, error budgets, incident response, toil reduction, and DR * Ability to design observability and alerting for critical systems using metrics, logs, traces, and golden signals * Experience writing maintainable automation software to reduce manual intervention * Experience with stateful production systems (relational DBs, caches, queues, streaming platforms) * Ability to balance reliability, performance, cost, and delivery speed pragmatically * Comfort in transitional environments introducing SRE practices while maintaining hands-on reliability support * Collaborates effectively with Engineering, Platform, Security, and Product stakeholders * Clear communication, documentation, and coaching to improve production ownership * Initiative, accountability, and driving improvements to completion Requisitos principales * Remote work flexibility * Health insurance (Alan) * Paid vacation (25 days) * Meal allowance and office meals * Home office gear provided * Language classes (Spanish)

Requirements

Internal Developer Platform features via APIs and Kubernetes operators * Improve deployment safety, rollbackability, and release observability * Enhance reliability of stateful systems (databases, caches, queues, streaming platforms) * Participate in on-call, troubleshoot, incident response, and blameless post-incident reviews * Execute disaster recovery drills and analyze cloud usage for cost/resource efficiency gains without compromising reliability Responsabilidades * 7+ years of production experience with Kubernetes-based platforms and cloud infrastructure * Strong grasp of SRE practices including SLIs, SLOs, error budgets, incident response, toil reduction, and DR * Ability to design observability and alerting for critical systems using metrics, logs, traces, and golden signals * Experience writing maintainable automation software to reduce manual intervention * Experience with stateful production systems (relational DBs, caches, queues, streaming platforms) * Ability to balance reliability, performance, cost, and delivery speed pragmatically * Comfort in transitional environments introducing SRE practices while maintaining hands-on reliability support * Collaborates effectively with Engineering, Platform, Security, and Product stakeholders * Clear communication, documentation, and coaching to improve production ownership * Initiative, accountability, and driving improvements to completion Requisitos principales * Remote work flexibility * Health insurance (Alan) * Paid vacation (25 days) * Meal allowance and office meals * Home office gear provided * Language classes (Spanish)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

1:07 min

Architecting the availability stack with Prometheus and Grafana

Gabriel Labachelerie · World Congress 2023

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel · World Congress 2025

Videos

See all

Related articles

See all