Site Reliability Engineer (Observability) - (Hybrid)

Edreams
Barcelona, Spain
22 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Continuous Integration Software Debugging Distributed Systems Python (Programming Language) Reliability Engineering Site Reliability Engineering Practices YAML Scripting Mttr Kubernetes
+3 more
Deployment Automation Terraform Golang

Job description

Experteer Overview In this hybrid Barcelona role, you will drive the reliability of our global platform by applying SRE practices to maximize uptime and observability.You’ll define and monitor SLIs/SLOs, reduce toil through automation, and lead incident response and post?mortem learning.You’ll collaborate with cross?functional teams to instrument code, build dashboards, and scale resilient infrastructure that underpins a world?leading travel subscription platform.Compensaciones / Beneficios* Lead incident response, triage, and troubleshooting in complex distributed systems* Design and implement automated remediation to reduce operational toil* Manage comprehensive observability (monitoring, alerts, logs) for ecosystem health* Facilitate blameless post?mortems and drive improvements from incidents* Act as internal consultant/evangelist; train product teams on instrumentation (OpenTelemetry/APM)* Support GitOps and automation to ensure stability (CI/CD, IaC, containerization)* Define and monitor SLIs/SLOs to align performance with business needs* Optimize infrastructure via code for scalability and availability* Create/manage automated deployment lifecycles* Collaborate with other teams to apply best practices in the development lifecycle* Implement CI, CD, and deployment methodologies to improve MTTD/MTTRResponsabilidades* Solid incident management experience* Advanced debugging in distributed systems* Experience with Infrastructure as Code (Terraform, Terragrunt)* Scripting skills (Python, YAML, Go)* Automation tools familiarity (Terraform, ArgoCD, Crossplane)* Kubernetes expertise* Experience with GCP (and cloud providers like AWS/Azure)* OpenTelemetry/APM instrumentation knowledge* Proactive, data?driven reliability mindset* CAN DO attitude and learning agilityRequisitos principales* Hybrid home?office model* Relocation support* Premium equipment and role?based options* Coursera access and ongoing training* Career development programs (eVOLVE)* Flexible benefits and performance?based bonuses

Requirements

  • Solid incident management experience
  • Advanced debugging in distributed systems
  • Experience with Infrastructure as Code (Terraform, Terragrunt)
  • Scripting skills (Python, YAML, Go)
  • Automation tools familiarity (Terraform, ArgoCD, Crossplane)
  • Kubernetes expertise
  • Experience with GCP (and cloud providers like AWS/Azure)
  • OpenTelemetry/APM instrumentation knowledge
  • Proactive, data?driven reliability mindset
  • CAN DO attitude and learning agility

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · WWC 2021

1:35 min

Centralizing configuration logic with native YAML block references

Matthieu Vincent Matthieu Vincent · Europe 2026 Virtual

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · WWC Europe 2026

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz · LIVE

Videos

See all

Related articles

See all