Site Reliability Engineer (Observability) - (Hybrid)
Role details
Job location
Tech stack
Job description
Experteer Overview In this hybrid Barcelona role, you will drive the reliability of our global platform by applying SRE practices to maximize uptime and observability.You'll define and monitor SLIs/SLOs, reduce toil through automation, and lead incident response and post?mortem learning.You'll collaborate with cross?functional teams to instrument code, build dashboards, and scale resilient infrastructure that underpins a world?leading travel subscription platform.Compensaciones / Beneficios* Lead incident response, triage, and troubleshooting in complex distributed systems* Design and implement automated remediation to reduce operational toil* Manage comprehensive observability (monitoring, alerts, logs) for ecosystem health* Facilitate blameless post?mortems and drive improvements from incidents* Act as internal consultant/evangelist; train product teams on instrumentation (OpenTelemetry/APM)* Support GitOps and automation to ensure stability (CI/CD, IaC, containerization)* Define and monitor SLIs/SLOs to align performance with business needs* Optimize infrastructure via code for scalability and availability* Create/manage automated deployment lifecycles* Collaborate with other teams to apply best practices in the development lifecycle* Implement CI, CD, and deployment methodologies to improve MTTD/MTTRResponsabilidades* Solid incident management experience* Advanced debugging in distributed systems* Experience with Infrastructure as Code (Terraform, Terragrunt)* Scripting skills (Python, YAML, Go)* Automation tools familiarity (Terraform, ArgoCD, Crossplane)* Kubernetes expertise* Experience with GCP (and cloud providers like AWS/Azure)* OpenTelemetry/APM instrumentation knowledge* Proactive, data?driven reliability mindset* CAN DO attitude and learning agilityRequisitos principales* Hybrid home?office model* Relocation support* Premium equipment and role?based options* Coursera access and ongoing training* Career development programs (eVOLVE)* Flexible benefits and performance?based bonuses
Requirements
- Solid incident management experience
- Advanced debugging in distributed systems
- Experience with Infrastructure as Code (Terraform, Terragrunt)
- Scripting skills (Python, YAML, Go)
- Automation tools familiarity (Terraform, ArgoCD, Crossplane)
- Kubernetes expertise
- Experience with GCP (and cloud providers like AWS/Azure)
- OpenTelemetry/APM instrumentation knowledge
- Proactive, data?driven reliability mindset
- CAN DO attitude and learning agility