Senior Site Reliability Engineer

TK Elevator
Madrid, Spain
21 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
5 years minimum
Working hours
Shift work
Languages
English

Tech stack

.NET Framework Application Performance Management Microsoft Azure C Sharp (Programming Language) Cloud Computing DevOps Distributed Systems Python (Programming Language) Log Analysis Windows PowerShell Reliability Engineering Kusto Query Language
+4 more
Scripting Grafana Azure Service Fabric Predix

Job description

Experteer Overview In this role you will own and evolve system health monitoring across our digital ecosystem, establishing observability standards and driving reliable platform practices.You will work with cross-functional teams to implement SLIs/SLOs and alerting, leveraging Azure observability tools to deliver proactive insights.Your work supports scaling and resilience for the MAX IoT Platform and related products.This is a high-visibility opportunity to shape reliability culture across a global organization.Compensaciones / Beneficios- Own and evolve System Health Monitoring across products and platforms- Define and govern observability standards, monitoring requirements, health models, and alerting strategies- Unify platform health views using Azure observability solutions, Log Analytics, Grafana, and DevOps monitoring tools- Design and optimize SLIs, SLOs, and error budgets- Promote reliability engineering practices including post-incident learning- Analyze incident trends and improve monitoring, alerting, testing, and resilience- Translate monitoring data into actionable insights and predictive analytics- Collaborate with Product, Architecture, DevOps, and Incident Operations teams on monitoring coverage and alert quality- Provide guidance and mentorship to engineering teams; maintain runbooks and incident procedures- Contribute knowledge to the global DevOps communityResponsabilidades- Min. 5 years in Site Reliability Engineering, DevOps, Cloud Operations, or related field- Strong expertise in Microsoft Azure and cloud-native tech- Deep knowledge of Azure Monitor, Log Analytics/KQL, Application Insights, Grafana- Experience defining/managing SLIs, SLOs, and reliability frameworks for large-scale systems- Understanding of distributed architectures, cloud platforms, and Azure PaaS services- Experience with incident management, post-mortems, and reliability improvements- Scripting/automation skills (C#/.NET, PowerShell, or Python)- Excellent analytical, problem-solving, and communication skills- Fluent English (written and spoken)Requisitos principales- health and safety programs- flexible working hours- remote working options- training and education programs- modern workplaces and IT equipment- subsidized meals and discounted transport tickets

Requirements

Min. 5 years in Site Reliability Engineering, DevOps, Cloud Operations, or related field

  • Strong expertise in Microsoft Azure and cloud-native tech
  • Deep knowledge of Azure Monitor, Log Analytics/KQL, Application Insights, Grafana
  • Experience defining/managing SLIs, SLOs, and reliability frameworks for large-scale systems
  • Understanding of distributed architectures, cloud platforms, and Azure PaaS services
  • Experience with incident management, post-mortems, and reliability improvements
  • Scripting/automation skills (C#/.NET, PowerShell, or Python)
  • Excellent analytical, problem-solving, and communication skills
  • Fluent English (written and spoken)Requisitos principales

Benefits & conditions

flexible working hours

  • remote working options
  • training and education programs
  • modern workplaces and IT equipment
  • subsidized meals and discounted transport tickets

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all