> Markdown version of [/jobs/ext/1461550-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1461550-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** TK Elevator - **Location:** Madrid, Spain (Remote available) - **Contract:** Permanent contract - **Skills:** .NET Framework, Application Performance Management, Microsoft Azure, C Sharp (Programming Language), Cloud Computing, DevOps, Distributed Systems, Python (Programming Language), Log Analysis, Windows PowerShell, Reliability Engineering, Kusto Query Language, Scripting, Grafana, Azure Service Fabric, Predix - **Published:** July 28, 2026 - **Apply:** https://www.buscojobs.com.es/senior-site-reliability-engineer-d-f-m-en-madrid-ID-364775308 ## About the Role Min. 5 years in Site Reliability Engineering, DevOps, Cloud Operations, or related field - Strong expertise in Microsoft Azure and cloud-native tech - Deep knowledge of Azure Monitor, Log Analytics/KQL, Application Insights, Grafana - Experience defining/managing SLIs, SLOs, and reliability frameworks for large-scale systems - Understanding of distributed architectures, cloud platforms, and Azure PaaS services - Experience with incident management, post-mortems, and reliability improvements - Scripting/automation skills (C#/.NET, PowerShell, or Python) - Excellent analytical, problem-solving, and communication skills - Fluent English (written and spoken)Requisitos principales ## Description Experteer Overview In this role you will own and evolve system health monitoring across our digital ecosystem, establishing observability standards and driving reliable platform practices.You will work with cross-functional teams to implement SLIs/SLOs and alerting, leveraging Azure observability tools to deliver proactive insights.Your work supports scaling and resilience for the MAX IoT Platform and related products.This is a high-visibility opportunity to shape reliability culture across a global organization.Compensaciones / Beneficios- Own and evolve System Health Monitoring across products and platforms- Define and govern observability standards, monitoring requirements, health models, and alerting strategies- Unify platform health views using Azure observability solutions, Log Analytics, Grafana, and DevOps monitoring tools- Design and optimize SLIs, SLOs, and error budgets- Promote reliability engineering practices including post-incident learning- Analyze incident trends and improve monitoring, alerting, testing, and resilience- Translate monitoring data into actionable insights and predictive analytics- Collaborate with Product, Architecture, DevOps, and Incident Operations teams on monitoring coverage and alert quality- Provide guidance and mentorship to engineering teams; maintain runbooks and incident procedures- Contribute knowledge to the global DevOps communityResponsabilidades- Min. 5 years in Site Reliability Engineering, DevOps, Cloud Operations, or related field- Strong expertise in Microsoft Azure and cloud-native tech- Deep knowledge of Azure Monitor, Log Analytics/KQL, Application Insights, Grafana- Experience defining/managing SLIs, SLOs, and reliability frameworks for large-scale systems- Understanding of distributed architectures, cloud platforms, and Azure PaaS services- Experience with incident management, post-mortems, and reliability improvements- Scripting/automation skills (C#/.NET, PowerShell, or Python)- Excellent analytical, problem-solving, and communication skills- Fluent English (written and spoken)Requisitos principales- health and safety programs- flexible working hours- remote working options- training and education programs- modern workplaces and IT equipment- subsidized meals and discounted transport tickets ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Azure-Well Architected Framework - designing mission critical workloads in practice](https://www.wearedevelopers.com/videos/1529-azure-well-architected-framework-designing-mission-critical-workloads-in-practice) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)