Senior Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+4 more
Job description
Experteer Overview In this role you will own and evolve system health monitoring across our digital ecosystem, establishing observability standards and driving reliable platform practices.You will work with cross-functional teams to implement SLIs/SLOs and alerting, leveraging Azure observability tools to deliver proactive insights.Your work supports scaling and resilience for the MAX IoT Platform and related products.This is a high-visibility opportunity to shape reliability culture across a global organization.Compensaciones / Beneficios- Own and evolve System Health Monitoring across products and platforms- Define and govern observability standards, monitoring requirements, health models, and alerting strategies- Unify platform health views using Azure observability solutions, Log Analytics, Grafana, and DevOps monitoring tools- Design and optimize SLIs, SLOs, and error budgets- Promote reliability engineering practices including post-incident learning- Analyze incident trends and improve monitoring, alerting, testing, and resilience- Translate monitoring data into actionable insights and predictive analytics- Collaborate with Product, Architecture, DevOps, and Incident Operations teams on monitoring coverage and alert quality- Provide guidance and mentorship to engineering teams; maintain runbooks and incident procedures- Contribute knowledge to the global DevOps communityResponsabilidades- Min. 5 years in Site Reliability Engineering, DevOps, Cloud Operations, or related field- Strong expertise in Microsoft Azure and cloud-native tech- Deep knowledge of Azure Monitor, Log Analytics/KQL, Application Insights, Grafana- Experience defining/managing SLIs, SLOs, and reliability frameworks for large-scale systems- Understanding of distributed architectures, cloud platforms, and Azure PaaS services- Experience with incident management, post-mortems, and reliability improvements- Scripting/automation skills (C#/.NET, PowerShell, or Python)- Excellent analytical, problem-solving, and communication skills- Fluent English (written and spoken)Requisitos principales- health and safety programs- flexible working hours- remote working options- training and education programs- modern workplaces and IT equipment- subsidized meals and discounted transport tickets
Requirements
Min. 5 years in Site Reliability Engineering, DevOps, Cloud Operations, or related field
- Strong expertise in Microsoft Azure and cloud-native tech
- Deep knowledge of Azure Monitor, Log Analytics/KQL, Application Insights, Grafana
- Experience defining/managing SLIs, SLOs, and reliability frameworks for large-scale systems
- Understanding of distributed architectures, cloud platforms, and Azure PaaS services
- Experience with incident management, post-mortems, and reliability improvements
- Scripting/automation skills (C#/.NET, PowerShell, or Python)
- Excellent analytical, problem-solving, and communication skills
- Fluent English (written and spoken)Requisitos principales
Benefits & conditions
flexible working hours
- remote working options
- training and education programs
- modern workplaces and IT equipment
- subsidized meals and discounted transport tickets
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Are The Top Skills Required For Azure Developers?
Where To Find Software Engineering Jobs
Find a Developer Job: 12 Best Job Sites For Developers
Top-Paying Tech Jobs (with Salaries)