Senior Engineer - Site Reliability Engineering
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+1 more
Job description
Experteer Overview As a Senior SRE, you will help shape reliability foundations across platforms within our Markets and Risk Intelligence division. You will work with Architecture, Engineering, Security, and Platform teams to bake reliability in from day one, while occasionally supporting major incidents. This hands-on role demands proactive leadership and ownership of platform reliability outcomes. You will drive observability, incident reduction, and secure, cost-conscious operations in a complex, multi-cloud environment. Compensation / Benefits * Establish SRE foundations for new projects, ensuring readiness, monitoring, and alerting from day one * Define and champion observability standards across metrics, logs, traces, and SLIs/SLOs * Design and evolve monitoring/alerting to improve visibility and reduce toil * Drive reliability improvements through incident reduction, performance tuning, and resilient patterns * Collaborate with Security to meet compliance, security, and risk-management expectations * Lead smooth handovers from delivery to BAU SRE operations with robust documentation and practices * Provide technical leadership and mentorship to engineers, shaping standards and fostering learning Tasks * 5+ years hands-on experience in SRE, Platform Engineering, Infrastructure, or related roles * Strong experience with Azure and services like AKS, Azure Container Apps, VMs, VNet, Entra ID, and managed services * Hands-on Kubernetes and containerized platforms * Proven experience designing/operating observability platforms (monitoring, logging, alerting) * Hands-on Datadog for metrics, logs, APM, and alerting * Solid understanding of SRE principles: SLOs, error budgets, incident management * Experience collaborating with security teams and understanding cloud security principles * Experience with cloud cost optimization strategies and tooling * Good to have: AWS, multi-cloud/hybrid environments, IaC (Terraform, CloudFormation) * Exposure to large-scale, regulated environments and AI/Observability integrations Key requirements * healthcare * retirement planning * paid volunteering days * wellbeing initiatives * sustainability involvement * flexible benefits
Requirements
FULL_TIME expectations * Lead smooth handovers from delivery to BAU SRE operations with robust documentation and practices * Provide technical leadership and mentorship to engineers, shaping standards and fostering learning Tasks * 5+ years hands-on experience in SRE, Platform Engineering, Infrastructure, or related roles * Strong experience with Azure and services like AKS, Azure Container Apps, VMs, VNet, Entra ID, and managed services * Hands-on Kubernetes and containerized platforms * Proven experience designing/operating observability platforms (monitoring, logging, alerting) * Hands-on Datadog for metrics, logs, APM, and alerting * Solid understanding of SRE principles: SLOs, error budgets, incident management * Experience collaborating with security teams and understanding cloud security principles * Experience with cloud cost optimization strategies and tooling * Good to have: AWS, multi-cloud/hybrid environments, IaC (Terraform, CloudFormation) * Exposure to large-scale, aaaaaaa in environments and AI/Observability integrations Key requirements * healthcare * retirement planning * paid volunteering days * wellbeing initiatives * sustainability involvement * flexible benefits
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Software Engineer Salary London
How Much Does a Software Engineer Make? Realistic Software Engineering Salaries
Where To Find Software Engineering Jobs