Senior Engineer - Site Reliability Engineering

London Stock Exchange Group
Raleigh, NC, United States
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Security Performance Tuning Reliability Engineering Datadog Data Logging Multi-Cloud Cloudformation Kubernetes Cloud Optimization
+1 more
Terraform

Job description

Experteer Overview As a Senior SRE, you will help shape reliability foundations across platforms within our Markets and Risk Intelligence division. You will work with Architecture, Engineering, Security, and Platform teams to bake reliability in from day one, while occasionally supporting major incidents. This hands-on role demands proactive leadership and ownership of platform reliability outcomes. You will drive observability, incident reduction, and secure, cost-conscious operations in a complex, multi-cloud environment. Compensation / Benefits * Establish SRE foundations for new projects, ensuring readiness, monitoring, and alerting from day one * Define and champion observability standards across metrics, logs, traces, and SLIs/SLOs * Design and evolve monitoring/alerting to improve visibility and reduce toil * Drive reliability improvements through incident reduction, performance tuning, and resilient patterns * Collaborate with Security to meet compliance, security, and risk-management expectations * Lead smooth handovers from delivery to BAU SRE operations with robust documentation and practices * Provide technical leadership and mentorship to engineers, shaping standards and fostering learning Tasks * 5+ years hands-on experience in SRE, Platform Engineering, Infrastructure, or related roles * Strong experience with Azure and services like AKS, Azure Container Apps, VMs, VNet, Entra ID, and managed services * Hands-on Kubernetes and containerized platforms * Proven experience designing/operating observability platforms (monitoring, logging, alerting) * Hands-on Datadog for metrics, logs, APM, and alerting * Solid understanding of SRE principles: SLOs, error budgets, incident management * Experience collaborating with security teams and understanding cloud security principles * Experience with cloud cost optimization strategies and tooling * Good to have: AWS, multi-cloud/hybrid environments, IaC (Terraform, CloudFormation) * Exposure to large-scale, regulated environments and AI/Observability integrations Key requirements * healthcare * retirement planning * paid volunteering days * wellbeing initiatives * sustainability involvement * flexible benefits

Requirements

FULL_TIME expectations * Lead smooth handovers from delivery to BAU SRE operations with robust documentation and practices * Provide technical leadership and mentorship to engineers, shaping standards and fostering learning Tasks * 5+ years hands-on experience in SRE, Platform Engineering, Infrastructure, or related roles * Strong experience with Azure and services like AKS, Azure Container Apps, VMs, VNet, Entra ID, and managed services * Hands-on Kubernetes and containerized platforms * Proven experience designing/operating observability platforms (monitoring, logging, alerting) * Hands-on Datadog for metrics, logs, APM, and alerting * Solid understanding of SRE principles: SLOs, error budgets, incident management * Experience collaborating with security teams and understanding cloud security principles * Experience with cloud cost optimization strategies and tooling * Good to have: AWS, multi-cloud/hybrid environments, IaC (Terraform, CloudFormation) * Exposure to large-scale, aaaaaaa in environments and AI/Observability integrations Key requirements * healthcare * retirement planning * paid volunteering days * wellbeing initiatives * sustainability involvement * flexible benefits

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:34 min

Managing token budgets and enterprise usage of coding agents

Chris Heilmann +2 · LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel · WWC 2025

2:33 min

Advocating for SRE practices within agency environments

Martin Beránek · LIVE

Videos

See all

Related articles

See all