> Markdown version of [/jobs/ext/1409469-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1409469-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** London Stock Exchange Group - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Linux, Distributed Systems, Redis, Runbook, Datadog, Data Processing, Grafana, Git, Infrastructure Automation Frameworks, Apache Flink, Vertica, Data Pipelines - **Published:** July 23, 2026 - **Apply:** https://www.dice.com/job-detail/28aac692-f86d-44de-b8c2-40fda5cbbe6b ## About the Role * Experience using observability data such as metrics, logs, traces, alerts, dashboards, or service health views to investigate issues. * Understanding of incident response, problem management, service readiness, or operational support processes. * Experience using automation or infrastructure-as-code practices to support repeatable delivery and operations. * Solid understanding of cloud, container, Linux, networking, or distributed system environments. * Ability to communicate technical information clearly to engineering, operations, and service stakeholders. * Experience writing or maintaining operational documentation such as runbooks, support guides, or onboarding material. * A practical approach to improving reliability, reducing manual work, and helping teams use shared platforms optimally. Desirable skills and experience: * Experience with observability, monitoring, telemetry, or data pipeline technologies such as OpenTelemetry, Grafana, ClickHouse, Cribl, Datadog, BigPanda, Redis, Flink, or similar tools. * Experience with GitOps workflows and tools such as Git, CI/CD pipelines, pull requests, environment promotion, or configuration-as-code. * Experience building or supporting internal platforms used by multiple engineering teams. * Experience defining or using SLOs, SLIs, error budgets, alert quality measures, or service health models. * Experience supporting telemetry pipelines, data routing, data filtering, retention, or cost management. * Experience working in a regulated, financial services, or large enterprise technology environment. * Experience helping engineering teams adopt shared standards, templates, or self-service platform capabilities. ## Description * Monitor platform health, investigate and diagnose issues, and support service recovery across production and non-production environments. * Contribute to service readiness through dashboards, SLOs, SLIs, runbooks, resilience checks, and performance validation. * Use metrics, logs, traces, and service health data to understand issues and guide practical decisions. * Incident response and continuous improvement: * Support incident response, triage, problem management, and root cause analysis. * Drive follow-up actions that reduce repeat issues and improve operational consistency. * Improve automation and GitOps-based workflows to reduce manual work across the team. * Enablement and documentation: * Maintain documentation, onboarding guides, and self-service tooling so teams can operate with less reliance on direct support. * Provide direct enablement on telemetry standards, alerting guidance, SLO practices, and platform workflows. ## Related Videos - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Serverless Observability: where SLOs meet transforms](https://www.wearedevelopers.com/videos/854-serverless-observability-where-slos-meet-transforms) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Best Companies to work for in London: Top 25 Companies in 2023](https://www.wearedevelopers.com/magazine/187-best-companies-to-work-for-in-london-top-25-companies-in-2023) - [Fullstack Developer Salary UK](https://www.wearedevelopers.com/magazine/251-fullstack-developer-salary-uk) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)