> Markdown version of [/jobs/ext/2468392-site-reliability-engineering-professional](https://www.wearedevelopers.com/jobs/ext/2468392-site-reliability-engineering-professional). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineering Professional - **Company:** BT Group - **Location:** Ipswich, UK - **Contract:** Permanent contract - **Skills:** Border Gateway Protocol, Continuous Integration, Ethernet, Monitoring of Systems, Multi-protocol Systems, Open Shortest Path First (OSPF), Reliability Engineering, Prometheus, Wide Area Networks, Web Platforms, Computer Network Operations, System Availability, Grafana, Splunk, Dynatrace - **Published:** August 12, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=54303dd262fc3f4b ## About the Role * Experience supporting business-critical services within a 24x7 Operations, NOC, Service Operations, or Site Reliability Engineering (SRE) environment. * Strong technical knowledge of IP and Optical networking and WAN technologies, including MPLS, BGP, OSPF, Ethernet, SD-WAN, Internet, and Cloud connectivity services. * Experience troubleshooting complex end-to-end services across network, platform, cloud, and application environments. * Knowledge of SRE principles, including reliability, availability, automation, observability, and operational resilience. * Hands-on experience with monitoring and observability tools such as Dynatrace, Splunk, Grafana, ELK, Prometheus, or similar platforms. * Strong working knowledge of Incident, Problem, Change, and Major Incident Management practices, including Root Cause Analysis (RCA). * Excellent analytical and problem-solving skills, with the ability to make decisions under pressure. * Strong communication and stakeholder management skills, with a customer-focused approach. * Ability to work collaboratively across global, cross-functional teams and adapt in a fast-paced operational environment. * Passion for continuous improvement, innovation, and operational transformation. ## Description BT International is transforming the way we operate, evolving from traditional network operations to a modern Site Reliability Engineering (SRE) and Platform Operations model. As a Site Reliability Engineering Professional, you will play a key role in the operational management of BT International's global core platforms, ensuring they are reliable, secure, scalable, and deliver an exceptional customer experience. Working within our 2nd Line Operations team, you will support critical services across our international network and digital platforms, driving operational excellence, rapid incident resolution, and continual service improvement. This is an exciting opportunity to work at the heart of BT International's next-generation platforms, collaborating with engineering, product, supplier, and operational teams to improve service reliability through automation, observability, and SRE best practices. What you'll be doing * Support the 24x7 operation of BT International's core network and platform services, ensuring high availability and performance. * Proactively monitor services, identify risks, and prevent customer-impacting incidents. * Lead technical service restoration activities during major incidents and act as an escalation point for complex operational issues. * Deliver against operational KPIs, SLAs, OLAs, and customer service targets. * Drive continuous improvement through root cause analysis, problem management, automation, and defect reduction. * Work closely with Engineering and Product teams to improve platform reliability, resilience, and operational readiness. * Build and maintain high-quality operational documentation, including runbooks, service maps, playbooks, and handover processes. * Develop and enhance monitoring, observability, and operational tooling capabilities. * Support the implementation of automation and CI/CD practices to improve operational efficiency. * Coach and support your colleagues and customer-facing operational teams, ensuring customer service excellence remains at the heart of everything we do. * Collaborate effectively with global operational teams, suppliers, and stakeholders to deliver outstanding service outcomes. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Keycloak case study: Making users happy with service level indicators and observability](https://www.wearedevelopers.com/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Best Companies to work for in London: Top 25 Companies in 2023](https://www.wearedevelopers.com/magazine/187-best-companies-to-work-for-in-london-top-25-companies-in-2023) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [UK Business Culture and Etiquette](https://www.wearedevelopers.com/magazine/326-uk-business-culture-and-etiquette) - [The Geometry of Incidents: Connecting User Impact to Architecture](https://www.wearedevelopers.com/magazine/764-the-geometry-of-incidents-connecting-user-impact-to-architecture)