> Markdown version of [/jobs/ext/1445176-sre-technical-lead-bristol](https://www.wearedevelopers.com/jobs/ext/1445176-sre-technical-lead-bristol). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE Technical Lead- Bristol - **Company:** FDM Group - **Location:** Bristol, UK - **Experience:** Expert - **Contract:** Temporary to permanent - **Skills:** Operational Data Store, Reliability Engineering, Site Reliability Engineering Practices, Data Logging, Containerization, Infrastructure Automation Frameworks, Dynatrace, Service Stack, Legacy Systems - **Published:** July 26, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=bf78d752fd54daf1 ## About the Role * Significant hands-on experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or a similar reliability-focused role. * Strong background in supporting and improving business-critical production environments. * Proven experience implementing SRE practices, including SLIs, SLOs, error budgets, and observability frameworks. * Experience working within large, complex enterprise environments. * Strong engineering mindset with practical experience delivering automation and operational improvements. * Deep understanding of incident management, problem management, and operational resilience principles. * Experience designing and implementing monitoring, logging, alerting, and distributed tracing solutions. * Strong troubleshooting and root cause analysis skills across complex technology stacks. * Ability to influence technical teams and stakeholders without formal line management responsibility. * Strong communication skills with the ability to explain complex technical concepts to both technical and non-technical audiences. * Comfortable working in evolving environments where processes and capabilities are still being established. * Pragmatic and outcome-focused approach to solving reliability and operational challenges. Desirable: * Experience helping establish or mature SRE capabilities within an organisation. * Exposure to large-scale technology transformation programmes. * Experience working with legacy platforms alongside cloud-native technologies. * Experience standardising observability practices across multiple teams and toolsets. * Familiarity with financial services or other highly regulated environments. * Experience building automation solutions using scripting and Infrastructure as Code practices. * Experience mentoring engineers or contributing to reliability communities of practice. * Knowledge of cloud platforms, containerisation, orchestration technologies, and modern platform engineering practices. ## Description FDM is a global business and technology consultancy seeking an SRE Technical Lead to support a major global financial services organisation as it establishes its first formal Site Reliability Engineering function. This is initially a 12-month contract with the potential to go permanent and will be a hybrid role based in Bristol. This role offers a unique opportunity to play a key technical leadership role within a newly formed SRE function. Reporting directly to the Head of SRE, you will be responsible for driving the adoption of reliability engineering practices across critical business services, helping to improve service stability, resilience, observability, and operational efficiency. As a senior individual contributor, you will provide technical leadership rather than people management. You will work closely with platform, infrastructure, engineering, and support teams to implement SRE principles, define reliability standards, reduce operational toil through automation, and establish meaningful service health measurements. You will help accelerate the organisation's transition from reactive production support towards a proactive, engineering-led reliability model, ensuring reliability is designed into services rather than addressed after incidents occur Responsibilities: * Partner with the Head of SRE to implement and embed the organisation's SRE strategy, operating model, and reliability standards. * Act as a technical authority for reliability engineering, providing guidance and expertise across application, platform, and infrastructure teams. * Define, implement, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Critical User Journeys (CUJs) to establish meaningful service reliability metrics. * Drive adoption of SLO-based decision making, supporting teams in balancing reliability, delivery velocity, and operational risk. * Identify opportunities to reduce operational toil through automation, including runbook automation, self-healing capabilities, deployment improvements, and recovery processes. * Design and implement observability best practices across logging, metrics, tracing, alerting, and dashboarding. * Support and improve incident management processes, participating in major incident response and post-incident reviews while driving root cause analysis and preventative actions. * Work with engineering teams to improve service resilience, availability, scalability, and recoverability through proactive engineering improvements. * Analyse reliability trends and operational data to identify systemic issues and recommend long-term solutions. * Contribute to the development of reliability standards, frameworks, and technical roadmaps across both legacy and modern technology environments. * Champion engineering excellence and reliability best practices through mentoring, knowledge sharing, and collaboration with technical teams. * Support technology transformation initiatives by ensuring operational resilience and reliability requirements are embedded throughout delivery programmes. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Introducing a Digital Service Catalog for speed and scale](https://www.wearedevelopers.com/videos/763-introducing-a-digital-service-catalog-for-speed-and-scale) - [Crypto-secure Data Management with In-Database Blockchain](https://www.wearedevelopers.com/videos/632-crypto-secure-data-management-with-in-database-blockchain) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [Building high performance and scalable architectures for enterprises](https://www.wearedevelopers.com/videos/19-building-high-performance-and-scalable-architectures-for-enterprises) - [Build Delightful Mobile Experiences with Kotlin, Realm, and Atlas Device Sync](https://www.wearedevelopers.com/videos/694-build-delightful-mobile-experiences-with-kotlin-realm-and-atlas-device-sync) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Fullstack Developer Salary UK](https://www.wearedevelopers.com/magazine/251-fullstack-developer-salary-uk) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)