SRE Technical Lead- Bristol

FDM Group
Bristol, UK
about 1 month ago
Apply on uk.indeed.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Operational Data Store Reliability Engineering Site Reliability Engineering Practices Data Logging Containerization Infrastructure Automation Frameworks Dynatrace Service Stack Legacy Systems

Job description

FDM is a global business and technology consultancy seeking an SRE Technical Lead to support a major global financial services organisation as it establishes its first formal Site Reliability Engineering function. This is initially a 12-month contract with the potential to go permanent and will be a hybrid role based in Bristol. This role offers a unique opportunity to play a key technical leadership role within a newly formed SRE function. Reporting directly to the Head of SRE, you will be responsible for driving the adoption of reliability engineering practices across critical business services, helping to improve service stability, resilience, observability, and operational efficiency.

As a senior individual contributor, you will provide technical leadership rather than people management. You will work closely with platform, infrastructure, engineering, and support teams to implement SRE principles, define reliability standards, reduce operational toil through automation, and establish meaningful service health measurements. You will help accelerate the organisation’s transition from reactive production support towards a proactive, engineering-led reliability model, ensuring reliability is designed into services rather than addressed after incidents occur Responsibilities:

  • Partner with the Head of SRE to implement and embed the organisation’s SRE strategy, operating model, and reliability standards.
  • Act as a technical authority for reliability engineering, providing guidance and expertise across application, platform, and infrastructure teams.
  • Define, implement, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Critical User Journeys (CUJs) to establish meaningful service reliability metrics.
  • Drive adoption of SLO-based decision making, supporting teams in balancing reliability, delivery velocity, and operational risk.
  • Identify opportunities to reduce operational toil through automation, including runbook automation, self-healing capabilities, deployment improvements, and recovery processes.
  • Design and implement observability best practices across logging, metrics, tracing, alerting, and dashboarding.
  • Support and improve incident management processes, participating in major incident response and post-incident reviews while driving root cause analysis and preventative actions.
  • Work with engineering teams to improve service resilience, availability, scalability, and recoverability through proactive engineering improvements.
  • Analyse reliability trends and operational data to identify systemic issues and recommend long-term solutions.
  • Contribute to the development of reliability standards, frameworks, and technical roadmaps across both legacy and modern technology environments.
  • Champion engineering excellence and reliability best practices through mentoring, knowledge sharing, and collaboration with technical teams.
  • Support technology transformation initiatives by ensuring operational resilience and reliability requirements are embedded throughout delivery programmes.

Requirements

  • Significant hands-on experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or a similar reliability-focused role.
  • Strong background in supporting and improving business-critical production environments.
  • Proven experience implementing SRE practices, including SLIs, SLOs, error budgets, and observability frameworks.
  • Experience working within large, complex enterprise environments.
  • Strong engineering mindset with practical experience delivering automation and operational improvements.
  • Deep understanding of incident management, problem management, and operational resilience principles.
  • Experience designing and implementing monitoring, logging, alerting, and distributed tracing solutions.
  • Strong troubleshooting and root cause analysis skills across complex technology stacks.
  • Ability to influence technical teams and stakeholders without formal line management responsibility.
  • Strong communication skills with the ability to explain complex technical concepts to both technical and non-technical audiences.
  • Comfortable working in evolving environments where processes and capabilities are still being established.
  • Pragmatic and outcome-focused approach to solving reliability and operational challenges.

Desirable:

  • Experience helping establish or mature SRE capabilities within an organisation.
  • Exposure to large-scale technology transformation programmes.
  • Experience working with legacy platforms alongside cloud-native technologies.
  • Experience standardising observability practices across multiple teams and toolsets.
  • Familiarity with financial services or other highly regulated environments.
  • Experience building automation solutions using scripting and Infrastructure as Code practices.
  • Experience mentoring engineers or contributing to reliability communities of practice.
  • Knowledge of cloud platforms, containerisation, orchestration technologies, and modern platform engineering practices.

Benefits & conditions

Pulled from the full job description

  • Annual leave
  • Company pension

About the company

FDM is an award-winning global leader in tech and business talent solutions, backed by more than 35 years of industry experience. We have centres across Europe, North America, and Asia-Pacific, and a global workforce of over 2500 employees. FDM has shown exponential growth throughout the years, firmly establishing itself as an award-winning employer, currently listed on the FTSE4Good Index and as a 2026 Financial Times UK ‘Best Employer’.

Diversity and Inclusion FDM Group is an equal opportunity employer, and all qualified applicants will receive consideration for employment without regard to race, colour, religion, sex, sexual orientation, national origin, age, disability, veteran status or any other status protected by federal, provincial or local laws.

Why join us

  • Career coaching, mentoring and access to upskilling throughout your entire FDM career
  • Assignments with global companies and opportunities to work abroad
  • Opportunity to re-skill and up-skill into new areas, develop non-linear career paths and build a skillset within your field
  • Annual leave and work-place pension

You must create an Indeed account before continuing to the company website to apply

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on uk.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Advocating for SRE practices within agency environments

Martin Beránek · LIVE

3:47 min

Solving the knowledge deficit in large language models

Alejandro Saucedo Alejandro Saucedo +3 · World Congress 2024

1:10 min

Exposing sensitive information through partial search logs

Dennis Schulz Dennis Schulz +1 · World Congress 2026 Europe

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · World Congress 2025

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

41 sec

Operating critical stack observability and service monitoring

Michael Cade · LIVE

Videos

See all

Related articles

See all