Site Reliability Engineer (SRE) - Financial Wellbeing

Lloyds Banking Group
London, UK
2 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
£70,929.0 - £78,810.0
Working hours
Regular working hours

Tech stack

Clean Code Principles Agile Methodology JIRA Microsoft Azure Cloud Computing Github Key Management Reliability Engineering Workflow Management Systems Scripting File Transfer Protocol (FTP) Cloud Platform System
+4 more
Containerization Terraform Docker Jenkins

Job description

The Recoveries & Application Management Lab are building new systems and solutions. The role offers the chance for someone to engage and share their experience, serving as a mentor in the team’s decision-making processes related to site reliability that are effective and efficient.

Once the SRE support is evaluated, the selected person will join the Platform Engineering team to establish the SRE role by managing the cloud environment, developing solution designs, explaining incident management practices, conducting root cause analysis, and handling business changes. The successful candidate will be required to balance the demands of the LBG business and both stability and scalability requirements to reach a successful and scalable outcome.

Day to Day

  • Create documentation that details the establishment of the SRE function within the platform, supported by procedures that outline the guidelines to be followed through the incorporation of existing documentation.
  • Provide a framework in which to operate the cloud systems.
  • Lead the transition to cloud infrastructure and improve observability across systems.
  • Identify and eliminate toil through automation.
  • Manage incidents and post-mortems to improve service reliability.
  • Mentor engineers and support team development.
  • Collaborate with Product Owners to balance operational and development priorities.

Requirements

  • Proven experience as a Site Reliability Engineer in cloud environments (GCP or AWS).
  • Understanding of SRE principles including SLIs, SLOs, error budgets, and toil reduction.
  • Strong scripting and infrastructure-as-code (IaaC) skills (Terraform, Harness, GitHub).
  • Demonstrable experience in the Agile ways of working that focuses on delivering customer value and applying the Agile mindset; familiarity with tools like Jira.
  • Ability to lead incident response and drive service improvements.
  • Strong collaboration and mentoring skills., * Azure cloud environment experience, including connectivity, data buckets, secrets management, migration, and governance challenges.
  • Familiarity with containerisation and orchestration tools like Docker, Jenkins, GitHub, and Terraform.
  • Secure programming practices and experience of secure file transfer protocols, risk remediation, and audit actions.
  • Technical operations and service engineering.

Benefits & conditions

  • A generous pension contribution of up to 15%
  • An annual performance-related bonus
  • Share schemes including free shares
  • Benefits you can adapt to your lifestyle, such as discounted shopping
  • 30 days’ holiday, with bank holidays on top
  • A range of wellbeing initiatives and generous parental leave policies

This is a once in a career opportunity to help shape your future as well as ours. Join us and grow with purpose.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:33 min

Advocating for SRE practices within agency environments

Martin Beránek · LIVE

5:47 min

Integrating user stories and test automation via Jira tools

Christoph Ruggenthaler · LIVE

Videos

See all

Related articles

See all