Site Reliability Engineer (SRE) - Financial Wellbeing
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+4 more
Job description
The Recoveries & Application Management Lab are building new systems and solutions. The role offers the chance for someone to engage and share their experience, serving as a mentor in the team’s decision-making processes related to site reliability that are effective and efficient.
Once the SRE support is evaluated, the selected person will join the Platform Engineering team to establish the SRE role by managing the cloud environment, developing solution designs, explaining incident management practices, conducting root cause analysis, and handling business changes. The successful candidate will be required to balance the demands of the LBG business and both stability and scalability requirements to reach a successful and scalable outcome.
Day to Day
- Create documentation that details the establishment of the SRE function within the platform, supported by procedures that outline the guidelines to be followed through the incorporation of existing documentation.
- Provide a framework in which to operate the cloud systems.
- Lead the transition to cloud infrastructure and improve observability across systems.
- Identify and eliminate toil through automation.
- Manage incidents and post-mortems to improve service reliability.
- Mentor engineers and support team development.
- Collaborate with Product Owners to balance operational and development priorities.
Requirements
- Proven experience as a Site Reliability Engineer in cloud environments (GCP or AWS).
- Understanding of SRE principles including SLIs, SLOs, error budgets, and toil reduction.
- Strong scripting and infrastructure-as-code (IaaC) skills (Terraform, Harness, GitHub).
- Demonstrable experience in the Agile ways of working that focuses on delivering customer value and applying the Agile mindset; familiarity with tools like Jira.
- Ability to lead incident response and drive service improvements.
- Strong collaboration and mentoring skills., * Azure cloud environment experience, including connectivity, data buckets, secrets management, migration, and governance challenges.
- Familiarity with containerisation and orchestration tools like Docker, Jenkins, GitHub, and Terraform.
- Secure programming practices and experience of secure file transfer protocols, risk remediation, and audit actions.
- Technical operations and service engineering.
Benefits & conditions
- A generous pension contribution of up to 15%
- An annual performance-related bonus
- Share schemes including free shares
- Benefits you can adapt to your lifestyle, such as discounted shopping
- 30 days’ holiday, with bank holidays on top
- A range of wellbeing initiatives and generous parental leave policies
This is a once in a career opportunity to help shape your future as well as ours. Join us and grow with purpose.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Software Engineer Salary London
Is Software Engineering Over-Saturated?
Best Companies to work for in London: Top 25 Companies in 2023
Why Upskilling And Reskilling is Important For Developers