> Markdown version of [/jobs/ext/2301503-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2301503-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** WALT Labs - **Location:** London, UK - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** JIRA, Bash Shell, Cloud Computing, Cloud Engineering, Monitoring of Systems, Python (Programming Language), Reliability Engineering, Datadog, Scripting, Google Cloud, Grafana, Multi-Cloud, Kubernetes, Gsuite, Terraform, Golang - **Published:** August 30, 2026 - **Apply:** https://www.collegerecruiter.com/job/2815120811-site-reliability-engineer ## About the Role * 3-5 years experience with Google Cloud Platform * Minimum 2 Google Cloud Professional certifications * Advanced Kubernetes knowledge and troubleshooting * Proficient in Infrastructure as Code (Terraform) * Strong scripting abilities (Python, Go, Bash) * Expert with monitoring tools (Grafana, Datadog) * Experience leading incident response * Excellent communication and mentoring skills * Proven track record of process improvement * Ability to manage multiple priorities effectively * 20 holiday days + bank holidays (earn 1.5 days every 3 years) ## Description This is a full-time on-site role 3 days a week minimum in Kings Cross London. We are seeking a skilled Site Reliability Engineer with a strong focus on Google Cloud Platform (GCP) to join our dynamic team. In this role, you'll be responsible for maintaining cloud infrastructure, managing incidents, and ensuring seamless operations for our clients. You'll use tools like incident.io and JIRA to manage and resolve support requests efficiently., * Serve as L2 on-call escalation point for complex technical issues requiring advanced troubleshooting * Lead response to critical incidents, coordinating multiple teams and ensuring effective communication * Provide expert-level support for GCP services including advanced networking, security, and architecture * Perform advanced Google Workspace administration including domain management, security policies, and integration * Use incident.io to manage escalated incidents, major incidents, and coordinate war room activities * Optimize support workflows in JIRA, creating automation rules and improving ticket routing * Monitor and tune infrastructure performance using advanced Grafana queries and custom metrics * Lead technical projects including migrations, upgrades, and new service implementations * Create comprehensive documentation including architectural diagrams, runbooks, and best practices guides * Achieve minimum 50% billable hours through complex Cloud Assist/Managed Cloud customers and consulting engagements * Mentor Cloud Support Engineers and juniors through formal and informal training sessions * Identify and implement process improvements to increase efficiency and reduce resolution time * Conduct thorough root cause analysis for recurring issues and implement permanent fixes * Present technical solutions and recommendations to customer stakeholders and management * Design and implement monitoring strategies for complex multi-cloud environments * Develop automation scripts and tools to improve team efficiency and reduce manual work * Participate in pre-sales activities providing technical expertise for solution design * Review and approve changes to production environments following change management procedures * Lead knowledge sharing sessions and technical deep-dives for the team * Coordinate with vendor support for complex issues requiring manufacturer assistance * Maintain expertise in multiple GCP services and stay current with new feature releases * Participation in business hours escalation rotation ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) ## Related Articles - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london) - [Best Companies to work for in London: Top 25 Companies in 2023](https://www.wearedevelopers.com/magazine/187-best-companies-to-work-for-in-london-top-25-companies-in-2023) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Fullstack Developer Salary UK](https://www.wearedevelopers.com/magazine/251-fullstack-developer-salary-uk)