> Markdown version of [/jobs/ext/3101586-systems-engineering-manager-sre-ml-compute](https://www.wearedevelopers.com/jobs/ext/3101586-systems-engineering-manager-sre-ml-compute). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Systems Engineering Manager, SRE, ML Compute - **Company:** Hackajob Ltd - **Location:** Leeds, UK - **Experience:** Experienced - **Salary:** £61,000.0 - £101,000.0 - **Contract:** Permanent contract - **Skills:** Systems Engineering, Build Automation, Cloud Computing, Data Centers, File Systems, Fault Tolerance, Network Topologies, Infrastructure as a Service (IaaS), Routing, Reliability Engineering, System Programming, TCP/IP, Graphics Processing Unit (GPU), Information Technology, Low Latency, Programming Languages - **Published:** September 27, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5899595299 ## About the Role * We require a bachelors degree in Computer Science or a related technical field, or equivalent practical experience. * We require 5 years of experience programming in one or more programming languages. * We require 3 years of people management experience. * We require 3 years of experience leading projects and working with administration (such as filesystems, inodes, and system calls) or networking (such as TCP/IP, routing, network topologies and hardware, and SDN). * We prefer a masters degree in Computer Science or a related technical field involving coding, such as physics or mathematics. * We prefer a track record of mentoring technical leads. * We prefer proven success leading and influencing multiple technical teams. ## Description * We expect this role to lead a team of software and systems engineers on user-focused projects and be directly responsible for uptime. * We expect this role to own end-to-end availability and performance of key services and build automation to prevent problems from recurring. * We expect this role to automate responses to all non-exceptional service conditions. * We expect this role to lead by example, mentor the team, and establish credibility through quality technical execution. * We expect this role to manage on-call rotations across continents using a follow-the-sun model. * We expect this role to design, write, and deliver software that improves the availability, scalability, latency, and efficiency of our services. Technologies: * Cloud * Hardware * IaaS * Support * Network * TCP/IP More: Site Reliability Engineering combines software and systems engineering to build and operate large-scale, distributed, fault-tolerant systems. At Google, we work to ensure our services have reliability and uptime appropriate to users needs while improving them quickly. We monitor system capacity and performance, optimize existing systems, build infrastructure, and eliminate work through automation. Our SRE culture values intellectual curiosity, problem solving, openness, collaboration, and risk-taking in a blame-free environment. We encourage self-direction on meaningful projects and provide support and mentorship to help people learn and grow. The ML Compute SRE team delivers ML compute infrastructure for all users, ensuring that TPUs and GPUs are supported across our Technical Infrastructure and Cloud Compute platforms and that ML jobs run efficiently, safely, and reliably. We support the hardware and low-level services that provide ML as an IaaS. Our Technical Infrastructure team builds and maintains data centers, networks, and platforms that make Googles products possible. In most instances, we conduct in-person interviews as part of the hiring process. ## Related Videos - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [Creating a routing app with Google Maps API from scratch](https://www.wearedevelopers.com/videos/831-creating-a-routing-app-with-google-maps-api-from-scratch) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Unleash the power of 5G in your code: transform your apps](https://www.wearedevelopers.com/videos/1567-unleash-the-power-of-5g-in-your-code-transform-your-apps) - [A Technical Introduction to Bitcoin's 2nd Layer- The Lightning Network](https://www.wearedevelopers.com/videos/15-a-technical-introduction-to-bitcoin-s-2nd-layer-the-lightning-network) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where to Find Entry-Level Software Engineering Jobs](https://www.wearedevelopers.com/magazine/397-where-to-find-entry-level-software-engineering-jobs) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers)