Systems Engineering Manager, SRE, ML Compute

Hackajob Ltd
Leeds, UK
2 days ago
Apply on www.adzuna.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
£61,000.0 - £101,000.0
Working hours
Regular working hours

Tech stack

Systems Engineering Build Automation Cloud Computing Data Centers File Systems Fault Tolerance Network Topologies Infrastructure as a Service (IaaS) Routing Reliability Engineering System Programming TCP/IP
+4 more
Graphics Processing Unit (GPU) Information Technology Low Latency Programming Languages

Job description

  • We expect this role to lead a team of software and systems engineers on user-focused projects and be directly responsible for uptime.
  • We expect this role to own end-to-end availability and performance of key services and build automation to prevent problems from recurring.
  • We expect this role to automate responses to all non-exceptional service conditions.
  • We expect this role to lead by example, mentor the team, and establish credibility through quality technical execution.
  • We expect this role to manage on-call rotations across continents using a follow-the-sun model.
  • We expect this role to design, write, and deliver software that improves the availability, scalability, latency, and efficiency of our services.

Technologies:

  • Cloud
  • Hardware
  • IaaS
  • Support
  • Network
  • TCP/IP

More:

Site Reliability Engineering combines software and systems engineering to build and operate large-scale, distributed, fault-tolerant systems. At Google, we work to ensure our services have reliability and uptime appropriate to users needs while improving them quickly. We monitor system capacity and performance, optimize existing systems, build infrastructure, and eliminate work through automation. Our SRE culture values intellectual curiosity, problem solving, openness, collaboration, and risk-taking in a blame-free environment. We encourage self-direction on meaningful projects and provide support and mentorship to help people learn and grow. The ML Compute SRE team delivers ML compute infrastructure for all users, ensuring that TPUs and GPUs are supported across our Technical Infrastructure and Cloud Compute platforms and that ML jobs run efficiently, safely, and reliably. We support the hardware and low-level services that provide ML as an IaaS. Our Technical Infrastructure team builds and maintains data centers, networks, and platforms that make Googles products possible. In most instances, we conduct in-person interviews as part of the hiring process.

Requirements

  • We require a bachelors degree in Computer Science or a related technical field, or equivalent practical experience.
  • We require 5 years of experience programming in one or more programming languages.
  • We require 3 years of people management experience.
  • We require 3 years of experience leading projects and working with administration (such as filesystems, inodes, and system calls) or networking (such as TCP/IP, routing, network topologies and hardware, and SDN).
  • We prefer a masters degree in Computer Science or a related technical field involving coding, such as physics or mathematics.
  • We prefer a track record of mentoring technical leads.
  • We prefer proven success leading and influencing multiple technical teams.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

5:48 min

Balancing delivery latency with stream reliability and scale

Phil Cluff · LIVE

2:04 min

Enhancing network privacy with routing fees and onion routing

Andreas M Antonopoulos · LIVE

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

3:37 min

Accessing API documentation and testing remote driving latency

Alexandru Ciinaru Alexandru Ciinaru +3 · World Congress 2025

1:51 min

Overview of the three Google Maps routing applications

Germán Álvarez · LIVE

Videos

See all

Related articles

See all