Systems Engineering Manager, SRE, ML Compute
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+4 more
Job description
- We expect this role to lead a team of software and systems engineers on user-focused projects and be directly responsible for uptime.
- We expect this role to own end-to-end availability and performance of key services and build automation to prevent problems from recurring.
- We expect this role to automate responses to all non-exceptional service conditions.
- We expect this role to lead by example, mentor the team, and establish credibility through quality technical execution.
- We expect this role to manage on-call rotations across continents using a follow-the-sun model.
- We expect this role to design, write, and deliver software that improves the availability, scalability, latency, and efficiency of our services.
Technologies:
- Cloud
- Hardware
- IaaS
- Support
- Network
- TCP/IP
More:
Site Reliability Engineering combines software and systems engineering to build and operate large-scale, distributed, fault-tolerant systems. At Google, we work to ensure our services have reliability and uptime appropriate to users needs while improving them quickly. We monitor system capacity and performance, optimize existing systems, build infrastructure, and eliminate work through automation. Our SRE culture values intellectual curiosity, problem solving, openness, collaboration, and risk-taking in a blame-free environment. We encourage self-direction on meaningful projects and provide support and mentorship to help people learn and grow. The ML Compute SRE team delivers ML compute infrastructure for all users, ensuring that TPUs and GPUs are supported across our Technical Infrastructure and Cloud Compute platforms and that ML jobs run efficiently, safely, and reliably. We support the hardware and low-level services that provide ML as an IaaS. Our Technical Infrastructure team builds and maintains data centers, networks, and platforms that make Googles products possible. In most instances, we conduct in-person interviews as part of the hiring process.
Requirements
- We require a bachelors degree in Computer Science or a related technical field, or equivalent practical experience.
- We require 5 years of experience programming in one or more programming languages.
- We require 3 years of people management experience.
- We require 3 years of experience leading projects and working with administration (such as filesystems, inodes, and system calls) or networking (such as TCP/IP, routing, network topologies and hardware, and SDN).
- We prefer a masters degree in Computer Science or a related technical field involving coding, such as physics or mathematics.
- We prefer a track record of mentoring technical leads.
- We prefer proven success leading and influencing multiple technical teams.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Where to Find Entry-Level Software Engineering Jobs
The Most Popular IT Jobs on the Market
Top-Paying Tech Jobs (with Salaries)
Why Upskilling And Reskilling is Important For Developers