Software Engineering Manager, Fault Tolerance Testing

Google LLC
Kirkland, United States of America
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Intermediate
Compensation
$ 301K

Job location

Kirkland, United States of America

Tech stack

API
Artificial Intelligence
Code Review
Data Compression
Disaster Recovery
Distributed Systems
Fault Tolerance
Design of User Interfaces
Information Retrieval
Natural Language Processing
Systems Development Life Cycle
Reliability Engineering
Software Engineering
Systems Architecture
Google Cloud Platform
Information Technology
gRPC

Job description

Like Google's own ambitions, the work of a Software Engineer goes beyond just Search. Software Engineering Managers have not only the technical expertise to take on and provide technical leadership to major projects, but also manage a team of Engineers. You not only optimize your own code but make sure Engineers are able to optimize theirs. As a Software Engineering Manager you manage your project goals, contribute to product strategy and help develop your team. Teams work all across the company, in areas such as information retrieval, artificial intelligence, natural language processing, distributed computing, large-scale system design, networking, security, data compression, user interface design; the list goes on and is growing every day. Operating with scale and speed, our exceptional software engineers are just getting started -- and as a manager, you guide the way.

With technical and leadership expertise, you manage engineers across multiple teams and locations, a large product budget and oversee the deployment of large-scale projects across multiple sites internationally.

In this role, you will drive the technical goal to transition Fault Tolerance Testing (FTT) from a set of manual compliance verification tools to a proactive, AI-driven, and autonomous resilience platform. This position sits at the intersection of large-scale distributed systems, developer velocity, and cloud reliability, offering immense visibility and the opportunity to directly safeguard Google Cloud Platform's (GCP) global infrastructure.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $207000 - $301000 (USD) + 20% bonus target + equity + benefits, * Set and communicate team priorities that support the broader organization's goals. Align strategy, processes, and decision-making across teams.

  • Set clear expectations with individuals based on their level and role and aligned to the broader organization's goals. Meet regularly with individuals to discuss performance and development and provide feedback and coaching.
  • Develop the mid-term technical goal and roadmap within the scope of your (often multiple) team(s). Evolve the roadmap to meet anticipated future requirements and infrastructure needs.
  • Design, guide and vet systems designs within the scope of the broader area, and write product or system development code to solve ambiguous problems.
  • Review code developed by other engineers and provide feedback to ensure best practices (e.g., style guidelines, checking code in, accuracy, testability, and efficiency).

Requirements

Experience owning outcomes and decision making, solving ambiguous problems and influencing stakeholders; deep expertise in domain., * Bachelor's degree, or equivalent practical experience.

  • 8 years of experience in software development.
  • 5 years of experience creating product roadmaps, and working with cross-functional teams.
  • 3 years of experience in reliability engineering.
  • 3 years of experience in a technical leadership role.
  • 2 years of experience in a people management or team leadership role., * Master's degree or PhD in Computer Science or related technical field.
  • Experience designing, building, or operating highly available, fault-tolerant distributed systems. Direct experience with chaos engineering, fault-injection testing frameworks, or large-scale disaster recovery simulations.
  • Background in designing developer-facing products, APIs, SDKs, or self-service automation tools, with a strong emphasis on reducing friction and improving developer velocity.
  • Experience with Google's server frameworks (Pod), gRPC/Stubby-based RPC layers, or container orchestrators is highly desirable.
  • Experience defining and driving organizational key metrics (SLIs/SLOs, adoption rates, platform health) to measure the success of infrastructure initiatives.

Benefits & conditions

XIn accordance with Washington state law, we are highlighting our comprehensive benefits package, which is available to all eligible US based employees.

Benefits for this role include:

  • Health, dental, vision, life, disability insurance
  • Retirement Benefits: 401(k) with company match
  • Paid Time Off: 20 days of vacation per year, accruing at a rate of 6.15 hours per pay period for the first five years of employment
  • Sick Time: 40 hours/year (increased to 69 hours/year for Seattle) including 5 discretionary sick days per instance
  • Maternity Leave (Short-Term Disability + Baby Bonding): 28-30 weeks
  • Baby Bonding Leave: 18 weeks
  • Holidays: 13 paid days per year

About the company

Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents-to-be, criminal histories consistent with legal requirements, or any other basis protected by law. See alsoGoogle's EEO Policy (https://www.google.com/about/careers/applications/eeo/) ,Know your rights: workplace discrimination is illegal (https://careers.google.com/jobs/dist/legal/EEOC_KnowYourRights_10_20.pdf) ,Belonging at Google (https://about.google/belonging/) , andHow we hire (https://careers.google.com/how-we-hire/) . If you have a need that requires accommodation, please let us know by completing ourAccommodations for Applicants form (https://goo.gl/forms/aBt6Pu71i1kzpLHe2) . Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting. To all recruitment agencies: Google does not accept agency resumes. Please do not forward resumes to our jobs alias, Google employees, or any other organization location. Google is not responsible for any fees related to unsolicited resumes.

Apply for this position