Site Reliability Engineer - Gemini App Product

Azaaki, LLC
Pittsburgh, PA, United States
2 months ago
Apply on indeed.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Compensation
$166,400.0 - $187,200.0
Working hours
Regular working hours
Job source

Tech stack

Architectural Patterns Software Bug Management Software Quality Code Review Computer Programming Data Structures Software Debugging Systems Development Life Cycle Reliability Engineering Software Engineering Software Troubleshooting Synthesizing Data

Job description

As a Site Reliability Engineer (SRE-SWE) within the DeepMind Gemini App Product business unit, you will deliver medium-sized projects from start to finish with minimal supervision. You will work closely with cross-functional teams to build solutions and tools designed to configure, maintain, and scale critical system infrastructures. In this role, you will produce initial project designs, execute high-impact technical blueprints, and proactively address system issues within a large-scale service area., * Production Excellence: Maintain and elevate standards for production excellence within the team, driving the adoption of converged SRE solutions for systems design.

  • Infrastructure & Monitoring: Set up or enhance testing, monitoring, and debugging frameworks to ensure ongoing system scalability, reliability, and efficiency.
  • Software Development: Write and optimize code in internal tools, libraries, and systems with a sharp focus on resilience and toil reduction.
  • Architecture & Design: Review technical designs with an emphasis on simplicity and maintainability, designing solutions to resolve reliability and technology-related incidents.
  • Code Quality: Simplify code utilizing appropriate algorithms, abstractions, and standard shared libraries while conducting meaningful code reviews for your peers.

Requirements

Do you have experience in Software troubleshooting?, * Strong proficiency in programming, data structures, algorithms, and systems thinking.

  • Proven background in the Software Development Lifecycle (SDLC), architecture reliability, and architectural patterns.
  • Demonstrated experience in reliability debugging, bug-fixing, risk analysis, and data synthesis.

Preferred Qualifications

  • Direct professional experience working with ML Serving architectures and workflows.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:39 min

Addressing code review surrender and process exploitation

Laura Tacho Laura Tacho · World Congress 2026 Europe

4:42 min

Building robust data structures with structs and bound functions

Rainer Stropek Rainer Stropek · World Congress 2021

1:48 min

Balancing code generation velocity with software quality standards

Lilia Gargouri Lilia Gargouri · Coffee With Developers

1:54 min

Applying reliability principles to software engineering leadership

Maxim Schepelin Maxim Schepelin · World Congress 2026 Europe

56 sec

The hidden costs of delayed peer code reviews

Tim Gilboy Tim Gilboy

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

Videos

See all

Related articles

See all