Senior Software Engineer, SRE, Cloud Incident Response

Google
London, UK
11 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Working hours
Regular working hours
Languages
English

Tech stack

Systems Engineering Cloud Computing Cloud Engineering Program Optimization Data Structures Software Debugging Distributed Systems Fault Tolerance Reliability Engineering Software Engineering Systems Architecture Google Cloud
+2 more
Information Technology Programming Languages

Job description

Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud’s services-both internally critical and externally visible-maintain reliability, uptime appropriate to customer needs, and a rapid rate of improvement. Additionally, SREs monitor system capacity and performance continuously.

Much of our development focuses on optimizing existing systems, building infrastructure, and automating tasks. On the SRE team, you’ll tackle the unique challenges of scale in Google Cloud, leveraging your expertise in coding, algorithms, and large-scale system design. Our culture emphasizes curiosity, problem solving, and openness, fostering collaboration and innovation in a supportive environment., * Ensure Google Cloud Platform (GCP) stability and reliability through incident support, driving customer outcomes, and cross-team collaboration.

  • Create training and processes for incident management, collaborating with Cloud Support leadership.
  • Develop systems and tools to improve incident visibility, issue detection, and communication with stakeholders.
  • Identify and escalate risks, reducing major incident probabilities through pragmatic approaches.
  • Support system scalability and reliability throughout their lifecycle via design consulting, platform development, capacity planning, and automation.

Requirements

  • Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
  • 5 years of experience with software development in one or more programming languages.
  • 5 years of experience with data structures or algorithms.
  • 3 years of experience in designing, analyzing, and troubleshooting distributed systems, and 2 years of experience leading projects and providing technical leadership.
  • Experience in SRE or incident management/response environments., * Experience working in computing, distributed systems, storage, or networking.
  • Experience in telemetry systems, incident and risk management.
  • Experience in designing, analyzing, and troubleshooting large-scale distributed systems.
  • Ability to debug, optimize code, and automate routine tasks.
  • Excellent problem-solving skills, with strong verbal and written communication abilities., As a global company, English proficiency is required for all roles unless otherwise specified.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

6:10 min

Unlocking free learning credits via Google Cloud Innovators

Asrar Asrar · World Congress 2024

2:27 min

Introduction to WebAssembly in a cloud computing context

Edo Edo · World Congress 2024

4:42 min

Building robust data structures with structs and bound functions

Rainer Stropek Rainer Stropek · World Congress 2021

2:06 min

High-paying roles driven by Google Cloud certifications

Asrar Asrar · World Congress 2024

4:42 min

Container hosting options available on Google Cloud Platform

Federico Fregosi · World Congress 2022

Videos

See all

Related articles

See all