Senior Software Engineer, SRE, Cloud Incident Response
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+2 more
Job description
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud’s services-both internally critical and externally visible-maintain reliability, uptime appropriate to customer needs, and a rapid rate of improvement. Additionally, SREs monitor system capacity and performance continuously.
Much of our development focuses on optimizing existing systems, building infrastructure, and automating tasks. On the SRE team, you’ll tackle the unique challenges of scale in Google Cloud, leveraging your expertise in coding, algorithms, and large-scale system design. Our culture emphasizes curiosity, problem solving, and openness, fostering collaboration and innovation in a supportive environment., * Ensure Google Cloud Platform (GCP) stability and reliability through incident support, driving customer outcomes, and cross-team collaboration.
- Create training and processes for incident management, collaborating with Cloud Support leadership.
- Develop systems and tools to improve incident visibility, issue detection, and communication with stakeholders.
- Identify and escalate risks, reducing major incident probabilities through pragmatic approaches.
- Support system scalability and reliability throughout their lifecycle via design consulting, platform development, capacity planning, and automation.
Requirements
- Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
- 5 years of experience with software development in one or more programming languages.
- 5 years of experience with data structures or algorithms.
- 3 years of experience in designing, analyzing, and troubleshooting distributed systems, and 2 years of experience leading projects and providing technical leadership.
- Experience in SRE or incident management/response environments., * Experience working in computing, distributed systems, storage, or networking.
- Experience in telemetry systems, incident and risk management.
- Experience in designing, analyzing, and troubleshooting large-scale distributed systems.
- Ability to debug, optimize code, and automate routine tasks.
- Excellent problem-solving skills, with strong verbal and written communication abilities., As a global company, English proficiency is required for all roles unless otherwise specified.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Data Engineer Salary UK
How Much Does a Software Engineer Make? Realistic Software Engineering Salaries
Is Software Engineering Over-Saturated?
Best Countries for Software Engineers