> Markdown version of [/jobs/ext/3341301-staff-software-engineer-site-reliability-engineering-google-cloud](https://www.wearedevelopers.com/jobs/ext/3341301-staff-software-engineer-site-reliability-engineering-google-cloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer, Site Reliability Engineering, Google Cloud - **Company:** Google - **Location:** München, Germany - **Experience:** Experienced - **Salary:** €156,000.0 - €159,000.0 - **Contract:** Permanent contract - **Skills:** Software Code Optimization, Software Debugging, Distributed Systems, Reliability Engineering, Software Engineering, Google Cloud, Information Technology, Programming Languages - **Published:** September 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=e455b4673d6f892a ## About the Role * Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience. * 8 years of experience with software development in one or more programming languages. * 3 years of experience leading projects. * 3 years of experience designing, analyzing, and troubleshooting distributed systems., * Experience working in computing, distributed systems, storage, or networking. * Experience effectively, efficiently, and responsibly applying AI tooling and workflows to engineering practices. * Expertise in designing, analyzing, and troubleshooting large-scale distributed systems. * Ability to debug, optimize code, and to automate routine tasks. * Systematic problem-solving approach, coupled with effective verbal and written communication skills. ## Description * Engage in and improve the whole lifecycle of services-from inception and design, through to deployment, operation, and refinement. * Support services before they go live through activities such as system design consulting, developing software platforms and frameworks, capacity planning, and launch reviews. * Maintain services once they are live by measuring and monitoring availability, latency, and overall system health. * Scale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve reliability and velocity. * Practice sustainable incident response and blameless postmortems.