> Markdown version of [/jobs/ext/2027719-senior-software-engineer-google-distributed-cloud-ai](https://www.wearedevelopers.com/jobs/ext/2027719-senior-software-engineer-google-distributed-cloud-ai). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer, Google Distributed Cloud AI - **Company:** Google LLC - **Location:** Sunnyvale, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cloud Computing, Computer Programming, Extract Transform Load (ETL), Distributed Systems, Machine Learning, Performance Tuning, Software Architecture, Software Engineering, Data Logging, Load Balancing, Large Language Models, Kubernetes, Hardware Acceleration, GPT - **Published:** August 11, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/senior-software-engineer-google-distributed-cloud-ai-sunnyvale-ca-usa-58901894 ## About the Role Google for disaggregated serving, speculative decoding, quantization, and model sharding across distributed hardware * Collaborate with teams building LLM frameworks, Kubernetes/GKE, networking, and hardware acceleration to deliver a cohesive platform Tasks * Bachelor's degree or equivalent practical experience * 5 years of software development experience * 3 years of testing, maintaining, or launching software products * 1 year of software design and architecture experience * 3 years of experience with large-scale infrastructure, distributed systems, or related compute/storage knowledge * Experience programming in Go for software development including AI/ML applications Key requirements * bonus target * equity * benefits * competitive compensation ## Description Experteer Overview In this role you will shape the LLM inference serving layer of Google Distributed Cloud (GDC), driving the design, development, and performance optimization of critical components. You will work with cross-functional teams to ensure scalable model lifecycle, data loading, request routing, and intelligent load balancing. You will advance serving capabilities for advanced LLM techniques and collaborate with core teams on Kubernetes, networking, and hardware acceleration. This role offers impact across the platform, from billing and observability to security and quota management, in a fast-paced, edge-to-hybrid cloud environment. Compensation / Benefits * Lead the design, development, and optimization of LLM inference serving components on GDC * Oversee model lifecycle management, efficient data loading, dynamic request routing, and load balancing * Drive horizontal integration with billing, logging, observability, security, and quota management * Enhance serving capabilities for disaggregated serving, speculative decoding, quantization, and model sharding across distributed hardware * Collaborate with teams building LLM frameworks, Kubernetes/GKE, networking, and hardware acceleration to deliver a cohesive platform Tasks * Bachelor's degree or equivalent practical experience * 5 years of software development experience * 3 years of testing, maintaining, or launching software products * 1 year of software design and architecture experience * 3 years of experience with large-scale infrastructure, distributed systems, or related compute/storage knowledge * Experience programming in Go for software development including AI/ML applications Key requirements * bonus target * equity * benefits * competitive compensation ## Related Videos - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Crypto-secure Data Management with In-Database Blockchain](https://www.wearedevelopers.com/videos/632-crypto-secure-data-management-with-in-database-blockchain) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Speak, Code, Deploy: Transforming Developer Experience with Voice Commands](https://www.wearedevelopers.com/videos/1159-speak-code-deploy-transforming-developer-experience-with-voice-commands) - [The Cloud is Calling: Answer with In-Demand Skills](https://www.wearedevelopers.com/videos/945-the-cloud-is-calling-answer-with-in-demand-skills) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)