> Markdown version of [/jobs/ext/125416-staff-software-engineer-gpu-performance](https://www.wearedevelopers.com/jobs/ext/125416-staff-software-engineer-gpu-performance). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer, GPU Performance - **Company:** Google LLC - **Location:** Sunnyvale, CA, United States - **Experience:** Experienced - **Salary:** $207,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Code Generation, Nvidia CUDA, Data Structures, Machine Learning, Software Architecture, Software Engineering, Large Language Models, Gpu Programming, Information Technology - **Published:** May 14, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=d8796ee005464140 ## About the Role Do you have experience in Solution architecture design?, Do you have a Bachelor's degree?, * Bachelor's degree or equivalent practical experience. * 8 years of experience in software development. * 5 years of experience testing, and launching software products, and 3 years of experience with software design and architecture. * Experience with modern GPU architectures (NVIDIA, AMD, or other AI accelerators), memory hierarchies, and performance bottlenecks. * Experience with modern LLMs and their deployment on AI accelerators. * Experience with low-level GPU programming (CUDA, Triton, CUTLASS, etc.) and performance engineering techniques., * Master's degree or PhD in Engineering, Computer Science, or a related technical field. * 8 years of experience with data structures and algorithms. * 3 years of experience in a technical leadership role leading project teams and setting technical direction. * 3 years of experience working in a structured organization involving cross-functional, or cross-business projects. * Experience with compiler optimization, code generation, and runtime systems for GPU architectures (OpenXLA, MLIR, Triton, etc.). ## Description * Identify and maintain LLM training and serving benchmarks, using them to identify performance opportunities, drive XLA:GPU/Triton performance toward XLA releases. * Engage with various teams, like DeepMind, to solve challenging ML model performance problems. * Run architecture-level simulations on GPU designs and perform roofline analysis to guide partner teams. * Analyze performance and efficiency metrics to identify bottlenecks and then design and implement solutions at Google fleet-wide scale. * Run performance benchmarks on GPU hardware using internal and external tools such as TRT-LLM, vLLM , and SGLang. Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form. ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Leapter: The Reinvention of Software Development? A Future Built On AI Generated Code.](https://www.wearedevelopers.com/videos/1663-leapter-the-reinvention-of-software-development-a-future-built-on-ai-generated-code) - [Phel, a native Lisp for PHP](https://www.wearedevelopers.com/videos/791-phel-a-native-lisp-for-php) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [The weekly developer show: Boosting Python with CUDA, CSS Updates & Navigating New Tech Stacks](https://www.wearedevelopers.com/videos/1293-the-weekly-developer-show-boosting-python-with-cuda-css-updates-navigating-new-tech-stacks) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Software Developer Salary in Germany [2023]](https://www.wearedevelopers.com/magazine/194-software-developer-salary-in-germany-2023) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)