> Markdown version of [/jobs/ext/2545460-staff-software-engineer-inference-performance-optimization-genai-deepmind](https://www.wearedevelopers.com/jobs/ext/2545460-staff-software-engineer-inference-performance-optimization-genai-deepmind). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind - **Company:** Google LLC - **Location:** Mountain View, CA, United States - **Experience:** Expert - **Salary:** $207,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Profiling, Computer Engineering, Software Debugging, Distributed Systems, Python (Programming Language), Open Source Technology, Performance Tuning, Software Engineering, Pytorch, Large Language Models, Information Technology, Codebase, TensorRT - **Published:** August 21, 2026 - **Apply:** https://dejobs.org/x/x/F152BC85580F4EE19478866B219E4D63/job/ ## About the Role * Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related technical field, or equivalent practical experience. * 8 years of experience in software development. * Experience in Python and C++, including navigating, debugging, and modifying serving codebases. * Experience with AI model execution constraints, throughput-latency tradeoffs, memory bandwidth limitations, and modern serving architectures., * Experience with real world LLM inference serving environments or direct contributions to modern open-source inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang, Dynamo). * Experience profiling workloads using standard ML profilers (e.g., PyTorch profiler) and internal trace analysis tools. * Experience with observability and reliability for large distributed systems. * Familiarity with GPU/TPU/accelerator performance concepts (e.g. memory bandwidth, quantization, collective communication, kernel), and can reason their implications to the overall inference serving performance. ## Description * Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure layers to methodically increase throughput-per-GPU and reduce latency. * Design and implement inference optimization techniques. * Investigate and resolve complex model inference performance bottlenecks across the stack. * Model the latency-to-cost impacts of system variables (such as batch-sizing and utilization goals) and translate these insights into actionable signals that drive production systems. * Develop investigative tools and metrics (e.g., compute/FLOPs funnels) that track where compute is spent across the fleet. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Getting to Know Your Legacy (System) with AI-Driven Software Archeology](https://www.wearedevelopers.com/videos/1437-getting-to-know-your-legacy-system-with-ai-driven-software-archeology) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Why your codebase lies to AI?](https://www.wearedevelopers.com/videos/100281-why-your-codebase-lies-to-ai) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)