> Markdown version of [/jobs/ext/3564378-staff-software-engineer-inference-performance-optimization-genai-deepmind](https://www.wearedevelopers.com/jobs/ext/3564378-staff-software-engineer-inference-performance-optimization-genai-deepmind). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind - **Company:** Google LLC - **Location:** Mountain View, CA, United States - **Experience:** Expert - **Salary:** $207,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Profiling, Computer Engineering, Software Debugging, Distributed Systems, Python (Programming Language), Open Source Technology, Performance Tuning, Software Engineering, Pytorch, Large Language Models, Discretization, Information Technology, SGLang, Codebase, TensorRT, Model Inference, DeepMind - **Published:** October 3, 2026 - **Apply:** https://dejobs.org/x/x/E98D2D14D3424D69BE3A435B3E014064/job/ ## About the Role * Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related technical field, or equivalent practical experience. * 8 years of experience in software development. * Experience in Python and C++, including navigating, debugging, and modifying serving codebases. * Experience with AI model execution constraints, throughput-latency tradeoffs, memory bandwidth limitations, and modern serving architectures., * Experience with real world LLM inference serving environments or direct contributions to modern open-source inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang, Dynamo). * Experience profiling workloads using standard ML profilers (e.g., PyTorch profiler) and internal trace analysis tools. * Experience with observability and reliability for large distributed systems. * Familiarity with GPU/TPU/accelerator performance concepts (e.g. memory bandwidth, quantization, collective communication, kernel), and can reason their implications to the overall inference serving performance. ## Description * Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure layers to methodically increase throughput-per-GPU and reduce latency. * Design and implement inference optimization techniques. * Investigate and resolve complex model inference performance bottlenecks across the stack. * Model the latency-to-cost impacts of system variables (such as batch-sizing and utilization goals) and translate these insights into actionable signals that drive production systems. * Develop investigative tools and metrics (e.g., compute/FLOPs funnels) that track where compute is spent across the fleet. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Fake or News: European Robotaxis & Lawyer GPT - Bramus Van Damme](https://www.wearedevelopers.com/videos/2167-fake-or-news-european-robotaxis-lawyer-gpt-bramus-van-damme) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [What’s New with Google Gemini? ](https://www.wearedevelopers.com/videos/1348-what-s-new-with-google-gemini) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)