> Markdown version of [/jobs/ext/3037469-software-engineer-model-inference-deepmind](https://www.wearedevelopers.com/jobs/ext/3037469-software-engineer-model-inference-deepmind). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Model Inference, DeepMind - **Company:** Google LLC - **Location:** Mountain View, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Bioinformatics, Profiling, Nvidia CUDA, Systems Analysis, Machine Learning, OpenCL, Tensorflow, Software Engineering, System Programming, Graphics Processing Unit (GPU), Pytorch, Large Language Models, Hardware Acceleration, Machine Learning Operations - **Published:** September 23, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3401611078&tx=TT4845TTD&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * Bachelor's degree or equivalent practical experience. * 8 years of experience in software development. * 2 years of experience in deploying and maintaining machine learning models in a live production environment. * Experience in profiling, configuring, or executing ML workloads directly on hardware accelerators (e.g., GPU or TPU). * Experience designing, building, or optimizing model serving infrastructure or inference backends. Preferred qualifications: * Experience with developing serving infrastructure. * Experience programming hardware accelerators (GPUs, TPUs) via ML frameworks (e.g., JAX, PyTorch) or low-level programming models (e.g., Pallas, CUDA, OpenCL). * Experience profiling software to identify performance bottlenecks. * Experience with distributed ML systems optimization and parallelism (e.g., data, model, or pipeline parallelism). * Familiarity with writing performance-optimized kernels. * Understanding of LLM architecture and inference performance dynamics (e.g., Transformer models, memory bandwidth and compute bounds, KV cache scaling). ## Description * Collaborate closely with Research teams to understand next generation modeling approaches, ensuring they are designed and implemented with production considerations in mind. * Work with infrastructure teams to deliver serving infrastructure that is designed for maximum efficiency and performance, addressing bottlenecks in speed, scale, and quality. * Identify opportunities to automate tasks, eliminate redundancies, build performant tests, and improve the overall velocity of model releases. * Gain a deep understanding of serving frameworks, pre-processing pipelines, caching mechanisms, and other relevant technologies. * Leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers (XLA), custom kernels (Pallas), and serving infrastructure on hardware accelerators (TPUs/GPUs). Information collected and processed as part of your Google Careers profile, and any job applications you choose to submit is subject to Google'sApplicant and Candidate Privacy Policy (./privacy-policy) . Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents-to-be, criminal histories consistent with legal requirements, or any other basis protected by law. See alsoGoogle's EEO Policy (https://www.google.com/about/careers/applications/eeo/) ,Know your rights: workplace discrimination is illegal (https://careers.google.com/jobs/dist/legal/EEOC_KnowYourRights_10_20.pdf) ,Belonging at Google (https://about.google/belonging/) , andHow we hire (https://careers.google.com/how-we-hire/) . If you have a need that requires accommodation, please let us know by completing ourAccommodations for Applicants form (https://goo.gl/forms/aBt6Pu71i1kzpLHe2) . Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting. To all recruitment agencies: Google does not accept agency resumes. Please do not forward resumes to our jobs alias, Google employees, or any other organization location. Google is not responsible for any fees related to unsolicited resumes. Equity is granted exclusively and discretionarily by Alphabet Inc. on the basis of an agreement concluded between you and Alphabet Inc. Alphabet Inc. is your sole contractual partner with respect to equity grants. GSU grants are not guaranteed, are discretionary, are subject to approval by the Alphabet Inc. board of directors or its delegate, the terms of the relevant Alphabet Inc. stock plan, and your grant agreement. They have no impact on statutory payments. Current or past grants do not confer an acquired right. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)