> Markdown version of [/jobs/ext/2634254-software-engineer-model-inference-deepmind](https://www.wearedevelopers.com/jobs/ext/2634254-software-engineer-model-inference-deepmind). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Model Inference, DeepMind - **Company:** DEEPMIND MATRIX - **Location:** Mountain View, CA, United States - **Experience:** Experienced - **Salary:** $207,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Profiling, Nvidia CUDA, Systems Analysis, Machine Learning, OpenCL, Tensorflow, Software Engineering, System Programming, Graphics Processing Unit (GPU), Pytorch, Large Language Models, Hardware Acceleration, Machine Learning Operations - **Published:** August 11, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=af78a2172b782398 ## About the Role * Bachelor's degree or equivalent practical experience. * 8 years of experience in software development. * 2 years of experience in deploying and maintaining machine learning models in a live production environment. * Experience in profiling, configuring, or executing ML workloads directly on hardware accelerators (e.g., GPU or TPU). * Experience designing, building, or optimizing model serving infrastructure or inference backends., * Experience with developing serving infrastructure. * Experience programming hardware accelerators (GPUs, TPUs) via ML frameworks (e.g., JAX, PyTorch) or low-level programming models (e.g., Pallas, CUDA, OpenCL). * Experience profiling software to identify performance bottlenecks. * Experience with distributed ML systems optimization and parallelism (e.g., data, model, or pipeline parallelism). * Familiarity with writing performance-optimized kernels. * Understanding of LLM architecture and inference performance dynamics (e.g., Transformer models, memory bandwidth and compute bounds, KV cache scaling). ## Description * Collaborate closely with Research teams to understand next generation modeling approaches, ensuring they are designed and implemented with production considerations in mind. * Work with infrastructure teams to deliver serving infrastructure that is designed for maximum efficiency and performance, addressing bottlenecks in speed, scale, and quality. * Identify opportunities to automate tasks, eliminate redundancies, build performant tests, and improve the overall velocity of model releases. * Gain a deep understanding of serving frameworks, pre-processing pipelines, caching mechanisms, and other relevant technologies. * Leverage roofline analysis, hardware-level profiling, and systems analysis to identify and eliminate performance bottlenecks across ML frameworks, compilers (XLA), custom kernels (Pallas), and serving infrastructure on hardware accelerators (TPUs/GPUs). Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form. ## Related Videos - [30 Golden Rules of Deep Learning Performance](https://www.wearedevelopers.com/videos/11-30-golden-rules-of-deep-learning-performance) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)