> Markdown version of [/jobs/ext/1173583-lead-machine-learning-inference-engineer](https://www.wearedevelopers.com/jobs/ext/1173583-lead-machine-learning-inference-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Machine Learning Inference Engineer - **Company:** Roku, Inc. - **Location:** San Jose, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $246,500.0 - **Contract:** Permanent contract - **Skills:** Distributed Systems, Machine Learning, Open Source Technology, High Performance Computing, Model Validation, Low Latency, Hardware Acceleration, Machine Learning Operations - **Published:** July 3, 2026 - **Apply:** https://www.weareroku.com/jobs/lead-machine-learning-inference-engineer-advertising-san-jose-california-united-states ## About the Role * M.S. or above in CS, ECE, or a related field * 10+ years of experience in developing and deploying large-scale, distributed systems, with at least 5 years in a leadership or technical lead role * Strong programming skills in high-performance languages * Deep understanding of inference frameworks and ML system deployment * Proven experience optimizing performance for large-scale machine learning systems, including a deep knowledge of SOTA model optimizations, hardware-software co-design, GPU acceleration, and HPC techniques * Excellent communication and collaboration skills * Experience leading teams working on high-throughput, low-latency ML serving systems * Experience collaborating with and leading global, cross-functional teams * Contributions to open-source ML or systems projects ## Description In this role, you will architect, design, and lead the development of a SOTA Inference platform that can handle Advertising-level low latencies, scale, throughput, and availability with optimizations that span across hardware, software, and models. We're looking for a strong technical leader with deep experience in ML serving, high-performance computing, and industry standard frameworks - someone excited to mentor engineers, innovate at scale, and shape the future of machine learning at Roku., * Lead the design and development of a SOTA Inference platform * Oversee the development of monitoring, observability, and other tooling to ensure system and model performance, reliability, and scalability of online inference services * Identify and resolve system inefficiencies, performance bottlenecks, and reliability issues, ensuring optimized end-to-end performance * Stay at the forefront of advancements in inference frameworks, ML hardware acceleration, and distributed systems, and incorporate innovations where and when they are impactful ## Related Videos - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [Introduction to Responsible AI: Balancing Value and Risk](https://www.wearedevelopers.com/videos/1972-introduction-to-responsible-ai-balancing-value-and-risk) - [How Machine Learning is turning the Automotive Industry upside down](https://www.wearedevelopers.com/videos/61-how-machine-learning-is-turning-the-automotive-industry-upside-down) - [What non-automotive Machine Learning projects can learn from automotive Machine Learning projects](https://www.wearedevelopers.com/videos/397-what-non-automotive-machine-learning-projects-can-learn-from-automotive-machine-learning-projects) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)