> Markdown version of [/jobs/ext/2916870-software-engineer-ml-infrastructure](https://www.wearedevelopers.com/jobs/ext/2916870-software-engineer-ml-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, ML Infrastructure - **Company:** Roku, Inc. - **Location:** UK (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Batch Processing, Software Quality, Continuous Integration, Cursor (Graphical User Interface Elements), Software Debugging, Distributed Systems, Python (Programming Language), Machine Learning, Node.Js, Search Technologies, SQL Databases, Management of Software Versions, Google Cloud, Large Language Models, Multi-Agent Systems, Apache Spark, Kubernetes, Information Technology, Low Latency, Apache Kafka, Machine Learning Operations, Presto, Stream Processing - **Published:** September 15, 2026 - **Apply:** https://startup.jobs/senior-software-engineer-ml-infrastructure-roku-10068994 ## About the Role * Strong production software-engineering experience, including designing, testing, operating and debugging distributed services. * Meaningful ownership of ML infrastructure or production model delivery, with clear evidence of your contribution and its operational or adoption impact. * A practical understanding of the ML lifecycle: training, evaluation, versioning, deployment, online validation, monitoring and safe rollback. * Experience designing for latency, resilience, observability and cost in production systems. * Experience with LLMs, embeddings, semantic search, retrieval or agent systems, including how you evaluate quality and control failure modes. * Experience with caching systems and the trade-offs involved in using caching for low-latency services. * A commitment to automation, CI/CD, code quality and evidence-led engineering decisions. * A collaborative approach and the ability to explain technical trade-offs to partners across disciplines. * A degree in Computer Science, Electrical Engineering or a related field, or equivalent practical experience. Experience with AWS or GCP, Kubernetes, SQL warehouse and batch-processing technologies such as Trino, Presto or Spark, Kafka or other streaming systems, and Go or Python would be useful. Node.js experience is a plus. We value transferable evidence over an exact match to every named technology. ## Description We are looking for a Senior Software Engineer with strong machine-learning infrastructure experience to help the Roku Entertainment Assistant team deliver reliable, high-quality conversational experiences at scale. You will own production systems across ranking, model delivery, evaluation and LLM-agent behaviour, working closely with machine-learning, product, data and platform partners. This is a hands-on engineering role for someone who can move comfortably between software architecture, ML lifecycle decisions and production operations. You will improve the paths the team uses to train, evaluate, deploy and observe models, while helping us manage latency, quality and cost as the product grows. How will I use AI at Roku? At Roku, we don't just use AI, we work with it. AI agents and smart tools help power drafts, analysis, and repetitive workflows, while our people bring direction, judgment, and accountability. We're looking for curious, adaptable builders who can show how they've used AI or automation to move faster, raise the bar, and scale their impact. We value your AI skills if you have built fluency across the agentic engineering toolchain - coding harnesses like Claude Code or Cursor, MCP servers, custom skills, or agent frameworks. And you can describe projects where you shipped real work with these tools. You know how to drive an agent, verify its output, and ramp on an unfamiliar codebase with an agent helping you. What will you be doing? * Own fulfilment-ranking pipelines from training orchestration and quality gates through deployment and online validation. * Build and operate an offline-evaluation platform, combining LLM-as-judge harnesses with deterministic checks for answer quality. * Develop the Roku Entertainment Assistant's LLM agent, including tool routing, retrieval, guardrails and answer caching. * Design caching as an intentional latency and cost lever for high-volume services. * Build observability for ML and LLM systems, including latency attribution, quality metrics, tracing and per-request cost. * Improve the reliability and operability of distributed systems, and lead the diagnosis of difficult cross-service failures. * Drive accountable AI-assisted engineering practices across the team, building on the agents already used in release workflows. * Work across machine-learning, product, data and platform boundaries to turn ambiguous problems into measurable delivery outcomes., Roku fosters an inclusive and collaborative environment where teams generally work in the office Monday through Thursday. Fridays are generally flexible for remote work, except for employees whose specific roles or assigned office location require five days' a week attendance. ## Related Videos - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [The state of MLOps - machine learning in production at enterprise scale](https://www.wearedevelopers.com/videos/369-the-state-of-mlops-machine-learning-in-production-at-enterprise-scale) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [Stop Using Node.js Like It’s 2020! - Alfonso Graziano](https://www.wearedevelopers.com/videos/1863-stop-using-node-js-like-it-s-2020-alfonso-graziano) - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)