Software Engineer, ML Infrastructure

Roku, Inc.
UK
3 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Shift work
Job source

Tech stack

Artificial Intelligence Amazon Web Services Batch Processing Software Quality Continuous Integration Cursor (Graphical User Interface Elements) Software Debugging Distributed Systems Python (Programming Language) Machine Learning Node.Js Search Technologies
+13 more
SQL Databases Management of Software Versions Google Cloud Large Language Models Multi-Agent Systems Apache Spark Kubernetes Information Technology Low Latency Apache Kafka Machine Learning Operations Presto Stream Processing

Job description

We are looking for a Senior Software Engineer with strong machine-learning infrastructure experience to help the Roku Entertainment Assistant team deliver reliable, high-quality conversational experiences at scale. You will own production systems across ranking, model delivery, evaluation and LLM-agent behaviour, working closely with machine-learning, product, data and platform partners.

This is a hands-on engineering role for someone who can move comfortably between software architecture, ML lifecycle decisions and production operations. You will improve the paths the team uses to train, evaluate, deploy and observe models, while helping us manage latency, quality and cost as the product grows.

How will I use AI at Roku?

At Roku, we don’t just use AI, we work with it. AI agents and smart tools help power drafts, analysis, and repetitive workflows, while our people bring direction, judgment, and accountability. We’re looking for curious, adaptable builders who can show how they’ve used AI or automation to move faster, raise the bar, and scale their impact.

We value your AI skills if you have built fluency across the agentic engineering toolchain - coding harnesses like Claude Code or Cursor, MCP servers, custom skills, or agent frameworks. And you can describe projects where you shipped real work with these tools. You know how to drive an agent, verify its output, and ramp on an unfamiliar codebase with an agent helping you.

What will you be doing?

  • Own fulfilment-ranking pipelines from training orchestration and quality gates through deployment and online validation.
  • Build and operate an offline-evaluation platform, combining LLM-as-judge harnesses with deterministic checks for answer quality.
  • Develop the Roku Entertainment Assistant’s LLM agent, including tool routing, retrieval, guardrails and answer caching.
  • Design caching as an intentional latency and cost lever for high-volume services.
  • Build observability for ML and LLM systems, including latency attribution, quality metrics, tracing and per-request cost.
  • Improve the reliability and operability of distributed systems, and lead the diagnosis of difficult cross-service failures.
  • Drive accountable AI-assisted engineering practices across the team, building on the agents already used in release workflows.
  • Work across machine-learning, product, data and platform boundaries to turn ambiguous problems into measurable delivery outcomes., Roku fosters an inclusive and collaborative environment where teams generally work in the office Monday through Thursday. Fridays are generally flexible for remote work, except for employees whose specific roles or assigned office location require five days’ a week attendance.

Requirements

  • Strong production software-engineering experience, including designing, testing, operating and debugging distributed services.
  • Meaningful ownership of ML infrastructure or production model delivery, with clear evidence of your contribution and its operational or adoption impact.
  • A practical understanding of the ML lifecycle: training, evaluation, versioning, deployment, online validation, monitoring and safe rollback.
  • Experience designing for latency, resilience, observability and cost in production systems.
  • Experience with LLMs, embeddings, semantic search, retrieval or agent systems, including how you evaluate quality and control failure modes.
  • Experience with caching systems and the trade-offs involved in using caching for low-latency services.
  • A commitment to automation, CI/CD, code quality and evidence-led engineering decisions.
  • A collaborative approach and the ability to explain technical trade-offs to partners across disciplines.
  • A degree in Computer Science, Electrical Engineering or a related field, or equivalent practical experience.

Experience with AWS or GCP, Kubernetes, SQL warehouse and batch-processing technologies such as Trino, Presto or Spark, Kafka or other streaming systems, and Go or Python would be useful. Node.js experience is a plus. We value transferable evidence over an exact match to every named technology.

About the company

Roku is a great place for people who want to work in a fast-paced environment where everyone is focused on the company’s success rather than their own. We try to surround ourselves with people who are great at their jobs, who are easy to work with, and who keep their egos in check. We appreciate a sense of humor. We believe a fewer number of very talented folks can do more for less cost than a larger number of less talented teams. We’re independent thinkers with big ideas who act boldly, move fast and accomplish extraordinary things through collaboration and trust. In short, at Roku you’ll be part of a company that’s changing how the world watches TV.

We have a unique culture that we are proud of. We think of ourselves primarily as problem-solvers, which itself is a two-part idea. We come up with the solution, but the solution isn’t real until it is built and delivered to the customer. That penchant for action gives us a pragmatic approach to innovation, one that has served us well since 2002.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · World Congress 2023

2:37 min

Optimizing technical profiles for AI sourcing and recruitment

Mina Golesorkhi Mina Golesorkhi · World Congress 2026 Europe

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · World Congress 2022

3:55 min

Identifying underlying Node.js runtime vulnerabilities using fuzzing tools

Sonya Moisset · World Congress 2023

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all