Software Engineer, ML Infrastructure
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+13 more
Job description
We are looking for a Senior Software Engineer with strong machine-learning infrastructure experience to help the Roku Entertainment Assistant team deliver reliable, high-quality conversational experiences at scale. You will own production systems across ranking, model delivery, evaluation and LLM-agent behaviour, working closely with machine-learning, product, data and platform partners.
This is a hands-on engineering role for someone who can move comfortably between software architecture, ML lifecycle decisions and production operations. You will improve the paths the team uses to train, evaluate, deploy and observe models, while helping us manage latency, quality and cost as the product grows.
How will I use AI at Roku?
At Roku, we don’t just use AI, we work with it. AI agents and smart tools help power drafts, analysis, and repetitive workflows, while our people bring direction, judgment, and accountability. We’re looking for curious, adaptable builders who can show how they’ve used AI or automation to move faster, raise the bar, and scale their impact.
We value your AI skills if you have built fluency across the agentic engineering toolchain - coding harnesses like Claude Code or Cursor, MCP servers, custom skills, or agent frameworks. And you can describe projects where you shipped real work with these tools. You know how to drive an agent, verify its output, and ramp on an unfamiliar codebase with an agent helping you.
What will you be doing?
- Own fulfilment-ranking pipelines from training orchestration and quality gates through deployment and online validation.
- Build and operate an offline-evaluation platform, combining LLM-as-judge harnesses with deterministic checks for answer quality.
- Develop the Roku Entertainment Assistant’s LLM agent, including tool routing, retrieval, guardrails and answer caching.
- Design caching as an intentional latency and cost lever for high-volume services.
- Build observability for ML and LLM systems, including latency attribution, quality metrics, tracing and per-request cost.
- Improve the reliability and operability of distributed systems, and lead the diagnosis of difficult cross-service failures.
- Drive accountable AI-assisted engineering practices across the team, building on the agents already used in release workflows.
- Work across machine-learning, product, data and platform boundaries to turn ambiguous problems into measurable delivery outcomes., Roku fosters an inclusive and collaborative environment where teams generally work in the office Monday through Thursday. Fridays are generally flexible for remote work, except for employees whose specific roles or assigned office location require five days’ a week attendance.
Requirements
- Strong production software-engineering experience, including designing, testing, operating and debugging distributed services.
- Meaningful ownership of ML infrastructure or production model delivery, with clear evidence of your contribution and its operational or adoption impact.
- A practical understanding of the ML lifecycle: training, evaluation, versioning, deployment, online validation, monitoring and safe rollback.
- Experience designing for latency, resilience, observability and cost in production systems.
- Experience with LLMs, embeddings, semantic search, retrieval or agent systems, including how you evaluate quality and control failure modes.
- Experience with caching systems and the trade-offs involved in using caching for low-latency services.
- A commitment to automation, CI/CD, code quality and evidence-led engineering decisions.
- A collaborative approach and the ability to explain technical trade-offs to partners across disciplines.
- A degree in Computer Science, Electrical Engineering or a related field, or equivalent practical experience.
Experience with AWS or GCP, Kubernetes, SQL warehouse and batch-processing technologies such as Trino, Presto or Spark, Kafka or other streaming systems, and Go or Python would be useful. Node.js experience is a plus. We value transferable evidence over an exact match to every named technology.
About the company
Roku is a great place for people who want to work in a fast-paced environment where everyone is focused on the company’s success rather than their own. We try to surround ourselves with people who are great at their jobs, who are easy to work with, and who keep their egos in check. We appreciate a sense of humor. We believe a fewer number of very talented folks can do more for less cost than a larger number of less talented teams. We’re independent thinkers with big ideas who act boldly, move fast and accomplish extraordinary things through collaboration and trust. In short, at Roku you’ll be part of a company that’s changing how the world watches TV.
We have a unique culture that we are proud of. We think of ourselves primarily as problem-solvers, which itself is a two-part idea. We come up with the solution, but the solution isn’t real until it is built and delivered to the customer. That penchant for action gives us a pragmatic approach to innovation, one that has served us well since 2002.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
What Are Large Language Models?
Navigating the AI Shift