Machine Learning Engineer, DevOps/SRE

Roku, Inc.
Austin, TX, United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Working hours
Shift work
Job source

Tech stack

Java (Programming Language) Artificial Intelligence Airflow Amazon Web Services Computer Programming Continuous Integration Data Stores Cursor (Graphical User Interface Elements) DevOps Disaster Recovery Python (Programming Language) Machine Learning
+17 more
NoSQL Prometheus Datadog Aerospike Feature Engineering Grafana Apache Spark Gitlab Kubernetes Infrastructure Automation Frameworks Information Technology Apache Flink Apache Kafka Machine Learning Operations Terraform GPT Jenkins

Job description

We are seeking a talented and experienced Senior Software Engineer, MLOps/DevOps, to join the Advertising Performance team and play a critical role in supporting and scaling our Machine Learning infrastructure. The ideal candidate has a strong background in DevOps/SRE practices, cloud infrastructure management, and MLOps tooling - with a passion for building platforms that accelerate ML experimentation and deployment at internet scale.

You will partner closely with ML Scientists and Engineers to streamline the end-to-end ML lifecycle across training, evaluation, deployment, and monitoring - on top of a modern, cloud-native stack running on GCP and AWS using Kubernetes, Apache Airflow, Spark, Ray, MLflow, Chronon, etc.

What you’ll be doing

  • Lead the design and operation of scalable, production-grade cloud infrastructure for ML workloads across AWS and GCP, including GPU/TPU-based training and inference environments
  • Architect and improve CI/CD systems for ML models and platform services to enable fast, reliable, and safe production releases
  • Own and evolve low-latency infrastructure for real-time model inference, including KV store and vector databases
  • Define and enforce observability standards for ML systems, including model performance monitoring, drift detection, capacity planning, and pipeline health metrics
  • Participate in on-call rotation, leading incident response and root-cause analysis for critical ML training and serving infrastructure
  • Partner with data scientists and ML engineers to improve platform usability, accelerate model iteration, and implement strong MLOps and SRE best practices
  • Champion operational excellence across ML infrastructure through automation, resilience engineering, disaster recovery planning, and continuous improvement, Roku fosters an inclusive and collaborative environment where teams generally work in the office Monday through Thursday. Fridays are generally flexible for remote work, except for employees whose specific roles or assigned office location require five days’ a week attendance.

Requirements

  • BS or MS in Computer Science, Engineering, or a related quantitative field
  • Ability to demonstrate AI tool (Claude, Cursor, ChatGPT etc.) coding efficiency at an advanced skill level
  • 8+ years of experience in DevOps, SRE, or ML infrastructure, including 4+ years supporting large-scale ML or AI systems
  • Strong programming skills in Python and/or Scala or Java for platform automation and tooling
  • Deep experience with Kubernetes and container orchestration on GCP (GKE) and/or AWS (EKS)
  • Expertise with NoSQL or low-latency data stores such as Aerospike or similar technologies
  • Hands-on experience with data and orchestration technologies such as Apache Spark, Apache Flink, Apache Airflow, and Kafka
  • Experience building and maintaining CI/CD systems using tools such as Jenkins or GitLab Runner
  • Familiarity with feature engineering platforms such as Chronon and model lifecycle tools such as MLflow
  • Strong infrastructure-as-code experience with Terraform or similar tooling
  • Experience with observability platforms such as Prometheus, Grafana, and Datadog
  • Excellent communication and cross-functional collaboration skills
  • Experience in the Advertising domain is a plus

About the company

Roku is a great place for people who want to work in a fast-paced environment where everyone is focused on the company’s success rather than their own. We try to surround ourselves with people who are great at their jobs, who are easy to work with, and who keep their egos in check. We appreciate a sense of humor. We believe a fewer number of very talented folks can do more for less cost than a larger number of less talented teams. We’re independent thinkers with big ideas who act boldly, move fast and accomplish extraordinary things through collaboration and trust. In short, at Roku you’ll be part of a company that’s changing how the world watches TV.

We have a unique culture that we are proud of. We think of ourselves primarily as problem-solvers, which itself is a two-part idea. We come up with the solution, but the solution isn’t real until it is built and delivered to the customer. That penchant for action gives us a pragmatic approach to innovation, one that has served us well since 2002.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · World Congress 2024

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler · LIVE

4:15 min

Introduction to artificial intelligence driven development

Natalie Pistunovich · LIVE

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky · World Congress 2025

Videos

See all

Related articles

See all