> Markdown version of [/jobs/ext/525229-senior-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/525229-senior-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Machine Learning Engineer - **Company:** Roku, Inc. - **Location:** Austin, TX, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Airflow, Amazon Web Services, Cloud Computing, Cloud Engineering, Computer Programming, Continuous Integration, Data Stores, DevOps, Disaster Recovery, Python (Programming Language), Machine Learning, NoSQL, Site Reliability Engineering Practices, Prometheus, Datadog, Aerospike, Feature Engineering, Grafana, Apache Spark, Gitlab, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Apache Flink, Apache Kafka, Machine Learning Operations, Terraform, Jenkins - **Published:** June 10, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=8e2b721d35d7a5a7 ## About the Role Do you have experience in Tooling?, Do you have a Bachelor's degree?, * BS or MS in Computer Science, Engineering, or a related quantitative field * 8+ years of experience in DevOps, SRE, or ML infrastructure, including 4+ years supporting large-scale ML or AI systems * Strong programming skills in Python and/or Scala or Java for platform automation and tooling * Deep experience with Kubernetes and container orchestration on GCP (GKE) and/or AWS (EKS) * Expertise with NoSQL or low-latency data stores such as Aerospike or similar technologies * Hands-on experience with data and orchestration technologies such as Apache Spark, Apache Flink, Apache Airflow, and Kafka * Experience building and maintaining CI/CD systems using tools such as Jenkins or GitLab Runner * Familiarity with feature engineering platforms such as Chronon and model lifecycle tools such as MLflow * Strong infrastructure-as-code experience with Terraform or similar tooling * Experience with observability platforms such as Prometheus, Grafana, and Datadog * Excellent communication and cross-functional collaboration skills * Experience in the Advertising domain is a plus ## Description We are seeking a talented and experienced Senior Software Engineer, MLOps/DevOps to join the Advertising Performance team and play a critical role in supporting and scaling our Machine Learning infrastructure. The ideal candidate has a strong background in DevOps/SRE practices, cloud infrastructure management, and MLOps tooling - with a passion for building platforms that accelerate ML experimentation and deployment at internet scale. You will partner closely with ML Scientists and Engineers to streamline the end-to-end ML lifecycle across training, evaluation, deployment, and monitoring - on top of a modern, cloud-native stack running on GCP and AWS using Kubernetes, Apache Airflow, Spark, Ray, MLflow, Chronon, etc., * Lead the design and operation of scalable, production-grade cloud infrastructure for ML workloads across AWS and GCP, including GPU/TPU-based training and inference environments * Architect and improve CI/CD systems for ML models and platform services to enable fast, reliable, and safe production releases * Own and evolve low-latency infrastructure for real-time model inference, including KV store and vector databases * Define and enforce observability standards for ML systems, including model performance monitoring, drift detection, capacity planning, and pipeline health metrics * Participate in on-call rotation, leading incident response and root-cause analysis for critical ML training and serving infrastructure * Partner with data scientists and ML engineers to improve platform usability, accelerate model iteration, and implement strong MLOps and SRE best practices * Champion operational excellence across ML infrastructure through automation, resilience engineering, disaster recovery planning, and continuous improvement ## Related Videos - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [The state of MLOps - machine learning in production at enterprise scale](https://www.wearedevelopers.com/videos/369-the-state-of-mlops-machine-learning-in-production-at-enterprise-scale) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [React Developer Salary [2023]](https://www.wearedevelopers.com/magazine/198-react-developer-salary-2023) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)