> Markdown version of [/jobs/ext/2229288-director-ml-engineering-infrastructure](https://www.wearedevelopers.com/jobs/ext/2229288-director-ml-engineering-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Director, ML Engineering & Infrastructure - **Company:** Tubi, Inc. - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Experienced - **Salary:** $292,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Big Data, Distributed Systems, Fault Tolerance, Monitoring of Systems, Machine Learning, Tensorflow, Feature Engineering, Data Ingestion, Pytorch, Deep Learning, Information Technology, Machine Learning Operations, Databricks - **Published:** August 25, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/peci8nq22e ## About the Role * 10+ years of industry experience spanning machine learning engineering and distributed systems. * 3+ years of leadership and management experience, with a proven ability to build and lead strong technical teams. * MSc or Ph.D. in Computer Science, Machine Learning, or related field, or equivalent practical experience. * Proven expertise in building and deploying end-to-end ML systems at scale, including recommendation and personalization systems. * Strong background in distributed systems architecture, including low-latency services, streaming platforms, and large-scale serving. * Hands-on experience with deep learning frameworks (e.g., TensorFlow, PyTorch) and ML infrastructure technologies. * Track record of delivering high-quality, scalable, and fault-tolerant systems. * Excellent communication skills and ability to influence product and technical strategy. * Proven experience deploying large-scale serving systems on AWS and demonstrated expertise in leveraging Databricks for large-scale data processing and ML workflows ## Description The Machine Learning team at Tubi drives the innovation behind personalized user experiences. With the largest inventory in the industry and hundreds of millions of viewers, we tackle problems in the space of recommendations, search, content understanding, and ads optimization that shape the future of streaming., We are seeking a Director of Machine Learning Engineering and Infrastructure to lead a hybrid team bridging advanced ML engineering with world-class infrastructure design. In this role, you will own the strategic direction and execution for scaling our machine learning capabilities while ensuring our distributed systems and infrastructure can support innovation at massive scale. You will combine technical depth with leadership excellence to guide teams that deliver both foundational ML systems and high-performance distributed services. What You'll Do: * Lead and manage high-performing teams across ML engineering and ML infrastructure, fostering a culture of innovation, collaboration, and growth. * Define and execute the strategic roadmap for ML systems, including recommendation, personalization, and ads optimization. * Oversee the design, development, and deployment of scalable ML pipelines: data ingestion, feature engineering, model training, evaluation, and serving. * Architect distributed systems to support ML workloads at scale, ensuring reliability, observability, and operational excellence. * Partner closely with Product, Engineering, and Content teams to align on business goals and deliver impactful ML-driven experiences. * Support best practices in experimentation, evaluation, and ML system monitoring. * Ensure cost efficiency, scalability, and performance in ML infrastructure investments. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [How We Built a Machine Learning-Based Recommendation System (And Survived to Tell the Tale)](https://www.wearedevelopers.com/videos/752-how-we-built-a-machine-learning-based-recommendation-system-and-survived-to-tell-the-tale) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Tech YouTube Channels for Developers in 2023](https://www.wearedevelopers.com/magazine/261-top-tech-youtube-channels-for-developers-in-2023) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)