> Markdown version of [/jobs/ext/93261-software-engineer-ml-infra-distributed-systems-staff-principal](https://www.wearedevelopers.com/jobs/ext/93261-software-engineer-ml-infra-distributed-systems-staff-principal). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, ML Infra & Distributed Systems (Staff & Principal) - **Company:** Tubi, Inc. - **Location:** San Francisco, CA, United States (Remote available) - **Salary:** $227,200.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), A/B Testing, Akka (Toolkit), Amazon Web Services, C++ (Programming Language), Content Analysis, Distributed Systems, Java Virtual Machine (JVM), Python (Programming Language), PostgreSQL, Machine Learning, Message Broker, NoSQL, Recommender Systems, Redis, Data Driven Tests, Scala (Programming Language), SQL Databases, Cloud Platform System, Erlang, Amazon ElastiCache, Large Language Models, Deep Learning, Caching, Backend, Build Management, Containerization, Kubernetes, Low Latency, Cassandra, Apache Kafka, Machine Learning Operations, Docker, Elixir, Golang, Microservices - **Published:** May 24, 2026 - **Apply:** https://arc.dev/remote-jobs/j/tubi-software-engineer-ml-infra-distributed-systems-staff-principal-orohhu0ynr ## About the Role * Experience designing and building scalable, distributed systems in any modern backend language (e.g., Scala, Java, Python, Go, C++); experience with Scala or JVM based language is a plus. * Strong experience with AWS or an equivalent cloud platform * Experience building online microservices at scale with low latency serving * Experience with both SQL (e.g. Postgres) and NoSQL databases (e.g. Cassandra), message brokers (e.g. Kafka), and caches (e.g. Redis) * Experience with containerization technologies, such as Docker or Kubernetes * Led the response and resolution efforts for multiple major, large-scale incidents, * Familiarity with the machine learning infrastructure like inference engines (e.g. torschserve, triton, vLLM), vector stores (e.g. LanceDB, FAISS), feature stores (e.g. Feast), ElastiCache, model training orchestration, etc. * Understanding of ML model training pipelines and model internals. Experience with Recommender Systems, Search, Autocomplete and Ads ML is a plus * Previous experience with Akka, Erlang, Elixir or Go * Proficient in data-driven analysis of complex A/B testing results ## Description As a Software Engineer on the ML Infrastructure team, you will collaborate closely with the Machine Learning and Product teams to build world-class machine learning inference platforms. These platforms power essential services like personalized recommendations, search, and content understanding across Tubi. A core responsibility of this team is developing and maintaining low-latency ML model serving systems that support Deep Learning, LLM, and Search models. This involves building self-service infrastructure and critical components such as the inference engine, feature store, vector store, and experimentation engine. You will improve the way we deploy and operate our services and even contribute to open-source projects. This role grants the architectural freedom to explore new frameworks, lead critical cross-functional projects, and transform the capabilities of our ML and Product teams., * Principal Software Engineer * Additional Details: As a Principal Engineer on the ML Infrastructure team, you will be a technical leader and visionary, driving the evolution of our machine learning platform. You will tackle the most complex and impactful technical challenges, shaping the architecture and technology choices that enable our ML capabilities to scale and deliver exceptional user experiences. You will be a key influencer, bridging the gap between engineering and product, and a mentor to senior engineers, fostering a culture of technical excellence and continuous improvement. Your work will be used by millions of users., * Design and build scalable, high throughput, and low latency distributed systems using Scala * Build reusable components and services that serve various ML applications like Personalization, Search, Ads and Exploration * Partner closely with ML engineers to understand their challenges and limitations and develop scalable solutions to address them. Proactively recommend solutions to keep our ML Inference stack state of the art. * Take a data driven approach to identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if necessary * Mentor other engineers on the team on system design, effective incident management, interviewing, leveraging LLMs for work, etc. * Collaborate with ML, Product, and cross functional engineering teams to define the long term vision and architecture for ML Infrastructure at Tubi. ## Related Videos - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Bringing digital education to refugee and host communities in remote regions of Africa](https://www.wearedevelopers.com/videos/644-bringing-digital-education-to-refugee-and-host-communities-in-remote-regions-of-africa) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [NoSQL Data Modeling for Front-end Developers](https://www.wearedevelopers.com/videos/297-nosql-data-modeling-for-front-end-developers) - [TiDB, One Layer at a Time: How Distributed SQL Became an Agentic AI Backbone](https://www.wearedevelopers.com/videos/100117-tidb-one-layer-at-a-time-how-distributed-sql-became-an-agentic-ai-backbone) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 122 - Cracks in the polyfill](https://www.wearedevelopers.com/magazine/457-dev-digest-122-cracks-in-the-polyfill) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)