> Markdown version of [/jobs/ext/2171785-machine-learning-engineer-infrastructure](https://www.wearedevelopers.com/jobs/ext/2171785-machine-learning-engineer-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer, Infrastructure - **Company:** Patreon, Inc. - **Location:** New York, NY, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Clean Code Principles, Code Review, Information Engineering, Software Debugging, Distributed Systems, Python (Programming Language), Software Reliability Testing, Systems Architecture, System Availability, Low Latency, Machine Learning Operations - **Published:** August 21, 2026 - **Apply:** https://www.dice.com/job-detail/3bd55787-b3a6-473d-935e-371e9fd0894a ## About the Role * You have deep experience building, deploying, and maintaining production-grade ML infrastructure at scale, specifically with low-latency live inference pipelines and feature store architectures. * You have a strong background in distributed systems and backend engineering, with the ability to write robust, maintainable code in Python. * You have a systematic approach to debugging complex, high-throughput systems and performance bottlenecks. * You are energized by building '0 to 1' infrastructure systems that stand the test of time and provide a reliable foundation for the team. * You possess strong communication skills and are effective at creating clear documentation for system architectures and infrastructure strategies. * You have a growth mindset, a keen eye for detail in code reviews, and a passion for empowering your teammates by improving developer velocity. We hire talented and passionate people from different backgrounds because workplace diversity and inclusion is critical to our ability to serve creators worldwide. If you're excited about a role but your past experience doesn't match with every bullet point outlined above, we strongly encourage you to apply anyway. If you're a creator at heart, are energized by our mission, and share our company values, we'd love to hear from you. ## Description You'll join the Relevance team, whose mission is to build the ML systems that power how fans discover creators and how content surfaces across Patreon. The team is responsible for search, feed ranking, and creator-fan matching. You'll work closely with a small, collaborative group of MLEs on shared infrastructure, code reviews, and roadmap alignment, while partnering cross-functionally with Product, Data Engineering, and Trust & Safety to deliver measurable impact across the platform., * Architect, scale, and maintain high-throughput, low-latency live inference infrastructure to support our relevance systems. * Own the end-to-end feature store lifecycle-from ingestion and transformation to production serving, ensuring high availability and consistency between online and offline features. * Design and implement observability, monitoring, and validation frameworks to detect performance gaps, latency spikes, and production drift. * Collaborate with cross-functional partners, such as product, data engineering, and trust and safety, to translate product requirements into robust, scalable infrastructure solutions. * Automate model deployment and reliability testing to improve developer velocity and ensure system stability. * Debug complex relevance systems when monitoring identifies performance bottlenecks or reliability issues. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Using AI Without Losing Your Skills](https://www.wearedevelopers.com/videos/2045-using-ai-without-losing-your-skills) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [Agentic employees in world's most downloaded FinTech app](https://www.wearedevelopers.com/videos/100123-agentic-employees-in-world-s-most-downloaded-fintech-app) - [What I learned as a developer from accidents in space](https://www.wearedevelopers.com/videos/642-what-i-learned-as-a-developer-from-accidents-in-space) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)