> Markdown version of [/jobs/ext/2115414-staff-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/2115414-staff-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Machine Learning Engineer - **Company:** reddit Inc. - **Location:** United States - **Experience:** Expert - **Salary:** $292,500.0 - **Contract:** Permanent contract - **Skills:** LTE (Telecommunication), Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Cloud Computing, Cloud Storage, Continuous Integration, Python (Programming Language), Machine Learning, Azure Machine Learning, Software Deployment, Management of Software Versions, Delivery Pipeline, Large Language Models, AI Platforms, Kubernetes, Machine Learning Operations, Terraform, Web Api, Programming Languages - **Published:** August 19, 2026 - **Apply:** https://job-boards.greenhouse.io/reddit/jobs/7772274 ## About the Role * 10+ years of experience in ML Engineering, AI Platform Engineering, or Cloud AI Deployment roles. * Have a track record of leading technical strategy and delivering AI platforms in cloud-based production environments at scale. * Demonstrate strong execution by turning strategy into action, driving complex initiatives end to end, and consistently delivering high-quality platform outcomes. * Bring deep experience operating Kubernetes and other orchestration systems in large-scale production environments. * Deep experience with cloud-based technologies for supporting an ML platform, including tools like AWS, Google Cloud Storage, infrastructure-as-code (Terraform), and more * Proficiency with the common programming languages and frameworks of ML, such as Go, Python, etc. * Excellent communication skills with the ability to articulate technical AI concepts to non-technical stakeholders * Strong focus on scalability, reliability, performance, and developer experience. You are an undying advocate for platform users and have a deep intuition for the genAI product development lifecycle. * Strong knowledge of model serving, inference pipelines, monitoring, and observability for AI systems is a plus ## Description As a Senior Staff Software Engineer, you will help define and lead the vision for Reddit's large-scale GenAI Platform, shaping the strategy, architecture, and operating model that enable teams across the company to build, deploy, and scale generative AI products with confidence. Contribute to the design, implementation, and maintenance of the LLM Gateway, focusing on features like unified API endpoints for internal/externally hosted LLM, rate/token limit management, and intelligent failover mechanisms to boost uptime and reliability. * Lead and execute the vision, strategy, and roadmap for Reddit's large-scale GenAI Platform. * Define the platform architecture and operating model that enable teams to build, deploy, and scale GenAI products reliably. * Drive the strategy for a unified LAG Gateway supporting internally and externally hosted LLMs through consistent APIs and abstractions. * Set the direction for core platform capabilities such as rate and token limit management, intelligent failover, and production resilience. * Shape Reddit's approach to an enterprise-grade RAG system * Establish the strategic direction for agentic AI workflows and tool-use patterns across the platform. * Own the end-to-end platform strategy from concept through production adoption and long-term evolution. * Drive MLOps and LLMOps standards across CI/CD, testing, versioning, evaluation, and lifecycle management. * Define best practices for observability, monitoring, governance, and operational excellence across GenAI systems. * Partner across engineering, product, and leadership to align platform investments with company priorities and user needs. * Champion platform thinking with a strong focus on scalability, reliability, performance, and developer experience. * Influence technical direction across teams by turning emerging AI capabilities into a scalable platform strategy. ## Related Videos - [Web APIs you might not know about](https://www.wearedevelopers.com/videos/281-web-apis-you-might-not-know-about) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Project Fugu: Extending the web](https://www.wearedevelopers.com/videos/832-project-fugu-extending-the-web) - [Machine Learning for Software Developers (and Knitters)](https://www.wearedevelopers.com/videos/154-machine-learning-for-software-developers-and-knitters) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)