> Markdown version of [/jobs/ext/2966780-staff-machine-learning-engineer-ml-platform](https://www.wearedevelopers.com/jobs/ext/2966780-staff-machine-learning-engineer-ml-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Machine Learning Engineer, ML Platform - **Company:** Braze - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $314,000.0 - **Contract:** Permanent contract - **Skills:** Cloud Computing, Code Review, Continuous Integration, Data Systems, Distributed Systems, Identity and Access Management, Python (Programming Language), MongoDB, RabbitMQ, Ruby on Rails, Redis, Azure Machine Learning, Kubernetes, Apache Kafka, Machine Learning Operations, Celery - **Published:** September 17, 2026 - **Apply:** https://job-boards.greenhouse.io/braze/jobs/8209438 ## About the Role * 8+ years building and operating distributed systems in production, with depth in deployment and operations. You have designed services for scale and reliability, owned CI/CD and infrastructure as code, and run what you built under production load * Hands-on experience with ML workloads in production. Training pipelines, model serving, feature systems, or ML platform tooling all count; deep modeling experience is a plus rather than a requirement * A technical leader who has owned direction for a team, led multi-quarter initiatives across team boundaries, and grown senior engineers, all while keeping a high personal output * Deep working knowledge of Kubernetes and cloud infrastructure, including identity and access management, networking, and the cost profile of what you run * An effective communicator, both verbal and written, whose designs and recommendations build consensus and drive forward decision making * Bonus: * Queueing and orchestration systems such as Celery, RabbitMQ, Kafka, or Ray * ML platform tooling such as MLflow or another model registry, feature stores, or ML observability * Experience in our stack (Python, Ruby on Rails, MongoDB, Redis, Kubernetes) * Operating under compliance regimes such as SOX or HIPAA * Customer engagement, personalization, or marketing technology domain experience ## Description * Identify and drive the transformative initiatives that change how the team runs ML in production, whether that's replatforming our queueing and orchestration, overhauling deployment and cloud identity, or retiring a generation of infrastructure * Build and ship at high velocity. Staff at Braze is a hands-on delivery role; you carry the most complex infrastructure initiatives yourself from design through production. Current examples include multi-region model serving fleets, the pipelines that keep hundreds of customer-specific models healthy, and the CI and deployment tooling that moves it all safely * Own the platform's technical vision and production quality bar. Set direction for how models are trained, deployed, served, and observed; lead incident response for ML systems; and drive the reliability and cost work that keeps the platform efficient at scale * Drive initiatives that span teams. Our platform builds on shared infrastructure, deployment tooling, and data systems owned with partner teams, and you carry the technical relationships with those teams * Raise the team's engineering quality through design review, code review, and production readiness for ML systems, and mentor other senior engineers and data scientists * Connect technical decisions to customer and business outcomes, and represent the team's technical perspective to product and engineering leadership ## Related Videos - [Celery on AWS ECS - the art of background tasks & continuous deployment](https://www.wearedevelopers.com/videos/561-celery-on-aws-ecs-the-art-of-background-tasks-continuous-deployment) - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [40 Minutes to Build a Serverless COVID-19 REST and GraphQL APIs](https://www.wearedevelopers.com/videos/208-40-minutes-to-build-a-serverless-covid-19-rest-and-graphql-apis) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [DevOps for Machine Learning](https://www.wearedevelopers.com/videos/179-devops-for-machine-learning) - [NoSQL Data Modeling for Front-end Developers](https://www.wearedevelopers.com/videos/297-nosql-data-modeling-for-front-end-developers) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)