> Markdown version of [/jobs/ext/2289602-engineering-manager-ml-platform](https://www.wearedevelopers.com/jobs/ext/2289602-engineering-manager-ml-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Engineering Manager, ML Platform - **Company:** DATA 206 LLC - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Build Automation, Fraud Prevention and Detection, Python (Programming Language), Machine Learning, Azure Machine Learning, Google Cloud, Feature Engineering, Apache Spark, Model Validation, Technical Debt, Kubernetes, Information Technology, Apache Flink, Apache Kafka, Machine Learning Operations, Docker, Databricks - **Published:** August 29, 2026 - **Apply:** https://www.dice.com/job-detail/6e816f5e-ae53-48f5-bc30-6f2be4f9add0 ## About the Role * 8+ years of overall hands-on engineering experience, including 4+ years managing software or machine learning engineering teams. * Deep technical fluency in machine learning systems: model training pipelines, feature engineering, model serving, and evaluation at production scale. * Proven track record leading technical customer engagements or POVs, including direct interaction with enterprise customers. * Demonstrated success reducing technical debt in a live, high-traffic production system without stalling feature delivery. * Experience designing or scaling evaluation frameworks (offline and/or online) for machine learning models. * Track record of identifying manual, repeatable engineering processes and driving their automation. * Experience hiring, mentoring, and developing engineering talent. * B.S. in Computer Science (or related technical discipline), or equivalent practical experience., * Experience with large-scale distributed ML infrastructure such as Spark, Flink, Databricks, or similar. * Familiarity with fraud detection, risk, or trust & safety domains. * Hands-on experience with Google Cloud Platform or AWS ML infrastructure. * Experience with streaming architectures (e.g., Kafka) and containerized/orchestrated deployments (Docker, Kubernetes). * Familiarity with using AI coding assistants (e.g., Claude Code) to accelerate development. ## Description We're hiring a Senior Engineering Manager to lead this team as a backfill for our outgoing lead. This isn't a maintenance role - it's a chance to modernize a foundational platform at a moment when the stakes are high: our biggest deals increasingly come down to who can win a competitive proof-of-value the fastest, and this team's tooling determines whether we win it. You're a manager who's inspiring and technical, and who knows how to bring focus to what matters now without losing sight of the long term. You value collaboration and transparency, operate with a get-stuff-done mindset, and bring the technical depth and bias for shipping to spot the manual, brittle, or duplicated work that's quietly slowing the team down. You build a culture of mentorship, give regular and constructive feedback, set clear goals, and grow your team by hiring effectively. Projects You Might Lead * Launch a unified model evaluation framework that gives Data Science fast, trustworthy, apples-to-apples comparisons before a model ever reaches production or shadow traffic. * Evolve core feature infrastructure - including a new global feature store - to improve accuracy and unlock faster experimentation. * Build the tooling and metrics that let Sift run faster, sharper customer proof-of-value engagements, online and offline, so we win competitive bake-offs instead of losing them to slow iteration. * Bring a fresh approach to model configuration, replacing tribal knowledge and manual gating with auditable, safely-controlled releases. * Introduce agentic, AI-assisted tooling into customer investigations, automating repetitive data pulls and validation so analysts spend their time on judgment calls, not manual digging. * Build automation that detects an active fraud attack, adjusts score calibration in real time, and cleanly reverts once it subsides. What You'll Do * Lead and grow the team: Own the roadmap, execution, and quality of the systems that train, evaluate, and serve Sift's ML models in production, leading a team of ML platform engineers and data scientists. * Stay technical: Review designs, unblock engineers on hard problems, and make credible calls on architecture and trade-offs. * Drive customer POVs: Partner directly with strategic customers and Sales/Solutions Engineering on technical proof-of-value engagements, translating customer requirements into platform capabilities. * Reduce technical debt: Drive a sustained, measurable reduction in technical debt across the ML platform, balancing new feature delivery with the health of existing systems. * Build evaluation frameworks: Mature the systems that give Data Science and ML Engineering fast, trustworthy signals on model quality before and after deployment. * Automate the ML lifecycle: Identify repeatable, manual processes across training, evaluation, deployment, and monitoring, and drive their automation. * Partner cross-functionally: Align platform investments with business priorities alongside Data Science, Core Infrastructure, Product, and Customer Success. Technical Stack Google Cloud Platform, AWS, Spark, Kafka, Kubernetes, Docker, Databricks, Python ## Related Videos - [Developer Experience, Platform Engineering and AI powered Apps](https://www.wearedevelopers.com/videos/990-developer-experience-platform-engineering-and-ai-powered-apps) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Microservices: how to get started with Spring Boot and Kubernetes](https://www.wearedevelopers.com/videos/242-microservices-how-to-get-started-with-spring-boot-and-kubernetes) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)