> Markdown version of [/videos/179-devops-for-machine-learning](https://www.wearedevelopers.com/videos/179-devops-for-machine-learning). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # DevOps for Machine Learning Ninety percent of machine learning models never reach production. Close the workflow gap between data science and engineering by adopting an evolutionary MLOps infrastructure. - **Speakers:** Hauke Brammer - **Event:** World Congress 2021 - **Published:** June 28, 2021 - **Duration:** 44:40 - **URL:** https://www.wearedevelopers.com/videos/179-devops-for-machine-learning ## Summary While 90% of machine learning models fail to reach production due to workflow gaps, bridging software engineering best practices with data science through MLOps can solve this operational divide. Unlike linear software development, ML requires an experimental, iterative approach where models naturally degrade over time due to concept drift. This necessitates a structured three-pillar infrastructure: data and feature pipelines supported by feature stores and versioning tools like DVC; robust experiment tracking utilizing MLflow and Kubeflow to eliminate duplicated efforts; and comprehensive deployment strategies capable of running centrally or at the edge with statistical monitoring via active learning. Organizations should avoid a big-bang rollout, instead adopting an evolutionary approach and fostering cross-functional teams that embrace a 'you build it, you run it' culture to sustainably operationalize machine learning. **Keywords:** mlops, machine learning lifecycle, production ml models, concept drift, feature stores, data version control, experiment tracking, model registry, active learning, statistical monitoring, shadow models, edge deployment strategies, kubeflow pipelines, workflow automation tools, cross-functional data teams, reproducible data science ## Chapters 1. **Defining machine learning operations and deployment failure rates** (00:02) — The transition from pure machine learning research to operationalizing models prevents massive production failure rates. 1. **Common life cycle challenges in machine learning projects** (03:49) — Recognizing duplicated efforts, lack of reproducibility, and poor monitoring highlights the need for robust operational processes. 1. **Differences between traditional software engineering and machine learning** (05:57) — The linear development of traditional software contrasts sharply with the experimental, degrading nature of machine learning models. 1. **Establishing goals and the machine learning operations lifecycle** (09:21) — Establishing reproducible pipelines and continuous evaluation environments controls the cyclical phases of algorithmic refinement. 1. **Building data pipelines and managing machine learning features** (10:26) — Centralized feature stores transform raw data into reusable and versioned elements for multiple project teams. 1. **Choosing data flow tools and dedicated feature stores** (14:09) — Modern infrastructure combines workflow managers with centralized stores to distribute reusable training inputs effectively. 1. **Managing machine learning experiments and collaborative research hubs** (18:00) — Shared experimental environments prevent configuration drift and safely coordinate thousands of concurrent model iterations. 1. **Tracking training metadata and automating model deployment pipelines** (20:25) — Versioning experiments and automatically provisioning distributed training environments captures every parameter essential for model replication. 1. **Evaluating central server APIs against edge deployment models** (23:56) — Serving models close to input sources solves acute privacy limitations and high-bandwidth network bottlenecks. 1. **Monitoring data distributions and implementing active learning feedback** (27:51) — Tracking stochastic input variations and capturing user corrections reveals conceptual drift missed by traditional metrics. 1. **Creating model transparency with continuous operational metrics logging** (32:11) — Documenting prediction dependencies and standardizing timeseries metric databases ensures accountability inside opaque machine learning algorithms. 1. **Fostering cross-functional collaboration and incremental process adoption** (33:54) — Blending data scientists with operational developers establishes shared accountability that eliminates siloed production failures. 1. **Question and answer session on tooling and scaling** (38:33) — Addressing participant inquiries clarifies integrating massive parallel computing pipelines, navigating data versioning limits, and identifying malicious feedback. ## Related Moments - [Defining MLOps and its role in production systems](https://www.wearedevelopers.com/videos/825-mlops-on-kubernetes-exploring-argo-workflows) (from "MLOps on Kubernetes: Exploring Argo Workflows") - [Bridging the gap between model management and devops](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) (from "AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment") - [Introduction to DevOps for AI and MLOps](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) (from "DevOps for AI: running LLMs in production with Kubernetes and KubeFlow") - [Solving application deployment complexities using LLMOps pipelines](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) (from "LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices") - [Defining machine learning operations in a fragmented ecosystem](https://www.wearedevelopers.com/videos/392-mlops-what-s-the-deal-behind-it) (from "MLOps - What’s the deal behind it?") - [Adopting a DevOps culture for machine learning pipelines](https://www.wearedevelopers.com/videos/369-the-state-of-mlops-machine-learning-in-production-at-enterprise-scale) (from "The state of MLOps - machine learning in production at enterprise scale") ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Why Your AI Tool Fails After the Demo](https://www.wearedevelopers.com/magazine/704-why-your-ai-tool-fails-after-the-demo) ## Related Jobs - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/588393-machine-learning-engineer) at **Twilio** - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/1597388-machine-learning-engineer) at **ZEISS Group** - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/1355348-machine-learning-engineer) at **TWILIO** - [Staff, Machine Learning Engineer (L4)](https://www.wearedevelopers.com/jobs/ext/1202639-staff-machine-learning-engineer-l4) at **Twilio** - [Data Scientist](https://www.wearedevelopers.com/jobs/ext/1351648-data-scientist) at **Almedia** - [Principal Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/1706410-principal-machine-learning-engineer) at **Almedia**