World Congress 2021 Jun 28, 2021

DevOps for Machine Learning

Hauke Brammer

Ninety percent of machine learning models never reach production. Close the workflow gap between data science and engineering by adopting an evolutionary MLOps infrastructure.

Pause
Mute Enter Fullscreen
#1 about 4 min

Defining machine learning operations and deployment failure rates

The transition from pure machine learning research to operationalizing models prevents massive production failure rates.

#2 about 3 min

Common life cycle challenges in machine learning projects

Recognizing duplicated efforts, lack of reproducibility, and poor monitoring highlights the need for robust operational processes.

#3 about 4 min

Differences between traditional software engineering and machine learning

The linear development of traditional software contrasts sharply with the experimental, degrading nature of machine learning models.

#4 about 2 min

Establishing goals and the machine learning operations lifecycle

Establishing reproducible pipelines and continuous evaluation environments controls the cyclical phases of algorithmic refinement.

#5 about 4 min

Building data pipelines and managing machine learning features

Centralized feature stores transform raw data into reusable and versioned elements for multiple project teams.

#6 about 4 min

Choosing data flow tools and dedicated feature stores

Modern infrastructure combines workflow managers with centralized stores to distribute reusable training inputs effectively.

#7 about 3 min

Managing machine learning experiments and collaborative research hubs

Shared experimental environments prevent configuration drift and safely coordinate thousands of concurrent model iterations.

#8 about 4 min

Tracking training metadata and automating model deployment pipelines

Versioning experiments and automatically provisioning distributed training environments captures every parameter essential for model replication.

#9 about 4 min

Evaluating central server APIs against edge deployment models

Serving models close to input sources solves acute privacy limitations and high-bandwidth network bottlenecks.

#10 about 5 min

Monitoring data distributions and implementing active learning feedback

Tracking stochastic input variations and capturing user corrections reveals conceptual drift missed by traditional metrics.

#11 about 2 min

Creating model transparency with continuous operational metrics logging

Documenting prediction dependencies and standardizing timeseries metric databases ensures accountability inside opaque machine learning algorithms.

#12 about 5 min

Fostering cross-functional collaboration and incremental process adoption

Blending data scientists with operational developers establishes shared accountability that eliminates siloed production failures.

#13 about 7 min

Question and answer session on tooling and scaling

Addressing participant inquiries clarifies integrating massive parallel computing pipelines, navigating data versioning limits, and identifying malicious feedback.

Matching moments

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer · WWC 2023

2:15 min

Bridging the gap between model management and devops

Joy Joy · WWC 2024

4:19 min

Introduction to DevOps for AI and MLOps

Aarno Aukia · LIVE

3:56 min

Solving application deployment complexities using LLMOps pipelines

Anshul Jindal Anshul Jindal · WWC 2025

2:48 min

Defining machine learning operations in a fragmented ecosystem

Nico Axtmann · WWC 2022

4:31 min

Adopting a DevOps culture for machine learning pipelines

Bas Geerdink · LIVE

Upcoming sessions on this topic

Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

It’s Alive! Taming the MLOps Franken-Stack: Write, Run, and Serve with Michelangelo

Eric Wang, Paul Zimmerman

Eric Wang
Paul Zimmerman
Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

Reinventing Testing Practices in the AI Era

Eric Deandrea

Java Champion & Senior Principal Software Engineer, IBM

Eric Deandrea
Open session

World Congress 2026 North America

Run your agents in Kubernetes: Build once, deploy anywhere. But really?

Michal Salanci

Senior Systems Engineer at ESET Cybersecurity

Michal Salanci
Open session

World Congress 2026 North America

Trust, But Verify: Continuous GPU Validation at Scale

Kyle Bell

VP of AI @ TensorWave

Kyle Bell