WeAreDevelopers LIVE Oct 2, 2024

DevOps for AI: running LLMs in production with Kubernetes and KubeFlow

Aarno Aukia

Are silent API updates secretly breaking your production AI? Regain control by self-hosting and autoscaling LLMs natively with Kubernetes, KServe, and proven DevOps principles.

Pause
Mute Enter Fullscreen
#1 about 5 min

Introduction to DevOps for AI and MLOps

Applying traditional automation principles to machine learning operations minimizes deployment risk and maximizes developer productivity.

#2 about 4 min

Evolution of machine learning and generative AI

Scaling deep learning algorithms with hardware acceleration enables modern systems to reliably generate novel content.

#3 about 4 min

Understanding tokenization and probability in large language models

Analyzing the tokenization layer and statistical rendering engines of language models explains why their outputs remain explicitly non-deterministic.

#4 about 5 min

Prompt engineering techniques and security vulnerabilities

Providing targeted instructions and context eliminates unnecessary model complications but introduces potential security risks like malicious prompt injections.

#5 about 4 min

Enhancing models with retrieval augmented generation and fine-tuning

Indexing proprietary documentation for retrieval prevents model hallucinations, while permanent fine-tuning adapts generic models to domain-specific terminology.

#6 about 6 min

Consuming cloud-hosted models through black box APIs

Relying on commercial algorithms requires robust observability and failover mechanisms to securely handle unpredictable service updates.

#7 about 2 min

Running local open source models using Ollama

Leveraging standardized containerization engines allows developers to seamlessly download and execute language models offline natively on local hardware.

#8 about 4 min

Serving production machine learning workloads with KubeFlow and KServe

Implementing specialized orchestration operators standardizes the machine learning lifecycle to serve scalable inference workloads flawlessly.

#9 about 4 min

Configuring horizontal autoscaling for language models with KServe

Defining orchestration resource limitations dynamically scales active container instances while effortlessly capturing diagnostic metrics for visualization.

#10 about 1 min

Choosing between standalone hardware and highly scalable infrastructure

Organizations must evaluate the reduced complexity of standalone hardware runtimes against the absolute operational maturity demanded by auto-scaling architectures.

Matching moments

4:15 min

Introduction to artificial intelligence driven development

Natalie Pistunovich · LIVE

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer · WWC 2023

2:14 min

Differences between traditional MLOps and GenAIOps

Maxim Salnikov Maxim Salnikov · WWC 2025

7:28 min

Accelerating product features using generative large language models

David Singleton David Singleton +1 · Coffee With Developers

4:57 min

Centralizing LLMOps workflows within Azure AI Foundry

Maxim Salnikov Maxim Salnikov · LIVE

Upcoming sessions on this topic

Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Legare Kerrison, Cedric Clyburn

Legare Kerrison
Cedric Clyburn
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

Trust, But Verify: Continuous GPU Validation at Scale

Kyle Bell

VP of AI @ TensorWave

Kyle Bell
Open session

World Congress 2026 North America

It’s Alive! Taming the MLOps Franken-Stack: Write, Run, and Serve with Michelangelo

Paul Zimmerman, Eric Wang

Paul Zimmerman
Eric Wang