WeAreDevelopers LIVE • Oct 2, 2024

DevOps for AI: running LLMs in production with Kubernetes and KubeFlow

Aarno Aukia

Are silent API updates secretly breaking your production AI? Regain control by self-hosting and autoscaling LLMs natively with Kubernetes, KServe, and proven DevOps principles.

Pause
Mute Enter Fullscreen
#1 about 5 min

Introduction to DevOps for AI and MLOps

Applying traditional automation principles to machine learning operations minimizes deployment risk and maximizes developer productivity.

#2 about 4 min

Evolution of machine learning and generative AI

Scaling deep learning algorithms with hardware acceleration enables modern systems to reliably generate novel content.

#3 about 4 min

Understanding tokenization and probability in large language models

Analyzing the tokenization layer and statistical rendering engines of language models explains why their outputs remain explicitly non-deterministic.

#4 about 5 min

Prompt engineering techniques and security vulnerabilities

Providing targeted instructions and context eliminates unnecessary model complications but introduces potential security risks like malicious prompt injections.

#5 about 4 min

Enhancing models with retrieval augmented generation and fine-tuning

Indexing proprietary documentation for retrieval prevents model hallucinations, while permanent fine-tuning adapts generic models to domain-specific terminology.

#6 about 6 min

Consuming cloud-hosted models through black box APIs

Relying on commercial algorithms requires robust observability and failover mechanisms to securely handle unpredictable service updates.

#7 about 2 min

Running local open source models using Ollama

Leveraging standardized containerization engines allows developers to seamlessly download and execute language models offline natively on local hardware.

#8 about 4 min

Serving production machine learning workloads with KubeFlow and KServe

Implementing specialized orchestration operators standardizes the machine learning lifecycle to serve scalable inference workloads flawlessly.

#9 about 4 min

Configuring horizontal autoscaling for language models with KServe

Defining orchestration resource limitations dynamically scales active container instances while effortlessly capturing diagnostic metrics for visualization.

#10 about 1 min

Choosing between standalone hardware and highly scalable infrastructure

Organizations must evaluate the reduced complexity of standalone hardware runtimes against the absolute operational maturity demanded by auto-scaling architectures.

Matching moments

4:15 min

Introduction to artificial intelligence driven development

Natalie Pistunovich · LIVE

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer · World Congress 2023

2:14 min

Differences between traditional MLOps and GenAIOps

Maxim Salnikov Maxim Salnikov · World Congress 2025

7:28 min

Accelerating product features using generative large language models

David Singleton David Singleton +1 · Coffee With Developers

4:57 min

Centralizing LLMOps workflows within Azure AI Foundry

Maxim Salnikov Maxim Salnikov · LIVE