> Markdown version of [/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow?t=914](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow?t=914). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # DevOps for AI: running LLMs in production with Kubernetes and KubeFlow Are silent API updates secretly breaking your production AI? Regain control by self-hosting and autoscaling LLMs natively with Kubernetes, KServe, and proven DevOps principles. - **Speakers:** Aarno Aukia - **Event:** WeAreDevelopers LIVE - **Published:** October 2, 2024 - **Duration:** 34:21 - **URL:** https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow ## Summary The current state of MLOps often mirrors the chaotic "firefighting" days of early web operations, but software teams can achieve stability by applying proven DevOps automation principles to artificial intelligence workflows. At their core, large language models (LLMs) operate as massive probabilistic engines, or "random token generators," making them poorly suited for applications requiring strictly deterministic outputs. Because they process inputs as tokens mapped against statistical probabilities, engineers must guide behavior using specific application patterns. While prompt engineering currently drives roughly 90% of commercial enterprise use cases, relying purely on third-party cloud APIs introduces severe backend risks. Opaque, unannounced algorithm updates from cloud vendors can silently break application logic in production, making deep LLM observability—including the meticulous logging of prompts, contexts, and raw output responses—a non-negotiable architectural requirement. To manage hallucination risks and reduce vendor reliance, engineering organizations are increasingly adopting Retrieval-Augmented Generation (RAG) alongside open-source models like Llama 3. Supplying customized, localized context immediately improves inference accuracy. For local iteration, tools like Ollama serve colloquially as the "Docker for LLMs," packaging hardware acceleration and model binaries into a simplified, container-like developer experience. However, as AI interfaces become highly exposed to external users, DevOps engineers must implement rigorous defensive architectures against prompt injection attacks designed to subvert system instructions and leak proprietary operational data. Migrating these workloads into scalable production environments demands sophisticated operational abstraction to manage expensive compute resources. Ecosystems like Kubernetes paired with KubeFlow and KServe dramatically reduce the friction of self-hosting AI. Through KServe, teams achieve hardware-agnostic model serving backed by Knative, allowing deployments to seamlessly autoscale or spin entirely down to zero upon inactivity. By providing built-in configurations for canary deployments, traffic shifting, and native Grafana dashboard tracking for inference latency, KServe ensures platform teams can deploy enterprise-ready generative capabilities without reinventing the fundamental wheels of distributed model operations. **Keywords:** devops for ai, mlops process maturity, kubernetes ai workloads, kubeflow integration, kserve model serving, knative serverless inference, retrieval-augmented generation, prompt injection defense, llm prompt engineering, cloud inference observability, ollama model deployment, llm hallucination mitigation, hardware-agnostic AI hosting, generative ai infrastructure ## Chapters 1. **Introduction to DevOps for AI and MLOps** (00:02) — Applying traditional automation principles to machine learning operations minimizes deployment risk and maximizes developer productivity. 1. **Evolution of machine learning and generative AI** (04:21) — Scaling deep learning algorithms with hardware acceleration enables modern systems to reliably generate novel content. 1. **Understanding tokenization and probability in large language models** (07:31) — Analyzing the tokenization layer and statistical rendering engines of language models explains why their outputs remain explicitly non-deterministic. 1. **Prompt engineering techniques and security vulnerabilities** (10:55) — Providing targeted instructions and context eliminates unnecessary model complications but introduces potential security risks like malicious prompt injections. 1. **Enhancing models with retrieval augmented generation and fine-tuning** (15:14) — Indexing proprietary documentation for retrieval prevents model hallucinations, while permanent fine-tuning adapts generic models to domain-specific terminology. 1. **Consuming cloud-hosted models through black box APIs** (18:35) — Relying on commercial algorithms requires robust observability and failover mechanisms to securely handle unpredictable service updates. 1. **Running local open source models using Ollama** (24:00) — Leveraging standardized containerization engines allows developers to seamlessly download and execute language models offline natively on local hardware. 1. **Serving production machine learning workloads with KubeFlow and KServe** (25:56) — Implementing specialized orchestration operators standardizes the machine learning lifecycle to serve scalable inference workloads flawlessly. 1. **Configuring horizontal autoscaling for language models with KServe** (29:32) — Defining orchestration resource limitations dynamically scales active container instances while effortlessly capturing diagnostic metrics for visualization. 1. **Choosing between standalone hardware and highly scalable infrastructure** (33:22) — Organizations must evaluate the reduced complexity of standalone hardware runtimes against the absolute operational maturity demanded by auto-scaling architectures. ## Related Moments - [Introduction to artificial intelligence driven development](https://www.wearedevelopers.com/videos/347-mlops-and-ai-driven-development) (from "MLOps and AI Driven Development") - [Essential engineering roles in the generative AI space](https://www.wearedevelopers.com/videos/844-enter-the-brave-new-world-of-genai-with-vector-search) (from "Enter the Brave New World of GenAI with Vector Search") - [Defining MLOps and its role in production systems](https://www.wearedevelopers.com/videos/825-mlops-on-kubernetes-exploring-argo-workflows) (from "MLOps on Kubernetes: Exploring Argo Workflows") - [Differences between traditional MLOps and GenAIOps](https://www.wearedevelopers.com/videos/1535-from-traction-to-production-maturing-your-genaiops-step-by-step) (from "From Traction to Production: Maturing your GenAIOps step by step") - [Accelerating product features using generative large language models](https://www.wearedevelopers.com/videos/100362-navigating-growth-scaling-challenges-and-office-expansions-with-david-singleton-cto-at-stripe) (from "Navigating Growth, Scaling Challenges, and Office Expansions with David Singleton, CTO at Stripe") - [Centralizing LLMOps workflows within Azure AI Foundry](https://www.wearedevelopers.com/videos/1250-from-traction-to-production-maturing-your-llmops-step-by-step) (from "From Traction to Production: Maturing your LLMOps step by step") ## Related Articles - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) ## Related Jobs - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/588393-machine-learning-engineer) at **Twilio** - [AI Software Engineer (Germany)](https://www.wearedevelopers.com/jobs/48317-ai-software-engineer-germany) at **Sunhat** - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/1355348-machine-learning-engineer) at **TWILIO** - [Staff, Machine Learning Engineer (L4)](https://www.wearedevelopers.com/jobs/ext/1202639-staff-machine-learning-engineer-l4) at **Twilio** - [Staff Software Engineer, Copilot Experiences](https://www.wearedevelopers.com/jobs/ext/164361-staff-software-engineer-copilot-experiences) at **GitHub** - [AI Operations Manager (all genders)](https://www.wearedevelopers.com/jobs/48263-ai-operations-manager-all-genders) at **envelio**