World Congress 2025 Aug 20, 2025 Session details

Self-Hosted LLMs: From Zero to Inference

Cedric Clyburn , Roberto Carratalá

Why risk data privacy with third-party APIs when you can run models natively? Discover how to deploy, quantize, and scale self-hosted LLMs for completely secure, offline development.

Pause
Mute Enter Fullscreen
#1 about 3 min

The growing trend of self-hosting AI models

Why developers are increasingly opting to self-host language models to reduce reliance on third-party web services.

#2 about 2 min

Why developers should run AI models locally

How self-hosting resolves organizational data privacy concerns and accelerates the inner loop of local application development.

#3 about 3 min

Open source tools for running and scaling models

How container-based execution engines assist in scaling and inferencing massive language models across varied hardware constraints.

#4 about 3 min

Selecting the right open source model for workloads

How to evaluate open source repository formats to locate models perfectly fine-tuned for specialized reasoning, language interpolation, or multimodal pipelines.

#5 about 3 min

Reducing memory footprints through AI model quantization

Using quantization techniques to compress active weighting constants into lower precision scales for highly performant execution on consumer components.

#6 about 2 min

Integrating local AI models with existing business data

How tightly coupling locally hosted open source models directly with active codebase logic and documentation solves generic contextual boundaries.

#7 about 3 min

Running an AI model locally using Podman AI Lab

How to pull and serve an instruct-tuned conversational model via an offline desktop containerized testing playground.

#8 about 4 min

Building local RAG architectures using the Anything LLM tool

How intertwining a lightweight document vector database with a locally served language model wholly eliminates false inferences during queries.

#9 about 5 min

Setting up a local AI code assistant workspace

How configuring an open source code extension rapidly generates functional Python endpoints using entirely private offline compute constraints.

#10 about 5 min

Developing agentic AI applications using Model Context Protocol

How implementing systemic communication protocols alongside a Python execution framework equips local models to resolve deep external calculation tasks.

#11 about 1 min

Replacing commercial AI APIs with self-hosted open source models

How exchanging proprietary cloud APIs for strictly self-hosted containerized infrastructure yields massive deployment flexibility and vendor autonomy.

Matching moments

1:56 min

Running local open source models using Ollama

Aarno Aukia · LIVE

54 sec

Running generative AI models in local environments

Cedric Clyburn Cedric Clyburn · WWC 2024

3:14 min

Building a community-governed LAMP stack for open AI

Raffi Krikorian Raffi Krikorian · WWC Europe 2026

1:45 min

Introduction to serving large language models locally

Patrick Koss Patrick Koss · WWC 2025

1:39 min

Addressing data privacy concerns with local language models

5:11 min

Developing a containerized AI code assistant locally

Cedric Clyburn Cedric Clyburn +1 · WWC 2025

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Legare Kerrison, Cedric Clyburn

Legare Kerrison
Cedric Clyburn
Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

Understanding LLM Architectures: Inside the Design of Modern Models

Jofia Jose Prakash

Enterprise AI Architect at American Chemical Society

Jofia Jose Prakash
Open session

World Congress 2026 North America

Building Stuff with GenAI - The Open Minded Workshop beyond OpenAI

Andreas Erben

CTO for Applied AI and Metaverse at daenet

Andreas Erben
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot