World Congress 2025

Self-Hosted LLMs: From Zero to Inference

July 10, 2025 14:50 – 15:20 · 30 min Stage 6 - Red Hat

What this session covers

As large language models (LLMs) become more accessible, running them locally unlocks exciting opportunities for developers, engineers, and privacy-focused users. Why rely on costly cloud AI services that share your data when you could deploy your own models tailored to your needs? In this session, we’ll dive into the advantages of local LLM deployment, from selecting the right open source model to optimizing performance on consumer hardware and integrating with your unique data. Let’s explore the journey to your own local stack for AI, and cover the important technical details such as model quantization, API integrations with IDE code assistants, and advanced methods like Retrieval-Augmented Generation (RAG) to connect your LLM to private data sources. Don’t miss out on the fun live demos that prove the bright future of open source AI is already here!

Related talks at this congress

Open session

World Congress 2025

July 11, 2025 · 13:55–14:05

Airstream 1

LLM Local Inference: Tools, Techniques, and Insights

Riccardo Zoncada

Software Engineer @ xtream

Riccardo Zoncada
Open session

World Congress 2025

July 10, 2025 · 10:50–11:20

Stage 11

Exploring LLMs across clouds

Tomislav Tipurić

Chief Technology Officer, Nephos

Tomislav Tipurić
Open session

World Congress 2025

July 11, 2025 · 13:00–13:30

Stage 7

Inside the Mind of an LLM

Emanuele Fabbiani

Head of AI at xtream

Emanuele Fabbiani
Open session

World Congress 2025

July 11, 2025 · 09:40–10:10

Stage 2

Adding knowledge to open-source LLMs

Harshita Seth, Sergio Perez

Harshita Seth
Sergio Perez
All sessions at this congress