World Congress 2025

LLM Local Inference: Tools, Techniques, and Insights

July 11, 2025 13:55 – 14:05 · 10 min Airstream 1
ai llms python open source

What this session covers

Over the past few months, I’ve had the chance to work on a research and development project exploring the feasibility of adopting solutions powered by Small and Large Language Models—with one unique twist: inference had to be performed at the edge.

This journey led me to scratch the surface of the vast, fascinating, and rapidly evolving world of LLMs, with a particular focus on inference servers. Along the way, I encountered a fair share of head-scratching moments and valuable insights that I’m excited to share with you.

In this talk, we’ll cover the key concepts behind LLM inference, untangle the tricky jargon, and give you a glimpse into the primary tools and solutions you can leverage if you ever need or want to explore LLM local inference.

Related talks at this congress

Open session

World Congress 2025

July 10, 2025 · 14:50–15:20

Stage 6 - Red Hat

Self-Hosted LLMs: From Zero to Inference

Cedric Clyburn, Roberto Carratalá

Cedric Clyburn
Roberto Carratalá
Open session

World Congress 2025

July 11, 2025 · 13:00–13:30

Stage 7

Inside the Mind of an LLM

Emanuele Fabbiani

Head of AI at xtream

Emanuele Fabbiani
Open session

World Congress 2025

July 10, 2025 · 10:50–11:20

Stage 11

Exploring LLMs across clouds

Tomislav Tipurić

Chief Technology Officer, Nephos

Tomislav Tipurić
Open session

World Congress 2025

July 11, 2025 · 12:00–14:00

M4 (40 Seats)

Model Compression Techniques for Efficient LLM Deployment

Harshita Seth, Lavinia Ghita, Sergio Perez, Ziv Ilan

Harshita Seth
Lavinia Ghita
Sergio Perez
Ziv Ilan
All sessions at this congress