World Congress 2025
July 11, 2025 · 13:55–14:05
Airstream 1
LLM Local Inference: Tools, Techniques, and Insights
Riccardo Zoncada
Software Engineer @ xtream
World Congress 2025
As large language models (LLMs) become more accessible, running them locally unlocks exciting opportunities for developers, engineers, and privacy-focused users. Why rely on costly cloud AI services that share your data when you could deploy your own models tailored to your needs? In this session, we’ll dive into the advantages of local LLM deployment, from selecting the right open source model to optimizing performance on consumer hardware and integrating with your unique data. Let’s explore the journey to your own local stack for AI, and cover the important technical details such as model quantization, API integrations with IDE code assistants, and advanced methods like Retrieval-Augmented Generation (RAG) to connect your LLM to private data sources. Don’t miss out on the fun live demos that prove the bright future of open source AI is already here!
World Congress 2025
July 11, 2025 · 13:55–14:05
Airstream 1
Riccardo Zoncada
Software Engineer @ xtream
World Congress 2025
July 10, 2025 · 10:50–11:20
Stage 11
Tomislav Tipurić
Chief Technology Officer, Nephos
World Congress 2025
July 11, 2025 · 13:00–13:30
Stage 7
Emanuele Fabbiani
Head of AI at xtream
World Congress 2025
July 11, 2025 · 09:40–10:10
Stage 2
Harshita Seth, Sergio Perez