World Congress 2025 Aug 20, 2025 Session details

Adding knowledge to open-source LLMs

Harshita Seth , Sergio Perez

How do you inject proprietary knowledge into open-source LLMs? Master the pipeline of continued pre-training, supervised fine-tuning, and modern preference alignment to build tailored, domain-specific AI.

Pause
Mute Enter Fullscreen
#1 about 2 min

The core stages of large language model training

Understanding the pre-training and alignment stages establishes a foundation for how language models process and refine information.

#2 about 2 min

Why open-source language models need continuous knowledge updates

Updating models with current affairs and domain-specific knowledge prevents information degradation and improves task relevance.

#3 about 5 min

Using continued pre-training for domain-specific model adaptation

Applying self-supervised learning during continued pre-training ensures foundational models adapt accurately to highly specialized industry domains.

#4 about 7 min

Adding reasoning capabilities through supervised fine-tuning tasks

Structuring instruction datasets with chain-of-thought traces teaches neural networks to formulate explainable reasoning steps before answering.

#5 about 9 min

Aligning behavior and human preferences with reinforcement learning

Implementing comparative reward models and direct preference optimization aligns generative textual outputs with human behavioral expectations.

#6 about 2 min

Orchestrating model alignment with the NeMo RL framework

Adopting specialized automation frameworks simplifies the integration of complex reinforcement learning loops without engineering custom training layouts.

Matching moments

5:01 min

Leveraging large language models for code optimization and development

Stephan Gillich Stephan Gillich +3 · WWC 2024

2:37 min

Understanding core parameters and mechanics of large language models

Julián Duque Julián Duque · WWC 2025

2:08 min

Applying large language models to infrastructure tasks

Alfonso Sandoval Rosas Alfonso Sandoval Rosas · Europe 2026 Virtual

2:05 min

Enhancing language models with retrieval-augmented generation

Mary Grygleski Mary Grygleski · LIVE

1:37 min

Extending large language models without expensive retraining

Damir Damir · WWC 2025

1:16 min

Maturing through the machine learning development lifecycle

Xavier Portilla Edo · LIVE

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Understanding LLM Architectures: Inside the Design of Modern Models

Jofia Jose Prakash

Enterprise AI Architect at American Chemical Society

Jofia Jose Prakash
Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Legare Kerrison, Cedric Clyburn

Legare Kerrison
Cedric Clyburn
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

Headroom: A Context Optimization Layer for LLM Applications

Tejas Chopra

Senior Software Engineer at Netflix

Tejas Chopra
Open session

World Congress 2026 North America

Building Stuff with GenAI - The Open Minded Workshop beyond OpenAI

Andreas Erben

CTO for Applied AI and Metaverse at daenet

Andreas Erben