World Congress 2025 Aug 20, 2025 Session details

Adding knowledge to open-source LLMs

Harshita Seth , Sergio Perez

How do you inject proprietary knowledge into open-source LLMs? Master the pipeline of continued pre-training, supervised fine-tuning, and modern preference alignment to build tailored, domain-specific AI.

Pause
Mute Enter Fullscreen
#1 about 2 min

The core stages of large language model training

Understanding the pre-training and alignment stages establishes a foundation for how language models process and refine information.

#2 about 2 min

Why open-source language models need continuous knowledge updates

Updating models with current affairs and domain-specific knowledge prevents information degradation and improves task relevance.

#3 about 5 min

Using continued pre-training for domain-specific model adaptation

Applying self-supervised learning during continued pre-training ensures foundational models adapt accurately to highly specialized industry domains.

#4 about 7 min

Adding reasoning capabilities through supervised fine-tuning tasks

Structuring instruction datasets with chain-of-thought traces teaches neural networks to formulate explainable reasoning steps before answering.

#5 about 9 min

Aligning behavior and human preferences with reinforcement learning

Implementing comparative reward models and direct preference optimization aligns generative textual outputs with human behavioral expectations.

#6 about 2 min

Orchestrating model alignment with the NeMo RL framework

Adopting specialized automation frameworks simplifies the integration of complex reinforcement learning loops without engineering custom training layouts.

Matching moments

5:01 min

Leveraging large language models for code optimization and development

Stephan Gillich Stephan Gillich +3 · World Congress 2024

2:37 min

Understanding core parameters and mechanics of large language models

Julián Duque Julián Duque · World Congress 2025

2:08 min

Applying large language models to infrastructure tasks

Alfonso Sandoval Rosas Alfonso Sandoval Rosas · Europe 2026 Virtual

2:05 min

Enhancing language models with retrieval-augmented generation

Mary Grygleski Mary Grygleski · LIVE

1:37 min

Extending large language models without expensive retraining

Damir Damir · World Congress 2025

1:16 min

Maturing through the machine learning development lifecycle

Xavier Portilla Edo · LIVE

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 24, 2026 · 13:30–14:00

Stage 9

Understanding LLM Architectures: Inside the Design of Modern Models

Jofia Jose Prakash

Director - AI & Governance at Humanity + AI, Inc

Jofia Jose Prakash
Open session

World Congress 2026 North America

September 24, 2026 · 17:30–18:00

Stage 6

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

September 24, 2026 · 16:10–16:40

Stage 9

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Legare Kerrison, Cedric Clyburn

Legare Kerrison
Cedric Clyburn
Open session

World Congress 2026 North America

September 25, 2026 · 11:40–12:10

Stage 9

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

September 25, 2026 · 14:10–14:40

Stage 4

Headroom: A Context Optimization Layer for LLM Applications

Tejas Chopra

Senior Software Engineer at Netflix

Tejas Chopra
Open session

World Congress 2026 North America

September 24, 2026 · 10:20–10:50

Stage 3

Reading the Mind of Your NGINX Fleet: A Hybrid Rule + ML Pipeline for NGINX Config Intelligence

Brandon LEE

AI Architect Lead at F5

Brandon LEE