World Congress 2026 Europe Jul 10, 2026 Session details

Fine-Tuning Small Language Models for Agentic AI

Björn Buchhold

Massive frontier models are bottlenecking your agentic AI workflows. Discover how fine-tuning specialized small language models on a single GPU slashes inference costs without sacrificing domain-specific performance.

Pause
Mute Enter Fullscreen
#1 about 7 min

The case for fine-tuning small models in agentic AI

Specialized small language models can replace large models for specific agentic subtasks to reduce costs and latency.

#2 about 4 min

Benchmarking text-to-SQL tasks with custom datasets

Custom training and evaluation datasets for manufacturing and hospital domains help benchmark agentic performance effectively.

#3 about 2 min

Comparing Claude Sonnet and Qwen for text-to-SQL generation

Claude Sonnet serves as the baseline large language model against candidate small model Qwen 4B.

#4 about 4 min

Methods for supervised fine-tuning and reinforcement learning

Data distillation enables supervised fine-tuning while GRPO offers reinforcement learning with verifiable rewards.

#5 about 3 min

Analyzing performance results on the manufacturing dataset

Supervised fine-tuning approaches the text-to-SQL quality of baseline large models on in-domain tasks.

#6 about 2 min

Implementing simple supervised fine-tuning with the Unsloth library

Specific training prompts and the Unsloth library streamline the process of fine-tuning quantized models.

#7 about 2 min

Evaluating model generalization on out-of-distribution schemas

Fine-tuned models excel narrowly in their trained domain but struggle to generalize to new schemas without reasoning skills.

#8 about 4 min

Designing effective reward functions for reinforcement learning algorithms

Evaluating SQL queries introduces nuance that makes building reliable reward functions for GRPO mathematically complex.

#9 about 3 min

Benchmark limitations and final recommendations for fine-tuning models

Despite evaluation limits, supervised fine-tuning offers clear return on investment for highly specialized sub-agent tasks.

#10 about 1 min

Answering questions on hardware requirements and GPU limitations

The Unsloth library enables local small language model training on a single A10 cloud GPU.

Matching moments

5:01 min

Leveraging large language models for code optimization and development

Stephan Gillich Stephan Gillich +3 · World Congress 2024

3:35 min

Moving beyond generalist models with local fine tuning

Cedric Clyburn Cedric Clyburn +1 · World Congress 2024

2:37 min

Understanding core parameters and mechanics of large language models

Julián Duque Julián Duque · World Congress 2025

1:59 min

Identifying when narrow tasks require custom model fine-tuning

Viktoria Semaan Viktoria Semaan · World Congress 2026 Europe

1:21 min

Cost considerations of fine-tuning large language models

Kevin Klues Kevin Klues

1:45 min

Introduction to serving large language models locally

Patrick Koss Patrick Koss · World Congress 2025

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 24, 2026 · 17:30–18:00

Stage 6

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

September 24, 2026 · 16:10–16:40

Stage 9

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Legare Kerrison, Cedric Clyburn

Legare Kerrison
Cedric Clyburn
Open session

World Congress 2026 North America

September 25, 2026 · 11:40–12:10

Stage 9

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

September 23, 2026 · 10:45–12:45

Stage 10

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 5

Edge AI: Running Agentic Intelligence Where Internet Can't Reach

Nitin Eusebius

AWS - Principal Solutions Architect

Nitin Eusebius
Open session

World Congress 2026 North America

September 24, 2026 · 13:30–14:00

Stage 9

Understanding LLM Architectures: Inside the Design of Modern Models

Jofia Jose Prakash

Director - AI & Governance at Humanity + AI, Inc

Jofia Jose Prakash