World Congress 2026 Europe Jul 10, 2026 Session details

Fine-Tuning Small Language Models for Agentic AI

Björn Buchhold

Massive frontier models are bottlenecking your agentic AI workflows. Discover how fine-tuning specialized small language models on a single GPU slashes inference costs without sacrificing domain-specific performance.

Pause
Mute Enter Fullscreen
#1 about 7 min

The case for fine-tuning small models in agentic AI

Specialized small language models can replace large models for specific agentic subtasks to reduce costs and latency.

#2 about 4 min

Benchmarking text-to-SQL tasks with custom datasets

Custom training and evaluation datasets for manufacturing and hospital domains help benchmark agentic performance effectively.

#3 about 2 min

Comparing Claude Sonnet and Qwen for text-to-SQL generation

Claude Sonnet serves as the baseline large language model against candidate small model Qwen 4B.

#4 about 4 min

Methods for supervised fine-tuning and reinforcement learning

Data distillation enables supervised fine-tuning while GRPO offers reinforcement learning with verifiable rewards.

#5 about 3 min

Analyzing performance results on the manufacturing dataset

Supervised fine-tuning approaches the text-to-SQL quality of baseline large models on in-domain tasks.

#6 about 2 min

Implementing simple supervised fine-tuning with the Unsloth library

Specific training prompts and the Unsloth library streamline the process of fine-tuning quantized models.

#7 about 2 min

Evaluating model generalization on out-of-distribution schemas

Fine-tuned models excel narrowly in their trained domain but struggle to generalize to new schemas without reasoning skills.

#8 about 4 min

Designing effective reward functions for reinforcement learning algorithms

Evaluating SQL queries introduces nuance that makes building reliable reward functions for GRPO mathematically complex.

#9 about 3 min

Benchmark limitations and final recommendations for fine-tuning models

Despite evaluation limits, supervised fine-tuning offers clear return on investment for highly specialized sub-agent tasks.

#10 about 1 min

Answering questions on hardware requirements and GPU limitations

The Unsloth library enables local small language model training on a single A10 cloud GPU.

Matching moments

5:01 min

Leveraging large language models for code optimization and development

Stephan Gillich Stephan Gillich +3 · WWC 2024

3:35 min

Moving beyond generalist models with local fine tuning

Cedric Clyburn Cedric Clyburn +1 · WWC 2024

2:37 min

Understanding core parameters and mechanics of large language models

Julián Duque Julián Duque · WWC 2025

1:59 min

Identifying when narrow tasks require custom model fine-tuning

Viktoria Semaan Viktoria Semaan · WWC Europe 2026

1:21 min

Cost considerations of fine-tuning large language models

Kevin Klues Kevin Klues

1:45 min

Introduction to serving large language models locally

Patrick Koss Patrick Koss · WWC 2025

Upcoming sessions on this topic

Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Legare Kerrison, Cedric Clyburn

Legare Kerrison
Cedric Clyburn
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

Understanding LLM Architectures: Inside the Design of Modern Models

Jofia Jose Prakash

Enterprise AI Architect at American Chemical Society

Jofia Jose Prakash
Open session

World Congress 2026 North America

AI Agents are Only as Smart as their Context: Building a Real-Time Context Engine at Intuit

Bharat Patel

Lead Software Engineer at Intuit

Bharat Patel