> Markdown version of [/videos/100352-fine-tuning-small-language-models-for-agentic-ai](https://www.wearedevelopers.com/videos/100352-fine-tuning-small-language-models-for-agentic-ai). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Fine-Tuning Small Language Models for Agentic AI Massive frontier models are bottlenecking your agentic AI workflows. Discover how fine-tuning specialized small language models on a single GPU slashes inference costs without sacrificing domain-specific performance. - **Speakers:** [Björn Buchhold](https://www.wearedevelopers.com/@bjorn-buchhold) - **Event:** World Congress 2026 Europe - **Published:** July 10, 2026 - **Duration:** 29:18 - **URL:** https://www.wearedevelopers.com/videos/100352-fine-tuning-small-language-models-for-agentic-ai ## Summary Agentic AI systems often rely on large, expensive frontier models for every subtask, creating latency, cost, and availability bottlenecks. However, many sub-agents perform highly narrow, predictable tasks—like translating natural language into SQL against a corporate data warehouse. By replacing large foundation models with specialized Small Language Models (SLMs) in the 4-billion parameter range, engineering teams can drastically reduce inference costs, maintain privacy over local hardware, and bypass broader chip shortages. Exploring fine-tuning strategies for these specialized agents highlights two primary approaches: data distillation via Supervised Fine-Tuning (SFT) and Reinforcement Learning from Verifiable Rewards (RLVR) utilizing algorithms like GRPO. Practical benchmarks on text-to-SQL tasks reveal that straightforward SFT on synthetic data generated by a "teacher" model easily achieves near-frontier quality on in-domain queries. The effort to execute this fine-tuning is exceptionally low, allowing developers to train quantized LoRA adapters securely on a single 24GB VRAM GPU using the Unsloth framework. While highly effective, SFT can cause a model to "unlearn" its intrinsic reasoning capabilities if not trained directly on thinking traces. Alternatively, RLVR successfully retains these reasoning strengths but relies on defining strictly verifiable custom reward functions, which can become complicated when parsing programmatic outputs like SQL queries without strict sequence rules. Ultimately, fine-tuning an SLM does not automatically yield a globally globally superior text-to-SQL master; instead, it delivers a profoundly optimized, domain-specific engine perfectly scaled for a specific segment of the enterprise architecture. **Keywords:** agentic AI, small language models, model fine-tuning, supervised fine-tuning, synthetic data distillation, RLVR, GRPO algorithm, text-to-SQL evaluation, unsloth framework, LoRA adapters, quantization, sub-agent architecture, domain-specific optimization, local LLM deployment, AI inference costs ## Chapters 1. **The case for fine-tuning small models in agentic AI** (00:52) — Specialized small language models can replace large models for specific agentic subtasks to reduce costs and latency. 1. **Benchmarking text-to-SQL tasks with custom datasets** (07:08) — Custom training and evaluation datasets for manufacturing and hospital domains help benchmark agentic performance effectively. 1. **Comparing Claude Sonnet and Qwen for text-to-SQL generation** (10:46) — Claude Sonnet serves as the baseline large language model against candidate small model Qwen 4B. 1. **Methods for supervised fine-tuning and reinforcement learning** (12:46) — Data distillation enables supervised fine-tuning while GRPO offers reinforcement learning with verifiable rewards. 1. **Analyzing performance results on the manufacturing dataset** (16:43) — Supervised fine-tuning approaches the text-to-SQL quality of baseline large models on in-domain tasks. 1. **Implementing simple supervised fine-tuning with the Unsloth library** (19:22) — Specific training prompts and the Unsloth library streamline the process of fine-tuning quantized models. 1. **Evaluating model generalization on out-of-distribution schemas** (20:37) — Fine-tuned models excel narrowly in their trained domain but struggle to generalize to new schemas without reasoning skills. 1. **Designing effective reward functions for reinforcement learning algorithms** (22:16) — Evaluating SQL queries introduces nuance that makes building reliable reward functions for GRPO mathematically complex. 1. **Benchmark limitations and final recommendations for fine-tuning models** (25:38) — Despite evaluation limits, supervised fine-tuning offers clear return on investment for highly specialized sub-agent tasks. 1. **Answering questions on hardware requirements and GPU limitations** (28:33) — The Unsloth library enables local small language model training on a single A10 cloud GPU. ## Related Moments - [Leveraging large language models for code optimization and development](https://www.wearedevelopers.com/videos/1106-the-future-of-computing-ai-technologies-in-the-exascale-era) (from "The Future of Computing: AI Technologies in the Exascale Era") - [Moving beyond generalist models with local fine tuning](https://www.wearedevelopers.com/videos/951-unlocking-the-power-of-ai-accessible-language-model-tuning-for-all) (from "Unlocking the Power of AI: Accessible Language Model Tuning for All") - [Understanding core parameters and mechanics of large language models](https://www.wearedevelopers.com/videos/1512-building-ai-applications-with-langchain-and-node-js) (from "Building AI Applications with LangChain and Node.js") - [Identifying when narrow tasks require custom model fine-tuning](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) (from "Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source") - [Cost considerations of fine-tuning large language models](https://www.wearedevelopers.com/videos/1170-from-foundation-model-to-hosted-ai-solution-in-minutes) (from "From foundation model to hosted AI solution in minutes") - [Introduction to serving large language models locally](https://www.wearedevelopers.com/videos/1619-unveiling-the-magic-scaling-large-language-models-to-serve-millions) (from "Unveiling the Magic: Scaling Large Language Models to Serve Millions") ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [WWC24 Talk - Scott Hanselman - AI: Superhero or Supervillain?](https://www.wearedevelopers.com/magazine/469-wwc24-talk-scott-hanselman-ai-superhero-or-supervillain) ## Related Jobs - [AI Software Engineer (Germany)](https://www.wearedevelopers.com/jobs/48317-ai-software-engineer-germany) at **Sunhat** - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/588393-machine-learning-engineer) at **Twilio** - [Senior AI Agent Software Engineer (Go, Python) (m/f/x)](https://www.wearedevelopers.com/jobs/48277-senior-ai-agent-software-engineer-go-python-m-f-x) at **Dynatrace** - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/1355348-machine-learning-engineer) at **TWILIO** - [Staff, Machine Learning Engineer (L4)](https://www.wearedevelopers.com/jobs/ext/1202639-staff-machine-learning-engineer-l4) at **Twilio** - [Head of AI Applications](https://www.wearedevelopers.com/jobs/ext/1231536-head-of-ai-applications) at **ZEISS Group**