World Congress 2026 Europe

Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source

July 9, 2026 16:10 – 16:40 · 30 min Stage 9

What this session covers

Let’s cut through the hype: most AI agents never make it past the demo stage. The gap between a working prototype and a production-grade system comes down to one thing—evaluation. Without reliable metrics, you’re guessing at what’s working, what needs fixing, and whether your agents are actually improving. You’ll learn how to: - Define custom metrics tailored to your use case - Calibrate LLM judges for cost-effective assessments - Track evaluation results over time to measure real progress

Whether you’re building LLM-powered apps or leading AI teams, you’ll leave with actionable tools to move from proof-of-concept to production—with the transparency and reliability enterprises demand.

Related talks at this congress

Open session

World Congress 2026 Europe

July 10, 2026 · 16:20–16:50

Stage 6 - powered by Microsoft

Fine-Tuning Small Language Models for Agentic AI

Björn Buchhold

Technology Evangelist at CID

Björn Buchhold
Open session

World Congress 2026 Europe

July 9, 2026 · 14:50–15:20

Stage 6 - powered by Microsoft

Rules, Heuristics, or LLMs? Lessons from Solving the Same Problem Twice

Artur Naumenko

Senior Software Engineer at Softeta

Artur Naumenko
Open session

World Congress 2026 Europe

July 10, 2026 · 15:00–15:30

Stage 6 - powered by Microsoft

LLMs in the wild: Building an AI agent that survives production

Steven Mi, Giampaolo Casolla

Steven Mi
Giampaolo Casolla
Open session

World Congress 2026 Europe

July 10, 2026 · 09:45–11:45

Room M8 (60 Seats)

From Hallucination to Justification: Hands-On Explainability for LLMs

Lucía Conde-Moreno, Tessel Haagen

Lucía Conde-Moreno
Tessel Haagen
All sessions at this congress