World Congress 2026 Europe Jul 9, 2026 Session details

Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source

Viktoria Semaan

Relying on proprietary LLMs is like hiring a PhD for an intern's task. Learn how fine-tuning compact open-source models maintains top-tier accuracy while slashing production costs by 86%.

Pause
Mute Enter Fullscreen
#1 about 2 min

The shift toward token efficiency and independent model layers

Decoupling the application layer from the model serving layer mitigates risks caused by the short life cycle of commercial models.

#2 about 2 min

Cursor case study on solving high task costs

Fine-tuning an open weight model explicitly for coding enabled an eighty-six percent reduction in token usage.

#3 about 3 min

Defining an AI moderation use case for evaluation setup

Establishing message classification and human review routing tasks demonstrates practical framework assessments.

#4 about 4 min

Trade-offs between proprietary, open weight, and fine-tuned models

Navigating model selection requires balancing organizational control, hosting responsibilities, and specific task complexities.

#5 about 5 min

Programmatic model evaluation and custom metrics via MLflow

Using an LLM as a judge alongside clear guidelines within MLflow creates an automated baseline for continuous improvement.

#6 about 3 min

Managing traffic and tracking costs with Databricks Unity Catalog

Implementing a centralized gateway controls API spending and enables intelligent traffic splitting between different parameter sizes.

#7 about 2 min

Analyzing evaluation dashboard results for optimal model selection

Comparing latency, benchmark accuracy, and operational costs reveals cases where smaller open-source models outclass proprietary peers.

#8 about 2 min

Identifying when narrow tasks require custom model fine-tuning

Determining when to transition from prompt engineering and retrieval-augmented generation to supervised fine-tuning reduces latency and API dependency.

#9 about 3 min

Executing LoRA fine-tuning using serverless Databricks AI runtimes

Utilizing low-rank adaptation on serverless GPUs minimizes infrastructure overhead while efficiently training smaller architectures like Qwen.

#10 about 2 min

Comparing resulting task performance across varied model classifications

While fine-tuned implementations dramatically lower classification costs, larger commercial deployments remain necessary for complex reasoning and context ranking.

#11 about 2 min

Key takeaways and accessing the Databricks developer toolkit

Adopting a standardized evaluation protocol prevents costly assumptions and is highly accessible via the provided community tier resources.

#12 about 2 min

Addressing system transparency when utilizing an LLM judge

Leveraging deep evaluation traces clarifies the logical steps and exact reasoning behind individual automated assessment scores.

Matching moments

3:53 min

Evaluating and observing large language model performance at scale

Krzysztof Cieślak Krzysztof Cieślak · WWC 2024

3:26 min

Introducing LLMs as judges for automated testing

Sebastian Messingfeld Sebastian Messingfeld · WWC Europe 2026

2:36 min

Addressing AI observability and cost transparency with Langfuse

Hellmar Becker Hellmar Becker · WWC Europe 2026

3:22 min

Evaluating advanced artificial intelligence platforms for daily recruitment

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

1:37 min

Managing AI costs with open-source models

Laurie Voss · Coffee With Developers

3:14 min

Balancing performance and costs with custom language models

Jemiah Sius Jemiah Sius · WWC 2024

Upcoming sessions on this topic

Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Legare Kerrison, Cedric Clyburn

Legare Kerrison
Cedric Clyburn
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

Understanding LLM Architectures: Inside the Design of Modern Models

Jofia Jose Prakash

Enterprise AI Architect at American Chemical Society

Jofia Jose Prakash
Open session

World Congress 2026 North America

AI ROI: The Hard Unit Economics of AI-Native Engineering

Manu Gurudatha

Manu Gurudatha, VP of Engineering at PagerDuty

Manu Gurudatha