World Congress 2025

Unveiling the Magic: Scaling Large Language Models to Serve Millions

July 11, 2025 15:40 – 16:10 · 30 min Stage 7
ai llms performance kubernetes scalability

What this session covers

Ever wondered what’s behind the curtain when you chat with AI like OpenAI’s ChatGPT?

At STACKIT, we’ve developed a model-serving service designed to host chat models, embedding models, and more, with a particular focus on supporting large language models.

In this talk, we’ll unveil the challenges and solutions of scaling large language models (LLMs) to serve millions. We’ll explore model acquisition, caching, building an OpenAI-compatible server, secure authentication, usage tracking for billing, autoscaling, and hosting multiple models under the domain.

Whether you’re a developer, engineer, or just curious, this talk will show you how every token counts in making AI accessible to millions.

Related talks at this congress

Open session

World Congress 2025

July 11, 2025 · 09:00–09:30

Stage 6 - Red Hat

One AI API to Power Them All

Roberto Carratalá

Principal AI Architect at Red Hat

Roberto Carratalá
Open session

World Congress 2025

July 10, 2025 · 16:10–16:40

Mainstage

How AI Models Get Smarter

Ankit Patel

Senior Director Developer Marketing at NVIDIA

Ankit Patel
Open session

World Congress 2025

July 10, 2025 · 10:50–11:20

Stage 11

Exploring LLMs across clouds

Tomislav Tipurić

Chief Technology Officer, Nephos

Tomislav Tipurić
Open session

World Congress 2025

July 10, 2025 · 13:30–14:00

Stage 5

Your Next AI Needs 10,000 GPUs. Now What?

Anshul Jindal, Martin Piercy

Anshul Jindal
Martin Piercy
All sessions at this congress