World Congress 2026 Europe

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

July 10, 2026 14:45 – 16:45 · 120 min Room M1 (60 Seats)

What this session covers

Every production agent today is renting its intelligence. You’re paying per token, sending your customer’s data to someone else’s servers, and hoping the provider doesn’t rate-limit you during your launch. For most teams, that’s fine. But for a growing number of teams in regulated industries, with high-volume products, latency-sensitive workloads, or rising token bills, it’s starting to look like a liability.

In this 120-minute hands-on workshop you’ll get a dedicated GPU and build an agent that runs on infrastructure you control. You’ll stand up vLLM, point your agent at it, and drive concurrent load through the stack until you can see batching, KV cache pressure, and throughput limits in the metrics. Then you’ll optimize the deployment to improve throughput while keeping per-request latency in line.

The focus isn’t agent frameworks. It’s the inference layer underneath them. You’ll leave with working code and a real understanding of continuous batching under real concurrency, KV cache tradeoffs, vLLM’s metrics, and the bottlenecks that only show up when you operate the inference server yourself.

Related talks at this congress

Open session

World Congress 2026 Europe

July 8, 2026 · 13:30–15:30

Room R3 (30 Seats)

Build a Production-Ready AI Agent in 90 Minutes

Tamas Piros

AI Consultant

Tamas Piros
Open session

World Congress 2026 Europe

July 9, 2026 · 13:00–15:00

Room R2 (30 Seats)

Build a Production-Ready AI Agent in 90 Minutes

Tamas Piros

AI Consultant

Tamas Piros
Open session

World Congress 2026 Europe

July 9, 2026 · 13:00–15:00

Room M2 (40 Seats)

Accelerating AI Inference at Scale: A Deep Dive Into NVIDIA Dynamo on Kubernetes

Anshul Jindal, Mohak Chadha

Anshul Jindal
Mohak Chadha
Open session

World Congress 2026 Europe

July 9, 2026 · 14:10–14:40

Stage 6 - powered by Microsoft

Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated

Christin Pohl

Global Black Belt Solution Engineer at Microsoft

Christin Pohl
All sessions at this congress