World Congress 2026 Europe

Faster Together: Train and Deploy a Speculative Decoding Model for Low-Latency LLM Inference

July 9, 2026 15:30 – 17:30 · 120 min Room M2 (40 Seats)

What this session covers

Dive deep into theory and practice of low-latency inference by deploying NVIDIA TensorRT-LLM with advanced speculative decoding techniques. You’ll train an Eagle-3 draft head to propose candidate tokens efficiently, serve it, and benchmark it using AIPerf to quantify how these strategies minimize latency.

Workshop Preparation: - Please bring your own laptop. - Please review the following document and prepare accordingly before the workshop: https://developer.nvidia.com/dli/getready

Related talks at this congress

Open session

World Congress 2026 Europe

July 9, 2026 · 14:10–14:40

Stage 6 - powered by Microsoft

Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated

Christin Pohl

Global Black Belt Solution Engineer at Microsoft

Christin Pohl
Open session

World Congress 2026 Europe

July 9, 2026 · 13:00–15:00

Room M2 (40 Seats)

Accelerating AI Inference at Scale: A Deep Dive Into NVIDIA Dynamo on Kubernetes

Anshul Jindal, Mohak Chadha

Anshul Jindal
Mohak Chadha
Open session

World Congress 2026 Europe

July 10, 2026 · 14:45–16:45

Room M1 (60 Seats)

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Senior Developer Advocate at Akamai

Duan Lightfoot
Open session

World Congress 2026 Europe

July 10, 2026 · 15:40–16:10

Stage 6 - powered by Microsoft

The LLM Evolution: From Sequence Imitation to Verifiable Reasoning

Kamen Petroff

Software Developer at ATOS

Kamen Petroff
All sessions at this congress