World Congress 2024

Efficient deployment and inference of GPU-accelerated LLMs​

July 18, 2024 11:30 – 12:00 · 30 min STAGE 11 (700)

What this session covers

NVIDIA TensorRT-LLM is an open-source software that delivers state-of-the-art performance for LLM serving using NVIDIA GPUs. It consists of the TensorRT deep learning compiler and includes optimized kernels, pre- and post-processing steps, and multi-GPU/multi-node communication primitives.​ During this session, I will present TensorRT-LLM features and capabilities and walk the audience through the steps needed to build and run a model in TensorRT-LLM on both single GPU and multi-GPUs. I will also show how to use TRT-LLM backend and Triton Inference Server for deployment.​

Related talks at this congress

Open session

World Congress 2024

July 18, 2024 · 14:50–15:20

Virtual Stage 1

Efficient Large Language Model Customization with NVIDIA NeMo Framework.

Miguel Martínez

Senior Deep Learning Data Scientist at NVIDIA

Miguel Martínez
Open session

World Congress 2024

July 19, 2024 · 14:20–14:50

STAGE 6 (120)

Intro to LLMs and recent developments

Christian Winkler

CEO of datanizing, research professor at TH Nürnberg

Christian Winkler
Open session

World Congress 2024

July 18, 2024 · 10:50–11:20

STAGE 6 (120)

Using LLMs in your Product

Daniel Töws

Software and GenAI Consultant at codecentric

Daniel Töws
Open session

World Congress 2024

July 18, 2024 · 16:00–18:00

Workshop room M4 (40)

Supercharge Inferencing of GenAI & LLM on AI PC

Adrian Boguszewski, Dmitriy Pastushenkov

Adrian Boguszewski
DP
All sessions at this congress