> Markdown version of [/events/world-congress-2024/sessions/73-efficient-deployment](https://www.wearedevelopers.com/events/world-congress-2024/sessions/73-efficient-deployment). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Efficient deployment and inference of GPU-accelerated LLMs​ - **Date:** Thursday, Jul 18, 2024 - **Time:** 11:30–12:00 (30 min) - **Room:** STAGE 11 (700) - **Event:** World Congress 2024 ## Recording [Watch recording](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Description NVIDIA TensorRT-LLM is an open-source software that delivers state-of-the-art performance for LLM serving using NVIDIA GPUs. It consists of the TensorRT deep learning compiler and includes optimized kernels, pre- and post-processing steps, and multi-GPU/multi-node communication primitives.​ During this session, I will present TensorRT-LLM features and capabilities and walk the audience through the steps needed to build and run a model in TensorRT-LLM on both single GPU and multi-GPUs. I will also show how to use TRT-LLM backend and Triton Inference Server for deployment.​ ## Speaker ### [Adolf Hohl](https://www.wearedevelopers.com/@adolf-hohl) Sr. Mgr. Solution Architects AUTO Enterprise ## Related talks at this congress - [Efficient Large Language Model Customization with NVIDIA NeMo Framework.](https://www.wearedevelopers.com/events/world-congress-2024/sessions/170-efficient-large) — Miguel Martínez - [Intro to LLMs and recent developments](https://www.wearedevelopers.com/events/world-congress-2024/sessions/397-intro-to-llms-and) — Christian Winkler - [Using LLMs in your Product](https://www.wearedevelopers.com/events/world-congress-2024/sessions/50-using-llms-in-your) — Daniel Töws - [Supercharge Inferencing of GenAI & LLM on AI PC](https://www.wearedevelopers.com/events/world-congress-2024/sessions/196-supercharge) — Adrian Boguszewski, Dmitriy Pastushenkov