World Congress 2024
July 18, 2024 · 14:50–15:20
Virtual Stage 1
Efficient Large Language Model Customization with NVIDIA NeMo Framework.
Miguel Martínez
Senior Deep Learning Data Scientist at NVIDIA
World Congress 2024
NVIDIA TensorRT-LLM is an open-source software that delivers state-of-the-art performance for LLM serving using NVIDIA GPUs. It consists of the TensorRT deep learning compiler and includes optimized kernels, pre- and post-processing steps, and multi-GPU/multi-node communication primitives. During this session, I will present TensorRT-LLM features and capabilities and walk the audience through the steps needed to build and run a model in TensorRT-LLM on both single GPU and multi-GPUs. I will also show how to use TRT-LLM backend and Triton Inference Server for deployment.
World Congress 2024
July 18, 2024 · 14:50–15:20
Virtual Stage 1
Miguel Martínez
Senior Deep Learning Data Scientist at NVIDIA
World Congress 2024
July 19, 2024 · 14:20–14:50
STAGE 6 (120)
Christian Winkler
CEO of datanizing, research professor at TH Nürnberg
World Congress 2024
July 18, 2024 · 10:50–11:20
STAGE 6 (120)
Daniel Töws
Software and GenAI Consultant at codecentric
World Congress 2024
July 18, 2024 · 16:00–18:00
Workshop room M4 (40)
Adrian Boguszewski, Dmitriy Pastushenkov