World Congress 2024 Aug 20, 2024 Session details

Efficient deployment and inference of GPU-accelerated LLMs​

Adolf Hohl

Stop choosing between strict data privacy and deployment speed. Deploy optimized, GPU-accelerated LLMs securely to your infrastructure using NVIDIA NIM and LoRA adapters.

Pause
Mute Enter Fullscreen
#1 about 3 min

Transitioning generative AI from experimentation to production

The evolution of AI models from early breakthroughs to widespread production deployment stages.

#2 about 3 min

Evaluating managed AI services versus self-hosted infrastructure

Trade-offs between fully managed endpoints and building custom deployment infrastructure for complete environmental control.

#3 about 3 min

Streamlining model deployment with containerized microservice architectures

Utilizing containerized microservices as an immediate plug-and-play replacement for standardized application programming interfaces.

#4 about 2 min

Comparing pre-packaged microservices against manual model deployments

How pre-packaged solutions reduce deployment time from weeks to minutes while enforcing reliable interfaces.

#5 about 2 min

Maximizing hardware throughput via reduced model precision

Applying variable precision formats to substantially extract higher inference throughput and better economic sustainability.

#6 about 2 min

Abstracting low-level infrastructure with end-to-end software platforms

Obscuring complex accelerator communications to easily maintain reliable and scalable generative cloud pipelines.

#7 about 3 min

Core libraries driving inference engines and multi-GPU networking

How specialized runtime components manage hardware heuristics and parameter distribution across multiple distributed accelerators.

#8 about 3 min

Analyzing container hardware discovery and request batching workflows

The automated boot process where containers detect device topology before directing requests through specialized execution logic.

#9 about 3 min

Minimizing memory footprint using simultaneous adapter representation layers

Mounting multiple lightweight parameter adapters atop a base network to execute distinct application tasks concurrently.

#10 about 3 min

Handling optimal pre-built model configurations in isolated networks

The automated decision tree for selecting robust engine configurations across disconnected or distinct physical architectures.

Matching moments

2:04 min

Optimizing and deploying containerized AI inference workloads

Ankit Patel Ankit Patel · World Congress 2024

3:10 min

Scaling customized inference models with NVIDIA NIM

Anshul Jindal Anshul Jindal · World Congress 2025

2:50 min

Executing LoRA fine-tuning using serverless Databricks AI runtimes

Viktoria Semaan Viktoria Semaan · World Congress 2026 Europe

2:39 min

Open source tools for running and scaling models

Cedric Clyburn Cedric Clyburn +1 · World Congress 2025

7:28 min

Accelerating product features using generative large language models

David Singleton David Singleton +1 · Coffee With Developers

2:41 min

Transitioning artificial intelligence infrastructure into scalable commodity cloud services

juarezjunior juarezjunior · World Congress 2024

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 24, 2026 · 16:10–16:40

Stage 9

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Legare Kerrison, Cedric Clyburn

Legare Kerrison
Cedric Clyburn
Open session

World Congress 2026 North America

September 23, 2026 · 10:45–12:45

Stage 10

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

September 24, 2026 · 17:30–18:00

Stage 6

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

September 25, 2026 · 13:30–14:00

Mainstage

A Hands-On Developer Guide to Inference Engineering

Ankit Patel, Philip Kiely

Ankit Patel
Philip Kiely
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 1

Anatomy of an AI Request: Where Latency and Cost Are Really Born

Dan Fu

VP of Kernels at Together AI

Dan Fu
Open session

World Congress 2026 North America

September 24, 2026 · 13:30–14:00

Stage 9

Understanding LLM Architectures: Inside the Design of Modern Models

Jofia Jose Prakash

Director - AI & Governance at Humanity + AI, Inc

Jofia Jose Prakash