World Congress 2024 Aug 20, 2024 Session details

Efficient deployment and inference of GPU-accelerated LLMs​

Adolf Hohl

Stop choosing between strict data privacy and deployment speed. Deploy optimized, GPU-accelerated LLMs securely to your infrastructure using NVIDIA NIM and LoRA adapters.

Pause
Mute Enter Fullscreen
#1 about 3 min

Transitioning generative AI from experimentation to production

The evolution of AI models from early breakthroughs to widespread production deployment stages.

#2 about 3 min

Evaluating managed AI services versus self-hosted infrastructure

Trade-offs between fully managed endpoints and building custom deployment infrastructure for complete environmental control.

#3 about 3 min

Streamlining model deployment with containerized microservice architectures

Utilizing containerized microservices as an immediate plug-and-play replacement for standardized application programming interfaces.

#4 about 2 min

Comparing pre-packaged microservices against manual model deployments

How pre-packaged solutions reduce deployment time from weeks to minutes while enforcing reliable interfaces.

#5 about 2 min

Maximizing hardware throughput via reduced model precision

Applying variable precision formats to substantially extract higher inference throughput and better economic sustainability.

#6 about 2 min

Abstracting low-level infrastructure with end-to-end software platforms

Obscuring complex accelerator communications to easily maintain reliable and scalable generative cloud pipelines.

#7 about 3 min

Core libraries driving inference engines and multi-GPU networking

How specialized runtime components manage hardware heuristics and parameter distribution across multiple distributed accelerators.

#8 about 3 min

Analyzing container hardware discovery and request batching workflows

The automated boot process where containers detect device topology before directing requests through specialized execution logic.

#9 about 3 min

Minimizing memory footprint using simultaneous adapter representation layers

Mounting multiple lightweight parameter adapters atop a base network to execute distinct application tasks concurrently.

#10 about 3 min

Handling optimal pre-built model configurations in isolated networks

The automated decision tree for selecting robust engine configurations across disconnected or distinct physical architectures.

Matching moments

2:04 min

Optimizing and deploying containerized AI inference workloads

Ankit Patel Ankit Patel · WWC 2024

3:10 min

Scaling customized inference models with NVIDIA NIM

Anshul Jindal Anshul Jindal · WWC 2025

2:50 min

Executing LoRA fine-tuning using serverless Databricks AI runtimes

Viktoria Semaan Viktoria Semaan · WWC Europe 2026

2:39 min

Open source tools for running and scaling models

Cedric Clyburn Cedric Clyburn +1 · WWC 2025

7:28 min

Accelerating product features using generative large language models

David Singleton David Singleton +1 · Coffee With Developers

2:41 min

Transitioning artificial intelligence infrastructure into scalable commodity cloud services

juarezjunior juarezjunior · WWC 2024

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Legare Kerrison, Cedric Clyburn

Legare Kerrison
Cedric Clyburn
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

Understanding LLM Architectures: Inside the Design of Modern Models

Jofia Jose Prakash

Enterprise AI Architect at American Chemical Society

Jofia Jose Prakash
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

Trust, But Verify: Continuous GPU Validation at Scale

Kyle Bell

VP of AI @ TensorWave

Kyle Bell