Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Maximizing open-source LLM throughput requires aggressively balancing compute and memory bottlenecks. Master the complete GPU optimization stack, from simple model quantization to sophisticated speculative decoding.
Matching moments
More from World Congress 2026 Europe
Related videos
Related articles
BB
Benedikt Bischof
LM
Luis Minvielle
DC
Daniel Cranney
BB
Benedikt Bischof
From learning to earning
Jobs that call for the skills explored in this talk.
about 1 month ago
•
Verified
LLM Training Engineer
Sciforium
San Francisco, United States
Expert
$155k–220k
Python
about 1 month ago
•
Verified
Lead Software Engineer, Model Serving Platform
Sciforium
San Francisco, United States
Expert
$230k–300k
Python
about 1 month ago
•
Verified
Distributed Training and Inference Engineer
Sciforium
San Francisco, United States
Expert
$190k–250k
Linux kernel
16 days ago
Principal Software Engineer, AI Inference Runtime
ARM
Seattle, WA, United States
Expert
$262k
Compilers
Low Latency
Concurrency
16 days ago
Principal Software Engineer, AI Inference Cloud
ARM
Seattle, WA, United States
Expert
$262k
Pytorch
TensorRT
Kubernetes
about 1 month ago
•
Verified
Model Implementation Engineer
Sciforium
San Francisco, United States
Expert
$165k–220k
Python