World Congress 2026 North America
September 24, 2026 · 13:30–14:00
Stage 9
Understanding LLM Architectures: Inside the Design of Modern Models
Jofia Jose Prakash
Director - AI & Governance at Humanity + AI, Inc
World Congress 2026 North America
World Congress 2026 North America
September 23–25, 2026 · San José, CA
Attend in person
Get ticketsWatch remotely
Pro
Can’t make it to San José? Watch this session live with Pro. You also get:
Running large language models at scale can get expensive fast, but the right optimizations can cut latency and GPU costs dramatically.
We’ll walk through how to serve models efficiently using vLLM, an open-source, high-performance inference engine. Then and generate and test quantized models, expose them through vLLM’s OpenAI-compatible API, and tune runtime flags to balance throughput, latency, and accuracy on different GPUs.
We’ll benchmark performance live, inspect token-throughput metrics, and discuss real-world deployment trade-offs.
World Congress 2026 North America
September 24, 2026 · 13:30–14:00
Stage 9
Jofia Jose Prakash
Director - AI & Governance at Humanity + AI, Inc
World Congress 2026 North America
September 24, 2026 · 14:10–14:40
Stage 1
Dan Fu
VP of Kernels at Together AI
World Congress 2026 North America
September 24, 2026 · 15:30–16:00
Stage 2
David vonThenen
AI/ML Leader | Keynote Speaker | OSS Engineer & Developer Advocate | Agentic AI, Deep Learning, Production AI | Python, Go, C++
World Congress 2026 North America
September 24, 2026 · 17:30–18:00
Stage 6
Emmanuel Acheampong
Senior Manager Developer Relations at Crusoe AI