World Congress 2026 North America
Understanding LLM Architectures: Inside the Design of Modern Models
Jofia Jose Prakash
Enterprise AI Architect at American Chemical Society
World Congress 2026 North America
World Congress 2026 North America
September 23–25, 2026 · San José, CA
Attend in person
Get ticketsWatch remotely
Pro
Can’t make it to San José? Watch this session live with Pro. You also get:
Running large language models at scale can get expensive fast, but the right optimizations can cut latency and GPU costs dramatically.
We’ll walk through how to serve models efficiently using vLLM, an open-source, high-performance inference engine. Then and generate and test quantized models, expose them through vLLM’s OpenAI-compatible API, and tune runtime flags to balance throughput, latency, and accuracy on different GPUs.
We’ll benchmark performance live, inspect token-throughput metrics, and discuss real-world deployment trade-offs.
World Congress 2026 North America
Jofia Jose Prakash
Enterprise AI Architect at American Chemical Society
World Congress 2026 North America
David vonThenen
AI/ML Leader | Keynote Speaker | OSS Engineer & Developer Advocate | Agentic AI, Deep Learning, Production AI | Python, Go, C++
World Congress 2026 North America
Emmanuel Acheampong
Senior Manager Developer Relations at Crusoe AI
World Congress 2026 North America
Duan Lightfoot
Sr. AI Engineer, Akamai