World Congress 2026 North America
KV Cache Is Not About Speed: It's About Surviving Inference Costs
David vonThenen
AI/ML Leader | Keynote Speaker | OSS Engineer & Developer Advocate | Agentic AI, Deep Learning, Production AI | Python, Go, C++
Chris Heilmann , Daniel Cranney , Raphael De Lio , Advocate At Redis
Stop flooding LLMs with redundant data. Leverage vector similarity search for semantic routing and caching to drastically reduce token costs while slashing processing times.
Jobs that call for the skills explored in this talk.