WeAreDevelopers LIVE Aug 27, 2025

WeAreDevelopers LIVE - Vector Similarity Search Patterns for Efficiency and more

Chris Heilmann , Daniel Cranney , Raphael De Lio , Advocate At Redis

Stop flooding LLMs with redundant data. Leverage vector similarity search for semantic routing and caching to drastically reduce token costs while slashing processing times.

Pause
Mute Enter Fullscreen
#1 about 3 min

Welcome and updates from the WeAreDevelopers team

Recent platform updates include publishing conference talks and improving video topic searchability.

#2 about 7 min

Building visibility through open source and community engagement

Sharing technical knowledge publicly often leads to unexpected career opportunities and streamlined interviews.

#3 about 4 min

Challenging marketing claims with technical analysis and benchmarking

Proving accepted narratives wrong through independent testing demonstrates strong engineering capabilities.

#4 about 9 min

Mitigating language model costs with vector search patterns

Using vector databases to handle semantic embeddings reduces repetitive token costs and execution delays.

#5 about 9 min

Classifying social media text using semantic classification patterns

Leveraging vector similarity to categorize posts efficiently avoids the latency of external API calls.

#6 about 6 min

Optimizing tool calling in chatbots with semantic routing

Mapping user intents directly to execution functions by vectorizing predefined query triggers prevents unnecessary language model invocations.

#7 about 6 min

Implementing semantic caching for repetitive chatbot user queries

Storing query vectors and their corresponding responses prevents redundant processing for semantically identical questions.

#8 about 10 min

Tuning vector search accuracy and parameter extraction strategies

Applying retrieval optimizers and self-improvement feedback loops significantly improves the precision of semantic matches.

#9 about 7 min

Handling stale data and caching errors in chatbots

Strategies like time-to-live expirations and routing blocks prevent systems from serving outdated or hallucinated information.

#10 about 6 min

Managing vector database architecture and data lifecycle ownership

Enterprise vector workloads require proper distributed clustering and ongoing management by dedicated data engineers.

#11 about 5 min

Balancing convenience and costs in applied artificial intelligence

Developers should utilize robust optimization patterns instead of blindly relying on expensive automated subscription platforms.

Matching moments

3:37 min

Demonstrating semantic latency reductions using Spring AI configurations

1:48 min

Reusing language model responses through semantic database caching

3:13 min

Using semantic caching and pre-generated audio

Nathaniel Okenwa Nathaniel Okenwa · WWC 2024

1:39 min

Enabling semantic search with automated query vectorization

Iulia Feroli Iulia Feroli · LIVE

2:38 min

Enabling contextual responses with retrieval-augmented generation and vector databases

David Leconte David Leconte +1 · WWC 2024

3:08 min

Scaling semantic search with Astra DB and Apache Cassandra

David Leconte David Leconte +1 · WWC 2024

Upcoming sessions on this topic

Open session

World Congress 2026 North America

KV Cache Is Not About Speed: It's About Surviving Inference Costs

David vonThenen

AI/ML Leader | Keynote Speaker | OSS Engineer & Developer Advocate | Agentic AI, Deep Learning, Production AI | Python, Go, C++

David vonThenen
Open session

World Congress 2026 North America

Headroom: A Context Optimization Layer for LLM Applications

Tejas Chopra

Senior Software Engineer at Netflix

Tejas Chopra
Open session

World Congress 2026 North America

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Cedric Clyburn, Legare Kerrison

Cedric Clyburn
Legare Kerrison
Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

AI Agents are Only as Smart as their Context: Building a Real-Time Context Engine at Intuit

Bharat Patel

Lead Software Engineer at Intuit

Bharat Patel
Open session

World Congress 2026 North America

Understanding LLM Architectures: Inside the Design of Modern Models

Jofia Jose Prakash

Enterprise AI Architect at American Chemical Society

Jofia Jose Prakash