World Congress 2024 Aug 22, 2024 Session details

Performant Architecture for a Fast Gen AI User Experience

Nathaniel Okenwa

Multi-second latency ruins real-time voice apps. Stop relying on naive batch processing. Discover how semantic chunking and edge models combine to deliver instantaneous generative AI experiences.

Pause
Mute Enter Fullscreen
#1 about 6 min

Building a real-time voice translator

A first attempt to create a live conversation translator using basic cloud integrations results in heavy processing delays.

#2 about 4 min

Upgrading the stack with generative APIs

Swapping legacy models for targeted speech generation platforms lowers base latency before deploying deeper architectural improvements.

#3 about 2 min

Lowering pipeline latency with data streaming

Piping inputs between services immediately instead of waiting for full responses eliminates long waterfall bottlenecks.

#4 about 3 min

Balancing speed and context via data chunking

Grouping streamed tokens by sentence boundaries preserves the linguistic context required for high-quality audio generation.

#5 about 6 min

Running edge models to eliminate network trips

Deploying lightweight language models directly alongside the primary processing server eradicates transatlantic API roundtrips.

#6 about 4 min

Using semantic caching and pre-generated audio

Storing pre-computed assets and querying natural language intent via vector databases avoids redundant processing for common interactions.

#7 about 2 min

Minimizing inference time with prompt optimization

Imposing structural constraints on output tokens and parallelizing auxiliary tasks squeezes out the last milliseconds of processing time.

#8 about 1 min

Designing responsive AI loading states

Informing users with progress indicators and transparent application states mitigates the perceived wait time for generative outputs.

#9 about 2 min

Architecture summary and handling output limits

A review of core generative architecture patterns is followed by insights on triggering programmatic buffer flushes.

Matching moments

1:22 min

Leveraging the comprehensive generative artificial intelligence stack

Duan Lightfoot Duan Lightfoot · WWC 2024

1:10 min

Architecting a web real-time communication stack for agents

Nathaniel Okenwa Nathaniel Okenwa · WWC 2025

5:28 min

Integrating OpenAI for real-time streaming responses

Chris Heilmann +3 · LIVE

4:03 min

Enhancing conversational output via generative AI models

Xavier Portilla Edo · LIVE

4:16 min

Addressing latency and architecture in voice agents

Chris Heilmann +2 · LIVE

1:38 min

Scaling bottlenecks in generative AI applications

Stan Girard Stan Girard · WWC 2024

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Building Stuff with GenAI - The Open Minded Workshop beyond OpenAI

Andreas Erben

CTO for Applied AI and Metaverse at daenet

Andreas Erben
Open session

World Congress 2026 North America

From Software Agents to Physical Devices: Inside the Agentic Hardware Stack

Michael Yuan, Vivian Hu

Michael Yuan
Vivian Hu
Open session

World Congress 2026 North America

AI Agents are Only as Smart as their Context: Building a Real-Time Context Engine at Intuit

Bharat Patel

Lead Software Engineer at Intuit

Bharat Patel
Open session

World Congress 2026 North America

When Humans Stop Writing Code: Rethinking Languages, Compilers, and Responsibility

Simon Auer

Organizer of flutter vienna meetup and CEO of marqably

Simon Auer
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong