Coffee With Developers Jun 17, 2026

Speech-to-Speech AI models - Marius Obert

Marius Obert

Marius Obert argues traditional transcription pipelines destroy essential conversational context. Discover how native speech-to-speech models preserve emotional nuances and deploy multimodal voice agents in under 70 lines of code.

Pause
Mute Enter Fullscreen
#1 about 2 min

Comparing transcription-based voice pipelines to native models

Routing audio through separate transcription and generation steps strips conversational nuances like hesitations.

#2 about 2 min

Advantages of using native speech-to-speech streaming models

Streaming audio directly through unified models improves latency and maintains consistent voices across languages.

#3 about 2 min

Loss of developer control in black box voice models

Unified voice models abstract away intermediary steps, making it difficult to inject custom logic or webhooks.

#4 about 2 min

Managing token costs and multimodal context in voice agents

Streaming cheaper text tokens alongside audio inputs allows developers to provide rich context efficiently.

#5 about 2 min

Detecting emotional intent directly from voice audio streams

Analyzing raw voice inputs enables models to identify subtle nuances like frustration that text transcriptions lose.

#6 about 2 min

Detecting answering machines and preventing automated voice fraud

Differentiating between human speakers and automated systems remains a complex challenge for continuous outbound calling.

#7 about 2 min

Integrating speech models using the Twilio real-time SDK

New tooling simplifies telephony integration by automatically handling codec translation and sample rate conversions.

Matching moments

7:39 min

Streaming text responses from an AI language model

Chris Heilmann +2 · LIVE

2:11 min

Analyzing voice interface research projects and technical limitations

Tobias Münch Tobias Münch · World Congress 2024

58 sec

Enabling live translations with AI voice-to-text pipelines

Alex Goldman Alex Goldman · Coffee With Developers

42 sec

Delivering dynamic voice interactions via ElevenLabs

Alexandra Mihai Alexandra Mihai +2 · Europe 2026 Virtual

1:44 min

Summarizing developer experience and artificial intelligence companions

Robert Hoffmann Robert Hoffmann +1 · LIVE

2:05 min

Challenges of managing voice architecture and latency

Chris Heilmann +3 · LIVE

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 24, 2026 · 11:00–11:30

Stage 4

From Voice Demo to Enterprise Production

Anahita Havewala, Anuj Gupta

Anahita Havewala
Anuj Gupta
Open session

World Congress 2026 North America

September 24, 2026 · 11:20–11:25

Outdoor Stage

Finding the Edges: Testing, Evaluating, and Monitoring Voice AI Agents Before Your Users Do

Matt Wyman

CEO at Okareo

Matt Wyman
Open session

World Congress 2026 North America

September 24, 2026 · 17:30–18:00

Stage 6

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

September 24, 2026 · 11:00–11:30

Stage 2

From Stateless to Stateful: Real-Time Voice & Messaging Agents with Twilio and AWS

Rishab Kumar

Staff Developer Evangelist @ Twilio

Rishab Kumar
Open session

World Congress 2026 North America

September 25, 2026 · 09:00–09:30

Stage 7

AI Agents are Only as Smart as their Context: Building a Real-Time Context Engine at Intuit

Bharat Patel

Lead Software Engineer at Intuit

Bharat Patel
Open session

World Congress 2026 North America

September 25, 2026 · 10:20–10:50

Stage 5

Physical AI: 5 Things You Can Build That Aren't Another Chatbot

Vini Senger

Senior Technical Evangelist for Startups

Vini Senger