Coffee With Developers Jun 17, 2026

Speech-to-Speech AI models - Marius Obert

Marius Obert

Marius Obert argues traditional transcription pipelines destroy essential conversational context. Discover how native speech-to-speech models preserve emotional nuances and deploy multimodal voice agents in under 70 lines of code.

Pause
Mute Enter Fullscreen
#1 about 2 min

Comparing transcription-based voice pipelines to native models

Routing audio through separate transcription and generation steps strips conversational nuances like hesitations.

#2 about 2 min

Advantages of using native speech-to-speech streaming models

Streaming audio directly through unified models improves latency and maintains consistent voices across languages.

#3 about 2 min

Loss of developer control in black box voice models

Unified voice models abstract away intermediary steps, making it difficult to inject custom logic or webhooks.

#4 about 2 min

Managing token costs and multimodal context in voice agents

Streaming cheaper text tokens alongside audio inputs allows developers to provide rich context efficiently.

#5 about 2 min

Detecting emotional intent directly from voice audio streams

Analyzing raw voice inputs enables models to identify subtle nuances like frustration that text transcriptions lose.

#6 about 2 min

Detecting answering machines and preventing automated voice fraud

Differentiating between human speakers and automated systems remains a complex challenge for continuous outbound calling.

#7 about 2 min

Integrating speech models using the Twilio real-time SDK

New tooling simplifies telephony integration by automatically handling codec translation and sample rate conversions.

Matching moments

7:39 min

Streaming text responses from an AI language model

Chris Heilmann +2 · LIVE

2:11 min

Analyzing voice interface research projects and technical limitations

Tobias Münch Tobias Münch · WWC 2024

58 sec

Enabling live translations with AI voice-to-text pipelines

Alex Goldman · Coffee With Developers

1:44 min

Summarizing developer experience and artificial intelligence companions

Robert Hoffmann Robert Hoffmann +1 · LIVE

2:05 min

Challenges of managing voice architecture and latency

Chris Heilmann +3 · LIVE

1:12 min

Enhancing conversational intent through modern large language models

Nathaniel Okenwa Nathaniel Okenwa · WWC 2025

Upcoming sessions on this topic

Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

AI Agents are Only as Smart as their Context: Building a Real-Time Context Engine at Intuit

Bharat Patel

Lead Software Engineer at Intuit

Bharat Patel
Open session

World Congress 2026 North America

Building Stuff with GenAI - The Open Minded Workshop beyond OpenAI

Andreas Erben

CTO for Applied AI and Metaverse at daenet

Andreas Erben
Open session

World Congress 2026 North America

From Software Agents to Physical Devices: Inside the Agentic Hardware Stack

Michael Yuan, Vivian Hu

Michael Yuan
Vivian Hu
Open session

World Congress 2026 North America

The Things Your AI Isn't Telling You

Desmond Lamptey

Lead Software Engineer @ Capital One

Desmond Lamptey
Open session

World Congress 2026 North America

When Agents Became Users: Rearchitecting Identity and Permissions for AI at Scale

Yoav Gal, Dor Cohen

Yoav Gal
Dor Cohen