Coffee With Developers • Jun 17, 2026

Speech-to-Speech AI models - Marius Obert

Marius Obert

Marius Obert argues traditional transcription pipelines destroy essential conversational context. Discover how native speech-to-speech models preserve emotional nuances and deploy multimodal voice agents in under 70 lines of code.

Pause
Mute Enter Fullscreen
#1 about 2 min

Comparing transcription-based voice pipelines to native models

Routing audio through separate transcription and generation steps strips conversational nuances like hesitations.

#2 about 2 min

Advantages of using native speech-to-speech streaming models

Streaming audio directly through unified models improves latency and maintains consistent voices across languages.

#3 about 2 min

Loss of developer control in black box voice models

Unified voice models abstract away intermediary steps, making it difficult to inject custom logic or webhooks.

#4 about 2 min

Managing token costs and multimodal context in voice agents

Streaming cheaper text tokens alongside audio inputs allows developers to provide rich context efficiently.

#5 about 2 min

Detecting emotional intent directly from voice audio streams

Analyzing raw voice inputs enables models to identify subtle nuances like frustration that text transcriptions lose.

#6 about 2 min

Detecting answering machines and preventing automated voice fraud

Differentiating between human speakers and automated systems remains a complex challenge for continuous outbound calling.

#7 about 2 min

Integrating speech models using the Twilio real-time SDK

New tooling simplifies telephony integration by automatically handling codec translation and sample rate conversions.

Matching moments

3:57 min

Advancements in synthetic speech and voice banking technology

Léonie Watson · A11y + AI

7:39 min

Streaming text responses from an AI language model

Chris Heilmann Chris Heilmann +2 · LIVE

2:11 min

Analyzing voice interface research projects and technical limitations

Tobias Münch Tobias Münch · World Congress 2024

58 sec

Enabling live translations with AI voice-to-text pipelines

Alex Goldman Alex Goldman · Coffee With Developers

49 sec

Voice cloning models and legacy programming languages

Andrew MacLean Andrew MacLean +2 · LIVE

42 sec

Delivering dynamic voice interactions via ElevenLabs

Alexandra Mihai Alexandra Mihai +2 · Europe 2026 Virtual