WeAreDevelopers LIVE Jun 17, 2020

Raise your voice!

Lee Boonstra

Ditch rigid wake words and consumer timeouts by building your own custom voice AI. Master the full architecture using Angular, Node.js, and Dialogflow to handle complex, enterprise-grade conversations.

Pause
Mute Enter Fullscreen
#1 about 6 min

Developing custom voice AI versus ecosystem platforms

Custom voice solutions enable specialized hardware and independent deployments without consumer platform restrictions.

#2 about 4 min

Voice AI architecture and interactive kiosk demo

The application processes frontend WebRTC streams through Google APIs to drive an interactive multilingual kiosk.

#3 about 5 min

Capturing browser microphone access with RecordRTC

The RecordRTC library normalizes cross-browser media quirks and establishes predictable stream sample rates.

#4 about 2 min

Streaming bi-directional audio via Socket.io

Socket.io-stream establishes realtime bidirectional endpoints to pass binary audio from Angular to Node.js.

#5 about 4 min

Converting continuous audio streams to text

The Google Speech-to-Text API requires strictly matched audio encoding and sample rates to process streamed chunks.

#6 about 4 min

Classifying conversational intents using Dialogflow models

Dialogflow applies natural language understanding to detect complex conversational intents, map entity parameters, and trigger automated webhook fulfillments.

#7 about 5 min

Reconciling multilingual voice inputs with Translate API

The dynamic translation client converts arbitrary foreign input back to baseline English to ensure accurate intent matching.

#8 about 2 min

Synthesizing human-like voice responses with Wavenet

The Text-to-Speech API transforms textual responses into playback-ready linear16 audio buffers.

#9 about 2 min

Playing synthesized backend audio objects in browsers

The Angular frontend parses returned audio strings into HTML5 AudioContext nodes to circumvent restrictive mobile playback rules.

#10 about 2 min

Securing application microphone access with HTTPS deployments

App Engine Flex provides the automated SSL certificates required by modern browsers to allow microphone recording endpoints.

Matching moments

42 sec

Delivering dynamic voice interactions via ElevenLabs

Alexandra Mihai Alexandra Mihai +2 · Europe 2026 Virtual

2:11 min

Analyzing voice interface research projects and technical limitations

Tobias Münch Tobias Münch · World Congress 2024

1:05 min

Building a personal assistant interface using web speech

Tobias Münch Tobias Münch · World Congress 2024

1:44 min

Summarizing developer experience and artificial intelligence companions

Robert Hoffmann Robert Hoffmann +1 · LIVE

1:10 min

Architecting a web real-time communication stack for agents

Nathaniel Okenwa Nathaniel Okenwa · World Congress 2025

7:39 min

Streaming text responses from an AI language model

Chris Heilmann +2 · LIVE

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 24, 2026 · 11:00–11:30

Stage 4

From Voice Demo to Enterprise Production

Anahita Havewala, Anuj Gupta

Anahita Havewala
Anuj Gupta
Open session

World Congress 2026 North America

September 24, 2026 · 11:20–11:25

Outdoor Stage

Finding the Edges: Testing, Evaluating, and Monitoring Voice AI Agents Before Your Users Do

Matt Wyman

CEO of Okareo

Matt Wyman
Open session

World Congress 2026 North America

September 24, 2026 · 16:10–16:40

Stage 5

From Software Agents to Physical Devices: Inside the Agentic Hardware Stack

Michael Yuan, Vivian Hu

Michael Yuan
Vivian Hu
Open session

World Congress 2026 North America

September 25, 2026 · 14:50–15:20

Stage 4

Intelligence in Motion: Building the Next Generation of AI-Powered Apps on Zoom's Developer Platform

Brendan Ittelson

Chief Ecosystem Officer of Zoom

Brendan Ittelson
Open session

World Congress 2026 North America

September 25, 2026 · 10:20–10:50

Stage 5

Physical AI: 5 Things You Can Build That Aren't Another Chatbot

Vini Senger

Senior Technical Evangelist for Startups

Vini Senger
Open session

World Congress 2026 North America

September 24, 2026 · 11:00–11:30

Stage 2

From Stateless to Stateful: Real-Time Voice & Messaging Agents with Twilio and AWS

Rishab Kumar

Staff Developer Evangelist @ Twilio

Rishab Kumar