World Congress 2024 • Aug 20, 2024 • Session details

Is the web ready for voice user interfaces?

Tobias Münch

Built-in browser speech APIs struggle with cloud privacy and complex state management. Can WebAssembly and local LLMs finally unlock offline, conversational web apps?

Pause
Mute Enter Fullscreen
#1 about 2 min

Overview of voice user interfaces on the web

Exploring existing W3C web speech standards and their potential application in modern web applications.

#2 about 2 min

Accessibility and practical use cases for voice interaction

Voice interfaces enable online participation for users with physical disabilities and streamline documentation for mobile workers.

#3 about 1 min

Core components of the web speech API ecosystem

The browser implementation divides voice functionality into distinct speech recognition and speech synthesis capabilities.

#4 about 3 min

Analyzing voice interface research projects and technical limitations

Experimental conversational web models highlight issues with recognition accuracy, latency, and required manual triggering mechanics.

#5 about 2 min

Building a personal assistant interface using web speech

Combining the standard web speech capability with large language models enables simple voice-driven chat applications.

#6 about 2 min

Configuring the browser speech recognition class and grammar

Instantiating the webkit speech recognition service requires specific parameter configurations and domain-specific grammar formatting.

#7 about 3 min

Extracting and mapping recognition results to application state

Handling multi-dimensional result arrays allows developers to extract text transcripts for dynamic interface updates.

#8 about 3 min

Privacy concerns and developer experience barriers to adoption

Low adoption rates stem from state management complexity alongside critical data privacy risks for sensitive workflows.

#9 about 3 min

Tracing the speech API architecture within the browser

The underlying Chromium implementation delegates voice data through multiple service layers before reaching an external processing engine.

#10 about 4 min

Decoupling state and voice commands with react genie

Semantic parsers map natural language inputs to backend application state to support complex offscreen operations hands-free.

#11 about 1 min

Quantifying developer experience with third-party software packages

Evaluating the hidden technical friction in external libraries helps inform better engineering tooling and methodology.

#12 about 2 min

The production readiness of current web voice interfaces

While experimental tools remain inconsistent for everyday tasks, emerging architectural concepts point toward a viable production future.

Matching moments

2:39 min

Current limitations of browser-based voice integrations

Ana Rodrigues · LIVE

58 sec

Uncensored voice assistants and neural controlled web browser accessibility

Chris Heilmann Chris Heilmann +2 · LIVE

2:37 min

Browser support and privacy concerns for speech recognition

Ana Rodrigues · LIVE

6:00 min

Expectations versus reality for voice and virtual interfaces

Léonie Watson · A11y + AI

1:57 min

Architectural stack for speech-driven UI compilation

Sami Ekblad Sami Ekblad · World Congress 2024

10:01 min

Why current voice recognition algorithms fail true accessibility standards

Chris Heilmann Chris Heilmann +2 · LIVE