WeAreDevelopers LIVE • Dec 16, 2024

From ML to LLM: On-device AI in the Browser

Nico Martin

Run large language models entirely in the browser without making a single cloud request. Discover how WebGPU, 4-bit quantization, and local RAG unlock fast, privacy-preserving AI applications.

Pause
Mute Enter Fullscreen
#1 about 3 min

Overcoming limitations of browser speech detection for local dialects

Attempting to transcribe unique regional languages highlights the limitations of standard built-in browser speech APIs.

#2 about 2 min

Accelerating machine learning in the browser using WebGPU

Processing heavy mathematical operations requires efficient backends like WebGPU over standard CPU or WebAssembly.

#3 about 2 min

Comparing hardware acceleration for face landmark detection

Drawing face meshes in real-time demonstrates the framerate differences between CPU, WebGL, and WebGPU execution.

#4 about 2 min

Building interactive browser extensions using hand gesture recognition

Tracking index and thumb finger positions allows for pinch gestures that map to scrolling and clicking on websites.

#5 about 3 min

Detecting custom speech commands locally with trained models

Recording and training audio samples entirely on the client side enables the detection of unsupported regional dialects.

#6 about 3 min

Mechanics of running large language models in browsers

Converting text input to numerical tokens and managing billions of model weights is necessary for local language model execution.

#7 about 1 min

Shrinking language models for browsers using four-bit quantization

Storing weights as four-bit floats instead of higher precision formats reduces model payload sizes to acceptable browser limits.

#8 about 2 min

Compiling language models to WebAssembly using web libraries

Utilizing open-source tools like WebLLM and Apache TVM converts standard architectures into browser-compatible WebAssembly engines.

#9 about 3 min

Integrating local language models into progressive web apps

Loading smaller parameter models in web applications processes input tokens without relying on external cloud providers.

#10 about 6 min

Fact-checking language models using retrieval-augmented generation

Parsing documents and creating vector embeddings directly in the browser ensures answers are based on verifiable local contexts.

#11 about 4 min

Adopting local AI as a privacy-preserving progressive enhancement

Optimizing for diverse user hardware requires leveraging upcoming web neural network APIs while keeping AI features optional.

Matching moments

7:17 min

Accelerating in-browser AI with WebGPU and WebNN

Christian Liebel Christian Liebel · World Congress 2024

2:46 min

The case for native AI in web browsers

Maxim Salnikov Maxim Salnikov · World Congress 2025

1:04 min

Running local large language models securely using paired web GPUs

Önder Ceylan Önder Ceylan · World Congress 2024

2:17 min

Accelerating local machine learning models via WebNN

Christian Liebel Christian Liebel · World Congress 2026 Europe

2:54 min

Leveraging Chrome AI and nano models for web applications

Raymond Camden · Perf + AI

1:33 min

Real-world implementations of browser-based machine learning

Håkan Silfvernagel · LIVE