> Markdown version of [/videos/1283-from-ml-to-llm-on-device-ai-in-the-browser](https://www.wearedevelopers.com/videos/1283-from-ml-to-llm-on-device-ai-in-the-browser). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # From ML to LLM: On-device AI in the Browser Run large language models entirely in the browser without making a single cloud request. Discover how WebGPU, 4-bit quantization, and local RAG unlock fast, privacy-preserving AI applications. - **Speakers:** Nico Martin - **Event:** WeAreDevelopers LIVE - **Published:** December 16, 2024 - **Duration:** 27:23 - **URL:** https://www.wearedevelopers.com/videos/1283-from-ml-to-llm-on-device-ai-in-the-browser ## Summary The evolution of on-device AI is moving rapidly from basic machine learning tasks to running large language models (LLMs) entirely within the browser. Early implementations relied on TensorFlow.js and CPU calculations, but modern browser capabilities like WebGPU drastically accelerate tasks such as real-time face landmark detection and custom gesture controls, pushing frame rates from sluggish to instantaneous. Bringing LLMs to the client side requires overcoming significant memory and compute constraints. Through 4-bit quantization and utilizing smaller, optimized models like Google's Gemma 2B, developers can compress massive weight files down to manageable sizes (around 1.4GB). Powered by WebAssembly, WebGPU, and libraries like WebLLM, these local models achieve impressive speeds—processing over 100 tokens per second on modern hardware—without ever making a cloud request. To mitigate LLM hallucinations, developers can implement Retrieval-Augmented Generation (RAG) directly in the browser. By generating local vector embeddings and matching query contexts against provided documents, applications can ground AI responses in verifiable facts and force models to admit when they lack information, all while keeping sensitive data strictly on the device. Ultimately, browser-based AI enables a privacy-preserving and indefinitely scalable architecture where users supply their own compute power. While current performance is heavily bottlenecked by hardware capabilities, emerging standards like the WebNN API will soon unlock direct browser access to native NPUs and TPUs. Until then, on-device AI serves as a powerful progressive enhancement, delivering zero-latency interactions without requiring complex installations or costly backend infrastructure. **Keywords:** on-device browser AI, webgpu machine learning, tensorflow.js optimization, LLM 4-bit quantization, gemma 2b local deployment, webllm integration, client-side RAG, local vector embeddings, teachable machine training, webnn API hardware access, AI progressive enhancement, privacy-preserving architecture, client-side compute scaling, LLM hallucination mitigation, zero-latency ML interactions ## Chapters 1. **Overcoming limitations of browser speech detection for local dialects** (00:02) — Attempting to transcribe unique regional languages highlights the limitations of standard built-in browser speech APIs. 1. **Accelerating machine learning in the browser using WebGPU** (02:32) — Processing heavy mathematical operations requires efficient backends like WebGPU over standard CPU or WebAssembly. 1. **Comparing hardware acceleration for face landmark detection** (03:54) — Drawing face meshes in real-time demonstrates the framerate differences between CPU, WebGL, and WebGPU execution. 1. **Building interactive browser extensions using hand gesture recognition** (05:54) — Tracking index and thumb finger positions allows for pinch gestures that map to scrolling and clicking on websites. 1. **Detecting custom speech commands locally with trained models** (07:22) — Recording and training audio samples entirely on the client side enables the detection of unsupported regional dialects. 1. **Mechanics of running large language models in browsers** (09:47) — Converting text input to numerical tokens and managing billions of model weights is necessary for local language model execution. 1. **Shrinking language models for browsers using four-bit quantization** (12:43) — Storing weights as four-bit floats instead of higher precision formats reduces model payload sizes to acceptable browser limits. 1. **Compiling language models to WebAssembly using web libraries** (13:20) — Utilizing open-source tools like WebLLM and Apache TVM converts standard architectures into browser-compatible WebAssembly engines. 1. **Integrating local language models into progressive web apps** (15:17) — Loading smaller parameter models in web applications processes input tokens without relying on external cloud providers. 1. **Fact-checking language models using retrieval-augmented generation** (17:34) — Parsing documents and creating vector embeddings directly in the browser ensures answers are based on verifiable local contexts. 1. **Adopting local AI as a privacy-preserving progressive enhancement** (23:24) — Optimizing for diverse user hardware requires leveraging upcoming web neural network APIs while keeping AI features optional. ## Related Moments - [Accelerating in-browser AI with WebGPU and WebNN](https://www.wearedevelopers.com/videos/953-generative-ai-power-on-the-web-making-web-apps-smarter-with-webgpu-and-webnn) (from "Generative AI power on the web: making web apps smarter with WebGPU and WebNN") - [The case for native AI in web browsers](https://www.wearedevelopers.com/videos/1572-privacy-first-in-browser-generative-ai-web-apps-offline-ready-future-proof-standards-based) (from "Privacy-first in-browser Generative AI web apps: offline-ready, future-proof, standards-based") - [Running local large language models securely using paired web GPUs](https://www.wearedevelopers.com/videos/1149-webassembly-revolution-elevating-javascript-s-reach-and-performance) (from "WebAssembly Revolution: Elevating JavaScript's Reach and Performance") - [Accelerating local machine learning models via WebNN](https://www.wearedevelopers.com/videos/100014-what-s-new-in-web-2026-edition) (from "What’s New in Web? 2026 Edition") - [Leveraging Chrome AI and nano models for web applications](https://www.wearedevelopers.com/videos/1770-generate-ai-in-the-browser-with-chrome-ai-raymond-camden) (from "Generate AI in the Browser with Chrome AI - Raymond Camden") - [Real-world implementations of browser-based machine learning](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) (from "Machine learning in the browser with TensorFlowjs") ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Dev Digest 138 - Are you secure about this?](https://www.wearedevelopers.com/magazine/486-dev-digest-138-are-you-secure-about-this) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [AI overspill Dec 2026: AI in a JAM, Blocking AI browsers, learning programming languages ](https://www.wearedevelopers.com/magazine/673-ai-overspill-dec-2026-ai-in-a-jam-blocking-ai-browsers-learning-programming-languages) ## Related Jobs - [LLM Training Engineer](https://www.wearedevelopers.com/jobs/48420-llm-training-engineer) at **Sciforium** - [AI Software Engineer (Germany)](https://www.wearedevelopers.com/jobs/48317-ai-software-engineer-germany) at **Sunhat** - [Principal Software Engineer, AI Inference Runtime](https://www.wearedevelopers.com/jobs/ext/2854958-principal-software-engineer-ai-inference-runtime) at **ARM** - [ML Engineer](https://www.wearedevelopers.com/jobs/48448-ml-engineer) at **Docker, Inc.** - [Principal Software Engineer, AI Inference Cloud](https://www.wearedevelopers.com/jobs/ext/2854957-principal-software-engineer-ai-inference-cloud) at **ARM** - [Senior AI/ML Engineer](https://www.wearedevelopers.com/jobs/48352-senior-ai-ml-engineer) at **PagerDuty**