> Markdown version of [/videos/953-generative-ai-power-on-the-web-making-web-apps-smarter-with-webgpu-and-webnn](https://www.wearedevelopers.com/videos/953-generative-ai-power-on-the-web-making-web-apps-smarter-with-webgpu-and-webnn). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Generative AI power on the web: making web apps smarter with WebGPU and WebNN Why rely on costly cloud providers when you can run generative AI locally? Discover how WebGPU and WebNN unlock near-native, client-side machine learning performance directly in the browser. - **Speakers:** [Christian Liebel](https://www.wearedevelopers.com/@christian-liebel) - **Event:** World Congress 2024 - **Published:** August 20, 2024 - **Duration:** 32:27 - **URL:** https://www.wearedevelopers.com/videos/953-generative-ai-power-on-the-web-making-web-apps-smarter-with-webgpu-and-webnn ## Summary Integrating generative AI into web applications traditionally relies on cloud providers, introducing network latency, subscription costs, and severe data privacy risks when handling sensitive user information. To bypass these limitations, developers can now leverage emerging web standards to run models—like Llama 3 or Stable Diffusion—entirely client-side. The primary hurdle for local execution has historically been model size, as multi-gigabyte files exhaust browser storage and violate the same-origin policy by forcing redundant downloads across different domains. However, the browser ecosystem is rapidly evolving to achieve near-native machine learning performance. Tools like WebLLM utilize WebAssembly and the newly shipped WebGPU API to tap into local graphics hardware, significantly accelerating on-device computations. Pushing performance even further, the experimental WebNN API allows web apps to communicate directly with native Neural Processing Units (NPUs). By offloading inference to dedicated hardware via OS-level frameworks like DirectML or CoreML, browsers can achieve over 80% of native execution speeds without data ever leaving the device. Looking forward, the web AI paradigm is shifting toward natively integrated, device-level models. Chrome's exploratory Prompt API aims to expose an on-device version of Gemini Nano to developers, eliminating redundant downloads and allowing cross-origin model sharing. While local inference currently yields lower throughput (roughly 20 tokens per second) compared to massive cloud deployments, leveraging small, specialized models—like those found in Transformers.js—offers a highly practical alternative. Ultimately, local generative AI is becoming a powerful architectural choice for developers prioritizing strict data privacy, offline availability, and specialized client-side tasks. **Keywords:** webgpu, webnn, browser-based llms, local ai inference, webllm, on-device machine learning, neural processing units, chrome prompt api, same-origin policy caching, offline data extraction, transformers.js, client-side ai privacy, webassembly acceleration, small language models, gemini nano integration ## Chapters 1. **Generative AI use cases for web applications** (00:03) — Generative AI enables image manipulation and unstructured data extraction directly within the browser. 1. **Challenges of cloud-based AI model providers** (04:40) — Cloud AI services present drawbacks including network dependency, latency, privacy risks, and subscription costs. 1. **Running large language models locally using WebLLM** (06:57) — Developers can execute downloaded large language models locally in the browser to process data offline. 1. **Accelerating in-browser AI with WebGPU and WebNN** (13:51) — Technologies like WebAssembly, WebGPU, and WebNN leverage system hardware to dramatically improve machine learning performance. 1. **Sharing models across origins with the Prompt API** (21:08) — Chrome's Prompt API mitigates large model downloads by sharing an integrated local model across web origins. 1. **Generating images locally with Stable Diffusion and WebSD** (28:45) — Web-based image generation is possible on commodity hardware using open-source models like Stable Diffusion. 1. **Evaluating trade-offs for local versus cloud AI models** (30:25) — While local models ensure privacy and offline availability, cloud models remain superior for high-performance tasks. ## Related Moments - [Leveraging Chrome AI and nano models for web applications](https://www.wearedevelopers.com/videos/1770-generate-ai-in-the-browser-with-chrome-ai-raymond-camden) (from "Generate AI in the Browser with Chrome AI - Raymond Camden") - [Accelerating local machine learning models via WebNN](https://www.wearedevelopers.com/videos/100014-what-s-new-in-web-2026-edition) (from "What’s New in Web? 2026 Edition") - [Running local large language models securely using paired web GPUs](https://www.wearedevelopers.com/videos/1149-webassembly-revolution-elevating-javascript-s-reach-and-performance) (from "WebAssembly Revolution: Elevating JavaScript's Reach and Performance") - [The case for native AI in web browsers](https://www.wearedevelopers.com/videos/1572-privacy-first-in-browser-generative-ai-web-apps-offline-ready-future-proof-standards-based) (from "Privacy-first in-browser Generative AI web apps: offline-ready, future-proof, standards-based") - [Exploring agentic browsers and artificial intelligence generation](https://www.wearedevelopers.com/videos/1723-wearedevelopers-live-graalvm-in-action-static-analysis-insights-and-more) (from "WeAreDevelopers LIVE - GraalVM in action, Static Analysis insights and more") - [Enhancing native browser experiences with on-device generative AI](https://www.wearedevelopers.com/videos/1743-wearedevelopers-live-ai-vs-the-web-ai-in-browsers) (from "WeAreDevelopers LIVE – AI vs the Web & AI in Browsers") ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Dev Digest 138 - Are you secure about this?](https://www.wearedevelopers.com/magazine/486-dev-digest-138-are-you-secure-about-this) - [AI overspill Dec 2026: AI in a JAM, Blocking AI browsers, learning programming languages ](https://www.wearedevelopers.com/magazine/673-ai-overspill-dec-2026-ai-in-a-jam-blocking-ai-browsers-learning-programming-languages) ## Related Jobs - [AI Software Engineer (Germany)](https://www.wearedevelopers.com/jobs/48317-ai-software-engineer-germany) at **Sunhat** - [LLM Training Engineer](https://www.wearedevelopers.com/jobs/48420-llm-training-engineer) at **Sciforium** - [Principal Software Engineer, AI Inference Cloud](https://www.wearedevelopers.com/jobs/ext/2854957-principal-software-engineer-ai-inference-cloud) at **ARM** - [Senior AI Developer](https://www.wearedevelopers.com/jobs/ext/2836034-senior-ai-developer) at **PwC** - [Principal Software Engineer, AI Inference Runtime](https://www.wearedevelopers.com/jobs/ext/2854958-principal-software-engineer-ai-inference-runtime) at **ARM** - [Principal Software Engineer, AI Compute Infrastructure](https://www.wearedevelopers.com/jobs/ext/2847709-principal-software-engineer-ai-compute-infrastructure) at **ARM**