> Markdown version of [/videos/1420-mobile-ai-just-got-faster-what-s-coming-for-developers-on-arm](https://www.wearedevelopers.com/videos/1420-mobile-ai-just-got-faster-what-s-coming-for-developers-on-arm). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Mobile AI Just Got Faster: What’s Coming for Developers on Arm Run heavy generative AI entirely on-device without rewriting your code. Arm's KleidiAI integrates natively into standard frameworks, unlocking a 6x speedup for seamless, offline mobile execution. - **Speakers:** [Gian Marco Iodice](https://www.wearedevelopers.com/@gian-marco-iodice) - **Event:** World Congress 2025 - **Published:** August 20, 2025 - **Duration:** 23:09 - **URL:** https://www.wearedevelopers.com/videos/1420-mobile-ai-just-got-faster-what-s-coming-for-developers-on-arm ## Summary The rapid evolution of mobile generative AI has aggressively expanded beyond simple text prediction, bringing rich, compute-heavy modalities—such as real-time audio generation and fully offline smart speaker logic—directly to the edge. Running complex inference models locally on the CPU bypasses cloud latency bottlenecks, ensures user data privacy, and removes the continuous overhead of API requests for iterative workflows. However, balancing mixed computational pipelines that seamlessly weave neural networks, digital signal processing, and multiple data types (such as FP32, FP16, and INT8) across fragmented AI frameworks remains a major hurdle for developers. Arm addresses this ecosystem friction by leveraging a unified "optimize once, deploy everywhere" approach, pairing hardware advancements with low-level software accessibility. To standardize high performance across mobile environments, the Arm KleidiAI library acts as a critical equalizer for cross-framework deployment. Built as a lightweight, C-based suite of micro-kernels with zero heavy dependencies, KleidiAI integrates natively into ubiquitous execution environments like ONNX Runtime, ExecuTorch, and LiteRT. At the silicon level, performance is further amplified by the Armv9 architecture’s Scalable Matrix Extension 2 (SME2). By decomposing heavy matrix multiplications into multi-precision outer product accumulate (FMOPA) operations, SME2 radically accelerates key generative workloads. This synergistic combination delivers up to a 6x speedup on models like Gemma 3 and Whisper, allowing complex tasks like synthesizing 11 seconds of audio to resolve in roughly two seconds entirely on-device. For Android engineers, this convergence of hardware instructions and open-source software libraries offers profound future-proofing. By simply relying on standard AI frameworks that already bundle KleidiAI support, development teams immediately inherit these advanced CPU execution optimizations without needing to rewrite runtime application code. This frictionless upgrade path empowers teams to build highly responsive, privacy-first mobile applications while the underlying architecture automatically resolves the computational intensity of next-generation local AI. **Keywords:** mobile generative ai, on-device audio generation, local machine learning execution, armv9 cpu architecture, scalable matrix extension 2, fmopa hardware instructions, arm kleidiai library, onnx runtime mobile integration, executorch ai frameworks, multi-precision matrix multiplication, offline smart speaker inference, hardware-accelerated matrix math, cross-framework ai deployment, edge computing latency reduction, mobile cpu pipeline optimization ## Chapters 1. **Generative AI use cases on mobile devices** (00:07) — How on-device generative AI works without internet access for tasks like group chat summarization. 1. **Running audio generation locally on smartphones** (03:01) — Overcoming cloud latency by generating high-quality stereophonic audio directly on a mobile CPU using AudioGen. 1. **Scalability, security, and performance of Arm processors** (04:05) — The primary benefits of deploying AI workloads on mobile CPUs including optimization scaling and security. 1. **Open-source community and machine learning frameworks** (07:13) — Leveraging open-source frameworks like ExecuTorch and ONNX Runtime for diverse AI model deployments. 1. **Optimizing AI routines with the KleidiAI library** (09:28) — Integrating a lightweight C-based micro-kernel library into popular frameworks to accelerate neural networks. 1. **AudioGen pipelines and mixed memory data types** (11:59) — How on-device generative AI reduces cloud latency for iterative audio production using flexible floating-point operations. 1. **Building private smart assistants without cloud dependencies** (15:07) — Running speech-to-text, large language models, and text-to-speech stages securely on-device without internet connectivity. 1. **Matrix multiplication with SME2 architecture instructions** (17:13) — Using the Scalable Vector Extension 2 and Matrix Outer Product Accumulate to speed up heavy computations. 1. **Performance benchmarks and Android developer adoption** (20:21) — Unlocking significant performance speedups on key AI models and preparing Android applications for automatic hardware acceleration. ## Related Moments - [Debunking common biases about limited mobile AI capabilities](https://www.wearedevelopers.com/videos/100264-future-of-mobile-ai-what-on-device-intelligence-means-for-app-developers) (from "Future of Mobile AI. What On-Device Intelligence Means for App Developers") - [Accelerating machine learning workloads using KleidiAI libraries](https://www.wearedevelopers.com/videos/940-unleashing-the-full-potential-of-the-arm-architecture-write-once-deploy-anywhere) (from "Unleashing the Full Potential of the Arm Architecture – Write Once, Deploy Anywhere") - [Reducing cloud dependency with on-device edge AI models](https://www.wearedevelopers.com/videos/100225-edge-ai-on-ios-beyond-the-cloud-designing-the-next-generation-of-intelligent-on-device-apps) (from "Edge AI on iOS: Beyond the Cloud, Designing the Next Generation of Intelligent On-Device Apps") - [Shifting artificial intelligence models to local smartphone hardware](https://www.wearedevelopers.com/videos/1790-fake-or-news-coding-on-a-phone-emotional-support-toasters-chatgpt-weddings-and-more-anselm-hannemann) (from "Fake or News: Coding on a Phone, Emotional Support Toasters, ChatGPT Weddings and more - Anselm Hannemann") - [Addressing mobile energy limits and AI processing tradeoffs](https://www.wearedevelopers.com/videos/100225-edge-ai-on-ios-beyond-the-cloud-designing-the-next-generation-of-intelligent-on-device-apps) (from "Edge AI on iOS: Beyond the Cloud, Designing the Next Generation of Intelligent On-Device Apps") - [The convergence of mobile engineering and machine learning](https://www.wearedevelopers.com/videos/100264-future-of-mobile-ai-what-on-device-intelligence-means-for-app-developers) (from "Future of Mobile AI. What On-Device Intelligence Means for App Developers") ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How we Build The Software of Tomorrow](https://www.wearedevelopers.com/magazine/120-how-we-build-the-software-of-tomorrow) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) ## Related Jobs - [Founding Mobile Engineer (iOS)](https://www.wearedevelopers.com/jobs/ext/1648532-founding-mobile-engineer-ios) at **Almedia** - [AI Software Engineer (Germany)](https://www.wearedevelopers.com/jobs/48317-ai-software-engineer-germany) at **Sunhat** - [AI & Machine Learning Engineer (all genders)](https://www.wearedevelopers.com/jobs/48217-ai-machine-learning-engineer-all-genders) at **msg** - [AI Full Stack Engineer](https://www.wearedevelopers.com/jobs/ext/1354435-ai-full-stack-engineer) at **Almedia** - [Principal Field Architect - AI Agents](https://www.wearedevelopers.com/jobs/ext/1442858-principal-field-architect-ai-agents) at **Twilio** - [Data Scientist](https://www.wearedevelopers.com/jobs/ext/1351648-data-scientist) at **Almedia**