> Markdown version of [/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence?t=5](https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence?t=5). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # From Model to Metal: An Open Source Stack for Accelerating Intelligence Are hardware silos bottlenecking your AI workflows? Discover how the open-source UXL stack lets you write once and deploy seamlessly across any CPU, GPU, or NPU architecture. - **Speakers:** [Andrew Wafaa](https://www.wearedevelopers.com/@andrew-wafaa) - **Event:** World Congress 2025 - **Published:** August 20, 2025 - **Duration:** 17:18 - **URL:** https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence ## Summary The AI and machine learning ecosystem suffers from deep software fragmentation, with developers frequently tied to vendor-specific libraries and proprietary hardware silos. Because "a good developer is a lazy developer," teams often rely on the quickest, hardcoded path to deployment—sacrificing portability across CPU, GPU, and NPU architectures. Addressing this bottleneck requires bridging the gap from model to metal seamlessly across edge environments and hyper-scaler data centers without duplicating engineering effort. The Unified Acceleration Library (UXL) Foundation introduces a vendor-neutral, open-source stack that resolves these scaling challenges. Built upon the oneAPI specification, UXL offers a suite of standardized frameworks—including oneDNN for deep neural networks, oneMath for optimized statistical functions, and oneDAL for data analytics pipelines. By dynamically dispatching operations to highly optimized micro-kernels across different platforms, UXL allows developers to write once and run anywhere. This approach integrates seamlessly with established frameworks like PyTorch and TensorFlow, eliminating the need to adopt distinct, proprietary tooling for varying hardware. By democratizing compute optimization, the open-source UXL ecosystem ensures software advancements remain in lockstep with new hardware releases. Leveraging specialized backend tools like oneTBB for efficient task scheduling and oneCCL for multi-node communication guarantees high throughput natively across diverse clusters. Ultimately, this foundational layer minimizes workflow fragmentation, enabling engineers to stop reinventing hardware interconnects and instead focus on scaling intelligence for both the shiny "short head" and the expansive "long tail" of specialized AI workloads. **Keywords:** UXL foundation, oneAPI specification, AI hardware acceleration, machine learning pipelines, oneDNN neural networks, oneMath optimization, oneDAL data analytics, oneTBB task scheduling, oneCCL node communication, cross-architecture deployments, vendor-neutral frameworks, dynamic compute dispatch, open-source hardware integration, compute resource fragmentation ## Chapters 1. **Arm architecture and growing compute requirements for machine learning** (00:05) — The widespread deployment of arm architecture highlights the huge compute and energy demands of modern artificial intelligence. 1. **Addressing software fragmentation in long-tail machine learning workloads** (01:35) — Fragmented software stacks and vendor-specific libraries cause duplicated effort and hardcoded shortcuts. 1. **The need for unified open tooling across hardware vendors** (03:06) — Standardizing application programming interfaces removes proprietary silos and bridges architectural differences across diverse processors. 1. **Introduction to the unified acceleration library and oneAPI specification** (04:15) — The UXL Foundation provides a vendor-neutral specification to enable code portability and near-native performance through dynamic dispatch. 1. **Optimizing neural networks and integrating frameworks with oneDNN** (06:24) — Using oneDNN maximizes operation throughput and automatically maps cross-framework requests to optimized routines. 1. **Consolidating mathematical function implementations using the oneMath library** (08:05) — A single flexible API handles complex mathematical and statistical functions without requiring multiple backend implementations. 1. **Handling task parallelism and scheduling consistently with oneTBB** (09:03) — The oneTBB library minimizes thread scheduling overhead and allows code to execute natively across differing hardware configurations. 1. **Scaling multi-node communications and building machine learning pipelines** (10:02) — Leveraging oneCCL ensures efficient cross-node scaling while oneDAL handles pre-processing and analytics pipelines. 1. **Integrating the UXL stack from hardware execution to frameworks** (11:21) — A unified layer architecture automatically maps heavy mathematical queries and data preparation requirements directly to inference workloads. 1. **Democratizing intelligent computing via an open hardware-agnostic stack** (13:11) — Creating a write-once, run-anywhere ecosystem ensures that software optimizations stay in lockstep with upcoming processor advancements. 1. **Contributing code and micro kernels to the UXL Foundation** (16:09) — Developers can help minimize duplicate engineering efforts by sharing micro kernels and workload performance data. ## Related Moments - [Building an open collaborative software stack for AI workloads](https://www.wearedevelopers.com/videos/1132-bringing-ai-everywhere) (from "Bringing AI Everywhere") - [Architecting CUDA and the AI software stack](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) (from "Building the Nervous System of AI - Michael Kagan (NVIDIA)") - [Open-source community and machine learning frameworks](https://www.wearedevelopers.com/videos/1420-mobile-ai-just-got-faster-what-s-coming-for-developers-on-arm) (from "Mobile AI Just Got Faster: What’s Coming for Developers on Arm") - [Building collaborative hardware architectures and developer startup ecosystems](https://www.wearedevelopers.com/videos/1106-the-future-of-computing-ai-technologies-in-the-exascale-era) (from "The Future of Computing: AI Technologies in the Exascale Era") - [Unifying toolchains across AI development pipelines](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) (from "Developing an AI.SDK") - [Accelerating AI development with software and pretrained models](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) (from "Trends, Challenges and Best Practices for AI at the Edge") ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) ## Related Jobs - [AI Full Stack Engineer](https://www.wearedevelopers.com/jobs/ext/1354435-ai-full-stack-engineer) at **Almedia** - [Principal Software Engineer, Enterprise AI Platform](https://www.wearedevelopers.com/jobs/ext/1467292-principal-software-engineer-enterprise-ai-platform) at **GitHub** - [Principal Engineer - AI Search & Vector Infrastructure](https://www.wearedevelopers.com/jobs/ext/319507-principal-engineer-ai-search-vector-infrastructure) at **Redis** - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/1597388-machine-learning-engineer) at **ZEISS Group** - [Principal Engineer - AI Search & Vector Infrastructure](https://www.wearedevelopers.com/jobs/ext/353953-principal-engineer-ai-search-vector-infrastructure) at **Redis** - [Principal Engineer - AI Search & Vector Infrastructure](https://www.wearedevelopers.com/jobs/ext/381484-principal-engineer-ai-search-vector-infrastructure) at **Redis**