World Congress 2025 Aug 20, 2025 Session details

From Model to Metal: An Open Source Stack for Accelerating Intelligence

Andrew Wafaa

Are hardware silos bottlenecking your AI workflows? Discover how the open-source UXL stack lets you write once and deploy seamlessly across any CPU, GPU, or NPU architecture.

Pause
Mute Enter Fullscreen
#1 about 2 min

Arm architecture and growing compute requirements for machine learning

The widespread deployment of arm architecture highlights the huge compute and energy demands of modern artificial intelligence.

#2 about 2 min

Addressing software fragmentation in long-tail machine learning workloads

Fragmented software stacks and vendor-specific libraries cause duplicated effort and hardcoded shortcuts.

#3 about 2 min

The need for unified open tooling across hardware vendors

Standardizing application programming interfaces removes proprietary silos and bridges architectural differences across diverse processors.

#4 about 3 min

Introduction to the unified acceleration library and oneAPI specification

The UXL Foundation provides a vendor-neutral specification to enable code portability and near-native performance through dynamic dispatch.

#5 about 2 min

Optimizing neural networks and integrating frameworks with oneDNN

Using oneDNN maximizes operation throughput and automatically maps cross-framework requests to optimized routines.

#6 about 1 min

Consolidating mathematical function implementations using the oneMath library

A single flexible API handles complex mathematical and statistical functions without requiring multiple backend implementations.

#7 about 1 min

Handling task parallelism and scheduling consistently with oneTBB

The oneTBB library minimizes thread scheduling overhead and allows code to execute natively across differing hardware configurations.

#8 about 2 min

Scaling multi-node communications and building machine learning pipelines

Leveraging oneCCL ensures efficient cross-node scaling while oneDAL handles pre-processing and analytics pipelines.

#9 about 2 min

Integrating the UXL stack from hardware execution to frameworks

A unified layer architecture automatically maps heavy mathematical queries and data preparation requirements directly to inference workloads.

#10 about 3 min

Democratizing intelligent computing via an open hardware-agnostic stack

Creating a write-once, run-anywhere ecosystem ensures that software optimizations stay in lockstep with upcoming processor advancements.

#11 about 2 min

Contributing code and micro kernels to the UXL Foundation

Developers can help minimize duplicate engineering efforts by sharing micro kernels and workload performance data.

Matching moments

1:57 min

Building an open collaborative software stack for AI workloads

Stephan Gillich Stephan Gillich · WWC 2024

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

2:15 min

Open-source community and machine learning frameworks

Gian Marco Iodice Gian Marco Iodice · WWC 2025

3:16 min

Building collaborative hardware architectures and developer startup ecosystems

Stephan Gillich Stephan Gillich +3 · WWC 2024

1:00 min

Unifying toolchains across AI development pipelines

Daniel Graff +1 · WWC 2021

4:36 min

Accelerating AI development with software and pretrained models

Ekaterina Sirazitdinova · WWC 2023

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Making Science Larger, not just Faster

Yuval Dvir

Commercial Executive, SandboxAQ

Yuval Dvir
Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

Compute for your AI model: GPUs, LPUs, TPUs and beyond..

Kushaagra Goyal

Tech Lead at Rubrik, ex-CTO at Gan.AI, ex-Databricks

Kushaagra Goyal
Open session

World Congress 2026 North America

Trust, But Verify: Continuous GPU Validation at Scale

Kyle Bell

VP of AI @ TensorWave

Kyle Bell
Open session

World Congress 2026 North America

It’s Alive! Taming the MLOps Franken-Stack: Write, Run, and Serve with Michelangelo

Eric Wang, Paul Zimmerman

Eric Wang
Paul Zimmerman
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan