World Congress 2025 Aug 20, 2025 Session details

From Model to Metal: An Open Source Stack for Accelerating Intelligence

Andrew Wafaa

Are hardware silos bottlenecking your AI workflows? Discover how the open-source UXL stack lets you write once and deploy seamlessly across any CPU, GPU, or NPU architecture.

Pause
Mute Enter Fullscreen
#1 about 2 min

Arm architecture and growing compute requirements for machine learning

The widespread deployment of arm architecture highlights the huge compute and energy demands of modern artificial intelligence.

#2 about 2 min

Addressing software fragmentation in long-tail machine learning workloads

Fragmented software stacks and vendor-specific libraries cause duplicated effort and hardcoded shortcuts.

#3 about 2 min

The need for unified open tooling across hardware vendors

Standardizing application programming interfaces removes proprietary silos and bridges architectural differences across diverse processors.

#4 about 3 min

Introduction to the unified acceleration library and oneAPI specification

The UXL Foundation provides a vendor-neutral specification to enable code portability and near-native performance through dynamic dispatch.

#5 about 2 min

Optimizing neural networks and integrating frameworks with oneDNN

Using oneDNN maximizes operation throughput and automatically maps cross-framework requests to optimized routines.

#6 about 1 min

Consolidating mathematical function implementations using the oneMath library

A single flexible API handles complex mathematical and statistical functions without requiring multiple backend implementations.

#7 about 1 min

Handling task parallelism and scheduling consistently with oneTBB

The oneTBB library minimizes thread scheduling overhead and allows code to execute natively across differing hardware configurations.

#8 about 2 min

Scaling multi-node communications and building machine learning pipelines

Leveraging oneCCL ensures efficient cross-node scaling while oneDAL handles pre-processing and analytics pipelines.

#9 about 2 min

Integrating the UXL stack from hardware execution to frameworks

A unified layer architecture automatically maps heavy mathematical queries and data preparation requirements directly to inference workloads.

#10 about 3 min

Democratizing intelligent computing via an open hardware-agnostic stack

Creating a write-once, run-anywhere ecosystem ensures that software optimizations stay in lockstep with upcoming processor advancements.

#11 about 2 min

Contributing code and micro kernels to the UXL Foundation

Developers can help minimize duplicate engineering efforts by sharing micro kernels and workload performance data.

Matching moments

1:57 min

Building an open collaborative software stack for AI workloads

Stephan Gillich Stephan Gillich · World Congress 2024

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

2:15 min

Open-source community and machine learning frameworks

Gian Marco Iodice Gian Marco Iodice · World Congress 2025

3:16 min

Building collaborative hardware architectures and developer startup ecosystems

Stephan Gillich Stephan Gillich +3 · World Congress 2024

1:00 min

Unifying toolchains across AI development pipelines

Daniel Graff +1 · World Congress 2021

4:36 min

Accelerating AI development with software and pretrained models

Ekaterina Sirazitdinova · World Congress 2023

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 1

Anatomy of an AI Request: Where Latency and Cost Are Really Born

Dan Fu

VP of Kernels at Together AI

Dan Fu
Open session

World Congress 2026 North America

September 24, 2026 · 11:00–11:30

Stage 3

Making Science Larger, not just Faster

Yuval Dvir

Commercial Executive, SandboxAQ

Yuval Dvir
Open session

World Congress 2026 North America

September 24, 2026 · 17:30–18:00

Stage 6

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 6

Compute for your AI model: GPUs, LPUs, TPUs and beyond..

Kushaagra Goyal

Tech Lead at Rubrik, ex-CTO at Gan.AI, ex-Databricks

Kushaagra Goyal
Open session

World Congress 2026 North America

September 25, 2026 · 15:30–16:00

Stage 7

Trust, But Verify: Continuous GPU Validation at Scale

Kyle Bell

VP of AI at TensorWave

Kyle Bell
Open session

World Congress 2026 North America

September 25, 2026 · 13:30–14:00

Stage 7

It’s Alive! Taming the MLOps Franken-Stack: Write, Run, and Serve with Michelangelo

Eric Wang, Paul Zimmerman

Eric Wang
Paul Zimmerman