World Congress 2026 North America

PrimaLabs: The Application-Specific AI Inference Stack

September 24, 2026 12:20 – 12:25 · 5 min Outdoor Stage

World Congress 2026 North America

September 23–25, 2026 · San José, CA

Attend in person

Get tickets

Watch remotely

Watch live with Pro

Pro

Can’t make it to San José? Watch this session live with Pro. You also get:

  • All full videos, bookmarks, and playlists
  • World Congress livestreams
See pricing

What this session covers

AI applications place very different demands on inference. Coding agents reuse long context, vision-language systems process variable multimodal inputs, and generative recommendation models operate under strict latency requirements. A generic inference stack cannot maximize performance across all three. PrimaLabs delivers an application-specific inference stack engineered around each production workload. The platform jointly optimizes runtimes, kernels, caching, batching, scheduling, parallelism, and GPU infrastructure, then operates the resulting stack in an isolated environment.

Related talks at this congress

Open session

World Congress 2026 North America

September 25, 2026 · 13:30–14:00

Mainstage

A Hands-On Developer Guide to Inference Engineering

Ankit Patel, Philip Kiely

Ankit Patel
Philip Kiely
Open session

World Congress 2026 North America

September 23, 2026 · 10:45–12:45

Stage 10

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 1

Anatomy of an AI Request: Where Latency and Cost Are Really Born

Dan Fu

VP of Kernels at Together AI

Dan Fu
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 5

Edge AI: Running Agentic Intelligence Where Internet Can't Reach

Nitin Eusebius

AWS - Principal Solutions Architect

Nitin Eusebius
All sessions at this congress