Aug 22, 2024

From foundation model to hosted AI solution in minutes

Kevin Klues

Ditch complex infrastructure and execute full RAG workflows in a single API call. Integrate secure, powerful open-source foundation models in minutes using your existing AI tooling.

Pause
Mute Enter Fullscreen
#1 about 1 min

Introduction to the hosted AI model hub

A unified machine learning inference service offers foundation models via a simple programming interface.

#2 about 3 min

Supported open-source foundation models and framework compatibility

Developers access open-source models like Meta Llama 3 and Mistral through interfaces compatible with OpenAI specifications.

#3 about 2 min

Language-agnostic authentication and simple API integration

Developers pass credentials via headers and supply prompts in JSON payloads without needing language-specific frameworks.

#4 about 2 min

Cost considerations of fine-tuning large language models

Fine-tuning requires massive datasets and costly specialized hardware which proves impractical for most organizational use cases.

#5 about 3 min

Retrieving semantic context using vector databases

Vector databases map unstructured text into multi-dimensional spaces to dynamically find highly relevant context for targeted queries.

#6 about 3 min

Implementing contextual answers with a single API request

Integrating custom collection references and query parameters constructs dynamic prompts without manually parsing document content.

#7 about 2 min

Ensuring compliance with European data sovereignty requirements

Managed container orchestration hosted exclusively in localized data centers guarantees strict data privacy for generative workflows.

#8 about 2 min

Unlocking direct GPU access within managed Kubernetes platforms

Deploying specialized operators bridges containerized workloads directly with underlying hardware acceleration for localized inference jobs.

#9 about 2 min

Core components driving the Nvidia GPU operator

Installing the deployment tool automatically provisions essential drivers, container toolkits, and device plugins across eligible cluster nodes.

#10 about 2 min

Deploying inference microservices via Kubernetes pod specifications

Defining accelerator resource requirements directly within standard declarative files launches production-ready inference endpoints deterministically.

#11 about 2 min

Automating model lifecycle management with the NIM operator

A dedicated automation pipeline handles the downloading and integration of packaged inference infrastructure across managed clusters.

Matching moments

4:36 min

Accelerating AI development with software and pretrained models

Ekaterina Sirazitdinova · World Congress 2023

53 sec

Creating an open ecosystem for artificial intelligence models

Chris Heilmann +2 · LIVE

2:31 min

Integrating generative AI into cloud-native applications

Cedric Clyburn Cedric Clyburn · World Congress 2024

58 sec

Replacing commercial AI APIs with self-hosted open source models

Cedric Clyburn Cedric Clyburn +1 · World Congress 2025

1:01 min

Understanding foundation models and generative AI capabilities

Timo Salm Timo Salm · World Congress 2025

54 sec

Running generative AI models in local environments

Cedric Clyburn Cedric Clyburn · World Congress 2024

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 25, 2026 · 15:30–16:00

Stage 7

Trust, But Verify: Continuous GPU Validation at Scale

Kyle Bell

VP of AI at TensorWave

Kyle Bell
Open session

World Congress 2026 North America

September 23, 2026 · 10:00–17:00

Stage 11

Building Stuff with GenAI - The Open Minded Workshop beyond OpenAI

Andreas Erben

CTO for Applied AI and Metaverse at daenet

Andreas Erben
Open session

World Congress 2026 North America

September 24, 2026 · 17:30–18:00

Stage 6

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 5

Edge AI: Running Agentic Intelligence Where Internet Can't Reach

Nitin Eusebius

AWS - Principal Solutions Architect

Nitin Eusebius
Open session

World Congress 2026 North America

September 23, 2026 · 10:45–12:45

Stage 10

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Khaja Omer, Sheilah Kirui

Khaja Omer
Sheilah Kirui
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 1

Anatomy of an AI Request: Where Latency and Cost Are Really Born

Dan Fu

VP of Kernels at Together AI

Dan Fu