Aug 22, 2024

From foundation model to hosted AI solution in minutes

Kevin Klues

Ditch complex infrastructure and execute full RAG workflows in a single API call. Integrate secure, powerful open-source foundation models in minutes using your existing AI tooling.

Pause
Mute Enter Fullscreen
#1 about 1 min

Introduction to the hosted AI model hub

A unified machine learning inference service offers foundation models via a simple programming interface.

#2 about 3 min

Supported open-source foundation models and framework compatibility

Developers access open-source models like Meta Llama 3 and Mistral through interfaces compatible with OpenAI specifications.

#3 about 2 min

Language-agnostic authentication and simple API integration

Developers pass credentials via headers and supply prompts in JSON payloads without needing language-specific frameworks.

#4 about 2 min

Cost considerations of fine-tuning large language models

Fine-tuning requires massive datasets and costly specialized hardware which proves impractical for most organizational use cases.

#5 about 3 min

Retrieving semantic context using vector databases

Vector databases map unstructured text into multi-dimensional spaces to dynamically find highly relevant context for targeted queries.

#6 about 3 min

Implementing contextual answers with a single API request

Integrating custom collection references and query parameters constructs dynamic prompts without manually parsing document content.

#7 about 2 min

Ensuring compliance with European data sovereignty requirements

Managed container orchestration hosted exclusively in localized data centers guarantees strict data privacy for generative workflows.

#8 about 2 min

Unlocking direct GPU access within managed Kubernetes platforms

Deploying specialized operators bridges containerized workloads directly with underlying hardware acceleration for localized inference jobs.

#9 about 2 min

Core components driving the Nvidia GPU operator

Installing the deployment tool automatically provisions essential drivers, container toolkits, and device plugins across eligible cluster nodes.

#10 about 2 min

Deploying inference microservices via Kubernetes pod specifications

Defining accelerator resource requirements directly within standard declarative files launches production-ready inference endpoints deterministically.

#11 about 2 min

Automating model lifecycle management with the NIM operator

A dedicated automation pipeline handles the downloading and integration of packaged inference infrastructure across managed clusters.

Matching moments

4:36 min

Accelerating AI development with software and pretrained models

Ekaterina Sirazitdinova · WWC 2023

53 sec

Creating an open ecosystem for artificial intelligence models

Chris Heilmann +2 · LIVE

2:31 min

Integrating generative AI into cloud-native applications

Cedric Clyburn Cedric Clyburn · WWC 2024

58 sec

Replacing commercial AI APIs with self-hosted open source models

Cedric Clyburn Cedric Clyburn +1 · WWC 2025

1:01 min

Understanding foundation models and generative AI capabilities

Timo Salm Timo Salm · WWC 2025

54 sec

Running generative AI models in local environments

Cedric Clyburn Cedric Clyburn · WWC 2024

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Trust, But Verify: Continuous GPU Validation at Scale

Kyle Bell

VP of AI @ TensorWave

Kyle Bell
Open session

World Congress 2026 North America

Building Stuff with GenAI - The Open Minded Workshop beyond OpenAI

Andreas Erben

CTO for Applied AI and Metaverse at daenet

Andreas Erben
Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

Autonomous Infrastructure: Building AI Agents for Global-Scale Capacity Efficiency

Tommy Tran, Gregoire Colin

Tommy Tran
Gregoire Colin
Open session

World Congress 2026 North America

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Legare Kerrison, Cedric Clyburn

Legare Kerrison
Cedric Clyburn