Aug 22, 2024

From foundation model to hosted AI solution in minutes

Kevin Klues

Ditch complex infrastructure and execute full RAG workflows in a single API call. Integrate secure, powerful open-source foundation models in minutes using your existing AI tooling.

Pause
Mute Enter Fullscreen
#1 about 1 min

Introduction to the hosted AI model hub

A unified machine learning inference service offers foundation models via a simple programming interface.

#2 about 3 min

Supported open-source foundation models and framework compatibility

Developers access open-source models like Meta Llama 3 and Mistral through interfaces compatible with OpenAI specifications.

#3 about 2 min

Language-agnostic authentication and simple API integration

Developers pass credentials via headers and supply prompts in JSON payloads without needing language-specific frameworks.

#4 about 2 min

Cost considerations of fine-tuning large language models

Fine-tuning requires massive datasets and costly specialized hardware which proves impractical for most organizational use cases.

#5 about 3 min

Retrieving semantic context using vector databases

Vector databases map unstructured text into multi-dimensional spaces to dynamically find highly relevant context for targeted queries.

#6 about 3 min

Implementing contextual answers with a single API request

Integrating custom collection references and query parameters constructs dynamic prompts without manually parsing document content.

#7 about 2 min

Ensuring compliance with European data sovereignty requirements

Managed container orchestration hosted exclusively in localized data centers guarantees strict data privacy for generative workflows.

#8 about 2 min

Unlocking direct GPU access within managed Kubernetes platforms

Deploying specialized operators bridges containerized workloads directly with underlying hardware acceleration for localized inference jobs.

#9 about 2 min

Core components driving the Nvidia GPU operator

Installing the deployment tool automatically provisions essential drivers, container toolkits, and device plugins across eligible cluster nodes.

#10 about 2 min

Deploying inference microservices via Kubernetes pod specifications

Defining accelerator resource requirements directly within standard declarative files launches production-ready inference endpoints deterministically.

#11 about 2 min

Automating model lifecycle management with the NIM operator

A dedicated automation pipeline handles the downloading and integration of packaged inference infrastructure across managed clusters.

Matching moments

4:36 min

Accelerating AI development with software and pretrained models

Ekaterina Sirazitdinova · World Congress 2023

53 sec

Creating an open ecosystem for artificial intelligence models

Chris Heilmann Chris Heilmann +2 · LIVE

2:31 min

Integrating generative AI into cloud-native applications

Cedric Clyburn Cedric Clyburn · World Congress 2024

58 sec

Replacing commercial AI APIs with self-hosted open source models

Cedric Clyburn Cedric Clyburn +1 · World Congress 2025

1:01 min

Understanding foundation models and generative AI capabilities

Timo Salm Timo Salm · World Congress 2025

54 sec

Running generative AI models in local environments

Cedric Clyburn Cedric Clyburn · World Congress 2024