> Markdown version of [/videos/1170-from-foundation-model-to-hosted-ai-solution-in-minutes](https://www.wearedevelopers.com/videos/1170-from-foundation-model-to-hosted-ai-solution-in-minutes). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # From foundation model to hosted AI solution in minutes Ditch complex infrastructure and execute full RAG workflows in a single API call. Integrate secure, powerful open-source foundation models in minutes using your existing AI tooling. - **Speakers:** [Kevin Klues](https://www.wearedevelopers.com/@kevin-klues) - **Event:** - **Published:** August 22, 2024 - **Duration:** 18:33 - **URL:** https://www.wearedevelopers.com/videos/1170-from-foundation-model-to-hosted-ai-solution-in-minutes ## Summary The video introduces the Ionos AI Model Hub, a new product designed to simplify the integration of open-source foundation models into production environments. Addressing the complexity and prohibitive costs of fine-tuning large language models from scratch, this service provides direct access to top-tier open-source models—including Llama 3, Mistral, Mixtral, and Stable Diffusion XL—through a highly intuitive REST API. By fully supporting the OpenAI API specification, the hub allows development teams to use their existing AI tooling and seamlessly swap out the backend URL, instantly migrating workloads to secure, European-hosted infrastructure that guarantees data sovereignty. A standout feature of the AI Model Hub is its approach to retrieval-augmented generation (RAG). Recognizing that customizing models via fine-tuning requires specialized GPU hardware and vast datasets, Ionos has embedded a vector database natively into their ecosystem. Instead of piecing together separate document retrieval and prompting steps, developers can execute a full RAG workflow in just a single API call. By storing context in the vector database and passing a unique collection ID within the payload, the platform automatically retrieves semantically relevant company data and injects it into the prompt, delivering hyper-relevant answers securely and efficiently. To elevate this infrastructure for advanced, custom workloads, the presentation highlights an upcoming integration with NVIDIA. While the AI Model Hub operates as a hosted service, users will soon be able to provision GPU access directly within Ionos managed Kubernetes clusters using the NVIDIA GPU Operator. This automates the underlying host-level driver installations and container toolkits. Furthermore, the newly introduced NIM Operator (NVIDIA Inference Microservices) will streamline the deployment of self-hosted, scalable inference services, empowering engineering teams to build, scale, and orchestrate localized AI architectures inside secure data centers. **Keywords:** ionos AI model hub, retrieval-augmented generation, open-source foundation models, vector database implementation, llama 3 deployment, mixtral language model, openai API compatibility, nvidia GPU operator, managed kubernetes AI infrastructure, AI data sovereignty, nvidia inference microservices, containerized GPU orchestration, NIM operator, AI hardware cost optimization, scalable AI inferencing ## Chapters 1. **Introduction to the hosted AI model hub** (00:30) — A unified machine learning inference service offers foundation models via a simple programming interface. 1. **Supported open-source foundation models and framework compatibility** (01:28) — Developers access open-source models like Meta Llama 3 and Mistral through interfaces compatible with OpenAI specifications. 1. **Language-agnostic authentication and simple API integration** (04:08) — Developers pass credentials via headers and supply prompts in JSON payloads without needing language-specific frameworks. 1. **Cost considerations of fine-tuning large language models** (05:26) — Fine-tuning requires massive datasets and costly specialized hardware which proves impractical for most organizational use cases. 1. **Retrieving semantic context using vector databases** (06:47) — Vector databases map unstructured text into multi-dimensional spaces to dynamically find highly relevant context for targeted queries. 1. **Implementing contextual answers with a single API request** (09:16) — Integrating custom collection references and query parameters constructs dynamic prompts without manually parsing document content. 1. **Ensuring compliance with European data sovereignty requirements** (11:37) — Managed container orchestration hosted exclusively in localized data centers guarantees strict data privacy for generative workflows. 1. **Unlocking direct GPU access within managed Kubernetes platforms** (12:47) — Deploying specialized operators bridges containerized workloads directly with underlying hardware acceleration for localized inference jobs. 1. **Core components driving the Nvidia GPU operator** (14:24) — Installing the deployment tool automatically provisions essential drivers, container toolkits, and device plugins across eligible cluster nodes. 1. **Deploying inference microservices via Kubernetes pod specifications** (15:28) — Defining accelerator resource requirements directly within standard declarative files launches production-ready inference endpoints deterministically. 1. **Automating model lifecycle management with the NIM operator** (17:00) — A dedicated automation pipeline handles the downloading and integration of packaged inference infrastructure across managed clusters. ## Related Moments - [Accelerating AI development with software and pretrained models](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) (from "Trends, Challenges and Best Practices for AI at the Edge") - [Creating an open ecosystem for artificial intelligence models](https://www.wearedevelopers.com/videos/1761-wearedevelopers-live-frontend-inspirations-web-standards-and-more) (from "WeAreDevelopers LIVE – Frontend Inspirations, Web Standards and more") - [Integrating generative AI into cloud-native applications](https://www.wearedevelopers.com/videos/950-supercharge-your-cloud-native-applications-with-generative-ai) (from "Supercharge your cloud-native applications with Generative AI") - [Replacing commercial AI APIs with self-hosted open source models](https://www.wearedevelopers.com/videos/1597-self-hosted-llms-from-zero-to-inference) (from "Self-Hosted LLMs: From Zero to Inference") - [Understanding foundation models and generative AI capabilities](https://www.wearedevelopers.com/videos/1554-java-meets-ai-empowering-spring-developers-to-build-intelligent-apps) (from "Java Meets AI: Empowering Spring Developers to Build Intelligent Apps") - [Running generative AI models in local environments](https://www.wearedevelopers.com/videos/950-supercharge-your-cloud-native-applications-with-generative-ai) (from "Supercharge your cloud-native applications with Generative AI") ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) ## Related Jobs - [AI Software Engineer (Germany)](https://www.wearedevelopers.com/jobs/48317-ai-software-engineer-germany) at **Sunhat** - [AI Operations Manager (all genders)](https://www.wearedevelopers.com/jobs/48263-ai-operations-manager-all-genders) at **envelio** - [AI & Machine Learning Engineer (all genders)](https://www.wearedevelopers.com/jobs/48217-ai-machine-learning-engineer-all-genders) at **msg** - [AI Full Stack Engineer](https://www.wearedevelopers.com/jobs/ext/1354435-ai-full-stack-engineer) at **Almedia** - [Principal Engineer - AI Search & Vector Infrastructure](https://www.wearedevelopers.com/jobs/ext/319507-principal-engineer-ai-search-vector-infrastructure) at **Redis** - [Principal Engineer - AI Search & Vector Infrastructure](https://www.wearedevelopers.com/jobs/ext/353953-principal-engineer-ai-search-vector-infrastructure) at **Redis**