World Congress 2024 Aug 20, 2024 Session details

Chatbots are going to destroy infrastructures and your cloud bills

Stan Girard

Is your new RAG chatbot secretly bankrupting your cloud budget? Monolithic architectures bottleneck servers and skyrocket costs. Learn how decoupling heavy GPU tasks saves your infrastructure.

Pause
Mute Enter Fullscreen
#1 about 1 min

Introduction to chatbot infrastructure and cloud challenges

How the rapid growth of an open-source chatbot project highlights the importance of avoiding structural infrastructure mistakes.

#2 about 2 min

Contrasting web developers and data scientists

Differences in resource constraints, scaling models, and deployment environments separate traditional software and data engineering roles.

#3 about 2 min

The new paradigm of AI engineers

AI engineering merges practices from both web and data science, introducing unique optimization and computing constraints.

#4 about 2 min

Understanding basic retrieval-augmented generation architectures in chatbots

A breakdown of RAG workflows entails document chunking, GPU-bound embeddings, and sequential model generation calls.

#5 about 2 min

Scaling bottlenecks in generative AI applications

The challenge of managing parallel execution, slow response times, and sequential steps complicates user-facing AI deployments.

#6 about 2 min

Building a lean initial chatbot prototype application

Using a minimal tech stack creates an efficient and lightweight artificial intelligence container application.

#7 about 2 min

Bloat and complexity from adding GPU-bound tasks

Introducing OCR and advanced embeddings drastically increases container image sizes and subsequent infrastructure scaling costs.

#8 about 2 min

Refactoring application architectures for complex technical tasks

Separating fast user-facing business logic from slow background workers handles complex orchestrations and integrations effectively.

#9 about 3 min

Infrastructure challenges of on-premise chatbot deployments

Hosting monolithic, multi-model AI applications on local organizational servers poses massive server tuning and scaling challenges.

#10 about 1 min

Architecting generative AI systems by targeted user load

Guidelines for splitting application components apply directly to managing continuous or spiky application traffic patterns.

#11 about 3 min

Best practices for utilizing large language models in production

Strategies favoring managed APIs and service-oriented architectures reduce the immense overhead of self-hosting models.

#12 about 2 min

Overlooked AI infrastructure and operational deployment barriers

Managing complex deployments like configuring GPUs on Kubernetes requires dedicated pipelines for backups and non-deterministic testing.

#13 about 3 min

Evaluating large language models using real-time video games

A hackathon approach to measuring model capability proves that smaller setups often outperform larger architectures in speed-sensitive tasks.

Matching moments

1:22 min

Leveraging the comprehensive generative artificial intelligence stack

Duan Lightfoot Duan Lightfoot · World Congress 2024

1:32 min

Architectural patterns for developing robust generative AI applications

Julián Duque Julián Duque · World Congress 2025

2:25 min

Progressing gracefully from generic chatbots to agentic workflows

Anshul Jindal Anshul Jindal +1 · World Congress 2026 Europe

3:25 min

Constructing scalable AI solutions using LangChain and LangGraph

Julián Duque Julián Duque · World Congress 2025

4:20 min

Evolution and challenges of building AI applications

Roberto Carratalá Roberto Carratalá · World Congress 2025

4:13 min

Architectural components of a generative AI agent

Dieter Flick · World Congress 2023

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 23, 2026 · 10:00–17:00

Stage 11

Building Stuff with GenAI - The Open Minded Workshop beyond OpenAI

Andreas Erben

CTO for Applied AI and Metaverse at daenet

Andreas Erben
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 1

Anatomy of an AI Request: Where Latency and Cost Are Really Born

Dan Fu

VP of Kernels at Together AI

Dan Fu
Open session

World Congress 2026 North America

September 24, 2026 · 16:10–16:40

Stage 1

Sandboxing the Swarm: Building Secure, Serverless AI Agents with Wasm

Lena Hall, Thorsten Hans

Lena Hall
Thorsten Hans
Open session

World Congress 2026 North America

September 23, 2026 · 10:45–12:45

Stage 10

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 5

Edge AI: Running Agentic Intelligence Where Internet Can't Reach

Nitin Eusebius

AWS - Principal Solutions Architect

Nitin Eusebius
Open session

World Congress 2026 North America

September 25, 2026 · 11:00–11:30

Stage 5

Managing GPUs by Just Asking, Infrastructure in the Age of MCP

Jessica Garson Beauchemin

Developer Relations Lead, Community at Runpod

Jessica Garson Beauchemin