World Congress 2024 Aug 20, 2024 Session details

Chatbots are going to destroy infrastructures and your cloud bills

Stan Girard

Is your new RAG chatbot secretly bankrupting your cloud budget? Monolithic architectures bottleneck servers and skyrocket costs. Learn how decoupling heavy GPU tasks saves your infrastructure.

Pause
Mute Enter Fullscreen
#1 about 1 min

Introduction to chatbot infrastructure and cloud challenges

How the rapid growth of an open-source chatbot project highlights the importance of avoiding structural infrastructure mistakes.

#2 about 2 min

Contrasting web developers and data scientists

Differences in resource constraints, scaling models, and deployment environments separate traditional software and data engineering roles.

#3 about 2 min

The new paradigm of AI engineers

AI engineering merges practices from both web and data science, introducing unique optimization and computing constraints.

#4 about 2 min

Understanding basic retrieval-augmented generation architectures in chatbots

A breakdown of RAG workflows entails document chunking, GPU-bound embeddings, and sequential model generation calls.

#5 about 2 min

Scaling bottlenecks in generative AI applications

The challenge of managing parallel execution, slow response times, and sequential steps complicates user-facing AI deployments.

#6 about 2 min

Building a lean initial chatbot prototype application

Using a minimal tech stack creates an efficient and lightweight artificial intelligence container application.

#7 about 2 min

Bloat and complexity from adding GPU-bound tasks

Introducing OCR and advanced embeddings drastically increases container image sizes and subsequent infrastructure scaling costs.

#8 about 2 min

Refactoring application architectures for complex technical tasks

Separating fast user-facing business logic from slow background workers handles complex orchestrations and integrations effectively.

#9 about 3 min

Infrastructure challenges of on-premise chatbot deployments

Hosting monolithic, multi-model AI applications on local organizational servers poses massive server tuning and scaling challenges.

#10 about 1 min

Architecting generative AI systems by targeted user load

Guidelines for splitting application components apply directly to managing continuous or spiky application traffic patterns.

#11 about 3 min

Best practices for utilizing large language models in production

Strategies favoring managed APIs and service-oriented architectures reduce the immense overhead of self-hosting models.

#12 about 2 min

Overlooked AI infrastructure and operational deployment barriers

Managing complex deployments like configuring GPUs on Kubernetes requires dedicated pipelines for backups and non-deterministic testing.

#13 about 3 min

Evaluating large language models using real-time video games

A hackathon approach to measuring model capability proves that smaller setups often outperform larger architectures in speed-sensitive tasks.

Matching moments

1:22 min

Leveraging the comprehensive generative artificial intelligence stack

Duan Lightfoot Duan Lightfoot · WWC 2024

1:32 min

Architectural patterns for developing robust generative AI applications

Julián Duque Julián Duque · WWC 2025

2:25 min

Progressing gracefully from generic chatbots to agentic workflows

Anshul Jindal Anshul Jindal +1 · WWC Europe 2026

3:25 min

Constructing scalable AI solutions using LangChain and LangGraph

Julián Duque Julián Duque · WWC 2025

4:20 min

Evolution and challenges of building AI applications

Roberto Carratalá Roberto Carratalá · WWC 2025

4:13 min

Architectural components of a generative AI agent

Dieter Flick · WWC 2023

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Building Stuff with GenAI - The Open Minded Workshop beyond OpenAI

Andreas Erben

CTO for Applied AI and Metaverse at daenet

Andreas Erben
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

The Broken Rung: How AI is Rebuilding Software Development from the Ground Up

Tomislav Tipurić

Chief Technology Officer, Nephos

Tomislav Tipurić
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

Designing APIs That Survive AI Agents at Scale

Phani Pendurthi

Mastercard, Principal Software Engineer

Phani Pendurthi