World Congress 2025 • Aug 20, 2025 • Session details

How to scrape modern websites to feed AI agents

Jan Curn

Feeding raw HTML to your LLM guarantees garbage outputs. Learn to bypass modern bot blockers to extract clean data. Transform your AI into a web-integrated operator.

Pause
Mute Enter Fullscreen
#1 about 1 min

Role of web scraping in training language models

Extensive public internet crawls remain the foundational mechanism enabling the vast reasoning capabilities of modern language models.

#2 about 4 min

Overcoming language model limitations with context engineering

Retrieval-augmented generation pipelines directly append external facts into prompt streams to prevent outdated intelligence without triggering costly parameter updates.

#3 about 2 min

Extracting dynamic web content using headless browsers

Rendering dynamic client-side applications necessitates migrating from naive HTTP requests to deploying scripted headless browsers.

#4 about 3 min

Bypassing anti-bot protections with proxies and emulation

Scraping systems rely on varied residential proxy rotations and hardware parameter emulation to remain undetected by aggressive firewall endpoints.

#5 about 2 min

Cleaning extraction payloads and scaling crawler infrastructure

Sanitizing retrieved documents with aggressive tag filtering heavily curtails token expenditures during the text embedding phase.

#6 about 2 min

Automating web extraction with cloud-based components

Pre-configured environment modules allow rapid configuration of deep site crawling routines that export directly into compatible markdown formats.

#7 about 4 min

Connecting crawler payloads directly to vector databases

Piping markdown outputs automatically into embedded index caches facilitates immediate deployment of custom generative bots.

#8 about 1 min

Integrating data extraction directly with orchestration frameworks

Packaging extraction routines into standardized containers unlocks immediate interoperability across popular orchestration libraries like Langchain.

#9 about 2 min

Building resilient agent architectures via model context protocols

Adopting dynamic connection protocols empowers autonomous systems to seamlessly adjust to changing endpoint schemas without static API maintenance.

#10 about 3 min

Implementing dynamic tool discovery within agent workflows

Equipping models with protocol-driven discovery endpoints allows them to intuitively locate and deploy specialized automation tools strictly as demanded.

#11 about 1 min

Monetizing community-built extraction scripts via an ecosystem

Exposing modular data scripts to broad developer ecosystems allows engineers to directly capture revenue by servicing niche aggregation needs.

Matching moments

2:21 min

Core concepts and mechanics behind Web MCP

Alex Nahas · Coffee With Developers

4:20 min

Automating tasks securely through agentic browser protocols

Christian Liebel Christian Liebel · World Congress 2026 Europe

5:54 min

The technical evolution of modern web scraping infrastructure

Chris Heilmann Chris Heilmann +4 · LIVE

3:35 min

Impact of AI data scraping on open web content

Léonie Watson · A11y + AI

5:54 min

Standardizing agent interactions with the Web MCP proposal

Chris Heilmann Chris Heilmann +2 · LIVE

1:48 min

Exploring AI integrations in modern agile development workflows

Markus Walker Markus Walker · World Congress 2023