World Congress 2025 Aug 20, 2025 Session details

Scrape, Train, Predict: The Lifecycle of Data for AI Applications

Vidas Bacevičius

Public datasets are too stale to fuel modern LLMs. Learn how engineers are using AI-driven adaptive parsing to bypass anti-bot walls and build resilient data extraction pipelines.

Pause
Mute Enter Fullscreen
#1 about 3 min

Understanding the basic mechanics of automated web scraping

Proxy servers mask scraper locations to safely automate HTTP requests for large-scale data extraction.

#2 about 3 min

Analyzing historical data and predicting future market trends

Continuous web scraping enables tracking search engine rankings, pricing strategies, and demand forecasting.

#3 about 3 min

Addressing the limitations of common public training datasets

Static public datasets lack the recency, scale, and multimodality needed for competitive AI training operations.

#4 about 2 min

Tailoring custom web scrapers for artificial intelligence training

Building customized in-house scrapers enables access to real-time, targeted, and multimodal web data streams.

#5 about 3 min

Enhancing generative models with real-time external data retrieval

Integrating fast web scraping with retrieval-augmented generation feeds language models with live external knowledge.

#6 about 4 min

Overcoming modern anti-bot mechanisms and network access restrictions

Sophisticated blocking methods, CAPTCHA challenges, and geolocation restrictions complicate modern automated web scraping workflows.

#7 about 4 min

Identifying deceptive server content and parsing messy HTML

Brittle XPath selectors and deceptive server responses make traditional manual HTML parsing difficult to maintain.

#8 about 2 min

Training classification models for resilient automated HTML parsing

Machine learning algorithms process labeled historical documents to automatically identify structural element patterns across websites.

#9 about 3 min

Validating website responses and adapting to layout transformations

Artificial intelligence feedback loops validate true server responses and dynamically extract targeted elements from evolving layouts.

#10 about 4 min

Automating data extraction pipelines using natural language assistants

AI assistants parse natural language prompts to automatically generate reliable extraction instructions for complex e-commerce platforms.

Matching moments

5:54 min

The technical evolution of modern web scraping infrastructure

Chris Heilmann +4 · LIVE

2:54 min

Technical workflow for generating self-healing scrapers

Ariel Shulman Ariel Shulman +1 · World Congress 2026 Europe

2:14 min

Assessing the future of AI in web performance optimization

Perf + AI

49 sec

Role of web scraping in training language models

Jan Curn Jan Curn · World Congress 2025

1:35 min

Hypocrisy in artificial intelligence data scraping practices

Chris Heilmann +2 · LIVE

4:28 min

Exploring artificial intelligence as a solution for web accessibility

Chris Heilmann +2 · LIVE

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 25, 2026 · 12:55–13:25

Stage 1

From Simulation to Reality: Overcoming the Data Scarcity Crisis in Physical AI

Mitesh Patel

NVIDIA Corporation, Developer Advocate -- Manager

Mitesh Patel
Open session

World Congress 2026 North America

September 25, 2026 · 15:30–16:00

Stage 6

Small LLM in your Browser: Huge Opportunities for Web Applications

Daniel Ostrovsky

UI/UX Architect at Payoneer | AI Architect | Full Cycle Development Expert | Public Speaker | Open Source Contributor |

Daniel Ostrovsky
Open session

World Congress 2026 North America

September 23, 2026 · 10:00–17:00

Stage 13

Building Pragmatic AI: 10 AI Features Your Users Actually Want

Jonathan "J." Tower

.NET Foundation Board | 12x Microsoft MVP | Founder & Consultant

Jonathan "J." Tower
Open session

World Congress 2026 North America

September 23, 2026 · 10:00–17:00

Stage 11

Building Stuff with GenAI - The Open Minded Workshop beyond OpenAI

Andreas Erben

CTO for Applied AI and Metaverse at daenet

Andreas Erben
Open session

World Congress 2026 North America

September 25, 2026 · 11:40–12:10

Stage 9

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

September 24, 2026 · 11:10–11:15

Outdoor Stage

Architecting the 100X SDLC: Building Production Trust into AI-Assisted Delivery

Ranjan Parthasarathy

Founder, CPTO/CEO at AXIOMSTUDIO.AI

Ranjan Parthasarathy