World Congress 2026 Europe - Virtual Stage Jul 2, 2026 Session details

Modern Web Data Extraction: Techniques, Tools, and Legal & Ethical Considerations

Domagoj Marić

Is your web scraper risking IP bans or lawsuits? Master strict politeness policies and modern Python tools to build resilient, legally compliant data extraction pipelines.

Pause
Mute Enter Fullscreen
#1 about 3 min

Core definitions of modern web data extraction

Core definitions distinguishing iterating through links from extracting specific site data.

#2 about 2 min

Common use cases for extracting external web data

How businesses utilize scraped external data for competitive analysis and model training.

#3 about 3 min

Legal precedents and the gray zone of scraping

Lessons learned from prominent lawsuits regarding unauthorized access and platform aggregation.

#4 about 1 min

Key policies for compliant and polite crawling

Important mechanisms for selecting appropriate targets without overloading server resources.

#5 about 3 min

Understanding robots exclusion protocol syntax and rules

Using text files to determine which server directories permit automated extraction.

#6 about 4 min

Terms of service and avoiding server overload

Mitigating block risks by managing request rates and configuring appropriate identification strings.

#7 about 3 min

Leveraging website sitemaps to optimize crawling frequency

Utilizing formatted structural files to schedule off-peak retrieval and identify modified content.

#8 about 3 min

Regulatory compliance and technical prerequisites for scraping

Safeguarding privacy with anonymization alongside necessary skills in document traversal and text analytics.

#9 about 3 min

Handling http network errors and anti-scraping challenges

Overcoming common blockers through protocol understanding, proxy rotation, and ongoing scraper maintenance.

#10 about 3 min

Exploring internal endpoints and alternative data services

Bypassing presentation layer extraction by interacting directly with backend request payloads or external providers.

#11 about 7 min

Choosing the appropriate python tools and frameworks

Selecting between static parsing libraries, heavy-duty frameworks, and browser automation utilities.

#12 about 2 min

Establishing a structured workflow for web scraping

A systematic methodology from checking permissions to inspecting elements and post-processing results.

#13 about 4 min

Practical extraction examples with static and dynamic methods

Step-by-step demonstrations extracting blog content using straightforward requests versus simulated human navigation.

#14 about 3 min

Summary of data extraction decisions and best practices

Final review of verifying legality, analyzing load strategies, and handling extracted artifacts responsibly.

Matching moments

5:54 min

The technical evolution of modern web scraping infrastructure

Chris Heilmann +4 · LIVE

1:19 min

Defining web scraping and recognizing proper use cases

Lars Kölker · WWC 2023

6:20 min

Securing competitive intelligence using automated web extraction crawlers

Coffee With Developers

1:19 min

Tailoring custom web scrapers for artificial intelligence training

Vidas Bacevičius Vidas Bacevičius · WWC 2025

2:14 min

Understanding the basic mechanics of automated web scraping

Vidas Bacevičius Vidas Bacevičius · WWC 2025

3:40 min

Overcoming modern anti-bot mechanisms and network access restrictions

Vidas Bacevičius Vidas Bacevičius · WWC 2025

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Small LLM in your Browser: Huge Opportunities for Web Applications

Daniel Ostrovsky

UI/UX Architect at Payoneer | AI Architect | Full Cycle Development Expert | Public Speaker | Open Source Contributor |

Daniel Ostrovsky
Open session

World Congress 2026 North America

AI-Driven API Design

Mike Amundsen

Author, Adviser, Trainer

Mike Amundsen
Open session

World Congress 2026 North America

Building Pragmatic AI: 10 AI Features Your Users Actually Want

Jonathan "J." Tower

.NET Foundation Board | 12x Microsoft MVP | Founder & Consultant

Jonathan "J." Tower
Open session

World Congress 2026 North America

Developer Liability in the AI Agent Era: Building Responsibly

Alla Barbalat

Freelancer Trade Show Spokesmodel and Tech Event Host

Alla Barbalat
Open session

World Congress 2026 North America

Building Stuff with GenAI - The Open Minded Workshop beyond OpenAI

Andreas Erben

CTO for Applied AI and Metaverse at daenet

Andreas Erben
Open session

World Congress 2026 North America

Fast by Design: A Masterclass in High-Performance Web Engineering

Aaron Grogg

Senior web developer, committed to improving the user experience by improving web performance

Aaron Grogg