World Congress 2026 Europe Jul 10, 2026 Session details

Garbage In, Garbage Out: Engineering Reliable AI Document Extraction Pipelines

Nazeer Saeed

AI doesn't fail on complex documents. It fails on brute-force data pipelines. Discover how on-device validation and structured JSON extraction drastically cut LLM token costs and eliminate hallucinations.

Pause
Mute Enter Fullscreen
#1 about 5 min

Introduction and the receipt extraction problem

How simple receipt imaging creates significant challenges for downstream automated data processing.

#2 about 2 min

The naive document extraction pipeline architecture

The common but flawed three-step process of scanning, performing OCR, and passing raw text directly to an LLM.

#3 about 3 min

Why raw OCR pipelines fail in reality

How blurry images, layout complexity, and a lack of structural context break standard document extraction flows.

#4 about 2 min

Business and privacy impacts of raw OCR pipelines

The financial, operational, and security risks of sending raw, token-heavy data to cloud language models.

#5 about 3 min

A five-stage architecture for reliable document pipelines

Designing a sequential flow for capturing, processing, extracting, and validating documents before LLM ingestion.

#6 about 2 min

Evaluating the performance of a naive OCR pipeline

A demonstration showing how processing a raw invoice image through Docling OCR yields high latencies and structural errors.

#7 about 2 min

Improving OCR performance through image quality preprocessing

Applying automatic cropping, perspective correction, and filtering to drastically reduce processing time and minimize recognition errors.

#8 about 2 min

Overcoming OCR limitations by extracting structural JSON

Transitioning from flat plaintext to structured output formats that preserve tables, data fields, and visual layouts.

#9 about 5 min

Extracting contextual JSON metadata for LLM processing

A side-by-side runtime comparison showcasing faster execution speeds and targeted component parsing with structured JSON outputs.

#10 about 2 min

Preventing document errors early at the capture stage

Implementing on-device quality controls to reject unreadable images before they enter the data extraction backend.

#11 about 3 min

Why better data inputs beat bigger language models

Feeding corrected, well-structured JSON payloads to AI interfaces ensures cheaper computations and highly predictable document parsing.

#12 about 2 min

Identifying the most critical extraction pipeline optimizations

Emphasizing the initial scan quality and contextual metadata structuring as the distinct stages yielding the highest data reliability improvements.

Matching moments

4:49 min

Handling mixed document quality and multilingual text formats

Etienne Bernard Etienne Bernard · World Congress 2026 Europe

3:13 min

Processing unstructured data through intelligent document understanding

Boris Krumrey +2 · LIVE

2:21 min

Real-world customization reducing error rates to human levels

Etienne Bernard Etienne Bernard · World Congress 2026 Europe

4:25 min

Overcoming AI hallucinations and restrictive content guardrails

Perf + AI

44 sec

Overcoming challenges in unstructured document layout ingestion

Alex Soto Alex Soto +1 · World Congress 2025

37 sec

Building core AI literacy and algorithmic transparency

Manjuri Sinha Manjuri Sinha · World Congress 2025

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 25, 2026 · 11:40–12:10

Stage 9

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

September 25, 2026 · 16:50–17:20

Stage 5

The Things Your AI Isn't Telling You

Desmond Lamptey

Lead Software Engineer @ Capital One

Desmond Lamptey
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 1

Anatomy of an AI Request: Where Latency and Cost Are Really Born

Dan Fu

VP of Kernels at Together AI

Dan Fu
Open session

World Congress 2026 North America

September 25, 2026 · 16:10–16:40

Stage 1

The Five Percent Club: The Culture and Technological Shift Behind Successful AI Deployments

Tara Hernandez

Tara Hernandez, VP of Developer Productivity at MongoDB

Tara Hernandez
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 5

Edge AI: Running Agentic Intelligence Where Internet Can't Reach

Nitin Eusebius

AWS - Principal Solutions Architect

Nitin Eusebius
Open session

World Congress 2026 North America

September 24, 2026 · 11:10–11:15

Outdoor Stage

Architecting the 100X SDLC: Building Production Trust into AI-Assisted Delivery

Ranjan Parthasarathy

Founder, CPTO/CEO at AXIOMSTUDIO.AI

Ranjan Parthasarathy