World Congress 2026 Europe Jul 10, 2026 Session details

Garbage In, Garbage Out: Engineering Reliable AI Document Extraction Pipelines

Nazeer Saeed

AI doesn't fail on complex documents. It fails on brute-force data pipelines. Discover how on-device validation and structured JSON extraction drastically cut LLM token costs and eliminate hallucinations.

Pause
Mute Enter Fullscreen
#1 about 5 min

Introduction and the receipt extraction problem

How simple receipt imaging creates significant challenges for downstream automated data processing.

#2 about 2 min

The naive document extraction pipeline architecture

The common but flawed three-step process of scanning, performing OCR, and passing raw text directly to an LLM.

#3 about 3 min

Why raw OCR pipelines fail in reality

How blurry images, layout complexity, and a lack of structural context break standard document extraction flows.

#4 about 2 min

Business and privacy impacts of raw OCR pipelines

The financial, operational, and security risks of sending raw, token-heavy data to cloud language models.

#5 about 3 min

A five-stage architecture for reliable document pipelines

Designing a sequential flow for capturing, processing, extracting, and validating documents before LLM ingestion.

#6 about 2 min

Evaluating the performance of a naive OCR pipeline

A demonstration showing how processing a raw invoice image through Docling OCR yields high latencies and structural errors.

#7 about 2 min

Improving OCR performance through image quality preprocessing

Applying automatic cropping, perspective correction, and filtering to drastically reduce processing time and minimize recognition errors.

#8 about 2 min

Overcoming OCR limitations by extracting structural JSON

Transitioning from flat plaintext to structured output formats that preserve tables, data fields, and visual layouts.

#9 about 5 min

Extracting contextual JSON metadata for LLM processing

A side-by-side runtime comparison showcasing faster execution speeds and targeted component parsing with structured JSON outputs.

#10 about 2 min

Preventing document errors early at the capture stage

Implementing on-device quality controls to reject unreadable images before they enter the data extraction backend.

#11 about 3 min

Why better data inputs beat bigger language models

Feeding corrected, well-structured JSON payloads to AI interfaces ensures cheaper computations and highly predictable document parsing.

#12 about 2 min

Identifying the most critical extraction pipeline optimizations

Emphasizing the initial scan quality and contextual metadata structuring as the distinct stages yielding the highest data reliability improvements.

Matching moments

4:49 min

Handling mixed document quality and multilingual text formats

Etienne Bernard Etienne Bernard · WWC Europe 2026

3:13 min

Processing unstructured data through intelligent document understanding

Boris Krumrey +2 · LIVE

2:21 min

Real-world customization reducing error rates to human levels

Etienne Bernard Etienne Bernard · WWC Europe 2026

4:25 min

Overcoming AI hallucinations and restrictive content guardrails

Perf + AI

44 sec

Overcoming challenges in unstructured document layout ingestion

Alex Soto Alex Soto +1 · WWC 2025

37 sec

Building core AI literacy and algorithmic transparency

Manjuri Sinha Manjuri Sinha · WWC 2025

Upcoming sessions on this topic

Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

The Things Your AI Isn't Telling You

Desmond Lamptey

Lead Software Engineer @ Capital One

Desmond Lamptey
Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

AI ROI: The Hard Unit Economics of AI-Native Engineering

Manu Gurudatha

Manu Gurudatha, VP of Engineering at PagerDuty

Manu Gurudatha
Open session

World Congress 2026 North America

Understanding LLM Architectures: Inside the Design of Modern Models

Jofia Jose Prakash

Enterprise AI Architect at American Chemical Society

Jofia Jose Prakash
Open session

World Congress 2026 North America

Engineering the Pivot: How Creative Strategy Solves the Hard Problems of AI Accuracy and Scale

Shruti Tiwari

AI/ML product manager, Dell

Shruti Tiwari