World Congress 2026 Europe Jul 10, 2026 Session details

Outclassing Frontier LLMs at Extracting Information

Etienne Bernard

A specialized 4B model reduced complex document extraction errors from 70% to just 3%. Discover how NuExtract3 outclasses massive frontier LLMs directly on a single GPU.

Pause
Mute Enter Fullscreen
#1 about 2 min

Introduction to specialized document extraction models

Specialized language models overcome traditional barriers by efficiently extracting information from complex business documents.

#2 about 3 min

Structured data extraction for automated data entry

Extracting specific information into JSON schemas enables predictable automated data entry from documents like identity cards and invoices.

#3 about 1 min

Industry applications and strict accuracy requirements

Regulated industries like banking and healthcare require highly accurate structured extraction processes to avoid critical downstream consequences.

#4 about 3 min

Transforming full document content into markdown for data retrieval

Converting entire documents into text-based formats like markdown provides essential accessible data for retrieval-augmented generation systems.

#5 about 4 min

Limitations of general-purpose language models in complex extraction

High operational costs and an inability to reliably convey confidence levels hinder general-purpose models in production extraction environments.

#6 about 2 min

Training specialized extraction models through supervised learning

Fine-tuning a general-purpose model with millions of diverse extraction examples enables efficient and precise domain-specific document understanding.

#7 about 4 min

Open-source extraction model capabilities and live demonstration

The open-source extraction model employs succinct reasoning techniques to efficiently understand complex visual layouts and correctly structure raw data.

#8 about 2 min

Benchmarking specialized extraction against general purpose models

Focused extraction models utilize optimized thinking tokens to significantly outperform much larger general-purpose variants in targeted structured data benchmarks.

#9 about 3 min

Enterprise capabilities of the professional extraction platform

The scaled-up professional version achieves frontier-level performance for customized private enterprise deployments while dramatically reducing core computational requirements.

#10 about 3 min

Real-world customization reducing error rates to human levels

Customizing extraction models for highly specific client documents drastically drops operational error percentages at a fraction of traditional manual processing costs.

#11 about 2 min

Optimizing human-in-the-loop workflows through uncertainty scoring

Providing extraction models with specialized mechanisms to express uncertainty dramatically improves human review efficiency within compliance-heavy sectors.

#12 about 5 min

Handling mixed document quality and multilingual text formats

Specialized extraction architectures natively manage low-resolution scans and complex localized scripts substantially better than legacy document automation solutions.

Matching moments

6:43 min

Replacing LLMs with specialized data extractors

Dmytro Kurets Dmytro Kurets · Europe 2026 Virtual

4:39 min

Enhancing legacy record extraction using machine learning

Maria Doina Irimias · LIVE

1:26 min

The naive document extraction pipeline architecture

Nazeer Saeed Nazeer Saeed · World Congress 2026 Europe

1:45 min

Evaluating document extraction performance with benchmarking tools

Luca Bianchi Luca Bianchi · World Congress 2026 Europe

3:13 min

Processing unstructured data through intelligent document understanding

Boris Krumrey +2 · LIVE

1:10 min

Business and privacy impacts of raw OCR pipelines

Nazeer Saeed Nazeer Saeed · World Congress 2026 Europe

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 24, 2026 · 17:30–18:00

Stage 6

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

September 25, 2026 · 11:40–12:10

Stage 9

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

September 24, 2026 · 16:10–16:40

Stage 9

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Legare Kerrison, Cedric Clyburn

Legare Kerrison
Cedric Clyburn
Open session

World Congress 2026 North America

September 24, 2026 · 13:30–14:00

Stage 9

Understanding LLM Architectures: Inside the Design of Modern Models

Jofia Jose Prakash

Director - AI & Governance at Humanity + AI, Inc

Jofia Jose Prakash
Open session

World Congress 2026 North America

September 25, 2026 · 14:10–14:40

Stage 4

Headroom: A Context Optimization Layer for LLM Applications

Tejas Chopra

Senior Software Engineer at Netflix

Tejas Chopra
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 5

Edge AI: Running Agentic Intelligence Where Internet Can't Reach

Nitin Eusebius

AWS - Principal Solutions Architect

Nitin Eusebius