World Congress 2026 Europe • Jul 10, 2026 • Session details

Outclassing Frontier LLMs at Extracting Information

Etienne Bernard

A specialized 4B model reduced complex document extraction errors from 70% to just 3%. Discover how NuExtract3 outclasses massive frontier LLMs directly on a single GPU.

Pause
Mute Enter Fullscreen
#1 about 2 min

Introduction to specialized document extraction models

Specialized language models overcome traditional barriers by efficiently extracting information from complex business documents.

#2 about 3 min

Structured data extraction for automated data entry

Extracting specific information into JSON schemas enables predictable automated data entry from documents like identity cards and invoices.

#3 about 1 min

Industry applications and strict accuracy requirements

Regulated industries like banking and healthcare require highly accurate structured extraction processes to avoid critical downstream consequences.

#4 about 3 min

Transforming full document content into markdown for data retrieval

Converting entire documents into text-based formats like markdown provides essential accessible data for retrieval-augmented generation systems.

#5 about 4 min

Limitations of general-purpose language models in complex extraction

High operational costs and an inability to reliably convey confidence levels hinder general-purpose models in production extraction environments.

#6 about 2 min

Training specialized extraction models through supervised learning

Fine-tuning a general-purpose model with millions of diverse extraction examples enables efficient and precise domain-specific document understanding.

#7 about 4 min

Open-source extraction model capabilities and live demonstration

The open-source extraction model employs succinct reasoning techniques to efficiently understand complex visual layouts and correctly structure raw data.

#8 about 2 min

Benchmarking specialized extraction against general purpose models

Focused extraction models utilize optimized thinking tokens to significantly outperform much larger general-purpose variants in targeted structured data benchmarks.

#9 about 3 min

Enterprise capabilities of the professional extraction platform

The scaled-up professional version achieves frontier-level performance for customized private enterprise deployments while dramatically reducing core computational requirements.

#10 about 3 min

Real-world customization reducing error rates to human levels

Customizing extraction models for highly specific client documents drastically drops operational error percentages at a fraction of traditional manual processing costs.

#11 about 2 min

Optimizing human-in-the-loop workflows through uncertainty scoring

Providing extraction models with specialized mechanisms to express uncertainty dramatically improves human review efficiency within compliance-heavy sectors.

#12 about 5 min

Handling mixed document quality and multilingual text formats

Specialized extraction architectures natively manage low-resolution scans and complex localized scripts substantially better than legacy document automation solutions.

Matching moments

6:43 min

Replacing LLMs with specialized data extractors

Dmytro Kurets Dmytro Kurets · Europe 2026 Virtual

4:39 min

Enhancing legacy record extraction using machine learning

Maria Doina Irimias · LIVE

1:26 min

The naive document extraction pipeline architecture

Nazeer Saeed Nazeer Saeed · World Congress 2026 Europe

1:45 min

Evaluating document extraction performance with benchmarking tools

Luca Bianchi Luca Bianchi · World Congress 2026 Europe

3:13 min

Processing unstructured data through intelligent document understanding

Boris Krumrey +2 · LIVE

1:10 min

Business and privacy impacts of raw OCR pipelines

Nazeer Saeed Nazeer Saeed · World Congress 2026 Europe