World Congress 2025

Breaking Language Barriers: GPU-Accelerated Corpus Curation for LLM Training/Customization

July 11, 2025 12:00 – 14:00 · 120 min M7 (18 Seats)

What this session covers

Creating large language models (LLMs) for non-English languages is frequently hindered by the scarcity or imbalance of available datasets. In this session, we introduce an advanced methodology for assembling high-quality, multilingual text corpora—spanning languages such as Spanish and French—by harnessing the computational power of GPU acceleration. We guide you through the corpus construction process, from foundational steps like text selection and cleaning to advanced approaches including semantic deduplication. The discussion will also highlight creative techniques for generating synthetic data, an essential measure for enriching datasets when authentic examples are limited.

Attendees will gain actionable knowledge on corpus curation workflows that can be adapted to a wide range of non-English languages, ultimately fostering the development of more inclusive and representative language technologies.

Related talks at this congress

Open session

World Congress 2025

July 11, 2025 · 09:40–10:10

Stage 2

Adding knowledge to open-source LLMs

Harshita Seth, Sergio Perez

Harshita Seth
Sergio Perez
Open session

World Congress 2025

July 11, 2025 · 12:00–14:00

M4 (40 Seats)

Model Compression Techniques for Efficient LLM Deployment

Harshita Seth, Lavinia Ghita, Sergio Perez, Ziv Ilan

Harshita Seth
Lavinia Ghita
Sergio Perez
Ziv Ilan
Open session

World Congress 2025

July 10, 2025 · 14:50–15:20

Stage 6 - Red Hat

Self-Hosted LLMs: From Zero to Inference

Cedric Clyburn, Roberto Carratalá

Cedric Clyburn
Roberto Carratalá
Open session

World Congress 2025

July 10, 2025 · 16:30–17:30

NVIDIA Lounge (Booth A16/17)

NVIDIA Expert Session: Sovereign AI in Practice: Building, Evaluating, and Scaling Multilingual LLMs

Sergio Perez, Ziv Ilan

Sergio Perez
Ziv Ilan
All sessions at this congress