> Markdown version of [/events/world-congress-2025/sessions/812-breaking-language](https://www.wearedevelopers.com/events/world-congress-2025/sessions/812-breaking-language). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Breaking Language Barriers: GPU-Accelerated Corpus Curation for LLM Training/Customization - **Date:** Friday, Jul 11, 2025 - **Time:** 12:00–14:00 (120 min) - **Room:** M7 (18 Seats) - **Event:** World Congress 2025 ## Description Creating large language models (LLMs) for non-English languages is frequently hindered by the scarcity or imbalance of available datasets. In this session, we introduce an advanced methodology for assembling high-quality, multilingual text corpora—spanning languages such as Spanish and French—by harnessing the computational power of GPU acceleration. We guide you through the corpus construction process, from foundational steps like text selection and cleaning to advanced approaches including semantic deduplication. The discussion will also highlight creative techniques for generating synthetic data, an essential measure for enriching datasets when authentic examples are limited. Attendees will gain actionable knowledge on corpus curation workflows that can be adapted to a wide range of non-English languages, ultimately fostering the development of more inclusive and representative language technologies. ## Speakers ### [Miguel Martínez](https://www.wearedevelopers.com/@miguel-martinez) Senior Deep Learning Data Scientist at NVIDIA ### [Roman](https://www.wearedevelopers.com/@roman) Solutions Architect at NVIDIA ## Related talks at this congress - [Adding knowledge to open-source LLMs](https://www.wearedevelopers.com/events/world-congress-2025/sessions/734-adding-knowledge-to) — Harshita Seth, Sergio Perez - [ Model Compression Techniques for Efficient LLM Deployment](https://www.wearedevelopers.com/events/world-congress-2025/sessions/811-model-compression) — Harshita Seth, Lavinia Ghita, Sergio Perez, Ziv Ilan - [Self-Hosted LLMs: From Zero to Inference](https://www.wearedevelopers.com/events/world-congress-2025/sessions/625-self-hosted-llms) — Cedric Clyburn, Roberto Carratalá - [NVIDIA Expert Session: Sovereign AI in Practice: Building, Evaluating, and Scaling Multilingual LLMs](https://www.wearedevelopers.com/events/world-congress-2025/sessions/676-nvidia-expert) — Sergio Perez, Ziv Ilan