World Congress 2025
July 11, 2025 · 09:40–10:10
Stage 2
Adding knowledge to open-source LLMs
Harshita Seth, Sergio Perez
World Congress 2025
Creating large language models (LLMs) for non-English languages is frequently hindered by the scarcity or imbalance of available datasets. In this session, we introduce an advanced methodology for assembling high-quality, multilingual text corpora—spanning languages such as Spanish and French—by harnessing the computational power of GPU acceleration. We guide you through the corpus construction process, from foundational steps like text selection and cleaning to advanced approaches including semantic deduplication. The discussion will also highlight creative techniques for generating synthetic data, an essential measure for enriching datasets when authentic examples are limited.
Attendees will gain actionable knowledge on corpus curation workflows that can be adapted to a wide range of non-English languages, ultimately fostering the development of more inclusive and representative language technologies.
World Congress 2025
July 11, 2025 · 09:40–10:10
Stage 2
Harshita Seth, Sergio Perez
World Congress 2025
July 11, 2025 · 12:00–14:00
M4 (40 Seats)
Harshita Seth, Lavinia Ghita, Sergio Perez, Ziv Ilan
World Congress 2025
July 10, 2025 · 14:50–15:20
Stage 6 - Red Hat
Cedric Clyburn, Roberto Carratalá
World Congress 2025
July 10, 2025 · 16:30–17:30
NVIDIA Lounge (Booth A16/17)
Sergio Perez, Ziv Ilan