AI Researcher - Multilingual Data

Jobgether
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English

Job location

Tech stack

Artificial Intelligence
Data Deduplication
Python
Machine Learning
Language Modeling
Natural Language Processing
Open Source Technology
PyTorch
Transfer Learning
Deep Learning
Machine Learning Operations

Job description

Join a cutting-edge research environment where multilingual AI innovation meets real-world impact. In this role, you will drive the development of high-quality multilingual datasets and research strategies that power next-generation language models across diverse languages and domains. Working at the intersection of research and engineering, you will transform scientific discoveries into scalable production solutions while contributing to state-of-the-art advancements in natural language processing. This position offers the opportunity to publish influential research, collaborate with highly skilled experts, and shape the future of multilingual AI in a fast-paced, innovation-driven environment. If you are passionate about solving complex language challenges and advancing machine learning research, this is an opportunity to make a meaningful global impact. Accountabilities

  • Design and conduct research focused on multilingual datasets, including data collection, filtering, deduplication, quality assessment, and optimization.
  • Develop innovative strategies for low-resource and long-tail languages through advanced sampling, data augmentation, and curriculum learning techniques.
  • Research and improve multilingual large language models by enhancing cross-lingual transfer, alignment, robustness, and representation learning.
  • Build, maintain, and refine multilingual evaluation benchmarks to measure model quality and performance across languages.
  • Collaborate closely with machine learning engineers and researchers to influence training pipelines, model architectures, and production deployment strategies.
  • Publish research findings at leading AI and NLP conferences while contributing to open-source initiatives when appropriate.
  • Translate research outcomes into practical improvements that enhance production-ready AI systems.

Requirements

  • Advanced background in Natural Language Processing, Machine Learning, Artificial Intelligence, or a closely related field.
  • Proven research experience in multilingual or cross-lingual language modeling with publications at recognized conferences or journals such as ACL, EMNLP, NeurIPS, ICML, or ICLR.
  • Hands-on experience working with large-scale multilingual text datasets and modern machine learning workflows.
  • Strong understanding of multilingual tokenization, vocabulary design, transfer learning, multilingual representation learning, dataset quality assessment, filtering techniques, and bias mitigation.
  • Proficiency in Python and modern deep learning frameworks such as PyTorch or JAX.
  • Ability to work independently, manage research initiatives, and deliver high-quality results in a fast-moving startup environment.
  • Experience with low-resource languages, non-Latin scripts, multilingual evaluation benchmarks (XTREME, FLORES, TyDi QA), open-source NLP projects, or large language model training is considered a strong advantage.

Benefits & conditions

  • Competitive compensation package.
  • Meaningful equity opportunity within an early-stage, high-growth company.
  • Significant ownership over research direction and technical decision-making.
  • Opportunity to balance academic research with real-world production impact.
  • Access to large-scale multilingual datasets, modern AI infrastructure, and rapid experimentation cycles.
  • Collaborative environment that values innovation, research excellence, and continuous learning.
  • Opportunity to publish at leading international AI and NLP conferences while contributing to impactful open-source initiatives.

Apply for this position