Data Scientist (LLM)
Role details
Job location
Tech stack
Job description
We are seeking a Senior Data Scientist with deep expertise in creating high-quality datasets for training and fine-tuning Large Language Models (LLMs). You will be responsible for designing and implementing scalable data pipelines and strategies to support all stages of LLM development: pretraining, supervised fine-tuning, and reinforcement learning with human feedback (RLHF). This role is critical to ensuring the robustness, safety, and alignment of our AI models. You will have the autonomy to explore innovative data sourcing and curation methods and the opportunity to directly influence the capabilities of state-of-the-art LLMs. As a Senior Data Scientist, you will
- Design and implement strategies for creating, sourcing, and augmenting datasets tailored for LLM training and fine-tuning.
- Develop scalable pipelines to collect, clean, filter, annotate, and validate large volumes of text data.
- Conduct data audits to ensure quality, diversity, ethical compliance, and bias mitigation.
- Collaborate with ML engineers and researchers to align datasets with training objectives and model evaluation needs.
- Use tools like Active Learning, synthetic data generation, and self-supervised learning to maximize dataset efficiency.
- Leverage human-in-the-loop (HITL) workflows for data labeling and validation where necessary.
- Contribute to building data documentation and metadata standards (e.g., Datasheets for Datasets).
- Keep up to date with research trends in dataset curation, LLM pretraining data, and benchmarking.
Requirements
- Bachelor's, Master's, or Ph.D. in Computer Science, AI, Data Science, or a related field.
- 3+ years of experience in data science, machine learning, or related roles, with demonstrated experience in dataset creation for NLP or LLMs.
- In-depth knowledge of the LLM lifecycle: pretraining, fine-tuning, alignment, and evaluation.
- Proficient in Python and data tooling ecosystems (Pandas, NumPy, spaCy, Hugging Face Datasets & Transformers).
- Hands-on experience with text data collection from diverse sources: web scraping, APIs, proprietary corpora, etc.
- Strong understanding of data quality metrics including bias detection, toxicity, and readability.
- Experience working with annotation tools (e.g., Prodigy, Label Studio) and managing annotation teams or workflows., + Experience building or contributing to datasets used in LLM pretraining or supervised fine-tuning.
- Familiarity with RLHF workflows and alignment techniques (e.g., preference modeling, reward modeling).
- Exposure to multilingual and low-resource language datasets.
- Contributions to open-source datasets, tools, or publications in dataset-centric research.
- Knowledge of ethical AI, data governance, privacy laws (e.g., GDPR), and responsible data use.
Benefits & conditions
- Indefinite contract.
- Equal pay guaranteed.
- Variable performance bonus.
- Signing bonus.
- We offer work visa sponsorship (If applicable).
- Relocation package (if applicable).
- Private health insurance.
- Eligibility for educational budget according to internal policy.
- Hybrid opportunity.
- Flexible working hours.
- Language classes and discounted lunch options
- Working in a high paced environment, working on cutting edge technologies.
- Career plan. Opportunity to learn and teach.
- Progressive Company. Happy people culture As an equal opportunity employer, Multiverse Computing is committed to building an inclusive workplace. The company welcomes people from all different backgrounds, including age, citizenship, ethnic and racial origins, gender identities, individuals with disabilities, marital status, religions and ideologies, and sexual orientations to apply. Inscribirse en esta oferta Recibir ofertas similares por correo electrónico Al crear una alerta, aceptas nuestros Términos y condiciones y Política de privacidad, y el uso de cookies.
About the company
Multiverse is a well-funded and fast-growing deep-tech company founded in 2019. We are one of the few companies working with Quantum Computing and the biggest Quantum Software company in the EU.
We provide hyper-efficient software to companies wanting to gain an edge with quantum computing and artificial intelligence. Our product, Singularity, is a software platform that contains quantum and quantum-inspired algorithms developed and patented through proof-of-concept trials we have been performing for industrial and service clients. We work in finance, energy, manufacturing, cybersecurity and many more industries.
Digital methods usually fail at efficiently tackling these problems. Quantum computing, however, provides us with a powerful toolbox to tackle these complex problems, such as outstanding optimization methods, software for quantum machine learning, and quantum enhanced Monte Carlo algorithms.
Multiverse Computing applies these cutting edge methods to provide software which is customized to your needs, giving companies a chance to derive value from the second quantum revolution.