> Markdown version of [/jobs/ext/2799222-ai-evaluation-data-scientist-9-month-contract-with-bonus-incentives](https://www.wearedevelopers.com/jobs/ext/2799222-ai-evaluation-data-scientist-9-month-contract-with-bonus-incentives). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Ai Evaluation Data Scientist (9-Month Contract With Bonus Incentives) - **Company:** Multiverse Computing - **Location:** Madrid, Spain - **Contract:** Temporary contract - **Skills:** Artificial Intelligence, Data Analysis, Big Data, Data Visualization, Python (Programming Language), Machine Learning, NumPy, Quantum Computing, Large Language Models, Model Validation, Pandas, Scikit Learn, Information Technology, Data Analytics, Machine Learning Operations, Software Version Control - **Published:** September 8, 2026 - **Apply:** https://www.buscojobs.com.es/ai-evaluation-data-scientist-9-month-contract-with-bonus-incentives-en-madrid-ID-369263447 ## About the Role Master's degree in Computer Science, Data Science, Mathematics, Engineering, or related field. 2+ years of experience in AI/ML model evaluation or large-scale data analysis. Strong proficiency in Python, and evaluation/data analysis libraries (NumPy, Pandas, scikit-learn, etc.). Direct experience creating evaluation pipelines for AI or data-driven systems, with a track record in metric definition and performance analysis. Familiarity with ML infrastructure (experiment tracking, model versioning, continuous evaluation). Experience working with large unstructured datasets and data visualization/reporting to communicate findings. Preferred Qualifications Experience evaluating generative AI agents, LLMs, or advanced SML components in production or R&D settings. Exposure to human annotation workflows and automated labeling tools. Knowledge of statistical analysis, fairness, bias, and robust metric development for model evaluation. Familiarity with ML experimentation platforms, cloud-based AI infrastructure, and advanced model validation techniques. ## Description We are looking to fill this roleimmediatelyand are reviewing applications daily. Expect a fast, transparent process with quick feedback.Why join us?We are a European deep-tech leader in quantum and AI, backed by major global strategic investors and strong EU support. Our groundbreaking technology is already transforming how AI is deployed worldwide - compressing large language models by up to 95% without losing accuracy and cutting inference costs by *****%.Joining us means working on cutting-edge solutions that make AI faster, greener, and more accessible - and being part of a company often described as a "quantum-AI unicorn in the making."We offerCompetitive annual salaryTwo unique bonuses: signing bonus at incorporation and retention bonus at contract completion.Relocation package (if applicable).Up to 9-month contract, ending on June ****.Hybrid role and flexible working hours.Be part of a fast-scaling Series B company at the forefront of deep tech.Equal pay guaranteed.International exposure in a multicultural, cutting-edge environment.As an AI Evaluation Data Scientist, you will:Build robust evaluation pipelines that accurately mirror real production use cases for AI agentic systems and SML components.Define and track success metrics for both agentic workflows and key model components, focusing on measurement aligned with practical production performance rather than pure benchmarks.Design annotation frameworks and automated evaluation tools to conduct rigorous, repeatable analysis of model/system output.Execute deep-dive data investigations, monitor success and failure cases, and deliver actionable feedback loops to guide model/system improvements.Collaborate with engineers and researchers to shape experiments, validate hypotheses, and iterate toward higher performing model deployments.Communicate evaluation methodologies, results, and insights to technical and non-technical stakeholders, supporting ongoing improvement and transparency.Required QualificationsMaster's degree in Computer Science, Data Science, Mathematics, Engineering, or related field.2+ years of experience in AI/ML model evaluation or large-scale data analysis.Strong proficiency in Python, and evaluation/data analysis libraries (NumPy, Pandas, scikit-learn, etc.).Direct experience creating evaluation pipelines for AI or data-driven systems, with a track record in metric definition and performance analysis.Familiarity with ML infrastructure (experiment tracking, model versioning, continuous evaluation).Experience working with large unstructured datasets and data visualization/reporting to communicate findings.Preferred QualificationsExperience evaluating generative AI agents, LLMs, or advanced SML components in production or R&D settings.Exposure to human annotation workflows and automated labeling tools.Knowledge of statistical analysis, fairness, bias, and robust metric development for model evaluation.Familiarity with ML experimentation platforms, cloud-based AI infrastructure, and advanced model validation techniques.About Multiverse ComputingFounded in ****, we are a well-funded, fast-growing deep-tech company with a team of 180+ employees worldwide. Recognized by CB Insights ***** & ****) as one of theTop 100 most promising AI companies globally, we are also the largest quantum software company in the EU.Our flagship products address critical industry needs:CompactifAI ? a groundbreaking compression tool for foundational AI models, reducing their size by up to 95% while maintaining accuracy, enabling portability across devices from cloud to mobile and beyond.Singularity ? a quantum and quantum-inspired optimization platform used by blue-chip companies in finance, energy, and manufacturing to solve complex challenges with immediate performance gains.You'll be working alongside world-leading experts in quantum computing and AI, developing solutions that deliver real-world impact for global clients. We are committed to an inclusive, ethics-driven culture that values sustainability, diversity, and collaboration - a place where passionate people can grow and thrive. Come and join us!As an equal opportunity employer, Multiverse Computing is committed to building an inclusive workplace. The company welcomes people from alldifferent backgrounds, including age, citizenship, ethnic and racial origins, gender identities, individuals with disabilities, marital status, religions and ideologies, and sexual orientations to apply. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [The State of WebDev AI 2025 Results: What Can We Learn?](https://www.wearedevelopers.com/magazine/581-the-state-of-webdev-ai-2025-results-what-can-we-learn)