> Markdown version of [/jobs/ext/3074247-senior-scientist-genai-evaluation](https://www.wearedevelopers.com/jobs/ext/3074247-senior-scientist-genai-evaluation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Scientist - GenAI Evaluation - **Company:** J&J Family of Companies - **Location:** Madrid, Spain - **Salary:** €55,400.0 - €87,860.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Data Analysis, Bioinformatics, Health Informatics, Computational Biology, Continuous Delivery, Data Visualization, Decision Support Systems, Python (Programming Language), Machine Learning, Retrieval-Augmented Generation, Large Language Models, Prompt Engineering, Generative AI, Information Technology, Machine Learning Operations - **Published:** September 25, 2026 - **Apply:** https://dejobs.org/x/x/23E76601658A41E8A81DD565B685ECF7/job/ ## About the Role Master's degree in AI/ML, Computer Science, Data Science, Computational Biology, Bioinformatics, Biomedical Engineering, Applied Mathematics, Biostatistics, or a related field required. PhD preferred., We are looking for someone with 6+ years of hands-on experience in AI/ML evaluation or data science (Master's) or 3+ years of industry experience (PhD). You should have experience designing and running evaluation frameworks, scientific benchmarks, or quality assessments for AI/ML systems, and hands-on work with generative AI - large language models, retrieval-augmented generation, agentic frameworks, and prompt engineering. We also value strong proficiency in Python and modern AI/ML tooling (evaluation harnesses, embedding models, vector databases, LLM APIs), the ability to translate expert scientific judgment into measurable criteria, rubrics, and reproducible protocols, and a collaborative, self-driven approach to working across multidisciplinary teams. PREFERRED: We would love to find someone who also brings experience with AI/ML evaluation in regulated environments (FDA, EMA, or equivalent), understanding of the drug development pipeline and biomedical data types, or domain expertise in oncology, immunology, or neuroscience. Experience designing or validating LLM-as-judge systems, implementing CI/CD evaluation pipelines, or publications in AI evaluation, NLP, or biomedical informatics are all a plus. OTHER: English proficiency is required (written and verbal). This is a hybrid role based in Madrid or Barcelona, with limited travel (<10%), primarily within Europe. Required Skills: Preferred Skills: Advanced Analytics, Business Intelligence (BI), Coaching, Collaboration, Critical Thinking, Data Analysis, Database Management, Data Privacy Standards, Data Reporting, Data Savvy, Data Science, Data Visualization, Econometric Models, Process Improvements, Technical Credibility, Technologically Savvy, Workflow Analysis ## Description At J&J we are building Generative AI solutions to support pharmaceutical R&D - literature review, evidence synthesis, document Q&A, therapeutic area knowledge search, translational science workflows, and R&D decision support. These systems need to be evaluated before teams rely on them in scientific workflows. In pharma, a useful AI response depends on the question, user, source material, therapeutic area, and risk of error - so quality must be measurable, repeatable, traceable, and scientifically defensible., We need someone who can build the evaluation assets our teams count on, and continuously improve them based on what we learn. Day to day, you will: * Design, build, and maintain automated evaluation pipelines for LLM quality, RAG performance, agent reliability, safety, and scientific accuracy. * Author evaluation rubrics and scoring criteria, curate golden and synthetic datasets with domain experts, and maintain our registry of reusable evaluation assets. * Validate AI judges against human expert agreement and run model, prompt, retriever, and agent benchmarks that produce standardized quality readouts. * Analyze failure patterns - hallucination, unsupported claims, weak traceability - and turn findings into actionable recommendations. * Develop therapeutic-area-specific evaluation criteria with scientific, clinical, and regulatory partners, refining them based on real-world feedback. * Design evaluation methods for scientific reasoning, evidence synthesis, and hypothesis quality - where generic benchmarks fall short. * Build evaluation tooling and reusable patterns that enable other teams to self-serve. ## Related Videos - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Your imaginations is (no longer) the limit: how Generative AI empowers people to be creative](https://www.wearedevelopers.com/videos/741-your-imaginations-is-no-longer-the-limit-how-generative-ai-empowers-people-to-be-creative) - [Developer Experience, Platform Engineering and AI powered Apps](https://www.wearedevelopers.com/videos/990-developer-experience-platform-engineering-and-ai-powered-apps) - [The shadows that follow the AI generative models](https://www.wearedevelopers.com/videos/624-the-shadows-that-follow-the-ai-generative-models) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)