Senior Scientific Data Curator
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+13 more
Job description
The Senior Scientific Data Curator will lead the systematic discovery, assessment, harmonization, and quality assurance of Lilly’s scientific datasets across the full breadth of TuneLab’s modeling domains-small-molecule ADME/ADMET, safety and secondary pharmacology, in vivo pharmacokinetics and toxicology, antibody and biologics developability, and clinical PK/PD-in support of a strategic, cross-modality data unification initiative. This role sits at the intersection of biological and pharmacological domain expertise and data science, translating decades of fragmented, heterogeneous datasets spanning discovery through the clinic into a unified, AI-ready data infrastructure. The curator will partner closely with computational scientists, DMPK scientists, pharmacometricians, antibody engineers, and external consortium collaborators to ensure that the data substrate underpinning TuneLab’s federated AI/ML models is comprehensive, well-documented, and scientifically sound., * Conduct comprehensive inventory of historical and ongoing datasets across TuneLab’s modeling domains-small-molecule ADME/ADMET, safety and secondary pharmacology, in vivo PK and toxicology, antibody and biologics developability, and clinical PK/PD-spanning therapeutic areas (oncology, immunology, metabolic diseases, neuroscience, etc.) and 20+ years of discovery, preclinical, and clinical data
- Assess and score data quality, completeness, and integration feasibility for each dataset, accounting for the distinct data structures of each domain, including assay and dose-response measurements, concentration-time profiles and dosing regimens, in vivo study readouts, sequence- and structure-derived features for biologics, and biomarker and clinical covariate data
- Map metadata gaps across legacy systems and source platforms, documenting study contexts, assay and protocol methods, protocol deviations, data quality flags, and provenance information
- Develop automated pipelines (including LLM-assisted extraction where appropriate) to identify and extract domain-relevant data from internal documents, assay databases, and study reports into standardized, model-ready formats
- Produce a prioritized data assessment report recommending which domains, therapeutic areas, and indications to integrate first, based on data volume, complexity, portfolio relevance, and model feasibility
Data Harmonization & Integration
- Design and implement standardized, extensible schemas for the integrated multi-domain database, working with computational partners to ensure AI/ML readiness across small-molecule and biologics modalities
- Build and maintain data harmonization pipelines: label normalization, unit and assay-condition standardization across studies and sources, time-point and dose alignment, sequence and structure normalization for biologics, and covariate encoding
- Apply domain-driven quality control practices-sequence validation, hidden duplicate detection, cross-source discrepancy resolution, and cross-species dataset integration using allometric scaling where applicable
- Develop and execute outlier detection protocols, flagging and adjudicating anomalous values in collaboration with clinical pharmacologists, DMPK scientists, toxicologists, and antibody engineers as appropriate to the domain
- Create reproducible data quality assurance workflows with documented acceptance criteria and audit trails
- Curate and enrich metadata to enable cross-study and cross-domain querying-linking compound and molecule identifiers, sequence and construct identifiers, assay methods, formulation details, and study design parameters
Cross-functional Partnership
- Serve as the primary data domain expert for external consortium partners working within Lilly’s controlled cloud environment
- Collaborate with pharmacometricians, DMPK scientists, toxicologists, and antibody engineers to validate harmonized datasets against legacy models and established analyses (e.g., NONMEM/Monolix outputs for the clinical PK/PD domain)
- Work with the TuneLab ML team to ensure curated datasets meet the input specifications for the platform’s multi-task ML models, representation and foundation-model embeddings, and mechanistic/hybrid PK/PD frameworks (e.g., Neural ODE, SINDy)
- Contribute to platform deployment by supporting the development of data dictionaries, user documentation, and training materials for internal and consortium end users
Requirements
- M.S. or Ph.D. in Computational Data Science, Pharmacometrics, Pharmaceutical Sciences, Computational Biology, Cheminformatics, Biomedical Informatics, or a related quantitative discipline
- 1+ year of hands-on experience curating, harmonizing, or building analysis-ready datasets from biological, chemical, or clinical data sources
- Demonstrated skill in scientific dataset construction with domain-driven QC: sequence validation, duplicate detection, cross-source discrepancy resolution, or equivalent rigor applied to noisy real-world data
- Proficiency in Python and/or R for data wrangling, transformation, and quality checks at scale
- Working knowledge of pharmacological or chemical data structures across one or more TuneLab domains-for example, ADME/ADMET assay data, in vivo PK and toxicology readouts, antibody and biologics developability measurements, or clinical concentration-time and covariate data
- Track record of producing clear data documentation, quality reports, and data dictionaries, * Experience integrating cross-species datasets (e.g., allometric scaling) or multi-source public/internal data to expand training sets for ML models
- Familiarity with cheminformatics and computational biology tooling, including biologics-specific tools (e.g., ANARCI, protein language model embeddings, molecular operating environment software)
- Exposure to LLM-assisted workflows for information extraction, document parsing, or automated data-pipeline development
- Experience with the data conventions of one or more TuneLab domains-population PK/PD modeling tools (NONMEM, Monolix, nlmixr) or CDISC standards (SDTM, ADaM) for clinical PK/PD; ADMET/DMPK assay conventions for small molecules; or developability assays for biologics
- Familiarity with cloud-based data infrastructure (AWS, Azure, or GCP) and version-controlled, reproducible analysis environments (Git, Docker, Conda)
- Prior experience providing curated data to federated learning or collaborative ML initiatives
Benefits & conditions
Actual compensation will depend on a candidate’s education, experience, skills, and geographic location. The anticipated wage for this position is
$132,000 - $244,200
Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities).Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly’s compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.
WeAreLilly
About the company
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work-but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
What Are Large Language Models?
MLOps – What’s the deal behind it?
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production