> Markdown version of [/jobs/ext/1010026-scientific-lead-scientific-data-engineer](https://www.wearedevelopers.com/jobs/ext/1010026-scientific-lead-scientific-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Scientific Lead - Scientific Data Engineer - **Company:** Eli Lilly and Company - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $166,500.0 - $266,200.0 - **Contract:** Permanent contract - **Skills:** Training Data, Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Business Logic, Audit Trail, Bioinformatics, Clinical Data Repository, Encodings, Databases, Continuous Integration, Data Architecture, Data Dictionary, Information Engineering, Data Governance, Data Integration, Extract Transform Load (ETL), Relational Databases, Database Queries, Python (Programming Language), Laboratory Information Management Systems, Machine Learning, Metadata Standards, Neo4j, Software Tools, Cloud Services, SPARQL, SQL Databases, Management of Software Versions, Data Processing, Scripting, Bioconductor, Large Language Models, Snowflake, Apache Spark, Generative AI, Build Management, Data Lakes, Information Technology, Data Lineage, Machine Learning Operations, Data Pipelines, Databricks - **Published:** June 7, 2026 - **Apply:** https://www.biospace.com/job/3053638/scientific-lead-scientific-data-engineer/ ## About the Role * Bachelors degree in Computer Science, Data Engineering, Bioinformatics, or a related field + 8 years data engineering experience OR Masters degree and 5 years data engineering experience, * Phd in data or related field * Demonstrated expertise in building data pipelines, ETL/ELT workflows, and data products that serve downstream AI/ML systems * Strong SQL skills and experience with complex relational database schemas (hundreds of tables, multi-level joins, domain-specific conventions) * Experience with modern data platform technologies, including at least one of: Databricks, Snowflake, or equivalent lakehouse platforms * Experience with modern data engineering tools: dbt, Spark, Airflow, or similar orchestration and transformation frameworks * Proficiency in Python for data processing, scripting, and pipeline development * Experience with cloud data platforms (AWS preferred: Redshift, Athena, Glue, S3, or similar) * Familiarity with at least one of: vector databases, embedding pipelines, or semantic layer tooling * Strong communication skills - you can work effectively with both engineers who think in schemas and scientists who think in biology * Experience with biomedical or scientific data: omics datasets (RNA-seq, proteomics, GWAS), clinical data, or laboratory information management systems * Experience in pharmaceutical, biotech, or life sciences environments * Familiarity with biomedical ontologies and controlled vocabularies (Gene Ontology, MeSH, ChEBI, HGNC) and their application to data integration * Experience building data products that serve AI/ML systems - feature stores, training datasets, evaluation benchmarks, or semantic annotations for text-to-SQL * Knowledge of data governance practices in regulated industries: data lineage, access controls, versioning, and auditability * Experience with knowledge graph technologies (Neo4j, Amazon Neptune, RDF/SPARQL) or graph-based data modeling * Deep experience with Databricks ecosystem: Unity Catalog for data governance, Delta Lake for ACID transactions, MLflow integration, and Databricks SQL for analytics workloads * Experience designing data architectures that bridge traditional bioinformatics workflows (Nextflow, R/Bioconductor) with modern lakehouse consumption patterns ## Description Data Harmonization and Lakehouse Architecture * Design and build the data architecture that transforms raw and processed omics data into harmonized, AI-consumable layers * Build and optimize ETL/ELT pipelines that produce denormalized views, pre-computed aggregations, embedding-ready text representations, and feature stores optimized for AI system consumption * Implement data quality monitoring, automated profiling, and validation checks across harmonization layers * Create versioned, reproducible data snapshots that support model training, evaluation, and audit requirements in a regulated environment * Partner with the teams to extend harmonization patterns as data modalities expand beyond genomics and proteomics into spatial transcriptomics, perturbational data (Perturb-Seq), single-cell, and digital pathology Semantic Layer and Schema Engineering * Design and maintain a semantic layer over Lilly's multi-omics databases that enables AI systems * Create comprehensive schema documentation: table descriptions, column-level annotations, relationship mappings, business logic rules, and domain-specific constraints (e.g., statistical thresholds, unit conventions, experimental design metadata) * Develop gold-standard question/SQL pairs for each major database, in collaboration with computational biologists and Generative AI Engineers, to serve as training data, few-shot examples, and evaluation benchmarks * Build and maintain a data dictionary and ontology mapping layer that translates how scientists think and speak about data (gene names, pathway terms, assay types) into how the data is physically stored AI-Ready Data Products * Build and manage vector embedding pipelines for scientific documents, study metadata, and structured data descriptions to power RAG-based retrieval * Build integration pipelines that connect heterogeneous data sources - omics databases, internal publications, electronic lab notebooks, assay results, and clinical annotations - into a unified, queryable layer * Develop and enforce metadata standards that ensure new data sources are AI-accessible from the point of ingestion, not retroactively * Design data products that serve multiple consumption patterns: direct SQL access for computational biologists, structured feeds for ML training pipelines, and semantic interfaces for LLM-powered tools ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Cyber Sleuth: Finding Hidden Connections in Cyber Data](https://www.wearedevelopers.com/videos/893-cyber-sleuth-finding-hidden-connections-in-cyber-data) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)