Scientific Lead - Scientific Data Engineer

Eli Lilly and Company
San Francisco, CA, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$166,500.0 - $266,200.0
Working hours
Regular working hours
Job source

Tech stack

Training Data Artificial Intelligence Airflow Amazon Web Services Amazon S3 Business Logic Audit Trail Bioinformatics Clinical Data Repository Encodings Databases Continuous Integration
+32 more
Data Architecture Data Dictionary Information Engineering Data Governance Data Integration Extract Transform Load (ETL) Relational Databases Database Queries Python (Programming Language) Laboratory Information Management Systems Machine Learning Metadata Standards Neo4j Software Tools Cloud Services SPARQL SQL Databases Management of Software Versions Data Processing Scripting Bioconductor Large Language Models Snowflake Apache Spark Generative AI Build Management Data Lakes Information Technology Data Lineage Machine Learning Operations Data Pipelines Databricks

Job description

Data Harmonization and Lakehouse Architecture

  • Design and build the data architecture that transforms raw and processed omics data into harmonized, AI-consumable layers
  • Build and optimize ETL/ELT pipelines that produce denormalized views, pre-computed aggregations, embedding-ready text representations, and feature stores optimized for AI system consumption
  • Implement data quality monitoring, automated profiling, and validation checks across harmonization layers
  • Create versioned, reproducible data snapshots that support model training, evaluation, and audit requirements in a regulated environment
  • Partner with the teams to extend harmonization patterns as data modalities expand beyond genomics and proteomics into spatial transcriptomics, perturbational data (Perturb-Seq), single-cell, and digital pathology

Semantic Layer and Schema Engineering

  • Design and maintain a semantic layer over Lilly’s multi-omics databases that enables AI systems
  • Create comprehensive schema documentation: table descriptions, column-level annotations, relationship mappings, business logic rules, and domain-specific constraints (e.g., statistical thresholds, unit conventions, experimental design metadata)
  • Develop gold-standard question/SQL pairs for each major database, in collaboration with computational biologists and Generative AI Engineers, to serve as training data, few-shot examples, and evaluation benchmarks
  • Build and maintain a data dictionary and ontology mapping layer that translates how scientists think and speak about data (gene names, pathway terms, assay types) into how the data is physically stored

AI-Ready Data Products

  • Build and manage vector embedding pipelines for scientific documents, study metadata, and structured data descriptions to power RAG-based retrieval
  • Build integration pipelines that connect heterogeneous data sources - omics databases, internal publications, electronic lab notebooks, assay results, and clinical annotations - into a unified, queryable layer
  • Develop and enforce metadata standards that ensure new data sources are AI-accessible from the point of ingestion, not retroactively
  • Design data products that serve multiple consumption patterns: direct SQL access for computational biologists, structured feeds for ML training pipelines, and semantic interfaces for LLM-powered tools

Requirements

  • Bachelors degree in Computer Science, Data Engineering, Bioinformatics, or a related field + 8 years data engineering experience OR Masters degree and 5 years data engineering experience, * Phd in data or related field
  • Demonstrated expertise in building data pipelines, ETL/ELT workflows, and data products that serve downstream AI/ML systems
  • Strong SQL skills and experience with complex relational database schemas (hundreds of tables, multi-level joins, domain-specific conventions)
  • Experience with modern data platform technologies, including at least one of: Databricks, Snowflake, or equivalent lakehouse platforms
  • Experience with modern data engineering tools: dbt, Spark, Airflow, or similar orchestration and transformation frameworks
  • Proficiency in Python for data processing, scripting, and pipeline development
  • Experience with cloud data platforms (AWS preferred: Redshift, Athena, Glue, S3, or similar)
  • Familiarity with at least one of: vector databases, embedding pipelines, or semantic layer tooling
  • Strong communication skills - you can work effectively with both engineers who think in schemas and scientists who think in biology
  • Experience with biomedical or scientific data: omics datasets (RNA-seq, proteomics, GWAS), clinical data, or laboratory information management systems
  • Experience in pharmaceutical, biotech, or life sciences environments

  • Familiarity with biomedical ontologies and controlled vocabularies (Gene Ontology, MeSH, ChEBI, HGNC) and their application to data integration
  • Experience building data products that serve AI/ML systems - feature stores, training datasets, evaluation benchmarks, or semantic annotations for text-to-SQL
  • Knowledge of data governance practices in regulated industries: data lineage, access controls, versioning, and auditability
  • Experience with knowledge graph technologies (Neo4j, Amazon Neptune, RDF/SPARQL) or graph-based data modeling
  • Deep experience with Databricks ecosystem: Unity Catalog for data governance, Delta Lake for ACID transactions, MLflow integration, and Databricks SQL for analytics workloads
  • Experience designing data architectures that bridge traditional bioinformatics workflows (Nextflow, R/Bioconductor) with modern lakehouse consumption patterns

Benefits & conditions

Actual compensation will depend on a candidate’s education, experience, skills, and geographic location. The anticipated wage for this position is $166,500 - $266,200

Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities).Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly’s compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.

WeAreLilly

About the company

At Lilly, we unite caring with discovery to make life better for people around the world. We are a global healthcare leader headquartered in Indianapolis, Indiana. Our employees around the world work to discover and bring life-changing medicines to those who need them, improve the understanding and management of disease, and give back to our communities through philanthropy and volunteerism. We give our best effort to our work, and we put people first. We’re looking for people who are determined to make life better for people around the world.

The Opportunity

We are building something unprecedented - an AI foundation that will push the frontier on what is possible today across drug discovery research, from target identification and disease biology through translational science.

The Applied Intelligence for Discovery (AI4D) team is a newly formed group within Lilly Research Laboratories that operates at the intersection of scientific delivery and core platform development. AI4D’s mission is connecting scientists to petabyte-scale data through natural language interfaces, automated analysis workflows, and intelligent search - and to convert early deployments into repeatable system standards and evaluation practices that scale across therapeutic areas.

As a Scientific Data Engineer, you will close that gap. You will build the semantic layer, data harmonization infrastructure, AI-ready data products, and lakehouse architecture that bridge how data is stored and how AI systems need to consume it. You will be working at the intersection of the data infrastructure team and the generative AI engineers who build the systems scientists interact with., Science has been our calling from the beginning. Colonel Eli Lilly founded the company in 1876 and charged employees to “take what you find here and make it better and better.” More than 147 years later, we remain committed to his vision through every aspect of our business and the people we serve, starting with discovering the best treatments for those who take our medicines and extending to health care professionals, employees and the communities in which we live. Moreover, you can also count on the team at Lilly to be incredibly civic-minded, supporting our communities through philanthropy, volunteerism, and a creative and innovative can-do spirit.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on biospace.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:24 min

Comparing Neo4j and GraphQL conceptual models

William Lyon · LIVE

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

3:30 min

Approaching data problems with an engineering and strategy mindset

Becky Gandillon · LIVE

3:30 min

Introduction to Neo4j and remote developer relations work

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · WWC 2024

Videos

See all

Related articles

See all