Data Architect, Data Foundry
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+34 more
Job description
We are seeking Data Architects at multiple levels to design and build the data infrastructure that makes AI-native drug discovery possible. You will create the schemas, ontologies, data models, knowledge graphs, and platform architectures that transform raw scientific data into machine-actionable, FAIR-compliant, insight-ready assets-serving both discovery scientists and autonomous AI agents.
This role is the foundation of Architecture4Insight. Everything the software engineering team builds-pipelines, APIs, prototypes-depends on the data models and platform architecture this team designs. You will work with deep knowledge of scientific data (chemical, biological, HTE, automation-generated) to create custom-fit solutions, then partner with Tech@Lillyto scale and maintain them. The role spans three focus areas depending on expertise: data modeling & ontologies, data platform & lakehouse architecture, and knowledge graph & specialized data systems. ResponsibilitiesData Modeling & Ontologies
- Design and implement data models, schemas, and ontologies for chemical, biological, and automation-generated data that serve discovery workflows across the portfolio.
- Define and maintain controlled vocabularies, metadata standards, and FAIR-compliant data frameworks in partnership with Preparedness4Insight.
- Implement semantic data standards (RDF, OWL, SPARQL) and ontology engineering practices to create interoperable, machine-readable scientific data.
Data Platform & Lakehouse Architecture
- Design and implement data lakehouse architecture using modern platforms (Databricks, Snowflake, or equivalent), including data storage patterns, partitioning strategies, and query optimization.
- Build and optimize ETL/ELT pipelines using Spark, dbt, or similar tools to transform raw scientific data into analytical and ML-ready formats.
- Implement real-time and streaming data integration (Kafka, Kinesis, event-driven patterns) connecting LIMS, instruments, and lab automation systems to the data infrastructure.
Knowledge Graph & Specialized Data Systems
- Design and implement knowledge graphs (Neo4j, Amazon Neptune, TigerGraph) that capture molecular, target, pathway, and experimental relationships across the discovery landscape.
- Architect specialized data solutions: array databases (TileDB) for genomics/imaging, document stores (MongoDB) for experimental records, and vector databases for embedding-based retrieval supporting ML and RAG workflows.
- Build query and traversal patterns that enable scientists and AI agents to ask relational questions across the entire data landscape.
Cross-Functional Partnership
- Partner with scientific software engineers to ensure data architectures are implementable, performant, and well-documented.
- Collaborate with Methods4Insight to design data structures that support analytical model training, deployment, and evaluation.
- Work with Tech@Lilly to define scaling strategies, ensure enterprise compliance, and transition data architectures to production-grade management.
- Contribute to build-versus-buy-versus-adopt decisions by evaluating commercial and open-source data platforms against Data Foundry requirements.
Requirements
- B.S. or M.S. in Computer Science, Data Science, Bioinformatics, Computational Biology, Information Science, or related STEM field; Ph.D. valued for ontology and knowledge graph roles.
- B.S. with 7+ years and M.S. with 5+ years of data architecture, data engineering, or scientific informaticsâ experience.
- SQL skills and experience in multiple database paradigms (relational, graph, document, columnar, key-value).
- Qualified applicants must be authorized to work in the United States on a full-time basis. Lilly will not provide support for or sponsor work authorization or visas for this role, including but not limited to F-1 CPT, F-1 OPT, F-1 STEM OPT, J-1, H-1B, TN, O-1, E-3, H-1B1, or L-1., * Expertise in at least one of: data modeling/ontologies, data platform engineering (Databricks, Snowflake, Spark), or graph/specialized databases (Neo4j, Neptune, MongoDB).
- Familiarity with cloud platforms (AWS, Azure, or GCP) and modern data integration patterns.
- Understanding of scientific data types and experimental workflows in life sciences or pharma (chemical, biological, HTE data).
- Strong communication skills with ability to translate data architecture concepts for both technical and scientific audiences.
- Pharmaceutical or biotech research industry experience, particularly in discovery data management or research informatics.
- Experience with semantic web technologies: RDF, OWL, SPARQL, ProtĂŠgĂŠ, or equivalent ontology engineering tools.
- Hands-on experience with graph databases (Neo4j, Neptune, TigerGraph) and knowledge graph design patterns for scientific data.
- Data lakehouse architecture experience: Databricks (Delta Lake, Unity Catalog), Snowflake, or equivalent; ETL/ELT with Spark, dbt.
- Experience with streaming/real-time data platforms (Kafka, Kinesis, Flink) and event-driven architectures.
- Familiarity with LIMS, ELN systems (e.g., Benchling), and laboratory instrument data integration.
- Experience with vector databases (Pinecone, Weaviate, pgvector) and embedding-based retrieval for ML/RAG applications.
- Array database experience (TileDB, Zarr) for genomics, imaging, or high-dimensional scientific data.
- Experience with bioinformatics data formats (FASTA, BAM/CRAM, VCF) and biological sequence databases; familiarity with NGS data pipelines and proteomics data management.
- FAIR data principles implementation experience and Data Readiness Level frameworks.
- Scientific data standards and controlled vocabularies in chemistry (InChI, SMILES) or biology (Gene Ontology, UniProt, pathway databases such as Reactome or KEGG).
Benefits & conditions
Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include: Africa, Middle East, Central Asia (AMECA), Black Employees at Lilly (BE@Lilly), Chinese Culture Network (CCN), EnAble, Evolve, Lilly Indian Network (LIN), Organization of Latinx at Lilly (OLA), Pride (LGBTQ+ Allies), Veterans Leadership Network (VLN) and Womenâs Initiative for Leading at Lilly (WILL).
Actual compensation will depend on a candidateâs education, experience, skills, and geographic location. The anticipated wage for this position is $132,000 - $193,600
Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities).Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lillyâs compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.
WeAreLilly
About the company
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work-but itâs work worth doing. If youâre driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us., Science has been our calling from the beginning. Colonel Eli Lilly founded the company in 1876 and charged employees to âtake what you find here and make it better and better.â More than 147 years later, we remain committed to his vision through every aspect of our business and the people we serve, starting with discovering the best treatments for those who take our medicines and extending to health care professionals, employees and the communities in which we live. Moreover, you can also count on the team at Lilly to be incredibly civic-minded, supporting our communities through philanthropy, volunteerism, and a creative and innovative can-do spirit.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.biospace.comGood distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?
Highest Paying Tech Companies for Developers
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
Data Engineer Salary UK