> Markdown version of [/jobs/ext/1759576-data-architect-data-foundry](https://www.wearedevelopers.com/jobs/ext/1759576-data-architect-data-foundry). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Architect, Data Foundry - **Company:** Eli Lilly and Company - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $132,000.0 - $193,600.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Bioinformatics, Cloud Computing, Encodings, Computational Biology, Databases, Data Architecture, Information Engineering, Data Governance, Data Infrastructure, Data Integration, Extract Transform Load (ETL), Data Structures, Data Systems, Software Design Patterns, Graph Database, Information Sciences, Laboratory Information Management Systems, Metadata Standards, MongoDB, Neo4j, Web Ontology Language, Query Optimization, Standard Sql, Semantic Web, Software Engineering, SPARQL, Data Streaming, Data Storage Technologies, Snowflake, Apache Spark, Event Driven Architecture, Build Management, Data Lakes, Information Technology, Apache Flink, Integration Frameworks, Real Time Data, Apache Kafka, Data Management, Data Lakehouse, Data Pipelines, Databricks - **Published:** July 31, 2026 - **Apply:** https://www.biospace.com/logon?PipelinedPage=%2Fjob%2F3064840%2Fdata-architect-data-foundry%3FAction%3DContinueJobApplication%23application-form ## About the Role * B.S. or M.S. in Computer Science, Data Science, Bioinformatics, Computational Biology, Information Science, or related STEM field; Ph.D. valued for ontology and knowledge graph roles. * B.S. with 7+ years and M.S. with 5+ years of data architecture, data engineering, or scientific informatics' experience. * SQL skills and experience in multiple database paradigms (relational, graph, document, columnar, key-value). * Qualified applicants must be authorized to work in the United States on a full-time basis. Lilly will not provide support for or sponsor work authorization or visas for this role, including but not limited to F-1 CPT, F-1 OPT, F-1 STEM OPT, J-1, H-1B, TN, O-1, E-3, H-1B1, or L-1., * Expertise in at least one of: data modeling/ontologies, data platform engineering (Databricks, Snowflake, Spark), or graph/specialized databases (Neo4j, Neptune, MongoDB). * Familiarity with cloud platforms (AWS, Azure, or GCP) and modern data integration patterns. * Understanding of scientific data types and experimental workflows in life sciences or pharma (chemical, biological, HTE data). * Strong communication skills with ability to translate data architecture concepts for both technical and scientific audiences. * Pharmaceutical or biotech research industry experience, particularly in discovery data management or research informatics. * Experience with semantic web technologies: RDF, OWL, SPARQL, Protégé, or equivalent ontology engineering tools. * Hands-on experience with graph databases (Neo4j, Neptune, TigerGraph) and knowledge graph design patterns for scientific data. * Data lakehouse architecture experience: Databricks (Delta Lake, Unity Catalog), Snowflake, or equivalent; ETL/ELT with Spark, dbt. * Experience with streaming/real-time data platforms (Kafka, Kinesis, Flink) and event-driven architectures. * Familiarity with LIMS, ELN systems (e.g., Benchling), and laboratory instrument data integration. * Experience with vector databases (Pinecone, Weaviate, pgvector) and embedding-based retrieval for ML/RAG applications. * Array database experience (TileDB, Zarr) for genomics, imaging, or high-dimensional scientific data. * Experience with bioinformatics data formats (FASTA, BAM/CRAM, VCF) and biological sequence databases; familiarity with NGS data pipelines and proteomics data management. * FAIR data principles implementation experience and Data Readiness Level frameworks. * Scientific data standards and controlled vocabularies in chemistry (InChI, SMILES) or biology (Gene Ontology, UniProt, pathway databases such as Reactome or KEGG). ## Description We are seeking Data Architects at multiple levels to design and build the data infrastructure that makes AI-native drug discovery possible. You will create the schemas, ontologies, data models, knowledge graphs, and platform architectures that transform raw scientific data into machine-actionable, FAIR-compliant, insight-ready assets-serving both discovery scientists and autonomous AI agents. This role is the foundation of Architecture4Insight. Everything the software engineering team builds-pipelines, APIs, prototypes-depends on the data models and platform architecture this team designs. You will work with deep knowledge of scientific data (chemical, biological, HTE, automation-generated) to create custom-fit solutions, then partner with Tech@Lillyto scale and maintain them. The role spans three focus areas depending on expertise: data modeling & ontologies, data platform & lakehouse architecture, and knowledge graph & specialized data systems. ResponsibilitiesData Modeling & Ontologies * Design and implement data models, schemas, and ontologies for chemical, biological, and automation-generated data that serve discovery workflows across the portfolio. * Define and maintain controlled vocabularies, metadata standards, and FAIR-compliant data frameworks in partnership with Preparedness4Insight. * Implement semantic data standards (RDF, OWL, SPARQL) and ontology engineering practices to create interoperable, machine-readable scientific data. Data Platform & Lakehouse Architecture * Design and implement data lakehouse architecture using modern platforms (Databricks, Snowflake, or equivalent), including data storage patterns, partitioning strategies, and query optimization. * Build and optimize ETL/ELT pipelines using Spark, dbt, or similar tools to transform raw scientific data into analytical and ML-ready formats. * Implement real-time and streaming data integration (Kafka, Kinesis, event-driven patterns) connecting LIMS, instruments, and lab automation systems to the data infrastructure. Knowledge Graph & Specialized Data Systems * Design and implement knowledge graphs (Neo4j, Amazon Neptune, TigerGraph) that capture molecular, target, pathway, and experimental relationships across the discovery landscape. * Architect specialized data solutions: array databases (TileDB) for genomics/imaging, document stores (MongoDB) for experimental records, and vector databases for embedding-based retrieval supporting ML and RAG workflows. * Build query and traversal patterns that enable scientists and AI agents to ask relational questions across the entire data landscape. Cross-Functional Partnership * Partner with scientific software engineers to ensure data architectures are implementable, performant, and well-documented. * Collaborate with Methods4Insight to design data structures that support analytical model training, deployment, and evaluation. * Work with Tech@Lilly to define scaling strategies, ensure enterprise compliance, and transition data architectures to production-grade management. * Contribute to build-versus-buy-versus-adopt decisions by evaluating commercial and open-source data platforms against Data Foundry requirements. ## Related Videos - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Putting the Graph In GraphQL With The Neo4j GraphQL Library](https://www.wearedevelopers.com/videos/257-putting-the-graph-in-graphql-with-the-neo4j-graphql-library) - [40 Minutes to Build a Serverless COVID-19 REST and GraphQL APIs](https://www.wearedevelopers.com/videos/208-40-minutes-to-build-a-serverless-covid-19-rest-and-graphql-apis) - [Cyber Sleuth: Finding Hidden Connections in Cyber Data](https://www.wearedevelopers.com/videos/893-cyber-sleuth-finding-hidden-connections-in-cyber-data) - [Blueprints for Success: Steering a Global Data & AI Architecture](https://www.wearedevelopers.com/videos/1577-blueprints-for-success-steering-a-global-data-ai-architecture) - [NoSQL Data Modeling for Front-end Developers](https://www.wearedevelopers.com/videos/297-nosql-data-modeling-for-front-end-developers) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)