Scientific Knowledge Engineer, Ontology & Data Modeling

Xebia
Barcelona, Spain
15 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
6 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Bioinformatics Encodings Information Engineering Data Governance Serialization Graph Database JSON Python (Programming Language) Knowledge Management
+13 more
Linked Data Neo4j Web Ontology Language Cloud Services Search Technologies Semantic Web SPARQL Talend Scripting Google Cloud Large Language Models Information Technology Programming Languages

Job description

Scientific Knowledge Engineer, Ontology & Data ModelingThis role is responsible for maximizing the value of our data assets over a lifetime to bring purpose to data by acting as translators of highly technical information from domain experts into an appropriate data model - complete with significant ontology and vocabulary - that can be utilized to effectively structure and index the data.Specifically, the engineer works with Product managers and R&D subject matter expertise to define the language (data models, ontology, standards, etc.) of science into data products by acting as the voice of the “Knowledge base” and the interoperability/value of the asset.Key ResponsibilitiesDefinition of schemas/ontology and data models of scientific information required for the creation of value?adding data products.This includes accountability for the quality control and mapping specifications to be industrialized by data engineering and maintained in platform?provisioned tooling.Accountable for the quality control (through validation and verification) of mapping specifications to be industrialized by data engineering and maintained in platform?provisioned tooling - e.g., models, schemas, controlled vocab.Working with Product managers/engineers confidently converting business needs into defined deliverable business requirements to enable the integration of large?scale biology data to predict, model, and stabilize therapeutically relevant protein complex and antigen conformations for drug and vaccine discovery.Collaborate with external groups to align data standards with industry/academic ontologies ensuring that data standards are defined with usage/analytics in mind.Provide bespoke subject?matter expertise for R&D data to translate deep science into data for actionable insights.Contribute to and maintain documentation of data standards, ontology decisions, and mapping rationale to support organizational knowledge transfer and auditability.Basic QualificationsMasters degree in Bioinformatics, Biomedical Science, Biomedical Engineering, Molecular Biology, or Computer Science (with a life science application focus).6+ years of relevant work experience.Specific experience contributing to Knowledge Graph development efforts, including entity modeling, relationship design, and schema governance.Hands?on experience with open?source ontology tools and languages: Protégé, SPARQL, OWL, SKOS, SHACL, RML, RDF/Turtle.Working knowledge of major life sciences ontologies: Gene Ontology (GO), OBO Foundry ontologies (CL, UBERON, HPO, MONDO, CHEBI, EFO, CLO), MeSH, SNOMED CT, UMLS.Familiarity with linked data principles and semantic web technologies.Experience with industry?standard tools for building data serialization protocols (e.g., JSON Schema, LinkML).Proficiency in at least one programming language - preferably Python - for scripting vocabulary mappings, building data models, automating QC, and prototyping pipelines.Preferred QualificationsExperience with data governance and data quality tooling (e.g., Ataccama, Informatica, Talend, OpenRefine, Great Expectations, dbt).Experience with at least one programming language - e.g., Python - for scripting vocabulary mappings, building data models, etc.Experience supporting LLM integration or AI?readiness workflows - including metadata enrichment, entity linking, embedding pipelines, or retrieval?augmented generation (RAG) architectures.Understanding of vector databases and their role in semantic search and knowledge retrieval (e.g., Weaviate, Chroma).Familiarity with cloud data platforms and infrastructure relevant to large?scale biological data (e.g., AWS, GCP, Azure).Familiarity with graph database technologies (e.g., Neo4j, Amazon Neptune, Stardog, GraphDB, TigerGraph).Equality, Diversity, and InclusionWe welcome all individuals and evaluate solely on the quality of their work and teamwork.#J-*****-Ljbffr

Requirements

Masters degree in Bioinformatics, Biomedical Science, Biomedical Engineering, Molecular Biology, or Computer Science (with a life science application focus). 6+ years of relevant work experience. Specific experience contributing to Knowledge Graph development efforts, including entity modeling, relationship design, and schema governance. Hands?on experience with open?source ontology tools and languages: Protégé, SPARQL, OWL, SKOS, SHACL, RML, RDF/Turtle. Working knowledge of major life sciences ontologies: Gene Ontology (GO), OBO Foundry ontologies (CL, UBERON, HPO, MONDO, CHEBI, EFO, CLO), MeSH, SNOMED CT, UMLS. Familiarity with linked data principles and semantic web technologies. Experience with industry?standard tools for building data serialization protocols (e.g., JSON Schema, LinkML). Proficiency in at least one programming language - preferably Python - for scripting vocabulary mappings, building data models, automating QC, and prototyping pipelines. Preferred Qualifications Experience with data governance and data quality tooling (e.g., Ataccama, Informatica, Talend, OpenRefine, Great Expectations, dbt). Experience with at least one programming language - e.g., Python - for scripting vocabulary mappings, building data models, etc. Experience supporting LLM integration or AI?readiness workflows - including metadata enrichment, entity linking, embedding pipelines, or retrieval?augmented generation (RAG) architectures. Understanding of vector databases and their role in semantic search and knowledge retrieval (e.g., Weaviate, Chroma). Familiarity with cloud data platforms and infrastructure relevant to large?scale biological data (e.g., AWS, GCP, Azure). Familiarity with graph database technologies (e.g., Neo4j, Amazon Neptune, Stardog, GraphDB, TigerGraph).

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:38 min

Structuring enterprise data utilizing knowledge graph databases

Michael Hunger Michael Hunger · WWC 2024

2:24 min

Comparing Neo4j and GraphQL conceptual models

William Lyon · LIVE

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

3:34 min

Augmenting codebases with semantic architectures

Zaak Chalal Zaak Chalal · WWC Europe 2026

3:30 min

Introduction to Neo4j and remote developer relations work

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · WWC 2025

Videos

See all

Related articles

See all