Remote Senior Scientific Data Engineer
Role details
Job location
Tech stack
Job description
The Senior Data Engineer will design, build, and deliver a new enterprise data product supporting the clients generative drug design and computational chemistry platforms. This role focuses on creating scalable, well-structured data architecture from the ground up, with long-term expansion and downstream AI/ML integration in mind. The ideal candidate combines strong data engineering expertise with an understanding of drug design, chemistry, and scientific data workflows.
-Design and implement a new enterprise data product, initially scoped as a standalone deliverable with future integration into broader AI-driven drug discovery platforms.
-Build scalable data pipelines, schemas, and storage models capable of supporting large, complex scientific and chemistry-derived datasets.
-Develop data solutions primarily on GCP / BigQuery, adhering to enterprise data engineering templates and standards.
-Implement data transformations and pipelines using Python, with a focus on data quality, traceability, and performance.
-Ensure the data architecture supports future expansion, additional datasets, and evolving analytical and computational needs.
-Collaborate closely with computational chemists, data scientists, and ML engineers to ensure data models align with generative design, molecular representations, and ML outputs.
-Apply an understanding of drug design and chemistry concepts (e.g., molecular properties, structure-activity data, experimental outputs) to inform data modeling and integration decisions.
-Provide technical guidance on data structure, scalability, and long-term maintainability in an enterprise environment.
Requirements
The Data Engineer will take end-to-end ownership of a new cloud-native data product on Google Cloud Platform, leveraging established in-house templates and standards to deliver a robust, scalable, and production-grade solution. The role involves designing and operating reliable ingestion pipelines for external data sources, integrating and harmonising key fields with parallel external data products, and delivering curated, analytics-ready datasets that are readable, updatable, and trusted by downstream users. The engineer will apply strong expertise in Python, SQL, columnar data warehouses (e.g. BigQuery), schema and data-model design, pipeline orchestration, and query optimisation, while embedding best practices around data quality, testing, metadata, and documentation. Operating as part of a cross-functional environment, the role requires a solid understanding of core cloud concepts, CI/CD, version control, and data reliability in production. Experience with scientific or R&D data-particularly chemistry or related life-science domains-would be highly advantageous, enabling effective standardisation, interpretation, and integration of complex domain-specific datasets., Chemistry, scientific, or life sciences educational background.
- Experience working with scientific and chemical datasets/ computational chemistry data
-Strong experience in data engineering, including database, schema, and data product design.
-Hands-on experience with GCP and BigQuery (Postgres familiarity a plus).
-Experience with ChemAxon, RDKit or Pipeline Pilot
-Proficiency in Python for building and maintaining data pipelines.
-CI/CD
-Experience working with large, complex datasets at scale, ideally in scientific or R&D contexts.
-Background in life sciences, pharma, or scientific data platforms. -Database Design
-Experience supporting downstream analytics, ML pipelines, or AI-driven platforms, particularly in R&D or discovery environments.