Sr. Data Expert, Data Engineer (Clinical...

Genentech
South San Francisco, CA, United States
26 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Compensation
$119,800.0 - $222,400.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Amazon Web Services Bioinformatics Catalyst (Software) Clinical Data Repository Collaborative Software Computational Biology Computer Simulation Data Sharing Data Systems Dicom
+16 more
Machine Learning Meta-Data Management Metadata Standards Cloud Services SQL Databases Parquet Data Processing Data Ingestion Fast Healthcare Interoperability Resources Git Pandas Information Technology Data Management Machine Learning Operations Data Delivery Data Pipelines

Job description

A healthier future. It’s what drives us to innovate. To continuously advance science and ensure everyone has access to the healthcare they need today and for generations to come. Creating a world where we all have more time with the people we love. That’s what makes us Roche.

Advances in AI, data and computational sciences are transforming drug discovery and development. Roche’s Research and Early Development organizations at Genentech (gRED) and Pharma (pRED) have demonstrated how these technologies accelerate R&D, leveraging data and novel computational models to drive impact. Seamless data sharing and access to models across gRED and pRED are essential to maximising these opportunities. The Computational Sciences Center of Excellence (CS CoE) is a strategic, unified group whose goal is to harness the transformative power of data and Artificial Intelligence (AI) to assist our scientists in both pRED and gRED to deliver more innovative and life-changing medicines for patients worldwide.

The Computational Sciences Center of Excellence (CS CoE) brings together data, AI, and computational expertise to accelerate innovation across gRED and pRED. Within CS CoE, the Data and Digital Catalyst (DDC) organization leads the modernization of our data ecosystem, enabling scalable, data-driven science.

The Data Capability organization within DDC is responsible for establishing foundational data capabilities, including data connectivity, data compliance, scientific content management and data ingestion, curation, integration, and delivery. The team ensures that high-quality, well-structured datasets are available to power analytics, AI/ML, and scientific discovery across Research and Early Development., We are seeking a Sr. Data Expert, Data Engineer to lead the integration and delivery of clinically anchored, multimodal scientific datasets spanning clinical, sequencing, imaging, proteomics, and other emerging data modalities.

In this role, you will:

  • Lead the integration and harmonization of clinical and multimodal scientific datasets, applying industry data standards and metadata frameworks to improve interoperability and scientific usability.

  • Own end-to-end data delivery by designing, validating, documenting, and delivering high-quality, analysis-ready datasets that support research, AI/ML, and computational biology initiatives.

  • Develop scalable data workflows that automate data ingestion, quality control, transformation, and metadata management across diverse scientific data sources.

  • Partner with computational scientists, bioinformaticians, and data engineers to understand scientific requirements and translate them into scalable, reusable data solutions.

  • Drive data quality and continuous improvement by implementing validation frameworks, metadata standards, and AI-assisted data curation practices that improve data discoverability and reuse.

  • Support emerging AI and foundation model initiatives by preparing interoperable, metadata-rich datasets optimized for downstream analytics and machine learning applications.

Requirements

  • You have a PhD with 2+ years, a Master’s degree with 3-5 years, or a Bachelor’s degree with 5+ years of experience in Bioinformatics, Data Science, Biomedical Engineering, Computer Science, Clinical Sciences, or a related discipline, with experience working with clinical, biomedical, or scientific datasets.

  • You have hands-on experience integrating clinical data with one or more scientific modalities, including sequencing, imaging, proteomics, or other omics datasets, and understand clinical data models and longitudinal patient data.

  • You are proficient in Python (Pandas), SQL, and scientific data processing, with experience working with scientific data formats such as FASTQ, BAM/CRAM, VCF, DICOM, AnnData, or Parquet, and familiarity with cloud data platforms (AWS or GCP).

  • You have experience developing or supporting automated data pipelines using workflow orchestration tools such as Airflow, Nextflow, Snakemake, or Prefect, and are comfortable using Git for collaborative software development.

  • You are a collaborative problem solver with a strong focus on data quality, metadata management, and scientific reproducibility, and enjoy partnering with multidisciplinary teams to deliver scalable data solutions.

Preferred Qualifications:

  • Experience with clinical data standards such as CDISC (SDTM/ADaM), OMOP, or FHIR.

  • Experience integrating multimodal datasets (e.g., clinical + genomics, imaging + transcriptomics, or multi-omics).

  • Familiarity with biomedical ontologies, controlled vocabularies, FAIR data principles, and metadata standards.

  • Experience preparing scientific datasets for AI/ML workflows or foundation model development.

  • Experience supporting translational research, biomarker discovery, or drug discovery programs.

Onsite presence, on our South San Francisco campus, is expected for at least 3 days a week.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:30 min

Scaling agile frameworks and data interoperability in healthcare

Leo Lindhorst · WWC 2022

3:03 min

Building an AI operating system for clinical diagnostics

Alexandre Guenoun Alexandre Guenoun +3 · WWC Europe 2026

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · WWC 2024

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

6:58 min

Analyzing production code coverage data using pandas

Markus Harrer Markus Harrer · WWC 2021

Videos

See all

Related articles

See all