Sr. Data Delivery Specialist (Public and purchased data collection)

Genentech
San Francisco, CA, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
0 years minimum
Compensation
$127,800.0 - $237,300.0
Working hours
Regular working hours
Job source

Tech stack

Microsoft Access Artificial Intelligence Amazon S3 Microsoft Azure Bash Shell Big Data Bioinformatics Health Informatics Catalyst (Software) Clinical Data Repository Computer Simulation Data Cleansing
+22 more
Information Engineering Data Governance Data Integration Data Security Data Sharing JSON Python (Programming Language) Metadata Standards DataOps SQL Databases Workflow Management Systems Jupyter Notebook Parquet Data Processing Cloud Platform System Data Ingestion Usage Tracking Pandas Information Technology Data Management Machine Learning Operations Data Delivery

Job description

A healthier future. It’s what drives us to innovate. To continuously advance science and ensure everyone has access to the healthcare they need today and for generations to come. Creating a world where we all have more time with the people we love. That’s what makes us Roche.

Advances in AI, data, and computational sciences are transforming drug discovery and development. Roche’s Research and Early Development organisations at Genentech (gRED) and Pharma (pRED) have demonstrated how these technologies accelerate R&D, leveraging data and novel computational models to drive impact. Seamless data sharing and access to models across gRED and pRED are essential to maximising these opportunities. The new computational sciences Center of Excellence (CoE) is a strategic, unified group whose goal is to harness this transformative power of data and Artificial Intelligence (AI) to assist our scientists in both pRED and gRED to deliver more innovative and transformative medicines for patients worldwide.

The Computational Sciences Center of Excellence (CS CoE) brings together data, AI, and computational expertise to accelerate innovation across gRED and pRED. Within CS CoE, the Data and Digital Catalyst (DDC) organization leads the modernization of our data ecosystem, enabling scalable, data-driven science.

The Data Capability organization within DDC is responsible for establishing foundational data capabilities, including data connectivity, data compliance, scientific content management and data ingestion, curation, integration, and delivery. The team ensures that high-quality, well-structured datasets are available to power analytics, AI/ML, and scientific discovery across Research and Early Development. THE OPPORTUNITY:

We are seeking an Associate Data Delivery Specialist to support the delivery and operationalization of real-world data (RWD) and clinical-genomic datasets sourced from external partnerships and public/purchased data collections.

In this entry-level role, you will contribute to the coordination, preparation, and delivery of multimodal, high-dimensional datasets, ensuring they are accessible, well-documented, and ready for use in research, analytics, and AI/ML workflows. You will also support interactions with external data providers and internal stakeholders to ensure efficient and compliant data usage.

You will work within a cross-functional environment spanning data engineering, data science, and research teams, helping to enable data-driven discovery across Roche’s R&D ecosystem.

  • RWD Data Operations & Delivery Support intake, tracking, and fulfillment of real-world data requests, including clinical-genomic and multimodal datasets. Assist in preparing datasets for delivery, ensuring completeness, quality, and documentation.

  • External Data Coordination Coordinate with external partners (e.g., Caris, FMI) to support data requests, query submissions, and data returns. Assist in managing communications, timelines, and deliverables.

  • Data Governance & Access Support Assist in managing data access workflows, ensuring appropriate approvals, training, and compliance with data usage agreements. Track data usage and maintain documentation.

  • High-Dimensional Data Handling Work with sequencing, imaging, and proteomics datasets, supporting standardized formatting, validation, and integration readiness. Contribute to handling emerging multimodal data types and evolving standards.

  • Data Delivery & Quality Control Perform quality checks, metadata validation, and documentation to ensure datasets are analysis-ready. Support troubleshooting of data delivery issues and escalate when necessary.

  • AI-Assisted Data Curation Support Contribute to early-stage efforts in AI-enabled data curation and harmonization, supporting improved scalability and efficiency in data delivery workflows.

  • Collaboration Across Teams Partner with internal teams (e.g., AIBT, CBM, gRED TM, pRED DTAs) to support data integration and delivery needs across diverse scientific use cases.

Requirements

Do you have experience in Technical troubleshooting support?, Do you have a Master’s degree?, * PhD and 0-2 years of experience, Master’s degree and 3-5 years of experience or a Bachelor’s degree and 4-7 years of experience in Data Science, Bioinformatics, Health Informatics, Biomedical Engineering, Computer Science, or a related field and experience working with real-world data, clinical data, or biomedical datasets

  • Basic understanding of RWD sources (e.g., EHR, claims, registries, clinical-genomic datasets)
  • Strong attention to detail and commitment to data quality and reliability
  • Strong organizational and communication skills, with the ability to support multiple stakeholders
  • You are someone who has the technical skills for: Programming: Python (Pandas) or SQL; familiarity with Bash is a plus. Data Formats: Experience with structured data (CSV, JSON, Parquet); exposure to scientific formats is a plus. Data Platforms: Exposure to cloud environments (AWS S3, GCS, or Azure). Tools: Familiarity with Jupyter notebooks, data portals, or workflow tools is beneficial

Preferred Qualifications:

  • Exposure to clinical-genomic or multimodal datasets (e.g., Caris, FMI, or similar)
  • Familiarity with data governance and compliance in healthcare or life sciences
  • Exposure to AI/ML workflows or data preparation for analytics
  • Understanding of FAIR data principles and metadata standards
  • Interest in working with external data partnerships and large-scale data ecosystems

Onsite presence, on our South San Francisco campus, is expected for at least 3 days a week.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · WWC 2024

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

3:33 min

Refactoring data science workflows using Rapids QDF and Pandas

Paul Graham Paul Graham · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all