> Markdown version of [/jobs/ext/576925-senior-data-engineer-ai-for-drug-discovery](https://www.wearedevelopers.com/jobs/ext/576925-senior-data-engineer-ai-for-drug-discovery). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Engineer, AI for Drug Discovery - **Company:** Genentech - **Location:** New York, NY, United States - **Experience:** Expert - **Salary:** $141,100.0 - $262,200.0 - **Contract:** Permanent contract - **Skills:** Agile Methodology, Artificial Intelligence, Amazon Web Services, Automation of Tests, Cloud Computing, Cloud Foundry, Computer Simulation, Databases, Information Engineering, Data Sharing, Experimental Data, Python (Programming Language), Search Algorithms, PostgreSQL, Machine Learning, Open Source Technology, Oracle (Applications), Software Deployment, Software Engineering, SQL Databases, Event Driven Architecture, Data Management, Docker, Databricks - **Published:** June 20, 2026 - **Apply:** https://www.juju.com/job/00000000g9newi ## About the Role + You have 7+ years of data engineering experience + You have expert knowledge of Postgres SQL and experience with Oracle + Skilled with at least one modern data toolkit (Glue, dbt, Databricks,...) + Experience with cloud platforms (preferably AWS) + Python programming skills + Strong testing practices and test automation + Understanding of CI/CD pipelines + Experience with agile development methodologies Preferred + Open source cheminformatics experience (e.g., RDKit, chemfp, Indigo, HELM toolkit) + Chemical database cartridge expertise + Familiarity with biological sequence alignment + Chemical & biological structure notation expertise + Familiarity with chemical structure canonicalization + Molecular structure searching algorithm expertise + Experience with scientific software development + Familiarity with Docker and Kubernetes + Experience with event-driven architectures + Knowledge of security best practices ## Description A healthier future. It's what drives us to innovate. To continuously advance science and ensure everyone has access to the healthcare they need today and for generations to come. Creating a world where we all have more time with the people we love. That's what makes us Roche. Advances in AI, data, and computational sciences are transforming drug discovery and development. Roche's Research and Early Development organisations at Genentech (gRED) and Pharma (pRED) have demonstrated how these technologies accelerate R&D, leveraging data and novel computational models to drive impact. Seamless data sharing and access to models across gRED and pRED are essential to maximising these opportunities. The new Computational Sciences Center of Excellence (CoE) is a strategic, unified group whose goal is to harness this transformative power of data and Artificial Intelligence (AI) to assist our scientists in both pRED and gRED to deliver more innovative and transformative medicines for patients worldwide. The Opportunity At Genentech and Roche, we're at the forefront of a revolutionary transformation in drug discovery powered by AI and machine learning. Our "lab in the loop" strategy processes massive quantities of experimental data to train AI models that accelerate the discovery of new medicines. To enable this vision, we're seeking an exceptional Senior Data Engineer to be part of the team building and maintaining our next-generation Therapeutic Molecule Registration (TMR) platform - a foundational component of our AI-driven drug discovery infrastructure, Lab-in-the-Loop ( https://www.youtube.com/watch?v=cN1PxxQWoEc ). This platform will serve as the central nervous system for managing and integrating molecular data across our global research organization, handling hundreds of billions of records and enabling unprecedented scale in virtual molecule design and testing. As the volume of AI-generated molecular designs grows exponentially, our TMR platform must evolve to become a high-performance, cloud-native system capable of supporting rapid iteration cycles between computational design and experimental validation. You will be instrumental in consolidating our molecule registration systems into a single, harmonized environment, unlocking the full potential of our data and accelerating the development of life-changing therapies. The ideal candidate has a proven record of standing up, migrating, and scaling databases with experience in chemical and/or biological registration systems. You will work on implementing scalable solutions for molecular data management and contribute to the architecture of our cloud-native platform. You will work closely with our machine learning for drug development team, Genentech Research & Early Development (gRED) Drug Discovery teams including the Antibody Engineering division, and other teams across the Roche family of companies to identify, strategize, and productionalize high-impact applications from across the drug discovery and development pipeline. Genentech provides a dynamic and challenging environment for cutting-edge, multidisciplinary research in AI and drug discovery including access to rich sources of data, close links to top academic institutions around the world, as well as internal Genentech and Roche partners and research units. In this role, you will: + Design and implement features of our TMR data model + Oversee cloud data migration to TMR and production deployment + Contribute to technical design discussions and architecture decisions ## Related Videos - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Blueprints for Success: Steering a Global Data & AI Architecture](https://www.wearedevelopers.com/videos/1577-blueprints-for-success-steering-a-global-data-ai-architecture) - [OLTP in the Lakehouse: Redefining Data for AI Workloads](https://www.wearedevelopers.com/videos/2038-oltp-in-the-lakehouse-redefining-data-for-ai-workloads) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [ Dev Digest 213: Petrol Prices, Agentic Workflows, AI Skills and CODE100!](https://www.wearedevelopers.com/magazine/718-dev-digest-213-petrol-prices-agentic-workflows-ai-skills-and-code100) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)