Lead Data Scientist

SR2
London, UK
7 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
£130,000.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Bioinformatics Computational Biology Distributed Systems Python (Programming Language) Machine Learning Tensorflow Software Deployment Reinforcement Learning Pytorch Deep Learning
+4 more
Model Validation Pyspark Dask Machine Learning Operations

Job description

  • Design, train, and deploy state-of-the-art machine learning and deep learning models to predict health outcomes and disease progression.
  • Apply advanced statistical and causal inference methods (e.g. survival analysis, time-to-event modelling, propensity scoring, Mendelian randomisation).
  • Analyse and integrate multi-omics and clinical datasets to uncover novel biomarkers and risk factors.
  • Build and productionise end-to-end ML pipelines, from research to deployment.
  • Collaborate with clinicians, engineers, and product teams to translate scientific findings into scalable tools.
  • Contribute to model evaluation, explainability, and validation across diverse data sources.

Requirements

  • PhD in Machine Learning, Computational Biology, Statistics, Bioinformatics, or a related quantitative field.
  • Background in cardiovascular, cardiometabolic, or precision medicine research.
  • Proven experience developing deep learning models using Python, PyTorch, or TensorFlow.
  • Strong understanding of statistical modelling, causal reasoning, and predictive analytics.
  • Demonstrated experience working with large-scale health, genomic, or biobank datasets (e.g. UK Biobank, All of Us, Our Future Health).
  • Exposure to production deployment and model lifecycle management (MLOps awareness a plus).
  • Strong communicator with the ability to operate between science and engineering teams., * Experience integrating multi-omic or imaging data with clinical outcomes.
  • Knowledge of cloud platforms (AWS, GCP, or Azure) and distributed computing tools (PySpark, Dask, or Ray).
  • Familiarity with reinforcement learning or causal ML for adaptive interventions.

Benefits & conditions

Up to £130k + equity

We’re partnering with a pioneering health technology company using machine learning and predictive analytics to transform how cardiovascular and metabolic diseases are detected, treated, and ultimately prevented. Their mission is to use AI and data science to extend global health span by identifying individuals at risk of disease before symptoms occur.

You’ll join as the first Data Science hire, building the foundation of a platform that integrates multi-modal biomedical data, deep learning models, and large-scale population datasets such as UK Biobank and Our Future Health. This is a rare opportunity to lead model development in a setting that bridges scientific rigour with production-grade engineering., * Join a company combining scientific excellence, AI innovation, and real-world health impact.

  • Work with world-leading clinicians and researchers.
  • Shape a greenfield data science function from day one.

If you’re passionate about applying advanced machine learning to improve cardiovascular and metabolic health at population scale, we’d love to hear from you.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:19 min

Scaling performance across multiple GPUs using specialized frameworks

Paul Graham Paul Graham · World Congress 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

Videos

See all

Related articles

See all