Data Scientist, Cancer Informatics and AI/ML, Remote, Grant Funded

Northwell Health
Lake Success, NY, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$75,020.0 - $126,250.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Data Analysis Health Informatics Clinical Data Management Clinical Data Repository Cloud Computing Encodings Data Cleansing Data Security Relational Databases Document Retrieval
+25 more
Health Information Technology Python (Programming Language) Machine Learning Natural Language Processing NumPy Search Technologies SQL Databases Jupyter Notebook Data Processing Feature Engineering Retrieval-Augmented Generation Large Language Models Prompt Engineering Apache Spark Model Validation Electronic Medical Records Git Pandas Scikit Learn Information Technology Optimization Algorithms Enterprise Integration Software Version Control Data Pipelines Databricks

Job description

The Data Scientist will work semi-independently in close collaboration with clinical investigators, informatics teams, biostatisticians, and other data science stakeholders to design, build, evaluate, and refine computational pipelines. The ideal candidate will have practical prior experience developing data science workflows in Python and using modern machine learning or LLM-based tools in real projects.

Job Responsibility

  • Develop, test, and maintain Python-based data pipelines for clinical research, quality improvement, and computational oncology projects.
  • Support cancer informatics projects involving natural language processing, machine learning, large language models, and structured extraction from unstructured clinical data.
  • Build workflows for processing clinical notes, pathology reports, radiology reports, treatment records, genomics reports, and other real-world healthcare data sources.
  • Implement and evaluate LLM-assisted workflows, including prompt engineering, structured output generation, model benchmarking, validation pipelines, and error analysis.
  • Assist with the development of retrieval-augmented generation workflows, vector search, embedding-based retrieval, and related approaches where appropriate.
  • Work with clinical subject matter experts to translate oncology-focused research questions into executable data science tasks.
  • Perform data cleaning, data wrangling, exploratory analysis, feature engineering, model development, and model performance evaluation.
  • Generate reproducible analyses, reports, dashboards, tables, and visualizations to communicate findings to clinical and operational stakeholders.
  • Maintain clear documentation of code, analytic decisions, model assumptions, validation methods, and project outputs.
  • Participate in model validation efforts, including comparison of computational outputs against clinician-reviewed reference standards.
  • Contribute to manuscript, abstract, grant, and presentation development through data analysis, figure generation, and methods documentation.
  • Work independently on assigned analytic tasks while communicating progress, limitations, and blockers clearly to project leadership.

Requirements

Do you have experience in Healthcare IT?, Do you have a Master’s degree?, * Bachelor’s Degree in Computer Science, Informatics, Statistics, Engineering, Data Science, or related field, required. Master’s Degree, preferred.

  • Minimum of two (2) years of post-graduate training or experience involving quantitative data analysis, required and working with clinical data, data science, and machine learning, preferred.
  • Working familiarity with basic medical and health information technology concepts, including standardized terminologies and ontologies and electronic health records, as well as Data Warehousing and Business Intelligence tools, required.
  • Expertise in working with SQL relational databases and statistical or general programming languages (e.g., Python, R), required.
  • Deep understanding of statistical and predictive modeling concepts, machine-learning approaches, clustering and classification techniques, and recommendation and optimization algorithms.

HIGHLY PREFERRED

  • Demonstrated prior experience building or implementing applied data science, machine learning, NLP, or LLM-based workflows. Completion of a short AI certificate, bootcamp, or introductory course alone is not sufficient for this role.
  • Strong practical experience with Python for data science, including pandas, NumPy, scikit-learn, Jupyter notebooks, and reproducible analytic workflows.
  • Prior experience applying machine learning, natural language processing, or large language models to real-world data problems.
  • Experience using off-the-shelf LLMs through APIs or enterprise platforms, including structured prompting, output parsing, evaluation, and workflow integration.
  • Experience with retrieval-augmented generation, vector databases, embeddings, semantic search, or document retrieval pipelines.
  • Experience working with clinical, biomedical, or electronic health record data.
  • Familiarity with oncology data, cancer registries, pathology reports, radiology reports, genomics reports, or clinical trial data.
  • Experience working in secure data environments, enterprise data warehouses, Databricks, Spark, SQL databases, or cloud-based analytic platforms.
  • Ability to write clean, maintainable, well-documented code and use version control such as Git.
  • Demonstrated ability to work semi-independently, manage multiple analytic tasks, and communicate technical concepts to non-technical clinical collaborators.
  • Prior experience contributing to academic research, abstracts, manuscripts, grant-funded projects, or healthcare quality improvement initiatives.
  • Understanding of model evaluation concepts including accuracy, precision, recall, F1 score, calibration, error analysis, and external validation.
  • Experience with prompt engineering alone is not sufficient; candidates should have substantive prior experience in data science, machine learning, computational research, or applied analytics.

  • Additional Salary Detail

The salary range and/or hourly rate listed is a good faith determination of potential base compensation that may be offered to a successful applicant for this position at the time of this job advertisement and may be modified in the future. When determining a team member’s base salary and/or rate, several factors may be considered as applicable (e.g., location, specialty, service line, years of relevant experience, education, credentials, negotiated contracts, budget and internal equity).

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Building an AI operating system for clinical diagnostics

Alexandre Guenoun Alexandre Guenoun +3 · World Congress 2026 Europe

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

8:51 min

Addressing technical strategies and interdisciplinary computing dynamics

Noah Weber · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all