> Markdown version of [/jobs/ext/2069591-data-scientist-ii](https://www.wearedevelopers.com/jobs/ext/2069591-data-scientist-ii). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist II - **Company:** Exel Inc. - **Location:** Rockville, MD, United States (Remote available) - **Salary:** $130,000.0 - $145,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Business Analytics Applications, Data Analysis, Bioinformatics, Health Informatics, Clinical Data Repository, Cloud Computing, Cloud Database, Cluster Analysis, Profiling, Databases, Data Governance, Data Infrastructure, Data Integration, Extract Transform Load (ETL), Data Mapping, Data Security, Relational Databases, Statistical Hypothesis Testing, Python (Programming Language), Machine Learning, Meta-Data Management, Open Source Technology, Scientific Computating, Unstructured Data, Web Applications, Scripting, High Performance Computing, Fast Healthcare Interoperability Resources, Large Language Models, Snowflake, Jupyter, Data Lakes, Information Technology, Data Lineage, Health Level Seven International, Slurm, Functional Programming, Streamlit Framework, Data Pipelines, Api Management, Databricks - **Published:** August 15, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/ofagzfomax ## About the Role * Education & Background: Bachelor's degree in Data Science, Bioinformatics, Computer Science, Biological Sciences, or a related field (advanced degree preferred), or equivalent experience. Demonstrated experience in a data-intensive role supporting biomedical research or scientific computing. * Data Science and Bioinformatics Expertise: Strong proficiency in Python and R for analysis, scripting, and visualization. Hands-on experience with at least two omics data types (e.g., bulk RNA-seq, scRNA-seq, spatial transcriptomics, proteomics, metagenomics, GWAS). * Analytical Skills: Solid understanding of statistical modeling, dimensionality reduction, clustering, differential expression, and pathway analysis. Ability to work with structured, semi-structured, and unstructured data across relational and data lake environments. * Collaboration & Communication: Strong problem-solving skills with the ability to communicate effectively across technical and non-technical audiences. Able to translate scientific needs into technical solutions and clearly articulate risks, assumptions, and limitations. * Domain Alignment: Genuine interest in biomedical and translational research. Ability to quickly learn domain-specific terminology and workflows, with awareness of data governance, privacy, and compliance requirements for clinical and research data., * Data Platform Experience: Experience building analytics solutions in platforms such as Snowflake, Databricks, or cloud data warehouses, with integrations across databases, APIs, dashboards, and application environments. * Bioinformatics Workflow Tooling: Experience with workflow and reproducibility tools used in Galaxy, Terra, Nextflow/WDL, Snakemake, Singularity, or CWL. Familiarity with the scverse Python ecosystem (Scanpy, Squidpy, SCIMAP, AnnData) and spatial single-cell analysis methods, including PhenoGraph, Louvain/Leiden clustering, UMAP, and Ripley's L statistic, is a plus. * Research and Application Enablement: Experience preparing curated datasets for dashboards, APIs, and web applications. Familiarity with Posit Connect, R/Shiny, Streamlit, Jupyter, or similar platforms is a plus. * Cloud, HPC, Storage, and Automation: Experience with AWS (EC2, S3, Lambda), object storage, relational databases, scheduled jobs, API integrations, and secure data movement. Familiarity with HPC environments, SLURM/SGE, or NIH Biowulf is preferred. * Biomedical Domain Knowledge: Background in biomedical research, clinical research, or healthcare analytics. Familiarity with standards such as HL7/FHIR, CDISC, or OMOP, and experience with clinical, genomic, or biospecimen data is a plus. * Governance and Reproducibility: Experience with metadata management, data lineage, open-source code release, containerized analyses, and secure handling of de-identified or access-controlled research datasets. * Training and Scientific Enablement: Experience creating documentation, training materials, or workshops for researchers and non-coder audiences. Ability to support tool adoption and explain workflows and results clearly is strongly preferred. Disclaimer: The above description is meant to illustrate the general nature of work and level of effort being performed by individuals assigned to this position or job description. This is not restricted as a complete list of all skills, responsibilities, duties, and/or assignments required. Individuals may be required to perform duties outside of their position, job description or responsibilities as needed. ## Description We are seeking a Data Scientist II to join our vibrant team supporting the National Cancer Institute (NCI) at the NIH in Rockville, MD. This role is embedded within NCI's Center for Biomedical Informatics and Information Technology (CBIIT), where you will directly advance cancer research by building the computational infrastructure that scientists depend on every day. You will support the full omics data lifecycle across a broad spectrum of modalities, including bulk RNA-seq, single-cell RNA-seq (scRNA-seq), spatial transcriptomics, Digital Spatial Profiling (DSP), whole genome and exome sequencing (WGS/WES), metagenomics, metabolomics, and proteomics, as well as clinical, imaging, and biospecimen data. A core part of this role involves developing workflows that integrate these modalities to support systems-level biological questions, cross-cohort studies, and NCI CBIIT initiatives. You will collaborate closely with NCI scientists, bioinformaticians, clinician-researchers, data engineers, software developers, and government stakeholders to ensure analytical infrastructure is FAIR-compliant, containerized, version-controlled, well-documented, and purpose-built for long-term reuse across the research community., * Bioinformatics Workflow and Data Pipeline Development: Design, build, and maintain reproducible pipelines for diverse biomedical data types - including genomic, transcriptomic, single-cell, spatial, proteomic, metagenomic, metabolomic, and clinical datasets. Develop reusable transformation logic and curated datasets supporting analytics, dashboards, APIs, notebooks, and downstream research workflows. * Multi-Omics Analysis: Support NCI CBIIT labs in their analysis workflows including bulk RNA-seq (QC, DEG, GSEA), single-cell RNA-seq (clustering, UMAP/t-SNE, cell type annotation, DEG), and Digital Spatial Profiling (annotation, QC, normalization, spatial deconvolution, volcano plots, heatmaps). * Data Integration and Lifecycle Support: Enable reliable data movement from source systems into structured, analysis-ready formats. Support ingestion, curation, metadata capture, source-to-target mapping, schema management, provenance tracking, and long-term maintainability of data products. * Statistical Modeling and Machine Learning: Apply statistical and ML methods - including hypothesis testing, regression, clustering, PCA, UMAP, t-SNE, and classification - to biomedical datasets. Incorporate AI/LLM-based extraction where appropriate, with clear validation and communication to stakeholders. * Researcher-Facing Applications and Visualization: Build and support interactive dashboards (Shiny, Streamlit), notebooks, reports, and APIs enabling researchers to explore multi-omics and clinical data. Support figure generation for QC, differential expression, pathway, and spatial analyses. * Collaboration: Partner with data scientists, bioinformaticians, researchers, developers, and government stakeholders to translate scientific needs into technical specifications, data models, and reusable workflows that accelerate biomedical research. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Kubernetes dev is fun, but setup and ops isn't! See a fun PaaS alternative to push any code, ipynbs or even just data!](https://www.wearedevelopers.com/videos/732-kubernetes-dev-is-fun-but-setup-and-ops-isn-t-see-a-fun-paas-alternative-to-push-any-code-ipynbs-or-even-just-data) - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [Hacking AI at the Edge of the Indian Ocean](https://www.wearedevelopers.com/videos/100177-hacking-ai-at-the-edge-of-the-indian-ocean) - [Blueprints for Success: Steering a Global Data & AI Architecture](https://www.wearedevelopers.com/videos/1577-blueprints-for-success-steering-a-global-data-ai-architecture) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)