> Markdown version of [/jobs/ext/3038407-healthcare-data-scientist](https://www.wearedevelopers.com/jobs/ext/3038407-healthcare-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # healthcare data scientist - **Company:** Dandelion Health Inc. - **Location:** United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Cerner, Adaptable Database Systems, Artificial Intelligence, Amazon Web Services, Data Analysis, Health Informatics, Computer Programming, Databases, Data Integrity, Data Transformation, Data Mining, Data Structures, Dicom, Python (Programming Language), Natural Language Processing, Tensorflow, SQL Databases, Pytorch, Large Language Models, Deep Learning, Model Validation, Electronic Medical Records, Git, Semi-structured Data, Information Technology, HuggingFace, Data Pipelines, Allscripts - **Published:** September 23, 2026 - **Apply:** https://startup.jobs/nlp-llm-data-scientist-dandelion-health-inc-10154088 ## About the Role * Advanced degree in a quantitative field (ex. Data Science, Biomedical Informatics, Computer Science, Biostatistics), or B.S. with at least 5 years of professional experience * At least 2 years of data science and machine learning experience, including building pipelines to extract and curate unstructured and semi-structured data by applying advanced machine learning and AI techniques. Prior experience with clinical and healthcare data is a strong bonus. * Fluency in Python and SQL, including fluency with ML/NLP libraries (PyTorch, Tensorflow, HuggingFace, etc.) * Familiarity with using modern applied LLM techniques on real-world data * Strong technical writing, editing, and communication skills, along with a collaborative mindset * Excellent organizational skills with an ability to embrace change and effectively manage multiple projects and consistently plan work to meet deadlines * Experience working in or with startups is a plus Technical Experiences and Skills We don't expect anyone to have all of the following skills or experiences, but we do seek candidates who are interested in growing their skill sets and working with healthcare data in all its glorious complexity. The Data Team works closely with our Engineering Team to put our work into production and meet client needs. * Git and version control * Familiarity with encryption methods * Prior experience querying EDWs or databases and creating reports or analytics for healthcare data * Familiarity with the data aspects of electronic medical records, ex. Epic, Cerner, Allscripts * Any medical ontology experience * Any experience working with DICOM or other imaging modalities * Experience with AWS * Experience with publishing work in peer-reviewed journals ## Description You are a healthcare data scientist who knows your way around clinical and text-based data. Your primary responsibility will be to leverage and build upon large language models (LLM) and other ML-based approaches for meaningful data abstraction from both unstructured and structured healthcare data. You will join a team of data scientists who own the creation and maintenance of AI-ready datasets for clients in our environment. You will work across the organization to curate datasets by identifying patient subpopulations or disease cohorts, developing methodology to abstract information from multimodal healthcare data sources to support patient phenotyping, and pooling data to enable rapid exploratory AI/ML analyses, model experimentation and/or model validation. You will use your data expertise, programming abilities, and critical thinking skills to support our technical product team, and develop your own analyses to derive insights and enhance our datasets based on the use case. Your team's ultimate goal is to deliver the highest-quality data possible to our clients, who are building products that improve patient health. You will report to the Data Science Manager, under the Head of Data. Responsibilities Your day-to-day responsibilities will include the following: * Develop Natural Language Processing (NLP), Large Language Model (LLM) and other ML-based pipelines to abstract relevant labels from text-based healthcare data and store them in scalable data models; * Query complex source systems in a range of health data sources (e.g., EMRs, semi-structured reports, free-text clinical provider notes) to identify key data elements and create and enrich high-quality datasets for real-world evidence analyses and training AI algorithms; * Own data extraction, wrangling, labeling and QC tasks to create analytical datasets that include abstracted clinical concepts and provide a range of solutions to support customers' AI activities; * Stay current on the latest in applied NLP and generative AI methods and proactively leverage these technologies where applicable; * Support the design, testing, validation, analysis, and merging of multimodal data structures from a wide variety of source systems; * Develop code and documentation to deliver high-quality and HIPAA-compliant data products on time to customers; * Identify and resolve problems using your knowledge, background, and troubleshooting skills; * Ensure accuracy, data integrity, and validity of data and analysis in all work; * Provide support for technical product team to advance development of the suite of data-related product offerings; * Summarize the complexity of abstraction methods, findings and recommendations into clear explanations and presentations for internal and external audiences that have a varying range of technical and clinical experience; You are not afraid to dig into massive, confusing, disorganized new datasets and get them under control. You are excited to learn new environments, languages, and skills. This is a small, early stage company with enormous ambitions and everyone pitches in across the team. ## Related Videos - [ZEISS & Microsoft - Building the Next Generation Medical Ecosystem in the Cloud](https://www.wearedevelopers.com/videos/424-zeiss-microsoft-building-the-next-generation-medical-ecosystem-in-the-cloud) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Data Science, ML & AI in the Oil and Gas Industry at NDT Global - Dr. Katja Träumner](https://www.wearedevelopers.com/videos/1308-data-science-ml-ai-in-the-oil-and-gas-industry-at-ndt-global-dr-katja-traumner) - [Stop Committing Your Secrets - GIt Hooks To The Rescue!](https://www.wearedevelopers.com/videos/573-stop-committing-your-secrets-git-hooks-to-the-rescue) ## Related Articles - [The Biggest German Tech Companies](https://www.wearedevelopers.com/magazine/424-the-biggest-german-tech-companies) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)