Data Scientist - Onshore (B)

V4C, LLC
United States
16 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Data Analysis Microsoft Azure Clinical Data Repository Data Architecture Data Validation Data Deduplication Information Engineering Data Governance Data Systems Interoperability Python (Programming Language)
+16 more
Machine Learning Power BI Standard Sql SQL Databases Tableau (Software) Feature Engineering Fast Healthcare Interoperability Resources Model Validation Data Lakes Pyspark Data Lineage Health Level Seven International Data Management Machine Learning Operations Data Pipelines Databricks

Job description

  • Develop and productionize machine learning models, statistical analyses, and predictive analytics using Python and Databricks.
  • Build and maintain scalable data science workflows using Databricks, PySpark, SQL, Delta Lake, and related cloud data technologies.
  • Work with MDM processes and frameworks to establish consistent, accurate, and trusted master data across multiple source systems.
  • Analyze and resolve data quality, duplication, matching, and entity-resolution issues across member, provider, patient, and other healthcare-related datasets.
  • Partner with Data Engineering, Product, Analytics, and business stakeholders to translate healthcare business problems into data science solutions.
  • Develop data validation, profiling, and quality-monitoring approaches to improve reliability of analytical datasets.
  • Perform exploratory data analysis and identify trends, patterns, and insights that can support member engagement and healthcare outcomes.
  • Contribute to feature engineering, model evaluation, experimentation, and deployment of data science solutions into production.
  • Ensure data solutions follow applicable healthcare data privacy, security, and governance requirements, including HIPAA where applicable.
  • Document models, datasets, assumptions, methodologies, and data lineage to support reproducibility and governance.

Requirements

  • 8+ years of experience in Data Science, Machine Learning, Advanced Analytics, or a related field.
  • Strong hands-on experience with Databricks and PySpark.
  • Advanced Python and SQL skills.
  • Experience developing and deploying machine learning or predictive models.
  • Strong understanding of Master Data Management (MDM) concepts, including:
  • Data matching and deduplication
  • Entity resolution
  • Golden/master records
  • Data standardization
  • Data quality
  • Reference/master data
  • Experience working with large-scale structured and semi-structured datasets.
  • Experience with Delta Lake / Lakehouse architecture.
  • Strong understanding of data governance, data quality, and data lineage.
  • Experience working with healthcare, payer, provider, patient, or other regulated data is preferred.
  • Experience working in a HIPAA-regulated environment is highly desirable.

Preferred Qualifications

  • Experience with healthcare member/patient data and healthcare data models.
  • Experience with MDM platforms such as Informatica MDM, Reltio, IBM MDM, or similar technologies.
  • Experience with cloud platforms such as Azure or AWS.
  • Experience with MLflow or similar model lifecycle management tools.
  • Experience with Power BI, Tableau, or other analytics/visualization platforms.
  • Experience building production-grade ML/data science pipelines.
  • Familiarity with healthcare interoperability standards such as FHIR, HL7, or claims data is a plus.

Core Skills Data Science: Python, Machine Learning, Statistics, Predictive Analytics Databricks: Databricks, PySpark, Delta Lake, MLflow Data: SQL, Data Quality, Data Governance, Data Lineage, Data Modeling MDM: Master Data Management, Entity Resolution, Matching, Deduplication, Golden Records Healthcare: Healthcare Data, HIPAA, Patient/Member Data, FHIR/HL7 Cloud: Azure/AWS

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:14 min

Executing Databricks jobs with built-in Airflow operators

Alan Mazankiewicz · LIVE

1:24 min

Moving the semantic layer upstream to avoid vendor lock-in

Piotr Menclewicz Piotr Menclewicz · Europe 2026 Virtual

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

8:27 min

Building generic custom operators for Databricks APIs

Alan Mazankiewicz · LIVE

3:48 min

Standardizing data access schemas with OData

Florian Bader Florian Bader · World Congress 2026 Europe

Videos

See all

Related articles

See all