Senior Data Scientist

Spirite Industries, Inc.
United States
3 days ago
Apply on arc.dev
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Big Data BigQuery Continuous Integration Data Deduplication Extract Transform Load (ETL) Python (Programming Language) Machine Learning Neo4j SQL Databases Snowflake Git Fastapi
+3 more
Pytest Docker Databricks

Job description

  • Extracting entities and attributes from heterogeneous sources including warehouse tables, registries, and documents
  • Designing and implementing probabilistic entity linking and deduplication across multiple large healthcare datasets
  • Owning the linking pipeline from end to end, including pairwise and graph-based scoring and calibrated match probabilities
  • Defining how link quality is measured through precision, recall, calibration, and error analysis
  • Linking entities and attributes to controlled medical vocabularies to ensure consistency across sources
  • Utilizing graph representations and graph ML techniques for linking, including clustering, centrality, and link prediction
  • Building and maintaining the data and graph estate, including warehouse transformation and monitoring
  • Improving the performance of existing algorithms in response to changing data sources and volumes

Requirements

  • Minimum 5 years of experience in shipping production ML or data-science work, including hands-on entity linking and resolution at scale
  • Deep expertise in Python, including production-grade, tested, and type-hinted code
  • Strong command of SQL and experience with FastAPI, Docker, pytest, Git, and CI/CD practices
  • Practical experience in probabilistic methods for match scoring, threshold selection, and evaluation
  • Comfort with graph algorithms and graph ML techniques, including embeddings and GNNs
  • Experience with data and ML engineering, including ETL processes over large datasets and data-quality monitoring
  • Familiarity with distributed SQL warehouses such as Snowflake, BigQuery, or Databricks, and graph technologies like Neo4j or Amazon Neptune
  • Advanced level of English

Benefits & conditions

Benefits For You

  • Great Place to Work
  • Solid financial situation
  • Contracts with the biggest brands
  • Centre of internal trainings
  • Many experts you can learn from
  • Open and accessible management team
  • Profit sharing
  • Passion Sponsorship program
  • Regular integration events and trips
  • Comfortable and well-equipped offices
  • MySii app
  • Medical care

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:24 min

Comparing Neo4j and GraphQL conceptual models

William Lyon · LIVE

3:05 min

Tagging and organizing execution scenarios with pytest markers

Florian Bruhin · World Congress 2021

1:34 min

Bringing diverse skills to industrial data science roles

Katja Träumner

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all