Senior Data Scientist
Spirite Industries, Inc.
United States
3 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on arc.dev
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Languages
English
Job source
Tech stack
Big Data
BigQuery
Continuous Integration
Data Deduplication
Extract Transform Load (ETL)
Python (Programming Language)
Machine Learning
Neo4j
SQL Databases
Snowflake
Git
Fastapi
+3 more
Pytest
Docker
Databricks
Job description
- Extracting entities and attributes from heterogeneous sources including warehouse tables, registries, and documents
- Designing and implementing probabilistic entity linking and deduplication across multiple large healthcare datasets
- Owning the linking pipeline from end to end, including pairwise and graph-based scoring and calibrated match probabilities
- Defining how link quality is measured through precision, recall, calibration, and error analysis
- Linking entities and attributes to controlled medical vocabularies to ensure consistency across sources
- Utilizing graph representations and graph ML techniques for linking, including clustering, centrality, and link prediction
- Building and maintaining the data and graph estate, including warehouse transformation and monitoring
- Improving the performance of existing algorithms in response to changing data sources and volumes
Requirements
- Minimum 5 years of experience in shipping production ML or data-science work, including hands-on entity linking and resolution at scale
- Deep expertise in Python, including production-grade, tested, and type-hinted code
- Strong command of SQL and experience with FastAPI, Docker, pytest, Git, and CI/CD practices
- Practical experience in probabilistic methods for match scoring, threshold selection, and evaluation
- Comfort with graph algorithms and graph ML techniques, including embeddings and GNNs
- Experience with data and ML engineering, including ETL processes over large datasets and data-quality monitoring
- Familiarity with distributed SQL warehouses such as Snowflake, BigQuery, or Databricks, and graph technologies like Neo4j or Amazon Neptune
- Advanced level of English
Benefits & conditions
Benefits For You
- Great Place to Work
- Solid financial situation
- Contracts with the biggest brands
- Centre of internal trainings
- Many experts you can learn from
- Open and accessible management team
- Profit sharing
- Passion Sponsorship program
- Regular integration events and trips
- Comfortable and well-equipped offices
- MySii app
- Medical care
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on arc.dev
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
EM
Eli McGarvie
over 3 years ago
EM
Eli McGarvie
Data Engineer Salary UK
over 3 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
over 2 years ago
MH
Michael Hunger
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?
8 months ago
DS
Dhannush Subramani
Top Big Data Technologies That You Need to Know
about 4 years ago
CH
Chris Heilmann
Dev Digest 121 - AI goes offline
over 2 years ago