Senior AI Data Scientist

Knowtion Health
United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Information Leak Prevention Data Stores Python (Programming Language) NumPy Standard Sql SQL Stored Procedures Systems Integration Management of Software Versions EHR Systems System Availability Deep Learning
+6 more
Git Pandas Scikit Learn Information Technology Machine Learning Operations Data Pipelines

Job description

Knowtion Health is seeking a Senior AI Data Scientist who owns systems end to end - framing the problem, building and validating the models, and standing up the data pipelines and monitoring that keep them running in production. The role is data science and engineering in equal measure: the statistical rigor to build a model that is sound, and the engineering craft to ship and operate it. Because our models touch payments and patients, they are held to a high bar: statistically defensible, operationally reliable, and defensible to an auditor. This is a senior individual-contributor role for someone who has both built models and kept them alive in production.

What’s Attractive to the Right Candidate?

  • Knowtion Health is a growing firm in a growing industry. Our status as a leader in this industry means that we have the resources to invest in the business and to innovate.
  • Our business is intensely competitive and is constantly evolving. We quickly identify new challenges and develop solutions, so you won’t simply be doing what was done last year. Our new employees are frequently pleased and surprised by how quickly we make decisions and adapt to market conditions.
  • Knowtion Health culture is inviting and competitive, embracing challenge and celebrating accomplishment; dedicated colleagues striving to provide quality results that have lasting impact.

The Opportunity:

  • Frame ambiguous revenue-cycle problems as tractable modeling problems and own them end to end - such as claim and invoice viability scoring, denial and underpayment prediction, revenue/ARR prediction, and work prioritization - from hypothesis to a validated, monitored result in production.
  • Build and validate models with statistical rigor: thoughtful feature design, sound handling of missingness and class imbalance, calibration, honest evaluation against simple baselines, and a clear treatment of uncertainty.
  • Design, build, and maintain the data pipelines feeding these models - integrating with source collections, billing, and EHR systems and internal data stores, and handling data collection, cleansing, field mapping, and payer/plan normalization - with a bias toward reliability and data quality.
  • Establish and own MLOps practice: model monitoring and drift detection (e.g., population stability), calibration and threshold selection, benchmarking against simple baselines, explainability (e.g., SHAP and interpretable coefficients), experiment and version tracking, and model documentation such as model cards.
  • Partner with subject-matter experts to encode business-rule and heuristic logic alongside statistical models where that produces a more accurate or more defensible result.
  • Gather and document business requirements for revenue-cycle initiatives, including workflows, decision points, stakeholder objectives, operational constraints, success criteria, data availability, and underlying business assumptions, and maintain traceability as those requirements evolve.

Requirements

  • Several years building and validating models and operating them in production, with clear ownership of at least one model that a product or business function depended on.
  • A degree in a quantitative field such as statistics, computer science, mathematics, or data science, or equivalent demonstrated ability.
  • A strong statistical foundation - experimental design, inference, evaluation of imbalanced real-world data, and knowing when a result is real versus an artifact of how it was measured.
  • Deep, genuine fluency in the fundamentals: train/test discipline, overfitting and regularization, data leakage, class imbalance, calibration, and how to evaluate a model honestly. These will be probed directly in the interview.
  • Strong production Python (pandas, NumPy, scikit-learn; deep-learning frameworks a plus) and strong SQL, including stored procedures and performance-aware queries against large tables, with comfort using Git and reproducible workflows.
  • Hands-on MLOps: monitoring, drift, experiment tracking, versioning, and what it takes to keep a model healthy in production.
  • Comfort owning the data pipeline, not just the model - source-system integration, scheduling, and data-quality controls.

The above statements are intended to provide the general nature and level of work being performed by most people assigned to the position. They are not intended to be an exhaustive list of all responsibilities, duties and requirements.

Benefits & conditions

Pulled from the full job description

  • 401(k)
  • Health insurance
  • Paid time off
  • Vision insurance
  • Dental insurance
  • Life insurance
  • Disability insurance, This position is remote and requires a dedicated, distraction-free work space at home. We offer a competitive benefits package including medical, dental, vision, life insurance, short term disability, long term disability, paid holidays, 401k, and a generous PTO policy.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:12 min

Using artificial intelligence in healthcare diagnostics and workforce surveillance

Chris Heilmann +2 · LIVE

Videos

See all

Related articles

See all