Senior Data Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+14 more
Job description
Caret Health is looking for a Senior Data Engineer to own the entire data foundation that our platform and AI systems run on. You will be responsible for the holistic data infrastructure - from raw ingestion across every source system to the clean, versioned, audit-ready datasets that power clinical programs, product analytics, and machine learning. This is a high-ownership IC role at the intersection of senior data engineering and AI enablement: you will set the architectural direction for our data platform, work directly with the AI team to support model training and evaluation, build dashboards and reporting layers for clinical and operational stakeholders, and enable LLM-driven querying over our data. Healthcare data is complex and high-stakes - we need someone who has done this before, treats data quality and governance as a first-class concern, and can make sound architectural decisions that scale., Data Infrastructure & Architecture
- Own the end-to-end data infrastructure at Caret Health - data lake, data warehouse, and the pipelines connecting them - serving as the single accountable engineer for data reliability, availability, and governance.
- Design and maintain ingestion pipelines aggregating data from all source systems: EHRs, clinical programs, payer and claims feeds, third-party APIs, and internal operational tools.
- Architect schema design, modeling standards, partitioning strategy, and access controls with HIPAA compliance and PHI handling built in by default
- Implement transformation and modeling layers (dbt or equivalent) with clear lineage, documentation, and version control so every dataset is traceable and reproducible.
Data Quality & Observability
- Define and enforce data quality standards across all datasets - validation rules, anomaly detection, freshness monitoring, and lineage documentation.
- Build observability into every pipeline so failures surface early and are traceable to source; treat data incidents with the same urgency as production outages.
- Document data models, field definitions, and lineage so clinical and product teams can self-serve without engineering support.
AI & ML Enablement
- Partner closely with the AI engineering team to provide clean, versioned, well-documented datasets for model training, evaluation, fine-tuning, and inference across all use cases.
- Prepare and maintain purpose-built datasets for specific AI use cases - including supervised training sets, evaluation benchmarks, and prompt datasets for LLM workflows.
- Build and maintain feedback loops that route outcome data back into the platform so models can continuously improve over time.
- Support LLM-driven data querying - enabling natural language interfaces over structured clinical and operational data so non-technical stakeholders can query data conversationally.
Analytics, Dashboards & Reporting
- Own the analytics and dashboarding layer, giving clinical, operational, and executive teams real-time visibility into patient engagement outcomes, program performance, and quality metrics.
- Build a self-serve reporting infrastructure so product and clinical teams can access and slice data without requiring engineering cycles.
- Partner with product and program management on ad hoc analyses, outcome reporting, and client-facing data deliverables.
Requirements
Do you have experience in System design?, * 5+ years building and operating production data pipelines, data lakes, and warehouses - with a track record of owning the full data platform, not just individual pipelines.
- Demonstrated experience making and defending architectural decisions: warehouse design, lake structure, modeling standards, and access controls at scale.
- Expert-level SQL and strong Python - you write clean, tested transformation code and apply software engineering discipline to data work.
- Deep hands-on experience with a cloud data warehouse (Snowflake strongly preferred) and a transformation framework (dbt preferred); you know how to model complex, evolving domains.
- Experience designing multi-layer data lake architecture - raw, curated, and consumption layers - with clear governance, lineage, and access controls.
- Proven experience supporting ML or AI teams: you understand training data requirements, dataset versioning, evaluation set construction, and the feedback loops models depend on.
- Experience enabling LLM-driven or natural language querying over structured data - semantic layers, query generation, or similar patterns.
- Hands-on experience with healthcare data formats and standards (EHR feeds, HL7, FHIR, claims, ADT events) and HIPAA technical requirements for PHI handling is a strong plus.
Benefits & conditions
2.52.5 out of 5 stars Remote $130,000 - $190,000 a year - Full-time, Pulled from the full job description
- Paid time off, Pay: $130,000.00 - $190,000.00 per year
Benefits:
- Paid time off
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on indeed.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Data Engineer Salary UK
Highest Paying Tech Companies for Developers
Top Big Data Technologies That You Need to Know
Data Analyst Salary in the UK