> Markdown version of [/jobs/ext/564430-senior-data-engineer](https://www.wearedevelopers.com/jobs/ext/564430-senior-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Engineer - **Company:** CARET HEALTH INC. - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $130,000.0 - $190,000.0 - **Contract:** Permanent contract - **Skills:** Training Data, Sql Data Warehouse, Artificial Intelligence, Data Analysis, Backup Devices, Information Engineering, Data Infrastructure, Data Warehousing, Document-Oriented Databases, Python (Programming Language), Machine Learning, Operational Data Store, Operational Databases, Software Engineering, SQL Databases, Management of Software Versions, Usage Analysis, Data Ingestion, Fast Healthcare Interoperability Resources, Large Language Models, Snowflake, Data Layers, Data Lakes, Health Level Seven International, Software Version Control, Web Api - **Published:** June 12, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=27bcc3d44a3fe6d3 ## About the Role Do you have experience in System design?, * 5+ years building and operating production data pipelines, data lakes, and warehouses - with a track record of owning the full data platform, not just individual pipelines. * Demonstrated experience making and defending architectural decisions: warehouse design, lake structure, modeling standards, and access controls at scale. * Expert-level SQL and strong Python - you write clean, tested transformation code and apply software engineering discipline to data work. * Deep hands-on experience with a cloud data warehouse (Snowflake strongly preferred) and a transformation framework (dbt preferred); you know how to model complex, evolving domains. * Experience designing multi-layer data lake architecture - raw, curated, and consumption layers - with clear governance, lineage, and access controls. * Proven experience supporting ML or AI teams: you understand training data requirements, dataset versioning, evaluation set construction, and the feedback loops models depend on. * Experience enabling LLM-driven or natural language querying over structured data - semantic layers, query generation, or similar patterns. * Hands-on experience with healthcare data formats and standards (EHR feeds, HL7, FHIR, claims, ADT events) and HIPAA technical requirements for PHI handling is a strong plus. ## Description Caret Health is looking for a Senior Data Engineer to own the entire data foundation that our platform and AI systems run on. You will be responsible for the holistic data infrastructure - from raw ingestion across every source system to the clean, versioned, audit-ready datasets that power clinical programs, product analytics, and machine learning. This is a high-ownership IC role at the intersection of senior data engineering and AI enablement: you will set the architectural direction for our data platform, work directly with the AI team to support model training and evaluation, build dashboards and reporting layers for clinical and operational stakeholders, and enable LLM-driven querying over our data. Healthcare data is complex and high-stakes - we need someone who has done this before, treats data quality and governance as a first-class concern, and can make sound architectural decisions that scale., Data Infrastructure & Architecture * Own the end-to-end data infrastructure at Caret Health - data lake, data warehouse, and the pipelines connecting them - serving as the single accountable engineer for data reliability, availability, and governance. * Design and maintain ingestion pipelines aggregating data from all source systems: EHRs, clinical programs, payer and claims feeds, third-party APIs, and internal operational tools. * Architect schema design, modeling standards, partitioning strategy, and access controls with HIPAA compliance and PHI handling built in by default * Implement transformation and modeling layers (dbt or equivalent) with clear lineage, documentation, and version control so every dataset is traceable and reproducible. Data Quality & Observability * Define and enforce data quality standards across all datasets - validation rules, anomaly detection, freshness monitoring, and lineage documentation. * Build observability into every pipeline so failures surface early and are traceable to source; treat data incidents with the same urgency as production outages. * Document data models, field definitions, and lineage so clinical and product teams can self-serve without engineering support. AI & ML Enablement * Partner closely with the AI engineering team to provide clean, versioned, well-documented datasets for model training, evaluation, fine-tuning, and inference across all use cases. * Prepare and maintain purpose-built datasets for specific AI use cases - including supervised training sets, evaluation benchmarks, and prompt datasets for LLM workflows. * Build and maintain feedback loops that route outcome data back into the platform so models can continuously improve over time. * Support LLM-driven data querying - enabling natural language interfaces over structured clinical and operational data so non-technical stakeholders can query data conversationally. Analytics, Dashboards & Reporting * Own the analytics and dashboarding layer, giving clinical, operational, and executive teams real-time visibility into patient engagement outcomes, program performance, and quality metrics. * Build a self-serve reporting infrastructure so product and clinical teams can access and slice data without requiring engineering cycles. * Partner with product and program management on ad hoc analyses, outcome reporting, and client-facing data deliverables. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Web APIs you might not know about](https://www.wearedevelopers.com/videos/281-web-apis-you-might-not-know-about) - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [Blueprints for Success: Steering a Global Data & AI Architecture](https://www.wearedevelopers.com/videos/1577-blueprints-for-success-steering-a-global-data-ai-architecture) - [Hacking AI at the Edge of the Indian Ocean](https://www.wearedevelopers.com/videos/100177-hacking-ai-at-the-edge-of-the-indian-ocean) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)