Lead Data & AI Engineer

THE PHOENIX
Phoenix, AZ, United States
7 days ago
Apply on www.juju.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Compensation
$104,000.0 - $124,800.0
Working hours
Regular working hours
Job source

Tech stack

Cerner Third Normal Form Application Programming Interfaces (APIs) Artificial Intelligence ASC X12 Standards Cloud Database Continuous Integration Information Engineering Extract Transform Load (ETL) Data Security Data Sharing Data Vault Modeling
+28 more
Digital Architecture Dimensional Modeling Fraud Prevention and Detection Github Healthcare Effectiveness Data and Information Set Python (Programming Language) Machine Learning Meta-Data Management Role-Based Access Control Power BI Tensorflow Azure Machine Learning Unstructured Data Management of Software Versions Parquet File Transfer Protocol (FTP) Sql Optimization Pytorch Fast Healthcare Interoperability Resources Snowflake Build Management Microsoft Fabric Scikit Learn Data Lineage Health Level Seven International Data Management Machine Learning Operations Physical Data Models

Job description

We’re looking for a Lead Data & AI Engineer to lead the design and delivery of secure, scalable data and AI solutions within complex healthcare environments. The position focuses on building modern data platforms, integrating diverse clinical and claims datasets, and operationalizing machine learning models that improve cost, quality, and patient outcomes.

Your role

· Design, implement, and optimize data platforms using Snowflake and Microsoft Fabric, including Lakehouses, Warehouses, OneLake, and engineering pipelines.

· Build and maintain scalable ingestion frameworks for batch and streaming data sources such as APIs, ADLS, SFTP, and event streams with full lineage and governance.

· Develop secure data environments that comply with HIPAA and PHI requirements using role-based access, masking, tokenization, and de-identification.

· Create conceptual, logical, and physical data models using dimensional, normalized, and data vault approaches.

· Transform and normalize structured and unstructured healthcare data including claims, eligibility, enrollment, provider, and clinical documentation.

· Integrate and harmonize data using FHIR, HL7, X12/EDI 837/835, NCPDP, and CMS standards across payer, provider, EHR, and HIE systems.

· Build and deploy machine learning pipelines for risk modeling, utilization forecasting, fraud detection, quality measurement, and care gap analysis.

· Operationalize models with strong MLOps practices including versioning, CI/CD, monitoring, and drift detection.

· Implement data cataloging, metadata management, lineage tracking, and quality validation using tools such as Microsoft Purview or equivalent.

· Monitor and optimize pipeline performance, cost, and reliability across Snowflake and Fabric environments.

· Collaborate with clinicians, actuaries, product teams, and analysts to translate business needs into scalable technical solutions.

· Document architecture, data mappings, and design standards while mentoring engineers and contributing to enterprise best practices.

Requirements

· 8+ years of experience in data engineering or analytics with at least 5 years of hands-on Snowflake expertise including virtual warehouses, tasks, streams, Snowpipe, RBAC, masking, and data sharing.

· 2+ years of experience with Microsoft Fabric including OneLake, Lakehouses, Warehouses, Dataflows Gen2, Notebooks, and Pipelines.

· Advanced SQL skills with strong experience in ETL/ELT development using Python, dbt, Dataflows, or Fabric/ADF pipelines.

· Deep knowledge of healthcare data standards including CMS datasets, FHIR, HL7, X12/EDI, provider data, eligibility, and claims processing.

· Strong data modeling experience including dimensional modeling, SCD types, surrogate keys, 3NF, and data vault methodologies.

· Experience building and deploying machine learning solutions using tools such as scikit-learn, PyTorch, TensorFlow, Azure ML, or Fabric ML.

· Practical experience managing HIPAA compliance, PHI handling, auditing, and secure access controls within cloud data environments.

· Experience working with both structured data formats such as Parquet and CSV and unstructured data such as clinical notes and PDFs.

· Strong communication skills with the ability to produce mapping specifications, lineage documentation, and present technical trade-offs clearly.

· Preferred: Experience with Epic or Cerner integrations, HEDIS or risk adjustment programs, MLOps tools such as MLflow or GitHub Actions, Power BI semantic modeling, and relevant Snowflake or Microsoft certifications.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

1:26 min

Summarizing 2024 milestones and introducing the Data Phoenix concept

Prashanth Chandrasekar Prashanth Chandrasekar · World Congress 2024

2:03 min

Introduction to open table formats built on Parquet

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

Videos

See all

Related articles

See all