> Markdown version of [/jobs/ext/3108048-lead-data-ai-engineer](https://www.wearedevelopers.com/jobs/ext/3108048-lead-data-ai-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Data & AI Engineer - **Company:** THE PHOENIX - **Location:** Phoenix, AZ, United States (Remote available) - **Experience:** Expert - **Salary:** $104,000.0 - $124,800.0 - **Contract:** Temporary to permanent - **Skills:** Cerner, Third Normal Form, Application Programming Interfaces (APIs), Artificial Intelligence, ASC X12 Standards, Cloud Database, Continuous Integration, Information Engineering, Extract Transform Load (ETL), Data Security, Data Sharing, Data Vault Modeling, Digital Architecture, Dimensional Modeling, Fraud Prevention and Detection, Github, Healthcare Effectiveness Data and Information Set, Python (Programming Language), Machine Learning, Meta-Data Management, Role-Based Access Control, Power BI, Tensorflow, Azure Machine Learning, Unstructured Data, Management of Software Versions, Parquet, File Transfer Protocol (FTP), Sql Optimization, Pytorch, Fast Healthcare Interoperability Resources, Snowflake, Build Management, Microsoft Fabric, Scikit Learn, Data Lineage, Health Level Seven International, Data Management, Machine Learning Operations, Physical Data Models - **Published:** September 27, 2026 - **Apply:** https://www.juju.com/job/16_9aa784166 ## About the Role · 8+ years of experience in data engineering or analytics with at least 5 years of hands-on Snowflake expertise including virtual warehouses, tasks, streams, Snowpipe, RBAC, masking, and data sharing. · 2+ years of experience with Microsoft Fabric including OneLake, Lakehouses, Warehouses, Dataflows Gen2, Notebooks, and Pipelines. · Advanced SQL skills with strong experience in ETL/ELT development using Python, dbt, Dataflows, or Fabric/ADF pipelines. · Deep knowledge of healthcare data standards including CMS datasets, FHIR, HL7, X12/EDI, provider data, eligibility, and claims processing. · Strong data modeling experience including dimensional modeling, SCD types, surrogate keys, 3NF, and data vault methodologies. · Experience building and deploying machine learning solutions using tools such as scikit-learn, PyTorch, TensorFlow, Azure ML, or Fabric ML. · Practical experience managing HIPAA compliance, PHI handling, auditing, and secure access controls within cloud data environments. · Experience working with both structured data formats such as Parquet and CSV and unstructured data such as clinical notes and PDFs. · Strong communication skills with the ability to produce mapping specifications, lineage documentation, and present technical trade-offs clearly. · Preferred: Experience with Epic or Cerner integrations, HEDIS or risk adjustment programs, MLOps tools such as MLflow or GitHub Actions, Power BI semantic modeling, and relevant Snowflake or Microsoft certifications. ## Description We're looking for a Lead Data & AI Engineer to lead the design and delivery of secure, scalable data and AI solutions within complex healthcare environments. The position focuses on building modern data platforms, integrating diverse clinical and claims datasets, and operationalizing machine learning models that improve cost, quality, and patient outcomes. Your role · Design, implement, and optimize data platforms using Snowflake and Microsoft Fabric, including Lakehouses, Warehouses, OneLake, and engineering pipelines. · Build and maintain scalable ingestion frameworks for batch and streaming data sources such as APIs, ADLS, SFTP, and event streams with full lineage and governance. · Develop secure data environments that comply with HIPAA and PHI requirements using role-based access, masking, tokenization, and de-identification. · Create conceptual, logical, and physical data models using dimensional, normalized, and data vault approaches. · Transform and normalize structured and unstructured healthcare data including claims, eligibility, enrollment, provider, and clinical documentation. · Integrate and harmonize data using FHIR, HL7, X12/EDI 837/835, NCPDP, and CMS standards across payer, provider, EHR, and HIE systems. · Build and deploy machine learning pipelines for risk modeling, utilization forecasting, fraud detection, quality measurement, and care gap analysis. · Operationalize models with strong MLOps practices including versioning, CI/CD, monitoring, and drift detection. · Implement data cataloging, metadata management, lineage tracking, and quality validation using tools such as Microsoft Purview or equivalent. · Monitor and optimize pipeline performance, cost, and reliability across Snowflake and Fabric environments. · Collaborate with clinicians, actuaries, product teams, and analysts to translate business needs into scalable technical solutions. · Document architecture, data mappings, and design standards while mentoring engineers and contributing to enterprise best practices. ## Related Videos - [Parquet, Delta, Iceberg & Ducklake - An introduction for developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [OLAP for AI Applications and why you should care](https://www.wearedevelopers.com/videos/100212-olap-for-ai-applications-and-why-you-should-care) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know)