> Markdown version of [/jobs/ext/1459098-data-engineer](https://www.wearedevelopers.com/jobs/ext/1459098-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Pyx Health Inc - **Location:** United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Airflow, Microsoft Azure, Software Quality, Code Review, Continuous Integration, Directed Acyclic Graph (Directed Graphs), Data Governance, Data Infrastructure, Extract Transform Load (ETL), Data Transformation, Software Debugging, Python (Programming Language), Microsoft SQL Server, Operational Databases, Azure DevOps Pipelines, Standard Sql, Azure Data Lake, SQL Stored Procedures, SQL Databases, Transact-SQL, Data Processing, Microsoft Power Automate, Azure Data Factory, Apache Spark, Change Data Capture, Git, Data Lakes, Pyspark, Information Technology, Data Analytics, Health Level Seven International, Software Version Control, Data Pipelines, Databricks - **Published:** July 27, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=8ea50071b5800db9 ## About the Role * 2-4 years of experience as a Data Engineer or in a closely related data role * Hands-on experience with Azure cloud services (ADLS, Databricks, or similar) * Working knowledge of SQL for scripting and data modeling including T-SQL, Spark, and Databricks SQL * Ability to contribute to technical projects with moderate oversight * Strong communication skills - comfortable asking questions and providing updates to cross-functional partners * Familiarity with CI/CD concepts, Git, and version control workflows * Solid problem-solving skills with a systematic approach to debugging * Familiarity with healthcare data standards and regulations (HIPAA, HL7, etc.) is a plus * Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field (or equivalent work experience) MUST HAVES * Databricks SQL (Mid): Delta Lake fundamentals, basic merge/upsert patterns, familiarity with CDC concepts. * Python (Mid): Pipeline logic, data transformation, SQL scripting; some PySpark/Spark DataFrame experience. * Databricks Spark Notebooks (Mid): Notebook-based development, basic cluster usage, Delta table operations. * Airflow Python Development (Mid): Ability to write and maintain DAGs; understanding of error handling and retry patterns. * Airflow Astro Configuration (Foundational): Familiarity with Astronomer or willingness to learn; basic Astro CLI usage. * Azure Ecosystem (Foundational): Working knowledge of ADLS Gen2, Key Vaults, and Azure DevOps. * T-SQL (Foundational-Mid): Basic stored procedures, SQL Server querying, and data manipulation. NICE TO HAVES * Azure Data Factory (Foundational): Basic familiarity with pipeline authoring and triggers. * Git + Azure DevOps CI/CD (Foundational): Experience with branching, pull requests, peer review processes, and version control best practices as well as deploying production data pipelines using CI/CD workflows. ## Description Pyx Health is looking for a motivated and technically solid Healthcare Data Engineer to join our growing Data & Analytics team. In this role, you will contribute to building and maintaining our data infrastructure on Azure, work alongside senior engineers to develop reliable data pipelines, and collaborate with analytics teams to engineer data foundations that support our healthcare solutions. You will work with Databricks, Airflow, Python, and SQL to implement and improve data workflows in a collaborative, high-growth environment. Only candidates residing in the USA may apply., Pipeline Development & Maintenance * Build and maintain data pipelines for ingesting, transforming, and cleaning healthcare data in the Azure cloud using Databricks, PySpark, and Delta Lake * Implement pipeline logic from defined specifications, with guidance from senior engineers on architectural decisions * Build reusable testing frameworks to ensure pipeline reliability * Monitor pipelines for failures and performance issues, escalating complex problems appropriately Orchestration & Automation * Develop and maintain Airflow DAGs with appropriate error handling and retry logic * Support deployment and configuration of pipelines via Astronomer on Azure using the Astro CLI * Contribute to improving pipeline reliability and reducing manual intervention Data Modeling & Storage * Implement data models for efficient storage and retrieval using Delta Lake, including merge/upsert patterns * Design and optimize Delta Lake tables for analytic workloads * Leverage Unity Catalog for data governance, organization, and security * Apply Change Data Capture (CDC) patterns under the direction of senior team members * Write and optimize T-SQL and Databricks SQL scripts, stored procedures, and ETL support Azure Infrastructure * Work within the Azure ecosystem including ADLS Gen2, Key Vaults, Logic Apps, and Azure DevOps * Follow established security and scalability standards when building data infrastructure * Support infrastructure tasks with guidance from senior engineers on architecture Data Quality & Monitoring * Ensure scripts and datasets are well-documented to support enterprise data governance * Design, implement and continuously improve automated data quality monitoring, reconciliation, and alerting to ensure reliable downstream reporting * Troubleshoot and resolve pipeline failures, documenting root causes and resolutions Collaboration & Documentation * Work closely with data scientists, analysts, and business stakeholders to understand reporting requirements, troubleshoot issues, and improve data assets * Participate in code reviews, incorporating feedback to improve code quality * Document pipelines, processes, and implementation decisions clearly and consistently ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk)