> Markdown version of [/jobs/ext/604953-senior-data-engineer](https://www.wearedevelopers.com/jobs/ext/604953-senior-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Engineer - **Company:** Start Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $130,000.0 - $145,000.0 - **Contract:** Permanent contract - **Skills:** Microsoft Azure, Clinical Trial Management Systems, Continuous Integration, Data Deduplication, Information Engineering, Data Governance, Data Infrastructure, Extract Transform Load (ETL), Filemaker, Monitoring of Systems, Apache Hive, Job Scheduling, Python (Programming Language), Metadata, Microsoft Software, SQL Azure, Netsuite, Operational Databases, Role-Based Access Control, Software Tools, Azure Data Lake, SQL Databases, Systems Integration, Management of Software Versions, Enterprise Data Management, Azure Service Bus, Data Classification, Data Ingestion, Azure Data Factory, Snowflake, Apache Spark, Boomi, Change Data Capture, Data Lakes, Pyspark, Data Lineage, Data Management, Hubspot, Data Pipelines, Databricks - **Published:** June 24, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=76fda9798e02d484 ## About the Role * 4+ years of experience as a data engineer, with at least 2 years on Azure Databricks or equivalent Spark-based platforms * Strong/Current proficiency in Azure Data Factory, ADLS Gen2, Databricks, Delta Lake, Databricks SQL, Azure SQL * Strong/Current proficiency in Python, SQL, PySpark, and Spark SQL * Experience with Data Lakes, Lake Houses and Warehouses * Experience building production data pipelines with proper error handling, retry logic, idempotency, and monitoring * Familiarity Azure and Azure ecosystem services * Experience with Unity Catalog, Purview or equivalent data governance/catalog tooling * Experience with Data Governance guidelines such as Data Classification, Retention, De-Identification, Tenancy, Sovereignty and Data Standards * Experience with CI/CD for data engineering workloads (Databricks Repos, Azure DevOps, or similar) Preferred Education and Experience: * Experience with MDM data pipelines (matching, deduplication, golden record logic) * Familiarity with ERP data models (NetSuite preferred) or clinical trial management systems (OnCore) * Experience with Snowflake (existing analytical layer we are integrating with) * Experience with Boomi or other iPaaS tools from a data engineering perspective * Background in healthcare, life sciences, or clinical research data * Experience building financial reconciliation or revenue recognition datasets ## Description We are hiring a Sr. Data Engineer to build the data infrastructure that powers our Enterprise Data Platform and integrated systems capabilities. You will work closely with our Enterprise Data Leader to Design and Implement ingestion pipelines, Medallion Tiers, construct Lakehouse/Relational/Warehouse data models, build data quality/lineage/access-control frameworks, and operationalize the analytical datasets that our data producing/consuming teams depend on. This is a build-phase role. We have a newly deployed Azure/Databricks environment and need someone who can be productive immediately. You will be responsible for turning architectural designs into working, reliable, production-grade data pipelines within a tight timeline., Data Pipeline Development * Build and maintain ingestion pipelines from source systems (OnCore, NetSuite, HubSpot, Microsoft Lists, Snowflake, FileMaker) into ADLS/Databricks * Implement incremental load patterns, change data capture, and idempotent pipeline design to ensure reliability * Design/Implement Metadata, Data Quality and Data Lineage capabilities * Design/Implement AccessControl/RBAC capabilities * Design/Implement DataRights/Licensing capabilities * Develop the ETL/ELT processes that feed Lakehouse/Relational/Warehouse modeling requirements (matching, deduplication, golden record assembly) as designed by the MDM Lead * Build publication pipelines that push canonical data models from the Data Platform back to spoke systems, coordinating with Integration Platform (Boomi based) and Messaging Platform (Azure Service Bus based) Teams MDM Hub Implementation * Implement Lakehouse/Relational/Warehouse tables for golden records across priority entities: Study/Protocol, Customer/Sponsor, Item/Charge Code, and Contract * Build the matching and survivorship logic based on rules defined by data model requirements and validated by business stakeholders * Implement versioning, lineage and audit trails using Delta Lake time travel capabilities for full traceability of master data changes * Configure Unity Catalog for data governance, access controls, and lineage tracking Data Quality and Monitoring * Build automated data quality checks at ingestion, transformation, and publication stages * Develop data quality dashboards and alerting (integrate with existing monitoring tools or build in Databricks SQL) * Implement reconciliation count checks between source systems and the hub to detect drift or sync failures * Create exception handling pipelines that surface records requiring manual review Analytical and Reconciliation Datasets * Build the data models supporting Revenue Cycle (e.g. OnCore-to-NetSuite) reconciliation (matching clinical events to financial transactions) * Develop the unbilled-vs-billed tracking datasets that compare recognized sales order lines against invoiced amounts * Create revenue accrual support datasets that feed the finance team's automated journal entry processes in NetSuite * Support pass-through item mapping and amendment pricing reconciliation data needs as the finance team defines requirements Platform Operations * Establish CI/CD patterns for Databricks notebooks and jobs (Repos integration, testing frameworks) * Configure and manage job scheduling, cluster policies, and cost optimization * Maintain dev/staging/production environment separation * Document all pipelines, data models, and operational procedures, * Occasional travel (0-5% annually) may be required for team meetings, company events, or client engagements, though most collaboration will occur virtually via video conferencing and online tools. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Integrate your Cognitive Assistant with 3rd-party DBs and software](https://www.wearedevelopers.com/videos/249-integrate-your-cognitive-assistant-with-3rd-party-dbs-and-software) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)