> Markdown version of [/jobs/ext/1338602-principal-data-engineer](https://www.wearedevelopers.com/jobs/ext/1338602-principal-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Data Engineer - **Company:** Codametrix Inc. - **Location:** Boston, MA, United States - **Experience:** Expert - **Salary:** $175,000.0 - $200,000.0 - **Contract:** Permanent contract - **Skills:** Training Data, Artificial Intelligence, Amazon Web Services, Amazon S3, Data Analysis, Audit Trail, Automation of Tests, Continuous Integration, Data Architecture, Information Engineering, Data Governance, Data Infrastructure, DevOps, Disaster Recovery, Failover, Github, Apache Hive, Identity and Access Management, Python (Programming Language), Machine Learning, Performance Tuning, Standard Sql, Data Streaming, Tableau (Software), Software Technical Review, User Provisioning Software, YAML, Fast Healthcare Interoperability Resources, Electronic Medical Records, Amazon Virtual Private Cloud (VPC), Data Strategy, Data Lakes, Pyspark, Information Technology, Health Level Seven International, Apache Kafka, Data Management, Terraform, Serverless Computing, Jenkins, Databricks - **Published:** July 18, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=01f2831abce85a26 ## About the Role * A degree in Computer Science or a related field (Bachelor's, Master's, or Ph.D.), or an equivalent combination of education and demonstrable professional experience * 8+ years of data engineering experience with progressive responsibility * 5+ years hands-on experience with the Databricks platform (Unity Catalog, Delta Lake, Structured Streaming, Spark SQL, cluster management, platform administration) * Expert proficiency in PySpark and Python; working knowledge of Scala * Expert experience with Terraform for infrastructure-as-code (state management, modular structures, multi-environment deployments, YAML-driven configuration) * Expert experience with Apache Kafka / AWS MSK (streaming ingestion, SASL/IAM auth, topic management, cluster migrations) * Deep understanding of medallion/lakehouse architecture patterns (bronze/silver/gold, SCD2, materialized views, slowly changing dimensions) * Proven track record building and maintaining CI/CD pipelines (Jenkins, GitHub Actions) for data platform deployments * Expert SQL skills - complex CTEs, window functions, performance optimization of large-scale queries across petabyte-scale datasets * Strong AWS experience (S3, IAM, Secrets Manager, VPC/Private Link, cross-region replication) * Demonstrated ability to operate in a HIPAA-regulated environment with PHI handling requirements * Experience with disaster recovery architecture - cross-region replication, failover procedures, read-only replicas * Experience using AI tools (e.g., Claude, Gemini, Codex) and agentic workflows to augment design, development, and testing processes. * Huge Plus: Prior experience working with healthcare data, such as medical claims, electronic health records (EHR), or billing code systems (ICD-10, CPT). Familiarity with healthcare data standards like HL7 or FHIR is highly desirable. Track record of influencing technical direction beyond immediate team - architecture reviews, establishing org-widestandards, technology selection. Strategic ability to make build-vs-buy decisions and evaluate emerging technologies for platform evolution ## Description The Principal Data Engineer is a member of the Data Platform team, reporting to the Director of Machine Learning Engineering and Data. The Data Platform team is responsible for executing the data strategy for the organization, ensuring high-quality external data is ingested into the Lakehouse and realized in powerful insights for internal and external customers, while ensuring ML/AI, BI and customer success teams have the data they need to develop and train their models, build insightful dashboards and design semantic layer. As a Principal Data Engineer (L4), you will serve as a key technical leader for CodaMetrix's Databricks-based data platform, supporting streaming and batch workloads across 30+ healthcare customers. You will own the platform's architecture, evolution, and operational excellence, including infrastructure-as-code, CI/CD automation, disaster recovery, cost optimization, data contracts, and build-vs-buy decisions, while enabling ML Engineering, Analytics, DevOps, and Product teams and influencing technical standards beyond the immediate team. This role operates with minimal oversight and is expected to influence technical standards beyond the immediate team., * Own the technical strategy and roadmap - Own the vision, architecture, and roadmap for the CodaMetrix data platform, ensuring scalability, reliability, regulatory alignment, and operational excellence. Lead key technology decisions, disaster recovery design, architecture reviews, and data engineering standards across teams. * Platform Engineering at Scale - Design, build, and maintain scalable streaming and batch data platforms Databricks using PySpark, Unity Catalog, Delta Lake, and Structured Streaming. Own Terraform-based infrastructure across environments, including jobs, catalogs, schemas, permissions, compute policies, volumes, and external locations. Build Jenkins CI/CD workflows for automated testing, tagging, and deployments, while optimizing training pipelines, materialized views, retention policies, and production performance. * Data Governance & Security - Implement and evolve least-privilege access controls across Databricks using Unity Catalog grants, YAML-driven group policies, and schema-level restrictions. Ensure HIPAA and SOC 2 compliance through PHI masking, audit logging, environment-level data segmentation, user provisioning, compute policies, and cost attribution. * Cost Optimization & Operational Excellence - Drive platform cost reduction through compute policy tuning, serverless optimization, reserved pools, materialized view improvements, and remediation of underutilized resources. Monitor Databricks/AWS spend (using tools like CloudZero and AWS Cost Explorer), own cost attribution and budgeting, resolve production incidents, and maintain runbooks to ensure platform SLAs for uptime and performance. * Cross-Functional Enablement & Mentorship - Enable ML, BI, Analytics, DevOps, and customer success and implementation teams with training data pipelines, model-ready datasets, feature store architecture, optimized views, dashboards, Tableau refreshes, infrastructure changes, and tenant onboarding. Mentor engineers through code and design reviews, establish quality standards, and serve as a subject matter expert for data platform engineering across the organization. ## Related Videos - [CI/CD with Github Actions](https://www.wearedevelopers.com/videos/856-ci-cd-with-github-actions) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk)