> Markdown version of [/jobs/ext/3080503-principal-data-engineer](https://www.wearedevelopers.com/jobs/ext/3080503-principal-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Data Engineer - **Company:** DATA ENGINEERING, LLC - **Location:** Charlotte, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Agile Methodology, Artificial Intelligence, Data Analysis, Automation of Tests, Microsoft Azure, Backup Devices, Cloud Storage, Code Review, Continuous Integration, Data Architecture, Data Validation, Information Engineering, Data Governance, Data Infrastructure, Extract Transform Load (ETL), Data Security, Data Systems, Relational Databases, Software Debugging, DevOps, Disaster Recovery, Fault Tolerance, Github, Data Intelligence, Key Management, Machine Learning, Meta-Data Management, SQL Azure, Operational Databases, Open Web Application Security, Performance Tuning, Scrum Methodology, Role-Based Access Control, Migration Manager, DataOps, Azure Data Lake, Software Engineering, SQL Stored Procedures, SQL Databases, Data Streaming, Parquet, Large Language Models, Data Build Tool (dbt), Database Optimization, Apache Spark, Caching, Technical Debt, Data Strategy, Git, Data Lakes, Pyspark, Information Technology, Low Latency, Real Time Data, Apache Kafka, Spark Streaming, Data Management, Machine Learning Operations, Virtual Agents, Event Sourcing, Cloud Optimization, Stream Processing, Software Version Control, Data Pipelines, Confluent, Databricks - **Published:** September 25, 2026 - **Apply:** https://startup.jobs/principal-data-engineer-avidxchange-inc-10177924 ## About the Role · Bachelor's degree in Computer Science, Engineering, or a related field with 10+ years of data engineering experience in a high-availability, business-critical environment. · Hands-on Databricks expertise: Delta Lake, Unity Catalog, Databricks Workflows, Databricks SQL, and cluster/job optimization. · Proven experience migrating from legacy Azure SQL Server (or other relational RDBMS) to a Databricks Lakehouse - schema translation, data validation, and cutover strategies. · Strong proficiency in Apache Kafka for streaming pipelines - producers/consumers, topic design, partitioning strategy, and Kafka Connect. · Expert-level PySpark and/or Scala Spark skills, including performance tuning, broadcasting, partitioning, and caching. · Deep understanding of data architecture patterns: medallion (bronze/silver/gold), Lambda/Kappa, event sourcing, and streaming-first designs. · Hands-on experience with Azure cloud services (ADLS Gen2, Azure Event Hubs, ADF, Azure Key Vault, Azure Monitor). · Strong knowledge of infrastructure components - networking, cloud storage (Delta/Parquet), and cloud cost optimization. Preferred · Databricks Certified Data Engineer Associate or Professional certification. · Experience with ksqlDB, Kafka Streams, or Spark Structured Streaming for stateful stream processing. · Confluent Platform experience: Schema Registry, Kafka Connect connectors, RBAC, and cluster management. · Advanced Azure SQL Server expertise (2018+/Azure SQL MI) - stored procedures, indexing strategies, query plan analysis - valuable for migration contexts. · Experience with dbt (data build tool) for SQL-based transformation layers on Databricks. · Proficiency with Git and CI/CD pipelines for data engineering (Azure DevOps, GitHub Actions). · Experience working in Agile environments (Scrum/Kanban). · Familiarity with data governance frameworks, data cataloging (Unity Catalog, Microsoft Purview), and data quality tooling (Great Expectations, Monte Carlo). · Familiarity with secure coding practices, including OWASP Top 10 and secrets management. · Experience implementing DataOps practices: automated testing, data observability, pipeline CI/CD, data contracts, and SLA monitoring across the Lakehouse. · Hands-on MLOps experience: model versioning (MLflow), experiment tracking, model serving, and integrating ML pipelines with production data workflows on Databricks. · Familiarity with Databricks Mosaic AI (formerly MLflow + Model Serving) for end-to-end MLOps lifecycle management. · Experience with real-time ML feature stores or Lakehouse-based ML pipelines is a plus., **Must be full-time for at least 3 months ***Must be full-time for at least one year ## Description · Lead the design and implementation of scalable, cloud-native data architectures on Databricks (Delta Lake, Unity Catalog, Lakehouse patterns). · Own and execute the migration strategy from legacy Azure SQL Server to Databricks, including schema translation, ETL/ELT re-platforming, data validation, and cutover planning. · Define data modeling standards (medallion architecture, star/snowflake schemas) and ensure consistency across all pipelines and domains. · Evaluate and recommend tools, frameworks, and platforms to support long-term data strategy and organizational goals. · Collaborate with Solution and Enterprise Architects to review and approve new data architecture designs. Streaming & Real-Time Data Engineering · Architect and implement Kafka-based streaming pipelines for real-time data ingestion, transformation, and delivery. · Design event-driven architectures and streaming topologies using Kafka Streams, ksqlDB, or Spark Structured Streaming on Databricks. · Establish patterns for schema management (Confluent Schema Registry), consumer group strategy, offset management, and dead-letter queuing. · Ensure streaming pipelines meet SLA requirements for latency, throughput, and fault tolerance. Optimization, Quality & Standards · Debug and optimize Spark jobs, Delta Lake tables, and SQL workloads for performance, cost efficiency, and maintainability. · Lead code reviews focused on senior engineers to enforce standards, best practices, and technical quality. · Manage pipeline quality, data models, and CI/CD delivery workflows; guide teams on continuous improvement. · Promote strong data management practices - data quality, lineage, observability, and governance. · Identify opportunities to improve service delivery methods, processes, and resource utilization. Leadership, Mentorship & Strategy · Mentor data engineers at all levels, with particular emphasis on developing senior talent. · Establish and evolve data engineering standards, best practices, and management of technical debt. · Stay current on data platform trends (Databricks releases, Kafka ecosystem, open table formats) and contribute to long-term architectural vision. · Develop plans for data security, disaster recovery, backup, business continuity, and archiving across the Lakehouse. AI, Agentic Capabilities & Intelligent Data Products · Design and enable AI and agentic capabilities on the Databricks platform, including Databricks Genie for natural language data exploration and self-service analytics. · Architect data foundations - clean, governed, well-documented Delta tables - that power Genie spaces, AI/BI dashboards, and LLM-driven data agents. · Collaborate with ML and AI teams to build and maintain feature stores, vector stores, and retrieval-augmented generation (RAG) pipelines on Databricks. · Evaluate and integrate emerging agentic frameworks (LangChain, Mosaic AI Agent Framework) to automate data workflows and enable intelligent data products. · Define governance and observability standards for AI-driven data pipelines, ensuring reliability, auditability, and responsible AI practices. Cross-Functional Collaboration · Partner with project managers and business leaders on initiatives involving enterprise data. · Collaborate across teams to influence and strengthen data engineering practices organization-wide. · Work with Analytics, ML, and product engineers to design and deliver end-to-end data solutions that meet business needs. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)