Principal Data Engineer

The Summit
Warrenville, IL, United States
2 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$140,000.0 - $160,000.0
Working hours
Regular working hours
Job source

Tech stack

Data Analysis Computing Platforms Audit Trail Microsoft Azure Encodings Information Engineering Data Governance Data Infrastructure Data Security DevOps Data Flow Control Github
+28 more
Global Positioning Systems (GPS) Identity and Access Management Intrusion Detection Systems Python (Programming Language) Routing Power BI Azure Active Directory Kusto Query Language SQL Databases Data Streaming Technical Data Management Systems User Provisioning Software Enterprise Data Management Data Ingestion Azure Data Factory Apache Spark Indexer Git Pandas Microsoft Fabric Data Lakes Pyspark Optimization Algorithms Performance Monitor Data Management Terraform Azure Synapse Analytics Databricks

Job description

We are seeking a highly technical Principal Data Engineer to design, scale, and evolve our enterprise data ecosystem. In this role, you will serve as the technical authority for our data infrastructure, driving the strategy for data warehouse design, real-time streaming, and robust governance.

As a Principal Engineer, you will balance hands-on technical leadership with strategic cross-functional collaboration, mentoring engineers, and aligning our data capabilities with long-term business goals. You are the strategic anchor and technical authority for our enterprise data ecosystem. In this role, you will design, govern, and optimize our Microsoft Fabric and Power BI design to transform complex, multi-modal data streams-including near-real-time telematics, operational routing, and financial ERP systems-into secure, scalable, and actionable intelligence.

You will bridge the gap between technical execution and business strategy, ensuring our data platform is robust, compliant with federal regulations (FERPA), and built for high-performance self-service analytics., Enterprise Data Design

  • Define Medallion Standards: Author, document, and defend the design boundaries and ingestion logic across the Bronze, Silver, and Gold layers within Fabric.
  • Tool & Pattern Selection: Establish clear organizational frameworks for when to deploy specific Fabric components (e.g., Mirroring vs. Dataflow Gen2 vs. Data Factory Pipelines).
  • Power BI Storage Strategy: Standardize dataset configurations, defining precise criteria for utilizing DirectLake , Import , or DirectQuery modes to balance cost, freshness, and performance.

Integration & Advanced Pipeline Engineering

  • Multi-Speed Ingestion Patterns: Design repeatable, hardened ingestion patterns for diverse data velocities:
  • Near Real-Time: GPS and IoT telematics streams.
  • Operational Batch: Dispatch, scheduling, and routing systems.
  • Transactional/Financial: Enterprise ERP data.
  • Resiliency & Frameworks: Own the end-to-end framework for data ingestion into OneLake , embedding enterprise-grade error handling, automated retry logic, and proactive monitoring/alerting systems.

Data Modeling & Master Data Management (MDM)

  • Gold Layer Definition: Design enterprise-wide dimensional models and star schemas within the Gold layer to serve as the single source of truth.
  • Semantic Layer Standardization: Establish strict governance for Power BI semantic models, including DAX measure naming conventions, hierarchy design, and Row-Level Security (RLS) structures.
  • Master Data Management: Spearhead the definition and enforcement of canonical master data entities across disparate systems, specifically governing unified definitions for Vehicle IDs, Driver IDs, and Student Records .

Data Governance, Security, and Compliance

  • Access Control & Lifecycle: Enforce Fabric workspace security utilizing Microsoft Entra ID / Security Groups for identity management, eliminating ad-hoc individual user provisioning.
  • Data Protection & Auditability: Implement Microsoft Purview sensitivity labels, robust Row-Level Security (RLS), and comprehensive audit trails, serving as the primary technical point of contact for data security audits.

Business Enablement & Analytics Delivery

  • Requirements Translation: Act as the primary conduit between operations, safety, finance, and district stakeholders, translating complex operational requests (e.g., “On-time arrival rate by district” ) into technical data products.
  • Productized Delivery: Lead the deployment of certified, user-friendly Power BI apps and dashboards, reducing friction for business users and eliminating the need for stakeholders to write direct SQL/KQL queries against the Gold layer.

Requirements

  • Platform Expertise: Deep, hands-on experience with Microsoft Fabric , OneLake , Power BI, and the Azure Data stack.
  • Data Modeling Mastery: Proven track record building enterprise Star Schemas, handling Slowly Changing Dimensions (SCDs), and establishing Master Data Management (MDM) protocols.
  • Security & Compliance: Direct experience securing highly sensitive data environments; familiarity with FERPA, HIPAA, or similar strict privacy frameworks is highly preferred.
  • Communication & Documentation: Exceptional ability to articulate complex technical decisions to executive leadership, document blueprints, and mentor engineering teams., * Experience: 10+ years of experience in data engineering, with at least 2+ years of hands-on experience designing enterprise solutions within Microsoft Fabric or advanced Azure Synapse/Databricks environments.
  • Fabric Ecosystem Mastery: Deep technical understanding of Fabric capacities, OneLake shortcuts, Managed Private Endpoints, and the interplay between Lakehouses and Warehouses.
  • Advanced Spark & Delta Lake: Expert-level knowledge of Apache Spark (PySpark) and the Delta Lake storage format, including optimization techniques like V-Order, Z-Order indexing, and liquid clustering.
  • Languages: Expert in SQL and Python (including libraries like Pandas, Polars, or PySpark).
  • Orchestration & DevOps: Experience with Fabric Git integration, Azure DevOps/GitHub Actions, and Infrastructure as Code (Terraform) for Fabric resource provisioning.

About the company

Summit School Services companies share a strong commitment to provide the highest level of transportation safety, quality transportation, outstanding customer service and positive employee relations. Our corporate headquarters, located in Warrenville, Illinois, houses the administrative and corporate support functions for the organization. Our 250+ local customer service centers (CSCs) are supported by regional operations teams located throughout North America., Summit School Services has a zero-tolerance policy on conduct that is incompatible with its policies and values, including sexual exploitation and abuse, harassment, abuse of authority, and discrimination. Summit School Services is committed to promoting the protection and safeguarding of all children and passengers.

We offer medical, dental, vision, basic life insurance coverage, holiday pay, and PTO accrual. Additionally, employees are able to enroll in a retirement savings plan.

At Summit School Services our goal is to be a diverse workforce that is representative of the communities we serve. All employment decisions are based on business needs, job requirements and individual qualifications, without regard to race, color, religion, sex, pregnancy (including childbirth, lactation and related medical conditions), national origin, age, physical and mental disability, marital status, sexual orientation, gender identity, gender expression, genetic information (including characteristics and testing), military and veteran status, and any other characteristic protected by applicable law.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:46 min

Traditional data architecture before Microsoft Fabric

Dr. Alexander Wachtel Dr. Alexander Wachtel +1 · World Congress 2025

3:33 min

Refactoring data science workflows using Rapids QDF and Pandas

Paul Graham Paul Graham · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all