> Markdown version of [/jobs/ext/2253686-data-engineer](https://www.wearedevelopers.com/jobs/ext/2253686-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Baker Group - **Location:** Ankeny, IA, United States - **Experience:** Experienced - **Salary:** $60,000.0 - $90,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Business Systems, Information Systems, Databases, Information Engineering, Data Governance, Data Infrastructure, Extract Transform Load (ETL), Data Transformation, Data Mining, Data Structures, Data Warehousing, DevOps, Dimensional Modeling, Human Resources Information System (HRIS), Python (Programming Language), Machine Learning, Meta-Data Management, Performance Tuning, Software Tools, Cloud Services, Software Engineering, SQL Databases, Technical Data Management Systems, Transact-SQL, Workflow Management Systems, Cloud Platform System, Data Classification, Azure Data Factory, Generative AI, Git, Microsoft Fabric, Pyspark, Information Technology, Data Lineage, Star Schema, Integration Frameworks, Data Pipelines, Programming Languages - **Published:** August 26, 2026 - **Apply:** https://www.careerjet.com/job/us12c95ef6b6a0359a39932bbf9979207b/eaa ## About the Role * Bachelor's degree in Computer Science, Data Engineering, Information Systems, or other relevant quantitative field * Three to five years of experience in data engineering, ETL/ELT development, or a related field * Proficiency with SQL and database technologies for data extraction, transformation, and loading * Experience with Microsoft Fabric, Azure Data Factory, or similar cloud ETL/orchestration tools * Experience with medallion architecture and modern data warehousing patterns * Experience with a programming language such as Python, PySpark, or T-SQL for data transformation * Familiarity with data modeling techniques (dimensional modeling, star schema) * Understanding of data governance, data quality, and metadata management practices * Experience preparing data for AI/ML consumption (e.g., vector embeddings, RAG architectures) is a plus * Business acumen and understanding of construction or related industries is a plus CERTIFICATES, LICENSES, REGISTRATIONS * No specific requirements; however, relevant certifications such as Microsoft Certified: Fabric Data Engineer Associate, Azure Data Engineer Associate, or similar cloud platform certifications are a plus, * Strong analytical and troubleshooting skills with the ability to diagnose and resolve complex pipeline and data quality issues * Excellent time and project management skills with the ability to prioritize across multiple pipeline and infrastructure projects * Current with industry trends in data engineering, cloud platforms, and integration best practices * Strong communication skills with the ability to translate technical data structures for non-technical stakeholders * Team player with strong collaboration skills, particularly with the Data Scientist, Data Analysts, and business system owners * Must be able to focus on complex technical problems and work independently with minimal supervision * Ability to work in a fast-paced environment and adapt to changing business priorities * Meticulous attention to detail and commitment to producing reliable, well-documented data infrastructure ENVIRONMENTAL ADAPTABILITY * Prolonged periods of sitting at a desk and working on a computer * Must be able to lift 10 pounds occasionally * May have occasional visits to a job site which would require periods of standing, walking and/or climbing stairs ## Description The Data Engineer is responsible for designing, building, and maintaining the data pipelines and infrastructure that power Baker Group's Microsoft Fabric data warehouse, serving as the organization's single source of truth. This role owns the ingestion, transformation, and orchestration of data from disparate internal systems (ERP, HRIS, MRP and other structured data sources) into governed, reliable data products used by Data Analysts and developers to deliver insights to executive and operational teams and ensures that data is structured to support both traditional reporting and emerging AI and machine learning use cases. The Data Engineer curates and maintains core datasets spanning employees, finance, construction and manufacturing projects, and service, and partners with the Data Scientist, Data Analyst, and Software Development roles to ensure data is trustworthy, well-structured, and fit for downstream use. ESSENTIAL FUNCTIONS AND RESPONSIBILITIES The following duties are typical for this job. These are not to be constructed as exclusive or all inclusive. Other duties may be required and assigned. * Designs, builds, and maintains ETL/ELT pipelines that ingest data from enterprise systems into Microsoft Fabric. * Architects and maintains the Fabric medallion Lakehouse structure (bronze, silver, gold layers) as Baker Group's single source of truth. * Develop and implement best practices for the data infrastructure and environment (e.g. Development/Test/Production environments, Git for version control). * Owns pipeline orchestration, scheduling, and monitoring to ensure reliable, timely, and accurate data availability. * Curates and maintains core datasets across employee, finance, project, service, and manufacturing domains. * Establishes and enforces data quality, validation, and reconciliation processes across all pipelines. * Designs and manages data models, schemas, and semantic layers that support Data Analyst reporting and Data Scientist modeling work. * Defines and maintains data ontologies and canonical business definitions (for example, what constitutes a "project," "employee," or "cost code") to ensure consistent meaning across systems and consumers. * Prepares and structures data to support AI and machine learning use cases, including feature-ready datasets, retrieval-augmented generation (RAG) pipelines, and vector embedding storage. * Manages Fabric capacity planning, workspace organization, and performance optimization. * Implements data governance practices, including access controls, lineage tracking, and metadata management, consistent with Baker Group's data classification standards. * Partners with business system owners (ERP, HRIS, MRP, etc.) to understand upstream data structures and manage change impacts. * Collaborates with the Data Scientist to ensure pipeline outputs support analytical and machine learning use cases. * Collaborates with Data Analysts to ensure data products support paginated reporting, dashboards, and self-service BI needs. * Collaborates with Software Development and DevOps Teams to ensure data products support application development needs. * Coordinates with 3rd party consultants when necessary to deliver data engineering projects and augment capacity for demanding business needs. * Develops and maintains documentation for pipelines, schemas, and integration logic. * Troubleshoots and resolves pipeline failures, latency issues, and data quality incidents. * Monitors and maintains data-specific infrastructure, including Fabric capacity, pipeline orchestration tools, and monitoring/alerting systems. * Evaluates and recommends new data engineering tools, patterns, and best practices. * Stays current on emerging trends in data engineering, cloud data platforms, and integration techniques. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story)