Data Engineer

Baker Group
Ankeny, IA, United States
11 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$60,000.0 - $90,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Business Systems Information Systems Databases Information Engineering Data Governance Data Infrastructure Extract Transform Load (ETL) Data Transformation Data Mining Data Structures Data Warehousing
+27 more
DevOps Dimensional Modeling Human Resources Information System (HRIS) Python (Programming Language) Machine Learning Meta-Data Management Performance Tuning Software Tools Cloud Services Software Engineering SQL Databases Technical Data Management Systems Transact-SQL Workflow Management Systems Cloud Platform System Data Classification Azure Data Factory Generative AI Git Microsoft Fabric Pyspark Information Technology Data Lineage Star Schema Integration Frameworks Data Pipelines Programming Languages

Job description

The Data Engineer is responsible for designing, building, and maintaining the data pipelines and infrastructure that power Baker Group’s Microsoft Fabric data warehouse, serving as the organization’s single source of truth. This role owns the ingestion, transformation, and orchestration of data from disparate internal systems (ERP, HRIS, MRP and other structured data sources) into governed, reliable data products used by Data Analysts and developers to deliver insights to executive and operational teams and ensures that data is structured to support both traditional reporting and emerging AI and machine learning use cases. The Data Engineer curates and maintains core datasets spanning employees, finance, construction and manufacturing projects, and service, and partners with the Data Scientist, Data Analyst, and Software Development roles to ensure data is trustworthy, well-structured, and fit for downstream use. ESSENTIAL FUNCTIONS AND RESPONSIBILITIES The following duties are typical for this job. These are not to be constructed as exclusive or all inclusive. Other duties may be required and assigned.

  • Designs, builds, and maintains ETL/ELT pipelines that ingest data from enterprise systems into Microsoft Fabric.
  • Architects and maintains the Fabric medallion Lakehouse structure (bronze, silver, gold layers) as Baker Group’s single source of truth.
  • Develop and implement best practices for the data infrastructure and environment (e.g. Development/Test/Production environments, Git for version control).
  • Owns pipeline orchestration, scheduling, and monitoring to ensure reliable, timely, and accurate data availability.
  • Curates and maintains core datasets across employee, finance, project, service, and manufacturing domains.
  • Establishes and enforces data quality, validation, and reconciliation processes across all pipelines.
  • Designs and manages data models, schemas, and semantic layers that support Data Analyst reporting and Data Scientist modeling work.
  • Defines and maintains data ontologies and canonical business definitions (for example, what constitutes a “project,” “employee,” or “cost code”) to ensure consistent meaning across systems and consumers.
  • Prepares and structures data to support AI and machine learning use cases, including feature-ready datasets, retrieval-augmented generation (RAG) pipelines, and vector embedding storage.
  • Manages Fabric capacity planning, workspace organization, and performance optimization.
  • Implements data governance practices, including access controls, lineage tracking, and metadata management, consistent with Baker Group’s data classification standards.
  • Partners with business system owners (ERP, HRIS, MRP, etc.) to understand upstream data structures and manage change impacts.
  • Collaborates with the Data Scientist to ensure pipeline outputs support analytical and machine learning use cases.
  • Collaborates with Data Analysts to ensure data products support paginated reporting, dashboards, and self-service BI needs.
  • Collaborates with Software Development and DevOps Teams to ensure data products support application development needs.
  • Coordinates with 3rd party consultants when necessary to deliver data engineering projects and augment capacity for demanding business needs.
  • Develops and maintains documentation for pipelines, schemas, and integration logic.
  • Troubleshoots and resolves pipeline failures, latency issues, and data quality incidents.
  • Monitors and maintains data-specific infrastructure, including Fabric capacity, pipeline orchestration tools, and monitoring/alerting systems.
  • Evaluates and recommends new data engineering tools, patterns, and best practices.
  • Stays current on emerging trends in data engineering, cloud data platforms, and integration techniques.

Requirements

  • Bachelor’s degree in Computer Science, Data Engineering, Information Systems, or other relevant quantitative field
  • Three to five years of experience in data engineering, ETL/ELT development, or a related field
  • Proficiency with SQL and database technologies for data extraction, transformation, and loading
  • Experience with Microsoft Fabric, Azure Data Factory, or similar cloud ETL/orchestration tools
  • Experience with medallion architecture and modern data warehousing patterns
  • Experience with a programming language such as Python, PySpark, or T-SQL for data transformation
  • Familiarity with data modeling techniques (dimensional modeling, star schema)
  • Understanding of data governance, data quality, and metadata management practices
  • Experience preparing data for AI/ML consumption (e.g., vector embeddings, RAG architectures) is a plus
  • Business acumen and understanding of construction or related industries is a plus

CERTIFICATES, LICENSES, REGISTRATIONS

  • No specific requirements; however, relevant certifications such as Microsoft Certified: Fabric Data Engineer Associate, Azure Data Engineer Associate, or similar cloud platform certifications are a plus, * Strong analytical and troubleshooting skills with the ability to diagnose and resolve complex pipeline and data quality issues
  • Excellent time and project management skills with the ability to prioritize across multiple pipeline and infrastructure projects
  • Current with industry trends in data engineering, cloud platforms, and integration best practices
  • Strong communication skills with the ability to translate technical data structures for non-technical stakeholders
  • Team player with strong collaboration skills, particularly with the Data Scientist, Data Analysts, and business system owners
  • Must be able to focus on complex technical problems and work independently with minimal supervision
  • Ability to work in a fast-paced environment and adapt to changing business priorities
  • Meticulous attention to detail and commitment to producing reliable, well-documented data infrastructure

ENVIRONMENTAL ADAPTABILITY

  • Prolonged periods of sitting at a desk and working on a computer
  • Must be able to lift 10 pounds occasionally
  • May have occasional visits to a job site which would require periods of standing, walking and/or climbing stairs

About the company

Baker Group is an Equal Opportunity Employer. In compliance with the Americans with Disabilities Act, Baker Group will consider reasonable accommodations for qualified individuals with disabilities and encourage prospective employees and incumbents to discuss potential accommodations with the Employer.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all