AWS Lakehouse Data Engineer

Guidehouse Inc.
Boulder, CO, United States
about 1 month ago
Apply on guidehouse.wd1.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$113,000.0 - $188,000.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Airflow Amazon Web Services Amazon S3 Data Analysis Apache HTTP Server Audit Trail Automation of Tests Continuous Integration Information Engineering Data Governance
+62 more
Data Infrastructure Extract Transform Load (ETL) Data Security Data Visualization Relational Databases Cursor DevOps Distributed Data Store Github Graph Database Identity and Access Management Python (Programming Language) Key Management Network Security Machine Learning Metadata Meta-Data Management Metadata Repositories Operational Data Store Office Suite Performance Tuning Systems Development Life Cycle Query Optimization Role-Based Access Control Power BI Search Technologies SQL Databases Data Streaming Tableau (Software) Parquet AWS Cdk Data Logging Cloud Platform System Data Classification Sql Optimization GitHub Copilot Delivery Pipeline AWS Lambda Change Data Capture Infrastructure as Code (IaC) Git Cloudformation SC Clearance Data Layers Database Migration Data Lakes Pyspark Information Technology Deployment Automation AWS Glue Maintaining Code Data Management Tools for Reporting Terraform GPT Data Pipelines Amazon Elastic Mapreduce (EMR) Docker Jenkins Custom Reports Amazon Redshift Databricks

Job description

We are seeking an AWS Lakehouse Data Engineer to design, implement, and operate the cloud-native data platform that powers AI/ML, analytics, reporting, and data visualization. You will build a modern lakehouse on Amazon S3 using AWS-native services and open table formats, providing Databricks-like capabilities while maintaining portability, strong governance, cost efficiency, and operational control. You will also develop scalable batch and streaming ingestion, Python and PySpark ETL/ELT pipelines, metadata and governance services, and automated cloud provisioning and CI/CD across environments.

This role is ideal for an engineer who enjoys platform building, automation, performance optimization, and enabling advanced analytics through trusted, secure, and well-governed data.

What You Will Do

Build and Operate Data Pipelines (Batch and Streaming)

  • Design and implement batch and streaming ingestion from APIs, relational databases, file drops, event streams, and external partners.

  • Implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce curated, analytics-ready datasets for reporting, visualization, and machine learning.

  • Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.

  • Improve pipeline reliability through automated testing, orchestration, monitoring, retry handling, and operational runbooks.

Deliver an AWS-Native Lakehouse Data Platform

  • Design and implement a Delta Lakehouse-style data platform using AWS-native services to provide Databricks-like capabilities for data engineering, analysis, and data visualization.

  • Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.

  • Implement SQL-like table reliability for data stored in Amazon S3, including ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities using Apache Iceberg.

  • Enable fast, interactive querying of lakehouse data using AWS-native query and compute services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift where appropriate.

  • Optimize performance and cost through partitioning, compaction, file sizing, statistics, caching, lifecycle policies, and efficient separation of compute and storage.

  • Establish standardized development, test, and production environments with consistent configuration and controlled promotion across stages.

Metadata, Governance, Access Control, Lineage, and Quality

  • Implement data governance and fine-grained access control using AWS-native services, including AWS Lake Formation, AWS Glue Data Catalog, AWS Identity and Access Management (IAM), AWS Key Management Service (KMS), and related security services.

  • Implement a managed metadata repository for dataset cataloging, ownership, business definitions, tagging, classification, and discoverability.

  • Enable end-to-end lineage from source through transformation and consumption to support auditability, impact analysis, and regulatory requirements.

  • Apply policy-based access, least-privilege permissions, row-, column-, and cell-level controls where required, data classification, retention, encryption, and secure data handling.

  • Build operational data quality checks for freshness, completeness, uniqueness, validity, consistency, and anomaly detection, and publish measurable SLAs/SLOs.

AWS Automation, CI/CD, and Operations

  • Implement automated AWS provisioning using Infrastructure as Code (IaC) to create consistent environments and secure-by-default baselines.

  • Build and enhance CI/CD for data pipelines and lakehouse components, including automated tests, security checks, validation gates, packaging, deployment, promotion, and rollback strategies.

  • Implement observability with centralized metrics, logs, traces, alerts, dashboards, runbooks, and incident-response procedures.

  • Continuously evaluate platform performance, scalability, reliability, security, and cost, and implement measurable improvements.

Cross-Team Collaboration and Documentation

  • Work closely with data, application, analytics, AI/ML, security, networking, and cloud platform teams to support mission needs and delivery timelines.

  • Maintain high-quality engineering documentation, including architecture diagrams, data models, SOPs, interface specifications, operational runbooks, and secure configuration baselines.

  • Present technical findings, trade-offs, risks, and recommendations clearly to technical and non-technical stakeholders.

What You Will Need, Senior level Senior level Fintech * Payments * Productivity * Software * Automation Own the SMB sales revenue target, pipeline coverage, forecasting, CRM discipline, and close rates. Build an AI-first sales motion, establish deal stages and reporting, hire and ramp at least two account executives, and coach the team on urgency, pipeline hygiene, qualification, and deal execution. Partner with Finance on forecasting and use data to improve win rates and conversion across the sales funnel. Top Skills: AICRMHubspot Wipfli

Assistant Controller

9 Minutes Ago Remote or Hybrid 74K-100K Annually Senior level 74K-100K Annually Senior level Cloud * Fintech * Software * Business Intelligence * Consulting * Financial Services Manage financial reporting, general ledgers, balance sheets, month-end close, reporting packages, work papers, and client procedures for nonprofit organizations. Review account classifications, ensure reporting accuracy, coordinate client workflows, manage accountant assignments, and support best-practice development. The role requires nonprofit federal grant processing and reporting experience, strong accounting knowledge, analytical ability, communication skills, and the capacity to manage multiple deadlines. Top Skills: Accounting SoftwareMicrosoft Office Suite

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Requirements

  • Bachelor’s degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or FOUR (4) years equivalent practical experience in leu of degree.

  • SIX (6) years of relevant experience.

  • Hands-on experience implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services such as AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.

  • Strong experience developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling.

  • Hands-on experience with Apache Iceberg, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization.

  • Advanced SQL skills and experience supporting analytical queries, semantic layers, reporting tools, and data visualization workloads.

  • Experience implementing metadata management and governance capabilities, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls.

  • Experience with AWS security fundamentals, including IAM and least privilege, KMS encryption, secrets management, network security, logging, and secure SDLC practices.

  • Experience provisioning AWS resources using IaC and operating data platforms across multiple environments.

  • Experience building or operating CI/CD pipelines for data workflows, including testing, packaging, deployment automation, environment promotion, and rollback.

  • Ability to troubleshoot distributed data-processing workloads and optimize performance, reliability, and cost.

What Would Be Nice to Have

  • Hands-on experience with Databricks, Delta Lake, or migrating Databricks workloads to AWS-native services and Apache Iceberg.

  • Experience with AWS Step Functions, Amazon Managed Workflows for Apache Airflow (MWAA), Amazon Kinesis, AWS Database Migration Service (DMS), AWS Lambda, Amazon MSK, or similar ingestion and orchestration services.

  • Experience with modern DevOps practices and tools such as Git, Terraform, AWS CloudFormation or AWS CDK, Jenkins, AWS CodePipeline, GitHub Actions, and Docker.

  • Experience integrating lakehouse data with business intelligence and visualization tools such as Amazon QuickSight, Tableau, or Power BI.

  • Experience using AI-assisted coding tools, such as GitHub Copilot, ChatGPT, Cursor, or Kiro, to accelerate implementation while maintaining code quality, testing, review, privacy, and security controls.

  • Knowledge graph and Graph RAG experience, including graph modeling, ontology and taxonomy alignment, entity resolution, relationship extraction, and hybrid retrieval that combines graph traversal with semantic or vector search.

Benefits & conditions

Designs and operates an AWS-native lakehouse supporting analytics, reporting, visualization, and AI/ML. Builds batch and streaming ingestion pipelines using Python and PySpark, Apache Iceberg data structures, governance and metadata services, quality controls, security, observability, and CI/CD automation. Optimizes performance, scalability, reliability, and cost while collaborating with cross-functional technical teams and documenting architecture, operations, and secure configurations. The summary above was generated by AI, The annual salary range for this position is $113,000.00-$188,000.00. Compensation decisions depend on a wide range of factors, including but not limited to skill sets, experience and training, security clearances, licensure and certifications, and other business and organizational needs.

What We Offer:

Guidehouse offers a comprehensive, total rewards package that includes competitive compensation and a flexible benefits package that reflects our commitment to creating a diverse and supportive workplace.

Benefits include:

  • Medical, Rx, Dental & Vision Insurance
  • Personal and Family Sick Time & Company Paid Holidays
  • Parental Leave
  • 401(k) Retirement Plan
  • Group Term Life and Travel Assistance
  • Voluntary Life and AD&D Insurance
  • Health Savings Account, Health Care & Dependent Care Flexible Spending Accounts
  • Transit and Parking Commuter Benefits
  • Short-Term & Long-Term Disability
  • Tuition Reimbursement, Personal Development, Certifications & Learning Opportunities
  • Employee Referral Program
  • Corporate Sponsored Events & Community Outreach
  • Care.com annual membership
  • Employee Assistance Program
  • Supplemental Benefits via Corestream (Critical Care, Hospital Indemnity, Accident Insurance, Legal Assistance and ID theft protection, etc.)
  • Position may be eligible for a discretionary variable incentive bonus

About Guidehouse

Guidehouse is an Equal Opportunity Employer-Protected Veterans, Individuals with Disabilities or any other basis protected by law, ordinance, or regulation.

Guidehouse will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of applicable law or ordinance including the Fair Chance Ordinance of Los Angeles and San Francisco., 95K-140K Annually Entry level 95K-140K Annually Entry level Fintech * Payments * Productivity * Software * Automation Own the full sales cycle for high-value small-business customers, including inbound and outbound prospecting, product demonstrations, lead qualification, objection handling, deal closing, and onboarding routing. Build and refine scalable outbound and sales processes, use AI tools to improve productivity, identify customer trends, and provide structured feedback to product teams. Top Skills: Ai-Assisted Selling ToolsClaudeHubspot HoneyBook

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on guidehouse.wd1.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · World Congress 2024

5:20 min

Evaluating the valuation of Cursor and developer ecosystem lock-in

Chris Heilmann Chris Heilmann +2 · LIVE

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all