Data Engineer

Guidehouse Inc.
McLean, VA, United States
27 days ago
Apply on find.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Amazon S3 Automation of Tests Big Data Software Documentation Continuous Integration Data as a Services Data Architecture Data Validation Information Engineering Data Governance
+27 more
Extract Transform Load (ETL) Data Masking Data Security Distributed Data Store Identity and Access Management Interoperability Python (Programming Language) Metadata Performance Tuning Role-Based Access Control Software Deployment SQL Databases Data Streaming Enterprise Data Management Business Intelligence Development Studio Amazon Virtual Private Cloud (VPC) Data Layers Data Lakes Pyspark Data Lineage Apache Kafka Video Streaming Stream Processing Data Pipelines Devsecops Legacy Systems Databricks

Job description

  • Develop and maintain scalable data ingestion, transformation, and curation pipelines using Databricks (Delta Lake, Delta Live Tables, Auto Loader) and AWS services, supporting consistent delivery of analytics-ready datasets across the Data Platform.
  • Implement standardized batch and near Real Time data pipelines that integrate Legacy systems and cloud-native data sources, contributing to enterprise-wide data access, reuse, and platform consistency.
  • Support full data pipeline lifecycle activities, including intake, requirements analysis, source profiling, and technical implementation, aligning development work with defined intake processes, SLAs, and governance.
  • Build production-ready data pipelines that meet defined technical and documentation standards, including participation in validation, testing, and release processes to enable reliable and compliant production deployments.
  • Apply data quality checks, validation rules, and observability practices to improve pipeline reliability, support monitoring, and contribute to platform stability and operational performance targets.
  • Integrate pipelines with AWS services (eg, S3, streaming frameworks, APIs) and enterprise data tools to support secure, scalable data movement and interoperability across the ecosystem.
  • Contribute to governed data engineering practices by implementing metadata capture, lineage tracking, and supporting access control patterns aligned to enterprise data governance standards.
  • Support analytics and reporting use cases by preparing curated datasets and enabling consumption through SQL-based access, dashboards, and enterprise BI tooling.

Requirements

  • Bachelor’s degree is required
  • Minimum FOUR (4) years of experience in data engineering, with hands-on development of data pipelines in cloud or distributed data environments.
  • Strong proficiency in Python, PySpark, and SQL for building and maintaining scalable ETL/ELT pipelines.
  • Experience working with Databricks and Delta Lake to support ingestion, transformation, and curated data layer development.
  • Working knowledge of AWS data services (eg, S3, IAM, VPC) and integration patterns supporting secure and scalable data architectures.
  • Experience implementing data quality checks, monitoring, and basic performance optimization techniques for pipeline efficiency and reliability.
  • Familiarity with data governance concepts, including metadata, lineage, and access control frameworks in regulated environments.
  • Experience working within Agile delivery environments and contributing to CI/CD-enabled development workflows.

What Would Be Nice To Have:

  • Experience supporting enterprise data platforms or federal data modernization initiatives, particularly in highly regulated environments.
  • Exposure to streaming technologies such as Kafka, Kinesis, or EventBridge for near Real Time data processing.
  • Familiarity with Databricks Unity Catalog and governance capabilities (RBAC, data masking, auditing).
  • Experience using metadata/catalog tools such as Informatica EDC or similar platforms.
  • Understanding of DevSecOps practices, including automated testing, deployment, and environment promotion across dev/test/prod.
  • Exposure to performance optimization techniques (eg, partitioning, Z-ordering, clustering) for large-scale data processing workloads.

Benefits & conditions

Guidehouse offers a comprehensive, total rewards package that includes competitive compensation and a flexible benefits package that reflects our commitment to creating a diverse and supportive workplace.

Benefits include:

  • Medical, Rx, Dental & Vision Insurance
  • Personal and Family Sick Time & Company Paid Holidays
  • Position may be eligible for a discretionary variable incentive bonus
  • Parental Leave and Adoption Assistance
  • 401(k) Retirement Plan
  • Basic Life & Supplemental Life
  • Health Savings Account, Dental/Vision & Dependent Care Flexible Spending Accounts
  • Short-Term & Long-Term Disability
  • Student Loan PayDown
  • Tuition Reimbursement, Personal Development & Learning Opportunities
  • Skills Development & Certifications
  • Employee Referral Program
  • Corporate Sponsored Events & Community Outreach
  • Emergency Back-Up Childcare Program
  • Mobility Stipend

About Guidehouse

Guidehouse is an Equal Opportunity Employer-Protected Veterans, Individuals with Disabilities or any other basis protected by law, ordinance, or regulation.

Guidehouse will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of applicable law or ordinance including the Fair Chance Ordinance of Los Angeles and San Francisco.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on find.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

1:47 min

Comparing Egeria to alternative open metadata solutions

Ferd Scheepers · World Congress 2022

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all