Data Engineer

Smart Synergies
Bethesda, United States of America
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior

Job location

Bethesda, United States of America

Tech stack

API
Agile Methodologies
Amazon Web Services (AWS)
Amazon Web Services (AWS)
Data analysis
Automation of Tests
Business Intelligence
Code Review
Information Systems
Continuous Integration
Data Architecture
Data Validation
Data Cleansing
Information Engineering
Data Governance
Data Integration
Data Systems
Data Warehousing
Dimensional Modeling
Distributed Data Store
Github
Python
Meta-Data Management
Performance Tuning
Scrum
Power BI
Shell Script
Software Deployment
SQL Databases
Delivery Pipeline
Spark
Cloudformation
Data Lake
PySpark
Core Data
Information Technology
Data Lineage
Collibra
Star Schema
Amazon Web Services (AWS)
Bitbucket
Oracle Ebusiness
Amazon Web Services (AWS)
Terraform
Data Pipelines

Job description

We are looking for a driven and hands-on Data Engineer to design, implement, and support scalable, production-level data pipelines within our AWS-based data ecosystem. This position is responsible for full lifecycle pipeline delivery-from ingesting data from enterprise systems to producing analytics-ready datasets-while collaborating closely with Data Architects, Analytics teams, and business partners. The role is critical in ensuring reliable, high-quality data is available across the organization to power analytics and informed decision-making. As a member of the core data platform team, you will also contribute to advancing our AWS Lakehouse architecture and establishing engineering best practices., * Pipeline Engineering: Develop and maintain scalable ELT pipelines using reusable ingestion frameworks that support both batch and event-driven processing across multiple data sources such as ERPs, APIs, vendor feeds, and relational systems.

  • AWS Data Platform Development: Design and enhance AWS-native data solutions aligned with a Medallion (Bronze/Silver/Gold) Lakehouse architecture, with an emphasis on performance optimization and cost management.
  • Infrastructure & DevOps: Build and manage AWS infrastructure using Terraform, and support CI/CD processes through GitOps methodologies, including automated testing, monitoring, alerting, and system recovery capabilities.
  • Data Modeling & Transformation: Create and maintain dimensional models and Gold-layer datasets using SQL, Python, and PySpark. Implement scalable ingestion processes with strong handling of schema evolution, auditing, and performance tuning techniques such as partitioning, clustering, and materialization.
  • Data Quality & Governance: Integrate automated data validation, anomaly detection, and lineage tracking into pipelines, while contributing to metadata management practices.
  • Reporting & BI Enablement: Diagnose and resolve complex data issues, ensuring pipelines deliver data optimized for Power BI. Work closely with BI developers to align data models with reporting needs and troubleshoot dashboard-related challenges.
  • Cross-Functional Collaboration: Partner with architecture and analytics teams to translate business requirements into technical solutions, and actively participate in code reviews, sprint planning, and architectural discussions.

Requirements

  • Bachelor's degree in Computer Science, Information Systems, Engineering, or a related discipline, or equivalent practical experience.
  • At least 7 years of experience in data engineering, with strong hands-on expertise in AWS services and distributed data technologies.
  • Advanced proficiency in Python and SQL, including experience with Spark/PySpark.
  • Proven experience building and maintaining AWS-based data pipelines using services such as Glue, Step Functions, Lambda, S3, Athena, SNS, SQS, and Redshift.
  • Experience with event-driven data architectures.
  • Hands-on experience processing and managing large-scale vendor data feeds.
  • Practical knowledge of Medallion architecture within a data lake or Lakehouse environment.
  • Experience using Terraform for infrastructure-as-code deployments.
  • Familiarity with CI/CD tools such as Bitbucket, GitHub, or AWS CodePipeline for pipeline deployment.
  • Strong understanding of data warehousing principles, including star schema design, dimensional modeling, and slowly changing dimensions (SCD).
  • Experience integrating data with Power BI or comparable business intelligence tools.
  • Working knowledge of UNIX/Linux environments, including shell scripting.
  • Experience supporting production systems, including monitoring, troubleshooting, and on-call responsibilities.
  • Familiarity with Agile development methodologies., * Experience integrating data from Oracle EBS.
  • Familiarity with data quality tools such as Great Expectations or dbt testing frameworks.
  • Experience with AWS CDK or CloudFormation alongside Terraform.
  • Knowledge of data cataloging and lineage tools such as Alation or Collibra.
  • AWS certifications (e.g., AWS Certified Data Engineer, AWS Certified Solutions Architect).

Apply for this position