AWS Data Engineer

Enexus Global
Austin, TX, United States
16 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Airflow Amazon S3 Information Engineering Data Governance Dimensional Modeling Operational Databases Performance Tuning Software Architecture SQL Databases Workflow Management Systems Pyspark Storage Technologies
+1 more
Data Pipelines

Requirements

  1. 5+ years of hands-on data engineering experience building production data pipelines, not just supporting or maintaining them.
  2. Proven experience designing and owning end-to-end data pipelines across ingestion, transformation, storage, and consumption layers.
  3. Strong experience with AWS-native data stack including S3, Glue (PySpark), Athena, and orchestration tools such as Airflow or Step Functions.
  4. Experience building automated data quality checks into pipelines, such as validation rules, reconciliation logic, and anomaly detection.
  5. Strong understanding of data modeling and storage design including partitioning strategies, file sizing, and handling of small files.
  6. Deep, hands-on data modeling expertise, with solid grounding in data modeling concepts (dimensional modeling, normalization, slowly changing dimensions, fact/dimension design) and the ability to independently apply them to design data models and schemas for new datasets, including resolving grain or structural differences between source and target systems.
  7. Direct, end-to-end ownership experience with data quality frameworks, defining validation rules, thresholds, and exception handling, not just operating within a framework someone else built.
  8. Demonstrated experience defining testing standards for data pipelines and leading their adoption across a team.
  9. Experience making architectural decisions and trade-offs, such as choosing between Glue, Lambda, EMR or container-based approaches based on use case.
  10. Solid experience with PySpark and SQL including performance tuning and optimization in distributed processing environments.
  11. Experience building reliable, production-grade pipelines including error handling, retries, and monitoring.
  12. Experience reconciling and validating data across source and target systems to identify discrepancies and drive sign-off before production release.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Loading talks and stories from around this role…