AWS Data Engineer
Enexus Global
Austin, TX, United States
16 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source
Tech stack
Airflow
Amazon S3
Information Engineering
Data Governance
Dimensional Modeling
Operational Databases
Performance Tuning
Software Architecture
SQL Databases
Workflow Management Systems
Pyspark
Storage Technologies
+1 more
Data Pipelines
Requirements
- 5+ years of hands-on data engineering experience building production data pipelines, not just supporting or maintaining them.
- Proven experience designing and owning end-to-end data pipelines across ingestion, transformation, storage, and consumption layers.
- Strong experience with AWS-native data stack including S3, Glue (PySpark), Athena, and orchestration tools such as Airflow or Step Functions.
- Experience building automated data quality checks into pipelines, such as validation rules, reconciliation logic, and anomaly detection.
- Strong understanding of data modeling and storage design including partitioning strategies, file sizing, and handling of small files.
- Deep, hands-on data modeling expertise, with solid grounding in data modeling concepts (dimensional modeling, normalization, slowly changing dimensions, fact/dimension design) and the ability to independently apply them to design data models and schemas for new datasets, including resolving grain or structural differences between source and target systems.
- Direct, end-to-end ownership experience with data quality frameworks, defining validation rules, thresholds, and exception handling, not just operating within a framework someone else built.
- Demonstrated experience defining testing standards for data pipelines and leading their adoption across a team.
- Experience making architectural decisions and trade-offs, such as choosing between Glue, Lambda, EMR or container-based approaches based on use case.
- Solid experience with PySpark and SQL including performance tuning and optimization in distributed processing environments.
- Experience building reliable, production-grade pipelines including error handling, retries, and monitoring.
- Experience reconciling and validating data across source and target systems to identify discrepancies and drive sign-off before production release.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this roleβ¦