Data Engineer (Senior Level) - AWS & Streaming

TUPPL Technology Inc
Austin, TX, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Profiling Data Validation Information Engineering Data Governance Data Integrity Extract Transform Load (ETL) DevOps Distributed Computing Environment Python (Programming Language) Cloud Services SQL Databases
+13 more
Data Streaming Data Processing AWS Lambda Cloudformation Data Lakes Pyspark Apache Flink AWS Glue Apache Kafka Data Management Terraform Stream Processing Data Pipelines

Job description

We are seeking a Mid-Senior Data Engineer with strong expertise in AWS-based data engineering, real-time streaming technologies, and enterprise-grade data quality frameworks. The ideal candidate will design, build, and optimize scalable batch and streaming data pipelines, implement robust data validation and monitoring processes, and support mission-critical analytics platforms., * Develop and maintain scalable ETL/ELT pipelines using AWS Glue, PySpark, and Python

  • Build event-driven workflows using AWS Lambda

  • Design and manage real-time streaming solutions using Kafka, KSQL, and Apache Flink

  • Implement and enforce comprehensive data quality frameworks, including validation, profiling, monitoring, and reconciliation

  • Optimize data processing performance, scalability, reliability, and cost in cloud environments

  • Collaborate with cross-functional teams to deliver reliable, production-grade data platforms and ensure data integrity across the pipeline

Requirements

  • Strong hands-on experience with Python and PySpark

  • Proven expertise in AWS Glue, Lambda, and other cloud-native data services

  • Solid experience with the Kafka ecosystem (topics, partitions, consumer groups, streaming patterns)

  • Demonstrated experience building and supporting data quality frameworks (validation rules, reconciliation checks, profiling, anomaly detection)

  • Strong understanding of distributed data processing and scalable architecture patterns

Good-to-Have Skills:

  • Experience with Apache Flink for real-time stream processing and stateful computations

  • Knowledge of KSQL or other streaming SQL engines

  • Exposure to CI/CD pipelines, IaC (Terraform/CloudFormation), and DevOps practices

  • Familiarity with data lake/lakehouse architectures and table formats such as Iceberg, Delta, or Hudi

  • Experience working in enterprise or financial data environments

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

55 sec

Validating data processing architectures via containerized events

Modood Alvi · WWC 2025

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

1:09 min

Evaluating mature stream processing frameworks for production systems

Soroosh Khodami Soroosh Khodami · WWC 2024

Videos

See all

Related articles

See all