> Markdown version of [/jobs/ext/2698469-aws-glue-data-engineer](https://www.wearedevelopers.com/jobs/ext/2698469-aws-glue-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AWS Glue Data Engineer - **Company:** TUPPL Technology Inc - **Location:** Fort Mill, SC, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon S3, Business Analytics Applications, Data Analysis, Apache HTTP Server, Application Frameworks, Batch Processing, Cloud Computing, Program Optimization, Information Systems, Continuous Integration, Data as a Services, Data Validation, Data Discovery, Information Engineering, Data Governance, Data Infrastructure, Extract Transform Load (ETL), Data Masking, Data Security, Data Transformation Services, DevOps, Github, Identity and Access Management, Python (Programming Language), PostgreSQL, Meta-Data Management, Performance Tuning, Role-Based Access Control, Amazon Simple Notification Service (SNS), SQL Databases, Systems Integration, Enterprise Data Management, Automated Data Processing (ADP), Data Processing, Data Ingestion, Delivery Pipeline, Snowflake, Apache Spark, Change Data Capture, Infrastructure as Code (IaC), Cloudformation, Data Layers, Data Lakes, Pyspark, Information Technology, Data Lineage, AWS Glue, AWS Data Analytics, Apache Kafka, Data Management, Cloud Migration, Cloudwatch, Terraform, Software Version Control, Data Pipelines, Amazon Elastic Mapreduce (EMR), Jenkins, Amazon Redshift - **Published:** September 3, 2026 - **Apply:** https://www.dice.com/job-detail/7a1d645b-b139-4634-aea3-fdea40fe7d96 ## About the Role * 12+ years of experience in Data Engineering, with at least 5+ years designing and implementing AWS cloud-based data platforms and enterprise-scale data lakes. * Strong hands-on expertise in AWS Glue, AWS DMS, Amazon S3, Amazon Redshift, Athena, Lambda, Step Functions, EventBridge, CloudWatch, SNS, and Glue Data Catalog. * Advanced proficiency in PySpark, Spark, Python, and SQL, with experience building reusable frameworks, ETL/ELT pipelines, and high-volume data transformation solutions. * Hands-on experience designing and implementing incremental and Change Data Capture (CDC) pipelines using AWS DMS, Glue, and related AWS services. * Strong experience with Apache Iceberg including partitioning strategies, compaction, schema evolution, metadata management, performance optimization, and Iceberg catalog implementation. * Extensive experience building and supporting enterprise Data Lake architectures using Bronze, Silver, and Gold data layers with strong focus on scalability, reliability, governance, and self-service analytics. * Experience implementing and orchestrating complex workflows using AWS Step Functions, Glue Workflows, Lambda, and EventBridge. * Strong knowledge of data modeling, including dimensional, star, snowflake, and lakehouse modeling techniques. * Experience optimizing AWS Glue and Spark workloads, including partitioning, job bookmarks, worker sizing, Spark configuration tuning, and cost optimization. * Hands-on experience with real-time and batch data processing architectures integrating Kafka, Kinesis, PostgreSQL, S3, and Redshift. * Experience implementing data quality, reconciliation, observability, lineage, monitoring, and alerting frameworks across enterprise data platforms. * Strong understanding of data security and governance, including Lake Formation, IAM, RBAC, encryption, masking, auditing, retention policies, and regulatory compliance requirements. * Experience with CI/CD and Infrastructure as Code (IaC) using Terraform, CloudFormation, AWS CodePipeline, CodeBuild, GitHub, Jenkins, or similar technologies. * Strong analytical, troubleshooting, and root cause analysis skills for resolving complex production issues across Spark, Glue, Iceberg, data pipelines, and cloud infrastructure. * Experience working in Agile/Scrum environments and collaborating with architecture, governance, DevOps, and business stakeholders., * Experience in Financial Services, Wealth Management, Brokerage, Capital Markets, or BFSI domains, preferably supporting regulatory and governed data environments. * AWS Certified Data Engineer - Associate certification required/preferred; additional AWS certifications in Analytics, Data Engineering, or Solutions Architecture are highly desirable. * Hands-on experience with AWS Glue, PySpark, Apache Iceberg, Lake Formation, Kafka, Kinesis, and Amazon EMR/Spark in large-scale enterprise implementations. * Experience building metadata-driven ingestion and ETL frameworks and platform accelerators for reusable data engineering patterns. * Experience implementing data cataloging, lineage, observability, and governance solutions using enterprise data management tools. * Familiarity with modern Lakehouse architectures, cloud migration initiatives, and data platform modernization programs. * Experience with Terraform, CloudFormation, GitHub Actions, Jenkins, CodePipeline, and DevOps practices for enterprise data platforms. * Master's degree in Computer Science, Information Systems, Engineering, or related discipline preferred. Education * Bachelor's degree in Computer Science, Information Technology, Engineering, or related field. * Master's degree preferred. ## Description We are seeking a highly skilled AWS Data Engineer to design, develop, and optimize large-scale data pipelines and ETL workflows on AWS. The ideal candidate will have strong expertise in AWS cloud-native data services, data modeling, and pipeline orchestration, with hands-on experience building robust and scalable data solutions for enterprise environments., * Design and implement incremental and CDC (Change Data Capture) data pipelines using AWS Glue, DMS, and Iceberg to support near real-time analytics. * Develop and maintain metadata-driven ETL frameworks to improve reusability, scalability, and operational efficiency across data platforms. * Create and manage partitioning, compaction, and optimization strategies for Iceberg datasets to reduce query latency and storage costs. * Build and orchestrate complex workflows using AWS Step Functions, EventBridge, Lambda, and Glue Workflows for automated data processing. * Perform performance tuning and cost optimization of AWS Glue jobs by optimizing Spark configurations, worker types, partitioning, and job bookmarks. * Implement CI/CD pipelines for data engineering solutions using AWS CodePipeline, CodeBuild, GitHub, Jenkins, or Terraform. * Develop and maintain data lake architecture following AWS best practices, ensuring scalability, reliability, and governance. * Automate data validation and reconciliation processes to ensure data accuracy, completeness, and consistency across multiple systems. * Create and maintain Athena external tables, Iceberg catalogs, and Glue Data Catalog metadata for efficient data discovery and querying. * Design and implement role-based access controls (RBAC), data masking, encryption, and audit mechanisms using Lake Formation and IAM policies. * Support real-time and batch processing architectures integrating Kafka, Kinesis, PostgreSQL, S3, and Redshift. * Monitor data pipelines using CloudWatch, SNS, AWS Glue Monitoring, and custom alerting mechanisms to ensure SLA compliance. * Work closely with enterprise architecture and governance teams to establish data standards, retention policies, and compliance frameworks. * Perform root cause analysis and resolve complex production issues involving Spark, Glue, Iceberg metadata, PostgreSQL connectivity, and permission models. * Enable self-service analytics by creating curated gold, silver, and bronze data layers within enterprise data lakes. * Manage schema evolution and version control for Iceberg datasets while maintaining backward compatibility for downstream consumers. * Develop reusable PySpark utilities, frameworks, and common libraries to standardize data ingestion and transformation patterns. * Participate in architecture reviews and recommend best practices for data lake modernization, cloud migration, and platform optimization initiatives. * Implement data lineage, cataloging, and observability solutions to improve data trust, discoverability, and governance. * Collaborate with DevOps and Infrastructure teams to provision and manage AWS resources using Terraform, CloudFormation, or Infrastructure as Code (IaC) methodologies ## Related Videos - [The Road to MLOps: How Verivox Transitioned to AWS](https://www.wearedevelopers.com/videos/1050-the-road-to-mlops-how-verivox-transitioned-to-aws) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Our GitOps approach for deploying an Identity Provider and an API Gateway in a SaaS company](https://www.wearedevelopers.com/videos/776-our-gitops-approach-for-deploying-an-identity-provider-and-an-api-gateway-in-a-saas-company) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know)