> Markdown version of [/jobs/ext/2060046-data-software-engineer](https://www.wearedevelopers.com/jobs/ext/2060046-data-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Software Engineer - **Company:** EPAM Systems, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Airflow, Amazon Web Services, Amazon S3, Big Data, Code Review, Continuous Integration, Information Engineering, Data Governance, Data Infrastructure, Extract Transform Load (ETL), Data Warehousing, Distributed Computing Environment, Identity and Access Management, Python (Programming Language), Object-Oriented Software Development, Performance Tuning, Data Logging, Large Language Models, Snowflake, Generative AI, Cloudformation, Pyspark, Integration Tests, AWS Glue, Data Programming, Functional Programming, Cloudwatch, Terraform, Data Pipelines - **Published:** August 14, 2026 - **Apply:** https://www.dice.com/job-detail/c87b25dd-9f2c-40b8-badc-1ed4990996d6 ## About the Role We are seeking a Senior Data Software Engineer to join a client-facing delivery team building and hardening cloud-native data pipelines on AWS as part of a data platform modernization program. The role involves ingesting and transforming large datasets with PySpark on AWS Glue and delivering curated, validated data into Snowflake, with a core focus on data quality, validation, and reconciliation for downstream analytics. This position is delivered at a Senior Consultant level with high autonomy and direct client stakeholder communication. Responsibilities Design, build, and optimize scalable batch and incremental ETL/ELT pipelines using PySpark on AWS Glue Configure Glue jobs, crawlers, triggers, connections, bookmarks, workflows, and the Glue Data Catalog Tune workers, partitioning, and shuffle behavior for cost and performance optimization Model and load curated datasets into Snowflake with staging, transformation, and publishing layers Implement automated data quality and validation frameworks, including schema/contract enforcement and null/uniqueness/referential checks Develop row-count and financial reconciliation processes, anomaly detection, and quarantine/reject handling Configure and extend Glue Data Quality (DQDL) rules per requirements Write clean, modular, testable Python with unit/integration tests and reusable libraries Integrate pipelines with AWS services such as S3, IAM, Lambda, Athena, CloudWatch, Step Functions, and Secrets Manager Instrument observability through logging, metrics, alerting, and pipeline SLA monitoring Participate in code reviews, CI/CD automation, and documentation Engage directly with client stakeholders in requirements refinement, design walkthroughs, status reporting, and act as technical advisor within the workstream Requirements 3+ years of experience with Python for production-level data engineering, including OOP and functional patterns Expertise in PySpark for distributed data processing and the DataFrame API Advanced proficiency in Snowflake, including data warehousing and staging/transformation layers Skills in AWS Glue, including job configuration, crawlers, Data Catalog, and DQDL Background in data quality engineering, including validation frameworks and reconciliation Proficiency in AWS services including S3, IAM, Lambda, Athena, and CloudWatch English proficiency at B2 level or higher Nice to have Familiarity with Generative AI / LLM concepts Knowledge of Airflow / Step Functions orchestration Familiarity with Great Expectations or similar data quality frameworks Knowledge of Terraform / CloudFormation ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [How we built an AI-powered code reviewer in 80 hours](https://www.wearedevelopers.com/videos/1511-how-we-built-an-ai-powered-code-reviewer-in-80-hours) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)