> Markdown version of [/jobs/ext/1249851-mid-aws-data-engineer-pyspark-glue-redshift](https://www.wearedevelopers.com/jobs/ext/1249851-mid-aws-data-engineer-pyspark-glue-redshift). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Mid-AWS Data Engineer (PySpark / Glue / Redshift) - **Company:** Vsg Business Solutions - **Location:** Malvern, PA, United States - **Experience:** Experienced - **Salary:** $104,000.0 - $114,400.0 - **Contract:** Temporary to permanent - **Skills:** Clean Code Principles, Adaptable Database Systems, Amazon Web Services, Amazon S3, Cloud Engineering, Information Systems, Information Engineering, Extract Transform Load (ETL), Database Queries, Distributed Systems, Apache Hive, Identity and Access Management, Python (Programming Language), SQL Databases, Tableau (Software), Data Processing, Apache Spark, State Machines, Git, Cloudformation, Pyspark, Information Technology, AWS Glue, AWS Data Analytics, Presto, Functional Programming, Data Pipelines, Amazon Redshift - **Published:** July 12, 2026 - **Apply:** https://www.careerjet.com/jobad/us2a12ddad0db4ce570e74727f8088ff83 ## About the Role AWS Core: Deep proficiency in AWS Glue, S3, Lambda, IAM, and Step Functions. Data Engineering: Expert-level knowledge of PySpark and Spark SQL. Querying & SQL: Extensive experience with Athena and Presto; strong SQL skills for complex data manipulation. Infrastructure as Code: Proven experience with CloudFormation for operational excellence. Development Mindset: Ability to write clean, maintainable code for automation and custom utility development. Dremio: Experience with Dremio implementation (reflections, views, using system tables)., 5-8 years of Data Engineering or IT experience. 4+ years of AWS cloud development experience. Expert Python programming skills. Advanced SQL development experience. Strong hands-on experience with Apache Spark / PySpark. Experience building cloud-native ETL pipelines. Strong understanding of distributed computing. Experience designing scalable data models. Experience with Git and CI/CD pipelines. Strong analytical and problem-solving skills. Build scalable and auto recoverable pipelines. Bachelor's degree in Computer Science, Engineering, Information Systems, or related discipline. Nice to Have Tableau: Familiarity with Tableau for end-user data consumption. ## Description We are seeking a talented Senior Data Engineer to join our infrastructure support team. This role is focused on designing and developing robust AWS utilities that enable our data engineers and analysts to seamlessly query across multiple data sources, specifically bridging the gap between S3-based tables and Amazon Redshift. As we transition to a multi-technology data environment, you will be responsible for building resilient, enterprise-grade solutions that prioritize reliability and quick recovery., Design & Development: Lead the design and development of custom AWS utilities that allow analysts to query disparate data sources (S3 and Redshift) with minimal effort and high productivity. Pipeline Engineering: Build and maintain scalable data pipelines using AWS Glue, PySpark, and Spark SQL to ingest and transform data across platforms. Resiliency & Scalability: Apply an "enterprise-grade" development mindset. Build solutions that are resilient, scalable, and capable of quick automated recovery in the event of failure. Utility Creation: Create tools to streamline the user experience for analysts, such as building automated processes to convert SQL outputs for multi-source tables. Infrastructure Ops: Utilize CloudFormation to manage and deploy infrastructure components. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)