> Markdown version of [/jobs/ext/638164-remote-cloud-data-engineer-must-have-experience-with-pyspark](https://www.wearedevelopers.com/jobs/ext/638164-remote-cloud-data-engineer-must-have-experience-with-pyspark). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Remote Cloud Data Engineer (Must have experience with Pyspark) - **Company:** ADVANTECH INC - **Location:** Fort Meade, MD, United States (Remote available) - **Experience:** Expert - **Salary:** $12,000.0 - **Contract:** Permanent contract - **Skills:** Agile Methodology, Amazon Web Services, Amazon S3, Microsoft Azure, Big Data, Cloud Database, Cloud Engineering, Information Systems, Databases, Continuous Delivery, Continuous Integration, Data as a Services, Data Governance, Extract Transform Load (ETL), Data Migration, Data Security, Database Queries, DevOps, Apache Hive, JSON, Python (Programming Language), Key Management, Cisco Nexus Switches, Oracle Databases, Performance Tuning, Query Optimization, Role-Based Access Control, Software Engineering, SQL Databases, Systems Integration, Unstructured Data, Data Logging, Data Ingestion, Azure Data Factory, Change Data Capture, SC Clearance, Pyspark, Information Technology, Deployment Automation, AWS Glue, Restful APIs, Azure Synapse Analytics, Data Pipelines, Amazon Redshift, Databricks - **Published:** June 25, 2026 - **Apply:** https://www.clearancejobs.com/jobs/8996175/remote-cloud-data-engineer-must-have-experience-with-pyspark ## About the Role Active Secret Clearance required Bachelor's degree in Computer Science, Data Science, Engineering, Information Systems, or related technical field 5+ years of recent experience designing and operating scalable, production-grade data pipelines using Azure Synapse Analytics and/or Databricks Strong hands-on experience with PySpark and Spark SQL for large-scale transformations and optimization Advanced proficiency in Python and SQL for data querying, automation, and analysis Experience ingesting and integrating structured and unstructured datasets from databases, flat files, REST APIs, and external systems Experience implementing full and incremental load strategies including CDC concepts and rerunnable pipeline architectures Experience with data quality controls including schema validation, enforcement, and schema drift handling Experience with pipeline monitoring, logging, alerting, and operational support Experience implementing CI/CD pipelines and automated deployment processes Knowledge of data governance and secure access management concepts including RBAC and credential management Hands-on experience with Azure and/or AWS cloud data services such as Azure Synapse, Azure Data Factory, Databricks, AWS Glue, Redshift, and S3 ## Description Advantech GS Enterprises is seeking a highly skilled Data Cloud Engineer to support the DISA NEXUS program at Fort Meade. This role will focus on designing, building, and maintaining scalable cloud-based data solutions supporting enterprise modernization and mission-critical analytics within secure DoD environments. The ideal candidate will have strong experience developing production-grade data pipelines within Azure and/or AWS cloud environments, with expertise in PySpark, Spark SQL, Python, and modern data engineering best practices. This position offers the opportunity to contribute to a long-term federal modernization effort centered around multi-cloud integration, secure data architecture, and advanced analytics capabilities., Design, develop, and maintain scalable cloud-based ETL/ELT pipelines using Azure Synapse Analytics, Databricks, AWS Glue, and related technologies Build and optimize large-scale data transformations using PySpark and Spark SQL, applying best practices for partitioning, query optimization, and performance tuning Develop and support data ingestion frameworks for both structured data (relational tables, CSV files) and unstructured data (JSON, nested structures, REST API integrations) Implement full and incremental data loading strategies, including change data capture (CDC), late-arriving record handling, and rerunnable pipelines Design and maintain cloud-based data lakes, warehouses, and analytics-ready datasets supporting enterprise reporting and operational decision-making Implement data quality and governance controls including schema validation, schema enforcement, schema drift handling, RBAC, lineage, cataloging, and credential management Monitor and troubleshoot pipelines for latency, failures, logging, alerting, and operational reliability Support CI/CD pipeline implementation for automated deployments, rollback strategies, and environment promotion processes Collaborate with cybersecurity and cloud engineering teams to ensure compliance with RMF, STIG, FedRAMP, and DoD security standards Utilize Oracle databases and cloud-native tools to support data migration, integration, and modernization initiatives Support Agile development efforts and collaborate with DevOps and software engineering teams across the program lifecycle ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)