> Markdown version of [/jobs/ext/3010408-data-engineer-with-gcp-and-pyspark](https://www.wearedevelopers.com/jobs/ext/3010408-data-engineer-with-gcp-and-pyspark). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer with GCP and Pyspark - **Company:** ApTask - **Location:** Charlotte, NC, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Audit Trail, Big Data, BigQuery, Code Review, Continuous Integration, Information Engineering, Data Governance, Data Transformation, Database Queries, Python (Programming Language), Machine Learning, Meta-Data Management, Operational Data Store, Cloudera, Simple Data Format, Data Streaming, Parquet, Apache Spark, Pyspark, Machine Learning Operations, Software Version Control, Data Pipelines - **Published:** September 20, 2026 - **Apply:** https://www2.jobdiva.com/portal/?a=4ujdnwqsdebu7m13em5f0pt5dw80o500d7dv9cbq5ebzngb7yk0n43mjtefnbx0d&compid=0/jobs/29326139#/jobs/29326139 ## About the Role * 3+ years (L3) / 6+ years (L4) hands-on data engineering in production environments. * Strong SQL skills and proficiency building data transformations at scale. * Practical experience with Spark (preferably PySpark) and big data file formats (Parquet; familiarity with Iceberg concepts). * Experience delivering reliable, scheduled pipelines with monitoring, alerting, and incident response. * Understanding of data cataloging, metadata, and lineage concepts, and why they matter for governance and reuse. * Exposure to GCP data services (e.g., Dataproc/Spark, BigQuery). Deep infra setup knowledge is not required; collaboration with central cloud teams is expected. * Solid software engineering practices: version control, CI/CD, testing, code reviews, and documentation. Preferred: * Experience building reusable pipeline frameworks or platform components adopted by multiple teams. * Familiarity with Starburst/Trino for federated query across hybrid environments. * Knowledge of data quality frameworks, reconciliation, and auditability in regulated or enterprise settings. * Python proficiency; Java experience a plus (willingness to work primarily in Python). * Background in operational data environments with time-bound SLAs and prod support rotations. * Exposure to AI/ML model lifecycle platforms or MLOps concepts (platform enablement rather than model authoring). Soft Skills * High ownership, bias for action, and clear communication with technical and non-technical partners. * Comfort operating in a fast-scaling organization with evolving scope. has context menu ## Description * Design, build, and harden production-grade data pipelines for ingest, curation, reconciliation, DQ monitoring, lineage, and notifications. * Contribute to and extend the "generic pipeline" framework so federated teams can onboard data consistently and at speed. * Engineer data transformations primarily using PySpark on GCP (Dataproc/Sparkflow) and BigQuery as the main query engine. * Implement best practices for schema management, performance, reliability, and cost efficiency across large-scale datasets (Parquet, Iceberg). * Integrate with centralized data cataloging and governance (e.g., Dataplex/Knowledge Catalog or equivalent), and support automated lineage harvesting. * Collaborate with platform, security, and central cloud teams; support hybrid access patterns using Starburst where needed. * Operate with an ownership mindset in an on-time, production SLAs environment, ensuring resilient, observable, and recoverable data flows. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Parquet, Delta, Iceberg & Ducklake - An introduction for developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)