Data Engineer with GCP and Pyspark

ApTask
Charlotte, NC, United States
8 days ago
Apply on www2.jobdiva.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) Artificial Intelligence Audit Trail Big Data BigQuery Code Review Continuous Integration Information Engineering Data Governance Data Transformation Database Queries Python (Programming Language)
+12 more
Machine Learning Meta-Data Management Operational Data Store Cloudera Simple Data Format Data Streaming Parquet Apache Spark Pyspark Machine Learning Operations Software Version Control Data Pipelines

Job description

  • Design, build, and harden production-grade data pipelines for ingest, curation, reconciliation, DQ monitoring, lineage, and notifications.
  • Contribute to and extend the ā€œgeneric pipelineā€ framework so federated teams can onboard data consistently and at speed.
  • Engineer data transformations primarily using PySpark on GCP (Dataproc/Sparkflow) and BigQuery as the main query engine.
  • Implement best practices for schema management, performance, reliability, and cost efficiency across large-scale datasets (Parquet, Iceberg).
  • Integrate with centralized data cataloging and governance (e.g., Dataplex/Knowledge Catalog or equivalent), and support automated lineage harvesting.
  • Collaborate with platform, security, and central cloud teams; support hybrid access patterns using Starburst where needed.
  • Operate with an ownership mindset in an on-time, production SLAs environment, ensuring resilient, observable, and recoverable data flows.

Requirements

  • 3+ years (L3) / 6+ years (L4) hands-on data engineering in production environments.
  • Strong SQL skills and proficiency building data transformations at scale.
  • Practical experience with Spark (preferably PySpark) and big data file formats (Parquet; familiarity with Iceberg concepts).
  • Experience delivering reliable, scheduled pipelines with monitoring, alerting, and incident response.
  • Understanding of data cataloging, metadata, and lineage concepts, and why they matter for governance and reuse.
  • Exposure to GCP data services (e.g., Dataproc/Spark, BigQuery). Deep infra setup knowledge is not required; collaboration with central cloud teams is expected.
  • Solid software engineering practices: version control, CI/CD, testing, code reviews, and documentation.

Preferred:

  • Experience building reusable pipeline frameworks or platform components adopted by multiple teams.
  • Familiarity with Starburst/Trino for federated query across hybrid environments.
  • Knowledge of data quality frameworks, reconciliation, and auditability in regulated or enterprise settings.
  • Python proficiency; Java experience a plus (willingness to work primarily in Python).
  • Background in operational data environments with time-bound SLAs and prod support rotations.
  • Exposure to AI/ML model lifecycle platforms or MLOps concepts (platform enablement rather than model authoring).

Soft Skills

  • High ownership, bias for action, and clear communication with technical and non-technical partners.

  • Comfort operating in a fast-scaling organization with evolving scope. has context menu

About the company

ApTask is a leading global provider of workforce solutions and talent acquisition services, dedicated to shaping the future of work. As an African American-owned and Veteran-owned company, ApTask offers a comprehensive suite of services, including staffing and recruitment solutions, managed services, IT consulting, and project management. With a focus on excellence, collaboration, and innovation, ApTask provides unparalleled opportunities for professional growth and development. As a member of the ApTask team, you will have the chance to connect businesses with top-tier professionals, optimize workforce performance, and drive success across diverse industries. Join us at ApTask and be part of our mission to empower organizations to thrive while fostering a diverse and inclusive work environment.

Applicants may be required to attend interviews in person or by video conference. In addition, candidates may be required to present their current state or government issued ID during each interview.

Candidate Data Collection Disclaimer: At ApTask, we prioritize safeguarding your privacy. As part of our recruitment process, certain Personally Identifiable Information (PII) may be requested by our clients for verification and application purposes. Rest assured, we strictly adhere to confidentiality standards and comply with all relevant data protection laws. Please note that we only collect the necessary information as specified by each client and do not request sensitive details during the initial stages of recruitment.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www2.jobdiva.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy Ā· LIVE

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 Ā· LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy Ā· LIVE

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff Ā· World Congress 2026 Europe

2:10 min

Why organizations combine big data and machine learning

Ayon Roy Ā· LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph Ā· LIVE

Videos

See all

Related articles

See all