GCP Data Engineer

AIVANTA TECHNOLOGIES LLC
New York, NY, United States
19 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$93,600.0 - $104,000.0
Working hours
Regular working hours
Job source

Tech stack

Microsoft Access Airflow Amazon S3 BigQuery Cluster Analysis Cyber Security Databases Data Fusion Data Governance Extract Transform Load (ETL) Data Masking Data Migration
+29 more
Data Warehousing IBM DB2 Data Flow Control Apache Hadoop Hadoop Distributed File System Apache Hive Identity and Access Management Python (Programming Language) Oracle (Applications) Zero Trust Network Access Cloudera Reverse Engineering SAP (Applications) SQL Databases Z/OS Parquet Google Cloud Data Ingestion Apache Spark Change Data Capture Ansi Sql Pyspark Core Data SAP S/4HANA Codebase Data Lakehouse Data Pipelines Serverless Computing Apache Beam

Job description

The Senior Cloud Data Engineer will serve as the technical anchor for the physical data migration and modern pipeline engineering required to build DLA’s GCP Data Fabric. This role will drive the critical transition of moving over 100 terabytes of uncompressed business data from distributed Cloudera Hadoop (HDFS/Hive) clusters and AWS S3/FSx data dumps into a serverless, optimized storage layer in BigQuery.

The Data Engineer will establish secure Medallion Architecture pathways (Bronze * Silver *Gold) using modern GCP data movement frameworks.

Key Roles and Responsibilities

  • BigQuery Data Lakehouse Engineering: Design, implement, and maintain scalable BigQuery data warehouses. Map legacy HDFS/Hive schemas and parquet/flat-file configurations directly into BigQuery tables.
  • ETL Reduction & SQL Modernization: Lead the massive migration effort to reverse-engineer and collapse legacy data pipelines. Analyze the 87% data-wrangling/ETL script footprint in DLA’s codebase and translate it into high-performance, native BigQuery SQL, Dataflow (Apache Beam), Dataproc (Serverless Spark), or dbt pipelines.
  • Authoritative Source Integration via SLT & Datastream: Architect secure ingestion pipelines to synchronize transactional tables directly from S/4HANA (EBS) database layers to GCP using SAP SLT via Pub/Sub and Private Google Access (PGA). Deploy Google Datastream for continuous Change Data Capture (CDC) replication from legacy non-SAP engines (Oracle 19c, IBM DB2 on z/OS).
  • Storage Performance Benchmarking & Optimization: Optimize BigQuery scan overhead and combat potential project bill shock by implementing Column Clustering and Partitioning strategies customized around target Mission Support Element (MSC) high-frequency join/filter columns. Conduct performance benchmarking on GCP Filestore deployments to satisfy the required 125 MB/s per core throughput metric for legacy file transformations.
  • Data Governance & Zero-Trust Persistence: Enforce structural file lifecycle management rules and fine-grained column-level and row-level access controls, utilizing Dataplex and BigQuery data masking parameters to satisfy stringent DoD IL5/IL6 security standards.

Requirements

  • Google Core Data Stack: Expert-level mastery of BigQuery (Dremel Engine, SQL, IAM controls), Dataflow, Dataproc, Cloud Composer (Airflow), Data Fusion, and Pub/Sub.
  • Distributed Computing Expertise: Advanced experience working with Apache Spark, Hive Metastore, PySpark pipelines, and data migration utilities (DistCp, Storage Transfer Service).
  • Languages: Superior proficiency in Python, PySpark, and advanced ANSI SQL.
  • Certifications: Google Cloud Professional Data Engineer (Highly Preferred); Must maintain DoD 8570 baseline compliance (e.g., Security+ CE)

Benefits & conditions

Pay: $45.00 - $50.00 per hour

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:03 min

Introduction to open table formats built on Parquet

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all