Senior GCP Data Engineer

Capgemini
United States
8 days ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$103,646.0 - $161,949.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) JavaScript (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Airflow Data Analysis ARM Architecture Automation of Tests Big Data BigQuery Software as a Service Cloud Computing
+33 more
Code Generation Code Review Computer Programming Data Architecture Information Engineering Data Governance Data Infrastructure Data Masking Data Warehousing Relational Databases Distributed Computing Environment Data Flow Control Apache Hadoop Apache Hive Identity and Access Management Python (Programming Language) Machine Learning NoSQL Standard Sql Cloudera SAP HANA SQL Databases Cloud Platform System System Availability Apache Spark Build Server Data Lineage Production Code Star Schema Data Management Terraform Data Pipelines Apache Beam

Job description

We are seeking a highly skilled Senior Data Engineer to lead the architectural design and hands-on implementation of our next-generation data platform on Cloud Platform (GCP). The ideal candidate is a subject matter expert in BigQuery and possesses deep experience building robust, scalable data pipelines using Dataform, Dataflow, and Dataproc.

In this role, you will bridge the gap between high-level architectural strategy and technical execution, ensuring our data ecosystem is performant, cost-effective, and capable of supporting advanced analytics and machine learning initiatives., * Design and implement end-to-end data architectures on GCP, ensuring high availability, security, and scalability.

  • Establish best practices for data modeling (Medallion architecture, Star Schema) specifically optimized for BigQuery’s columnar storage.
  • Evaluate emerging GCP technologies and provide strategic recommendations for platform evolution.
  • Mentor junior engineers and conduct rigorous code reviews to maintain high engineering standards.

Data Pipeline Engineering

  • Develop complex ELT workflows using Dataform to manage SQL-based transformations, ensuring data lineage and quality through automated testing.
  • Build and maintain large-scale batch and streaming data processing pipelines using Dataflow (Apache Beam).
  • Preferred to have Dataproc knowledge for distributed processing of massive datasets using Spark or Hadoop where specialized compute is required.
  • Integrate data from various sources (SaaS APIs, RDBMS, NoSQL) into the central BigQuery data warehouse.

Performance & Governance

  • Optimize BigQuery performance through partitioning, clustering, and materialized views while managing cost through efficient slot utilization.
  • Implement CI/CD pipelines for data infrastructure and code deployments using Cloud Build or similar tools.
  • Ensure data governance standards are met by implementing Identity and Access Management (IAM), data masking, and encryption.

Technical Requirements

  • Data Warehousing Advanced BigQuery (Storage optimization, BQML, Analytics Functions)

  • Transformation Expert-level Dataform (SQLX, JavaScript, Project Configuration), understand Cortex CDC and Merge logical code logic and flow

  • Stream & Batch Hands-on Dataflow (Apache Beam in Java or Python)

  • Big Data Compute Dataproc (Spark, Hive, Hadoop ecosystem)

  • Programming Proficiency in Python, Java, or Go; Expert SQL skills

  • Orchestration Dataform code generation and custom namespace configuration. Cloud Composer (Airflow) or Vertex AI Pipelines

  • Infrastructure Terraform for GCP Resource Provisioning (IaC)

Requirements

  • Experience: data engineering, with at least 3 years focused heavily on Cloud Platform.
  • Architecture: Proven track record of designing data platforms from the ground up that support both BI and Data Science workloads.
  • Hands-on Skills: Ability to write clean, maintainable, and production-ready code.
  • Soft Skills: Excellent communication skills with the ability to explain complex technical concepts to non-technical stakeholders.

Preferred Certifications

  • Professional Data Engineer
  • Professional Cloud Architect
  • Knowledge of SAP Data model and SAP HANA required / preferred.

Benefits & conditions

Pulled from the full job description Retirement plan Vision insurance Dental insurance, The pay range that the employer in good faith reasonably expects to pay for this position is $49.83/hour - $77.86/hour. Our offered benefits include medical, dental, vision and retirement benefits. Applications will be accepted on an ongoing basis.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:36 min

Building a research database prototype from scratch

Markus Dreseler · WWC 2023

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

3:04 min

Identifying mapping overheads within in-memory database clusters

Markus Kett Markus Kett · WWC 2023

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · WWC 2024

Videos

See all

Related articles

See all