> Markdown version of [/jobs/ext/456169-gcp-data-engineer](https://www.wearedevelopers.com/jobs/ext/456169-gcp-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # GCP Data Engineer - **Company:** Recutify Inc. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $93,600.0 - $104,000.0 - **Contract:** Permanent contract - **Skills:** Airflow, Application Frameworks, Big Data, BigQuery, Cloud Computing, Computer Programming, Data Validation, Information Engineering, Data Governance, Data Integrity, Extract Transform Load (ETL), Data Security, Data Systems, Data Warehousing, Database Design, Data Flow Control, Apache Hadoop, Hadoop Distributed File System, Apache Hive, Python (Programming Language), Performance Tuning, Query Optimization, Cloudera, Software Construction, SQL Stored Procedures, SQL Databases, Database Engines, Data Streaming, Teradata SQL, Google Cloud, Apache Yarn, Sql Optimization, Informatica Powercenter, Event Driven Architecture, Database Migration, Data Lakes, Pyspark, Deployment Automation, Google Bigquery, Data Management, Data Pipelines, Software Library - **Published:** June 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=20490252a7282e6d ## About the Role Do you have experience in System tuning?, Professional Experience: 5+ years of professional experience in a data engineering role, with a proven track record of building and maintaining large-scale data systems. Framework Design & Development: Demonstrable experience designing, building, and promoting the adoption of reusable data engineering frameworks such as Ingestion, Transformation, Validation etc On-Premises Data Warehouse/Big Data Expertise: Deep, hands-on experience with at least one of the following on-premise ecosystems: Teradata: Strong understanding of the Teradata architecture, utilities (BTEQ, TPT, FastLoad/MultiLoad), and advanced SQL/Stored Procedure development. Hadoop: Experience with the Hadoop ecosystem (HDFS, YARN, Hive) and hands-on proficiency with PySpark for large-scale data processing. Enterprise ETL: Demonstrable experience designing and building complex workflows in an enterprise ETL tool like Informatica PowerCenter. Google Cloud Platform (GCP) Proficiency: Demonstrable hands-on experience designing, building, and operating solutions with a comprehensive set of GCP data services, including: Core Data Processing & Warehousing: Google BigQuery (including data modeling, performance tuning, cost management), Google Cloud Storage (GCS), Cloud Dataflow, and Cloud Dataproc. Orchestration & Event-Driven Architecture: Cloud Composer (Managed Airflow) for complex workflow orchestration, and Pub/Sub and Cloud Functions for building streaming and event-driven data pipelines. Data Governance & Management: Practical experience using Dataplex for unified data management, security, and governance across data lakes and warehouses. Core Engineering & Migration Skills: Expert-level proficiency in SQL, including complex joins, window functions, and performance tuning across different database engines. Strong programming skills in Python, applying software engineering best practices. Hands-on experience with Google's migration assessment tools (e.g., BigQuery Migration Service, Database Migration Service) to analyze on-premise workloads and accelerate migration. Deep understanding of data warehousing concepts, ETL/ELT patterns, data modeling, and database design. Preferred Qualifications (Nice-to-Haves): Proven Migration Experience: Direct, hands-on experience successfully completing at least one large-scale on-premise (Teradata, Hadoop, etc.) to Google Cloud Platform. Certifications: A Google Cloud Professional Data Engineer certification is a plus. ## Description Architect & Design: Design and implement robust, scalable, and cost-effective data solutions on Google Cloud, serving as the target architecture for migrated workloads. Develop Reusable Frameworks & Accelerators: Design, build, and maintain reusable frameworks, templates, and code libraries to standardize and accelerate data engineering work. This includes creating boilerplate pipeline structures, generic data validation modules, and automated deployment patterns that other engineers will leverage. Migrate & Modernize: Lead the hands-on migration of data and processes from on-premises systems like Teradata and Hadoop to Google Cloud services, with a primary focus on BigQuery, Google Cloud Storage (GCS), Dataflow, and Dataproc. ETL/ELT Transformation: Analyze, deconstruct, and translate complex legacy ETL logic from tools like Informatica and Teradata BTEQ/Stored Procedures into modern, cloud-native pipelines, leveraging the frameworks/tooling you help create. Pipeline Development: Build and automate new data pipelines for batch and streaming data using Python, SQL, and GCP's core services, ensuring all new development contributes to and benefits from our shared engineering frameworks. Performance & Cost Optimization: Proactively optimize BigQuery performance through effective partitioning, clustering, and query tuning. Data Validation & Governance: Develop and implement rigorous data validation frameworks to ensure data integrity and accuracy post-migration. Collaborate with governance teams to apply data security, lineage, and cataloging using tools like Google Cloud Data Catalog and Dataplex. Collaboration & Mentorship: Work closely with on-premises data experts, business analysts, and other engineers to understand requirements, ensure a smooth transition, and act as a subject matter expert and mentor for GCP and internal framework best practices. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)