Google Cloud Data Architect - IAM Data Modernization

VYTWO TECHNOLOGIES INC.
Dallas, TX, United States
11 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Clean Code Principles Sql Data Warehouse Application Programming Interfaces (APIs) Airflow Data Analysis Apache HTTP Server Build Automation Big Data BigQuery Cloud Computing Cloud Engineering Cloud Storage
+73 more
Cluster Analysis Configuration Management Cyber Security Information Systems Computer Programming Information Technology Consulting Databases Continuous Delivery Continuous Integration Data Architecture Data Deduplication Information Engineering Data Files Data Governance Extract Transform Load (ETL) Data Transformation Data Migration Data Security Data Systems Data Warehousing Database Design DevOps Distributed File Systems File Systems Distributed Computing Environment Distributed Systems Data Flow Control Apache Hadoop Hadoop Distributed File System MapReduce Apache Hive Identity and Access Management Python (Programming Language) Metadata Meta-Data Management OpenShift Performance Tuning Scrum Methodology Systems Development Life Cycle Query Optimization Cloudera Runbook Server Administration Software Engineering SQL Databases Sqoop Data Streaming Systems Integration Management of Software Versions Parquet Enterprise Application Integration Data Logging Data Processing Scripting Google Cloud Data Ingestion System Availability Apache Spark Apache Pig Git Data Layers Containerization Data Lakes Pyspark Information Technology Data Lineage Deployment Automation Avro Data Analytics Data Management Software Version Control Data Pipelines Apache Beam

Requirements

*Must be a US Citizen/ GC only About Position: Identity & Access Management (IAM) Data Modernization - migration of an on-premises SQL data warehouse to a target-state Data Lake on Google Cloud (GCP), enabling metrics & reporting, advanced analytics, and GenAI use cases (natural language querying, accelerated summarization, cross-domain trend analysis) leveraging PySpark-based processing, cloud-native DevOps CI/CD pipelines, and containerized deployments on OpenShift (OCP) to deliver scalable, secure, and high-performance data solutions. What You’ll Do: DevOps / CI-CD

  • Experience implementing CI/CD pipelines for data and analytics workloads
  • Familiarity with Git-based source control, build automation, and deployment strategies

Containers & Platform

  • Experience with OpenShift Container Platform (OCP) for deploying data workloads and services
  • Understanding of containerized architecture, scaling, and environment management
  • Proven ability to build CI/CD pipelines for data and infrastructure workloads
  • Experience managing secrets securely using GCP Secret Manager
  • Ownership of observability, SLOs, dashboards, alerts, and runbooks
  • Proficiency in logging, monitoring, and alerting for data pipelines and platform reliability

Big Data & Processing

  • Hands-on experience with PySpark for ETL/ELT, data transformation, and performance optimization
  • Solid understanding of distributed data processing concepts

Data & Cloud Architecture

  • Strong experience designing data platforms on Google Cloud Platform (GCP)
  • Experience with Data Lakes, data warehousing, and large-scale migration programs

Data Lake Architecture & Storage

  • Proven experience designing and implementing data lake architectures (e.g., Bronze/Silver/Gold or layered models).
  • Strong knowledge of Cloud Storage (GCS) design, including bucket layout, naming conventions, lifecycle policies, and access controls

· Experience with Hadoop/HDFS architecture, distributed file systems, and data locality principles

  • Hands-on experience with columnar data formats (Parquet, Avro, ORC) and compression techniques
  • Expertise in partitioning strategies, backfills, and large-scale data organization
  • Ability to design data models optimized for analytics and BI consumption

Data Ingestion & Orchestration · Experience building batch and streaming ingestion pipelines using GCP-native services · Knowledge of Pub/Sub-based streaming architectures, event schema design, and versioning · Strong understanding of incremental ingestion and CDC patterns, including idempotency and deduplication · Hands-on experience with workflow orchestration tools (Cloud Composer / Airflow) · Ability to design robust error handling, replay, and backfill mechanisms Data Processing & Transformation · Experience developing scalable batch and streaming pipelines using Dataflow (Apache Beam) and/or Spark (Dataproc) · Strong proficiency in BigQuery SQL, including query optimization, partitioning, clustering, and cost control. · Hands-on experience with Hadoop MapReduce and ecosystem tools (Hive, Pig, Sqoop) · Advanced Python programming skills for data engineering, including testing and maintainable code design · Experience managing schema evolution while minimizing downstream impact Analytics & Data Serving · Expertise in BigQuery performance optimization and data serving patterns · Experience building semantic layers and governed metrics for consistent analytics · Familiarity with BI integration, access controls, and dashboard standards · Understanding of data exposure patterns via views, APIs, or curated datasets Data Governance, Quality & Metadata · Experience implementing data catalogs, metadata management, and ownership models · Understanding of data lineage for auditability and troubleshooting · Strong focus on data quality frameworks, including validation, freshness checks, and alerting · Experience defining and enforcing data contracts, schemas, and SLAs Good to have Security, Privacy & Compliance · Hands-on experience implementing fine-grained access controls for BigQuery and GCS · Experience with Sprint planning and helping team technically. · Strong stakeholder communication and solution-architecture skills Expertise You’ll Bring:

  • Experience: [10-14]+ years in DevOps and Data Architecture, 5+ years designing on Pyspark/GCP/OCP at scale; prior on-prem cloud migration a must.
  • Education: Bachelor’s/Master’s in Computer Science, Information Systems, or equivalent experience.
  • Certifications:Google Cloud Professional Cloud Architect/DevOps/OCP (required or within 3 months). Plus: Professional Data Engineer, Security Engineer, Access Control, Apache, Apache Avro, Apache Hadoop, Apache Hive, Apache Pig, Apache Sqoop, Application Programming Interface (API), Automation, Big Data, Business Intelligence, Business Services, Business Transformation, Centers for Disease Control and Prevention (CDC), Cloud Architecture, Cloud Computing, Cloud Storage, Communication Skills, Computer Programming, Computer Science, Consulting, Continuous Deployment/Delivery, Continuous Integration, Cost Control, Data Analysis, Data Formats, Data Lake, Data Management, Data Migration, Data Modeling, Data Processing, Data Quality, Data Sets, Data Warehousing, DataArchitect Data Modeling Tool, Database Design, Database Extract Transform and Load (ETL), DevOps, Distributed Computing, Ecosystems, Enterprise Application Integration (EAI), Environmental Management, Error Handling, File Systems, GCP (Good Clinical Practices), Git, HDFS (Hadoop Distributed File System), Identify Issues, Identity Data Management, Information Technology & Information Systems, Information Technology Consulting, Information/Data Security (InfoSec), Management Strategy, MapReduce, Metadata, Metrics, Performance Tuning/Optimization, Python Programming/Scripting Language, Query Optimization, Reporting Dashboards, SQL (Structured Query Language), Scalable System Development, Service Level Agreement (SLA), Software Administration, Software Development, Software Engineering, Source Code/Configuration Management (SCM), Sprint Planning, System Integration (SI), Testing, Trend Analysis, United States Citizen, Use Cases, Validation Testing, eBusiness

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

3:02 min

Audience Q&A on data formats and engine tradeoffs

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:52 min

Customizing block storage tiers and formats

Ricardo Sueiras Sueiras · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all