> Markdown version of [/jobs/ext/2297332-data-engineer](https://www.wearedevelopers.com/jobs/ext/2297332-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Capgemini - **Location:** United States - **Experience:** Experienced - **Salary:** $75,000.0 - $96,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Bash Shell, Big Data, BigQuery, Cloud Computing, Cloud Storage, Information Engineering, Extract Transform Load (ETL), Data Warehousing, Query Languages, Data Flow Control, Identity and Access Management, Python (Programming Language), Machine Learning, Metadata, Cloudera, Search Technologies, Shell Script, Software Engineering, SQL Databases, Data Streaming, Unstructured Data, Data Logging, Data Processing, Google Cloud, Data Storage Technologies, Data Ingestion, Cloud Monitoring, Large Language Models, Prompt Engineering, Apache Spark, Event Driven Architecture, Data Lakes, Pyspark, Deployment Automation, Google Cloud Functions, Data Analytics, Virtual Agents, Terraform, Data Pipelines - **Published:** August 29, 2026 - **Apply:** https://www.capgemini.com/jobs/538596-en_US_SAPBTP/x/ ## About the Role * 5+ years of overall data engineering or software engineering experience * 2+ years of hands-on Google Cloud Platform experience * 2+ years of Python development * 2+ years of experience building data pipelines (batch and streaming) Preferred Qualifications * Experience with Dataproc (Spark/PySpark) for large-scale processing * Familiarity with event-driven architectures * Knowledge of Terraform or Infrastructure as Code * Understanding of cost optimization (FinOps) Nice to Have * Google Cloud Professional Data Engineer Certification * Experience supporting AI/ML data pipelines ## Description We are seeking a highly skilled GCP Data Engineer with strong Python expertise to design, build, and optimize scalable data solutions on Google Cloud Platform (GCP). The ideal candidate will have hands-on experience developing batch and real-time data pipelines, working with large-scale datasets, and enabling analytics and AI/ML use cases. Key Responsibilities Data Engineering & Pipeline Development * Develop and optimize ETL/ELT workflows for structured and unstructured data processing using GCP services such as Dataflow, Dataproc, and Pub/Sub * Implement event-driven data processing using Cloud Functions and Pub/Sub * Build and manage data ingestion frameworks for streaming and batch data sources Data Storage & Processing * Design and optimize data lakes and data warehouses using BigQuery and Cloud Storage * Develop efficient data models to support analytics, reporting, and machine learning workloads * Optimize performance and cost of data pipelines and queries Development & Automation * Develop solutions using Python * Automate workflows and orchestration using Cloud Composer (Airflow) * Implement CI/CD pipelines and deployment automation Collaboration & Support * Collaborate with analytics, AI/ML, and business teams for data consumption needs * Troubleshoot data issues and perform root cause analysis * Continuously improve pipeline reliability, scalability, and performance Required Technical Skills * GCP Services: BigQuery, Dataflow, Dataproc, Pub/Sub, Cloud Functions (Gen2), Cloud Composer (Airflow), Cloud Storage, Cloud SQL * Programming: Python (advanced) * Query Language: SQL (advanced) * Data Processing: Batch & Streaming architectures * Scripting: Bash/Shell scripting * Concepts: Data Warehousing, ETL/ELT, Data Lake / Lakehouse architectures * Insurance Domain Experience|Certification Preferred * Google Cloud Platform (GCP) o Vertex AI o Cloud Run o BigQuery o Cloud Storage o IAM o Cloud Logging o Cloud Monitoring * AI & Machine Learning o Large Language Models (LLMs) o Gemini Models o Agentic AI Frameworks o Prompt Engineering o RAG Architectures o Vector Search Concepts o Semantic Retrieval o AI Evaluation Frameworks * Data & Analytics o SQL o BigQuery o Metadata Modeling o Data Pipelines ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)