> Markdown version of [/jobs/ext/1809906-data-engineer-lead](https://www.wearedevelopers.com/jobs/ext/1809906-data-engineer-lead). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer Lead - **Company:** Capgemini - **Location:** Atlanta, GA, United States - **Experience:** Expert - **Salary:** $62,000.0 - $72,000.0 - **Contract:** Permanent contract - **Skills:** Query Performance, Application Programming Interfaces (APIs), Airflow, Big Data, BigQuery, Cloud Database, Cloud Engineering, Cloud Storage, Computer Programming, Continuous Integration, Information Engineering, Data Governance, Data Integration, Extract Transform Load (ETL), Data Warehousing, Software Debugging, DevOps, Distributed Computing Environment, Distributed Systems, Data Flow Control, Python (Programming Language), Machine Learning, Cloudera, SQL Databases, Data Streaming, Workflow Management Systems, Google Cloud, Cloud Platform System, Apache Spark, Git, Data Lakes, Pyspark, Semi-structured Data, Deployment Automation, Integration Frameworks, Apache Kafka, Machine Learning Operations, Data Pipelines - **Published:** July 14, 2026 - **Apply:** https://www.capgemini.com/jobs/516212-en_US_SAPBTP/x/ ## About the Role Strong data engineering mindset (not just scripting) Experience with large-scale datasets (TB-level) Comfortable working in cloud-native, distributed environments Proactive in debugging and optimizing data pipelines' ## Description We are seeking a Data Engineer with strong expertise in Google Cloud Platform (GCP) and PySpark to design, build, and optimize scalable data pipelines. The role focuses on processing large-scale datasets, enabling analytics, and supporting data-driven decision-making within a cloud-native ecosystem. The ideal candidate will have hands-on experience with distributed data processing, cloud data services, and ETL/ELT frameworks, with a strong engineering mindset toward performance, scalability, and reliability. Core Responsibilities 1. Data Pipeline Development Design and develop scalable data pipelines using PySpark Process and transform large datasets (batch and streaming) Build reusable data processing frameworks 2. GCP Data Engineering Work extensively with GCP services, including: BigQuery (data warehousing) Cloud Storage (data lake) Dataflow / Dataproc (processing) Optimize ingestion, storage, and retrieval of datasets in GCP 3. ETL / ELT Engineering Develop and maintain end-to-end ETL/ELT pipelines Ensure: Data quality Data consistency Schema evolution handling 4. Performance & Optimization Optimize PySpark jobs for: Large-scale distributed processing Memory and execution efficiency Tune query performance in BigQuery 5. Data Integration & Collaboration Collaborate with: Data scientists Analysts Application teams Enable datasets for analytics, reporting, and ML workflows 6. DevOps & Automation Implement CI/CD for data pipelines Use Git for version control Automate deployments and monitoring of data workflows Required Technical Skills Programming & Processing Strong Python + PySpark Experience with Spark (RDD/DataFrame APIs) Cloud Platform Hands-on Google Cloud Platform (GCP): BigQuery Cloud Storage Dataproc / Dataflow Pub/Sub (preferred) Data Engineering ETL/ELT pipeline development Data modeling (structured & semi-structured data) SQL (advanced) Tools & Frameworks Airflow / Composer (workflow orchestration) CI/CD pipelines Nice-to-Have Skills Streaming frameworks (Kafka / Pub-Sub streaming) Delta Lake / Iceberg / Lakehouse patterns Machine Learning pipeline exposure Data governance and lineage tools ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)