> Markdown version of [/jobs/ext/3538635-data-engineer](https://www.wearedevelopers.com/jobs/ext/3538635-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Capgemini - **Location:** United States - **Experience:** Experienced - **Salary:** $86,129.0 - $127,189.0 - **Contract:** Permanent contract - **Skills:** LangGraph Framework, Artificial Intelligence, BigQuery, Cloud Storage, Continuous Integration, Data Deduplication, Information Engineering, Electronic Data Interchange (EDI), Identity and Access Management, Python (Programming Language), SQL Databases, Data Lakes, LangSmith, Data Lineage, Google Cloud Functions, Data Pipelines, User Identification, Databricks - **Published:** October 1, 2026 - **Apply:** https://www.capgemini.com/jobs/566918-en_US_SAPBTP/x/ ## About the Role * 5-7 years of data engineering experience with 2+ years of production Databricks work: Delta Lake, Unity Catalog, DLT/Lakeflow, structured streaming, and SQL warehouses * Strong Python for transformation logic, data-quality automation, and pipeline orchestration * Expert-level SQL; comfort designing schemas for both analytical and serving workloads * GCP fluency: BigQuery, Cloud Storage, Cloud Run, IAM - Helm runs on GCP and Databricks runs within it * Hands-on experience with at least one clean room or privacy-preserving data exchange technology * PII governance knowledge in a regulated industry context: data residency, consent frameworks, GLBA, Fair Lending, UDAAP, * Databricks AI Functions (ai_query) or Vertex AI used inside pipelines for document extraction and normalization * dbt or a comparable transformation and lineage framework * Identity resolution concepts at the pipeline layer: deterministic matching, RampID or UID2 linkage, match-rate monitoring * LiveRamp, AMC, GMP, or XMi clean room connector experience * Familiarity with LangSmith or LangGraph as a data consumer - understanding what the agentic layer needs from the serving layer ## Description * Design, build, and operate the full medallion stack on Databricks: raw landing (Bronze), identity-resolved and feature-engineered Silver, and the governed serving layer (Gold) that the Helm agent layer queries on every conversational turn * Implement Delta Live Tables (DLT) and Lakeflow pipelines for batch and structured-streaming ingestion of transaction signals, behavioral events, and third-party identity data * Own schema enforcement and evolution at the Bronze boundary so upstream change surfaces as a controlled event, not a downstream failure * Build the unstructured processing tier (Bronze * Silver) using Databricks AI Functions or ai_query to convert PDFs, brand guidelines, compliance rule sets, and creative assets into governed Silver tables Unity Catalog & Governance * Own Unity Catalog as the governance layer: access control, lineage tracking, and per-client tenant isolation across Helm's financial institution client base * Implement PII governance controls at the pipeline layer - redacted ID egress, consent signal propagation, and guardrail alignment to GLBA, Fair Lending, and UDAAP * Build and maintain data-quality expectations frameworks: reconciliation, deduplication, and lineage checks that reduce manual investigation cycles Serving Layer & Platform Operations * Build and operate the Gold serving layer behind Helm's twelve analytical models: MTA, MMM, incrementality, LTV:CAC, reach/frequency, cohorts, cobrand overlap, channel viability, card activation, product holdings, scenario reallocation, and channel universe * Meet the serving contract the agentic layer depends on: fixed signal shapes, live query on every turn, predictable latency, and per-client isolation * Own CI/CD, infrastructure-as-code, cost management, and pipeline observability for the Databricks environment ## Related Videos - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Fully Orchestrating Databricks from Airflow](https://www.wearedevelopers.com/videos/336-fully-orchestrating-databricks-from-airflow) - [Building AI Applications with LangChain and Node.js](https://www.wearedevelopers.com/videos/1512-building-ai-applications-with-langchain-and-node-js) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Data Governance in the Era of AI](https://www.wearedevelopers.com/videos/1622-data-governance-in-the-era-of-ai) - [Making Data Warehouses fast. A developer's story.](https://www.wearedevelopers.com/videos/302-making-data-warehouses-fast-a-developer-s-story) ## Related Articles - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j)