> Markdown version of [/jobs/ext/2619215-databricks-data-engineer](https://www.wearedevelopers.com/jobs/ext/2619215-databricks-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Databricks Data Engineer - **Company:** Capgemini - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $125,008.0 - $195,312.0 - **Contract:** Temporary contract - **Skills:** Amazon Web Services, Amazon S3, Big Data, Cloud Computing, Continuous Integration, Data Architecture, Data Validation, Information Engineering, Data Governance, Extract Transform Load (ETL), Software Debugging, DevOps, Fault Tolerance, Identity and Access Management, Performance Tuning, Cloud Services, Salesforce.Com, SAP Sales and Distribution, SQL Databases, Systems Integration, Management of Software Versions, Data Ingestion, Software Troubleshooting, Data Lakes, Pyspark, Data Lineage, AWS Glue, Data Management, Cloudwatch, Software Version Control, Data Pipelines, Databricks - **Published:** August 10, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=dc07515a51a0557e ## About the Role Core Technical Skills: * 8+ years in Data Engineering / Data Architecture roles * 4+ years hands-on experience with Databricks * Strong expertise in PySpark & Spark SQL * Strong expertise in dbt * Deep experience with Delta Lake * Strong knowledge of AWS cloud services (S3, IAM, Glue, CloudWatch) Data Engineering Expertise: * Sales data domain experience (orders, revenue, pricing, customers, products) Strong understanding of: * Fact & dimension modeling * Slowly Changing Dimensions (SCD Type 1 / 2) * Large-scale data processing patterns * Experience handling high-volume, high-velocity datasets Platform & Operational Skills: * Databricks job orchestration and scheduling * Cluster sizing and performance tuning * CI/CD for data platforms * Strong troubleshooting and debugging skills Nice-to-Have: * Experience with enterprise OneData / Data Mesh programs * Exposure to real-time or near-real-time ingestion patterns * Experience integrating CRM / Sales systems (e.g., Salesforce, SAP Sales data) * AWS certifications or Databricks certifications ## Description Remote Contract (7 months 25 days) Published 12 hours ago AWS certifications data governance aws cloud data modeling performance optimization pyspark Troubleshooting & Debugging SQL DevOps & CI/CD ETL/ELT pipelines * We are seeking a Databricks Engineer to lead the design and implementation of a scalable Sales Data Platform as part of the OneData initiative on Databricks running on AWS. * The role is hands-on and architecture-driven, focused on data ingestion, transformation, modeling, and optimization using modern lakehouse patterns. * The architect will work closely with data engineers, source system teams, and downstream consumers to deliver high-quality, governed, and performance-optimized sales datasets., Architecture & Design: * Define end-to-end lakehouse architecture on Databricks (AWS) for Sales data domains * Design medallion architecture (Bronze / Silver / Gold) aligned with OneData standards * Establish data modeling standards for Sales facts, dimensions, hierarchies, and aggregations * Define scalable ingestion patterns for batch and incremental loads * Drive performance, scalability, and cost optimization best practices Data Engineering & Implementation: * Build and guide development of PySpark-based data pipelines in Databricks * Implement Delta Lake features: * ACID transactions * Schema evolution & enforcement * Time travel & versioning * Design and optimize large-scale joins, aggregations, and window functions * Implement CDC and incremental processing using watermarking and change detection * Ensure idempotent, restartable, and fault-tolerant pipelines AWS & Platform Integration: * Architect solutions using AWS services: * Amazon S3 (data lake storage) * IAM (security & access control) * AWS Glue / Glue Catalog * CloudWatch (monitoring & logging) * Optimize Databricks cluster configurations (job vs all-purpose clusters) * Implement secrets management and secure connectivity patterns Data Quality, Governance & Reliability: * Define and implement data quality checks and validations * Ensure data lineage and metadata capture * Implement error handling, auditing, and reconciliation frameworks * Support data governance and access control requirements DevOps & Operational Excellence: * Implement CI/CD pipelines for Databricks notebooks and jobs * Enforce code versioning, reviews, and deployment standards * Design monitoring, alerting, and SLA tracking for pipelines * Support production stabilization and performance tuning Collaboration & Leadership: * Act as technical lead / mentor for Databricks data engineers Collaborate with: * Source system teams (Sales, CRM, ERP) * Data consumers (analytics, downstream apps) * Cloud/platform teams * Translate business requirements into robust technical designs ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk)