Databricks Data Engineer

Capgemini
United States
24 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$125,008.0 - $195,312.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon S3 Big Data Cloud Computing Continuous Integration Data Architecture Data Validation Information Engineering Data Governance Extract Transform Load (ETL) Software Debugging DevOps
+20 more
Fault Tolerance Identity and Access Management Performance Tuning Cloud Services Salesforce.Com SAP Sales and Distribution SQL Databases Systems Integration Management of Software Versions Data Ingestion Software Troubleshooting Data Lakes Pyspark Data Lineage AWS Glue Data Management Cloudwatch Software Version Control Data Pipelines Databricks

Job description

Remote Contract (7 months 25 days) Published 12 hours ago AWS certifications data governance aws cloud data modeling performance optimization pyspark Troubleshooting & Debugging SQL DevOps & CI/CD ETL/ELT pipelines

  • We are seeking a Databricks Engineer to lead the design and implementation of a scalable Sales Data Platform as part of the OneData initiative on Databricks running on AWS.
  • The role is hands-on and architecture-driven, focused on data ingestion, transformation, modeling, and optimization using modern lakehouse patterns.
  • The architect will work closely with data engineers, source system teams, and downstream consumers to deliver high-quality, governed, and performance-optimized sales datasets., Architecture & Design:
  • Define end-to-end lakehouse architecture on Databricks (AWS) for Sales data domains
  • Design medallion architecture (Bronze / Silver / Gold) aligned with OneData standards
  • Establish data modeling standards for Sales facts, dimensions, hierarchies, and aggregations
  • Define scalable ingestion patterns for batch and incremental loads
  • Drive performance, scalability, and cost optimization best practices

Data Engineering & Implementation:

  • Build and guide development of PySpark-based data pipelines in Databricks
  • Implement Delta Lake features:
  • ACID transactions
  • Schema evolution & enforcement
  • Time travel & versioning
  • Design and optimize large-scale joins, aggregations, and window functions
  • Implement CDC and incremental processing using watermarking and change detection
  • Ensure idempotent, restartable, and fault-tolerant pipelines

AWS & Platform Integration:

  • Architect solutions using AWS services:
  • Amazon S3 (data lake storage)
  • IAM (security & access control)
  • AWS Glue / Glue Catalog
  • CloudWatch (monitoring & logging)
  • Optimize Databricks cluster configurations (job vs all-purpose clusters)
  • Implement secrets management and secure connectivity patterns

Data Quality, Governance & Reliability:

  • Define and implement data quality checks and validations
  • Ensure data lineage and metadata capture
  • Implement error handling, auditing, and reconciliation frameworks
  • Support data governance and access control requirements

DevOps & Operational Excellence:

  • Implement CI/CD pipelines for Databricks notebooks and jobs
  • Enforce code versioning, reviews, and deployment standards
  • Design monitoring, alerting, and SLA tracking for pipelines
  • Support production stabilization and performance tuning

Collaboration & Leadership:

  • Act as technical lead / mentor for Databricks data engineers

Collaborate with:

  • Source system teams (Sales, CRM, ERP)
  • Data consumers (analytics, downstream apps)
  • Cloud/platform teams
  • Translate business requirements into robust technical designs

Requirements

Core Technical Skills:

  • 8+ years in Data Engineering / Data Architecture roles
  • 4+ years hands-on experience with Databricks
  • Strong expertise in PySpark & Spark SQL
  • Strong expertise in dbt
  • Deep experience with Delta Lake
  • Strong knowledge of AWS cloud services (S3, IAM, Glue, CloudWatch)

Data Engineering Expertise:

  • Sales data domain experience (orders, revenue, pricing, customers, products)

Strong understanding of:

  • Fact & dimension modeling
  • Slowly Changing Dimensions (SCD Type 1 / 2)
  • Large-scale data processing patterns
  • Experience handling high-volume, high-velocity datasets

Platform & Operational Skills:

  • Databricks job orchestration and scheduling
  • Cluster sizing and performance tuning
  • CI/CD for data platforms
  • Strong troubleshooting and debugging skills

Nice-to-Have:

  • Experience with enterprise OneData / Data Mesh programs
  • Exposure to real-time or near-real-time ingestion patterns
  • Experience integrating CRM / Sales systems (e.g., Salesforce, SAP Sales data)
  • AWS certifications or Databricks certifications

Benefits & conditions

Pulled from the full job description Health insurance Retirement plan Vision insurance Dental insurance, The pay range that the employer in good faith reasonably expects to pay for this position is $60.10/hour - $93.90/hour. Our benefits include medical, dental, vision and retirement benefits. Applications will be accepted on an ongoing basis.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

Videos

See all

Related articles

See all