Data Engineer - Databricks & dbt

Capgemini
United States
29 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$96,158.0 - $150,238.0
Working hours
Regular working hours
Job source

Tech stack

Application Frameworks Code Review Computer Programming Continuous Integration Data Architecture Information Engineering Data Governance Data Infrastructure Data Transformation DevOps Apache Hive Metadata
+19 more
Meta-Data Management Performance Tuning SQL Databases Software Technical Review Planning Software Data Logging Data Processing Macros Data Build Tool (dbt) Git Data Lakes Pyspark Kubernetes Deployment Automation Data Management Software Version Control Data Pipelines Databricks Control M

Job description

technical documentation Logging & Monitoring pyspark data engineering SQL DevOps & CI/CD dbt Certification metadata management databricks

  • We are looking for an experienced Data Engineer / Architect with strong hands-on expertise in Databricks and dbt (Data Build Tool) to support the implementation of enterprise data pipelines on a modern cloud-based data platform.
  • The role will be responsible for defining and implementing scalable data engineering patterns using Databricks and dbt, establishing development standards and reusable frameworks, and providing hands-on technical guidance to the data engineering team.
  • The candidate will work closely with architecture, DevOps, Data Governance, and delivery teams to ensure solutions are scalable, maintainable, and aligned with enterprise standards., * Define the technical architecture and implementation patterns for dbt-based data transformations on Databricks.
  • Design and develop scalable curated and consumption-layer data pipelines using Databricks and dbt.
  • Establish dbt project structure, model organization, dependencies, naming conventions, coding standards, and reusable development patterns.
  • Define and implement appropriate usage of dbt models, sources, tests, macros, snapshots, incremental models, packages, and documentation.
  • Design efficient transformation patterns leveraging Databricks, Delta Lake, Spark SQL, and PySpark.
  • Establish reusable and standardized pipeline development patterns for adoption across the engineering team.
  • Work with architecture teams to ensure implementation aligns with enterprise architecture, governance, security, and data-platform standards.
  • Provide hands-on development support and technical guidance to Data Engineers.
  • Review existing data transformation and pipeline patterns and determine the appropriate implementation using Databricks and dbt.
  • Define data quality, reconciliation, testing, logging, monitoring, and error-handling approaches.
  • Design and implement appropriate incremental data-processing strategies.
  • Support performance optimization and troubleshooting of dbt models and Databricks workloads.
  • Work with DevOps teams to integrate dbt and Databricks development into the enterprise CI/CD framework.
  • Support orchestration and scheduling integration with Databricks Workflows and enterprise scheduling platforms.
  • Conduct technical design reviews and code reviews and ensure adherence to defined engineering standards.
  • Support testing, deployment, production readiness, and troubleshooting activities.
  • Develop reusable templates, frameworks, utilities, and implementation guidelines to accelerate development.
  • Prepare technical documentation and provide knowledge transfer to engineering and support teams.
  • Mentor Data Engineers and provide technical leadership throughout design, development, testing, and deployment.

Requirements

  • 8+ years of experience in Data Engineering, Data Architecture, or related areas.
  • Strong hands-on experience with Databricks.
  • Strong hands-on experience implementing enterprise data solutions using dbt.
  • Strong understanding of dbt models, sources, macros, tests, snapshots, incremental models, packages, and documentation.
  • Strong programming and development skills in SQL, Spark SQL, and PySpark.
  • Strong experience with Delta Lake and Databricks data-engineering capabilities.
  • Experience designing and implementing Medallion Architecture / multi-layered data platforms.
  • Strong understanding of logical and physical data modeling concepts.
  • Experience building scalable, reusable, and metadata-driven data-engineering solutions.
  • Experience implementing data-quality, validation, reconciliation, lineage, monitoring, and observability capabilities.
  • Strong understanding of CI/CD and DevOps practices for dbt and Databricks.
  • Experience with Git-based source control and automated deployment processes.
  • Strong understanding of performance optimization and troubleshooting of large-scale data pipelines.
  • Ability to translate architecture standards and business requirements into practical engineering solutions.
  • Experience conducting technical design and code reviews.
  • Ability to mentor and provide technical direction to Data Engineers.
  • Strong analytical, problem-solving, communication, and stakeholder-management skills.

Preferred Skills:

  • Databricks certification.
  • dbt certification or significant production implementation experience.
  • Experience with AWS-based Databricks environments.
  • Experience with enterprise data governance, metadata management, lineage, and data-product concepts.
  • Experience with Control-M or other enterprise workload scheduling/orchestration platforms.
  • Experience working on large-scale enterprise data-platform transformation programs.

Benefits & conditions

The pay range that the employer in good faith reasonably expects to pay for this position is $46.23/hour - $72.23/hour. Our benefits include medical, dental, vision and retirement benefits. Applications will be accepted on an ongoing basis.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:05 min

Enhancing Databricks tooling for software engineering workflows

Alan Mazankiewicz · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:03 min

Exploring declarative and procedural macro subtypes in Rust environments

Mykhailo Maidan · LIVE

1:22 min

Analyzing differences between mobile and traditional backend DevOps

Mete Baydar Mete Baydar · World Congress 2025

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all