Databricks Data Engineer

EXL SERVICE
United States
13 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Compensation
$65,000.0 - $87,000.0
Working hours
Regular working hours

Tech stack

Unity 3d Agile Methodology Artificial Intelligence Data Analysis Microsoft Azure Code Review Continuous Integration Information Engineering Data Governance Data Integration Extract Transform Load (ETL) Data Security
+28 more
Data Warehousing Apache Hive JSON Python (Programming Language) Key Management Meta-Data Management Performance Tuning Systems Development Life Cycle Role-Based Access Control Azure Data Lake SQL Databases Data Streaming Extensible Markup Language (XML) Azure Data Factory Delivery Pipeline Snowflake Apache Spark Caching Ab Initio Data Lakes Pyspark Information Technology Deployment Automation Restful APIs Azure Synapse Analytics Data Pipelines Databricks Web Api

Job description

EXL Service is seeking an accomplished Senior Databricks Data Engineer with 10-12 years of experience designing, developing, and optimizing enterprise data platforms and large-scale ETL solutions. The ideal candidate brings deep expertise in Databricks, Spark, Delta Lake, Azure Data Platform, and modern data engineering practices.

This role is responsible for building scalable data pipelines, implementing cloud-native data solutions, improving platform performance, and delivering reliable data products for analytics, reporting, and AI/ML initiatives.

Base Compensation Range: 65,000 - 87,000, Data Engineering & Platform Development

  • Design, develop, and optimize enterprise-scale data pipelines using Databricks, Spark, Delta Lake, and Azure Data Lake Storage.
  • Build batch and incremental ETL/ELT pipelines, real-time streaming architectures, and reusable data integration frameworks.
  • Develop scalable ingestion frameworks supporting structured, semi-structured, and API-based data sources.
  • Build reusable data engineering frameworks for ingestion, transformation, validation, reconciliation, and publishing.
  • Develop Delta Lake solutions using partitioning, optimization, Z-Ordering, Liquid Clustering, and performance tuning techniques.
  • Implement CDC, incremental processing, merge strategies, and data synchronization across enterprise platforms.
  • Develop Databricks Workflows, notebooks, SQL Jobs, and automation for production workloads.
  • Integrate external REST APIs and process JSON/XML data into analytics-ready datasets.
  • Implement robust data quality checks, audit frameworks, monitoring, and error handling.
  • Optimize Spark jobs for performance, scalability, and cost efficiency.

Databricks & Azure Development

  • Develop solutions using Azure Data Lake Storage Gen2, Databricks, Unity Catalog, Azure Key Vault, Azure DevOps, and Azure Synapse.
  • Build secure data pipelines using Unity Catalog, RBAC, service principals, and managed identities.
  • Develop reusable notebook frameworks using PySpark and Spark SQL.
  • Implement CI/CD deployment pipelines using Azure DevOps.
  • Manage environment promotion across Development, QA, UAT, and Production.
  • Troubleshoot production issues and optimize workloads for reliability and scalability.

Data Integration & Analytics

  • Design enterprise data models supporting reporting, analytics, and downstream applications.
  • Develop healthcare and financial data integration pipelines supporting multiple source systems.
  • Build reusable metadata-driven ETL frameworks.
  • Integrate third-party APIs including NLP, terminology normalization, and identity resolution services.
  • Support data governance, lineage, and metadata management initiatives., Design, build, and migrate large-scale batch data pipelines into Databricks. Develop and optimize production-grade Python/Spark/SQL solutions, support parallel legacy and cloud systems during migration, lead technical design discussions, use AI tools to improve workflows, ensure pipeline reliability through testing/monitoring/tuning, and participate in a rotating on-call schedule. Top Skills: Anthropic ClaudeSparkAzure Data FactoryAzure DevopsCi/CdCosmos DbDatabricksDelta LakeEvent HubsGitGithub CopilotJenkinsNoSQLPythonSQLSynapse AnalyticsTeradata Velera

Requirements

Technical Expertise

  • 10-12 years of experience in Data Engineering, ETL Development, and Data Warehousing.
  • Hands-on Databricks development experience.
  • Strong experience with Spark, Delta Lake, Unity Catalog, Databricks Workflows, and SQL Warehouses.
  • Strong experience with PySpark, Spark SQL, SQL, and Python.
  • Experience building enterprise ETL/ELT pipelines using Databricks and Azure Data Platform.
  • Experience with Azure Data Lake Storage (ADLS Gen2), Azure Synapse, Azure Key Vault, and Azure DevOps.
  • Experience implementing CDC, SCD, incremental loading, and data quality frameworks.
  • Experience integrating REST APIs and processing JSON/XML datasets.
  • Experience with Git, CI/CD, release management, and deployment automation.
  • Strong knowledge of performance tuning, partitioning, caching, broadcast joins, and Spark optimization.
  • Experience working with healthcare data platforms is preferred.
  • Microsoft Azure certifications (e.g., DP-203 Azure Data Engineer Associate) or Databricks certifications (Data Engineer Associate/Professional).
  • Experience with Ab Initio suite of products is also preferred.

Benefits & conditions

113K-188K Annually Mid level 113K-188K Annually Mid level Consulting Design, build, and maintain scalable batch and streaming data pipelines on Databricks using PySpark, Delta Lake, Unity Catalog, and cloud platforms. Troubleshoot performance and quality issues, implement data governance, collaborate with cross-functional teams, and document pipelines and data models to deliver reliable, production-ready data products. Top Skills: AWSAzureCi/CdDatabricksDelta LakeGCPPysparkPythonSparkSQLTerraformUnity Catalog Sparq, Remote USA 110K-143K Annually Expert/Leader 110K-143K Annually Expert/Leader Fintech * Payments * Financial Services Design, develop, test, maintain, and support secure, scalable data solutions using Snowflake and Databricks. Lead technical design, author specifications, review code, mentor engineers, ensure SDLC, audit/security compliance, and drive technical strategy while collaborating with architects and business partners. Participate in Agile ceremonies, be on-call as needed, and travel occasionally. Top Skills: AzureDatabricksEltETLPysparkPythonSnowflakeSQL

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on fa-ewjt-saasfaprod1.fa.ocs.oraclecloud.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

9:56 min

Expanding browser capabilities with modern web APIs

Ire Aderinokun · JS Congress

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · WWC 2025

Videos

See all

Related articles

See all