AWS Databricks Data Engineer

Palni Inc
United States
10 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Airflow Amazon Web Services Amazon S3 Cloud Computing Cloud Storage Computer Programming Continuous Integration Information Engineering Data Files Data Governance Github Identity and Access Management
+14 more
Python (Programming Language) Medicaid Management Information Systems (MMIS) SQL Databases Data Streaming Apache Spark Gitlab Build Management Data Lakes Pyspark AWS Glue Data Lakehouse Data Pipelines Serverless Computing Databricks

Job description

6months Contract/Contract to Hire Must be authorized to work in the U.S. without sponsorship. No EAD’’s An AWS Databricks Data Engineer to design, build, and maintains scalable data pipelines and lakehouse architectures using Apache Spark, Python, and SQL. Optimize cloud storage, ensure data quality, and integrate seamlessly with other AWS services

Core Responsibilities

  • Pipeline Development: Design and build batch and streaming data pipelines using PySpark, Delta Lake, Autoloader and Delta Live Tables (DLT) to ingest data sets into Databricks.
  • AWS Integration: Build cloud-native data solutions leveraging AWS services (e.g., S3 for storage, IAM for access management, and AWS Glue or Lambda for serverless tasks
  • Data Governance & Security: Configure and manage access controls using Databricks Unity Catalog to ensure compliance and monitor lineage
  • Orchestration & CI/CD: Automate pipeline deployments using Databricks Workflows, Apache Airflow, and CI/CD tools (e.g., GitHub, GitLab).
  • Cross-Functional Collaboration: Partner closely with stakeholders to build robust feature stores and prepare datasets for various consumption needs

Requirements

  • Industry: Healthcare Payer industry experience. Have worked on MMIS data sets - claims, provider, member enrollment and similar data sets
  • Experience: 3-5 years of hands-on data engineering experience, specifically on Databricks.
  • Programming: High proficiency in Python (specifically PySpark) and advanced SQL.
  • Big Data & Cloud: Strong understanding of Apache Spark, Data Lakehouse architecture, and working within a production AWS environment.
  • Databricks Ecosystem: Familiarity with the Databricks platform ecosystem, including notebooks, Delta Lake, and Unity Catalog.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler · LIVE

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

1:43 min

AWS infrastructure stack and data flow pipeline overview

Artem Volk Artem Volk +1 · World Congress 2024

4:54 min

Implementing geographic salary tiers for compensation equity and fairness

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all