Lead Data Engineer

Sovereign Technologies
Eagan, MN, United States
2 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Agile Methodology Artificial Intelligence Airflow Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Big Data Information Engineering Extract Transform Load (ETL) Database Queries Distributed Computing Environment Distributed Systems
+13 more
Healthcare Effectiveness Data and Information Set Python (Programming Language) Scrum Methodology Standard Sql SAS (Software) Scripting Freeform SQL Generative AI Git Pyspark Software Version Control Data Pipelines Databricks

Job description

We are seeking a Senior/Lead Data Engineer to modernize and own enterprise data pipelines by migrating legacy SAS-based workflows to Python/PySpark on Databricks. The engineer will partner with SMEs during the transition period and eventually take ownership of a critical healthcare analytics pipeline supporting HEDIS reporting., * Migrate legacy SAS pipelines to Python/PySpark on Databricks.

  • Design, build, and maintain scalable ETL/ELT pipelines.
  • Develop distributed data processing solutions using Databricks.
  • Create and optimize complex SQL queries.
  • Schedule, automate, and monitor data pipelines.
  • Work with AWS services including S3, Lambda, Glue, and EC2.
  • Manage Databricks notebooks, workflows, and clusters.
  • Collaborate with business stakeholders and SMEs to transition pipeline ownership.
  • Follow Agile methodologies and Git-based development practices.

Requirements

  • 8+ years of Data Engineering experience.
  • Strong Python programming and scripting skills.
  • Hands-on PySpark development.
  • Extensive Databricks experience (Notebooks, Workflows, Cluster Management).
  • Experience processing large datasets in distributed environments.
  • Strong SQL skills.
  • AWS experience with:
  • S3
  • Glue
  • Lambda
  • EC2
  • Experience building and automating ETL/ELT pipelines.
  • Git/version control experience.
  • Agile/Scrum experience.

Preferred Skills

  • Experience with SAS-to-Python migration.
  • AI/Automation experience.
  • Healthcare or HEDIS domain experience.
  • Lead/Principal Data Engineering experience.

Must-Have Skills

  • Python
  • PySpark
  • Databricks
  • SQL
  • AWS (S3, Glue, Lambda, EC2)
  • ETL/ELT Pipeline Development
  • Distributed Data Processing
  • Git
  • Agile/Scrum

Nice-to-Have Skills

  • SAS Modernization
  • AI/Generative AI
  • Healthcare/HEDIS
  • Airflow or other workflow orchestration tools
  • Leadership/Pipeline Ownership experience

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all