Data Engineer (Redshift/Databricks) - INTL India

Insight Global
Plano, TX, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon S3 Microsoft Azure Code Review Continuous Integration Information Engineering Extract Transform Load (ETL) Github Apache Hive Performance Tuning SQL Stored Procedures SQL Databases Apache Spark
+7 more
Caching Data Lakes Pyspark AWS Glue Terraform Code Restructuring Databricks

Job description

Insight Global is looking for a Data Engineer to help lead a migration from Glue/Redshift to Databricks

What you will do:

Own the end-to-end migration roadmap: discovery, assessment, design, build, test, cutover, and decommissioning of legacy AWS Glue jobs, Redshift clusters, and associated tooling.

Inventory and prioritize existing pipelines, stored procedures, materialized views, and Redshift workloads; define a wave-based migration plan with clear success criteria for each wave.

Architect the target Databricks Lakehouse: workspace and account topology, Unity Catalog design, medallion (bronze/silver/gold) layering, storage layout on S3, cluster and SQL Warehouse strategy, and cost guardrails.

Re-platform pipelines from Glue (PySpark / Spark SQL) and Redshift SQL into Databricks using PySpark, Spark SQL, Delta Live Tables, and/or Databricks Workflows; refactor where it improves performance, maintainability, or cost.

Define and enforce data parity, reconciliation, and regression-testing strategy between legacy Redshift and the new Lakehouse during dual-run periods.

Plan and execute cutover for downstream consumers (BI tools, reverse ETL, ML feature stores, application reads), minimizing disruption.

Set technical direction and standards

Establish engineering standards for the new platform: repository structure, CI/CD (e.g., GitHub Actions / Azure DevOps with Databricks Asset Bundles or Terraform), environment promotion, code review, and testing (unit, data quality, integration).

Define and roll out data quality, observability, and lineage practices (e.g., Great Expectations / DLT expectations, Unity Catalog lineage, monitoring and alerting).

Govern access, security, and compliance via Unity Catalog: catalogs, schemas, row/column-level security, and audit.

Drive cost optimization: cluster sizing and policies, photon usage, autoscaling, job vs. all-purpose compute decisions, and SQL Warehouse right-sizing.

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global’s Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.

Requirements

8+ years of professional data engineering experience, with at least 2+ years in a tech-lead or lead engineer capacity.

-5-6 years of experience in each of the following areas:

Deep, hands-on experience with Apache Spark (PySpark and Spark SQL), including performance tuning (partitioning, shuffles, skew, caching, file sizing).

-Hands-on production experience with Databricks Delta Lake, Unity Catalog, Workflows, and either Delta Live Tables or a comparable declarative pipeline framework.

-Strong production experience with AWS Glue (jobs, crawlers, Data Catalog

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · WWC Europe 2026

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

47 sec

Building modern data pipelines for legacy exports

Dr. Alexander Wachtel Dr. Alexander Wachtel +1 · WWC 2025

Videos

See all

Related articles

See all