Remote Data Engineer - Python

Insight Global
Plano, TX, United States
2 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Airflow BigQuery Continuous Integration Directed Acyclic Graph (Directed Graphs) Information Engineering Data Governance Python (Programming Language) Reference Data Standard Sql Teradata SQL Datadog
+7 more
Macros Snowflake Technical Debt Git Code Restructuring Amazon Redshift Databricks

Job description

Our product catalog work is moving out of a legacy project and into a new one. You will own moving pipelines and code across, which in many cases means rewriting rather than porting. This role suits an engineer who is comfortable working in an unfamiliar codebase, figuring out intent from the code itself, and making judgment calls about what to preserve and what to rebuild.

What you’ll do

  1. Audit pipelines, jobs, and code in the legacy project: what runs, what it produces, who consumes it, and what is actually dead.

  2. Decide, with the team, what gets ported as-is, what gets rewritten to current standards, and what gets retired.

  3. Rewrite pipelines to the new project’s patterns, conventions, and directory structure, including modular models, tests, and documented dependencies.

  4. Refactor code that carries accumulated shortcuts: hardcoded values, duplicated logic, undocumented business rules, missing error handling.

  5. Migrate orchestration, scheduling, alerting, and access controls, not just the transformation code.

  6. Build validation to prove migrated pipelines produce equivalent output, and document any intentional differences.

  7. Manage cutover for each pipeline with downstream consumers, including communication and rollback.

  8. Leave behind documentation and diagrams for what you moved so the team can maintain it after the engagement.

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global’s Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.

Requirements

  1. 5+ years in data engineering, including at least one migration or major refactor of an existing production codebase.

  2. Strong SQL and strong Python- heavy code rewriting

  3. Experience with dbt or a comparable transformation framework, including project structure, macros, testing, and documentation.

  4. Hands-on Snowflake or comparable cloud warehouse experience.

  5. Orchestration experience: Airflow or similar, including migrating DAGs between environments or projects.

  6. Solid Git and CI/CD practice. You should be comfortable in a repo with existing standards and reviewers.

  7. Ability to read code without a spec and work out what the business rule was meant to be, then ask the right questions before changing it. 1. Experience with product catalog, item master, or reference data domains.

  8. Background in tech debt reduction or platform consolidation work.

  9. Experience with API or feed integrations to external partners.

  10. Familiarity with data quality frameworks and observability tooling.

  11. Teradata, BigQuery, RedShift, Databricks

  12. CPG/Consumer packaged good experience

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:03 min

Exploring declarative and procedural macro subtypes in Rust environments

Mykhailo Maidan · LIVE

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

1:59 min

Navigating the disadvantages and maintenance costs of procedural macros

Mykhailo Maidan · LIVE

Videos

See all

Related articles

See all