Data Engineer

Insight Global
Warsaw, United States of America
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Intermediate
Compensation
$ 64K

Job location

Warsaw, United States of America

Tech stack

Agile Methodologies
Application Frameworks
Business Logic
Unit Testing
Big Data
Continuous Integration
Data Validation
Information Engineering
ETL
Data Transformation
DevOps
Distributed Computing Environment
Github
Python
Scrum
Regression Testing
Software Reliability Testing
Standard Sql
SQL Databases
Tableau
Data Ingestion
Azure
Spark
Code Comments
PySpark
Software Coding
Software Version Control
Data Pipelines
Legacy Systems
Alteryx
Databricks

Job description

  • Develop and maintain ETL/ELT pipelines in Databricks using Python (PySpark) and SQL

  • Build data pipelines aligned to medallion architecture (Bronze, Silver, Gold layers)

  • Implement transformations, aggregations, and data enrichment logic to support downstream analytics

  • Ensure all pipelines are scalable, efficient, and production-ready Workflow Migration (Alteryx * Databricks)

  • Convert Alteryx workflows into Databricks-native pipelines

  • Translate legacy business logic into optimized Spark-based transformations

  • Support migration of both: o Simple workflows (standard transformations) o Complex workflows (multi-step logic, flat-file dependencies)

  • Validate output to ensure high parity with legacy systems Data Ingestion & Processing

  • Build and maintain ingestion pipelines for: o Structured data sources o Flat files and semi-structured datasets

  • Implement robust ingestion frameworks to handle schema changes and automate data loads

  • Optimize ingestion processes for performance, cost, and reliability Testing, Quality & Optimization

  • Execute unit testing, SIT, and regression testing for pipelines

  • Support data validation by comparing legacy vs. new system outputs

  • Identify and resolve defects, data discrepancies, and performance issues

  • Optimize pipeline performance (runtime, cost, scalability) Documentation & Best Practices

  • Write clear technical documentation (code comments, runbooks, README files)

  • Follow established coding standards, version control, and CI/CD processes

  • Contribute to continuous improvement of development standards and reusable frameworks Collaboration & Agile Delivery

  • Work within an Agile Scrum environment (2-week sprints)

  • Partner with Data Architects, Product teams, and SMEs to understand requirements

  • Participate in sprint ceremonies (standups, planning, retrospectives)

  • Support backlog refinement and provide technical input on implementation approaches

Requirements

  • 4-8+ years of experience in data engineering / ETL development

  • Strong hands-on experience with: o Databricks (Spark / PySpark) o Azure data ecosystem (ADLS, compute, storage)

  • Proficiency in: o Python and SQL for data engineering o Building and optimizing data pipelines

  • Experience working with: o Large datasets and distributed processing frameworks o Data ingestion from multiple source types (including flat files)

  • Strong understanding of: o Data transformation logic, ETL patterns, and data quality validation

  • Experience working in Agile environments

Nice to Have Skills & Experience

  • Experience converting workflows from Alteryx or similar ETL tools
  • Familiarity with medallion/layered data architectures
  • Exposure to Tableau or BI tools
  • Experience with CI/CD pipelines (GitHub, DevOps tools)
  • Prior experience in supply chain or analytics platforms

Benefits & conditions

Benefit packages for this role will start on the 1st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.

Apply for this position