Mid Data Engineer

Wizeline
Barcelona, Spain
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
3 years minimum
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Adobe Analytics Query Performance Artificial Intelligence Airflow Clickstream Directed Acyclic Graph (Directed Graphs) Software Debugging Google Analytics Apache Hive Identity and Access Management Python (Programming Language) Operational Databases
+8 more
Reverse Engineering SQL Databases Web Analytics Snowflake Apache Spark Data Lakes Pyspark Databricks

Job description

Existing platform (Databricks)

  • Keep production pipelines running: ingestion, transformation, and delivery to downstream consumers.
  • Diagnose and resolve pipeline failures and data quality issues, often without documentation to fall back on.
  • Reverse-engineer and document existing transformation logic and business rules - this is the input the migration depends on.
  • Migrate legacy tables from Hive Metastore to Unity Catalog.
  • Maintain Iceberg-enabled table sharing between Databricks and Snowflake.

New development (Snowflake, dbt, Airflow)

  • Build and test dbt models, including incremental materializations and data tests.
  • Develop and maintain Airflow DAGs for orchestration.
  • Validate that migrated pipelines produce output equivalent to the Databricks versions.
  • Contribute to Snowflake modeling, performance, and cost decisions.

Across both

  • Work directly with client stakeholders on technical topics, alongside the team lead.

Technical Requirements

Databricks

  • PySpark and SQL - able to read, debug, and modify existing pipelines. Deep Spark tuning is not required.
  • Delta Lake: MERGE/upsert patterns, table properties, OPTIMIZE, partitioning.
  • Databricks Workflows, cluster configuration, job troubleshooting.
  • Unity Catalog: catalogs, schemas, grants, lineage, and the metastore model.

Snowflake

  • Warehouses, roles and grants, and the general operating model.
  • Query performance and an awareness of how compute cost behaves.

Dbt

  • Models, sources, tests, and incremental materializations.
  • Project structure and how dbt fits into a deployment workflow.

Airflow

  • Writing and maintaining DAGs, operators, scheduling, and dependency management.
  • Understanding retries, backfills, and idempotent task design.

Requirements

  • 3+ years operating production data pipelines.
  • Strong SQL - window functions, complex joins, reading transformation logic written by someone else.
  • Python for scripting, automation, and API integration.
  • Incremental loading patterns, idempotency, late-arriving data, reprocessing.
  • AWS: S3, IAM basics. Basic working knowledge of Redshift and its role in the wider architecture.

Ways of working

  • Fluent English - client-facing role with stakeholders based abroad.
  • Self-directed. Able to make progress on an unfamiliar codebase without a structured onboarding path, and comfortable asking good questions when context is missing.
  • Clear communicator: can explain a production incident to a non-technical stakeholder and give a realistic ETA.

Nice-to-have:

  • Experience with an actual platform migration, not only greenfield work.
  • Open table formats, particularly Iceberg and cross-platform sharing.
  • Clickstream or web analytics data (Adobe Analytics, Google Analytics, Segment).
  • Experience taking over an undocumented system and stabilizing it.
  • AI Tooling Proficiency: Leverage one or more AI tools to optimize and augment day-to-day work, including drafting, analysis, research, or process automation. Provide recommendations on effective AI use and identify opportunities to streamline workflows.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

1:33 min

Integrating internal APIs and maintaining data sovereignty

Mahran Meißner Mahran Meißner · World Congress 2026 Europe

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:46 min

Transforming data architecture from on-premise to cloud

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all