> Markdown version of [/jobs/ext/2799607-mid-data-engineer](https://www.wearedevelopers.com/jobs/ext/2799607-mid-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Mid Data Engineer - **Company:** Wizeline - **Location:** Barcelona, Spain - **Contract:** Permanent contract - **Skills:** Adobe Analytics, Query Performance, Artificial Intelligence, Airflow, Clickstream, Directed Acyclic Graph (Directed Graphs), Software Debugging, Google Analytics, Apache Hive, Identity and Access Management, Python (Programming Language), Operational Databases, Reverse Engineering, SQL Databases, Web Analytics, Snowflake, Apache Spark, Data Lakes, Pyspark, Databricks - **Published:** September 8, 2026 - **Apply:** https://startup.jobs/mid-data-engineer-barcelona-hybrid-wizeline-7566333 ## About the Role * 3+ years operating production data pipelines. * Strong SQL - window functions, complex joins, reading transformation logic written by someone else. * Python for scripting, automation, and API integration. * Incremental loading patterns, idempotency, late-arriving data, reprocessing. * AWS: S3, IAM basics. Basic working knowledge of Redshift and its role in the wider architecture. Ways of working * Fluent English - client-facing role with stakeholders based abroad. * Self-directed. Able to make progress on an unfamiliar codebase without a structured onboarding path, and comfortable asking good questions when context is missing. * Clear communicator: can explain a production incident to a non-technical stakeholder and give a realistic ETA. Nice-to-have: * Experience with an actual platform migration, not only greenfield work. * Open table formats, particularly Iceberg and cross-platform sharing. * Clickstream or web analytics data (Adobe Analytics, Google Analytics, Segment). * Experience taking over an undocumented system and stabilizing it. * AI Tooling Proficiency: Leverage one or more AI tools to optimize and augment day-to-day work, including drafting, analysis, research, or process automation. Provide recommendations on effective AI use and identify opportunities to streamline workflows. ## Description Existing platform (Databricks) * Keep production pipelines running: ingestion, transformation, and delivery to downstream consumers. * Diagnose and resolve pipeline failures and data quality issues, often without documentation to fall back on. * Reverse-engineer and document existing transformation logic and business rules - this is the input the migration depends on. * Migrate legacy tables from Hive Metastore to Unity Catalog. * Maintain Iceberg-enabled table sharing between Databricks and Snowflake. New development (Snowflake, dbt, Airflow) * Build and test dbt models, including incremental materializations and data tests. * Develop and maintain Airflow DAGs for orchestration. * Validate that migrated pipelines produce output equivalent to the Databricks versions. * Contribute to Snowflake modeling, performance, and cost decisions. Across both * Work directly with client stakeholders on technical topics, alongside the team lead. Technical Requirements Databricks * PySpark and SQL - able to read, debug, and modify existing pipelines. Deep Spark tuning is not required. * Delta Lake: MERGE/upsert patterns, table properties, OPTIMIZE, partitioning. * Databricks Workflows, cluster configuration, job troubleshooting. * Unity Catalog: catalogs, schemas, grants, lineage, and the metastore model. Snowflake * Warehouses, roles and grants, and the general operating model. * Query performance and an awareness of how compute cost behaves. Dbt * Models, sources, tests, and incremental materializations. * Project structure and how dbt fits into a deployment workflow. Airflow * Writing and maintaining DAGs, operators, scheduling, and dependency management. * Understanding retries, backfills, and idempotent task design. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [How Cisco embraced a DevOps culture within its network engineering team](https://www.wearedevelopers.com/videos/99-how-cisco-embraced-a-devops-culture-within-its-network-engineering-team) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 139 - Soft and hard queries](https://www.wearedevelopers.com/magazine/487-dev-digest-139-soft-and-hard-queries) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)