Mid Data Engineer

Jobtailor
Madrid, Spain
12 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
3 years minimum
Working hours
Regular working hours
Languages
English

Tech stack

Query Performance Application Programming Interfaces (APIs) Airflow Amazon S3 Directed Acyclic Graph (Directed Graphs) Data Transformation Database Queries Software Debugging Apache Hive Identity and Access Management Python (Programming Language) Operational Databases
+10 more
Performance Tuning Standard Sql Reverse Engineering SQL Databases Scripting Snowflake Data Lakes Pyspark Data Pipelines Databricks

Job description

Keep production Databricks pipelines running, including ingestion, transformation, and delivery to downstream consumersDiagnose and resolve pipeline failures and data quality issuesReverse-engineer and document existing transformation logic and business rulesMigrate legacy tables from Hive Metastore to Unity CatalogMaintain Iceberg-enabled table sharing between Databricks and SnowflakeBuild and test dbt models, including incremental materializations and data testsDevelop and maintain Airflow DAGs for orchestrationValidate migrated pipelines against Databricks outputsContribute to Snowflake modeling, performance, and cost decisionsWork directly with client stakeholders on technical topics alongside the team leadRequirements3+ years operating production data pipelinesPySpark and SQL - able to read, debug, and modify existing pipelinesDelta Lake: MERGE/upsert patterns, table properties, OPTIMIZE, partitioningDatabricks Workflows, cluster configuration, job troubleshootingUnity Catalog: catalogs, schemas, grants, lineage, and the metastore modelSnowflake warehouses, roles and grants, and general operating modelSnowflake query performance and awareness of compute cost behaviordbt models, sources, tests, and incremental materializationsdbt project structure and deployment workflowAirflow DAGs, operators, scheduling, and dependency managementAirflow retries, backfills, and idempotent task designStrong SQL, including window functions, complex joins, and reading transformation logicPython for scripting, automation, and API integrationIncremental loading patterns, idempotency, late-arriving data, and reprocessingAWS S3 and IAM basicsBasic working knowledge of Redshift and its role in wider architectureFluent EnglishSelf-directed and able to progress on an unfamiliar codebase without structured onboardingAble to explain production incidents to non-technical stakeholders and provide realistic ETAsCore Competencies Demonstrates expertise in managing production data pipelines using Databricks, including ingestion, transformation, and delivery processes.Proficient in SQL, PySpark, and dbt for building and testing data models, with a strong understanding of Snowflake and Airflow for orchestration and performance optimization.Highest-signal resume keywordsDatabricks Pipeline ManagementSQL ProficiencyPySpark DevelopmentAirflow DAG DevelopmentDbt Model BuildingHard SkillsSQLPySparkDbtAirflowDelta LakeUnity CatalogSnowflakeAWS S3Incremental Loading PatternsData Quality DiagnosisSoft SkillsSelf-DirectedEffective CommunicationStakeholder EngagementIndustry KeywordsData PipelineData TransformationData QualityData ModelingOrchestrationTools & TechnologiesDatabricksSnowflakeAirflowHive MetastoreRedshift#J-*****-Ljbffr

Requirements

3+ years operating production data pipelines PySpark and SQL - able to read, debug, and modify existing pipelines, Airflow DAGs, operators, scheduling, and dependency management Airflow retries, backfills, and idempotent task design Strong SQL, including window functions, complex joins, and reading transformation logic Python for scripting, automation, and API integration Incremental loading patterns, idempotency, late-arriving data, and reprocessing AWS S3 and IAM basics Basic working knowledge of Redshift and its role in wider architecture Fluent English Self-directed and able to progress on an unfamiliar codebase without structured onboarding Able to explain production incidents to non-technical stakeholders and provide realistic ETAs Core Competencies Demonstrates expertise in managing production data pipelines using Databricks, including ingestion, transformation, and delivery processes. Proficient in SQL, PySpark, and dbt for building and testing data models, with a strong understanding of Snowflake and Airflow for orchestration and performance optimization. Highest-signal resume keywords Databricks Pipeline Management SQL Proficiency PySpark Development Airflow DAG Development Dbt Model Building Hard Skills SQL PySpark Dbt Airflow Delta Lake Unity Catalog Snowflake AWS S3 Incremental Loading Patterns Data Quality Diagnosis Soft Skills Self-Directed Effective Communication Stakeholder Engagement Industry Keywords Data Pipeline Data Transformation Data Quality Data Modeling

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann Chris Heilmann +3 · LIVE

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all