Software Engineer - Data - Data Mesh & Lakehouse

Lawrence Harvey
Cantabria, Spain
15 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Airflow Apache HTTP Server Microsoft Azure BigQuery Cloud Computing Cloud Storage Code Review Continuous Integration Data Architecture Information Engineering Data Governance Data Sharing
+25 more
Database Applications DevOps Distributed Data Store Machine Learning Operational Databases Role-Based Access Control Software Engineering SONAR (Symantec) Data Streaming Data Processing Cloud Platform System Data Ingestion Apache Spark Change Data Capture Gitlab Build Management Kubernetes Apache Flink Data Analytics Apache Kafka Data Management Data Delivery Confluent Databricks Artifactory

Job description

Software Engineer - Data Data Mesh & Lakehouse About the Role We are looking for a Software Engineer - Data to join a central Data Delivery team building the foundation of a global Data Mesh platform based on a Lakehouse architecture.You will work on the engineering layer responsible for ingesting, processing and provisioning source-aligned data products into central data catalogs, enabling teams across the organisation to consume trusted data for analytics, reporting, machine learning and other data-driven applications.A key part of the platform is enabling scalable data consumption through zero-copy data sharing, while maintaining strong governance, security and data quality standards.What You’ll Be Working On Design and build scalable batch and streaming data pipelines.Develop ingestion solutions for source-aligned data products.Work with Apache Spark and Databricks within a modern Lakehouse environment.Implement data ingestion strategies including Full Loads, Delta Loads and Change Data Capture (CDC).Build and maintain streaming pipelines using technologies such as Kafka, Flink or Confluent.Manage datasets stored across Google Cloud Storage and Azure Blob Storage.Enable secure zero-copy data sharing through technologies such as Databricks Unity Catalog and Big Query.Implement data governance and access-control models including RBAC and attribute-based access control.Design solutions capable of handling complex schema evolution, including backward and forward compatibility.Tech Environment Data & Processing: Apache Spark, Databricks, Big Query Streaming & Ingestion: Kafka, Flink, Confluent, Airbyte Cloud & Storage: GCP, Google Cloud Storage, Azure, Azure Blob Storage Orchestration & Platform: Airflow, Kubernetes Governance: Databricks Unity Catalog, RBAC, ABAC/CBAC, Data Contracts CI/CD: Git Lab, Azure Dev Ops, JFrog Artifactory Quality & Security: Sonar Qube, Snyk What We’re Looking For Strong professional experience in Data Engineering or Software Engineering focused on data platforms.Hands-on experience with Databricks and Apache Spark.Experience designing distributed data pipelines in cloud environments.Strong understanding of Lakehouse architectures.Experience or strong knowledge of Data Mesh principles and Data Products.Experience with batch and streaming ingestion patterns.Knowledge of CDC and incremental data processing strategies.Experience dealing with schema evolution in production data pipelines.Understanding of modern data governance and access-control models.Experience with Airflow, Kubernetes and CI/CD.Comfortable working collaboratively through code reviews and technical design discussions.

Requirements

Tech Environment Data & Processing: Apache Spark, Databricks, Big Query Streaming & Ingestion: Kafka, Flink, Confluent, Airbyte Cloud & Storage: GCP, Google Cloud Storage, Azure, Azure Blob Storage Orchestration & Platform: Airflow, Kubernetes Governance: Databricks Unity Catalog, RBAC, ABAC/CBAC, Data Contracts CI/CD: Git Lab, Azure Dev Ops, JFrog Artifactory Quality & Security: Sonar Qube, Snyk What We’re Looking For Strong professional experience in Data Engineering or Software Engineering focused on data platforms. Hands-on experience with Databricks and Apache Spark. Experience designing distributed data pipelines in cloud environments. Strong understanding of Lakehouse architectures. Experience or strong knowledge of Data Mesh principles and Data Products. Experience with batch and streaming ingestion patterns. Knowledge of CDC and incremental data processing strategies. Experience dealing with schema evolution in production data pipelines. Understanding of modern data governance and access-control models. Experience with Airflow, Kubernetes and CI/CD. Comfortable working collaboratively through code reviews and technical design discussions.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Testing and environment management in GitLab CI

Martin Beránek · LIVE

1:22 min

Analyzing differences between mobile and traditional backend DevOps

Mete Baydar Mete Baydar · World Congress 2025

2:19 min

Introduction to Apache Airflow for advanced orchestration

Alan Mazankiewicz · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:30 min

Feature requests for future GitLab CI versions

Martin Beránek · LIVE

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all