Software Engineer - Data - Data Mesh & Lakehouse

Lawrence Harvey
Zaragoza, Spain
19 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Airflow Apache HTTP Server Microsoft Azure Cloud Storage Code Review Continuous Integration Information Engineering Data Governance Data Sharing Distributed Data Store Machine Learning Operational Databases
+16 more
Role-Based Access Control Software Engineering SonarQube Data Streaming Data Processing Data Ingestion Apache Spark Change Data Capture Gitlab Kubernetes Apache Flink Data Analytics Apache Kafka Data Delivery Confluent Databricks

Job description

Software Engineer - Data Data Mesh & LakehouseĀæTodo listo para enviar su solicitud?Por favor, lea la descripción al menos una vez antes de hacer clic en ā€œSolicitarā€.About the RoleWe are looking for a Software Engineer - Data to join a central Data Delivery team building the foundation of a global Data Mesh platform based on a Lakehouse architecture.You will work on the engineering layer responsible for ingesting, processing and provisioning source-aligned data products into central data catalogs, enabling teams across the organisation to consume trusted data for analytics, reporting, machine learning and other data-driven applications.A key part of the platform is enabling scalable data consumption through zero-copy data sharing, while maintaining strong governance, security and data quality standards.What You’ll Be Working OnDesign and build scalable batch and streaming data pipelines.Develop ingestion solutions for source-aligned data products.Work with Apache Spark and Databricks within a modern Lakehouse environment.Implement data ingestion strategies including Full Loads, Delta Loads and Change Data Capture (CDC).Build and maintain streaming pipelines using technologies such as Kafka, Flink or Confluent.Manage datasets stored across Google Cloud Storage and Azure Blob Storage.Enable secure zero-copy data sharing through technologies such as Databricks Unity Catalog and BigQuery.Implement data governance and access-control models including RBAC and attribute-based access control.Design solutions capable of handling complex schema evolution, including backward and forward compatibility.Tech EnvironmentData & Processing: Apache Spark, Databricks, BigQueryStreaming & Ingestion: Kafka, Flink, Confluent, AirbyteCloud & Storage: GCP, Google Cloud Storage, Azure, Azure Blob StorageOrchestration & Platform: Airflow, KubernetesGovernance: Databricks Unity Catalog, RBAC, ABAC/CBAC, Data ContractsCI/CD: GitLab, Azure DevOps, JFrog ArtifactoryQuality & Security: SonarQube, SnykWhat We’re Looking ForStrong professional experience in Data Engineering or Software Engineering focused on data platforms.Hands-on experience with Databricks and Apache Spark.Experience designing distributed data pipelines in cloud environments.Strong understanding of Lakehouse architectures.Experience or strong knowledge of Data Mesh principles and Data Products.Experience with batch and streaming ingestion patterns.Knowledge of CDC and incremental data processing strategies.Experience dealing with schema evolution in production data pipelines.Understanding of modern data governance and access-control models.Experience with Airflow, Kubernetes and CI/CD.xsgfvud Comfortable working collaboratively through code reviews and technical design discussions.

Requirements

Build and maintain streaming pipelines using technologies such as Kafka, Flink or Confluent.Manage datasets stored across Google Cloud Storage and Azure Blob Storage.Enable secure zero-copy data sharing through technologies such as Databricks Unity Catalog and BigQuery.Implement data governance and access-control models including RBAC and attribute-based access control.Design solutions capable of handling complex schema evolution, including backward and forward compatibility.Tech EnvironmentData & Processing: Apache Spark, Databricks, BigQueryStreaming & Ingestion: Kafka, Flink, Confluent, AirbyteCloud & Storage: GCP, Google Cloud Storage, Azure, Azure Blob StorageOrchestration & Platform: Airflow, KubernetesGovernance: Databricks Unity Catalog, RBAC, ABAC/CBAC, Data ContractsCI/CD: GitLab, Azure DevOps, JFrog ArtifactoryQuality & Security: SonarQube, SnykWhat We’re Looking ForStrong professional experience in Data Engineering or Software Engineering focused on data platforms.Hands-on experience with Databricks and Apache Spark.Experience designing distributed data pipelines in cloud environments.Strong understanding of Lakehouse architectures.Experience or strong knowledge of Data Mesh principles and Data Products.Experience with batch and streaming ingestion patterns.Knowledge of CDC and incremental data processing strategies.Experience dealing with schema evolution in production data pipelines.Understanding of modern data governance and access-control models.Experience with Airflow, Kubernetes and CI/CD. xsgfvud Comfortable working collaboratively through code reviews and technical design discussions.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:09 min

Evaluating mature stream processing frameworks for production systems

Soroosh Khodami Soroosh Khodami Ā· World Congress 2024

1:50 min

Building a data producer using the Confluent Python library

Lucia Cerchie Ā· LIVE

2:19 min

Introduction to Apache Airflow for advanced orchestration

Alan Mazankiewicz Ā· LIVE

3:14 min

Testing and environment management in GitLab CI

Martin BerÔnek · LIVE

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou Ā· Coffee With Developers

6:21 min

Audience questions on brokers and monitoring tools

Developersteve Ā· LIVE

Videos

See all

Related articles

See all