Data Engineer

Launchmetrics
Madrid, Spain
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Unit Testing Software as a Service Software Quality Data Architecture Data Cleansing Data Infrastructure Distributed Computing Environment Python (Programming Language) Object-Oriented Software Development Pytest Data Lakes
+4 more
Pyspark Data Analytics Data Pipelines Databricks

Job description

Experteer Overview Asegúrese de enviar su solicitud rápidamente para maximizar sus posibilidades de ser considerado para una entrevista.Lea la descripción completa del puesto a continuación.In this role you will design and build batch and near-real-time data pipelines on a Databricks-based Lakehouse to enable reliable enrichment and AI-driven insights.You will work within the Data Platform and Data Enrichment team, contributing to data trust and customer-focused data products that power Discover and internal tooling.The role combines hands-on engineering with cross-pod collaboration to scale data infrastructure and improve pipeline reliability.You will be part of a mission-driven, pod-based culture that values ownership and continuous learning.This is a chance to shape scalable data foundations and enable impactful analytics for a global brand ecosystem.Compensaciones / Beneficios* Design and implement batch and near-real-time data pipelines using PySpark and Databricks across Bronze/Silver/Gold layers* Architect efficient Delta Lake table schemas, including partitioning, liquid clustering, schema evolution, and enrichment workflows* Collaborate with product, QA, and other data engineers to translate enrichment and search requirements into reliable pipelines* Own code quality with structured PySpark jobs, unit tests (pytest), and team conventions* Improve pipeline reliability and cost efficiency through scheduling optimization, retry logic, and concurrency management* Contribute to cross-pod initiatives within the data platformResponsabilidades* 3+ years of relevant work experience in a SaaS environment with distributed data processing* Strong Python and PySpark experience* Hands-on experience with Lakehouse architectures (Databricks, Delta Lake, xqbhyrx or equivalents)* Familiarity with Bronze/Silver/Gold data design patterns and schema evolution* Ability to reason about code, understand complex logic, and work with both procedural and object-oriented code* Self-motivated, adaptable, and able to thrive in a fast-paced, results-oriented setting* Fluent EnglishRequisitos principales* learning and development allowance* flexible working arrangements* remote-friendly with home office support* location-based benefits* opportunity for growth* pod autonomy

Requirements

This is a chance to shape scalable data foundations and enable impactful analytics for a global brand ecosystem.Compensaciones / Beneficios* Design and implement batch and near-real-time data pipelines using PySpark and Databricks across Bronze/Silver/Gold layers* Architect efficient Delta Lake table schemas, including partitioning, liquid clustering, schema evolution, and enrichment workflows* Collaborate with product, QA, and other data engineers to translate enrichment and search requirements into reliable pipelines* Own code quality with structured PySpark jobs, unit tests (pytest), and team conventions* Improve pipeline reliability and cost efficiency through scheduling optimization, retry logic, and concurrency management* Contribute to cross-pod initiatives within the data platformResponsabilidades* 3+ years of relevant work experience in a SaaS environment with distributed data processing* Strong Python and PySpark experience* Hands-on experience with Lakehouse architectures (Databricks, Delta Lake, xqbhyrx or equivalents)* Familiarity with Bronze/Silver/Gold data design patterns and schema evolution* Ability to reason about code, understand complex logic, and work with both procedural and object-oriented code* Self-motivated, adaptable, and able to thrive in a fast-paced, results-oriented setting* Fluent EnglishRequisitos principales* learning and development allowance* flexible working arrangements* remote-friendly with home office support* location-based benefits* opportunity for growth* pod autonomy

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

3:24 min

The governance failures of centralized data lakes

Mario Meir-Huber · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:05 min

Tagging and organizing execution scenarios with pytest markers

Florian Bruhin · WWC 2021

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

6:24 min

Distributed data lakes and containerized computing clusters

Ulrich Wurstbauer +1 · LIVE

Videos

See all

Related articles

See all