> Markdown version of [/jobs/ext/317906-data-engineer](https://www.wearedevelopers.com/jobs/ext/317906-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Neurons Lab - **Location:** Valencia, Spain (Remote available) - **Contract:** Permanent contract - **Skills:** Airflow, Amazon Web Services, Amazon S3, Data Analysis, Big Data, Encodings, Information Engineering, Data Integration, Dimensional Modeling, Python (Programming Language), Standard Sql, SQL Databases, Parquet, Apache Spark, Data Lineage, AWS Data Analytics - **Published:** June 21, 2026 - **Apply:** https://es.indeed.com/viewjob?jk=0fa275ee0f34bd75 ## About the Role Do you have experience in Spark?, * Strong SQL and Python for large-scale data processing * AWS data stack: S3, Glue, Lake Formation, Athena / Redshift, EMR / Spark, Step Functions / Airflow * Data modeling & semantic layer (dbt or equivalent); dimensional modeling * Entity resolution / record linkage across heterogeneous sources * Data-quality & testing frameworks (Great Expectations, dbt tests) and data lineage * Anonymization / pseudonymization techniques and their analytical trade-offs * Big-data processing (Spark) with performance and cost optimization at scale * Clear written / verbal English; documents for handover and works well with a distributed team, * GDPR fundamentals as applied to anonymized / pseudonymized financial data and UK / EU data residency * AWS Well-Architected (Analytics, Security) for BFSI * Awareness of credit / risk data structures and what downstream modeling consumers need - a plus, * 4+ years in data engineering, with strong AWS + Spark / SQL at scale * Demonstrated experience harmonizing / integrating data across multiple source systems * Experience building validated, reproducible pipelines in a regulated environment (BFSI, healthcare, government) - strong plus * Comfortable stepping into a messy, partly-built data estate and bringing it up to standard * Comfortable as the sole or lead data engineer on a small (3-4 person) delivery pod ## Description * Profile and reconcile differing source schemas across acquired entities: map differing field names, types, encodings and business definitions for the same concept into one conformed model. * Build dbt staging intermediate mart models with tests; codify the harmonized definitions the Data Science Lead specifies. * Write Great Expectations suites (null / range / uniqueness / referential checks) and wire them into the pipeline so bad data fails loudly rather than silently corrupting analysis. * Implement entity / identity resolution (deterministic + fuzzy matching) where there is no clean shared key for the same customer or account across sources. * Implement and verify anonymization / pseudonymization (hashing / tokenization / k-anonymity) and evidence that re-identification risk is controlled for the client's IT / compliance team. * Optimize Spark / Glue jobs over tens of millions of rows - partitioning, file formats (Parquet), incremental loads, cost control. * Orchestrate with Airflow / Step Functions; build repeatable, scheduled pipelines rather than one-off scripts. * Prepare clean, documented, feature-ready datasets for the PD / delinquency models. * Document runbooks so the offshore team can operate the pipelines and handover takes days, not weeks; help scope onboarding of the remaining (Ireland + additional) sources. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Parquet, Delta, Iceberg & Ducklake - An introduction for developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Implementing continuous delivery in a data processing pipeline](https://www.wearedevelopers.com/videos/73-implementing-continuous-delivery-in-a-data-processing-pipeline) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)