> Markdown version of [/jobs/ext/1620838-data-engineer](https://www.wearedevelopers.com/jobs/ext/1620838-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Atolls - **Location:** Madrid, Spain - **Contract:** Permanent contract - **Skills:** Airflow, Amazon Web Services, Amazon S3, Apache HTTP Server, BigQuery, Code Review, Continuous Integration, Data as a Services, Directed Acyclic Graph (Directed Graphs), Information Engineering, Data Systems, Software Debugging, Dimensional Modeling, Fault Tolerance, Python (Programming Language), Online Analytical Processing, Operational Databases, Ansible, SQL Databases, Data Streaming, Systems Integration, Tableau (Software), Parquet, Business Intelligence Development Studio, Sql Optimization, Apache Spark, Backend, Git, Containerization, Data Lakes, Pyspark, Apache Kafka, Vertica, Terraform, Looker Analytics, Data Pipelines, Docker - **Published:** July 22, 2026 - **Apply:** https://www.jobleads.com/es/job/e8f4b0dcfa87e2c32a4d063f906477145 ## About the Role We're looking for someone with data engineering experience, who is dedicated to creating exceptional user experiences and driving innovation. Must have: * 4+ years building and running production data pipelines. * Python - strong, idiomatic, tested. You write pipelines that other people can read and easily debug. * SQL - advanced. Window functions, CTEs, query plans, and the instinct to know why a query got slow. * Data modeling & warehousing - dimensional modeling, incremental vs. full-refresh strategies, slowly changing dimensions, and the trade-offs between them. * dbt - building, testing, and documenting models in a real project. * Apache Airflow - authoring DAGs, plus the operational side: backfills, retries, SLAs, and debugging a failed run. * AWS - hands-on experience with S3 and the surrounding data services (Athena, Kinesis, EKS, Lambda). * Docker & Kubernetes - you can containerize a job and understand what happens when a pod won't schedule. * Git and CI/CD - branching, code review, and pipelines that gate merges. Nice to have: * Knowledge of ClickHouse or another columnar OLAP engine (BigQuery, Redshift) table engines, partitioning, and MergeTree tuning are a big plus. * Good experience with streaming data ingestion technologies such as Kafka, Kinesis, or similar. * Familiarity with infrastructure as code (IaC) - Terraform, Helm, Ansible. * Experience with data lake architectures - Parquet, Iceberg, or Delta Lake, and an understanding of layered lake design (bronze * silver * gold). * Experience with BI tooling - Apache Superset, Looker, Tableau, or equivalent. * Strong experience integrating third-party APIs - handling inconsistent schemas, rate limits, and unpredictable failure modes at scale. * Working knowledge of Apache Spark - PySpark or Scala. Soft skills: * Attention to detail - You care about correctness and you build the checks that prove it. In data engineering, a silently wrong number is worse than a loud failure. * Ownership - you run what you build, and you improve it when needed. * Clear communication - you can effectively explain in a pipeline to an analyst and a trade-off to a stakeholder. * Pragmatism - you can tell the difference between the right long-term design and the short-term solution, and you know when each one applies. * Collaboration - you work well with data engineers, analysts, and other product and backend teams. Language requirements: * English - fluent, written and spoken. ## Description In this role, you will: * Design, Build and optimize batch and streaming data pipelines with strong performance, fault tolerance, and observability using Python, Spark, and Clickhouse. * Develop and operate workflow orchestration with Apache Airflow to schedule, monitor, and manage data pipelines and transformations. * Manage, model and document critical data systems for analytics using SQL and dbt to support business intelligence and reporting workloads. * Implement infrastructure-as-code (e.g., Terraform) to provision and manage cloud-based data platform components. * Containerize and deploy services using Docker and Kubernetes (and related tooling such as Helm). * Collaborate with analysts, application teams and other stakeholders to turn requirements into technical designs and delivered solutions. ## Related Videos - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)