> Markdown version of [/jobs/ext/2949530-senior-data-engineer](https://www.wearedevelopers.com/jobs/ext/2949530-senior-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Engineer - **Company:** Holcim - **Location:** Madrid, Spain - **Contract:** Permanent contract - **Skills:** Query Performance, Java (Programming Language), Agile Methodology, Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Automation of Tests, Big Data, Cloud Computing, Program Optimization, Software Quality, Continuous Integration, Data Infrastructure, Extract Transform Load (ETL), Data Systems, Data Visualization, Data Warehousing, DevOps, Distributed Computing Environment, Distributed Systems, Github, Identity and Access Management, Python (Programming Language), OpenShift, Scrum Methodology, SQL Databases, Data Streaming, Parquet, Qliksense, Data Storage Technologies, Real Time Systems, System Availability, Apache Spark, Git, Cloudformation, Data Lakes, Information Technology, Low Latency, Apache Flink, Avro, AWS Glue, Data Analytics, Qlikview, AWS Data Analytics, Real Time Data, Apache Kafka, Terraform, Data Pipelines, Serverless Computing, Amazon Elastic Mapreduce (EMR), Docker, Jenkins, Amazon Redshift, Databricks - **Published:** September 16, 2026 - **Apply:** https://careers.holcimgroup.com/talentcommunity/apply/1356730257/?locale=en_US ## About the Role * Experience: Minimum 4+ years of hands-on experience in active Big Data environments and 2+ years specializing in Data Analytics within AWS. + Compute & Processing: Amazon EMR: Architecting and managing Spark clusters for large-scale distributed processing. o AWS Glue: Developing serverless ETL jobs, managing the Data Catalog, and implementing Glue Crawlers., + Streaming (Advantage): Amazon Kinesis or MSK (Managed Streaming for Kafka) for real-time data ingestion. * Core Engineering: Expert-level proficiency in Spark, Python, and SQL. * Infrastructure & Tooling: Proven experience with Airflow for orchestration and Docker/ECS for containerization. * Good knowledge in Databricks and data mesh architectures. Good understanding in how to implement and maintain Lakehouse data models (bronze / silver / gold layers) using Delta Lake for reliability, ACID transactions, time travel and schema evolution. * Solid software engineering practices: Git, CI/CD for data pipelines, automated testing, code quality and documentation. * Communication: Excellent written and oral English communication skills, with the ability to explain complex technical concepts to non-technical audiences. * Degree in Computer Science, Engineering, Mathematics or related field, or equivalent practical experience., * Real-time Processing: Experience with streaming and distributed messaging applications like Flink and Kafka. * Core Tech: Java programming. * Industrialise ML use cases * Data Visualization: Experience with QlikView or QlikSense to support BI initiatives. * Agile: Experience working in a fast-paced Scrum or Kanban environment. * Certifications: AWS Certified Data Engineer - Associate/Professional or AWS Certified Solutions Architect, Databricks Data engineer (Associated/Professional) certification * DevOps: Experience with Openshift, Github Actions or Jenkins for CI/CD of data workflows. ## Description We are seeking a seasoned Senior Data Engineer to design, build, and optimize our next-generation data platform. You will be responsible for architecting scalable data pipelines, managing large-scale distributed systems, and ensuring our data infrastructure in AWS and Databricks is robust and efficient. The ideal candidate is a Spark expert with a deep understanding of the AWS ecosystem and a passion for automation., * Pipeline Architecture: Design and implement complex batch and streaming ETL/ELT pipelines using Python, SQL, and Spark to process massive datasets. * Cloud Infrastructure: Leverage AWS Data Analytics services to build scalable, secure, and cost-effective data solutions. * Orchestration & DevOps: Manage and automate data workflows using Airflow, while utilizing Docker and ECS for containerized application deployment. * System Optimization: Monitor and tune the performance of distributed systems (Spark Cluster) to ensure high availability and low latency. * Infrastructure as Code: Utilize AWS CloudFormation or Terraform to manage data infrastructure, ensuring repeatable and version-controlled environments. * Cost Optimization: Monitor and optimize AWS spend by selecting appropriate instance types (Spot vs. On-Demand) and refining data storage strategies. * Security & Compliance: Implement IAM roles, bucket policies, and encryption (KMS) to ensure data is secure at rest and in transit. * Collaboration: Work within an Agile framework to deliver iterative value, collaborating closely with Data Scientists and Stakeholders to translate business needs into technical reality. JOB DIMENSIONS List of direct reports: * Up to 2 Direct Reports, and around 15 externals Key interfaces, stakeholders and relationships: * Internal: + GDS: product manager, application manager, data & analytics & AI team + Country business stakeholders * External : 3rd party vendors, o Amazon S3: Implementing "Data Lake" best practices, including partitioning, compression (Parquet/Avro), and lifecycle policies. o Amazon Redshift: Designing star/snowflake schemas and optimizing query performance for high-volume data warehousing. o Amazon Athena: Performing ad-hoc SQL analysis directly on S3 data. o Experience with open table formats (iceberg/delta) + Orchestration & Integration: o Amazon MWAA (Managed Workflows for Apache Airflow): Deploying and scaling Airflow environments. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [From event streaming to event sourcing 101](https://www.wearedevelopers.com/videos/91-from-event-streaming-to-event-sourcing-101) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) - [Introducing JSON Structure](https://www.wearedevelopers.com/videos/100219-introducing-json-structure) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best Coding Boot Camps in Germany](https://www.wearedevelopers.com/magazine/237-best-coding-boot-camps-in-germany) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)