> Markdown version of [/jobs/ext/2269437-data-engineer](https://www.wearedevelopers.com/jobs/ext/2269437-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Shape Your Future with Us - **Location:** Frankfurt am Main, Germany - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Big Data, Continuous Integration, Data Architecture, Information Engineering, Data Governance, Data Transformation, Data Structures, Data Systems, DevOps, Python (Programming Language), Performance Tuning, Search Technologies, Data Processing, Data Ingestion, Retrieval-Augmented Generation, Delivery Pipeline, Large Language Models, Apache Spark, Generative AI, Data Lakes, Pyspark, Semi-structured Data, Data Management, Data Pipelines, Databricks - **Published:** August 27, 2026 - **Apply:** https://www.adzuna.de/details/5856680141 ## About the Role * Strong hands-on experience with Databricks in production environments. * In-depth knowledge of Apache Spark and PySpark, including performance tuning and optimization. * Strong Data Engineering background, with experience building production-grade data pipelines. * Solid understanding of data modeling, data structures, and modern data architectures. * Proven experience processing large-scale datasets. * Experience working with unstructured and semi-structured data. * Practical experience preparing and transforming data for AI/ML and RAG use cases. * Strong Python skills, particularly for data engineering and PySpark development. * Experience with data ingestion, transformation, orchestration, and pipeline automation. * Ability to work independently in a fast-paced banking/enterprise environment. * Candidate must be located within the EU. Nice-to-Have: * Experience with Generative AI / LLM / RAG architectures. * Knowledge of vector search, embeddings, chunking, and document-processing pipelines. * Experience with Delta Lake / Delta tables and modern lakehouse architectures. * Experience with cloud platforms such as Azure, AWS, or GCP. * Experience in banking or other regulated financial-services environments. * Knowledge of data governance, security, lineage, and compliance requirements. * Experience with CI/CD and DevOps practices for data platforms. ## Description We are looking for a strong, hands-on Data Engineer with deep Databricks and Apache Spark experience to support a Frankfurt-based banking client. The ideal candidate will have extensive experience designing and implementing scalable data engineering solutions using Databricks, Spark/PySpark, with a strong understanding of data modeling and modern data architectures. Experience working with unstructured and semi-structured data for AI/ML and RAG use cases is highly desirable., * Design, develop, and optimize data engineering pipelines and data processing solutions using Databricks and Apache Spark. * Build scalable and reliable data pipelines using PySpark and related Spark technologies. * Work with large and complex datasets across structured, semi-structured, and unstructured data sources. * Design and implement effective data models to support analytics, AI/ML, and downstream data consumption. * Process and transform unstructured and semi-structured data for AI-driven use cases, including RAG (Retrieval-Augmented Generation). * Develop data ingestion, transformation, cleansing, and enrichment workflows. * Optimize Spark jobs and Databricks workloads for performance, scalability, and reliability. * Work closely with data scientists, ML/AI engineers, architects, and business stakeholders to deliver production-ready data solutions. * Apply strong engineering practices around data quality, testing, monitoring, and operational reliability. * Contribute to the design and evolution of modern cloud-based data platforms. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) ## Related Articles - [Data Analyst Salary Germany](https://www.wearedevelopers.com/magazine/277-data-analyst-salary-germany) - [The Biggest German Tech Companies](https://www.wearedevelopers.com/magazine/424-the-biggest-german-tech-companies) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Backend Developer Salary in Germany [2023]](https://www.wearedevelopers.com/magazine/196-backend-developer-salary-in-germany-2023)