Data Engineer

Shape Your Future with Us
Frankfurt am Main, Germany
7 days ago
Apply on www.adzuna.de
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Big Data Continuous Integration Data Architecture Information Engineering Data Governance Data Transformation Data Structures Data Systems DevOps
+16 more
Python (Programming Language) Performance Tuning Search Technologies Data Processing Data Ingestion Retrieval-Augmented Generation Delivery Pipeline Large Language Models Apache Spark Generative AI Data Lakes Pyspark Semi-structured Data Data Management Data Pipelines Databricks

Job description

We are looking for a strong, hands-on Data Engineer with deep Databricks and Apache Spark experience to support a Frankfurt-based banking client.

The ideal candidate will have extensive experience designing and implementing scalable data engineering solutions using Databricks, Spark/PySpark, with a strong understanding of data modeling and modern data architectures. Experience working with unstructured and semi-structured data for AI/ML and RAG use cases is highly desirable., * Design, develop, and optimize data engineering pipelines and data processing solutions using Databricks and Apache Spark.

  • Build scalable and reliable data pipelines using PySpark and related Spark technologies.
  • Work with large and complex datasets across structured, semi-structured, and unstructured data sources.
  • Design and implement effective data models to support analytics, AI/ML, and downstream data consumption.
  • Process and transform unstructured and semi-structured data for AI-driven use cases, including RAG (Retrieval-Augmented Generation).
  • Develop data ingestion, transformation, cleansing, and enrichment workflows.
  • Optimize Spark jobs and Databricks workloads for performance, scalability, and reliability.
  • Work closely with data scientists, ML/AI engineers, architects, and business stakeholders to deliver production-ready data solutions.
  • Apply strong engineering practices around data quality, testing, monitoring, and operational reliability.
  • Contribute to the design and evolution of modern cloud-based data platforms.

Requirements

  • Strong hands-on experience with Databricks in production environments.
  • In-depth knowledge of Apache Spark and PySpark, including performance tuning and optimization.
  • Strong Data Engineering background, with experience building production-grade data pipelines.
  • Solid understanding of data modeling, data structures, and modern data architectures.
  • Proven experience processing large-scale datasets.
  • Experience working with unstructured and semi-structured data.
  • Practical experience preparing and transforming data for AI/ML and RAG use cases.
  • Strong Python skills, particularly for data engineering and PySpark development.
  • Experience with data ingestion, transformation, orchestration, and pipeline automation.
  • Ability to work independently in a fast-paced banking/enterprise environment.
  • Candidate must be located within the EU.

Nice-to-Have:

  • Experience with Generative AI / LLM / RAG architectures.
  • Knowledge of vector search, embeddings, chunking, and document-processing pipelines.
  • Experience with Delta Lake / Delta tables and modern lakehouse architectures.
  • Experience with cloud platforms such as Azure, AWS, or GCP.
  • Experience in banking or other regulated financial-services environments.
  • Knowledge of data governance, security, lineage, and compliance requirements.
  • Experience with CI/CD and DevOps practices for data platforms.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.de
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

Videos

See all

Related articles

See all