Senior AI Data Engineer

TechBiz Global GmbH
France
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Artificial Intelligence Airflow Amazon Web Services Microsoft Azure Cloud Database Information Systems Data Cleansing Information Engineering Extract Transform Load (ETL) Data Security Data Systems
+20 more
Data Warehousing Data Flow Control Python (Programming Language) Machine Learning Search Technologies Software Engineering SQL Databases Data Storage Management Real Time Systems Large Language Models Apache Spark Event Driven Architecture Kubernetes Information Technology Apache Flink Apache Kafka Spark Streaming Data Management Stream Processing Data Pipelines

Job description

  • Design, build, and scale robust ETL/ELT pipelines optimized for AI workloads, including RAG, fine-tuning, and batch inference.
  • Transform unstructured data sources such as PDFs, logs, and transcripts into structured and vectorized formats suitable for LLM consumption.
  • Maintain and automate the data-to-model lifecycle, ensuring AI knowledge bases remain synchronized with changing business data.
  • Develop and maintain real-time feature pipelines that support low-latency AI and machine learning applications.
  • Integrate data platforms with Kafka and other event-driven systems to enable real-time processing and AI-driven responses.
  • Manage and optimize Feature Stores to ensure consistency between model training and production environments.
  • Implement automated data quality controls and validation processes to ensure the reliability and accuracy of AI training and inference data.
  • Establish and maintain data lineage frameworks to provide traceability, auditability, and regulatory compliance across data workflows.
  • Enforce data security, privacy, and governance standards, including PII protection and compliance with industry regulations.
  • Manage data movement and synchronization across on-premises systems, cloud platforms, and data warehouses.
  • Optimize data storage and retrieval strategies for Vector Databases to support high-performance RAG and AI search workloads.
  • Collaborate with Data Scientists, ML Engineers, Software Engineers, and business stakeholders to deliver scalable AI data solutions.

Requirements

Do you have experience in Spark?, Do you have a Bachelor’s degree?, 10+ years of experience in Data Engineering or Backend Engineering with a strong focus on data platforms and pipelines.

  • 2+ years of hands-on experience supporting AI/ML data pipelines, including data preparation for machine learning and generative AI applications.
  • Expert-level proficiency in Python and SQL; experience with Java or Scala is an advantage.
  • Strong experience building and maintaining real-time data streaming solutions using Apache Kafka, Flink, or Spark Streaming.
  • Hands-on experience with modern data orchestration and transformation tools such as Airflow, dbt, and Prefect.
  • Experience working with Vector Databases and Feature Stores to support AI and machine learning workloads.
  • Strong knowledge of cloud-based data services on AWS, Azure, or GCP, including services such as Glue, Kinesis, Data Factory, or Dataflow.
  • Experience deploying and managing data workloads in Kubernetes (K8s) environments.
  • Proven experience handling sensitive data within regulated industries such as Fintech, Healthcare, or other compliance-driven environments.
  • Strong understanding of data quality, governance, security, and privacy best practices.
  • Bachelor’s degree in Computer Science, Software Engineering, Information Systems, or a related technical field. Equivalent practical experience will also be considered.
  • Excellent problem-solving skills and the ability to collaborate effectively with cross-functional engineering, data, and AI teams.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

2:04 min

Comparing offline data analytics with online stream processing

Artem Volk Artem Volk +1 · WWC 2024

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · WWC 2022

Videos

See all

Related articles

See all