> Markdown version of [/jobs/ext/1362475-ai-data-engineer](https://www.wearedevelopers.com/jobs/ext/1362475-ai-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Data Engineer - **Company:** Multiverse Computing - **Location:** Madrid, Spain - **Contract:** Temporary contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Microsoft Azure, Continuous Integration, Data Cleansing, Information Engineering, Data Governance, Extract Transform Load (ETL), Python (Programming Language), Machine Learning, Operational Databases, Software Tools, Cloud Services, DataOps, SQL Databases, Unstructured Data, Data Processing, Google Cloud, Data Storage Technologies, Data Ingestion, Large Language Models, Apache Spark, Generative AI, Information Technology, Data Pipelines, Data Generation - **Published:** July 21, 2026 - **Apply:** https://www.adzuna.es/contact-us.html ## About the Role Bachelor's degree in Computer Science, Data Engineering, Data Science, or a related technical field. 2+ years of experience designing and operating data pipelines for AI/ML products, including ETL, data cleaning, and integration. Proficiency in Python and SQL for data manipulation and pipeline development. Hands-on experience with data engineering tools and frameworks (e.g., Apache Airflow, Spark, dbt). Familiarity with cloud data platforms (AWS, Azure, GCP) and data storage solutions. Experience preparing synthetic datasets or simulating environments for ML when needed. Preferred Qualifications Experience supporting LLM or generative AI use cases with data engineering best practices. Working knowledge of vector databases and unstructured data management. Familiarity with data observability, data governance, and security in production environments. Experience collaborating with AI teams in research or enterprise settings. ## Description We are looking to fill this role immediately and are reviewing applications daily. Expect a fast, transparent process with quick feedback. Why join us? We are a European deep-tech leader in quantum and AI, backed by major global strategic investors and strong EU support. Our groundbreaking technology is already transforming how AI is deployed worldwide - compressing large language models by up to 95% without losing accuracy and cutting inference costs by 50-80%. Joining us means working on cutting-edge solutions that make AI faster, greener, and more accessible - and being part of a company often described as a "quantum-AI unicorn in the making.", Prepare and manage high-quality datasets for machine learning training and evaluation, ensuring they are organized, accessible, and reliable. Construct synthetic datasets and simulate task environments when production data is limited or unavailable, leveraging creative data generation techniques. Build and maintain efficient, scalable data pipelines that support continuous integration and flow of new data into production and research AI systems. Collaborate closely with data scientists and engineers to enable seamless data prep, transfer, and storage for ML experimentation and deployment. Monitor, validate, and optimize dataset quality and pipeline performance, introducing automation and best practices for data ingestion and transformation. Document data sources, pipeline logic, and data management decisions for transparency and team knowledge-sharing. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Let's Get Aggregated: Custom UDAFs in Spark ](https://www.wearedevelopers.com/videos/1649-let-s-get-aggregated-custom-udafs-in-spark) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)