> Markdown version of [/jobs/ext/1221377-senior-data-engineer-ai-ml-platform](https://www.wearedevelopers.com/jobs/ext/1221377-senior-data-engineer-ai-ml-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Engineer, AI/ML Platform - **Company:** Roamler - **Location:** Amsterdam, Netherlands - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Automation of Tests, Code Review, Data Validation, Web Scraping, Extract Transform Load (ETL), Identity and Access Management, Python (Programming Language), Operational Databases, Cloud Services, Azure Machine Learning, Software Engineering, SQL Databases, Systems Integration, Large Language Models, Prompt Engineering, Apache Spark, Build Management, Information Technology, Playwright, Virtual Agents, Amazon Simple Queue Service (SQS), Terraform, Data Pipelines - **Published:** July 10, 2026 - **Apply:** https://nl.indeed.com/viewjob?jk=5b9fc687486eab17 ## About the Role * Bachelor's and/or Master's degree in computer science or a related field. * 5+ years of experience as a Data Engineer with strong software engineering skills, including ownership of production data pipelines and/or web scraping and crawling systems. * Design, build, and operate scalable data pipelines on AWS that bring in data from external web sources and turn it into clean, queryable datasets with Playwright and/or BeautifulSoup. * Experience with Airflow, Python, SQL, and Spark. * Hands-on, practical experience integrating LLMs into production systems: prompt design as part of system design, evaluation, and cost/latency tradeoffs, not just experimenting in a notebook. * Experience with AI agent frameworks (for example browser automation agents) or a strong interest in and aptitude for learning them. * Manage infrastructure as code with Terraform on AWS (ECS, EMR, Glue, S3, SQS/SNS, IAM) * Develop and maintain Spark jobs on EMR for batch ETL and enrichment at scale. * Some exposure to classical ML/NLP concepts (classification, embeddings, matching) is a plus, since our data science team's pipeline uses these heavily. * Result-driven and hands-on attitude. * Proactive and courage to speak up. * Ability to communicate in English, verbally and in writing. ## Description * Own and evolve our data ingestion pipelines and our web scraping engine as production systems: architecture, reliability, scalability, and performance, not just individual scripts. * Design and build AI agent based extraction workflows (for example browser automation agents that navigate outlet websites, menus, and listing platforms) where they outperform traditional scraping and parsing approaches. * Build evaluation frameworks and quality checks so AI-driven extraction and enrichment can be trusted at the same bar as our existing rule-based systems. * Utilize AWS & Azure services effectively across both the traditional pipeline (Airflow, Spark, Glue) and the AI-driven components (model hosting, agent orchestration). * Enhance our automated testing, data validation, and monitoring capabilities, including for AI components, which introduce new failure modes and cost profiles compared to deterministic scraping and rule-based logic. * Collaborate with business and data stakeholders, including our data science team, to align requirements and expectations across the full pipeline. * Drive engineering projects end to end and ensure successful outcomes. * Drive continuous improvement in our engineering setup. * Foster a collaborative and enjoyable team environment through code reviews, knowledge sharing, and daily interactions. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [WeAreDevelopers LIVE - CSS is DOOMed](https://www.wearedevelopers.com/videos/1838-wearedevelopers-live-css-is-doomed) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [7 Most Popular Web Developer Jobs in Europe](https://www.wearedevelopers.com/magazine/163-7-most-popular-web-developer-jobs-in-europe) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)