Ai Data Engineer

Pst Ag
Madrid, Spain
5 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Continuous Integration Data Validation Web Scraping Data Mining DevOps JSON Python (Programming Language) NoSQL Open Source Technology SQL Databases XPath
+6 more
Large Language Models Rate Limiting Data Lakes Information Technology Playwright Docker

Job description

Descripción del trabajo Key Responsibilities Specification-Driven Extraction EngineeringDesign and maintain declarative extraction specifications-using Pydantic models, JSON schemas, or domain-specific languages-that describe exactly which fields to capture, their types, and validation rules.Implement pipelines that translate these specifications into executable extraction plans, leveraging both classical (Scrapy, Playwright) and AI-augmented (LLM-based semantic parsing) backends.Build reusable specification libraries for recurring data types (product prices, tariff codes, regulatory texts) to accelerate onboarding of new sources.Design and implement autonomous data extraction agents that can make decisions about source selection, retry logic, and parsing strategies Autonomous & Self-Healing SystemsDeploy self-healing spiders that automatically detect website layout changes and repair themselves using Model Context Protocol (MCP) servers (e.g., Scrapy MCP Server, Playwright MCP).Integrate semantic extraction (Scrapy-LLM, custom LLM pipelines) to eliminate selector brittleness-spiders rely on field descriptions, not fragile XPaths.Hands-on experience building AI agents and orchestration systems.Orchestrate complex, multi-step browsing workflows with agentic frameworks (BMAD/TEA, AutoGPT-like agents) that reason about page state, adapt to anti-bot measures, and correct their own behaviour in real time.Platform Thinking & ReusabilityMove beyond one-off scrapers: build a component-based extraction platform where selectors, login handlers, and pagination logic are shared, versioned, and tested.Implement monitoring, alerting, and automatic rollback for failed extraction runs.Champion ethical crawling by design-rate limiting, robots.txt respect, and compliance with GDPR/CCPA are built into the specification layer, not retrofitted.Collaboration & Continuous InnovationPartner with data scientists and domain experts to refine extraction specifications for complex, unstructured domains (e.g., legal texts, tariff classifications).Evaluate and pilot emerging tools to push automation coverage beyond 90%.Document and evangelise specification-driven best practices across the engineering organisation.QualificationBachelor’s degree in Computer Science 3+ years of experience in web scraping or data extraction Required SkillsProficiency with Python Experience with specification-Driven Extraction Experience with LangChain, LangGraph, LlamaIndex, AutoGen Hands on use of Scrapy LLM, Scrapy MCP Server, or similar systems that decouple field definitions from page structure Familiarity with frameworks that give LLMs browser control (Playwright + MCP, BMAD/TEA) to handle complex, non deterministic crawling tasks.Classical Scraping Fundamentals Data Validation & Storage - Ability to define validation rules within specifications and land clean data into SQL/NoSQL databases or data lake Basic API integration and authentication flows.HTTP, DOM, XPath, CSS.Nice to HavesContributions to open-source scraping or AI-automation projects.Contributions to open-source scraping or AI-automation projects.Familiarity with data privacy engineering (GDPR, CCPA) baked into specification design.DevOps light - Docker, CI/CD for testing extraction specifications.Ofertas similaresDiseñar, implementar y mantener pipelines de datos eficientes y escalables que den soporte a reporting, BI, IA y ML.Garantizar la calidad, integridad, disponibilidad y seguridad de los datos median…Crear una alerta de empleo para esta búsquedaLast month, haddock users consumed.This isn’t AI hype, it’s our AI workforce, running 24/7 for thousands of restaurants across Spain.And we’re just getting started.We’re looking for an AI Engineer …#J-*****-Ljbffr

Requirements

Bachelor’s degree in Computer Science 3+ years of experience in web scraping or data extraction Required Skills Proficiency with Python Experience with specification-Driven Extraction Experience with LangChain, LangGraph, LlamaIndex, AutoGen Hands on use of Scrapy LLM, Scrapy MCP Server, or similar systems that decouple field definitions from page structure Familiarity with frameworks that give LLMs browser control (Playwright + MCP, BMAD/TEA) to handle complex, non deterministic crawling tasks. Classical Scraping Fundamentals Data Validation & Storage - Ability to define validation rules within specifications and land clean data into SQL/NoSQL databases or data lake Basic API integration and authentication flows.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

3:18 min

Identifying deceptive server content and parsing messy HTML

Vidas Bacevičius Vidas Bacevičius · World Congress 2025

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

3:19 min

Automating data extraction pipelines using natural language assistants

Vidas Bacevičius Vidas Bacevičius · World Congress 2025

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes · LIVE

Videos

See all

Related articles

See all