AI Data Enablement Engineer

Xenon7
Spain
7 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow ARM Architecture Computer Vision Audit Trail Clinical Data Repository Information Engineering Data Governance Data Infrastructure Extract Transform Load (ETL) Data Warehousing Python (Programming Language)
+20 more
Meta-Data Management Query Optimization Role-Based Access Control Cloud Services Standard Sql Salesforce.Com SAP (Applications) Microsoft SharePoint SQL Databases Unstructured Data Large Language Models Multi-Agent Systems Caching Data Layers Build Management Pyspark Pure Data Streamlit Framework Data Pipelines Databricks

Job description

  • Design and build AI-ready data products on Databricks - trusted datasets with well-defined business semantics, KPIs, hierarchies, and business glossary alignment
  • Implement semantic layers and governed datasets that support both traditional BI consumption and natural-language querying by business users
  • Deploy and operate Databricks Genie spaces with Unity Catalog, tuning them for accuracy, adoption, and business relevance
  • Build RAG pipelines and conversational analytics applications grounded in governed enterprise data - including Streamlit or Databricks Apps that let business users query data without writing SQL
  • Engineer robust ETL/ELT pipelines (dbt, Airflow, PySpark) that produce and maintain the trusted data these AI experiences depend on
  • Implement data governance - RBAC, row/column-level security, masking, lineage, auditability, catalog and metadata management - in a regulated pharma environment
  • Optimize cost and performance on both the data platform side (warehouse sizing, cluster tuning, query optimization) and the AI side (token usage, caching, model routing)
  • Partner with Finance business stakeholders to translate domain requirements into semantic models and governed data products they can trust, * Pure Data Engineers who list Cortex or Genie as a skill but haven’t shipped it in production
  • AI/GenAI engineers whose center of gravity is LangChain agents or RAG-over-documents, without a strong governed data platform foundation
  • Computer vision, NLP model builders, or multi-agent orchestration specialists - wrong shape for this role

Requirements

Must-Have Experience

  • 5+ years hands-on data engineering on cloud data platforms - Databricks demonstrated in real project delivery, not skill-list-only
  • Direct hands-on experience with Databricks Genie - you have built, configured, and tuned these in production or advanced pilots, with specific reference to the flavors used (Genie spaces with semantic models)
  • Semantic layer / trusted data product delivery - you have built governed datasets that business users can rely on, with KPI definitions, hierarchies, and business glossary alignment
  • dbt, PySpark, SQL, Python - strong across the modern data stack
  • Orchestration with Airflow, Databricks Workflows, or equivalent
  • Data governance in regulated environments - RBAC, RLS, masking, lineage, auditability
  • Experience integrating structured and unstructured data (PDFs, SharePoint/Teams content, enterprise knowledge sources) into AI-enablement workflows

Nice to Have

  • Pharma, life sciences, or regulated financial services domain experience
  • Veeva CRM, IQVIA, SAP, or clinical data source integration
  • Streamlit or Databricks Apps for business-facing analytics
  • Databricks Data Engineer Professional certification
  • LangChain, LlamaIndex, or equivalent RAG frameworks
  • Cost optimization on both compute (warehouse/cluster) and LLM (tokens/caching/routing) dimensions

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:19 min

Introduction to Apache Airflow for advanced orchestration

Alan Mazankiewicz · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:11 min

Deploying and running Airflow in cloud environments

Alan Mazankiewicz · LIVE

1:59 min

Evolving roles in AI driven software teams

Ignacio Riesgo Ignacio Riesgo +1 · World Congress 2024

Videos

See all

Related articles

See all