RAG / AI Data Engineer

Diagonal recruitment
Greater London, UK
1 day ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Automated Storage and Retrieval Systems Data Architecture Extract Transform Load (ETL) Python (Programming Language) Machine Learning Standard Sql Unstructured Data Large Language Models Generative AI Build Management Data Pipelines

Job description

  • Design and build data pipelines to support AI applications
  • Develop Retrieval-Augmented Generation (RAG) architectures
  • Create and maintain vector databases and knowledge repositories
  • Structure, clean and prepare data for AI consumption
  • Improve retrieval accuracy, relevance and performance
  • Build scalable data foundations for AI products and agents
  • Collaborate with architects, engineers and product teams to enable AI delivery

Requirements

  • Python
  • SQL
  • Vector databases (Pinecone, Weaviate, Qdrant, Chroma or similar)
  • Embedding models and retrieval frameworks
  • LangChain, LlamaIndex or equivalent
  • Data pipeline and ETL tooling, * Building data pipelines for AI or ML applications
  • Designing or implementing RAG architectures
  • Working with vector databases
  • Managing structured and unstructured datasets
  • Optimising retrieval quality and search performance
  • Agent-based AI systems and workflows, * Strong understanding of data architecture and retrieval systems
  • Able to balance accuracy, performance and scalability
  • Comfortable working across structured and unstructured datasets
  • Interested in practical AI implementation rather than theoretical research
  • Focused on creating reliable foundations for AI systems
  • Strong problem-solving and analytical skills

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

38 sec

Introducing OpenRAG for custom data pipelines

Phil Nash · Coffee With Developers

4:06 min

Using ClickHouse as a foundation for fast analytics

Hellmar Becker Hellmar Becker · World Congress 2026 Europe

2:37 min

Tracing the evolution from early AI to generative AI

Mike Mike · World Congress 2025

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff · World Congress 2024

1:50 min

Simplifying generative AI deployments using the RagStack opinionated framework

David Leconte David Leconte +1 · World Congress 2024

1:56 min

Creating immutable blockchain tables using standard SQL syntax

Wei Hu Wei Hu · World Congress 2023

Videos

See all

Related articles

See all