> Markdown version of [/jobs/ext/2140241-ai-data-engineer](https://www.wearedevelopers.com/jobs/ext/2140241-ai-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Data Engineer - **Company:** Capgemini - **Location:** London, UK (Remote available) - **Salary:** £67,263.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Business Analytics Applications, Automation of Tests, Unit Testing, Cloud Storage, Code Review, Continuous Integration, Information Engineering, Extract Transform Load (ETL), DevOps, Distributed Computing Environment, Python (Programming Language), Search Technologies, Software Engineering, SQL Databases, Data Streaming, Data Logging, Large Language Models, Git, SC Clearance, Apache Kafka, Spark Streaming, Data Management, Virtual Agents, Software Version Control, Data Pipelines - **Published:** August 20, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5847552208 ## About the Role * Hands-on experience delivering data pipelines and data platforms in production environments. * Strong Python and SQL, with sound software engineering practice (Git, code review, unit testing, CI/CD) and the ability to troubleshoot production issues. * Experience with distributed data processing and modern lakehouse/warehouse patterns, with good data modelling and performance instincts. * Experience with cloud data services (storage, compute and orchestration) on at least one major cloud. * Practical use of AI coding assistants in a real engineering workflow, with an understanding of data confidentiality and secure, responsible use. * A continuous-learning mindset and clear communication with technical and non-technical colleagues. Nice to have * Exposure to GenAI/agentic building blocks: RAG, embeddings and vector search, LLM orchestration (e.g. LangGraph, LlamaIndex, Semantic Kernel), or LLM evaluation/observability (e.g. LangSmith, Ragas). * Streaming and event-driven experience (e.g. Kafka, Spark Structured Streaming). * Relevant cloud and/or data engineering certifications. ## Description To obtain SC clearance, the successful applicant must have resided continuously within the United Kingdom for the last 5 years, along with other criteria and requirements. Throughout the recruitment process, you will be asked questions about your security clearance eligibility such as, but not limited to, country of residence and nationality. Some posts are restricted to sole UK Nationals for security reasons; therefore, you may be asked about your citizenship in the application process. The Focus of Your Role You'll build and run the data products that modern AI depends on. That means classic, high-quality data engineering (batch and streaming pipelines, lakehouse and warehouse patterns, well-governed and auditable datasets), and the newer discipline of engineering data for AI and agents: retrieval pipelines, embeddings and vector stores, feature and context preparation, and the observability needed when the systems consuming your data are non-deterministic and can't simply be unit tested. You'll build and run the data products that modern AI depends on. That means classic, high-quality data engineering (batch and streaming pipelines, lakehouse and warehouse patterns, well-governed and auditable datasets), and the newer discipline of engineering data for AI and agents: retrieval pipelines, embeddings and vector stores, feature and context preparation, and the observability needed when the systems consuming your data are non-deterministic and can't simply be unit tested. What you'll be doing: * You'll work within agreed security and compliance boundaries throughout, which matters in the regulated and public sector environments we operate in. * Build and maintain data pipelines and data products. Use appropriate ETL/ELT and distributed processing to ingest, transform and curate trusted data in cloud storage and analytics platforms. * Engineer data for AI and agentic use cases. Prepare curated, governed datasets and features for analytics and ML; build retrieval and embedding pipelines (vector stores, chunking, metadata) that serve RAG and agent workflows; and help expose data to AI systems through patterns such as APIs and the Model Context Protocol (MCP). * Apply data modelling, quality and governance. Develop scalable models, apply validation and quality checks, and maintain lineage and documentation so data products are reliable and auditable. This matters even more when an AI agent, not just a human, is acting on them. * Build in observability for non-deterministic systems. Implement logging, alerting and SLAs for production pipelines, and contribute to evaluation and monitoring of AI-facing data and outputs (e.g. drift, retrieval quality, anomaly detection) with clear human review. * Use AI-assisted engineering responsibly. Accelerate development, testing and documentation with coding assistants, applying good prompt hygiene, protecting confidential data, and owning the results. * Collaborate across teams. Work with business analysts, platform engineers, data scientists and DevOps to deliver secure, well-tested solutions in an agile environment, communicating your design decisions clearly. * Keep raising the bar. Use version control, code review, automated testing and CI/CD; share knowledge, contribute to accelerators and standards, and keep current with modern data and AI engineering practice. What You Will Bring and Experience Needed You'll bring solid, hands-on experience delivering data engineering in production, and a genuine interest in how AI and agentic systems change what good data engineering looks like. You don't need to have done all of the AI-specific work below already, but you should be eager to, and able to show the engineering fundamentals that make it possible. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [AI Killed DevOps... What Now? - Lee Faus](https://www.wearedevelopers.com/videos/1759-ai-killed-devops-what-now-lee-faus) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)