> Markdown version of [/jobs/ext/2951886-data-engineer](https://www.wearedevelopers.com/jobs/ext/2951886-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Solve Intelligence - **Location:** London, UK - **Salary:** £83,016.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Systems Engineering, Big Data, C++ (Programming Language), Databases, Elasticsearch, Python (Programming Language), PostgreSQL, Metadata, NoSQL, Operational Databases, Query Optimization, Raw Data, SQL Databases, Apache Spark, Indexer, Gatsby, Data Lakes, AI Platforms, Information Technology, Low Latency - **Published:** September 17, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5887444329 ## About the Role * Strong Python and SQL, with experience designing and operating production databases. * Solid experience building and operating production data pipelines over large, messy datasets. * Expertise with running search systems over large document collections. * End-to-end ownership from raw data to user-facing functionality. * A good understanding of schema design, indexing and query optimisation. * A track record of diagnosing and fixing performance bottlenecks in live systems through profiling and measurement. Nice to Have: * Experience with PostgreSQL/pgvector, OpenSearch (or Elasticsearch), Spark/Delta Lake, AWS, NoSQL databases, or Rust/C++ is useful. You'll partner with a founding team of AI PhDs and elite systems engineers: * Sanj (CRO): PhD in AI (Gatsby Unit, UCL), ex-Huawei R&D, former lead at Magic Carpet AI (acquired). * Chris (CEO): PhD in AI (UCL), published researcher, ex-Dyson and Alan Turing Institute. * Angus (CTO): MEng Computer Science, ex-Qualcomm and Coremont (Brevan Howard). ## Description * Traction: 20-30% MoM revenue growth; selling to 700+ global IP teams (DLA Piper, tech boutiques, and global enterprises). * Proven Value: Users report 50-90% efficiency gains using our AI platform. * Backing: Recently featured in Sifted following our $40M Series B announcement, bringing our total funding to $55M from elite investors including Y Combinator, 20VC, Visionaries, Microsoft and Thomson Reuters. We're hiring a data engineer to build the ingestion and search systems behind Solve Intelligence's AI products. Our sources include global patent literature, case law, technical standards and contributions, scientific databases, academic papers, and content from across the web. The data spans structured records, documents, images, audio and video. You'll work across bulk ingestion and on-demand retrieval, making this information searchable and useful in our products You'll own systems from source acquisition through to serving queries. The work includes: * Large-scale ingestion. Build and operate high-throughput, resumable pipelines for large datasets, with efficient incremental updates, monitoring and recovery from failures. * Document processing and data quality. Extract useful content from complex documents and other formats. Handle malformed records and changing schemas, and validate outputs while preserving structure and metadata. * Search and serving. Build keyword, vector and structured search, and design schemas, indexes and partitioning for fast queries over tens to hundreds of millions of records. * Connecting information across sources. Link patents, scientific records and supporting documents, preserve dates and versions, and make results traceable to their original sources. * Performance engineering. Profile parsing, ingestion, database builds and queries throughout development, testing against representative datasets at realistic scale. Diagnose CPU, memory and storage I/O bottlenecks, and tune jobs and infrastructure for throughput, latency and cost. You'll work closely with our AI researchers and product engineers, with substantial freedom to choose the approach and build the systems yourself. ## Related Videos - [How Gatsby Cloud's real-time streaming architecture drives <5 second builds](https://www.wearedevelopers.com/videos/418-how-gatsby-cloud-s-real-time-streaming-architecture-drives-5-second-builds) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Optimizing Discovery: PostgreSQL's Role in Transforming GetYourGuide's Search](https://www.wearedevelopers.com/videos/1647-optimizing-discovery-postgresql-s-role-in-transforming-getyourguide-s-search) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [TiDB, One Layer at a Time: How Distributed SQL Became an Agentic AI Backbone](https://www.wearedevelopers.com/videos/100117-tidb-one-layer-at-a-time-how-distributed-sql-became-an-agentic-ai-backbone) - [Dynamic Entities in .NET: Building Low-Code Systems on Top of Entity Framework Core](https://www.wearedevelopers.com/videos/100218-dynamic-entities-in-net-building-low-code-systems-on-top-of-entity-framework-core) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)