Data Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+8 more
Job description
through to serving queries. The work includes:Large-scale ingestion. Build and operate high-throughput, resumable pipelines for large datasets, with efficient incremental updates, monitoring and recovery from failures.Document processing and data quality. Extract useful content from complex documents and other formats. Handle malformed records and changing schemas, and validate outputs while preserving structure and metadata.Search and serving. Build keyword, vector and structured search, and design schemas, indexes and partitioning for fast queries over tens to hundreds of millions of records.Connecting information across sources. Link patents, scientific records and supporting documents, preserve dates and versions, and make results traceable to their original sources.Performance engineering. Profile parsing, ingestion, database builds and queries throughout development, testing against representative datasets at realistic scale. Diagnose CPU, memory and storage I/O bottlenecks, and tune
Requirements
jobs and infrastructure for throughput, latency and cost.You’ll work closely with our AI researchers and product engineers, with substantial freedom to choose the approach and build the systems yourself.What you bringMust haves:Strong Python and SQL, with experience designing and operating production databases.Solid experience building and operating production data pipelines over large, messy datasets.Expertise with running search systems over large document collections.End-to-end ownership from raw data to user-facing functionality.A good understanding of schema design, indexing and query optimisation.A track record of diagnosing and fixing performance bottlenecks in live systems through profiling and measurement.Nice to Have:Experience with PostgreSQL/pgvector, OpenSearch (or Elasticsearch), Spark/Delta Lake, AWS, NoSQL databases, or Rust/C++ is useful.The FoundersYou’ll partner with a founding team of AI PhDs and elite systems engineers:Sanj (CRO): PhD in AI (Gatsby Unit, UCL), ex-Huawei R&D, former lead at Magic Carpet AI (acquired).Chris (CEO): PhD in AI (UCL), published researcher, ex-Dyson and Alan Turing Institute.Angus (CTO): MEng Computer Science, ex-Qualcomm and Coremont (Brevan Howard).What we offerCompetitive Salary + Significant Equity: We want you to have true ownership in the success of the company.Founding Impact: You’ll have a direct hand in how we build out the data infrastructure the rest of the product depends on.Support: Full visa sponsorship and private medical insurance.The Environment: Free meals and a seat at the table with an incredibly smart, ambitious team. #J-18808-Ljbffr
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 121 - AI goes offline
Dev Digest 120 - Apple and peers
How to Become an AI Engineer
Where To Find Software Engineering Jobs