Founding Data Engineer

Davis AI
Paris, France
17 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Data Deduplication Distributed Data Store Python (Programming Language) PostgreSQL Regression Testing Management of Software Versions HuggingFace Data Generation

Job description

Davis is an AI-native real estate company accelerating early-stage development and architectural design. Today developers coordinate four to five fragmented stakeholders over weeks or months. Soon they will need only one: Davis.

We turn every input that shapes a development decision into decision-ready outputs: investor-grade feasibility studies, investment analysis, and architect-certified designs, delivered in days. Every stage pairs our proprietary AI systems with expert review, so velocity never comes at the cost of reliability.

We closed a $5.5M pre-seed co-led by Heartcore Capital and Balderton Capital, with Yellow, Evantic and Entrepreneur First, alongside angels from the founding teams of Spacemaker, Black Forest Labs, Hugging Face, Supabase, Cleo and Spore Bio. We already work with leading developers and expect to support hundreds of projects over the coming year, deepening our research, our hiring, and our coverage of the development process end to end., You will own the data our foundation model learns from, a model we train from scratch to generate buildings as geometric graphs. Part of the corpus comes from real floorplans as images and PDFs that have to become clean, standardized graphs. A large part will be synthetic, procedurally generated building graphs, geometry and rendered floorplans, with controlled variation in style, scan noise, annotations and furniture, each kept with its ground-truth graph automatically., * Corpus from raw sources. Turn real floorplans (images, PDFs, scans) into clean, standardized building graphs, with the geometry and semantics that make them trainable.

  • Synthetic data generation. Explore strategies to expand the dataset with synthetic data.
  • Canonical representation and curation. Define the standardized representation, then filter, deduplicate, quality-score and validate at scale, with versioning, provenance and lineage.
  • Pretraining data and mixtures. Assemble the training datasets, design the mixture and the curriculum, and blend synthetic and real data for large-scale pretraining.
  • Data ablations. Train models to learn which data actually helps, read the results, and feed them back into the generator and the mixture.
  • Evaluation. Build the eval harness (datasets, metrics, regression tests, monitoring) that tracks data and model quality over time., * You have built the dataset, not just trained on it. You have personally built or generated the data used for a large pretraining run, from raw or synthetic sources, rather than only training on a dataset someone handed you.

Requirements

  • Senior and deeply hands-on. At least 5 years of strong experience, senior enough to architect the data stack and set strategy, but still coding the pipelines, running the experiments and doing the ablations yourself.
  • Data as a first-class problem. A track record where the data itself is the object: curation, filtering, deduplication, quality scoring, mixtures, synthetic generation.
  • Strong engineering. Deep Python, clean and typed code, async and concurrency, distributed data pipelines, TDD culture.

Benefits & conditions

  • Own the data end to end, from the generator to the pretraining mixture, as the person who defines what the model learns from.
  • Shape a foundation model from scratch, on a structured representation no one else is training on.
  • Work on genuinely hard problems, image to graph, synthetic worlds and large-scale pretraining, with your work reaching clients within days.
  • Competitive salary and meaningful equity, at a founding level, in an early-stage company.
  • Join a world-class team, a mix of AI researchers, engineers and architects backed by world-class VCs.

About the company

  • Computer vision and document understanding, images to structured output, OCR, layout extraction, segmentation, geometry extraction, vectorization, raster to vector, image to scene graph, 3D or CAD.
  • Graph and structured scientific data, molecular, protein or scene graphs, meshes, CAD, BIM, 3D geometry, or relational and structured world data.
  • Synthetic worlds and simulation, a structured state to a simulator to a renderer to synthetic images with perfect labels, then a perception model, with an eye on the sim-to-real gap.
  • Foundation models from scratch, real involvement in a large pretraining run, not only fine-tuning.
  • Public evidence of data ownership, lead on a dataset, a Hugging Face release, a dataset card, a technical blog on your pipeline, or a talk on data curation or synthetic data.
  • GIS and geometry, parcels, zoning layers, projections, computational geometry or constrained optimization.
  • Multi-country data, heterogeneous sources, localization and varying rules.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Loading talks and stories from around this role…