NLP Post-doc / Engineer for Information Mining in Historical Data

Inria
Paris, France
14 days ago
Apply on jobs.inria.fr
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
French
Job source

Tech stack

Human-Computer Interaction Information Retrieval Natural Language Processing Systems Integration Large Language Models Data Pipelines

Job description

This position is part of the ANR project DECIDON (Débats parlementaires et Espace médiatique (1870-1940) : comprendre la CIrculation du discours politique grâce à des méthodes à forte intensité de DONnées), which is a collaboration between multiple French team, led at Epita Paris by Marie Puren, and including the Inria ALMAnaCH (Inria Paris center) project team as a work package lead for natural language processing techniques. The objective of the project is to enable a better understanding of how the parliamentary debates are being set up in the Third Republic and how thematics can appear and disappear across the newspaper and the parliamentary debate proceedings. The ALMAnaCH team role specifically focus on:

  • Development of an annotation interface to support the creation of thematic datasets across the corpus of parliamentary debates and, potentially, the press.
  • Development of an easily deployable approach for robust information retrieval across the corpus, targeting topics whose vocabulary may have evolved over time and diverged from contemporary French.
  • Evaluation of information retrieval methods for topic-based search on a historically situated corpus (including RAG).

The employee will be supervised by Thibault Clérice, permanent researcher at Inria, and will work in close collaboration with other members of the team and project, namely Florian Cafiero and Marie Puren (EPITA)

They will work at Inria, within the ALMAnaCH project-team. Within the team, they will find researchers connected to the topic outside of the project itself, including Cecilia Graiff, a PhD student in the team, supervised by Benoît Sagot and Chloé Clavel, working on multilingual and cross-cultural automatic analysis of argumentation structures in political debates.

Participation to national meetings and national/international conference are to be expected.

Mission confiée

The post-doc / NLP engineer will design and train models to support historians in constructing focused sub-corpora from large, noisy, OCR’d historical text collections. The work involves:

  • Setting up an annotation workflow (true/false positive labeling of keyword occurrences in context) in collaboration with historians;
  • Training and evaluating small-scale classifiers - ranging lightweight models to larger pretrained models - capable of distinguishing relevant from irrelevant occurrences of ambiguous or polysemous terms in diachrony (e.g., distinguishing an “anti-parliamentary” use of réforme de l’État from a routine administrative reform);
  • Integrating active learning so that model performance improves iteratively as historians annotate, minimizing labeling effort while maximizing corpus quality;
  • Prioritizing model deployability: given the scale of the corpus (at least dozens of millions of tokens across noisy OCR output) and the need for the tool to run efficiently and reproducibly within a web interface used directly by non-specialist historians, the research should focus on enabling this on lightway models: fast at inference, and easy to retrain/update rather than relying on large LLM inference at scale;
  • Benchmarking against and complementing RAG-based exploration (T5.5), providing a transparent, low-cost alternative for corpus-scoping that historians can audit and reproduce before moving to more exploratory or generative tasks.

The research component focuses on efficient, low-resource sequence classification for historical/noisy text - including handling class imbalance, domain-shift across a ~70-year corpus, and figurative/contextual language. Publishable outputs (methods, benchmarks, and the resulting tool) are expected as part of the project’s open, FAIR-data pipeline.

Requirements

  • Python
  • NLP Model Training and evaluation, beyond just LLM

Languages :

  • French (Reading understanding to interact with the data)
  • English

Relational skills :

  • Good organizational skills.
  • Good interpersonal skills.

Additional skills considered an asset:

  • Knowledge of the French 3rd Republic or similar political systems.
  • Interefest in information retrieval
  • A strong interest in open science., * Feeling comfortable in an interdisciplinary environment, as well as a willingness to learn and listen, are essential qualities for success in this role.
  • Interest in open science issues.
  • A PhD thesis or master’s dissertation focusing on historical data or on information retrieval is an asset.
  • Interest in small models, beyond LLMs

Benefits & conditions

  • carrying out research on the topic outlined above, both in the development of new ideas, positioning with respect to related work and validation of the methodology via experiments and analysis
  • Producing a solution that will be deployable for post-lexical information retrieval
  • Interact with the project’s historians as well as the researchers from the work package on natural language processing
  • the presentation of work both internally to colleagues and externally in the form of conference/journal/workshop papers
  • interacting and exchanging with colleagues on NLP topics, * Subsidized meals
  • Partial reimbursement of public transport costs
  • Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in working hours) + possibility of exceptional leave (sick children, moving home, etc.)
  • Possibility of teleworking and flexible organization of working hours
  • Professional equipment available (videoconferencing, loan of computer equipment, etc.)
  • Social, cultural and sports events and activities
  • Access to vocational training
  • Social security coverage

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.inria.fr
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff · World Congress 2024

3:00 min

Designing data ingestion architecture with system integration

Eldert Grootenboer +1 · World Congress 2023

2:16 min

Enhancing code coverage and information retrieval with automation

Andrew Boyagi Andrew Boyagi · World Congress 2025

2:08 min

Applying large language models to infrastructure tasks

Alfonso Sandoval Rosas Alfonso Sandoval Rosas · Europe 2026 Virtual

47 sec

Building modern data pipelines for legacy exports

Dr. Alexander Wachtel Dr. Alexander Wachtel +1 · World Congress 2025

3:09 min

Assisting manual license research with artificial intelligence tools

Uwe Korn Uwe Korn · World Congress 2026 Europe

Videos

See all

Related articles

See all