> Markdown version of [/jobs/ext/2991737-nlp-post-doc-engineer-for-information-mining-in-historical-data](https://www.wearedevelopers.com/jobs/ext/2991737-nlp-post-doc-engineer-for-information-mining-in-historical-data). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # NLP Post-doc / Engineer for Information Mining in Historical Data - **Company:** Inria - **Location:** Paris, France (Remote available) - **Contract:** Permanent contract - **Skills:** Human-Computer Interaction, Information Retrieval, Natural Language Processing, Systems Integration, Large Language Models, Data Pipelines - **Published:** September 19, 2026 - **Apply:** https://jobs.inria.fr/public/classic/fr/offres/2026-10490 ## About the Role * Python * NLP Model Training and evaluation, beyond just LLM Languages : * French (Reading understanding to interact with the data) * English Relational skills : * Good organizational skills. * Good interpersonal skills. Additional skills considered an asset: * Knowledge of the French 3rd Republic or similar political systems. * Interefest in information retrieval * A strong interest in open science., * Feeling comfortable in an interdisciplinary environment, as well as a willingness to learn and listen, are essential qualities for success in this role. * Interest in open science issues. * A PhD thesis or master's dissertation focusing on historical data or on information retrieval is an asset. * Interest in small models, beyond LLMs ## Description This position is part of the ANR project DECIDON (Débats parlementaires et Espace médiatique (1870-1940) : comprendre la CIrculation du discours politique grâce à des méthodes à forte intensité de DONnées), which is a collaboration between multiple French team, led at Epita Paris by Marie Puren, and including the Inria ALMAnaCH (Inria Paris center) project team as a work package lead for natural language processing techniques. The objective of the project is to enable a better understanding of how the parliamentary debates are being set up in the Third Republic and how thematics can appear and disappear across the newspaper and the parliamentary debate proceedings. The ALMAnaCH team role specifically focus on: * Development of an annotation interface to support the creation of thematic datasets across the corpus of parliamentary debates and, potentially, the press. * Development of an easily deployable approach for robust information retrieval across the corpus, targeting topics whose vocabulary may have evolved over time and diverged from contemporary French. * Evaluation of information retrieval methods for topic-based search on a historically situated corpus (including RAG). The employee will be supervised by Thibault Clérice, permanent researcher at Inria, and will work in close collaboration with other members of the team and project, namely Florian Cafiero and Marie Puren (EPITA) They will work at Inria, within the ALMAnaCH project-team. Within the team, they will find researchers connected to the topic outside of the project itself, including Cecilia Graiff, a PhD student in the team, supervised by Benoît Sagot and Chloé Clavel, working on multilingual and cross-cultural automatic analysis of argumentation structures in political debates. Participation to national meetings and national/international conference are to be expected. Mission confiée The post-doc / NLP engineer will design and train models to support historians in constructing focused sub-corpora from large, noisy, OCR'd historical text collections. The work involves: * Setting up an annotation workflow (true/false positive labeling of keyword occurrences in context) in collaboration with historians; * Training and evaluating small-scale classifiers - ranging lightweight models to larger pretrained models - capable of distinguishing relevant from irrelevant occurrences of ambiguous or polysemous terms in diachrony (e.g., distinguishing an "anti-parliamentary" use of réforme de l'État from a routine administrative reform); * Integrating active learning so that model performance improves iteratively as historians annotate, minimizing labeling effort while maximizing corpus quality; * Prioritizing model deployability: given the scale of the corpus (at least dozens of millions of tokens across noisy OCR output) and the need for the tool to run efficiently and reproducibly within a web interface used directly by non-specialist historians, the research should focus on enabling this on lightway models: fast at inference, and easy to retrain/update rather than relying on large LLM inference at scale; * Benchmarking against and complementing RAG-based exploration (T5.5), providing a transparent, low-cost alternative for corpus-scoping that historians can audit and reproduce before moving to more exploratory or generative tasks. The research component focuses on efficient, low-resource sequence classification for historical/noisy text - including handling class imbalance, domain-shift across a ~70-year corpus, and figurative/contextual language. Publishable outputs (methods, benchmarks, and the resulting tool) are expected as part of the project's open, FAIR-data pipeline. ## Related Videos - [Outclassing Frontier LLMs at Extracting Information](https://www.wearedevelopers.com/videos/100303-outclassing-frontier-llms-at-extracting-information) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [AI Agents & Agentic AI](https://www.wearedevelopers.com/videos/2017-ai-agents-agentic-ai) - [New AI-Centric SDLC: Rethinking Software Development with Knowledge Graphs](https://www.wearedevelopers.com/videos/1417-new-ai-centric-sdlc-rethinking-software-development-with-knowledge-graphs) - [OLAP for AI Applications and why you should care](https://www.wearedevelopers.com/videos/100212-olap-for-ai-applications-and-why-you-should-care) - [Inside the Mind of an LLM](https://www.wearedevelopers.com/videos/1617-inside-the-mind-of-an-llm) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)