Post-Doctorant F/H

Inria
Montbonnot-Saint-Martin, France
22 days ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Compensation
€33,456.0
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Artificial Intelligence Computer Vision Machine Learning Video Editing Transfer Learning Large Language Models Variational Autoencoders

Job description

Recent advancements in generative AI, and in particular diffusion models [1,2], have significantly enhanced the capabilities of text-to-video (T2V) models [3,4], allowing users to produce richly varied and imaginative scenes from natural language descriptions. These systems demonstrate strong scene diversity and flexibility, making them attractive for applications in entertainment, simulation, and human-computer interaction.

However, a persistent limitation lies in their inability to enforce fine-grained conditioning and maintain strict consistency for specific visual elements. For example, while a T2V model can generate a “person walking in a park,” it struggles to ensure the persistent appearance of a specific object, character identity, or detailed attribute (such as a specific garment [5]) across complex poses and dynamic environmental interactions.

In contrast, highly specialized image and video editing systems-such as those designed for virtual try-on [5], face swapping, or precise object insertion-excel at fine-grained conditioning on target individuals or objects. They can adapt elements to morphology, pose, and texture details with remarkable realism. Yet, these specialized approaches generally operate in isolation, lacking the scene diversity and broader contextual awareness that foundational T2V models offer.

Bridging these two paradigms offers a powerful opportunity: to synthesize realistic, precisely controllable subjects and objects embedded within richly described, dynamic environments. To achieve this, novel alignment and editing techniques are required. Specifically, post-training with Reinforcement Learning (RL) presents a highly promising methodology to overcome these limitations. By leveraging RL during the post-training phase, foundation T2V models can be explicitly optimized to follow complex conditioning signals, enforce temporal consistency, and align with specific human-defined objectives for fine-grained editing tasks without sacrificing their generative diversity., The Postdoctoral Research Fellow will be responsible for the following main tasks. They will engage in Model Design and Development by designing and implementing novel architectures (e.g., Diffusion Models, Transformers, VAEs) specifically tailored for high-resolution, temporally consistent, and controllable video generation. A key focus is to develop conditional generation techniques to guide the Text-to-Video process using various complex inputs beyond a simple text prompt, such as image references, motion skeletons, semantic masks, or detailed scene descriptions. They will extensively research Video Editing and Manipulation, developing methods for high-fidelity post-generation video editing, allowing for non-destructive modification of generated videos (e.g., object replacement, style transfer, background alteration) while maintaining strong temporal consistency. Furthermore, they will investigate in-context editing mechanisms that enable precise changes to specific segments or objects within a generated video based on new text or image prompts. A core part of the role is Addressing Key T2V Challenges. This includes tackling the fundamental challenge of temporal coherence and consistency, ensuring that generated videos do not suffer from “flickering” or object identity changes across frames, and developing strategies to improve semantic fidelity, resolving issues where models misinterpret complex text prompts. They will also explore methods for efficient training and inference to manage the significant computational cost associated with high-resolution, long-duration video generation, and address the difficulties of data scarcity and bias through techniques like data augmentation or cross-modal transfer learning. Finally, they will perform Evaluation and Benchmarking, establishing rigorous quantitative and qualitative metrics to assess the quality, editability, and controllability of the developed models. The fellow is expected to prioritize Dissemination and Collaboration, which involves documenting research findings and publishing high-quality papers in top-tier machine learning and computer vision venues, actively participating in departmental seminars, and contributing to collaborative projects.

Requirements

Compétences techniques et niveau requis :We are seeking a motivated PhD candidate with a strong background in one or more the following areas :

  • speech processing, computer vision, machine learning,
  • solid programmming skills
  • interest in connecting AI with human cognition Prior experience with LLM, SpeechLMs, RL algorithms, or robotic platforms is a plus, but not mandatory

Langues : Anglais

Benefits & conditions

  • Restauration subventionnĂ©e
  • Transports publics remboursĂ©s partiellement
  • CongĂ©s: 7 semaines de congĂ©s annuels + 10 jours de RTT (base temps plein) + possibilitĂ© d’autorisations d’absence exceptionnelle (ex : enfants malades, dĂ©mĂ©nagement)
  • PossibilitĂ© de tĂ©lĂ©travail 90 jours/an fixes ou flottants et amĂ©nagement du temps de travail
  • Équipements professionnels Ă  disposition (visioconfĂ©rence, prĂȘts de matĂ©riels informatiques, etc.)
  • Prestations sociales, culturelles et sportives (Association de gestion des Ɠuvres sociales d’Inria)
  • AccĂšs Ă  la formation professionnelle
  • Participation Protection Sociale ComplĂ©mentaire sous conditions, Les candidatures doivent ĂȘtre dĂ©posĂ©es en ligne sur le site Inria.

Le traitement des candidatures adressĂ©es par d’autres canaux n’est pas garanti.

SĂ©curitĂ© dĂ©fense : Ce poste est susceptible d’ĂȘtre affectĂ© dans une zone Ă  rĂ©gime restrictif (ZRR), telle que dĂ©finie dans le dĂ©cret n°2011-1425 relatif Ă  la protection du potentiel scientifique et technique de la nation (PPST). L’autorisation d’accĂšs Ă  une zone est dĂ©livrĂ©e par le chef d’établissement, aprĂšs avis ministĂ©riel favorable, tel que dĂ©fini dans l’arrĂȘtĂ© du 03 juillet 2012, relatif Ă  la PPST. Un avis ministĂ©riel dĂ©favorable pour un poste affectĂ© dans une ZRR aurait pour consĂ©quence l’annulation du recrutement.

Politique de recrutement : Dans le cadre de sa politique diversité, tous les postes Inria sont accessibles aux personnes en situation de handicap.

About the company

A propos du centre ou de la direction fonctionnelle

Le centre de recherche Inria de l’UniversitĂ© Grenoble Alpes regroupe un peu moins de 600 personnes rĂ©parties au sein de 27 Ă©quipes de recherche et 8 services support Ă  la recherche.

Son effectif est distribuĂ© sur 3 campus Ă  Grenoble, en lien Ă©troit avec les laboratoires et les Ă©tablissements de recherche et d’enseignement supĂ©rieur (UniversitĂ© Grenoble Alpes, CNRS, CEA, INRAE, 
), mais aussi avec les acteurs Ă©conomiques du territoire.

PrĂ©sent dans les domaines du calcul et grands systĂšmes distribuĂ©s, logiciels sĂ»rs et systĂšmes embarquĂ©s, la modĂ©lisation de l’environnement Ă  diffĂ©rentes Ă©chelles et la science des donnĂ©es et intelligence artificielle, Inria Grenoble - RhĂŽne-Alpes participe au meilleur niveau Ă  la vie scientifique internationale par les rĂ©sultats obtenus et les collaborations tant en Europe que dans le reste du monde., Inria est l’institut national de recherche dĂ©diĂ© aux sciences et technologies du numĂ©rique. Il emploie 2600 personnes. Ses 215 Ă©quipes-projets agiles, en gĂ©nĂ©ral communes avec des partenaires acadĂ©miques, impliquent plus de 3900 scientifiques pour relever les dĂ©fis du numĂ©rique, souvent Ă  l’interface d’autres disciplines. L’institut fait appel Ă  de nombreux talents dans plus d’une quarantaine de mĂ©tiers diffĂ©rents. 900 personnels d’appui Ă  la recherche et Ă  l’innovation contribuent Ă  faire Ă©merger et grandir des projets scientifiques ou entrepreneuriaux qui impactent le monde. Inria travaille avec de nombreuses entreprises et a accompagnĂ© la crĂ©ation de plus de 200 start-up. L’institut s’eïŹ€orce ainsi de rĂ©pondre aux enjeux de la transformation numĂ©rique de la science, de la sociĂ©tĂ© et de l’économie.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.inria.fr

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:37 min

Optimizing technical profiles for AI sourcing and recruitment

Mina Golesorkhi Mina Golesorkhi · WWC Europe 2026

2:36 min

Applying supervised machine learning for practical rule extraction

Katja TrÀumner

5:55 min

Practical applications and use cases for computer vision

Flo Pachinger · LIVE

8:05 min

Processing the accessibility impacts of aggressive video editing styles

Chris Heilmann +2 · LIVE

2:11 min

Overcoming software challenges and future computer vision project integrations

Iulia Feroli Iulia Feroli · WWC Europe 2026

7:20 min

Artificial intelligence bypassing technical job interviews

Videos

See all

Related articles

See all