F/M Visualization of the Plausibility and Bias for Data Resources

Inria
Rocquencourt, France
3 days ago
Apply on jobs.inria.fr
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
€27,600.0
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Artificial Intelligence Data Visualization Machine Learning

Job description

Typical requirements analysis-design-implementation-evaluation approach that is used in most of HCI and visualization. Expected result : Prototypical implementations in the context and within the framework of the EDT project (https://edtlab.fr/en/), study results, scientific publications. Collaborations will be established within the context of the EDT project (https://edtlab.fr/en/), depending on the specific development of the project. Generally, members of AVIZ are well embedded into the international research community and have active collaborations with many researchers worldwide.

Requirements

  • well-motivated student with interest in interactive data visualization and analysis
  • good written and spoken communication in English to be able to interact within the research team, to be able collaborate with domain scientists, and to be able to disseminate the results in scientific publications
  • excellent skills in actual manual (non-AI-based) programming
  • experience in the use of machine learning and artificial intelligence
  • experience (but not necessarily academic) in data visualization
  • experience in HCI and empirical evaluation preferred but not required
  • past publications preferred but not required

English level : Advanced

Benefits & conditions

  • analysis and visualization of the plausibility of crowd-sourced and/or heterogenous geographic data
  • use of artificial intelligence to automate the analysis of the data errors and biases Digital twins [6] are virtual representations of real-world products, systems, or processes, enabling simulation, integration, testing, monitoring, and maintenance. They play a pivotal role in optimizing complex systems across a wide range of domains, from industrial manufacturing and energy to environmental monitoring and healthcare.

  • The intellectual property is shared among all collaborators and the scientific results will be published in international journals and/or at international conferences. For the remaining context please see the abstract.

More detailed description : https://www.edtlab.fr/en/join-us/phd-pc5-phd2-estimatingbiases, * Subsidized meals

  • Partial reimbursement of public transport costs
  • Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in working hours) + possibility of exceptional leave (sick children, moving home, etc.)
  • Possibility of teleworking and flexible organization of working hours
  • Professional equipment available (videoconferencing, loan of computer equipment, etc.)
  • Social, cultural and sports events and activities
  • Access to vocational training
  • Social security coverage

About the company

The Inria Saclay-ÃŽle-de-France Research Centre was established in 2008. It has developed as part of the Saclay site in partnership with Paris-Saclay University and with the Institut Polytechnique de Paris .

The centre has 39 project teams , 27 of which operate jointly with Paris-Saclay University and the Institut Polytechnique de Paris; Its activities occupy over 600 people, scientists and research and innovation support staff, including 44 different nationalities.

Contexte et atouts du poste

The project will be hosted by Inria’s AVIZ team and supervised by : Maria Lobo-Gunther This advertised topic targets students focused on visualization and data analysis, with a background in artificial intelligence. It is not focused purely on AI. Also, please note that we only consider applications sent in English.

In the field of visualization, questions on the visualization of data uncertainty have been a core part of the research to date [e.g., 1]. Yet past work largely often silently considered data uncertainty to arise (primarily) from measurement uncertainty or as connected to some statistical analysis of captured data values. Only relatively recently have questions of implicit notions of error [4] or data hunches [3] entered the discussions of this general problem-these cover aspects of data imprecision that can arise from various sources: e.g., different forms of measurement or data recording, biases in how people access data, and even intentionally introduced errors. For example, spatial geographic data may not only come from systematically controlled sources but also from, for instance, from multiple sets of data that come from different institutions with different data collection policies, from contributions from the general public, or even by sourcing data from social media. Naturally, this introduces a variety of different levels of data quality and plausibility for spatial data. Sometimes experts are aware of these data caveats (that exist even for professionally collected datasets) and we can try to visualize them [3-5], yet in other cases the data collection happens in a way that this knowledge is not available-e.g., when data is retro-actively collected from sources that were not (primarily) created for recording this data in the first place-and we have to find ways of recovering such information retroactively and without access to the ground truth [2]. For instance, image collection sites such as Flickr partially show geo-located images, from which locations of certain points of interest can be extracted. In this specific example we ourselves recovered species distribution data from images posted on photo sharing platforms, and found that such geographic data is subject to many biases and errors. Yet this data analysis so far [2] has been a manual process and also only recovers anecdotal information about the existence of data errors and biases. In this PhD research project we will investigate if those results generalize to other kinds of spatial data (e.g., automatically detected features in remote sensing imagery, other volunteered geographic information, other forms of 2D and 3D spatial data). Then we will investigate ways to automate the error and bias identification process as well as to quantify the existence of such biases and errors in some form of plausibility measure, to be visualized in the digital twin [6]. The funding for this thesis comes from the EDT project (https://edtlab.fr/en/).

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.inria.fr
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:37 min

Audience questions on AI tools, demographics, and competence

Robindro Ullah Robindro Ullah · World Congress 2026 Europe

2:36 min

Applying supervised machine learning for practical rule extraction

Katja Träumner

2:46 min

Visualizing statistical patterns mathematically using the ggplot package

Mihailo Joksimovic · World Congress 2022

3:46 min

Core terminology and audiences for interpretable artificial intelligence

Karol Przystalski · LIVE

3:16 min

Bridging university concepts and industry needs via visual analytics

Johanna Schmidt · LIVE

1:57 min

Evolution of machine learning algorithms and computing hardware

Alexandra Waldherr · LIVE

Videos

See all

Related articles

See all