> Markdown version of [/jobs/ext/75969-phd-position-f-m-distributed-training-of-machine-learning-models-with-malicious-clients](https://www.wearedevelopers.com/jobs/ext/75969-phd-position-f-m-distributed-training-of-machine-learning-models-with-malicious-clients). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # PhD Position F/M Distributed Training of Machine Learning Models with Malicious Clients - **Company:** Inria - **Location:** Rocquencourt, France (Remote available) - **Salary:** €27,600.0 - **Contract:** Temporary contract - **Skills:** Computer Programming, Distributed Computing Environment, Information Theory, Python (Programming Language), Machine Learning, Tensorflow, Pytorch, Distributed Learning, Software Library - **Published:** May 21, 2026 - **Apply:** https://fr.indeed.com/viewjob?jk=4bf1fd03585b666e ## About the Role Do you have a Master's degree?, The candidate should have a solid mathematical background, in particular in probability, optimization, information theory, or statistical machine learning. A strong interest in privacy and security for machine learning is expected. Good programming skills are required, preferably in Python. Previous experience with PyTorch, TensorFlow, JAX, or federated learning libraries is a plus. Knowledge of adversarial machine learning, privacy attacks, robust learning, or distributed optimization would be appreciated but is not mandatory. The candidate should be able to work both theoretically and experimentally, and should be motivated by the design of rigorous models that can lead to practical insights for real distributed AI systems. Fluency in English is expected. ## Description Federated Learning (FL) enables a large number of devices, such as smartphones, sensors, or connected objects, to collaboratively train a shared machine learning model while keeping their data local [mcmahan17, li20]. This paradigm reduces the need to centralize sensitive data and has already been deployed in real-world applications, for example for mobile keyboard prediction [hard18]. However, decentralizing the training process does not automatically guarantee security or privacy. In federated and distributed learning systems, the participating clients are no longer passive data holders: they compute model updates and influence the training dynamics. As a consequence, a malicious client may try to corrupt the final model, slow down convergence, or extract private information from other participants. Several attacks have already shown that malicious participants can severely affect FL systems. Model-poisoning attacks can degrade performance by flipping labels, manipulating gradients, or crafting malicious updates that bypass robust aggregation rules [baruch19, fang20, xie20, xie25]. Backdoor attacks can implant hidden behaviors in the trained model, which are activated only by specific inputs chosen by the attacker [wang20, lyu23]. Beyond integrity attacks, malicious clients may also exploit the collaborative training process to infer sensitive information about other participants. For instance, they may infer whether a particular property is present in another client's dataset [melis19], or train local generators to reconstruct class-level information from a target participant [hitaj17]. The central question of this PhD is the following: how much private information can a malicious participant extract while keeping its poisoned update sufficiently small or stealthy to avoid detection? Answering this question is essential for designing distributed learning systems that are not only robust to performance degradation, but also resilient to privacy attacks carried out by active participants. Research objectives The goal of this PhD is to study privacy vulnerabilities in federated and distributed training systems in the presence of malicious clients, and to design principled defenses against them. A first objective will be to advance existing privacy attacks in distributed learning. The candidate will investigate how a malicious participant can manipulate its local training objective or model update in order to extract richer private information from honest clients. The focus will go beyond recovering simple class-level representations, with the goal of understanding what information can be inferred about individual samples, data properties, or local data distributions. A second objective will be to study stealthy model-poisoning attacks. In practical systems, malicious updates are often constrained by anomaly detection, robust aggregation, clipping, or validation mechanisms. The thesis will therefore consider bounded poisoning models, where the attacker is restricted to a neighborhood of legitimate updates. This will make it possible to analyze attacks that are powerful enough to leak private information, but sufficiently small to remain hard to detect. A third objective will be to design and evaluate defenses. The thesis will study how existing mechanisms such as robust aggregation, clipping, anomaly detection, regularization, privacy-preserving training, or client-side validation can mitigate privacy leakage induced by malicious participants. When existing defenses are insufficient, the candidate will propose new methods that explicitly account for the trade-off between privacy protection, robustness, and final model accuracy. The work will combine theoretical analysis with experimental validation. The candidate will implement attacks and defenses in controlled simulation environments for federated and distributed learning, using standard machine learning datasets and relevant threat models. The empirical study will evaluate both the utility of the trained model and the effectiveness of the attacks and defenses, with particular attention to stealthiness and practical detectability. ## Related Videos - [A hundred ways to wreck your AI - the (in)security of machine learning systems](https://www.wearedevelopers.com/videos/715-a-hundred-ways-to-wreck-your-ai-the-in-security-of-machine-learning-systems) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [Making neural networks portable with ONNX](https://www.wearedevelopers.com/videos/301-making-neural-networks-portable-with-onnx) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 191: Malware interviews, EU ❤️ Open Source and Skilled Agents](https://www.wearedevelopers.com/magazine/645-dev-digest-191-malware-interviews-eu-open-source-and-skilled-agents) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 208: 4 Hours Code a Day, WebMCP Insights, PyTorch for Beginners](https://www.wearedevelopers.com/magazine/698-dev-digest-208-4-hours-code-a-day-webmcp-insights-pytorch-for-beginners) - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing)