[Data - FR] Senior Machine Learning Engineer - Orchestration

DOCTOLIB SAS
Paris, France
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Artificial Intelligence Computer Vision Distributed Systems Mobile Application Software Python (Programming Language) Machine Learning TypeScript Large Language Models Multi-Agent Systems Swift (Programming Language) Kotlin
+3 more
Information Technology React Native GPT

Job description

As a Senior/Staff Machine Learning Engineer, you’ll play a key role in designing, implementing, and scaling the evaluation framework that ensures our AI Health Companion behaves safely, reliably, and helpfully for millions of patients and practitioners.

You’ll join a cross-functional team of Machine Learning Engineers, Product Engineers, and Medical Experts to build robust evaluation pipelines for agentic AI systems - models capable of reasoning, planning, and interacting with complex healthcare data.

Your responsibilities include, but are not limited to:

  • Define and own the evaluation strategy for our AI agentic system - metrics, protocols, datasets, and tooling
  • Implement and maintain automated evaluation pipelines to monitor model quality, safety, and alignment across iterations
  • Run systematic experiments to assess reasoning, factuality, robustness, and user experience
  • Collaborate closely with model developers and research scientists to provide insights and drive iterative improvement
  • Contribute to research and internal knowledge sharing on LLM evaluation methodologies and best practices

About our tech environment

  • Our solutions are built on a single fully cloud-native platform that supports web and mobile app interfaces, multiple languages, and is adapted to the country and healthcare specialty requirements. To address these challenges, we are modularizing our platform run in a distributed architecture through reusable components
  • Our stack is composed of Rails, TypeScript, Java, Python, Kotlin, Swift, and React Native
  • We leverage AI ethically across our products to empower patients and health professionals. Discover our AI vision here!

Requirements

  • MSc or PhD in Computer Science, Machine Learning, Data Science, or related field
  • 7+ years of hands-on experience working with large language models (e.g., GPT, Claude, Llama, or BERT-like architectures)
  • Proven experience in evaluating agentic or reasoning systems (e.g., autonomous agents, tool-using LLMs, dialogue systems, or task-oriented assistants)
  • Strong track record in experiment design, metric definition, and evaluation automation
  • Ability to bridge research and production, influencing modeling and product decisions
  • Excellent communication skills and a collaborative mindset

Now it would be fantastic if:

  • You have experience in the clinical or medical domain and sensitivity to ethical or regulatory challenges in healthcare AI

Benefits & conditions

  • Free mental health and coaching services through our partner Moka.care
  • For caregivers and workers with disabilities, a package including an adaptation of the remote policy, extra days off for medical reasons, and psychological support
  • Work from EU countries and the UK for up to 10 days per year, thanks to our flexibility days policy
  • Work Council subsidy to refund part of sport club membership or creative class
  • Up to 14 days of RTT
  • Lunch voucher with Swile card

About the company

At Doctolib, we’re on a mission to transform how healthcare is delivered by harnessing the power of AI.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on fr.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · WWC 2024

3:09 min

Understanding Kotlin Multiplatform and its compiler targets

Petar Marijanović · LIVE

1:00 min

Misconceptions about TypeScript safety capabilities

Simone Sanfratello · JS Congress

3:13 min

Addressing language barriers and the data science talent deficit

Markus Hacker Markus Hacker +3 · WWC 2024

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky · WWC 2025

Videos

See all

Related articles

See all