Engineering Manager - Site Reliability & Observability

DOCTOLIB SAS
Paris, France
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Languages
English

Tech stack

Cloud Computing Elasticsearch Key Management Reliability Engineering Prometheus Software Engineering Datadog Data Logging Technical Debt Infrastructure Automation Frameworks Terraform

Job description

Experteer Overview As Engineering Manager for the SRE team, you will lead a group of Site Reliability Engineers to ensure Doctolib’s platform is reliable, scalable, and resilient at European scale. You will set reliability and observability direction and collaborate with product and engineering teams to enable safe, fast delivery. You’ll drive large-scale reliability initiatives, SLOs, and incident prevention across hundreds of applications. Your leadership will foster a culture of operational excellence and continuous improvement while aligning with the company mission to improve healthcare access. Pay / Benefits * Lead and grow a team of Site Reliability Engineers and shape their technical growth * Define and evolve reliability and observability strategy (infrastructure automation, logging, metrics, tracing, alerting) * Drive roadmap for large-scale reliability initiatives, including SLOs and error budgets * Own transversal services (secrets management, infrastructure as code tooling) * Improve developer experience by reducing technical debt and enabling scalable infrastructure * Own on-call experience and contribute to incident response and postmortems * Collaborate with Product, Engineering, and architecture teams to align reliability with product goals * Represent the team in engineering leadership forums and architectural reviews * Foster partnerships with software engineering to embed reliability early in the lifecycle Tasks * 5+ years software engineering or SRE experience in cloud-native environments * 3+ years engineering management experience * Strong observability tooling knowledge (OpenTelemetry, Prometheus, Datadog, Elasticsearch) * Experience with infrastructure as code (Terraform) and secrets management * Ability to mentor engineers, review designs, and guide architecture * Fluent in English Key requirements * health insurance * 25 days vacation + RTT * mental health and coaching services * flexibility days (work from abroad) * lunch vouchers * transport subsidy reimbursement 50% of public transport subscription

Requirements

and * Improve developer experience by reducing technical debt and enabling scalable infrastructure * Own on-call experience and contribute to incident response and postmortems * Collaborate with Product, Engineering, and architecture teams to align reliability with product goals * Represent the team in engineering leadership forums and architectural reviews * Foster partnerships with software engineering to embed reliability early in the lifecycle Tasks * 5+ years software engineering or SRE experience in cloud-native environments * 3+ years engineering management experience * Strong observability tooling knowledge (OpenTelemetry, Prometheus, Datadog, Elasticsearch) * Experience with infrastructure as code (Terraform) and secrets management * Ability to mentor engineers, review designs, and guide architecture * Fluent in English Key requirements * health insurance * 25 days vacation + RTT * mental health and coaching services * flexibility days (work from abroad) * lunch vouchers * transport subsidy reimbursement 50% of public transport subscription

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on eu.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:34 min

Managing token budgets and enterprise usage of coding agents

Chris Heilmann +2 · LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

1:05 min

Measuring system availability utilizing Prometheus and straightforward PromQL

Alexander Schwartz Alexander Schwartz · WWC 2025

3:52 min

Comparing French work-life balance principles with international engineering practices

Chris Heilmann +2 · LIVE

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel · WWC 2025

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

Videos

See all

Related articles

See all