Engineering Manager - Site Reliability & Observability
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Experteer Overview As Engineering Manager for the SRE team, you will lead a group of Site Reliability Engineers to ensure Doctolib’s platform is reliable, scalable, and resilient at European scale. You will set reliability and observability direction and collaborate with product and engineering teams to enable safe, fast delivery. You’ll drive large-scale reliability initiatives, SLOs, and incident prevention across hundreds of applications. Your leadership will foster a culture of operational excellence and continuous improvement while aligning with the company mission to improve healthcare access. Pay / Benefits * Lead and grow a team of Site Reliability Engineers and shape their technical growth * Define and evolve reliability and observability strategy (infrastructure automation, logging, metrics, tracing, alerting) * Drive roadmap for large-scale reliability initiatives, including SLOs and error budgets * Own transversal services (secrets management, infrastructure as code tooling) * Improve developer experience by reducing technical debt and enabling scalable infrastructure * Own on-call experience and contribute to incident response and postmortems * Collaborate with Product, Engineering, and architecture teams to align reliability with product goals * Represent the team in engineering leadership forums and architectural reviews * Foster partnerships with software engineering to embed reliability early in the lifecycle Tasks * 5+ years software engineering or SRE experience in cloud-native environments * 3+ years engineering management experience * Strong observability tooling knowledge (OpenTelemetry, Prometheus, Datadog, Elasticsearch) * Experience with infrastructure as code (Terraform) and secrets management * Ability to mentor engineers, review designs, and guide architecture * Fluent in English Key requirements * health insurance * 25 days vacation + RTT * mental health and coaching services * flexibility days (work from abroad) * lunch vouchers * transport subsidy reimbursement 50% of public transport subscription
Requirements
and * Improve developer experience by reducing technical debt and enabling scalable infrastructure * Own on-call experience and contribute to incident response and postmortems * Collaborate with Product, Engineering, and architecture teams to align reliability with product goals * Represent the team in engineering leadership forums and architectural reviews * Foster partnerships with software engineering to embed reliability early in the lifecycle Tasks * 5+ years software engineering or SRE experience in cloud-native environments * 3+ years engineering management experience * Strong observability tooling knowledge (OpenTelemetry, Prometheus, Datadog, Elasticsearch) * Experience with infrastructure as code (Terraform) and secrets management * Ability to mentor engineers, review designs, and guide architecture * Fluent in English Key requirements * health insurance * 25 days vacation + RTT * mental health and coaching services * flexibility days (work from abroad) * lunch vouchers * transport subsidy reimbursement 50% of public transport subscription
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on eu.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Best Companies to Work For in Paris: Top 25 Companies in 2023Â
DevOps Engineer Salary [2023]
Dev Digest 120 - Apple and peers
Find a Developer Job: 12 Best Job Sites For Developers