Domain Manager - Saas Sre H/F

Propertyvalue
Paris, France
7 days ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Software as a Service Cloud Computing Cloud Engineering Continuous Integration Monitoring of Systems Reliability Engineering Reliability of Systems Devsecops

Job description

A leading company is seeking a Domain Manager to ensure service reliability, operational excellence, and performance compliance across environments by embedding Site Reliability Engineering (SRE) practices within the Agile Release Train and the product delivery lifecycle.

Missions

  • Guarantee the stability, performance, and availability of services in both production and non-production environments, fostering a reliability-driven culture across delivery teams.
  • Act as the gatekeeper of product evolution towards production, ensuring quality always matches customer expectations.
  • Collaborate with Product, Tech, and Platform teams to maintain the right balance between innovation, velocity, and operational robustness.
  • Define, monitor, and report Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets across environments to ensure measurable reliability per application domain.
  • Ensure robust observability, monitoring, and alerting frameworks are implemented and continuously improved.
  • Oversee operational readiness of each release, ensuring production stability through cross-functional coordination with Product and Tech teams.
  • Veto product delivery when quality measured is not in line with customer expectations.
  • Manage incident response, root cause analysis, and post-mortem reviews to ensure accountability and continuous improvement per application domain.
  • Collaborate with Core Platform and Observability & FinOps teams to strengthen system resilience, optimize cost efficiency, and maintain platform performance.
  • Report reliability status, risks, and improvement actions to Agile Release Managers and domain leadership to ensure alignment across Agile Release Trains (ARTs).
  • Actively participate in the Agile Release Train as the voice of reliability and operations, supporting delivery cadence and quality., Profile wantedStrong expertise in Site Reliability Engineering (SRE) within SaaS or cloud-native environmentsDeep understanding of system observability, automation, and monitoring frameworksExperience defining and managing SLOs, SLIs, and error budgets in collaboration with engineering teamsProficiency in DevSecOps, CI/CD pipelines, and continuous monitoring practicesSolid experience in incident management, post-mortem analysis, and operational readinessProven ability to coordinate reliability initiatives across Product, Tech, and Platform domainsStrong focus on performance metrics, root cause prevention, and operational governanceData-driven mindset and analytical rigor in reliability tracking

Mais aussi…

Daily rate:Salary according to profile

False

EducationalOccupationalCredential postgraduate degree

Requirements

Strong expertise in Site Reliability Engineering (SRE) within SaaS or cloud-native environments.

  • Deep understanding of system observability, automation, and monitoring frameworks.
  • Experience defining and managing SLOs, SLIs, and error budgets in collaboration with engineering teams.
  • Proficiency in DevSecOps, CI/CD pipelines, and continuous monitoring practices.

Functional

  • Solid experience in incident management, post-mortem analysis, and operational readiness.
  • Proven ability to coordinate reliability initiatives across Product, Tech, and Platform domains.
  • Strong focus on performance metrics, root cause prevention, and operational governance.

Soft Skills

  • Data-driven mindset and analytical rigor in reliability tracking., 1. Strong expertise in Site Reliability Engineering (SRE) within SaaS or cloud-native environments 2. Deep understanding of system observability, automation, and monitoring frameworks 3. Experience defining and managing SLOs, SLIs, and error budgets in collaboration with engineering teams 4. Proficiency in DevSecOps, CI/CD pipelines, and continuous monitoring practices 5. Solid experience in incident management, post-mortem analysis, and operational readiness 6. Proven ability to coordinate reliability initiatives across Product, Tech, and Platform domains 7. Strong focus on performance metrics, root cause prevention, and operational governance 8. Data-driven mindset and analytical rigor in reliability tracking

Benefits & conditions

Remote work conditions: no remote work for the first 3 months of the assignment.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.hellowork.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:15 min

Key lessons learned from implementing automated mobile DevSecOps

Moataz Nabil Moataz Nabil · LIVE

2:27 min

Introduction to WebAssembly in a cloud computing context

Edo Edo · WWC 2024

1:23 min

Closing thoughts and educational resources for edge network engineering

Austin Gil · LIVE

1:06 min

Developer experience and project variety at scale

Alexandra Petri · WWC 2023

2:09 min

Shifting security left using the DevSecOps approach

Aarno Aukia · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all