Site Reliability Engineer for Observability

Nn.
The Hague, Netherlands
yesterday

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English

Job location

The Hague, Netherlands

Tech stack

Artificial Intelligence
Amazon Web Services (AWS)
Azure
Bash
Cloud Engineering
Continuous Integration
Data Retention
DevOps
Python
Reliability Engineering
Prometheus
Datadog
Grafana
Containerization
Kubernetes
Terraform
Amazon Web Services (AWS)
Databricks
Go

Job description

As a Site Reliability Engineer for Observability (Freelance) at NN, you design and embed a scalable observability foundation: build an OpenTelemetry collector layer (Azure Databricks, AWS/Azure Kubernetes, AWS/Azure), define SLIs/SLOs/SLAs, and enable teams to adopt it. Direct solliciteren Neem contact op, Role purpose: Design, implement, and operate an observability platform that improves service reliability, accelerates incident response, and enables data-driven performance and availability improvements across production systems., * Build and maintain end-to-end observability for distributed systems: metrics, logs, traces, and synthetic monitoring.

  • Define and operationalize SLIs/SLOs, error budgets, alerting strategies, and on-call readiness.
  • Develop dashboards and alerts that reduce noise and improve detection, triage, and root-cause analysis.
  • Automate incident response workflows, runbooks, postmortems, and continuous improvement actions.
  • Partner with engineering teams to instrument services, standardize telemetry, and improve reliability patterns.
  • Optimize monitoring cost, data retention, sampling, and performance of observability pipelines., The assignment focuses on setting up an OpenTelemetry collector layer for sources including Azure Databricks, AWS Kubernetes, Azure Kubernetes, AWS and Azure. You will define and implement SLIs, SLOs and SLAs for the Portable Stack domain and AI domain, and enable teams to use the new observability capabilities effectively., You will work closely with the newly formed observability team, the domain architect and the Principal Engineer of the Kubernetes domain. You will also align with other teams in the domain to deliver a practical, scalable solution that raises SRE maturity across NN.

Requirements

  • SRE/DevOps experience supporting production systems and incident management.
  • Strong knowledge of observability tooling (e.g., Prometheus, Grafana, OpenTelemetry, ELK/EFK, Datadog, New Relic).
  • Proficiency in Linux, networking fundamentals, and troubleshooting distributed systems.
  • Experience with cloud platforms and infrastructure as code (e.g., AWS/GCP/Azure; Terraform).
  • Scripting and automation skills (e.g., Python, Go, Bash) and CI/CD familiarity.

Key outcomes

  • Improved uptime and performance through measurable SLOs and actionable telemetry.
  • Faster MTTR via high-signal alerts, clear dashboards, and effective runbooks.

The Engineering Experience & Platforms domain aims to make life easier for engineers across NN. To strengthen reliability and improve observability across key platforms, NN is looking for an interim Site Reliability Engineer to design, implement and embed a scalable observability foundation., You are an experienced Site Reliability Engineer with strong production experience in Kubernetes and containerized workloads. You have hands-on cloud engineering experience in Azure and/or AWS, including infrastructure as code and GitOps-based deployments. You bring deep observability expertise across metrics, logs and traces, using tools such as Grafana, Prometheus, Loki, Tempo or similar. You are comfortable defining and managing SLOs, SLIs and error budgets, and you have a structured approach to incident management, root cause analysis and reliability improvements. Automation skills in Python, Bash or Go are expected, as well as solid knowledge of CI/CD and safe deployment practices. You communicate clearly and work effectively with development teams to embed reliability into the software delivery lifecycle.

Benefits & conditions

Our people are the driving force behind our organisation. We value the knowledge and expertise you bring. We believe that your temporary commitment can take our organisation to a higher level. We offer you:

  • Competitive hourly rate depending on your knowledge and experience
  • Project starts 1st of August 2026 and has a duration of 6 to 9 months
  • Hybrid way of working, partly from home and partly from the office
  • International working environment with loads of knowledge sharing

Apply for this position