> Markdown version of [/jobs/ext/1361432-site-reliability-engineer-for-observability](https://www.wearedevelopers.com/jobs/ext/1361432-site-reliability-engineer-for-observability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer for Observability - **Company:** Nn. - **Location:** Den Haag, Netherlands - **Contract:** Temporary contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Bash Shell, Cloud Engineering, Continuous Integration, Data Retention, DevOps, Python (Programming Language), Reliability Engineering, Prometheus, Datadog, Grafana, Containerization, Kubernetes, Terraform, AWS EKS, Databricks, Golang - **Published:** July 21, 2026 - **Apply:** https://www.nn-careers.com/vacature/3468/site-reliability-engineer-for-observability-freelance?apply=1&utm_source=careerguide ## About the Role * SRE/DevOps experience supporting production systems and incident management. * Strong knowledge of observability tooling (e.g., Prometheus, Grafana, OpenTelemetry, ELK/EFK, Datadog, New Relic). * Proficiency in Linux, networking fundamentals, and troubleshooting distributed systems. * Experience with cloud platforms and infrastructure as code (e.g., AWS/GCP/Azure; Terraform). * Scripting and automation skills (e.g., Python, Go, Bash) and CI/CD familiarity. Key outcomes * Improved uptime and performance through measurable SLOs and actionable telemetry. * Faster MTTR via high-signal alerts, clear dashboards, and effective runbooks. The Engineering Experience & Platforms domain aims to make life easier for engineers across NN. To strengthen reliability and improve observability across key platforms, NN is looking for an interim Site Reliability Engineer to design, implement and embed a scalable observability foundation., You are an experienced Site Reliability Engineer with strong production experience in Kubernetes and containerized workloads. You have hands-on cloud engineering experience in Azure and/or AWS, including infrastructure as code and GitOps-based deployments. You bring deep observability expertise across metrics, logs and traces, using tools such as Grafana, Prometheus, Loki, Tempo or similar. You are comfortable defining and managing SLOs, SLIs and error budgets, and you have a structured approach to incident management, root cause analysis and reliability improvements. Automation skills in Python, Bash or Go are expected, as well as solid knowledge of CI/CD and safe deployment practices. You communicate clearly and work effectively with development teams to embed reliability into the software delivery lifecycle. ## Description As a Site Reliability Engineer for Observability (Freelance) at NN, you design and embed a scalable observability foundation: build an OpenTelemetry collector layer (Azure Databricks, AWS/Azure Kubernetes, AWS/Azure), define SLIs/SLOs/SLAs, and enable teams to adopt it. Direct solliciteren Neem contact op, Role purpose: Design, implement, and operate an observability platform that improves service reliability, accelerates incident response, and enables data-driven performance and availability improvements across production systems., * Build and maintain end-to-end observability for distributed systems: metrics, logs, traces, and synthetic monitoring. * Define and operationalize SLIs/SLOs, error budgets, alerting strategies, and on-call readiness. * Develop dashboards and alerts that reduce noise and improve detection, triage, and root-cause analysis. * Automate incident response workflows, runbooks, postmortems, and continuous improvement actions. * Partner with engineering teams to instrument services, standardize telemetry, and improve reliability patterns. * Optimize monitoring cost, data retention, sampling, and performance of observability pipelines., The assignment focuses on setting up an OpenTelemetry collector layer for sources including Azure Databricks, AWS Kubernetes, Azure Kubernetes, AWS and Azure. You will define and implement SLIs, SLOs and SLAs for the Portable Stack domain and AI domain, and enable teams to use the new observability capabilities effectively., You will work closely with the newly formed observability team, the domain architect and the Principal Engineer of the Kubernetes domain. You will also align with other teams in the domain to deliver a practical, scalable solution that raises SRE maturity across NN. ## Related Videos - [Hosting a modern justice system](https://www.wearedevelopers.com/videos/332-hosting-a-modern-justice-system) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Effortlessly Scale Prometheus With The Telemetry Data Platform – And Keep your Grafana Dashboards, Too!](https://www.wearedevelopers.com/magazine/3-effortlessly-scale-prometheus-with-the-telemetry-data-platform-and-keep-your-grafana-dashboards-too)