> Markdown version of [/videos/753-monitoring-as-code-managing-your-dashboards-at-scale?t=127](https://www.wearedevelopers.com/videos/753-monitoring-as-code-managing-your-dashboards-at-scale?t=127). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Monitoring as Code - Managing your dashboards at scale Manual dashboard configuration fails at scale. Learn how to programmatically build, test, and deploy Grafana dashboards across massive microservices architectures using Jsonnet and CI/CD pipelines. - **Speakers:** Gabriel Labachelerie - **Event:** World Congress 2023 - **Published:** October 6, 2023 - **Duration:** 32:47 - **URL:** https://www.wearedevelopers.com/videos/753-monitoring-as-code-managing-your-dashboards-at-scale ## Summary Managing observability across a massive microservices architecture—like a system handling billions of daily travel transactions—requires more than manual UI configurations. To maintain system stability and provide immediate frontline support, engineering teams need an industrialized approach to Grafana dashboard creation. Monitoring as code solves the scalability and reliability bottlenecks of manual dashboarding by allowing engineers to programmatically define, test, and deploy observability assets. By leveraging Jsonnet to generate JSON definitions for Grafana, teams can build reusable panel templates and dynamically adapt PromQL queries based on context. A custom Golang-based CLI tool orchestrates the entire workflow, fetching dependencies, auto-formatting code, and deploying changes. This framework enables powerful automation, such as programmatically pre-computing thousands of dedicated dashboards—one for each customer or service—eliminating the friction of navigating complex dropdown filters during critical incidents. Crucially, treating dashboards as actual code means enforcing rigorous quality standards through CI/CD pipelines. Engineers can run unit tests to validate dashboard layouts and integrate promtool to simulate time-series data, ensuring PromQL queries execute correctly before deployment. Integrated with Jenkins for pull request diffs and extending to Prometheus alert configurations, this code-driven methodology ensures a safe, fast, and highly reliable developer experience for observability at scale. **Keywords:** monitoring as code, grafana dashboard automation, jsonnet templating, promql query validation, prometheus alert deployment, observability at scale, dashboard as code, time-series data simulation, thanos metric federation, CI/CD observability integration, golang CLI tooling, microservices monitoring strategy, incident response dashboards, promtool data testing, grafana JSON rendering ## Chapters 1. **Managing monitoring and observability at a large scale** (00:03) — Handling massive transaction volumes requires immediate and reliable observability tools. 1. **Architecting the availability stack with Prometheus and Grafana** (02:07) — Combining Prometheus instances via Thanos federation simplifies dashboard querying in Grafana. 1. **Transitioning from manual to automated dashboard creation** (03:15) — Industrializing dashboard creation ensures testability, quality, and faster deployment cycles. 1. **Leveraging Jsonnet for generating dashboard configurations** (04:38) — Using Jsonnet and Grafana's standard libraries allows engineers to define JSON-based dashboards as code. 1. **Generating an initial baseline dashboard with code** (07:23) — Compiling source files through a custom CLI tool outputs a foundational dashboard structure. 1. **Constructing panel templates with dynamic PromQL parameters** (09:50) — Declaring variables within query templates enables context-aware metric filtering across different dashboards. 1. **Instantiating configured panels into dashboard layouts** (13:18) — Calling a predefined template within the dashboard code automatically maps filters and generates queries. 1. **Extending metrics and visualizing error code variations** (15:16) — Adding new instances of a template allows for side-by-side comparison of success and failure metrics. 1. **Injecting dynamic template variables for dashboard interactivity** (18:01) — Defining query variables in code creates dropdown selectors that automatically update underlying PromQL expressions. 1. **Scaling production dashboards through loop iterations** (19:39) — Iterating over configuration lists programmatically generates thousands of dedicated rows or isolated dashboards. 1. **Implementing unit tests for dashboard structural integrity** (21:38) — Treating dashboards as code permits automated validation of panel counts and layout structures prior to deployment. 1. **Validating PromQL expressions using simulated data sets** (22:39) — Integrating promtool ensures generated queries behave correctly against mocked time series data. 1. **Automating deployments with Jenkins continuous integration pipelines** (24:04) — Reviewing JSON diffs on pull requests enables teams to understand dashboard modifications before merging. 1. **Leveraging auxiliary libraries for formatting and linking** (24:50) — Utilizing ecosystem tooling handles auto-formatting, linting, and maintaining context-aware links between panels. 1. **Enhancing developer experience with localized command line tools** (27:29) — Providing a single Go-based CLI ensures consistent testing and deployment workflows across local workstations and CI servers. 1. **Addressing practical deployment questions and incident resolution** (29:54) — Pre-computing customer-specific dashboards reduces time to resolution during critical production incidents. ## Related Moments - [Visualizing Prometheus open metrics using custom Grafana dashboards](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) (from "5 steps for running a Kubernetes environment at scale") - [Automating frontend performance metrics with Google Lighthouse](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) (from "Automate everything via NodeJS and Puppeteer") - [Limitations of infrastructure and dashboards as code approaches](https://www.wearedevelopers.com/videos/1618-planet-scale-dashboards) (from "Planet-Scale Dashboards") - [Shifting from reactive observability to runtime intelligence models](https://www.wearedevelopers.com/videos/100277-what-production-knows-closing-the-loop-between-ai-agents-and-the-systems-they-build) (from "What Production Knows: Closing the Loop Between AI Agents and the Systems They Build") - [Understanding observability through a practical dashboard analogy](https://www.wearedevelopers.com/videos/598-why-shifting-left-is-so-important-for-software-developers) (from "Why shifting left is so important for software developers") - [Configuring Prometheus remote write and connecting Grafana dashboards](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) (from "All your telemetry data from any source in one place") ## Related Articles - [Effortlessly Scale Prometheus With The Telemetry Data Platform – And Keep your Grafana Dashboards, Too!](https://www.wearedevelopers.com/magazine/3-effortlessly-scale-prometheus-with-the-telemetry-data-platform-and-keep-your-grafana-dashboards-too) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) ## Related Jobs - [GPU Cluster Engineer, Systems & Platform](https://www.wearedevelopers.com/jobs/48411-gpu-cluster-engineer-systems-platform) at **Sciforium** - [Software Engineer, Infrastructure Platform](https://www.wearedevelopers.com/jobs/ext/1940513-software-engineer-infrastructure-platform) at **Docker, Inc.** - [DevOps Engineer (Observability)](https://www.wearedevelopers.com/jobs/ext/2568619-devops-engineer-observability) at **Twilio** - [Software Engineer, Infrastructure Platform](https://www.wearedevelopers.com/jobs/ext/2004010-software-engineer-infrastructure-platform) at **Docker, Inc.** - [Cloud Engineer (German)](https://www.wearedevelopers.com/jobs/ext/2908194-cloud-engineer-german) at **Stackit** - [Sr. Solution Architect](https://www.wearedevelopers.com/jobs/48433-sr-solution-architect) at **Dynatrace**