> Markdown version of [/videos/1633-practical-ai-with-machine-learning-for-observability-in-netdata?t=266](https://www.wearedevelopers.com/videos/1633-practical-ai-with-machine-learning-for-observability-in-netdata?t=266). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Practical AI with Machine Learning for Observability in Netdata Are centralized observability models making real-time machine learning prohibitively expensive? Discover how Netdata distributes incredibly lightweight AI directly to the edge to mathematically eliminate false positives. - **Speakers:** [Costa](https://www.wearedevelopers.com/@costa) - **Event:** World Congress 2025 - **Published:** August 20, 2025 - **Duration:** 25:32 - **URL:** https://www.wearedevelopers.com/videos/1633-practical-ai-with-machine-learning-for-observability-in-netdata ## Summary Traditional observability models centralize data, making real-time, high-cardinality machine learning prohibitively expensive. Netdata flips this paradigm by distributing the code directly to edge agents, enabling per-second metric resolution across thousands of services. This distributed architecture resolves a critical flaw in modern AI observability assistants: the lack of operational context. By processing data locally, the system maintains a comprehensive understanding of infrastructure behavior without the latency and prohibitive cost of centralized aggregation. At the core of this approach is an extraordinarily lightweight, edge-based machine learning engine. Netdata trains models on user data every three hours, utilizing behavioral vectors—akin to measuring acceleration rather than speed—to detect anomalies as they occur. To guarantee reliability and eliminate noise, the system maintains 18 concurrent model versions that must reach full consensus to flag an anomaly. The efficiency is staggering, requiring only a fraction of a CPU core to train hundreds of thousands of models, 2.4 kilobytes of memory per metric, and zero additional storage footprint thanks to custom floating-point embeddings. These edge-computed anomalies feed directly into a specialized scoring engine and the Anomaly Advisor tool, fundamentally changing root cause analysis. Instead of relying on speculative troubleshooting based on assumed dependency maps, engineers can highlight a specific time window to receive a dynamically ordered list of correlated anomalous metrics. By detecting host-level anomalies when just 1% of metrics spike concurrently, the platform mathematically eliminates false positives and surfaces agnostic, infrastructure-wide insights that reveal the true source of outages in real time. **Keywords:** netdata observability, distributed observability agents, edge machine learning, real-time anomaly detection, automated root cause analysis, observability scoring engine, infrastructure troubleshooting, metric cardinality, contextual ai assistants, host-level anomalies, time series monitoring, machine learning consensus, dependency agnostic troubleshooting, cncf observability tools, custom floating point embeddings ## Chapters 1. **Introduction to NetData's distributed observability architecture** (00:05) — Distributing code to smart agents enables high-resolution metrics and scalable observability. 1. **The evolution of machine learning in observability** (01:24) — Overcoming past skepticism allows teams to leverage machine learning for complex infrastructure monitoring. 1. **Why AI assistants struggle without infrastructure context** (02:33) — Standard AI troubleshooting processes fail because they rely on isolated alerts without system context. 1. **Providing time-window context to AI assistants** (04:26) — Providing specific time windows enables AI tools to detect significant infrastructure changes. 1. **Training continuous machine learning models per metric** (05:08) — Generating continuous behavioral models at the edge accurately detects outliers in real-time metrics. 1. **Ensuring reliable anomaly detection with model consensus** (08:24) — Requiring agreement across multiple trained models mathematically eliminates anomaly false positives. 1. **Detecting host-level anomalies through metric clustering** (09:40) — Tracking concurrent behavioral shifts across multiple system metrics identifies major infrastructure outages. 1. **Internal agent architecture for real-time anomaly detection** (12:32) — Internal agent software paths stream metric discovery and machine learning training to database alerts. 1. **Resource efficiency of running machine learning at the edge** (13:55) — Custom data handling keeps CPU, memory, and storage footprints minimal while training thousands of models. 1. **Understanding the limits of anomaly detection workloads** (16:51) — Machine learning struggles to detect anomalies in short-lived cron jobs and immediately crashed services. 1. **Integrating anomaly insights directly into user interface charts** (18:29) — Anomaly ribbons and the needle framework help users instantly comprehend complex dashboard data. 1. **Ranking infrastructure issues with a dedicated scoring engine** (21:20) — Sorting metrics by anomaly rate provides AI assistants and users with prioritized issue lists. 1. **Flipping troubleshooting with automated root cause analysis** (22:34) — The anomaly advisor reveals correlated issues instead of relying on manual assumption testing. 1. **Summary of NetData's open-source observability platform** (24:53) — NetData provides unsupervised anomaly detection and automated root cause analysis for the open-source community. ## Related Moments - [Using AI for incident summaries and root cause analysis](https://www.wearedevelopers.com/videos/100086-unlocking-the-ai-black-box-building-trust-in-the-era-of-agentic-production) (from "Unlocking the AI Black Box: Building Trust in the Era of Agentic Production") - [Adapting observability strategies for long-running enterprise AI agents](https://www.wearedevelopers.com/videos/100166-shipping-with-confidence-observability-and-quality-at-scale) (from "Shipping with Confidence: Observability and Quality at Scale") - [Implementing monitoring and observability for AI software deployments](https://www.wearedevelopers.com/videos/1383-the-state-of-genai-machine-learning-in-2025) (from "The State of GenAI & Machine Learning in 2025") - [Monitoring enterprise AI workloads for continuous observability](https://www.wearedevelopers.com/videos/1535-from-traction-to-production-maturing-your-genaiops-step-by-step) (from "From Traction to Production: Maturing your GenAIOps step by step") - [Implementing semi-automated anomaly detection with human oversight](https://www.wearedevelopers.com/videos/1308-data-science-ml-ai-in-the-oil-and-gas-industry-at-ndt-global-dr-katja-traumner) (from "Data Science, ML & AI in the Oil and Gas Industry at NDT Global - Dr. Katja Träumner") - [Leveraging generative AI for application observability and security](https://www.wearedevelopers.com/videos/598-why-shifting-left-is-so-important-for-software-developers) (from "Why shifting left is so important for software developers") ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) ## Related Jobs - [Senior AI/ML Engineer](https://www.wearedevelopers.com/jobs/48352-senior-ai-ml-engineer) at **PagerDuty** - [Founding Data Scientist](https://www.wearedevelopers.com/jobs/ext/2731830-founding-data-scientist) at **Almedia** - [ML Engineer](https://www.wearedevelopers.com/jobs/48448-ml-engineer) at **Docker, Inc.** - [Staff ML Engineer](https://www.wearedevelopers.com/jobs/48463-staff-ml-engineer) at **Docker, Inc.** - [Data Scientist](https://www.wearedevelopers.com/jobs/ext/2725456-data-scientist) at **Almedia** - [Principal Software Engineer, AI Compute Infrastructure](https://www.wearedevelopers.com/jobs/ext/2847709-principal-software-engineer-ai-compute-infrastructure) at **ARM**