World Congress 2025 • Aug 20, 2025 • Session details

Practical AI with Machine Learning for Observability in Netdata

Costa

Are centralized observability models making real-time machine learning prohibitively expensive? Discover how Netdata distributes incredibly lightweight AI directly to the edge to mathematically eliminate false positives.

Pause
Mute Enter Fullscreen
#1 about 2 min

Introduction to NetData's distributed observability architecture

Distributing code to smart agents enables high-resolution metrics and scalable observability.

#2 about 2 min

The evolution of machine learning in observability

Overcoming past skepticism allows teams to leverage machine learning for complex infrastructure monitoring.

#3 about 2 min

Why AI assistants struggle without infrastructure context

Standard AI troubleshooting processes fail because they rely on isolated alerts without system context.

#4 about 1 min

Providing time-window context to AI assistants

Providing specific time windows enables AI tools to detect significant infrastructure changes.

#5 about 4 min

Training continuous machine learning models per metric

Generating continuous behavioral models at the edge accurately detects outliers in real-time metrics.

#6 about 2 min

Ensuring reliable anomaly detection with model consensus

Requiring agreement across multiple trained models mathematically eliminates anomaly false positives.

#7 about 3 min

Detecting host-level anomalies through metric clustering

Tracking concurrent behavioral shifts across multiple system metrics identifies major infrastructure outages.

#8 about 2 min

Internal agent architecture for real-time anomaly detection

Internal agent software paths stream metric discovery and machine learning training to database alerts.

#9 about 3 min

Resource efficiency of running machine learning at the edge

Custom data handling keeps CPU, memory, and storage footprints minimal while training thousands of models.

#10 about 2 min

Understanding the limits of anomaly detection workloads

Machine learning struggles to detect anomalies in short-lived cron jobs and immediately crashed services.

#11 about 3 min

Integrating anomaly insights directly into user interface charts

Anomaly ribbons and the needle framework help users instantly comprehend complex dashboard data.

#12 about 2 min

Ranking infrastructure issues with a dedicated scoring engine

Sorting metrics by anomaly rate provides AI assistants and users with prioritized issue lists.

#13 about 3 min

Flipping troubleshooting with automated root cause analysis

The anomaly advisor reveals correlated issues instead of relying on manual assumption testing.

#14 about 1 min

Summary of NetData's open-source observability platform

NetData provides unsupervised anomaly detection and automated root cause analysis for the open-source community.

Matching moments

2:21 min

Using AI for incident summaries and root cause analysis

Jemiah Sius Jemiah Sius +1 · World Congress 2026 Europe

1:48 min

Adapting observability strategies for long-running enterprise AI agents

Christian Heilmann Christian Heilmann +3 · World Congress 2026 Europe

1:05 min

Implementing monitoring and observability for AI software deployments

Alejandro Saucedo Alejandro Saucedo · World Congress 2025

1:13 min

Monitoring enterprise AI workloads for continuous observability

Maxim Salnikov Maxim Salnikov · World Congress 2025

1:41 min

Implementing semi-automated anomaly detection with human oversight

Katja Träumner

1:21 min

Leveraging generative AI for application observability and security

Jemiah Sius Jemiah Sius · World Congress 2023