<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Machine Learning on Netdata</title><link>https://www.netdata.cloud/tags/machine-learning/</link><description>Recent content in Machine Learning on Netdata</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Tue, 23 Jun 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://www.netdata.cloud/tags/machine-learning/index.xml" rel="self" type="application/rss+xml"/><item><title>SNMP &amp; Network Device Monitoring</title><link>https://www.netdata.cloud/features/network/network-device-monitoring/</link><pubDate>Tue, 23 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/network/network-device-monitoring/</guid><description>Real-time SNMP device monitoring with auto-matched vendor profiles, ML anomaly detection, and zero-config discovery.</description></item><item><title>The best AI-powered observability platforms in 2026</title><link>https://www.netdata.cloud/resources/best-ai-powered-observability-platforms/</link><pubDate>Sat, 30 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/resources/best-ai-powered-observability-platforms/</guid><description/></item><item><title>NVIDIA DCGM Collector: Deep GPU Monitoring For AI</title><link>https://www.netdata.cloud/blog/nvidia-dcgm-monitoring/</link><pubDate>Mon, 04 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/nvidia-dcgm-monitoring/</guid><description>&lt;p>&lt;img src="../images/dcgm-collector.svg" alt="NVIDIA DCGM Collector: Deep GPU Monitoring for Data Center and AI Infrastructure">&lt;/p>
&lt;p>GPU infrastructure is expensive and increasingly central to production workloads. Whether you&amp;rsquo;re running ML training jobs, inference serving, video transcoding, or HPC workloads, understanding what your GPUs are actually doing, and what&amp;rsquo;s going wrong when performance degrades, is not optional. The problem is that NVIDIA&amp;rsquo;s Data Center GPU Manager (DCGM) exposes an enormous amount of telemetry, but getting that data into a monitoring system in a useful, organized way has traditionally required significant setup and custom dashboarding work.&lt;/p></description></item><item><title>AI Co-Engineer For Instant Root Cause Insights</title><link>https://www.netdata.cloud/features/aiml/ai-co-engineer/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/aiml/ai-co-engineer/</guid><description>Netdata&amp;rsquo;s AI Co-Engineer combines edge-native machine learning with flexible AI integration, providing instant expert-level insights while keeping your data sovereign and secure.</description></item><item><title>AIOps Platform: Edge-Native ML, Zero Configuration</title><link>https://www.netdata.cloud/features/aiml/aiops/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/aiml/aiops/</guid><description>Enterprise AIOps intelligence without enterprise complexity. Edge-native ML, automated insights, and transparent pricing deliver operational excellence from day one.</description></item><item><title>Algorithmic Dashboards For Every Metric</title><link>https://www.netdata.cloud/features/architecture/algorithmic-dashboards/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/architecture/algorithmic-dashboards/</guid><description>Netdata&amp;rsquo;s algorithmic dashboards transform observability from a months-long configuration project into instant, comprehensive visibility. Every metric visualized automatically, every relationship revealed through intelligent point-and-click analysis, every engineer productive from day one.</description></item><item><title>Anomaly Advisor: Root Cause In Seconds, Not Hours</title><link>https://www.netdata.cloud/features/aiml/anomaly-advisor/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/aiml/anomaly-advisor/</guid><description>Netdata Anomaly Advisor uses edge-native machine learning with 18-model consensus to eliminate 99% of false positives while surfacing root causes in the top 30-50 metrics from thousands collected. Get sub-2-second correlation analysis at any scale without configuration, training delays, or specialist expertise.</description></item><item><title>Anomaly Detection Software: 99% Fewer False Positives</title><link>https://www.netdata.cloud/features/aiml/anomaly-detection/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/aiml/anomaly-detection/</guid><description>Production-ready ML anomaly detection from installation. 18 consensus models per metric achieve 99% false positive reduction in anomaly detection while catching issues competitors miss. Zero configuration. Zero false promises.</description></item><item><title>Anomaly Detection With 99% Fewer False Positives</title><link>https://www.netdata.cloud/features/aiml/machine-learning/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/aiml/machine-learning/</guid><description>Netdata trains 18 independent ML models per metric at the edge, achieving 99% false positive reduction in anomaly detection through unanimous consensus - all included at no additional cost.</description></item><item><title>Application Performance Monitoring Software</title><link>https://www.netdata.cloud/solutions/use-cases/application-performance/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/application-performance/</guid><description>Netdata revolutionizes application performance monitoring with distributed edge intelligence, zero-configuration deployment, and AI-powered troubleshooting that delivers enterprise-grade observability at a fraction of traditional APM costs.</description></item><item><title>Data Center Monitoring Software With Real-Time Metrics</title><link>https://www.netdata.cloud/solutions/use-cases/datacenters/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/datacenters/</guid><description>Transform data center operations with Netdata&amp;rsquo;s real-time monitoring platform. Per-second metrics, automated ML anomaly detection, and AI-powered troubleshooting deliver 80% faster MTTR at 90% lower cost than traditional solutions.</description></item><item><title>DevOps Monitoring Software With ML Anomaly Detection</title><link>https://www.netdata.cloud/solutions/built-for/devops/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/built-for/devops/</guid><description>Netdata delivers complete observability for DevOps teams with per-second metrics, zero-configuration deployment, and AI-powered root cause analysis. Monitor everything from bare metal to Kubernetes with a single platform that replaces 7 tools while reducing costs by 90%.</description></item><item><title>High Cardinality Protection For Unlimited Metrics</title><link>https://www.netdata.cloud/features/architecture/extreme-cardinality/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/architecture/extreme-cardinality/</guid><description>Automated multi-layer protection handles extreme cardinality at the edge, enabling unlimited observability without manual tuning or cost explosions.</description></item><item><title>Hybrid Cloud Monitoring Solution At 90% Lower Cost</title><link>https://www.netdata.cloud/solutions/use-cases/hybrid-cloud-observability/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/hybrid-cloud-observability/</guid><description>Transform hybrid cloud monitoring with Netdata&amp;rsquo;s distributed edge-native architecture. Get per-second visibility, ML-powered insights, and 90% cost savings without pipelines, sampling, or vendor lock-in.</description></item><item><title>Microservices Monitoring Tool Without Sampling</title><link>https://www.netdata.cloud/solutions/use-cases/microservices-observability/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/microservices-observability/</guid><description>Netdata delivers comprehensive microservices observability with per-second metrics, ML-based anomaly detection, and AI-powered troubleshooting - all without the complexity, cost overruns, or tool sprawl that plague traditional solutions.</description></item><item><title>Monitoring Agents For Accurate Production Insights</title><link>https://www.netdata.cloud/product/netdata-agents/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/product/netdata-agents/</guid><description>Discover how Netdata Agents revolutionize infrastructure monitoring with edge-native intelligence, per-second granularity, and zero-configuration deployment. Complete observability in 60 seconds.</description></item><item><title>Netdata vs Amazon CloudWatch | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/cloudwatch/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/cloudwatch/</guid><description>Netdata provides 90% cost reduction vs CloudWatch with superior real-time monitoring (1-second vs 1-60 minutes), comprehensive ML on all metrics, and zero-configuration simplicity. Deploy in 60 seconds and eliminate CloudWatch&amp;rsquo;s unpredictable costs while gaining better visibility.</description></item><item><title>Netdata vs BetterStack | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/betterstack/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/betterstack/</guid><description>Netdata and BetterStack serve different observability needs. Netdata excels at deep infrastructure monitoring with per-second metrics and ML intelligence, while BetterStack focuses on incident coordination and uptime monitoring. Learn which solution fits your technical requirements.</description></item><item><title>Netdata vs CardinalHQ | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/cardinalhq/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/cardinalhq/</guid><description>Comprehensive comparison of Netdata&amp;rsquo;s edge-native monitoring platform versus CardinalHQ&amp;rsquo;s AI-powered observability optimization layer. Learn which solution fits your infrastructure needs.</description></item><item><title>Netdata vs Guance | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/guance/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/guance/</guid><description/></item><item><title>Netdata vs Icinga | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/icinga/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/icinga/</guid><description>Comprehensive comparison of Netdata and Icinga monitoring platforms, highlighting real-time performance analysis, automated deployment, ML-based anomaly detection, and modern observability capabilities.</description></item><item><title>Netdata vs Monit | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/monit/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/monit/</guid><description>Compare Netdata&amp;rsquo;s real-time observability platform with Monit&amp;rsquo;s lightweight process supervision. Learn how Netdata provides per-second monitoring, ML anomaly detection, AI troubleshooting, and unlimited scalability versus Monit&amp;rsquo;s basic service checks and 1,000-host ceiling.</description></item><item><title>Netdata vs New Relic | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/newrelic/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/newrelic/</guid><description/></item><item><title>Netdata vs Prometheus: No PromQL Complexity</title><link>https://www.netdata.cloud/comparisons/prometheus/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/prometheus/</guid><description>Netdata vs Prometheus comparison: Real-time monitoring with zero configuration, built-in ML, and 90% lower costs. See how Netdata solves Prometheus pain points while maintaining enterprise-grade capabilities.</description></item><item><title>Netdata vs Zabbix | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/zabbix/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/zabbix/</guid><description/></item><item><title>Open Source</title><link>https://www.netdata.cloud/open-source/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/open-source/</guid><description>The world&amp;rsquo;s most popular open source monitoring platform. Deploy complete infrastructure observability in 60 seconds with zero configuration. Trusted by millions of engineers worldwide.</description></item><item><title>Real-Time GPU Monitoring For AI Infrastructure</title><link>https://www.netdata.cloud/solutions/industries/ai/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/industries/ai/</guid><description>Monitor AI training clusters, inference APIs, and ML workloads with Netdata&amp;rsquo;s edge-native platform. Per-second visibility, ML-powered insights, predictable pricing.</description></item><item><title>Real-Time Observability For Platform Engineers</title><link>https://www.netdata.cloud/solutions/built-for/platform-engineers/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/built-for/platform-engineers/</guid><description>Empower platform engineering teams with Netdata&amp;rsquo;s distributed observability platform. Get per-second visibility, automated dashboards, ML anomaly detection, and predictable costs - all without query languages or complex pipelines.</description></item><item><title>Real-Time Observability For SRE Teams</title><link>https://www.netdata.cloud/solutions/built-for/sre/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/built-for/sre/</guid><description>Transform SRE operations with Netdata&amp;rsquo;s edge-native observability platform. Get per-second visibility, ML anomaly detection on every metric, and AI-powered troubleshooting at 90% lower cost than traditional solutions.</description></item><item><title>Real-Time Troubleshooting With Sub-2-Second Latency</title><link>https://www.netdata.cloud/features/visualization/troubleshooting/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/visualization/troubleshooting/</guid><description>Interactive debugging with per-second precision and full historical context. Netdata transforms troubleshooting through edge-native ML, automated correlation, and AI-powered analysis.</description></item><item><title>Red Hat OpenShift Monitoring Tool At Scale</title><link>https://www.netdata.cloud/solutions/technologies/redhat-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/technologies/redhat-monitoring/</guid><description>Netdata delivers real-time, AI-powered observability for Red Hat OpenShift with 90% cost savings, per-second metrics, ML anomaly detection on every metric, and zero-configuration deployment. Eliminate Prometheus memory exhaustion, Loki query timeouts, and RHACM complexity.</description></item><item><title>Root Cause Analysis (RCA) With 80% MTTR Reduction</title><link>https://www.netdata.cloud/features/aiml/root-cause-analysis/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/aiml/root-cause-analysis/</guid><description>Netdata&amp;rsquo;s edge-native ML detects anomalies as they happen, automated correlation surfaces root causes in the top 30-50 results, and AI explains incidents in plain English - achieving 80% MTTR reduction at 90% lower cost.</description></item><item><title>Windows Event Log Monitoring Software</title><link>https://www.netdata.cloud/solutions/use-cases/windows-event-logs/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/windows-event-logs/</guid><description>Netdata revolutionizes Windows Event Log monitoring by combining infrastructure metrics with native log access, ML anomaly detection, and AI troubleshooting - delivering comprehensive observability at a fraction of traditional SIEM costs.</description></item><item><title>Our first ML based anomaly alert</title><link>https://www.netdata.cloud/blog/our-first-ml-based-anomaly-alert/</link><pubDate>Wed, 13 Sep 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/our-first-ml-based-anomaly-alert/</guid><description>&lt;p>Over the last few years we have slowly and methodically been building out the &lt;a href="https://learn.netdata.cloud/docs/ml-and-troubleshooting/">ML based capabilities&lt;/a> of the Netdata agent, dogfooding and iterating as we go. To date, these features have mostly been somewhat reactive and tools to aid once you are already troubleshooting.&lt;/p>
&lt;p>Now we feel we are ready to take a first gentle step into some more proactive use cases, starting with a &lt;a href="https://github.com/netdata/netdata/pull/14687">simple node level anomaly rate alert&lt;/a>.&lt;/p></description></item><item><title>Anomaly Rate By Type</title><link>https://www.netdata.cloud/blog/anomaly-rate-by-type/</link><pubDate>Wed, 30 Aug 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/anomaly-rate-by-type/</guid><description>&lt;p>We have &lt;a href="https://github.com/netdata/netdata/pull/15856">recently added&lt;/a> a more detailed anomaly rate chart to Netdata that breaks out the overall &lt;a href="https://learn.netdata.cloud/docs/ml-and-troubleshooting/machine-learning-ml-powered-anomaly-detection#node-anomaly-rate">node anomaly rate&lt;/a> by type, this lets you more easily see what parts of your infrastructure might be experiencing an uptick in anomalies when you see the overall node anomaly rate increase.&lt;/p>
&lt;h2 id="what-is-type">What is &lt;code>type&lt;/code>?&lt;/h2>
&lt;p>&lt;code>type&lt;/code> is generally the prefix of the chart id in Netdata and controls where charts live within the menu on the overview page, for example the &lt;code>mem.available&lt;/code> chart has a type of &lt;code>mem&lt;/code> which in part controls why it lives under the &amp;ldquo;Memory&amp;rdquo; section of the menu.&lt;/p></description></item><item><title>Netdata &amp; Ansible example: ML demo room</title><link>https://www.netdata.cloud/blog/ml-demo-ansible-configuration-management/</link><pubDate>Fri, 07 Jul 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/ml-demo-ansible-configuration-management/</guid><description>&lt;p>We are always trying to lower the barrier to entry when it comes to monitoring and observability and one place we have consistently witnessed some pain from users is around adopting and approaching &lt;a href="https://www.atlassian.com/microservices/microservices-architecture/configuration-management">configuration management&lt;/a> tools and practices as your infrastructure grows and becomes more complex.&lt;/p>
&lt;p>To that end, we have begun recently publishing our own &lt;a href="https://github.com/netdata/community/tree/main/configuration-management/ansible-ml-demo">little example ansible project&lt;/a> used to maintain and manage the servers used in our public &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/rooms/machine-learning/overview">Machine Learning Demo room&lt;/a>.&lt;/p></description></item><item><title>How Netdata's ML-based Anomaly Detection Works</title><link>https://www.netdata.cloud/blog/how-netdatas-ml-based-anomaly-detection-works/</link><pubDate>Tue, 23 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/how-netdatas-ml-based-anomaly-detection-works/</guid><description>&lt;p>&lt;img src="../2023-05-23-how-netdatas-ml-based-anomaly-detection-works/img/img.png" alt="title image">&lt;/p>
&lt;p>How does Netdata&amp;rsquo;s &lt;a href="https://learn.netdata.cloud/docs/troubleshooting-and-machine-learning/machine-learning-ml-powered-anomaly-detection">machine learning (ML) based anomaly detection&lt;/a> actually work? Read on to find out!&lt;/p>
&lt;!--truncate-->
&lt;h2 id="design-considerations">Design considerations&lt;/h2>
&lt;p>Lets first start with some of the key design considerations and principles of Netdata&amp;rsquo;s anomaly detection (&lt;em>and some comments in parenthesis along the way&lt;/em>):&lt;/p>
&lt;ol>
&lt;li>We don&amp;rsquo;t have any labels or examples of previous anomalies. This means we are in an &lt;a href="https://en.wikipedia.org/wiki/Unsupervised_learning">unsupervised setting&lt;/a> (&lt;em>best we can try to do is learn what &amp;ldquo;normal&amp;rdquo; data looks like assuming the collected data is &amp;ldquo;mostly&amp;rdquo; normal&lt;/em>).&lt;/li>
&lt;li>Needs to be lightweight and run on the agent (&lt;em>or a parent&lt;/em>).
&lt;ul>
&lt;li>Need to be very careful of impact on CPU overhead when training and scoring (&lt;em>lots of cheap models are better than a few expensive and heavy ones&lt;/em>).&lt;/li>
&lt;li>Models themselves need to be small so as to not drastically increase the agents memory footprint (&lt;em>model objects need to be small for storage&lt;/em>).&lt;/li>
&lt;li>This has implications for the ML formulation (&lt;em>sorry - no deep learning models yet) and its implementation (we need to be surgical and optimized&lt;/em>).&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Needs to scale for thousands of metrics and score in realtime every second as metrics are collected.
&lt;ul>
&lt;li>Typical Netdata nodes have thousands of metrics and we want to be able to score every metric every second with minimal latency overhead (&lt;em>we need to use sensible approaches to training like spreading the training cost over a wide training window&lt;/em>).&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Needs to be able to handle a wide variety of metrics.
&lt;ul>
&lt;li>There is no single perfect model or approach for all types of metrics so we need a good all rounder that can work well enough across any and all different types of time seres metrics (&lt;em>for any given metric of course you could handcraft a better model but thats not feasible here, we need something like a &amp;ldquo;weak learners&amp;rdquo; approach of lots of generally useful models adding up to &amp;ldquo;more than the sum of their parts&amp;rdquo;&lt;/em>).&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Needs to be written in C or C++ as that is the language of the Netdata agent (&lt;em>We are using &lt;a href="https://github.com/davisking/dlib">dlib&lt;/a> for the current implementation&lt;/em>).&lt;/li>
&lt;li>We Need to be very careful about taking big or complex dependencies if using third party libraries.
&lt;ul>
&lt;li>We want to be able to easily build and deploy Netdata on any Linux system without having to worry about installing or managing complex dependencies (&lt;em>we need to be careful of more complex algorithms that would have larger dependencies and potentially limit where Netdata can run&lt;/em>).&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ol>
&lt;p>The above considerations are important and useful to keep in mind as we explore the system in more detail.&lt;/p></description></item><item><title>Transform Monitoring With A Machine Learning Approach</title><link>https://www.netdata.cloud/blog/transform-monitoring-ml-first-approach/</link><pubDate>Thu, 11 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/transform-monitoring-ml-first-approach/</guid><description>&lt;p>Unlocking the full potential of monitoring through ML integration, anomaly detection, and innovative scoring engines.&lt;/p>
&lt;!--truncate-->
&lt;p>Machine Learning has been making waves in various industries, but its adoption in the monitoring and observability space has been slower than expected. Many “ML” features remain gimmicky and do not provide actual real world value to users that encourages their further use.&lt;/p>
&lt;p>At Netdata, we firmly believe that ML is crucial for monitoring, and we&amp;rsquo;ve taken an ML-first approach to provide users with powerful tools and insights. In this blog post, we&amp;rsquo;ll discuss the reasons behind our belief in ML, how we&amp;rsquo;ve integrated ML into our charts and visualizations, our query engines, the scoring engine we&amp;rsquo;ve built, and how these innovations enable metrics correlations and anomaly advisor.&lt;/p></description></item><item><title>Netdata's AI Insights &amp; Rapid Diagnostics</title><link>https://www.netdata.cloud/blog/netdata-ai-insights-rapid-diagnostics/</link><pubDate>Wed, 19 Apr 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-ai-insights-rapid-diagnostics/</guid><description>&lt;p>Introduction to Netdata&amp;rsquo;s new visualisation providing AI Insights, supporting Rapid Diagnostics.
&lt;img src="https://user-images.githubusercontent.com/96257330/233125254-f93c9520-0a3f-4844-8d43-1f3202a5e411.png" alt="logo">&lt;/p>
&lt;!--truncate-->
&lt;h2 id="a-new-era-in-monitoring-systems-dashboards">A New Era in Monitoring Systems Dashboards&lt;/h2>
&lt;p>We&amp;rsquo;re thrilled to share an important upgrade to Netdata: &lt;strong>AI Insights &amp;amp; Rapid Diagnostics&lt;/strong>, a technology aiming to redefine what we expect from a monitoring system.&lt;/p>
&lt;h2 id="challenges-with-traditional-monitoring-dashboards">Challenges with Traditional Monitoring Dashboards&lt;/h2>
&lt;p>Traditional monitoring systems rely on a query language to help engineers create dashboards and alerts. While these languages offer power and flexibility, they come with several challenges that make monitoring and troubleshooting more complex and time-consuming:&lt;/p></description></item><item><title>Anomaly Rates in the Menu!</title><link>https://www.netdata.cloud/blog/anomaly-rates-in-the-menu/</link><pubDate>Wed, 29 Mar 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/anomaly-rates-in-the-menu/</guid><description>&lt;p>The menu (on the &lt;a href="https://learn.netdata.cloud/docs/getting-started/monitor-your-infrastructure/home-overview-and-single-node-view#overview-and-single-node-view">overview or single node tab&lt;/a>) now has an &lt;a href="https://learn.netdata.cloud/docs/troubleshooting-and-machine-learning/machine-learning-ml-powered-anomaly-detection#anomaly-rate">anomaly rate&lt;/a> button built into it that, for the entire visible window or a highlighted time range, shows the maximum chart anomaly rate within each section.&lt;/p>
&lt;p>Read on to learn more about this new feature!&lt;/p>
&lt;iframe width="560" height="315" src="https://www.youtube.com/embed/PgVh_MFHMb0?si=F2Mq6wIxHJWaHykZ" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen>&lt;/iframe>
&lt;h2 id="wait-what-is-an-anomaly-rate">Wait, what is an anomaly rate?&lt;/h2>
&lt;p>Netdata is the only monitoring agent that natively (for every metric, with zero config and sane defaults) produces anomaly rates in addition to just collecting raw metrics.&lt;/p></description></item><item><title>Anomaly detection on Prometheus metrics</title><link>https://www.netdata.cloud/blog/anomaly-detection-on-prometheus-metrics/</link><pubDate>Wed, 01 Mar 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/anomaly-detection-on-prometheus-metrics/</guid><description>&lt;p>&lt;img src="../2023-03-01-anomaly-detection-on-prometheus-metrics/img/img.png" alt="img">&lt;/p>
&lt;p>We have recently extended the native machine learning (ML) based anomaly detection &lt;a href="https://learn.netdata.cloud/guides/monitor/anomaly-detection">capabilities&lt;/a> of Netdata to &lt;a href="https://github.com/netdata/netdata/issues/14218">support all metrics&lt;/a>, regardless on their collection frequency (&lt;code>update every&lt;/code>).&lt;/p>
&lt;p>Previously only metrics collected every second were supported, but now Netdata can run anomaly detection out of the box with zero config on metrics with any collection frequency.&lt;/p>
&lt;p>This post will illustrate an example of what this means using &lt;a href="https://prometheus.io/">Prometheus&lt;/a> metrics (via the &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/prometheus#gsc.tab=0">Netdata Prometheus collector&lt;/a>) since they typically have a default collection frequency of 10 seconds.&lt;/p></description></item><item><title>Extending Netdata's anomaly detection training window</title><link>https://www.netdata.cloud/blog/extending-anomaly-detection-training-window/</link><pubDate>Thu, 02 Feb 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/extending-anomaly-detection-training-window/</guid><description>&lt;p>We have been busy at work under the hood of the Netdata agent to introduce new capabilities that let you extend the &amp;ldquo;training window&amp;rdquo; used by Netdata&amp;rsquo;s &lt;a href="https://learn.netdata.cloud/docs/nightly/setup/configure-machine-learning-ml-powered-anomaly-detection">native anomaly detection capabilities&lt;/a>.&lt;/p>
&lt;p>This blog post will discuss one of these improvements to help you reduce &amp;ldquo;&lt;a href="https://en.wikipedia.org/wiki/False_positives_and_false_negatives#False_positive_error">false positives&lt;/a>&amp;rdquo; by essentially extending the training window by using the new (beautifully named) &lt;code>number of models per dimension&lt;/code> configuration parameter.&lt;/p>
&lt;h2 id="background">Background&lt;/h2>
&lt;p>One of the most important considerations of our native anomaly detection capabilities is the overhead of running the training and scoring computations required to train thousands of models (one per metric) and produce &lt;a href="https://learn.netdata.cloud/docs/nightly/setup/configure-machine-learning-ml-powered-anomaly-detection#anomaly-bit">anomaly bits&lt;/a> every second based on those trained models.&lt;/p></description></item><item><title>Data Collection Strategies For Infrastructure</title><link>https://www.netdata.cloud/blog/data-collection-strategies/</link><pubDate>Tue, 06 Sep 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/data-collection-strategies/</guid><description>&lt;!--truncate-->
&lt;p>Monitoring and troubleshooting; unfortunately, these terms are still used interchangeably, which can lead to misunderstandings about data collection strategies.&lt;/p>
&lt;p>In this article we aim to clarify some important definitions, processes, and common data collection strategies for monitoring solutions. We will specify the limitations of the described strategies, as well as key benefits which can potentially be also used for troubleshooting needs.&lt;/p>
&lt;p>&lt;strong>IT infrastructure monitoring&lt;/strong> is a business process of collecting and analyzing data over a period of time to improve business results.&lt;/p></description></item><item><title>How Netdata’s Machine Learning works</title><link>https://www.netdata.cloud/blog/how-netdatas-machine-learning-works/</link><pubDate>Thu, 01 Sep 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/how-netdatas-machine-learning-works/</guid><description>&lt;p>Following on from the &lt;a href="https://www.netdata.cloud/blog/introducing-anomaly-advisor-unsupervised-anomaly-detection-in-netdata" target="_blank" rel="noopener">recent launch&lt;/a> of our &lt;a href="https://learn.netdata.cloud/docs/cloud/insights/anomaly-advisor" target="_blank" rel="noopener">Anomaly Advisor&lt;/a> feature, and in keeping with &lt;a href="https://www.netdata.cloud/blog/our-approach-to-machine-learning/" target="_blank" rel="noopener">our approach to machine learning&lt;/a>, &lt;a href="https://github.com/netdata/netdata/blob/master/ml/notebooks/netdata_anomaly_detection_deepdive.ipynb" target="_blank" rel="noopener">here&lt;/a> is a detailed Python notebook outlining exactly how the machine learning powering the Anomaly Advisor actually works under the hood.&lt;/p>
&lt;!--truncate-->
&lt;p>Or if you&amp;rsquo;d rather watch a video walkthrough of the notebook then check out below.&lt;/p>
&lt;iframe width="560" height="315" src="https://www.youtube.com/embed/L1xleckyuDQ?si=rptYzWE-eLlhSL9x" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen>&lt;/iframe>
&lt;p>Try it for yourself, &lt;a href="https://learn.netdata.cloud/docs/cloud/get-started" target="_blank" rel="noopener">get started&lt;/a> by &lt;a href="https://app.netdata.cloud/?utm_source=blog&amp;amp;utm_content=how_netdata_ml_works" target="_blank" rel="noopener">signing in to Netdata&lt;/a> and connecting a node. Once initial models have been trained (usually after the agent has about one hour of data, zero configuration needed), you&amp;rsquo;ll be able to start exploring in the &lt;a href="https://learn.netdata.cloud/docs/cloud/insights/anomaly-advisor" target="_blank" rel="noopener">Anomaly Advisor&lt;/a> tab of Netdata.&lt;/p></description></item><item><title>Anomaly rate in every chart</title><link>https://www.netdata.cloud/blog/anomaly-rate-in-every-chart/</link><pubDate>Thu, 23 Jun 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/anomaly-rate-in-every-chart/</guid><description>&lt;p>A month ago, we introduced unsupervised ML &amp;amp; Anomaly Detection in Netdata, the &lt;a href="https://www.netdata.cloud/blog/introducing-anomaly-advisor-unsupervised-anomaly-detection-in-netdata/">Anomaly Advisor&lt;/a>. Today, we’re happy to announce that we’re bringing anomaly rates to every chart in Netdata Cloud. Anomaly information is no longer limited to the Anomalies tab and will be accessible to you from the Overview and Single Node View tabs as well. This will make your troubleshooting journey easier, as you will have the anomaly rates for any metric available with a single click. Whichever metric or chart you&amp;rsquo;re exploring will be instant.&lt;/p></description></item><item><title>Metric Correlations on the Agent</title><link>https://www.netdata.cloud/blog/metric-correlations-on-the-agent/</link><pubDate>Wed, 15 Jun 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/metric-correlations-on-the-agent/</guid><description>&lt;p>As of &lt;a href="https://github.com/netdata/netdata/releases/tag/v1.35.0" target="_blank" rel="noopener">&lt;code>v1.35.0&lt;/code>&lt;/a> the Netdata Agent can now run &lt;a href="https://learn.netdata.cloud/docs/cloud/insights/metric-correlations" target="_blank" rel="noopener">Metric Correlations&lt;/a> (MC) itself. This means that, for nodes with MC enabled, the Metric Correlations feature just got a whole lot faster!&lt;/p>
&lt;!--truncate-->
&lt;p>The Netdata Metric Correlations feature uses a &lt;a href="https://en.wikipedia.org/wiki/Kolmogorov%E2%80%93Smirnov_test#Two-sample_Kolmogorov%E2%80%93Smirnov_test" target="_blank" rel="noopener">Two Sample Kolmogorov-Smirnov test&lt;/a> to look for which metrics have a significant distributional change around a highlighted window of interest. This can be useful when you are interested in short term &amp;ldquo;&lt;a href="https://en.wikipedia.org/wiki/Change_detection" target="_blank" rel="noopener">change detection&lt;/a>&amp;rdquo; and want to try answer the question &amp;ldquo;what else changed around this time?&amp;rdquo;.&lt;/p></description></item><item><title>Anomaly Advisor: Unsupervised Anomaly Detection</title><link>https://www.netdata.cloud/blog/introducing-anomaly-advisor-unsupervised-anomaly-detection-in-netdata/</link><pubDate>Thu, 26 May 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/introducing-anomaly-advisor-unsupervised-anomaly-detection-in-netdata/</guid><description>&lt;p>Today we are excited to launch one of our flagship ML assisted troubleshooting features in Netdata – the Anomaly Advisor.&lt;/p>
&lt;p>The Anomaly Advisor builds on earlier work to introduce unsupervised &lt;a href="https://github.com/netdata/netdata/blob/master/ml/README.md">anomaly detection&lt;/a> capabilities into the &lt;a href="https://www.netdata.cloud/agent/">Netdata Agent&lt;/a> from &lt;a href="https://github.com/netdata/netdata/releases/tag/v1.32.0">v1.32.0&lt;/a> onwards.&lt;/p>
&lt;h2 id="getting-started">Getting Started&lt;/h2>
&lt;p>Once you &lt;a href="https://learn.netdata.cloud/docs/configure/machine-learning#configuration">enable ML&lt;/a> on your nodes, each node will begin producing an &amp;ldquo;&lt;a href="https://learn.netdata.cloud/docs/configure/machine-learning#anomaly-bit---100--anomalous-0--normal">Anomaly Bit&lt;/a>&amp;rdquo; every second in addition to raw metric values. This anomaly bit will be 1 when the trained ML models consider recent raw data for a metric to look anomalous or 0 when things look &amp;rsquo;normal&amp;rsquo;. The Anomaly Advisor leverages this information to enable seamless space or room level anomaly detection out of the box with minimal configuration.&lt;/p></description></item><item><title>CNCF Live: Machine Learning Anomaly Detection</title><link>https://www.netdata.cloud/blog/cncf-live-power-up-your-machine-learning-automated-anomaly-detection/</link><pubDate>Wed, 27 Apr 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/cncf-live-power-up-your-machine-learning-automated-anomaly-detection/</guid><description>&lt;h2 id="join-us-live-to-talk-ml">Join Us Live to Talk ML&lt;/h2>
&lt;p>Join ML Lead Andrew Maguire and Product Manager Shyam Sreevalsan on the 23rd of June at 5pm UTC for the Netdata Machine Learning Meetup, which will be livestreamed on Youtube. In this session, we will demo and preview new and future features, along with a Q&amp;amp;A, and discuss the following topics:&lt;/p>
&lt;ul>
 	&lt;li>The role of Machine Learning in DevOps Infrastructure Monitoring &amp;amp; Troubleshooting&lt;/li>
 	&lt;li>The Netdata way of approaching Machine Learning&lt;/li>
 	&lt;li>Challenges of building Macvhine learning solutions that are useful and user friendly.&lt;/li>
&lt;/ul>
Feel free to ask questions or share ideas on &lt;a href="https://discord.gg/ZyeDHQTdaW">our Community Discord&lt;/a>, or &lt;a href="https://(https://www.meetup.com/netdata-infrastructure-monitoring-meetup-group/events/286243158/">RSVP to the event&lt;/a>. We look forward to seeing you then!
&lt;h2 id="cncf-live-power-up-your-machine-learning---automated-anomaly-detection">CNCF Live: Power up your machine learning - Automated anomaly detection&lt;/h2>
&lt;p>Our Analytics &amp;amp; ML lead Andrew Maguire recently had a chance to share our new &lt;a href="https://community.netdata.cloud/t/anomaly-advisor-beta-launch/2717">Anomaly Advisor&lt;/a> feature with the wider CNCF community. In his demonstration he did some light chaos engineering (using &lt;a href="https://www.gremlin.com/">Gremlin&lt;/a> and &lt;a href="https://wiki.ubuntu.com/Kernel/Reference/stress-ng">stress-ng&lt;/a>) to generate some real anomalies on his infrastructure and watch how it all played out in the Anomaly Advisor in Netdata Cloud.&lt;/p></description></item><item><title>Our Approach to Machine Learning</title><link>https://www.netdata.cloud/blog/our-approach-to-machine-learning/</link><pubDate>Fri, 25 Mar 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/our-approach-to-machine-learning/</guid><description>&lt;p>There is a lot of buzz in the world of machine learning (ML) and as a layperson it can be hard to keep up with it all. Therefore, we decided to write down some of our thoughts and musings on how &lt;b>we&lt;/b> are approaching ML at Netdata.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="our-approach-to-machine-learning-ml">Our Approach to Machine Learning (ML)&lt;/h2>
&lt;p>We’ll touch on the current state of applied ML in industry in general, and zoom in on ML in the monitoring industry. We’ll discuss how we can leverage “good honest ML” to punch above our weight and add some useful and novel features for our users over the next few years.&lt;/p></description></item><item><title>Root cause analysis using Metric Correlations</title><link>https://www.netdata.cloud/blog/root-cause-analysis-using-metric-correlations/</link><pubDate>Fri, 03 Sep 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/root-cause-analysis-using-metric-correlations/</guid><description>&lt;!--truncate-->
&lt;figure class="wp-block-image size-large">&lt;img src="../wp-archive/uploads/2022/03/Screen-Shot-2021-09-03-at-1.43.32-PM-1-1200x608.png" alt="" class="wp-image-16297"/>&lt;/figure>
&lt;p>As complexity of systems and applications continue to evolve and change, the number of metrics that need to be monitored grows in parallel. Whether you’re on a DevOps team, an SRE, or a developer building the code yourself, many of these components may be fragmented across your infrastructure, making it increasingly difficult to identify the root cause when experiencing downtime or abnormal behavior. To help solve this challenge, we built the &lt;a href="https://learn.netdata.cloud/docs/cloud/insights/metric-correlations">Metric Correlations&lt;/a> feature – an automated analysis tool that evaluates all your metrics to identify which have changed the most within a given period of interest.&lt;/p></description></item><item><title>Metric Correlations: Detect Patterns &amp; Anomalies</title><link>https://www.netdata.cloud/blog/netdata-cloud-metric-correlations/</link><pubDate>Wed, 16 Sep 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-cloud-metric-correlations/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-large wp-image-16623" src="../wp-archive/uploads/2022/03/Cloud-Correlations@2x-1200x826.png" alt="" width="1200" height="826" />
&lt;p>Today, we are excited to launch our first Netdata Cloud Insights feature, Metric Correlations, developed for discovering underlying issues more quickly and identifying the root cause more efficiently. Read on to learn more about our approach to developing this new feature, how it works, and the many benefits you’ll find incorporating this into your team’s troubleshooting workflow.&lt;/p>
&lt;h2>Some background&lt;/h2>
Let’s start with a bit of a disclaimer. It seems machine learning (ML) (or “Artificial Intelligence,” if you are looking for more LinkedIn likes) has gone mainstream in the last few years, and we are probably by now somewhere near the “Peak of Inflated Expectations” on the &lt;a title="hype cycle" href="https://en.wikipedia.org/wiki/Hype_cycle" target="_blank" rel="noopener noreferrer">hype cycle&lt;/a>. It is in this context that we want to be clear about what our goals are in this space and our approach to releasing data-driven features that draw on techniques from statistics and ML. In short, we want to be clear, open, realistic, and avoid buzzwords at all costs!
&lt;p>Over the next 12 months, we are hoping to begin building a layer of intelligence&lt;sup>&lt;a href="https://staging-www.netdata.cloud/blog/netdata-cloud-metric-correlations/#1">1&lt;/a>&lt;/sup> throughout Netdata (both Cloud and Agent) to assist with “&lt;a title="human in the loop" href="https://hai.stanford.edu/blog/humans-loop-design-interactive-ai-systems" target="_blank" rel="noopener noreferrer">human in the loop&lt;/a>” troubleshooting, mainly to help users more easily surface slowdowns, anomalies, or other issues and lower your &lt;a title="cognitive load" href="https://en.wikipedia.org/wiki/Cognitive_load" target="_blank" rel="noopener noreferrer">cognitive load&lt;/a>&lt;sup>&lt;a href="https://staging-www.netdata.cloud/blog/netdata-cloud-metric-correlations/#2">2&lt;/a>&lt;/sup> as you troubleshoot using Netdata. Simply put, we’re working to streamline your mean time to resolution (MTTR).&lt;/p></description></item><item><title>Contribute to Netdata’s machine learning efforts!</title><link>https://www.netdata.cloud/blog/contribute-machine-learning/</link><pubDate>Mon, 16 Mar 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/contribute-machine-learning/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-full wp-image-16783" src="../wp-archive/uploads/2022/03/contribute-machine-learning.png" alt="" width="991" height="1072" />
&lt;p>Netdata contributors have greatly influenced the growth of our company and are essential to our success. The time and expertise that contributors volunteer are fundamental to our goal of helping you build extraordinary infrastructures. We highly value end-user feedback during product development, which is why we’re looking to involve you in progressing our machine learning (ML) efforts! &lt;span id="more-2975">&lt;/span>As we are continually looking for ways to improve and enhance Netdata, we are starting to explore how we can leverage machine learning to introduce new product features. Our main focus at the moment is around automated &lt;a href="https://en.wikipedia.org/wiki/Anomaly_detection">anomaly detection&lt;/a>. This is a really interesting and challenging problem (high volume, high dimensional data, lack of ground truth labels, and so on), but we should be able to use some of the metrics monitored by Netdata to deliver new, awesome product features and user experiences (AI is the &lt;a href="https://www.gsb.stanford.edu/insights/andrew-ng-why-ai-new-electricity">new electricity&lt;/a>, after all 😃). However, developing ML-driven product features is quite different than traditional software development (see steps 1 to 7 in the picture above). Mainly, this is because you never really know what specific data transformations, problem formulation, and sets of algorithms will work best in advance. (&lt;a href="https://www.kdnuggets.com/2019/09/no-free-lunch-data-science.html">Here&lt;/a> is a good article explaining things, and if you really want to go down a rabbit hole, check out this &lt;a href="https://ai.stackexchange.com/questions/15650/what-are-the-implications-of-the-no-free-lunch-theorem-for-machine-learning">Stack Overflow question&lt;/a> and this &lt;a href="https://www.quora.com/What-does-the-No-Free-Lunch-theorem-mean-for-machine-learning-In-what-ways-do-popular-ML-algorithms-overcome-the-limitations-set-by-this-theorem">Quora thread&lt;/a>). Ideally, you first need to prototype your solution “in the lab” on some data you have already collected and do a few iterations of data → problem formulation → prototype. This process gives you a level of confidence in what you are doing (and some data to back it up) to move on to the even-more-complicated step of going from prototype to production. At Netdata, we are currently trying to get to step 4, where we can first prototype some solutions on real-world data and come up with ways to measure progress.&lt;/p></description></item><item><title>Netdata vs Chronosphere | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/chronosphere/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/chronosphere/</guid><description>Netdata provides edge-native observability with per-second granularity, automatic ML anomaly detection, and transparent pricing—eliminating PromQL learning curves and SaaS-only limitations that challenge Chronosphere users.</description></item></channel></rss>