<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>ML on Netdata</title><link>https://www.netdata.cloud/tags/ml/</link><description>Recent content in ML on Netdata</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 18 Dec 2025 00:00:00 +0000</lastBuildDate><atom:link href="https://www.netdata.cloud/tags/ml/index.xml" rel="self" type="application/rss+xml"/><item><title>Infrastructure Monitoring Software With AI &amp; ML</title><link>https://www.netdata.cloud/solutions/use-cases/infrastructure-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/infrastructure-monitoring/</guid><description>Transform infrastructure monitoring with Netdata&amp;rsquo;s edge-native platform. Get per-second visibility, ML anomaly detection on every metric, and AI-powered troubleshooting - all while keeping your data sovereign and reducing costs by 90%.</description></item><item><title>Netdata vs Cribl | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/cribl/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/cribl/</guid><description>Netdata provides real-time infrastructure monitoring with ML-based anomaly detection and built-in dashboards. Cribl optimizes log pipelines and routes data for cost efficiency. Learn how these complementary solutions work together for complete observability.</description></item><item><title>Real-Time Monitoring &amp; ML Detection For Sysadmins</title><link>https://www.netdata.cloud/solutions/built-for/sysadmins/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/built-for/sysadmins/</guid><description>Purpose-built observability for sysadmins: per-second metrics, automatic dashboards, ML-powered anomaly detection, and predictable pricing. Monitor everything from bare metal to Kubernetes without learning query languages or building dashboards.</description></item><item><title>Netdata Best Practices</title><link>https://www.netdata.cloud/blog/netdata-best-practices/</link><pubDate>Fri, 03 Nov 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-best-practices/</guid><description>&lt;p>Effective &lt;strong>system monitoring&lt;/strong> is non-negotiable in today&amp;rsquo;s complex IT environments. Netdata offers real-time performance and health monitoring with precision and granularity. But the key to harnessing its full potential lies in the optimization of your setup. Let’s ensure you are not just collecting data, but doing it in the most optimal way while gaining actionable insights from it.&lt;/p>
&lt;p>The starting point for optimization is a robust setup. Netdata is engineered for minimal footprint and can run on a wide range of hardware—from IoT devices to powerful servers. Time for a deep dive into each of these key areas and what the best practices you should follow, if you are serious about monitoring and optimizing your Netdata monitoring setup:&lt;/p></description></item><item><title>Discover The New Netdata!</title><link>https://www.netdata.cloud/blog/discover-the-new-netdata/</link><pubDate>Fri, 27 Oct 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/discover-the-new-netdata/</guid><description>&lt;p>Missed the last &lt;strong>Netdata&lt;/strong> updates? Here is what is new:&lt;/p>
&lt;h2 id="explore-your-systemd-journal-logs-with-netdata">Explore your systemd-journal logs with Netdata&lt;/h2>
&lt;p>&lt;img src="https://github.com/netdata/blog/assets/139226121/7d2779c9-0efb-4491-8fe3-aedce1dc72fb" alt="systemd-journal-logs">&lt;/p>
&lt;p>Netdata &lt;a href="https://learn.netdata.cloud/docs/logs/systemd-journal/?utm_source=IL&amp;amp;utm_medium=internallinking&amp;amp;utm_campaign=new_netada">got a &lt;code>systemd&lt;/code>-journal logs explorer&lt;/a> to analyze your &lt;code>systemd&lt;/code>-journal logs, directly on their sources. By just installing &lt;strong>Netdata&lt;/strong> on any systemd based system, Netdata automatically finds all the &lt;strong>journal sources&lt;/strong> and presents a powerful dashboard to explore, search, filter and analyze your &lt;strong>logs&lt;/strong>. It works on both individual servers and journal centralization servers.&lt;/p></description></item><item><title>Our first ML based anomaly alert</title><link>https://www.netdata.cloud/blog/our-first-ml-based-anomaly-alert/</link><pubDate>Wed, 13 Sep 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/our-first-ml-based-anomaly-alert/</guid><description>&lt;p>Over the last few years we have slowly and methodically been building out the &lt;a href="https://learn.netdata.cloud/docs/ml-and-troubleshooting/">ML based capabilities&lt;/a> of the Netdata agent, dogfooding and iterating as we go. To date, these features have mostly been somewhat reactive and tools to aid once you are already troubleshooting.&lt;/p>
&lt;p>Now we feel we are ready to take a first gentle step into some more proactive use cases, starting with a &lt;a href="https://github.com/netdata/netdata/pull/14687">simple node level anomaly rate alert&lt;/a>.&lt;/p></description></item><item><title>Anomaly Rate By Type</title><link>https://www.netdata.cloud/blog/anomaly-rate-by-type/</link><pubDate>Wed, 30 Aug 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/anomaly-rate-by-type/</guid><description>&lt;p>We have &lt;a href="https://github.com/netdata/netdata/pull/15856">recently added&lt;/a> a more detailed anomaly rate chart to Netdata that breaks out the overall &lt;a href="https://learn.netdata.cloud/docs/ml-and-troubleshooting/machine-learning-ml-powered-anomaly-detection#node-anomaly-rate">node anomaly rate&lt;/a> by type, this lets you more easily see what parts of your infrastructure might be experiencing an uptick in anomalies when you see the overall node anomaly rate increase.&lt;/p>
&lt;h2 id="what-is-type">What is &lt;code>type&lt;/code>?&lt;/h2>
&lt;p>&lt;code>type&lt;/code> is generally the prefix of the chart id in Netdata and controls where charts live within the menu on the overview page, for example the &lt;code>mem.available&lt;/code> chart has a type of &lt;code>mem&lt;/code> which in part controls why it lives under the &amp;ldquo;Memory&amp;rdquo; section of the menu.&lt;/p></description></item><item><title>Netdata Assistant: Your AI-Powered Troubleshooting Sidekick</title><link>https://www.netdata.cloud/blog/netdata-assistant/</link><pubDate>Fri, 14 Jul 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-assistant/</guid><description>&lt;p>Hey there! We&amp;rsquo;re excited to share a new troubleshooting feature we have added to Netdata, the Netdata Assistant. We&amp;rsquo;ve built this tool to help you troubleshoot more effectively and with less stress. Let&amp;rsquo;s dive in.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="whats-the-netdata-assistant">What&amp;rsquo;s the Netdata Assistant?&lt;/h2>
&lt;p>The Netdata Assistant is an AI tool that uses large language models and our community&amp;rsquo;s knowledge to guide you during troubleshooting.&lt;/p>
&lt;p>Here&amp;rsquo;s a scenario. It&amp;rsquo;s 3 am and you get an alert. Instead of scrambling to Google what&amp;rsquo;s going on, you can just click on the assistant button. The Netdata Assistant will give you the lowdown on the alert, why it&amp;rsquo;s happening, and why you should care. It&amp;rsquo;ll also guide you on how to troubleshoot it and even offer some handy web links for more info, if you&amp;rsquo;re interested.&lt;/p></description></item><item><title>Netdata &amp; Ansible example: ML demo room</title><link>https://www.netdata.cloud/blog/ml-demo-ansible-configuration-management/</link><pubDate>Fri, 07 Jul 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/ml-demo-ansible-configuration-management/</guid><description>&lt;p>We are always trying to lower the barrier to entry when it comes to monitoring and observability and one place we have consistently witnessed some pain from users is around adopting and approaching &lt;a href="https://www.atlassian.com/microservices/microservices-architecture/configuration-management">configuration management&lt;/a> tools and practices as your infrastructure grows and becomes more complex.&lt;/p>
&lt;p>To that end, we have begun recently publishing our own &lt;a href="https://github.com/netdata/community/tree/main/configuration-management/ansible-ml-demo">little example ansible project&lt;/a> used to maintain and manage the servers used in our public &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/rooms/machine-learning/overview">Machine Learning Demo room&lt;/a>.&lt;/p></description></item><item><title>How Netdata's ML-based Anomaly Detection Works</title><link>https://www.netdata.cloud/blog/how-netdatas-ml-based-anomaly-detection-works/</link><pubDate>Tue, 23 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/how-netdatas-ml-based-anomaly-detection-works/</guid><description>&lt;p>&lt;img src="../2023-05-23-how-netdatas-ml-based-anomaly-detection-works/img/img.png" alt="title image">&lt;/p>
&lt;p>How does Netdata&amp;rsquo;s &lt;a href="https://learn.netdata.cloud/docs/troubleshooting-and-machine-learning/machine-learning-ml-powered-anomaly-detection">machine learning (ML) based anomaly detection&lt;/a> actually work? Read on to find out!&lt;/p>
&lt;!--truncate-->
&lt;h2 id="design-considerations">Design considerations&lt;/h2>
&lt;p>Lets first start with some of the key design considerations and principles of Netdata&amp;rsquo;s anomaly detection (&lt;em>and some comments in parenthesis along the way&lt;/em>):&lt;/p>
&lt;ol>
&lt;li>We don&amp;rsquo;t have any labels or examples of previous anomalies. This means we are in an &lt;a href="https://en.wikipedia.org/wiki/Unsupervised_learning">unsupervised setting&lt;/a> (&lt;em>best we can try to do is learn what &amp;ldquo;normal&amp;rdquo; data looks like assuming the collected data is &amp;ldquo;mostly&amp;rdquo; normal&lt;/em>).&lt;/li>
&lt;li>Needs to be lightweight and run on the agent (&lt;em>or a parent&lt;/em>).
&lt;ul>
&lt;li>Need to be very careful of impact on CPU overhead when training and scoring (&lt;em>lots of cheap models are better than a few expensive and heavy ones&lt;/em>).&lt;/li>
&lt;li>Models themselves need to be small so as to not drastically increase the agents memory footprint (&lt;em>model objects need to be small for storage&lt;/em>).&lt;/li>
&lt;li>This has implications for the ML formulation (&lt;em>sorry - no deep learning models yet) and its implementation (we need to be surgical and optimized&lt;/em>).&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Needs to scale for thousands of metrics and score in realtime every second as metrics are collected.
&lt;ul>
&lt;li>Typical Netdata nodes have thousands of metrics and we want to be able to score every metric every second with minimal latency overhead (&lt;em>we need to use sensible approaches to training like spreading the training cost over a wide training window&lt;/em>).&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Needs to be able to handle a wide variety of metrics.
&lt;ul>
&lt;li>There is no single perfect model or approach for all types of metrics so we need a good all rounder that can work well enough across any and all different types of time seres metrics (&lt;em>for any given metric of course you could handcraft a better model but thats not feasible here, we need something like a &amp;ldquo;weak learners&amp;rdquo; approach of lots of generally useful models adding up to &amp;ldquo;more than the sum of their parts&amp;rdquo;&lt;/em>).&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Needs to be written in C or C++ as that is the language of the Netdata agent (&lt;em>We are using &lt;a href="https://github.com/davisking/dlib">dlib&lt;/a> for the current implementation&lt;/em>).&lt;/li>
&lt;li>We Need to be very careful about taking big or complex dependencies if using third party libraries.
&lt;ul>
&lt;li>We want to be able to easily build and deploy Netdata on any Linux system without having to worry about installing or managing complex dependencies (&lt;em>we need to be careful of more complex algorithms that would have larger dependencies and potentially limit where Netdata can run&lt;/em>).&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ol>
&lt;p>The above considerations are important and useful to keep in mind as we explore the system in more detail.&lt;/p></description></item><item><title>Transform Monitoring With A Machine Learning Approach</title><link>https://www.netdata.cloud/blog/transform-monitoring-ml-first-approach/</link><pubDate>Thu, 11 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/transform-monitoring-ml-first-approach/</guid><description>&lt;p>Unlocking the full potential of monitoring through ML integration, anomaly detection, and innovative scoring engines.&lt;/p>
&lt;!--truncate-->
&lt;p>Machine Learning has been making waves in various industries, but its adoption in the monitoring and observability space has been slower than expected. Many “ML” features remain gimmicky and do not provide actual real world value to users that encourages their further use.&lt;/p>
&lt;p>At Netdata, we firmly believe that ML is crucial for monitoring, and we&amp;rsquo;ve taken an ML-first approach to provide users with powerful tools and insights. In this blog post, we&amp;rsquo;ll discuss the reasons behind our belief in ML, how we&amp;rsquo;ve integrated ML into our charts and visualizations, our query engines, the scoring engine we&amp;rsquo;ve built, and how these innovations enable metrics correlations and anomaly advisor.&lt;/p></description></item><item><title>Anomaly detection on Prometheus metrics</title><link>https://www.netdata.cloud/blog/anomaly-detection-on-prometheus-metrics/</link><pubDate>Wed, 01 Mar 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/anomaly-detection-on-prometheus-metrics/</guid><description>&lt;p>&lt;img src="../2023-03-01-anomaly-detection-on-prometheus-metrics/img/img.png" alt="img">&lt;/p>
&lt;p>We have recently extended the native machine learning (ML) based anomaly detection &lt;a href="https://learn.netdata.cloud/guides/monitor/anomaly-detection">capabilities&lt;/a> of Netdata to &lt;a href="https://github.com/netdata/netdata/issues/14218">support all metrics&lt;/a>, regardless on their collection frequency (&lt;code>update every&lt;/code>).&lt;/p>
&lt;p>Previously only metrics collected every second were supported, but now Netdata can run anomaly detection out of the box with zero config on metrics with any collection frequency.&lt;/p>
&lt;p>This post will illustrate an example of what this means using &lt;a href="https://prometheus.io/">Prometheus&lt;/a> metrics (via the &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/prometheus#gsc.tab=0">Netdata Prometheus collector&lt;/a>) since they typically have a default collection frequency of 10 seconds.&lt;/p></description></item><item><title>How Netdata’s Machine Learning works</title><link>https://www.netdata.cloud/blog/how-netdatas-machine-learning-works/</link><pubDate>Thu, 01 Sep 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/how-netdatas-machine-learning-works/</guid><description>&lt;p>Following on from the &lt;a href="https://www.netdata.cloud/blog/introducing-anomaly-advisor-unsupervised-anomaly-detection-in-netdata" target="_blank" rel="noopener">recent launch&lt;/a> of our &lt;a href="https://learn.netdata.cloud/docs/cloud/insights/anomaly-advisor" target="_blank" rel="noopener">Anomaly Advisor&lt;/a> feature, and in keeping with &lt;a href="https://www.netdata.cloud/blog/our-approach-to-machine-learning/" target="_blank" rel="noopener">our approach to machine learning&lt;/a>, &lt;a href="https://github.com/netdata/netdata/blob/master/ml/notebooks/netdata_anomaly_detection_deepdive.ipynb" target="_blank" rel="noopener">here&lt;/a> is a detailed Python notebook outlining exactly how the machine learning powering the Anomaly Advisor actually works under the hood.&lt;/p>
&lt;!--truncate-->
&lt;p>Or if you&amp;rsquo;d rather watch a video walkthrough of the notebook then check out below.&lt;/p>
&lt;iframe width="560" height="315" src="https://www.youtube.com/embed/L1xleckyuDQ?si=rptYzWE-eLlhSL9x" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen>&lt;/iframe>
&lt;p>Try it for yourself, &lt;a href="https://learn.netdata.cloud/docs/cloud/get-started" target="_blank" rel="noopener">get started&lt;/a> by &lt;a href="https://app.netdata.cloud/?utm_source=blog&amp;amp;utm_content=how_netdata_ml_works" target="_blank" rel="noopener">signing in to Netdata&lt;/a> and connecting a node. Once initial models have been trained (usually after the agent has about one hour of data, zero configuration needed), you&amp;rsquo;ll be able to start exploring in the &lt;a href="https://learn.netdata.cloud/docs/cloud/insights/anomaly-advisor" target="_blank" rel="noopener">Anomaly Advisor&lt;/a> tab of Netdata.&lt;/p></description></item><item><title>Anomaly rate in every chart</title><link>https://www.netdata.cloud/blog/anomaly-rate-in-every-chart/</link><pubDate>Thu, 23 Jun 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/anomaly-rate-in-every-chart/</guid><description>&lt;p>A month ago, we introduced unsupervised ML &amp;amp; Anomaly Detection in Netdata, the &lt;a href="https://www.netdata.cloud/blog/introducing-anomaly-advisor-unsupervised-anomaly-detection-in-netdata/">Anomaly Advisor&lt;/a>. Today, we’re happy to announce that we’re bringing anomaly rates to every chart in Netdata Cloud. Anomaly information is no longer limited to the Anomalies tab and will be accessible to you from the Overview and Single Node View tabs as well. This will make your troubleshooting journey easier, as you will have the anomaly rates for any metric available with a single click. Whichever metric or chart you&amp;rsquo;re exploring will be instant.&lt;/p></description></item><item><title>Metric Correlations on the Agent</title><link>https://www.netdata.cloud/blog/metric-correlations-on-the-agent/</link><pubDate>Wed, 15 Jun 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/metric-correlations-on-the-agent/</guid><description>&lt;p>As of &lt;a href="https://github.com/netdata/netdata/releases/tag/v1.35.0" target="_blank" rel="noopener">&lt;code>v1.35.0&lt;/code>&lt;/a> the Netdata Agent can now run &lt;a href="https://learn.netdata.cloud/docs/cloud/insights/metric-correlations" target="_blank" rel="noopener">Metric Correlations&lt;/a> (MC) itself. This means that, for nodes with MC enabled, the Metric Correlations feature just got a whole lot faster!&lt;/p>
&lt;!--truncate-->
&lt;p>The Netdata Metric Correlations feature uses a &lt;a href="https://en.wikipedia.org/wiki/Kolmogorov%E2%80%93Smirnov_test#Two-sample_Kolmogorov%E2%80%93Smirnov_test" target="_blank" rel="noopener">Two Sample Kolmogorov-Smirnov test&lt;/a> to look for which metrics have a significant distributional change around a highlighted window of interest. This can be useful when you are interested in short term &amp;ldquo;&lt;a href="https://en.wikipedia.org/wiki/Change_detection" target="_blank" rel="noopener">change detection&lt;/a>&amp;rdquo; and want to try answer the question &amp;ldquo;what else changed around this time?&amp;rdquo;.&lt;/p></description></item><item><title>Anomaly Advisor: Unsupervised Anomaly Detection</title><link>https://www.netdata.cloud/blog/introducing-anomaly-advisor-unsupervised-anomaly-detection-in-netdata/</link><pubDate>Thu, 26 May 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/introducing-anomaly-advisor-unsupervised-anomaly-detection-in-netdata/</guid><description>&lt;p>Today we are excited to launch one of our flagship ML assisted troubleshooting features in Netdata – the Anomaly Advisor.&lt;/p>
&lt;p>The Anomaly Advisor builds on earlier work to introduce unsupervised &lt;a href="https://github.com/netdata/netdata/blob/master/ml/README.md">anomaly detection&lt;/a> capabilities into the &lt;a href="https://www.netdata.cloud/agent/">Netdata Agent&lt;/a> from &lt;a href="https://github.com/netdata/netdata/releases/tag/v1.32.0">v1.32.0&lt;/a> onwards.&lt;/p>
&lt;h2 id="getting-started">Getting Started&lt;/h2>
&lt;p>Once you &lt;a href="https://learn.netdata.cloud/docs/configure/machine-learning#configuration">enable ML&lt;/a> on your nodes, each node will begin producing an &amp;ldquo;&lt;a href="https://learn.netdata.cloud/docs/configure/machine-learning#anomaly-bit---100--anomalous-0--normal">Anomaly Bit&lt;/a>&amp;rdquo; every second in addition to raw metric values. This anomaly bit will be 1 when the trained ML models consider recent raw data for a metric to look anomalous or 0 when things look &amp;rsquo;normal&amp;rsquo;. The Anomaly Advisor leverages this information to enable seamless space or room level anomaly detection out of the box with minimal configuration.&lt;/p></description></item></channel></rss>