The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / smartctl-disk-monitoring / smartctl-first-observation-baseline ▌

Operations Guides

First-observation baselining: don't page on lifetime counters at rollout

When you deploy SMART monitoring on an existing fleet for the first time, every in-service drive carries accumulated history in its lifetime counters. Offline Uncorrectable sectors, NVMe Media and Data Integrity Errors, Power-On Hours, Unsafe Shutdowns, Reallocated Sector Count, UDMA CRC Error Count. These counters started incrementing the moment the drive left the factory and never reset.

If your alerting fires on any non-zero absolute value, every drive with any history pages within minutes of enabling monitoring. A drive with 3 reallocated sectors from factory QA, a drive with 200 uncorrectable errors from a past thermal event three years ago, a drive with 5 unsafe shutdowns from a UPS failure. All produce identical alerts to a naive greater-than-zero rule.

The fix is straightforward: capture a baseline at first observation and alert on growth from that baseline, not on the absolute value. This is the difference between “something happened at some point in this drive’s life” and “something happened since you started watching.”

Why absolute values are useless at rollout

Cumulative SMART counters are monotonically increasing by design. They tell you that an event occurred at some point in the drive’s lifetime, not when. A drive that accumulated 50 Offline Uncorrectable sectors during shipping damage in 2021 and has been stable since is in a fundamentally different state than one that gained 50 sectors last week. Without a baseline, both drives produce the same alert.

New drives are not exempt. Factory burn-in routinely leaves non-zero counts on several attributes:

  • Power-On Hours: brand-new drives may show 1-2 hours from factory testing
  • Total Data Written: factory testing, firmware installation, and QA write data to the drive
  • Reallocated Sector Count: some SSDs ship with a handful of remapped blocks from manufacturing QA
  • NVMe Unsafe Shutdowns: the drive went through testing and burn-in at the factory. A single-digit count is expected
  • UDMA CRC Error Count: may show a small count from initial cable seating during integration

None of these indicate problems in your environment. Without a deployment-time snapshot, you cannot distinguish “the factory did this” from “production did this.”

How baselining works

The mechanism has three stages: capture, compare, and alert on delta.

flowchart TD
    A["Monitoring sees drive"] --> B{"Baseline recorded?"}
    B -->|"No"| C["Record current values as baseline"]
    C --> D["No alert this cycle"]
    B -->|"Yes"| E["Compare current vs baseline"]
    E --> F{"Counter changed?"}
    F -->|"No change"| G["Normal operation"]
    F -->|"Increased"| H["Alert: new event since baseline"]
    F -->|"Decreased"| I["Anomalous: drive swap or reset"]

Capture the baseline

At first scrape, record the full SMART attribute set for every drive. This snapshot is both the reference point for growth alerts and a forensic record of the drive’s state at deployment.

# Capture full SMART snapshot at deployment for archival
mkdir -p /var/lib/smart-baselines
smartctl -x -j /dev/sdX > /var/lib/smart-baselines/sdX-$(date +%Y%m%d).json

The -x flag prints all SMART and non-SMART information. The -j flag produces structured JSON output (available since smartmontools 7.0) for programmatic parsing. For fleet-wide capture, run this against every device before enabling alerting. Store the output somewhere that survives host rebuilds and monitoring system reinstallation.

Distinguish two moments:

  1. First-observation baseline: what the monitoring system records when it first scrapes a drive. If monitoring comes online after drives have been in production, this baseline includes production history.
  2. Deployment-time snapshot: what the operator captures when physically installing a drive. This is the true zero-time reference, useful for later triage.

When monitoring is deployed onto an existing fleet, the first-observation baseline is your starting point. It absorbs all prior history into a single reference. From that moment forward, growth is the signal.

Compare against baseline

Each subsequent scrape compares the current value to the baseline value:

  • Current equals baseline: no change, no alert
  • Current greater than baseline: growth detected. This is the signal that matters
  • Current less than baseline: anomalous. Counters are cumulative and should never decrease. Investigate drive swap, firmware update, or counter reset

For rate-of-change analysis, track the timestamp of each observation alongside the value. The growth rate (sectors per day, media errors per week) is more diagnostic than the absolute delta. A drive that gains 1 reallocated sector per month is in a different category than one that gains 10 per day.

Alert on growth, not absolute value

The alerting logic transforms from “is this counter above zero?” to “did this counter increase since the last observation?” This single change eliminates the rollout flood while preserving sensitivity to new events. A drive that had 17 Offline Uncorrectable sectors at baseline and now has 18 triggers an alert. A drive that had 17 at baseline and still has 17 does not.

The recommendation for severity: page only when the count has increased since baseline. Historical non-zero values get a ticket (investigate, but do not assume it just happened).

Which counters need baselining

Not all SMART signals need baseline treatment. Some are threshold-based or represent current state rather than accumulated history. The counters that require first-observation baselining share one property: they are monotonically increasing lifetime accumulators.

| Signal | Source | Why it accumulates | |—|—| | Offline Uncorrectable (ID 198) | ATA SMART | Confirmed sector-level data loss. Never resets, even if the sector is later reallocated | | NVMe Media and Data Integrity Errors | NVMe Health Log | Cumulative integrity failures. Monotonically increasing | | Reallocated Sector Count (ID 5) | ATA SMART | Consumed spare pool entries. Some SSDs ship non-zero from factory | | UDMA CRC Error Count (ID 199) | ATA SMART | Interface CRC errors. Cumulative, never resets. Drive carries history across hosts | | Power-On Hours (ID 9) | ATA SMART / NVMe | Drive age. Includes factory testing time | | Unsafe Shutdowns (NVMe) | NVMe Health Log | Unexpected power loss count. Includes factory burn-in events | | Total LBAs Written / Data Units Written | ID 241 / NVMe | Cumulative host write volume. Includes factory QA writes | | Power Cycle Count (ID 12) | ATA SMART | Power-on cycles. Includes factory testing | | ATA Error Log entry count | ATA SMART | Error counter is cumulative even though the log buffer is circular |

Signals that represent current state do not need this treatment:

SignalWhy no baseline needed
Current Pending Sector (ID 197)Can decrease (sectors resolve on rewrite). Alert on sustained non-zero, not growth from baseline
Drive TemperatureInstantaneous reading, not cumulative
NVMe Available SpareCurrent percentage. Alert on threshold crossing
NVMe Percentage UsedCurrent estimate, not a counter. Alert on threshold
NVMe Critical Warning bitsCurrent bitmask state, not cumulative

Implementation in common tools

smartd configuration

For operators using smartd, directive syntax controls whether alerts fire on absolute value or growth:

  • -U 198 reports whenever attribute 198 is non-zero on a check cycle
  • -U 198+ reports only when the value has increased since the last check cycle

By default, the smartd -a directive sets -U 198; drive-database presets using -v 198,increasing instead select -U 198+ unless you specify another -U directive. Verify the effective preset with smartctl -P show /dev/sdX.

The + suffix is the built-in mechanism for growth-based alerting. For an existing fleet with historical values, use the + variant to avoid re-alerting on pre-existing counts.

smartd can maintain state across daemon restarts using state files. A state file stores previous attribute values, temperature min/max, and error log entry counts. Upstream builds default to no default state-file prefix; packagers can enable one with --with-savestates. Verify your build with smartd --help | grep -i savestates and inspect its packaged service configuration. Without state persistence, smartd loses track of previous values on restart and may re-alert on already-known values.

Prometheus and smartctl_exporter

The Prometheus community smartctl_exporter exposes raw SMART counter values as metrics. It does not perform baselining or delta calculation. You must implement growth-based alerting in PromQL.

A naive rule like smartctl_device_smart_attributes_raw_value > 0 fires on every drive with any history at first scrape. Use a time-windowed comparison instead:

# Alert when Offline Uncorrectable increased in the last hour
delta(smartctl_device_smart_attributes_raw_value{attribute_id="198"}[1h]) > 0

This naturally implements baselining: the delta over the time window only fires on growth within that window, not on historical accumulated values. At first scrape, there is no previous data point in the window, so no alert fires. The window length defines your detection latency: a 1-hour window catches growth within the last hour.

For slower-moving counters, use a longer window or combine with changes() to detect any modification. The principle is the same: alert on the rate of change, never on the absolute value.

Deployment-time snapshots

The baseline serves double duty. Beyond alerting, it is a forensic record. When a drive develops 3 reallocated sectors, the deployment snapshot confirms whether those sectors were present at install time or appeared in production. Without it, you cannot distinguish factory reallocations from production ones.

For new drive installations, capture a full SMART snapshot at physical install time. This gives you a true zero-time reference even if the monitoring system was not running yet.

Common mistakes

  • Alerting on absolute non-zero at first deployment: Every drive with any history pages. This floods operators at the exact moment they are validating the monitoring system, training them to ignore SMART alerts before they have seen a real one.
  • Losing the baseline on monitoring system rebuild: If the monitoring system is rebuilt or state files are lost, the new “first observation” becomes the baseline. Growth that occurred between the old baseline and the rebuild is silently absorbed. Store baselines in persistent storage that survives host rebuilds.
  • Assuming new drives have zero counters: Factory testing writes data, spins up drives, and may cause unsafe shutdowns during burn-in. Brand-new drives routinely show small non-zero values for Power-On Hours, Total LBAs Written, Reallocated Sector Count, and Unsafe Shutdowns.
  • Ignoring counter decreases: A cumulative counter that decreases is anomalous. It may indicate a drive swap (new drive with lower counters), a firmware update that reset counters, or manufacturer diagnostics that forced sector reallocation.
  • Treating first-observation baseline as deployment-time truth: If monitoring was deployed after drives entered production, the first observation already includes production history. Growth from that baseline is valid, but the baseline itself is not a clean install-time record. Capture deployment-time snapshots separately for forensic purposes.
  • Using the same alert logic for all counters: Current Pending Sector (ID 197) can decrease when sectors resolve on rewrite. Alerting on growth from baseline for ID 197 is less useful than alerting on sustained non-zero across multiple polls. Match the alerting strategy to the counter’s behavior.

Correlating SMART growth with Netdata

SMART counter growth rarely tells the full story in isolation. Netdata’s value here is correlation, not just collection.

  • Per-second collection means growth is detected on the next collection cycle, not on the next 5-minute scrape. For a counter like Offline Uncorrectable that may jump in bursts, tighter polling reduces the window between event and detection.
  • Correlated timelines let you confirm whether SMART counter growth represents active failure. View Offline Uncorrectable growth alongside I/O latency (await, svctm, iowait), kernel I/O errors from dmesg, and drive temperature in a single view. Growth that coincides with await spikes and UNC errors in the kernel log is a confirmed active failure. Growth without corroborating signals may warrant investigation but not an immediate page.
  • ML-based anomaly detection flags unexpected changes in SMART counter values without explicit threshold rules. The anomaly model learns each drive’s normal behavior rather than checking against a fixed threshold, so drives with non-zero baselines from factory or pre-monitoring history do not generate false positives.
  • Historical retention lets you scroll back to when monitoring started and verify whether specific sectors or errors were present at first observation or appeared later. This provides the deployment-time snapshot automatically, as long as monitoring was active from install.