The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / smartctl-disk-monitoring / smartctl-health-passed-but-failing ▌

Operations Guides

SMART says PASSED but the drive is failing: why the health check lies

The PASSED verdict from smartctl -H tells you one narrow thing: no pre-fail SMART attribute has crossed its vendor-defined threshold. It does not mean the drive is healthy, that data is safe, or that the drive will survive the week.

A drive can report PASSED while hundreds of sectors are pending reallocation, dozens have already been remapped, I/O latency is spiking to seconds, and the spare pool is burning down. Google’s 2007 study of over 100,000 drives found that 36% of failed drives had zero prior SMART warnings. Backblaze’s fleet analysis showed that 23.3% of failed drives had no non-zero values across the five attributes they consider most predictive (IDs 5, 187, 188, 197, 198). The health check is a last-resort binary, not a health indicator.

What the PASSED verdict actually checks

The SMART Overall Health Self-Assessment is a binary pass/fail judgment from the drive firmware. On ATA/SATA drives, smartctl -H issues the SMART RETURN STATUS command. The drive compares its internal attribute values against vendor-defined thresholds and reports whether any pre-fail attribute has crossed its threshold.

SMART attributes carry a TYPE column in smartctl -A output: Pre-fail or Old_age. Only Pre-fail attributes participate in the health verdict. Old_age attributes (like Power-On Hours) never trigger FAILED regardless of their value. Even among Pre-fail attributes, the verdict only fires when the normalized VALUE drops to or below the THRESH value. Vendors set these thresholds so that FAILED triggers only at an advanced state of degradation. A drive can accumulate significant media damage and still have every normalized value sitting comfortably above its threshold.

This matters because most monitoring setups, naive scripts, and even some commercial tools treat PASSED as proof of health. They check the exit code or the verdict string and move on. By the time SMART says FAILED, the drive has been dying for weeks or months, and you are at the point of maximum damage and minimum recovery time.

How it works

The threshold comparison

For ATA drives, each SMART attribute has three key fields in the smartctl -A output:

  • VALUE: The normalized current value, on a vendor-defined scale (typically 1-253, where higher is better).
  • WORST: The lowest normalized value the attribute has ever reached.
  • THRESH: The vendor-defined threshold below which the attribute is considered failing.

The health check logic is straightforward: for each Pre-fail attribute, if VALUE is at or below THRESH, the drive reports FAILED. Vendors set THRESH low. A Reallocated_Sector_Ct (ID 5) attribute might have a threshold so low that the raw reallocated count can reach into the hundreds before the normalized value drops far enough to trip the verdict.

The exit code bitmask

smartctl encodes diagnostic information in its exit status as a bitmask. The relevant bits for health assessment:

  • Bit 3: SMART RETURN STATUS check returns “DISK FAILING.”
  • Bit 4: A prefail attribute is at or below its threshold.
  • Bit 5: SMART status check returned “DISK OK” but some usage or pre-fail attributes were at or below threshold in the past.
  • Bit 6: The device error log contains errors.
  • Bit 7: The device self-test log contains errors.

A clean PASSED with exit code 0 means bits 3 and 4 are not set: the drive is not reporting “DISK FAILING” and no prefail attribute is currently at or below threshold. This says nothing about attributes trending downward, pending sectors, error logs, or any signal that has not yet crossed the vendor’s line.

A reported case illustrates the gap: smartctl -H returned exit code 0 on a failing drive, while smartctl -a returned exit code 192 (bits 6 and 7 set: error log and self-test log both contain errors). A monitoring script checking only -H saw nothing.

Note: With smartctl -H, smartmontools 7.3 and later no longer set bit 2 solely because ATA attributes are available. Bit 2 still reports a failed SMART/ATA command or a checksum error. Do not treat bit 2 as a catch-all “something is wrong” signal.

NVMe: a different mechanism

NVMe drives do not use ATA SMART attribute thresholds. The smartctl -H health assessment for NVMe reads the Critical Warning byte from the SMART/Health Information Log (Log Page 02h). Each bit represents a specific condition:

  • Bit 0: Available spare below threshold
  • Bit 1: Temperature is above or below its threshold
  • Bit 2: NVM subsystem reliability degraded
  • Bit 3: Media placed in read-only mode
  • Bit 4: Volatile memory backup device failed
  • Bit 5: Persistent memory region is read-only or unreliable

If no Critical Warning bit is set, the drive reports PASSED. This catches some conditions the ATA check misses, but it still does not catch gradual wear. An NVMe drive can report 188% Percentage Used with the Critical Warning byte at 0x00 and still return PASSED. The wear is extreme, the drive is far beyond its rated endurance, but no threshold has been crossed because Available Spare has not yet dropped below the vendor threshold and no media errors have occurred.

flowchart TD
    A["smartctl -H /dev/sdX"] --> B{"ATA or NVMe?"}
    B -->|"ATA/SATA"| C["SMART RETURN STATUS\ncommand"]
    B -->|"NVMe"| D["Read Critical Warning\nbyte from Log Page 02h"]
    C --> E{"Any Pre-fail attribute\nVALUE at or below THRESH?"}
    E -->|"No"| F["PASSED"]
    E -->|"Yes"| G["FAILED"]
    D --> H{"Any Critical\nWarning bit set?"}
    H -->|"No"| F
    H -->|"Yes"| G
    F --> I["Drive can still have:\nhundreds of pending sectors\ndozens of reallocated sectors\n188% Percentage Used\nrising error log entries"]
    G --> J["Drive firmware declares\nfailure. Evacuate immediately."]

Where it shows up in production

The monitoring script that only checks exit code 0

The most common pattern: a cron job or monitoring check runs smartctl -H /dev/sdX, greps for “PASSED,” and alerts only on “FAILED.” This setup is blind to everything except the vendor’s final verdict. Drives with rapidly growing reallocated sectors, climbing pending counts, and degrading performance sail through with a clean bill of health. When the drive finally fails, the monitoring system had nothing to say for weeks.

USB enclosures and the attribute check fallback

USB-to-SATA bridges often do not pass the ATA SMART RETURN STATUS command through to the drive. When this happens, smartctl -H falls back to an attribute-based check and prints a warning: “This result is based on an Attribute check.” The result may still say PASSED, but it is a best-effort guess, not the drive’s own firmware verdict. Drives behind USB bridges are effectively unmonitored for health unless you use the correct -d sat or bridge-specific flag and verify the data path works.

NVMe drives past rated endurance

NVMe Percentage Used is explicitly allowed to exceed 100% per the NVMe specification. Real-world reports show drives at 123%, 188%, and higher, all reporting PASSED because the Critical Warning byte remains 0x00. The Available Spare threshold is vendor-specific. As long as the spare pool has not dropped below that threshold, the drive considers itself fine. It is operating well beyond warranty, NAND retention is degrading, and the next bad block could exhaust spares, but the health check says PASSED.

Hardware RAID controllers masking everything

Behind hardware RAID controllers (MegaRAID, Smart Array, Adaptec), standard SMART queries fail or return virtual device data. If your monitoring runs smartctl -H /dev/sda where /dev/sda is a RAID virtual disk, you are not checking physical drive health at all. You must use controller-specific passthrough:

smartctl -a -d megaraid,0 /dev/sda

Without the correct -d type, every physical drive could be failing and the health check would still say PASSED on the virtual device.

Counter resets and tampered drives

SMART counters are cumulative and should never decrease. A drive that shows Reallocated_Sector_Ct dropping to zero, or Power-On Hours decreasing, has either been swapped, had a firmware update that reset counters, or been tampered with. In 2025, used Seagate enterprise drives with 15,000-50,000 hours were sold as new with SMART values reset to near zero. The fraud was initially caught because Seagate drives keep a second log (FARM) that scammers did not modify. Later variants compromised FARM values too. If you are monitoring secondhand drives, a PASSED verdict with suspiciously clean counters is a red flag, not reassurance.

Common misuses

Treating PASSED as proof of health. PASSED means “no pre-fail attribute has crossed the vendor threshold yet.” A drive can be actively degrading for weeks before the verdict changes.

Alerting only on FAILED. By the time the drive firmware says FAILED, you have lost recovery time. Individual attributes provide warnings days to weeks earlier.

Assuming exit code 0 means no problems. The exit code bitmask includes bits for error log entries (bit 6) and self-test log errors (bit 7). A drive can return PASSED with a non-zero exit code from those bits while the health verdict string still says PASSED. Scripts that check only for the string “PASSED” miss this.

Using the same thresholds for all drive types. A 60C temperature reading is critical for an HDD but within spec for an enterprise NVMe. Seagate raw values for IDs 1 and 7 can be in the billions and be normal. Alert thresholds must be drive-type-aware and vendor-aware.

Ignoring rate of change. A drive with 5 reallocated sectors accumulated over 5 years is stable. A drive with 5 reallocated sectors gained this week is actively dying. Most monitoring checks absolute values only, missing the acceleration signal that distinguishes historical damage from active failure.

Signals to watch in production

SignalWhy it mattersWarning sign
Reallocated_Sector_Ct (ID 5)Drive is consuming its finite spare pool. Growth rate is the key metric.Any increase from baseline, especially if accelerating.
Current_Pending_Sector (ID 197)Sectors the drive cannot reliably read. Each one can cause multi-second I/O latency spikes.Any non-zero value. Sustained non-zero across polls indicates active deterioration.
Offline_Uncorrectable (ID 198)Permanent data loss at the media level. Data in these sectors is gone.Any increase from baseline. Verify the attribute name, not just the ID: some WD models use ID 198 for a different purpose.
UDMA_CRC_Error_Count (ID 199)Physical layer problem between drive and host. Usually a cable or backplane issue, not the drive.Any increase. Commonly misdiagnosed as drive failure, leading to unnecessary replacement.
NVMe Available SpareRemaining capacity to handle bad blocks. When it hits 0%, the next bad block means permanent data loss.Below vendor threshold. Any declining trend.
NVMe Percentage UsedEndurance consumption estimate. Can exceed 100% per spec.Above 90%: plan replacement. Above 100%: operating beyond rated life.
NVMe Critical Warning bitsDrive firmware flagging specific critical conditions via a bitmask.Any non-zero value. Each bit maps to a distinct condition (spare, temperature, reliability, read-only, PLP).
NVMe Media and Data Integrity ErrorsUncorrectable data errors from the NAND media. The NVMe equivalent of Offline_Uncorrectable.Any increase from baseline.
ATA Error Log entriesHistory of I/O errors with type, LBA, and timing relative to power cycle.Any non-zero count, especially UNC (uncorrectable) errors. The summary log holds only 5 entries; old errors are overwritten.
Kernel I/O errors (dmesg)Host-level perspective on drive health. Catches failure modes SMART cannot self-report.“I/O error”, “medium error”, “timeout”, or “reset” messages for the specific device.
Self-test resultsActive surface scanning that finds latent bad sectors before production I/O hits them.Any result other than “Completed without error”. Requires scheduled tests; drives do not run them automatically.

How Netdata helps

Netdata collects the individual signals that predict failure, rather than relying on the PASSED/FAILED verdict as the primary health indicator.

  • Per-second collection of SMART attributes means rate-of-change tracking is continuous. You see when Reallocated_Sector_Ct or Current_Pending_Sector starts climbing in real time, alongside whatever workload triggered the degradation.
  • Correlation between SMART attributes and I/O latency surfaces the zombie drive pattern: pending sectors causing await spikes in disk metrics while the health check still reports PASSED.
  • NVMe health log collection including Available Spare, Percentage Used, Critical Warning bits, and Media and Data Integrity Errors provides the full endurance and reliability picture. A drive at 150% Percentage Used with declining Available Spare gets attention before the Critical Warning byte flips.
  • Anomaly detection on SMART attribute trends flags unusual changes that static thresholds miss, such as slow reallocated sector accumulation that has not yet crossed any configured alert line.
  • Error log and self-test monitoring catches UNC errors and self-test failures that the health check ignores.
  • Kernel I/O error correlation provides the host-side perspective that complements drive-side SMART data, catching firmware bugs and controller issues that SMART cannot self-report.