The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / smartctl-disk-monitoring / smartctl-nvme-volatile-memory-backup-failed ▌

Operations Guides

NVMe volatile memory backup device failed: CriticalWarning bit 4 and lost power-loss protection

You run smartctl -H /dev/nvme0n1 and get FAILED. The detail line reads - volatile memory backup device has failed. The Critical Warning byte shows 0x10. The drive is still serving reads and writes at full speed, latency is normal, Available Spare is fine, and Percentage Used is well within spec.

This is NVMe Critical Warning bit 4. It does not mean the NAND is failing or the controller is dying. It means the drive’s power-loss protection (PLP) hardware, typically supercapacitors on enterprise NVMe SSDs, has failed. Under stable power, the drive operates normally. But if power drops unexpectedly, data sitting in the drive’s volatile write buffer will be lost because the capacitor bank can no longer flush it to NAND.

The drive continues to serve I/O until the next power event, at which point the PLP safety net that justified deploying enterprise drives is already gone. This article covers what bit 4 means, how to verify it is not a false positive, how to correlate it with unsafe shutdown count, and what remediation options exist.

What this means

The NVMe SMART/Health Information Log (Log Page 02h) contains a 1-byte field called Critical Warning. Each bit maps to a specific condition the drive firmware considers critical:

BitMaskMeaning
00x01Available spare below threshold
10x02Temperature threshold exceeded
20x04NVM subsystem reliability degraded
30x08Media in read-only mode
40x10Volatile memory backup device failed

When bit 4 is set, smartctl reports the overall health as FAILED and appends the specific condition text. Any non-zero Critical Warning byte triggers this FAILED declaration, regardless of which bit is set. This is expected smartctl behavior: it prints FAILED whenever any Critical Warning bit is present, whether or not that bit indicates an immediate media failure.

The NVMe specification defines this field as valid only if the controller has a volatile memory backup solution. On enterprise NVMe drives with PLP, this means the supercapacitor bank that provides power to flush the write buffer during an unexpected power loss has failed self-test. On drives without PLP hardware, bit 4 should not be set, but firmware bugs or spec misinterpretation on some consumer drives can produce false positives.

The operational impact is specific: the drive cannot guarantee data integrity across a power-loss event. The failure is latent, waiting for a power event to manifest as data loss.

Common causes

CauseWhat it looks likeFirst thing to check
Supercapacitor aging or failureEnterprise drive, bit 4 set, drive model known to ship with PLP hardwareConfirm drive model has PLP, check warranty status
Consumer drive firmware bugConsumer NVMe drive with no PLP hardware, bit 4 set, possibly since firmware update or first bootVerify the drive model lacks PLP hardware; check vendor forums for known firmware issues
Physical or electrical damageBit 4 set after a server power event, thermal incident, or physical impactCheck dmesg for related events around the time bit 4 appeared
Firmware update side effectBit 4 appeared immediately after a firmware updateCompare firmware version before and after; check vendor advisory

Quick checks

All commands below are read-only and safe to run on production drives.

# Check the Critical Warning byte value
smartctl -A /dev/nvme0n1 | grep "Critical Warning"

# Get the full health assessment text
smartctl -H /dev/nvme0n1

# Get the full SMART/Health output for context
smartctl -a /dev/nvme0n1

# Check Unsafe Shutdowns count
smartctl -A /dev/nvme0n1 | grep "Unsafe Shutdowns"

# Identify the drive model and firmware
smartctl -i /dev/nvme0n1

# Check for kernel-level events around the time the bit appeared
dmesg | grep -iE "nvme|power|shutdown|reset" | tail -30

# Confirm Available Spare is not also degraded (broader failure)
smartctl -A /dev/nvme0n1 | grep "Available Spare"

# Check Media and Data Integrity Errors to rule out NAND failure
smartctl -A /dev/nvme0n1 | grep "Media and Data Integrity Errors"

How to diagnose it

Step 1: Confirm the bit and decode the byte

Read the Critical Warning value. If it shows 0x10, only bit 4 is set. If the value is different (for example 0x14), multiple bits are active and you have a compound problem. Decode the full byte before acting on bit 4 in isolation.

Step 2: Verify the drive actually has PLP hardware

This is the most important diagnostic step. The NVMe spec says bit 4 is only valid if the controller has a volatile memory backup solution. If you are running a consumer NVMe SSD with no PLP capacitors, bit 4 may be a firmware false positive.

Check the drive model against the vendor datasheet. Enterprise NVMe drives commonly ship with PLP capacitors. If the datasheet does not mention power-loss protection or capacitors, the drive likely has no PLP hardware and bit 4 should be treated as a model/firmware anomaly. Do not rely on community model lists; verify the exact model and capacity variant.

Step 3: Correlate with Unsafe Shutdowns

# Check current Unsafe Shutdowns count
smartctl -A /dev/nvme0n1 | grep "Unsafe Shutdowns"

The NVMe 2.1 field is named Unexpected Power Losses and was previously named Unsafe Shutdowns; it increments when main power is lost while the controller has not reported readiness for power-off. If PLP has failed, every future unexpected power loss is a potential data-in-flight loss event. How firmware recovery affects NAND wear is controller-specific; judge the risk from the rising count and workload, not from a universal FTL-rebuild assumption.

Step 4: Rule out firmware bugs

Check whether the bit appeared after a firmware update:

# Current firmware version
smartctl -i /dev/nvme0n1 | grep -i firmware

Compare against known-good firmware versions. Check vendor advisories and community forums for reports of bit 4 false positives on this specific model and firmware combination.

Step 5: Assess the full SMART picture

Confirm that the rest of the drive’s health indicators are normal. If Available Spare is also below threshold (bit 0 set) or Media and Data Integrity Errors are non-zero and increasing, the drive may have broader age-related or hardware failures beyond the PLP capacitor.

flowchart TD
    A["Critical Warning = 0x10"] --> B{"Drive has PLP hardware?"}
    B -->|Yes| C["PLP capacitor failure
Safety net is gone"] B -->|No| D["Likely firmware false positive
Suppress or update firmware"] C --> E{"Unsafe Shutdowns
count rising?"} E -->|Yes| F["High risk: data loss
has likely occurred"] E -->|No| G["Latent risk: data loss
on next power event"] F --> H["Replace drive
or add UPS redundancy"] G --> H D --> I["Use smartd.conf to suppress
if supported by your version"]

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Critical Warning bytePrimary indicator of bit 4 statusAny value where bit 4 is set (0x10 in the byte)
Unsafe ShutdownsEach power event without PLP risks data lossCount increasing while bit 4 is set
Available SpareConfirms bit 4 is isolated, not part of broader failureAlso below threshold, meaning compound failure with bit 0
Percentage UsedContext for whether the drive is near end of lifeAbove 100% alongside bit 4 suggests multiple age-related failures
Media and Data Integrity ErrorsRules out NAND-level data corruptionNon-zero and increasing
Composite TemperatureCapacitor degradation can correlate with thermal stressSustained high temperature or past overheating events

Fixes

Replace the drive (enterprise with PLP requirement)

If the drive was deployed specifically because PLP is a data integrity requirement (database write cache, journal device, write-heavy workload without reliable UPS redundancy), the only correct fix is replacement. The drive cannot be repaired in the field. The capacitor bank is not a serviceable component.

Tradeoff: replacement requires a maintenance window, data migration, and potential RAID rebuild. But continuing to operate a PLP drive without PLP defeats the purpose of deploying enterprise drives.

Suppress the alert for consumer drives without PLP

If the drive does not actually have PLP hardware and bit 4 is a firmware false positive, you can configure smartd to ignore the bit.

smartmontools 7.5 and later support -H MASK for NVMe. The hexadecimal mask lists the Critical Warning bits smartd checks. To ignore bit 4 while checking the other warning bits, clear bit 4 from the mask:

# smartd.conf: check all warning bits except bit 4 (0x10)
/dev/nvme0n1 -H 0xef -l error

0xff is equivalent to plain -H; 0xef ignores bit 4 only. This feature is experimental upstream, so validate it after upgrading.

If bitmask masking is not supported in your version, the alternative is -d ignore in smartd.conf, which disables all SMART monitoring for the device. This is strongly discouraged because it creates a complete monitoring blind spot for every other failure mode. A better approach may be to exclude the specific drive from smartd and monitor it with a custom script that checks the Critical Warning byte and masks bit 4 before alerting.

Tradeoff: suppressing the bit is appropriate only for drives that genuinely lack PLP hardware. Suppressing it on an enterprise drive with actual PLP hides a real hardware failure.

If the server has reliable UPS backup with generator failover, the probability of an unsafe shutdown may be low enough that some teams accept the risk temporarily while awaiting a replacement. This is a business decision, not a technical fix. Document the acceptance explicitly, set a hard deadline for replacement, and track the Unsafe Shutdowns count closely during the interim.

Prevention

  • Track PLP health across the fleet: Monitor Critical Warning bit 4 on all enterprise NVMe drives. A PLP failure is silent until a power event exposes it.
  • Monitor Unsafe Shutdowns trend: A rising count on any drive indicates power infrastructure issues. On a drive with failed PLP, each increment is a potential data loss event.
  • Maintain UPS redundancy: PLP is defense in depth, not the primary protection. Reliable power infrastructure is the first line of defense against data loss from unexpected shutdowns.
  • Track firmware versions: Some bit 4 false positives are triggered by specific firmware versions on consumer drives. Maintain a firmware inventory and watch vendor advisories.
  • Baseline at deployment: Capture the Critical Warning byte value and Unsafe Shutdowns count when a drive is first deployed. Alert on change from baseline, not on absolute values, to avoid false alarms from historical counts.

How Netdata helps

  • Per-second metric collection: Netdata collects NVMe SMART attributes including the Critical Warning byte at per-second granularity, so you see the moment bit 4 transitions from 0 to 1 without waiting for the next polling cycle.
  • Correlation with Unsafe Shutdowns: When bit 4 is set, the Unsafe Shutdowns counter becomes the critical correlated signal. Netdata’s dashboard lets you overlay both metrics on the same timeline to assess whether power events have already occurred without PLP protection active.
  • Anomaly detection: Netdata’s ML-based anomaly detection can flag the transition of Critical Warning from 0x00 to 0x10 as anomalous even without a configured static threshold.
  • Fleet-level visibility: Across multiple servers, Netdata aggregates SMART data so you can see whether bit 4 is isolated to one drive or affecting multiple drives of the same model, which may indicate a batch defect or firmware issue.