The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nvme / nvme-unsafe-shutdowns-increasing ▌

Operations Guides

NVMe unsafe shutdowns increasing: power-loss events and silent corruption risk

The unsafe_shutdowns field in the NVMe SMART log is a lifetime counter. It increments every time the drive loses power without first receiving a shutdown notification (CC.SHN) from the host. A clean reboot increments power_cycles but not unsafe_shutdowns. A power cut, kernel panic, or someone holding the power button increments both.

The counter never goes down, so the absolute number is history. What matters is the rate of change: every new increment means something cut power to the drive unexpectedly. On enterprise drives with power-loss protection (PLP) capacitors, in-flight writes in DRAM get flushed to NAND before the power dies, and the event is a footnote. On consumer drives without PLP, every increment is a data-loss roll of the dice: writes the host believes were completed may never have reached NAND.

The damage is usually silent. The server comes back up, the journal replays, the application starts, and nobody connects the corrupted database page or the checksum mismatch three weeks later to the power event that caused it. Teams ignore this counter precisely because “the server came back up fine.”

What this means

On a clean shutdown, the NVMe driver sends a shutdown notification by setting CC.SHN in the controller configuration register. The controller then has a defined window to flush volatile caches and quiesce internal state before power is removed. When power disappears without that notification, the controller records an unsafe shutdown and increments the counter on next power-up.

The consequences depend on what was in flight and whether the drive has PLP:

  • Drive with PLP: Capacitors hold enough charge for the controller to flush its DRAM write cache to NAND. The counter increments, but acknowledged writes are durable.
  • Drive without PLP: Data sitting in the DRAM write cache is gone. The host already got the completion acknowledgment. The filesystem journal replays on mount, but the journal can only recover what actually reached the media. Anything acknowledged-but-never-persisted is silently lost.
  • Either way: The increment is evidence of a power-infrastructure problem: PSU fault, UPS failure, PDU flapping, kernel panic, hard reset, or a watchdog-triggered power cut. The drive is the witness, not the suspect.

There is also a growing class of false positives where the counter increments without any real power loss, usually tied to firmware quirks or aggressive platform power management. Distinguishing real events from firmware noise is the first diagnostic fork.

flowchart TD
  A[unsafe_shutdowns increment] --> B{Did power actually get cut?}
  B -->|dmesg panic, watchdog, outage| C[Real power-loss event]
  B -->|clean shutdown, suspend only| D[False positive: firmware or power-state quirk]
  C --> E{Does the drive have PLP?}
  E -->|yes, bit 4 clear| F[Writes protected: investigate power infrastructure]
  E -->|no, or bit 4 set| G[Silent corruption risk: verify filesystem, app integrity, backups]
  D --> H[Check kernel version, firmware, suspend behavior]

Common causes

CauseWhat it looks likeFirst thing to check
PSU, UPS, or PDU failureCounter increments coincide with unexpected host reboots; may affect multiple machines on the same feedHost power/event logs, IPMI SEL, PDU logs
Kernel panic or hard hangIncrement lines up with a crash; host rebooted without clean shutdownjournalctl --list-boots, kdump/pstore remnants, dmesg from previous boot
Human action: power button held, PDU outlet cycledSingle increment, no crash evidence, often during maintenance windowsOut-of-band console logs, ticketing/maintenance records
Firmware false positive on suspend or clean shutdownCounter increments on every s2idle sleep or every clean shutdown, with no crash evidenceCompare power_cycles vs unsafe_shutdowns growth; check kernel and firmware versions
PLP capacitor failed (bit 4)critical_warning bit 4 set; each subsequent unsafe shutdown now carries full data-loss risknvme smart-log critical_warning field
Hypervisor or runtime not issuing clean NVMe shutdownIncrements tied to VM stop/start or host-managed power eventsWhether the platform issues shutdown notification on stop

Controller resets are a different thing. A reset is the host driver recovering an unresponsive controller in software: power was never cut, unsafe_shutdowns does not increment, and you will see a ... timeout, reset controller line in the kernel log instead. If you are seeing resets rather than unsafe shutdowns, that is a firmware or PCIe problem, not a power problem. See NVMe controller reset timeout and NVMe controller reset loop.

Quick checks

All of these are read-only and safe on a production host.

# Read the SMART log: unsafe shutdowns, power cycles, critical warning
nvme smart-log /dev/nvme0 | grep -Ei "unsafe_shutdowns|power_cycles|critical_warning|media_errors"

# Same data via smartmontools if nvme-cli is not installed
smartctl -a /dev/nvme0 | grep -i "unsafe"

# Critical warning bitmask: bit 4 (0x10) means PLP backup has failed
nvme smart-log /dev/nvme0 | grep critical_warning
# Boot history: correlate increments with crashes or unclean reboots
journalctl --list-boots

# Kernel log evidence of panic, watchdog, or power event around the last boot boundary
journalctl -k -b -1 -n 200 --no-pager
# Rule out controller resets masquerading as power events
dmesg | grep -iE "nvme.*(reset|timeout)"
# Firmware version, for checking vendor advisories
nvme id-ctrl /dev/nvme0 | grep -i "^fr"

How to diagnose it

  1. Establish the rate. Record the current counter value with a timestamp. Check again in 24 hours and after the next reboot. A counter that increments on every clean shutdown or every suspend is a false-positive pattern, not a power problem. A counter that increments only when the host crashed or lost power is real.

  2. Compare against power_cycles. A clean shutdown increments power_cycles only; an unsafe one increments both. If power_cycles grew by 10 over a week and unsafe_shutdowns grew by 10 too, every power-off was ungraceful, which points at the platform or the shutdown path, not random outages. If power_cycles grew and unsafe_shutdowns did not, your shutdown path is clean and the historical count is old history.

  3. Correlate with host evidence. For each suspected increment window, find the cause: panic traces in the previous boot’s journal, watchdog messages, IPMI/BMC event log entries, outage reports, or maintenance records. If the host evidence says “clean reboot” but the counter incremented, you are in false-positive territory.

  4. Determine PLP status. SMART does not tell you directly whether a drive has PLP. The one SMART signal you get is negative: critical_warning bit 4 means a drive that has PLP has lost it. Otherwise, identify the drive model (nvme id-ctrl /dev/nvme0 | grep -i mn) and treat any consumer-class M.2 drive in a write-intensive role as unprotected by default. For fleet planning, confirm each model’s PLP capability against the vendor datasheet when you provision it.

  5. If the events are real and the drive lacks PLP: verify data integrity. The corruption, if it happened, is at the filesystem and application layer, not in SMART. Run application-level consistency checks (database page verification, checksum validation), verify backup integrity, and schedule a filesystem check at the next maintenance window. Do not run fsck on a mounted filesystem.

  6. If the events are false positives: collect the evidence (kernel version, firmware version, suspend behavior) and check vendor firmware updates and kernel changelogs. Community reports document several patterns: counters incrementing on every proper shutdown on certain drive and kernel combinations, and counters climbing during s2idle sleep on some laptop platforms, with workarounds involving NVMe power-state kernel parameters. Treat these as community-reported correlations, not guaranteed fixes, and verify against your specific hardware before applying kernel-parameter changes in production.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
unsafe_shutdowns rate (nvme.device_unsafe_shutdowns_count)Each increment is a power-loss event; the rate is the actionable signalAny increment during normal operation
power_cycles (nvme.device_power_cycles_count)Baseline for distinguishing graceful from ungraceful power-offsunsafe_shutdowns growing in lockstep with power_cycles
critical_warning bit 4PLP backup has failed; unsafe shutdowns now carry data-loss risk on a drive that was protectedAny assertion
media_errors rate (nvme.device_media_errors_rate)A cluster of media errors after an unsafe shutdown can be one-time power-loss corruption rather than ongoing degradationNew errors clustered after an unsafe shutdown event
Host reboot/crash evidenceAttributes each increment to a causeBoot with no clean shutdown marker in the journal

Fixes

Real power-loss events

Fix the power path. The drive is reporting a symptom of your infrastructure. Work the chain: PSU health and redundancy, UPS battery state and runtime tests, PDU behavior, and whether the server is on the feed you think it is. If increments correlate with kernel panics, treat the panic as the incident and the unsafe shutdown as collateral.

Protect the data path going forward. For drives without PLP in write-intensive roles, the honest options are: move the workload to enterprise drives with PLP, reduce the exposure window (mount options, application fsync behavior, disabling volatile write cache where the latency cost is acceptable), and make backups frequent enough that silent loss is recoverable. Baseline PLP capability at provisioning time: PLP absence is invisible in SMART, so it must be a procurement and inventory fact, not a discovered one.

If bit 4 is set, treat the drive as a no-PLP drive immediately regardless of its enterprise pedigree, and plan replacement. PLP capacitors degrade with age.

False positives

Update firmware first. Several reported false-positive patterns are firmware-side. Check the current revision against vendor advisories.

Check the kernel side. Some reported patterns were resolved by kernel changes; others by limiting how deep NVMe power states go during suspend (for example, power-state latency parameters floated in community reports). Verify against your hardware before applying, and stage it: a wrong power-state setting can trade a cosmetic counter for real latency or stability problems.

Do not “fix” it by ignoring the counter. If you confirm a firmware false positive, annotate the monitoring baseline for that drive model so the rate alert still means something for the rest of the fleet.

Prevention

  • Alert on the rate, not the value. Any increment of unsafe_shutdowns during normal operation warrants a ticket. The absolute count is archaeology.
  • Baseline PLP capability at provisioning. Record per-model whether drives have PLP. Consumer drives in write-heavy production roles are a standing risk decision, not a default.
  • Monitor bit 4 on enterprise drives. A failed PLP capacitor silently converts a protected drive into an unprotected one with zero performance symptoms.
  • Correlate power events with host telemetry. Every unsafe shutdown should be attributable to a named cause (panic, outage, maintenance) within one investigation cycle. Unattributed increments are either missed outages or false positives, and both matter.
  • Verify clean shutdown behavior in your stack. Some hypervisors and runtimes do not issue clean NVMe shutdown during stop. Test it: cleanly stop a guest or host and confirm power_cycles increments while unsafe_shutdowns does not.
  • Run integrity verification after events. After any real unsafe shutdown on a no-PLP drive, application-level checks and backup verification are part of the incident, not optional follow-up.

How Netdata helps

  • Netdata collects unsafe_shutdowns as nvme.device_unsafe_shutdowns_count alongside power_cycles, so the rate of change is visible as a trend rather than something you have to remember to poll by hand.
  • The nvme.device_critical_warnings_state chart decodes the critical warning bitmask per bit, so a PLP backup failure (bit 4) surfaces as its own signal instead of being buried in a nonzero aggregate.
  • Because media_errors is tracked as an incremental rate on the same device, you can see whether a cluster of media errors followed a specific unsafe shutdown event, which changes the diagnosis from “drive is dying” to “one-time power-loss corruption.”
  • Per-second host metrics around the event window let you confirm what actually happened at the host: a panic, a hard hang, or a clean reboot, which is the fork between a real power event and a firmware false positive.
  • Fleet-wide visibility makes patterns visible: if every host on one PDU increments together, you have a power-infrastructure incident; if one drive model increments on every clean shutdown across the fleet, you have a firmware bug.