The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nvme / nvme-media-errors-increasing ▌

Operations Guides

NVMe media_errors increasing: uncorrectable data-integrity errors on NAND

Your monitoring shows the media_errors counter on an NVMe drive going up. In nvme smart-log output the field is labelled media_errors (smartctl calls it “Media and Data Integrity Errors”), and it is the one SMART field you should never explain away: each increment is a read or write where the controller could not maintain data integrity even after its internal ECC and retry mechanisms were exhausted.

The absolute value is almost meaningless on its own. A drive with 3 lifetime errors after four years of service can be perfectly healthy; a drive that went from 0 to 12 errors this week is failing. The rate of change, and what changes alongside it, is the diagnostic. A single new error during a backup run that scanned cold data can be one latent bad block finally being touched. A steady climb combined with critical_warning bit 2 (NVM subsystem reliability degraded) is the drive telling you it is dying.

This guide covers how to read the counter, how to classify what you are seeing, and how to decide between “investigate this shift” and “page someone now.”

What this means

NAND flash cells degrade. The controller fights this constantly with ECC, read retries at adjusted voltage thresholds, and background relocation of marginal blocks into the spare pool. media_errors increments only when all of that fails: the data at some LBA could not be recovered by the drive itself. The NVMe spec includes uncorrectable ECC failures, CRC checksum failures, and LBA tag mismatches in this counter.

Two things follow from that definition:

  • The counter is a lifetime total maintained by the controller. It never decreases, and the host does not write to it. Monitoring tools (including Netdata, which exposes it as the nvme.device_media_errors_rate chart) present it as a rate so that increments are visible instead of buried in a large historical number.
  • An increment does not automatically mean the host lost data. Depending on the event, the controller may still have returned correct data from a redundant internal copy, or the read may have failed up the stack and been served by your RAID mirror or replica. Either way, the primary copy on that NAND was unreadable, and that block will be retired.

The counter interacts with the rest of the SMART picture in specific ways. media_errors can climb while percentage_used is low and available_spare is at 100% (a bad NAND batch or a sudden failure, not wear). It can sit at a nonzero value for years without impact (isolated latent defects on an old drive). And it is a different counter from num_err_log_entries: the error log counts all error events, including non-media errors like invalid admin commands, so it can increment while media_errors stays flat. Do not conflate them.

Common causes

CauseWhat it looks likeFirst thing to check
NAND wear-out (end of life)Steadily accelerating error rate, percentage_used high or past 100%, available_spare declining toward thresholdnvme smart-log: compare percentage_used and available_spare against their trajectories
Latent bad block surfaced by a scanOne or two increments during a backup, scrub, or crash-recovery read of cold data; rate returns to zero afterwardsTiming: did the increment coincide with a full-volume read?
Retention failure or read disturb on cold dataErrors concentrated on reads of data written long ago; device otherwise healthynvme error-log: are the failing LBAs in old, rarely-touched regions?
Unsafe shutdown aftermathA cluster of errors appearing shortly after unsafe_shutdowns incremented; may be a one-time eventCompare unsafe_shutdowns history against the error timeline
Manufacturing defect / bad NAND batchErrors on a young drive with low percentage_used and full available_spareDrive age (power_on_hours) and fleet cohort: are same-batch drives failing too?
PCIe transport or firmware issue masquerading as media errorsErrors correlate with link retraining, AER counters climbing, or a known-problematic firmwareAER counters in sysfs, current vs max link speed, nvme fw-log

Quick checks

All of these are read-only.

# Current media error count and the corroborating SMART fields
nvme smart-log /dev/nvme0 | grep -E "media_errors|num_err_log_entries|critical_warning|available_spare|percentage_used|unsafe_shutdowns"

# Individual error log entries: status codes, namespaces, LBAs
nvme error-log /dev/nvme0

# Drive age and lifetime write volume, for context
nvme smart-log /dev/nvme0 | grep -E "power_on_hours|data_units_written"

# PCIe transport health: in a healthy system these are all zero
cat /sys/class/nvme/nvme0/device/aer_dev_correctable
cat /sys/class/nvme/nvme0/device/aer_dev_fatal
cat /sys/class/nvme/nvme0/device/aer_dev_nonfatal

# Link negotiated down? current should equal max
cat /sys/class/nvme/nvme0/device/current_link_speed
cat /sys/class/nvme/nvme0/device/max_link_speed
cat /sys/class/nvme/nvme0/device/current_link_width
cat /sys/class/nvme/nvme0/device/max_link_width

# Kernel-side view: media errors surface as I/O errors with status codes
dmesg | grep -i nvme | grep -iE "error|sct"

On the kernel side, uncorrectable media errors surface in dmesg with Status Code Type 2 (media and data integrity). A line carrying sct 0x2 / sc 0x81 is an unrecovered read error: the kernel is reporting the same class of event the SMART counter is accumulating. See reading NVMe I/O errors in the kernel log for decoding these lines.

How to diagnose it

flowchart TD
  A[media_errors incremented] --> B{critical_warning bit 2 set?}
  B -->|Yes| C[PAGE: active degradation
verify redundancy, replace drive] B -->|No| D{Rate accelerating?} D -->|Yes, or errors frequent| E{TICKET: check percentage_used
and available_spare trajectory} D -->|Single isolated increment| F{Coincided with cold-data scan
or unsafe shutdown?} F -->|Yes| G[TICKET: likely latent defect
monitor rate, plan replacement] F -->|No| H{TICKET: check AER counters,
link speed, firmware, drive age} E --> I[End-of-life pattern:
procure and schedule swap] H --> J[Transport/firmware cause:
fix path, not the drive]

Work through it in this order:

  1. Confirm the increment and its timing. Pull the current media_errors value and compare against your monitoring history. When did it move, and by how much? A drive that has been at 5 for two years and is still at 5 has no active problem.

  2. Check critical_warning bit 2 immediately. This is the severity fork. Bit 2 means the controller itself has assessed its reliability as degraded. Bit 2 combined with a rising media_errors rate is the PAGE condition: active, confirmed degradation. See decoding the SMART critical warning bitmask and critical warning bit 2.

  3. Read the error log for localization. nvme error-log /dev/nvme0 shows status codes, namespace IDs, and LBAs for individual events. Errors clustered in one LBA region point to localized damage or cold-data retention issues. Errors scattered randomly across the namespace point to general media degradation. Act fast: the log is a circular buffer, and a high error rate overwrites the oldest entries, which are usually the root cause.

  4. Correlate with wear indicators. Pull percentage_used and available_spare. High and rising percentage_used plus declining spare plus rising media errors is the classic end-of-life pattern. Media errors on a drive with 3% used and 100% spare is a defect story, not a wear story.

  5. Check for an unsafe-shutdown trigger. If unsafe_shutdowns incremented shortly before the media errors appeared, the cluster may be corruption from an incomplete write flush rather than ongoing degradation. This matters for the prognosis: one-time event versus progressive failure.

  6. Rule out the transport layer. Check AER counters and link speed/width. PCIe link instability produces data-path errors that can present alongside or instead of genuine NAND errors, and reseating a connector is a very different fix from replacing a drive. Zero AER counters and full link speed close this branch.

  7. Classify and set the response. Isolated increment during a cold-data scan: same-shift investigation, watch the rate. Accelerating rate with corroborating wear signals: replacement pipeline. Bit 2 plus rising errors: treat as an active incident.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
media_errors rate (nvme.device_media_errors_rate)Each increment is an unrecovered integrity failureAny rate above zero on a previously stable drive
critical_warning bit 2The controller’s own reliability assessmentAsserted while media errors are rising: page
available_spare vs spare_threshRemaining bad-block replacement runwayDeclining toward threshold alongside rising errors
percentage_usedEndurance consumed; context for whether errors are wear-drivenHigh or past 100% with errors accelerating
num_err_log_entries rateBroader error activity beyond media errorsRising without media errors: firmware/driver path, not NAND
unsafe_shutdownsPower-loss events that can cause one-time error clustersIncrement shortly before a media error cluster
PCIe AER countersTransport-layer integrity that can masquerade as media errorsAny sustained nonzero rate
Read latencyRetry passes on marginal blocks show up as sporadic read latency before hard errorsSporadic read spikes on cold data with rising media errors

Fixes

End-of-life wear (high percentage_used, declining spare, rising errors)

Replace the drive. There is no remediation for worn NAND. Verify RAID or replication health first, force a rewrite of at-risk data so it lands on fresh blocks (or on the surviving mirror), and schedule immediate replacement. Drives can operate past 100% percentage_used, but once media errors accelerate, the trajectory only goes one way. See available spare below threshold and percentage used at 100%.

Isolated latent defect (single increment during a scan)

No immediate action on the drive itself. The controller has already retired the bad block and mapped in a spare. Record the event, keep the rate alert in place, and confirm your redundancy would have covered the affected LBA if the read had failed upward. If isolated increments keep recurring on successive scans, treat it as the early stage of the wear pattern instead.

Unsafe shutdown aftermath

Fix the power path, not the drive: PSU, UPS, PDU, or whatever caused the ungraceful power loss. Then verify data integrity at the filesystem or application layer, because in-flight writes on a drive without power-loss protection may have been acknowledged but not persisted. Monitor the media error rate over the following days; if it returns to zero, it was a one-time event.

Young drive with errors (defect or batch issue)

Treat as a warranty case. Capture the full SMART log, error log, and firmware version (nvme fw-log /dev/nvme0) before contacting the vendor. Check fleet-wide: if other drives from the same batch and age show similar counters, escalate procurement for the whole cohort.

Transport or firmware cause

Reseat the drive, inspect connectors and cables, and confirm the link trains to full speed and width. Check for a firmware update addressing data-integrity issues before condemning the hardware. If the errors stop after the physical or firmware fix and the SMART rate stays flat, the NAND was never the problem.

Prevention

  • Alert on the rate, not the value. Any media_errors increment within a monitoring window should open a ticket. Reserve the page for the corroborated condition: rising errors plus critical_warning bit 2.
  • Trend available_spare and percentage_used continuously. Media errors rarely arrive without warning on a wearing drive; spare consumption and endurance rate are the leading indicators that give you weeks of runway.
  • Track unsafe_shutdowns and fix power infrastructure. Each ungraceful power loss on a drive without PLP is a dice roll on write-cache data.
  • Baseline PCIe health at provisioning. AER counters at zero and full negotiated link speed should be a deployment gate, so later deviations are visible.
  • Track firmware versions fleet-wide. Data-integrity bugs are fixed in firmware; you cannot act on a vendor advisory if you do not know which drives run the affected version.
  • Scrub cold data periodically. Regular full reads (or filesystem-level scrubs) surface latent blocks while your redundancy can still repair them, and let the controller retire marginal blocks on its own schedule instead of during an incident.

How Netdata helps

  • Netdata collects the NVMe SMART log per device and exposes media_errors as an incremental rate chart (nvme.device_media_errors_rate), so a single new error is visible as a spike instead of being lost in a lifetime total.
  • The critical warning bitmask is broken out per bit, so you can alert on bit 2 (reliability degraded) separately from bit 0 or bit 1 instead of firing one blanket alert on any nonzero value.
  • Available spare, endurance consumed, and unsafe shutdowns sit on the same per-device dashboard as media errors, which is exactly the correlation set this diagnosis depends on.
  • Error log entries have their own rate chart, making the “non-media errors rising, media errors flat” firmware-driver case distinguishable from genuine NAND failure at a glance.
  • Per-second collection means the timing correlation that classifies the event (backup window, power event, link retraining) is preserved rather than averaged away.