The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / zfs / zfs-permanent-errors-detected ▌

Operations Guides

ZFS permanent errors have been detected in the following files: recovering from data loss

You ran zpool status -v and at the bottom, under the errors section, you see:

errors: Permanent errors have been detected in the following files:

        tank/data@daily-2026-07-18:backups/db.dump
        tank/data/logs/app.log

This is the most severe per-file integrity message ZFS produces. It means one or more data blocks failed checksum validation and ZFS could not repair them from redundancy. The data in those blocks is gone. This is not a warning, not a transient condition, and not something a reboot or zpool clear will fix. Irrecoverable data loss has already occurred.

The message is also your recovery checklist. ZFS is telling you exactly which files or objects are damaged, which is more than most storage stacks will ever give you. Your job now is to quantify the loss, stop it from growing, restore what you can, and find the hardware that caused it before it takes more data with it.

What this means

Every block ZFS writes carries a checksum. On every read, and during every scrub, ZFS re-verifies the block against that checksum. When a block fails verification, ZFS tries to repair it from a redundant copy: the other side of a mirror, or RAIDZ parity. If that repair succeeds, you get a correctable error (the CKSUM counter increments, the bad copy is rewritten, and no data is lost). If there is no good copy to repair from, the error is uncorrectable and the affected file or object is added to the permanent error list you are looking at.

Key properties of this state:

  • It is binary and cannot false-fire. No workload, idle period, backup job, or batch process can produce a permanent error report. ZFS attempted repair and failed. Treat it as a page-level event.
  • The pool usually stays ONLINE. ZFS isolates the damage to the affected blocks. The rest of the pool keeps serving I/O, which is why this often goes unnoticed until someone reads zpool status -v or a scrub completes.
  • The listed files are unrecoverable from the pool. Redundancy already failed to save them. Recovery means restoring from backups, replicas, or snapshots taken before the corruption.
  • The list is scrub-scoped. zpool status -v reports data errors since the last complete pool scrub; zpool clear clears device error counters, but is not the mechanism that empties the file list.

The entry format matters for scoping the damage:

  • A plain path like tank/data/logs/app.log is a live file in a mounted dataset.
  • A path containing @, like tank/data@daily-2026-07-18:backups/db.dump, is a file as it exists inside a snapshot. You cannot simply rm it; you have to deal with the snapshot.
  • A hex object reference such as <0x2f1d1c> means ZFS could not map the damaged block back to a path, typically because the file was deleted but the object is still referenced (for example by a snapshot or an open file handle).
  • Entries of the form <metadata>:<0x...> indicate damage in pool metadata rather than file contents. These are the most serious, because metadata corruption can make entire datasets or snapshots inaccessible, and there is no file to delete and restore.
flowchart TD
  A[Block fails checksum] --> B{Redundant copy available?}
  B -->|yes| C[Repair from mirror or parity
CKSUM counter increments
no data loss] B -->|no| D[Uncorrectable error] D --> E[File or object added to
permanent error list] E --> F{Recovery path} F --> G[Restore file from backup or replica] F --> H[Roll back or destroy affected snapshot] F --> I[Accept loss and delete damaged file]

Common causes

Permanent errors are an outcome, not a root cause. The root cause is whatever destroyed the data on every redundant copy at once.

CauseWhat it looks likeFirst thing to check
Redundancy exhausted before repairPool was DEGRADED (or has no redundancy at all) when corruption hit; every permanent error on a single-disk stripe is irrecoverablezpool status vdev tree for DEGRADED, FAULTED, or UNAVAIL devices
No scrubs for monthsCorruption accumulated silently; a single later failure turned correctable errors into permanent oneszpool status scan line: last scrub date and result
RAM corruption written to diskCKSUM errors on multiple unrelated devices simultaneouslyECC error logs; run memtest86
Controller or cable/backplane faultREAD/WRITE errors alongside CKSUM on devices sharing a controller or pathdmesg for SATA/SAS resets, timeouts, task aborts
Dying disk that went unactionedOne device with growing CKSUM/READ counts and elevated latency over weeksSMART data (Reallocated_Sector_Ct, Current_Pending_Sector)
Snapshot proliferationThe same bad block referenced by many snapshots appears as many list entrieszfs list -t snapshot -o name,used,refer -s used -r <pool>

Quick checks

All read-only. Run these before changing anything.

# Full damage report: vdev states, error counters, and the permanent error list
zpool status -v <pool>

# Last scrub result and whether one is running now
zpool status <pool> | grep -A 5 "scan:"

# Same status, but with full device paths instead of shortened names
zpool status -p <pool>

# Pool-level health summary
zpool list -H -o name,health,cap,frag <pool>

# Which snapshots reference the affected datasets
zfs list -t snapshot -o name,used,refer -s used -r <pool>

# Kernel-level hardware errors that correlate with ZFS error counters
dmesg | grep -i -E "ata|sas|reset|timeout|uncorrect"

# SMART health of the devices showing non-zero error counters
smartctl -A /dev/sdX

Two notes on interpretation:

  • Error counters (READ, WRITE, CKSUM) are cumulative since the last zpool clear. A device showing old errors may have already been replaced; check timestamps and history before blaming current hardware.
  • There is a known, still-open OpenZFS defect (#11545) where checksum errors repaired after a previous repair may not increment the per-vdev CKSUM counter. Do not treat CKSUM=0 as proof that nothing was ever wrong.

How to diagnose it

Work through this in order. The goal is to answer three questions: what exactly was lost, is the loss still growing, and what hardware caused it.

  1. Capture the full error list for the incident record. Save zpool status -v <pool> output to a file before you touch anything. Once a complete scrub replaces the error-log view, this evidence is gone, and you will want it for restore validation and post-incident review.
  2. Classify each entry. Split the list into live files, snapshot-referenced files (paths with @), unresolvable object IDs (<0x...>), and metadata entries. The recovery action differs per class.
  3. Check pool and vdev state. zpool status -v: is the pool ONLINE, DEGRADED, or FAULTED? Are error counters concentrated on one device or spread across several? One device points to that disk. Several unrelated devices with simultaneous CKSUM errors point to RAM or the controller, not the disks.
  4. Check whether the loss is still growing. Re-run zpool status after some I/O or after the next scrub and compare counters. Rising counts mean an active failure; static counts mean the damage may be historical (for example from a device that already failed and was replaced).
  5. Correlate with hardware telemetry. SMART data on the devices with errors, plus dmesg for link resets and timeouts. If CKSUM errors appear on multiple devices at once with no disk-level evidence, schedule a memory test: bad RAM produces bad data with valid checksums written to every disk, and ZFS faithfully stores the corruption.
  6. Verify your backups before restoring. Confirm that your backup or replica actually contains a good copy of each listed file, from before the corruption window. Restoring a corrupted backup over the top just re-imports the damage.
  7. Determine the detection gap. Compare the last scrub date against the likely corruption window. If scrubs were not running, the corruption sat undetected and a single later failure converted correctable damage into permanent loss. This gap is a process failure to fix, not just a hardware one.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Permanent error list (zpool status -v errors section)Direct evidence of data loss; the recovery checklistNon-empty. Page immediately.
Scrub result (scan line)Tells you whether errors were correctable or uncorrectable, and when integrity was last verified“with N errors” where N > 0 and unrepaired; no scrub in 30+ days
Per-vdev CKSUM counterPinpoints which device returned corrupt dataAny non-zero value; growth over time
Per-vdev READ/WRITE countersTransport-level failures that often precede or accompany corruptionAny non-zero value; rising counts
Pool health stateDEGRADED means redundancy is already consumed and the next error may be permanentAny state other than ONLINE
SMART reallocated/pending sectorsDrive-level confirmation of media failure behind ZFS errorsRising counts on the device with CKSUM errors
Memory ECC errorsExplains simultaneous corruption across unrelated devicesAny uncorrected ECC events

Fixes

There is no repair for the damaged blocks themselves. Every action below is about restoring data, stopping the bleed, and clearing the record.

Restore the affected files from backup

This is the primary recovery path. For each live file in the list, restore a known-good copy from backup, replica, or a snapshot that predates the corruption. After restoring, verify the file (application-level check, hash comparison against the backup source, or a test restore for databases). Do not assume a restore succeeded because the command exited 0.

Handle snapshot-referenced corruption

You cannot delete a file out of a snapshot. If the damaged block only exists inside snapshots, your options are:

  • Roll back the dataset to a good snapshot (zfs rollback), accepting the loss of everything written since. Disruptive; confirm the target snapshot is itself clean first.
  • Destroy the affected snapshots if they are expendable. Note that a block shared across many snapshots produces many list entries, and every snapshot referencing it must go before the entry can clear.
  • Leave read-only snapshots in place if they are retention-critical and the damaged content is tolerable, and document the decision.

Handle unresolvable object IDs and metadata entries

For <0x...> entries with no path, the file is already deleted from the live filesystem. The reference is usually held by a snapshot or an open handle; destroying the referencing snapshot or closing the handle releases it. For <metadata> entries, there is no user-facing file to fix. If the pool still imports and serves I/O, plan a migration: zfs send | zfs recv the healthy datasets to a new pool, then retire the damaged one. Metadata corruption does not heal.

Replace the failing hardware

Do this before or alongside data restoration, not after. If one device shows concentrated errors, replace it (zpool replace) and let the resilver complete. If errors point to RAM or the controller, fix that first; replacing disks against a corrupting memory path accomplishes nothing. Until the root cause is removed, every restored file is at risk of being corrupted again.

Clear the error record

Once the data is restored or the loss is accepted, and the hardware is fixed:

# Reset device error counters after the root cause is fixed
zpool clear <pool>

# Verify integrity end to end and confirm no new damage
zpool scrub <pool>

Do not run zpool clear while counters are still climbing. It resets the odometer while the failure is in progress and destroys your ability to quantify the problem. The scrub after the clear is not optional: it is the only way to confirm the pool is clean going forward. On OpenZFS 2.2.0 and later, zpool scrub -e scrubs only the blocks in the error log, which is much faster than a full scrub for validating previously damaged regions; it requires head_errlog and at least one prior scrub.

Prevention

  • Run scrubs on a schedule and alert on the result. Production pools should complete a scrub every 7 to 14 days. Alerting on “scrub completed with errors” is not enough; also alert on “no scrub completed in 30 days.” Zero errors without a recent scrub means zero errors detected, not zero errors present.
  • Treat correctable errors as hardware tickets. Every repaired CKSUM error is a device telling you it is failing. Replacing that device in business hours is what prevents the next scrub from reporting permanent errors.
  • Respond to DEGRADED same-day. A DEGRADED pool is one failure away from exactly this article. Permanent errors are what DEGRADED turns into when you wait.
  • Never run production data without redundancy. On a single-disk stripe, ZFS can detect corruption but can never repair it. Every checksum error is data loss by definition.
  • Use ECC memory. Non-ECC RAM can hand ZFS corrupted data with a valid checksum, which ZFS then writes to every redundant copy faithfully.
  • Keep tested, restorable backups. ZFS redundancy protects against device failure, not against corruption that reaches all copies. The permanent error list is the moment you find out whether your backups actually work.

How Netdata helps

  • Scrub and error-state visibility: Netdata surfaces pool health state, per-vdev READ/WRITE/CKSUM counters, and scrub status continuously, so a permanent error condition shows up on a dashboard and in alerts minutes after a scrub or read detects it, not weeks later when someone happens to run zpool status -v.
  • Correctable versus uncorrectable trend lines: Historical CKSUM counter graphs show whether you are in the “redundancy is absorbing damage” phase or the “damage is now permanent” phase, which is the difference between a hardware ticket and a data-loss incident.
  • Hardware correlation: ZFS error counters charted alongside disk SMART attributes and system-level I/O errors let you tie a specific failing device to the corruption window, instead of guessing which of twelve disks to replace.
  • Scrub recency as an alertable signal: Alerting when a pool has not completed a scrub within your policy window closes the detection gap that turns correctable errors into permanent ones.
  • DEGRADED state paging: Pool state transitions out of ONLINE page immediately, giving you the chance to restore redundancy before the next error becomes unrecoverable.