The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / zfs / zfs-scrub-not-running ▌

Operations Guides

ZFS scrub not running: 'CKSUM 0' means nothing without regular scrubs

The pool looks healthy. zpool status shows every device ONLINE, and the READ, WRITE, and CKSUM columns are all zero. No scrub errors have ever been reported. So the data is safe, right?

Not necessarily. Zero checksum errors means zero errors detected, not zero corruption. ZFS only discovers silent corruption when it reads a block and verifies its checksum, and most blocks on a typical pool are read rarely or never. The scrub is the only mechanism that systematically reads every allocated block and checks it. A pool that has not scrubbed in six months has unknown integrity.

The failure mode that makes this dangerous: scrubs are usually scheduled by cron or systemd timers, and those schedules silently stop working. The timer gets disabled, the cron job is lost in a migration, the pool was renamed and the per-pool timer no longer matches, or a scrub was paused and never resumed. Because “no scrub errors” and “no scrub at all” look identical in the error columns, the first sign is often a disk failure months later, when the resilver surfaces corruption that regular scrubs would have caught and repaired while redundancy was still intact.

This article covers how to tell whether scrubs are actually running, why they stop, and how to alert on scrub execution rather than scrub results.

What this means

A scrub walks the allocated block tree of the pool, reads every block, and verifies it against the checksum stored in metadata. When verification fails and redundancy exists (mirror, RAIDZ), ZFS repairs the block from the good copy. Without redundancy, the error becomes permanent and the affected file lands in the permanent error list.

The scan: line in zpool status is the authoritative record of scrub activity. It has three states:

  • In progress: scan: scrub in progress since <timestamp> with bytes scanned, rate, percent done, and ETA.
  • Completed: scan: scrub repaired 0B in 5h4m with 0 errors on Sun Feb 4 17:09:01 2018.
  • Never run: scan: none requested.

If the completed timestamp is months old, or the line says none requested, your “0 errors” status is meaningless. Corruption sits undetected on blocks nobody has read. When a device finally fails, you discover it during the resilver, which is the worst possible time: the pool is already DEGRADED and on a RAIDZ1 or two-way mirror there is no redundant copy left to repair from.

Two things make this worse than it looks:

  • There is a known OpenZFS issue (#11545, “Checksum errors may not be counted,” still open as of 2026-09-01) where scrub-repaired checksum errors do not always increment the per-vdev CKSUM counter, so even the CKSUM column is not a complete record.
  • Error counters are cumulative since the last zpool clear, and may reset on pool export/import. A historical zpool clear erases the evidence.

The only trustworthy statement about pool integrity is: “a scrub completed on this date, with this result.” Everything else is inference.

Common causes

CauseWhat it looks likeFirst thing to check
Scrub schedule never configuredscan: none requested on a pool in production for monthsLook for cron entries and systemd timers; nothing exists
Systemd timers shipped but never enabledTimer units exist on disk but systemctl list-timers shows nothing for zfs-scrubsystemctl list-timers --all | grep zfs
Cron job lost or overriddenDebian/Ubuntu ships /etc/cron.d/zfsutils-linux, but it may have been removed, or the host was rebuilt without the package configcat /etc/cron.d/zfsutils-linux
Pool renamed or recreatedPer-pool timer unit (zfs-scrub-monthly@oldname.timer) no longer matches any poolCompare enabled timers against zpool list -H -o name
Scrub paused and never resumedscan: line shows a paused scrub; pause state survives reboot and pool exportzpool status | grep -i pause
Scrub preempted by resilver repeatedlyScan line shows resilver activity, scrub never completesCheck zpool status and zpool history for resilver events
Scrub started but never finishingSame scrub “in progress” for weeks on a large or busy poolCheck progress rate in the scan: line over time

Quick checks

All read-only and safe to run any time.

# 1. Current scan state for every pool: last scrub date, in-progress, or none
zpool status | grep -E "pool:|state:|scan:"

# 2. Full detail for one pool, including permanent errors at the bottom
zpool status -v tank

# 3. Confirm the property everyone reaches for does NOT exist on OpenZFS
zpool get last_scrub_time tank
# expected: bad property list: invalid property 'last_scrub_time'

# 4. On recent OpenZFS: TXG up to which the last scrub completed (0 = never)
zpool get last_scrubbed_txg tank

# 5. Are the systemd timers actually active?
systemctl list-timers --all | grep -i zfs

# 6. Is the Debian/Ubuntu cron job present?
cat /etc/cron.d/zfsutils-linux

# 7. History of scrub invocations, useful when the scan line is ambiguous
zpool history tank | grep -i scrub | tail -n 10

Notes on what you will see:

  • Check 3 fails on OpenZFS. last_scrub_time is an Oracle Solaris property. On OpenZFS you parse the scan: line, or read last_scrubbed_txg where available, which is a transaction group number, not a timestamp. Zero means no scrub has completed since the property became available. The property was introduced in OpenZFS 2.3.0.
  • Check 5 commonly returns nothing even on systems where the timer unit files are installed. The units zfs-scrub-weekly@.timer and zfs-scrub-monthly@.timer are shipped by OpenZFS packaging but are not active until enabled per pool. “Vendor preset” state on the template does not mean an instance for your pool is running; check the instantiated unit state.
  • The Debian/Ubuntu cron job scrubs all pools monthly (second Sunday of the month). That cadence exists by default on that packaging, but it is easy to lose in a host rebuild or containerization.

How to diagnose it

Work through these in order. The goal is to answer one question: when did a scrub last complete, and what is supposed to start the next one?

flowchart TD
    A[Read scan: line in zpool status] --> B{What does it say?}
    B -->|none requested| C[No scrub ever run: schedule missing]
    B -->|completed, recent| D[Scrub ran: verify the schedule exists for the next one]
    B -->|completed, 30+ days ago| E[Schedule broken or disabled]
    B -->|in progress for weeks| F[Scrub stalled or I/O starved]
    B -->|paused| G[Paused scrub: persists across reboots]
    C --> H[Check cron and systemd timers]
    E --> H
    G --> I[Resume with zpool scrub]
    F --> J[Check progress rate and pool I/O load]
  1. Establish the last completed scrub. Parse the scan: line from zpool status <pool>. If it says none requested, no scrub has ever run on this pool and integrity is completely unknown. Treat that as a finding in itself.

  2. Compute days since completion. Over 30 days on a production pool with redundancy is a ticket-level gap. The target for production pools is a completed scrub every 7-14 days.

  3. Find the mechanism that is supposed to schedule scrubs. Check systemctl list-timers --all | grep zfs, /etc/cron.d/zfsutils-linux, any site-specific cron or Ansible-managed jobs, and tools like sanoid if deployed. If you find nothing, that is the root cause: scrubs were never scheduled.

  4. Verify the schedule targets this pool. Per-pool systemd timers are instantiated with the pool name (zfs-scrub-monthly@tank.timer). If the pool was renamed, recreated, or migrated from another host, the enabled instance may point at a pool that no longer exists. Compare enabled timer instances against zpool list -H -o name.

  5. Check for a paused or preempted scrub. A paused scrub stays paused across reboots and pool export/import; resume it with zpool scrub <pool>. Scrubs and resilvers are mutually exclusive per pool; a replacement that starts a resilver cancels an in-progress scrub, so a pool with recurring device problems may never complete one. zpool history <pool> | grep -iE 'scrub|resilver' shows the sequence.

  6. If a scrub is in progress but never finishing, sample the scan: line twice, an hour apart, and compare bytes scanned. A scrub crawling for weeks on a live pool usually means heavy production I/O plus default throttling, or a slow device. Check zpool iostat -v 1 for a vdev that lags its peers.

  7. Check the permanent error list. zpool status -v lists files and objects with uncorrectable damage at the bottom. If this is non-empty, data loss has already occurred and the question shifts from “run a scrub” to damage assessment and restore. This list is persistent pool metadata, not a volatile counter: zpool clear does not repair it or erase it. Reassess it with a normal or error scrub after the underlying block has been repaired, removed, or restored.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Days since last completed scrub (parsed from scan:)The only meaningful integrity-freshness metric30+ days on any production pool
Scrub result (errors, bytes repaired)Whether corruption was found and whether redundancy saved itAny repaired errors (ticket); any uncorrectable errors or non-empty permanent error list (page)
Scrub duration trendA scrub that took 4 hours and now takes 12 on similar data volume signals degrading devices or growing fragmentationDuration up 50%+ at similar used capacity
Schedule execution (timer last-run, cron logs)Scrubs fail silently; the result only exists if the schedule firesTimer inactive, cron entry missing, per-pool instance name mismatch
Per-vdev CKSUM countersNon-zero pinpoints the device returning corrupt dataAny non-zero value; sustained growth means failing hardware
Scrub paused statePause survives reboot and export, easy to forgetAny scrub in paused state

One parsing gotcha for monitoring scripts: the scan: line format is not stable. When a scrub runs longer than a day, zpool switches the duration to “N days HH:MM:SS,” which breaks naive awk parsers that assume fixed column positions. Match on keywords (scrub repaired, with N errors on) rather than column indexes, and test your parser against both a completed-scrub line and none requested.

Fixes

No schedule exists: create one

Pick one mechanism, not both, or you will get double scrubs.

On systemd-based systems, enable the shipped per-pool timer explicitly:

# Enable and start the monthly scrub timer for pool "tank"
systemctl enable --now zfs-scrub-monthly@tank.timer

# Verify it is actually scheduled
systemctl list-timers zfs-scrub-monthly@tank.timer

Do this for every pool. The timer is a template unit; enabling the template does nothing for pools you do not instantiate.

On Debian/Ubuntu, confirm /etc/cron.d/zfsutils-linux exists and survived any host rebuilds. If your provisioning system manages the host, put the cron file or the timer enablement into configuration management so the next rebuild does not silently drop it.

Scrub overdue right now: run one manually

# Start a scrub immediately
zpool scrub tank

# Watch progress
zpool status tank | grep -A2 scan:

A scrub is I/O intensive and competes with production traffic; on a large pool it can run for many hours or days under default throttling. Prefer starting it in a low-load window. If you need to interrupt it, zpool scrub -s <pool> stops it and zpool scrub -p <pool> pauses it, but a paused scrub survives reboots and an aborted scrub provides no integrity guarantee for the data it did not reach.

Expect the first scrub after a long gap to find errors. That is the point. If it reports correctable errors, the redundancy did its job and you now have a hardware investigation: check the per-vdev CKSUM counters and SMART data on the implicated device. If it reports uncorrectable errors, zpool status -v lists the damaged files and you are in restore-from-backup territory.

Paused or preempted scrubs

Resume a paused scrub with zpool scrub <pool>. If scrubs keep being canceled by device replacement/resilver activity, fix the underlying device problem first; a pool that resilvers monthly has a hardware or cabling issue, and the scrub gap is a symptom.

Prevention

  • Alert on scrub freshness, not scrub results. Page on uncorrectable errors and non-empty permanent error lists. Ticket on any completed scrub with repaired errors, and on any pool where no scrub has completed in 30+ days. Most teams only alert on errors, which is exactly the gap that lets “no scrub ever ran” pass unnoticed.
  • Monitor that the schedule executes. Track the timer’s last-run timestamp or the cron job’s log output. A schedule that exists but never fires produces the same silent outcome as no schedule.
  • Manage scheduling in configuration management. Timers enabled by hand on a Friday afternoon do not survive the next host rebuild.
  • Keep scrub cadence at 7-14 days for production pools, scheduled in low-load windows. Monthly (the Debian cron default) is a floor, not a target.
  • Trend scrub duration. A lengthening scrub is an early indicator of device degradation and fragmentation, visible months before anything faults.
  • Export zpool status into time series. Point-in-time checks miss transitions, and error counters can be cleared. Continuous collection gives you the “days since scrub” metric for free.

How Netdata helps

  • ZFS pool state collection: Netdata’s ZFS collector charts pool health states and error counters over time, so pool transitions and error-counter growth become time-series data instead of a command someone has to remember to run.
  • Scrub state visibility: where the collector exposes zpool status scan data, scrub completion timestamps, errors found, and bytes repaired are retained historically, which enables alerting on scrub age (30+ days without a completed scrub) rather than only on scrub errors.
  • Per-device error counters: READ, WRITE, and CKSUM counts are charted over time, so a slowly dying disk shows up as counter growth between scrubs, not as a surprise during a resilver.
  • Correlation with pool health: scrub findings and error counters can be viewed alongside pool state (ONLINE/DEGRADED), capacity, and I/O latency on the same dashboard, which makes it obvious when errors coincide with a degrading device.