The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / zfs / zfs-cannot-destroy-dataset-busy ▌

Operations Guides

ZFS cannot destroy dataset is busy: clones, holds, and mounted filesystems

You run zfs destroy on a dataset or snapshot and get one of two errors back:

cannot destroy 'tank/data': dataset is busy

or:

cannot destroy 'tank/data@daily-2026-07-20': snapshot has dependent clones

Both mean the same thing at a mechanical level: something still references the object you are trying to remove, and ZFS refuses to guess which reference you are willing to lose. The frustrating part is that the error tells you almost nothing about what holds the reference. The job here is to find the dependent, understand why it exists, and remove it in the right order.

One distinction up front: this article covers EBUSY-style failures where a dependent or a reference blocks destruction. A destroy that fails or hangs because the pool is completely full is a different failure mode (COW metadata updates themselves need free space) and belongs in capacity triage, not here.

What this means

ZFS datasets, snapshots, and clones form a dependency graph:

  • A snapshot is a read-only point-in-time view of a dataset.
  • A clone is a writable dataset created from a snapshot. The clone depends on its origin snapshot until you promote or destroy the clone. You cannot destroy a snapshot that has clones.
  • A hold is a user-created tag pinned on a snapshot, typically placed by backup or replication tools so retention jobs cannot destroy a snapshot out from under an incremental send. Destroying a held snapshot returns EBUSY.
  • A mounted or open dataset is busy in the kernel sense: a mountpoint, an open file handle, an NFS export, or a zvol mapped by a hypervisor all count as references.

The destroy path checks all of these and stops at the first blocker. Your job is to enumerate the blockers instead of blindly retrying with bigger flags.

Common causes

CauseWhat it looks likeFirst thing to check
Clone depends on the snapshot“snapshot has dependent clones”zfs get clones <snapshot> and zfs get origin on suspects
Hold tag pins the snapshot“dataset is busy” on a snapshot destroyhold audit over the pool’s snapshots, or the userrefs property
Dataset still mounted or in use“dataset is busy” on a filesystem destroyzfs get mounted, then mount tables
Mount inside a container namespaceHost shows nothing mounted, destroy still failsgrep <dataset> /proc/*/mountinfo
Lazy unmount left a referenceMountpoint gone from /proc/mounts, still busyRecent umount -l in shell history or scripts
NFS export holds the datasetNo local mount, still busyexportfs -v
VM process holds a zvol open“dataset is busy” on a zvol destroyfuser -am /dev/zvol/<pool>/<dataset>

Quick checks

These are all read-only and safe to run on a production pool.

# Is the dataset itself mounted right now?
zfs get mounted,mountpoint tank/data

# Does this snapshot have dependent clones?
zfs get clones tank/data@daily-2026-07-20

# Which snapshot is this dataset cloned from?
zfs get origin tank/data-copy

# List all holds on every snapshot in the pool
zfs list -H -t snapshot -r -o name tank | xargs -r -n1 zfs holds

# How many user holds does this snapshot have?
zfs get userrefs tank/data@daily-2026-07-20

# Search every process mount namespace for the dataset name
grep -l "tank/data" /proc/*/mountinfo 2>/dev/null

# Check for active NFS exports
exportfs -v

# Find processes holding a zvol open
fuser -am /dev/zvol/tank/vm-100-disk-0

# See space committed to snapshots and children per dataset
zfs list -o space -r tank

The grep /proc/*/mountinfo check is the single most valuable one in container environments. A mount that exists only inside a container’s mount namespace is invisible in the host’s /proc/mounts, and ZFS still sees the dataset as busy. This is the classic trap when deleting LXD/LXC containers backed by ZFS.

How to diagnose it

Work through the blockers in this order. The error message usually hints at the branch: “dependent clones” means clone path, plain “dataset is busy” means holds or mounts.

flowchart TD
  A[zfs destroy fails] --> B{Error text}
  B -->|has dependent clones| C[zfs get clones on snapshot]
  B -->|dataset is busy| D[list snapshot holds]
  C --> E[promote clone or destroy it]
  D -->|holds found| F[zfs release tag]
  D -->|no holds| G{Snapshot or filesystem?}
  G -->|filesystem| H[check mounts, namespaces, NFS, zvol users]
  G -->|snapshot| I[recheck holds and userrefs]
  H --> J[umount, exportfs -u, or stop the process]
  1. Read the error carefully. “snapshot has dependent clones” is definitive: go to step 2. Bare “dataset is busy” on a snapshot is almost always a hold: go to step 3. On a filesystem or volume, it is a mount or open handle: go to step 4.

  2. Map the clone dependency. Run zfs get clones <snapshot> to see exactly which clones depend on the snapshot, and zfs get origin <clone> on each to confirm the relationship. Decide per clone: is it still in use, or is it leftover from an old test, template, or container?

  3. Enumerate holds. Audit the snapshots: zfs list -H -t snapshot -r -o name <pool> | xargs -r -n1 zfs holds. Then inspect a specific snapshot with zfs holds <snapshot> and zfs get userrefs <snapshot>. A non-zero userrefs with no output from zfs holds on that snapshot usually means you looked at the wrong snapshot in the chain.

  4. Hunt hidden mounts. If the object is a filesystem and zfs get mounted says no (or you already unmounted it), check mount namespaces with grep -l "<dataset>" /proc/*/mountinfo. If a PID comes back, that process has the dataset mounted inside its own namespace. For zvols, check fuser -am /dev/zvol/<pool>/<name> for hypervisor processes holding the device open.

  5. Check exports and lazy unmounts. exportfs -v shows NFS exports, which hold a reference even with no client connected. If someone ran umount -l earlier, the mountpoint looks gone but the reference lingers until the last user exits; the cleanest resolution is usually to finish whatever holds it or reboot.

  6. Resolve, then retry the destroy. Apply the matching fix below, then re-run the same zfs destroy command. Do not escalate to -R until you know exactly what the dependents are.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
usedbysnapshots per datasetClone and snapshot chains pin space you think you deletedSnapshot space dominating a dataset’s usage
userrefs on snapshotsReveals hold accumulation from backup toolingHolds with no owner you can identify
Pool freeing propertyShows async reclaim in progress after destroysLarge freeing backlog that never drains
Snapshot count and ageRetention jobs that create but never prune lead to this errorSnapshot count growing monotonically week over week
Clone count per poolEvery clone is a future “dependent clones” errorClones surviving past the task they were created for

Fixes

Clone dependencies: promote or destroy

You have three real options, and they are not equivalent:

Destroy the clone. If the clone was temporary (a test environment, a container layer, a sandbox), destroy it first, then destroy the snapshot:

# Destroys data in the clone. Verify it is expendable first.
zfs destroy tank/data-copy
zfs destroy tank/data@daily-2026-07-20

Promote the clone. zfs promote <clone> reverses the dependency: the clone becomes the origin and the original dataset becomes a clone of it. After promotion you can destroy the original dataset, but understand what happened: promote does not break the dependency, it swaps parent and child. The blocks are still shared; you have not made the clone independent. If the clone diverged heavily and you want a fully independent copy, the only route is zfs send | zfs receive into a new dataset.

Force the recursive destroy. zfs destroy -R <snapshot> recursively destroys all dependents, including clones outside the target hierarchy. This is destructive across dataset boundaries and the list of casualties may be longer than you expect. Run it with the dry-run flag first (-n -v) to see exactly what would be destroyed. Never use -R as a first response to an error you have not diagnosed.

Holds: release the tag

List the tags, confirm which tool created them (the tag name usually tells you: backup software, replication tooling, sanoid-style retention), then release:

zfs holds tank/data@daily-2026-07-20
# Destructive in intent: removes the protection the hold provides
zfs release <tag> tank/data@daily-2026-07-20

Only release a hold when you know why it exists. Holds placed by replication tooling protect the common snapshot that the next incremental send depends on. Releasing and destroying it can force the next replication into a full send.

If you want the snapshot gone but it is still held, zfs destroy -d <snapshot> marks it for deferred destruction: it is automatically destroyed when the last hold is released and the last clone goes away. This is the clean answer when a retention script will release the hold on its own schedule.

Mounted or in-use datasets: unmount at the right layer

For an ordinary mount:

zfs umount tank/data
# or
umount /tank/data

For a mount trapped in a container namespace, find the PID from /proc/*/mountinfo and unmount inside that namespace:

# <PID> holds the dataset in its mount namespace
nsenter -t <PID> -m -- umount <mountpoint>

For NFS, remove the export before destroying:

exportfs -u *:/tank/data

For zvols held by a VM, shut the VM down cleanly rather than killing the hypervisor process; a killed qemu process can leave the device in a state that still reports busy. When every diagnostic comes up empty and the reference is a stale kernel artifact (lazy unmount leftovers, a dead namespace), a reboot clears it.

Prevention

  • Track holds like snapshots. Holds are invisible in zfs list. Periodically run zfs list -H -t snapshot -r -o name <pool> | xargs -r -n1 zfs holds to catch accumulation from tooling before you hit EBUSY during an emergency cleanup.
  • Alert on snapshot space and count. Runaway snapshot retention is what turns a routine destroy into a dependency maze. Watch usedbysnapshots per dataset and the total snapshot count; the failure almost always starts as a pruning job that silently stopped.
  • Name hold tags after their owner. When every tool uses a distinct tag, zfs holds output is self-documenting and you know exactly who to ask before releasing.
  • Treat clones as temporary by default. Give clones created for tests and templates an explicit owner and expiry. Long-lived clones should be promoted deliberately or replaced with send/recv copies.
  • Keep destroys out of capacity emergencies. Destroying snapshots on a nearly full pool competes with foreground I/O for async reclaim. Prune on a schedule so the destroy path is never your first response to a full pool.

How Netdata helps

  • Netdata’s ZFS collectors chart pool capacity, per-dataset usage, and snapshot space over time, so you see snapshot and clone accumulation weeks before a destroy fails during a capacity incident.
  • The pool freeing trend shows whether async reclaim after destroys is draining or backing up, which distinguishes “destroy is slow” from “destroy never happened”.
  • Correlating snapshot count growth with pool capacity growth on the same dashboard tells you whether a capacity problem is live data or retention, which determines whether you will be destroying snapshots at all.
  • Alerting on capacity growth rate gives you runway to clean up clones and holds calmly, instead of discovering them at 3 a.m. with a pool at 96%.