The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / zfs / zfs-resilver-in-progress ▌

Operations Guides

ZFS resilver in progress: the reduced-redundancy window after a disk replace

You replaced a failed disk, zpool status now shows scan: resilver in progress, and the pool is DEGRADED. This is normal recovery, but it is not a safe state. Until the resilver completes, the affected vdev is running with reduced redundancy, and every hour in this window is an hour where one more failure can take the pool down.

The playbook rule: a resilver in progress is a TICKET, not a page, but it is a ticket you actively manage. You do not close it because the replacement is physically seated. You close it when the scan finishes, the vdev is back at full redundancy, and the error counters on the surviving disks are still zero.

This article covers how to read the progress output, why the ETA is unreliable, what the real risk is for your topology, and what to watch on the surviving disks while the rebuild runs.

What this means

A resilver reconstructs the data that belongs on the replaced device by reading the surviving devices in the same vdev. Two properties matter operationally:

  • Redundancy is reduced for the entire duration. On a mirror, you are down one side. On RAIDZ1, you have zero parity left: a second failure in the same vdev during the resilver is pool loss. On RAIDZ2 you drop to single-parity tolerance; RAIDZ3 drops to double-parity. The pool stays online and serves I/O, but the safety margin you sized the pool for is partially or fully gone.
  • The rebuild hammers the surviving disks. The resilver reads from exactly the devices you are now depending on. This is the worst time to discover that a second disk in the group was also dying.
flowchart TD
  A[Disk fails or is replaced] --> B[Pool DEGRADED]
  B --> C[zpool replace triggers resilver]
  C --> D[Resilver in progress: reduced redundancy]
  D --> E{Resilver completes?}
  E -->|yes| F[Full redundancy restored]
  E -->|second failure in same vdev| G[RAIDZ1: data loss
RAIDZ2/Z3: tolerance consumed] D -->|sequential resilver| H[Auto scrub verifies checksums] F --> H H --> I[Pool back to healthy baseline]

RAIDZ and mirror resilvers are different workloads

The workload difference changes how you plan:

  • RAIDZ resilver (healing resilver) walks the block tree. It reconstructs only allocated blocks, but the walk produces random read I/O across the surviving disks. On HDDs this is IOPS-bound, not bandwidth-bound, and it gets slower as the pool fills and fragments. Multi-day resilvers on large RAIDZ pools with spinning disks are common.
  • Mirror resilver is sequential. It copies the device, with sequential I/O proportional to disk size. Much faster and much more predictable.
  • Sequential resilver (zpool replace -s) speeds up mirrors and dRAID; the flag exists in OpenZFS 2.0, while dRAID itself was introduced in OpenZFS 2.1. It is not supported for RAIDZ. When a sequential resilver finishes, ZFS automatically starts a scrub to verify checksums, so expect continued background I/O after the “resilver done” line appears.

The -s flag is not the default; if you want sequential resilver on a mirror you must pass it explicitly at replace time.

Why the ETA lies

The scan: line shows bytes done, rate, percentage, and an ETA. Treat the ETA as a rough guess, especially in the first hour:

  • The rate is an average since the resilver started. Early numbers include warmup and metadata-heavy phases and do not reflect steady-state throughput.
  • RAIDZ healing resilvers are random I/O. The rate varies with where the block tree walk currently is, how fragmented the pool is, and how much production I/O is competing.
  • ZFS throttles scan I/O to protect production traffic by default, so the rate drops when your workload is busy and recovers when it is idle. An ETA computed during a quiet night will be wildly optimistic about a busy day.

Use zpool iostat -v for actual per-device throughput instead of trusting the scan line’s rate column.

Quick checks

# Current resilver status and progress
zpool status <pool> | grep -A5 "scan:"

# Watch progress once a minute
watch -n 60 'zpool status <pool> | grep -A5 "scan:"'

# Real per-vdev throughput: is the rebuild actually moving?
zpool iostat -v <pool> 5

# Latency per vdev: are surviving disks struggling?
zpool iostat -l <pool> 5

# Queue depths: is production I/O piling up behind resilver I/O?
zpool iostat -q -v <pool> 5

# Error counters on the SURVIVING devices: this is the check that matters most
zpool status -v <pool>

# Kernel log for transport-level trouble on surviving disks
dmesg | grep -i -E "ata|sas|reset|timeout"

All of these are read-only and safe to run at any interval.

What to watch on the surviving disks

The most important monitoring target during a resilver is not the new disk. It is the error counters and latency of the disks the resilver is reading from. A resilver is effectively a full-surface read test of every surviving device in the vdev, and it routinely surfaces latent failures.

  • READ/WRITE/CKSUM counters on surviving devices. Zero is the only acceptable value. A counter that starts incrementing during the resilver means the rebuild is reading a disk that is also failing. On RAIDZ1, that is the beginning of the worst-case scenario: stop non-essential I/O and escalate immediately.
  • Per-vdev latency divergence. If one surviving disk is consistently slower than its peers in zpool iostat -l, it is a replacement candidate. A slow disk both extends the window and signals distress.
  • Scrub interaction. Scrubs and resilvers are mutually exclusive per pool; the resilver preempts any scrub that was running. Do not try to force a scrub mid-resilver. If you used sequential resilver, the verification scrub starts automatically afterward.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
scan: progress and rateTracks window length; stall detectionZero progress for >10 minutes, or rate far below device capability
Surviving-device READ/WRITE/CKSUMEarly warning of a second failure in the groupAny non-zero value, especially if incrementing
Per-vdev latency (zpool iostat -l)Finds a struggling survivor before it faultsOne vdev consistently 3x slower than peers
Queue depth (zpool iostat -q)Distinguishes resilver contention from a hangPending » active sustained on data vdevs
Pool state (zpool status -x)The window closes only when redundancy is restoredDEGRADED persisting with no resilver active
Application-facing write latencyResilver competes with production I/OSustained latency >2x baseline during the window

Operating during the window

These are judgment calls, but they are the decisions a resilver forces on you:

  • Do not defer it. DEGRADED with no resilver running is worse than resilver in progress. If the pool is DEGRADED and no scan is active, find out why: hot spare did not activate, replacement was never issued, or the replace failed.
  • Reduce non-essential I/O if the topology is thin. On RAIDZ1 or a two-way mirror, the marginal value of shortening the window is high. Postponing backups, scrubs on other pools sharing the enclosure, and bulk jobs is a reasonable trade.
  • Do not stack maintenance. This is not the time for firmware updates, controller work, or replacing a second disk in the same vdev. One change at a time, and the resilver finishes first.
  • Expect elevated latency and plan around it. Elevated latency during resilver is expected behavior, not a new incident. Alert thresholds that fire on latency alone will false-fire for the entire window; correlate with queue depth and error counters instead.
  • Throttle tradeoff. Default scan throttling prioritizes production I/O over resilver speed. Raising resilver priority shortens the risk window but taxes production latency. The tunables and their defaults vary by OpenZFS version, so check the module parameter documentation for the version you run before changing anything.
  • If it stalls. A resilver that makes zero progress for more than 10 minutes is abnormal. Check dmesg for device resets, check per-vdev latency for a disk that stopped responding, and check that the replacement device itself is healthy.

Prevention

You cannot prevent resilvers, but you can control how long and how dangerous they are:

  • Baseline your resilver time. Record how long the first resilver on a large pool takes at your current fill level so you can size the risk window before the next disk fails.
  • Prefer topology with margin. RAIDZ1 on large HDDs produces multi-day windows with zero redundancy. RAIDZ2/Z3 or mirrors cost capacity but buy you tolerance during the rebuild. This decision is made at pool creation; there is no retrofit.
  • Keep pools below the performance cliff. Healing resilvers on RAIDZ slow down as pools fill and fragment. Capacity discipline (see the capacity planning guide below) shortens future resilver windows.
  • Scrub on schedule. Regular scrubs surface dying disks before they fail outright, which means replacements happen on your schedule instead of during an incident, and the surviving disks are in known-good shape when the rebuild starts.
  • Use sequential resilver where supported. On mirrors (OpenZFS 2.0+) and dRAID pools (OpenZFS 2.1+), zpool replace -s meaningfully shortens the window. RAIDZ has no sequential path; plan accordingly.

How Netdata helps

  • Pool and vdev state over time. Netdata tracks the transition into DEGRADED and back, so the ticket opens when the disk fails and closes only when redundancy is actually restored, not when someone racks the replacement.
  • Scan progress as a trend. Resilver rate and completion percentage as time series make stalls obvious in a way that repeated zpool status runs do not.
  • Per-vdev latency and throughput during the window. Correlating zpool iostat-style latency and IOPS per device surfaces the slow surviving disk that is both extending the rebuild and signaling the next failure.
  • Error counter alerting. Incrementing READ/WRITE/CKSUM on surviving devices during an active resilver is exactly the escalation condition, and it is easy to miss in point-in-time zpool status output.
  • Production impact correlation. Overlaying application-facing latency with resilver progress tells you whether the rebuild is hurting users or just running quietly in the background.