The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / zfs / zfs-cksum-errors-multiple-devices ▌

Operations Guides

ZFS checksum errors on multiple devices: suspect RAM or the controller, not the disks

zpool status shows non-zero CKSUM counters, and not on one disk. Two, four, or every device in the vdev has them. The instinct is to start RMA-ing drives. Stop. Independent disks do not fail in the same way at the same time. When checksum errors appear on multiple unrelated devices simultaneously, the failure is almost always upstream of the disks: something is corrupting data between the application and the platters, and every disk downstream of it is faithfully recording the damage.

The usual suspects, in rough order of probability: bad RAM corrupting data before it reaches the disks, a failing HBA or RAID controller, or a shared backplane or cabling path. ZFS cannot tell you which, because ZFS only sees what it was handed. The error pattern plus a few hardware-level checks separates them quickly.

This article covers the triage: confirming the multi-device pattern, discriminating RAM from controller from cabling, and handling the data written while corruption was active.

What this means

ZFS checksums every block and verifies that checksum on read. A CKSUM error means the bytes the device returned do not match the checksum stored in metadata. On a single device, that points at the device: dying media, bad sectors, firmware returning stale data. The per-device failure modes are covered in ZFS checksum errors (CKSUM).

When several devices report CKSUM errors in the same window, the probability of multiple independent disks developing checksum-visible faults simultaneously is negligible. What is not negligible:

  • One bad DIMM corrupting every write. ZFS computes checksums over data in memory. If the data is already corrupt in RAM, ZFS computes a perfectly valid checksum over corrupt data and writes both to disk; later reads may also be corrupted in memory before verification. Every device in the pool receives its share of poisoned blocks. This is why ECC RAM is so strongly recommended for ZFS: ZFS trusts what memory hands it.
  • One HBA or controller corrupting I/O in flight. Every device attached to that controller sits behind the same failing component. A dying HBA can produce CKSUM errors on all twelve drives behind it while every drive is healthy.
  • A shared backplane or cabling fault. Less common, but a degraded backplane or a batch of marginal cables affects every device routed through it.

The critical consequence: if the root cause is RAM, the corruption is not limited to blocks that failed checksums. Data corrupted in memory and written with a matching checksum is invisible to scrubs. The CKSUM counter only catches the subset where corruption happened after checksumming, or where verification itself read back different bytes. Treat the counter as a symptom of a much larger problem.

flowchart TD
  A[CKSUM errors on multiple devices] --> B{Errors only on devices behind one HBA or backplane?}
  B -->|Yes| C{SMART UDMA_CRC errors or SATA/SAS resets in dmesg?}
  C -->|Yes| D[Suspect cabling or backplane]
  C -->|No| E[Suspect HBA or controller]
  B -->|No, spread across controllers| F[Suspect RAM]
  F --> G[Run memtest86, check ECC/EDAC logs]
  B -->|Identical counts on mirror members| F

Common causes

CauseWhat it looks likeFirst thing to check
Bad RAM (non-ECC)CKSUM errors across devices on different controllers, often symmetric counts on mirror members, no SMART anomalies, no bus resetsBoot memtest86, run multiple passes
Failing HBA / controllerCKSUM (and often READ/WRITE) errors on every device attached to one controllerdmesg for SAS/ATA errors; controller firmware level
Cabling or backplaneCKSUM plus non-zero UDMA_CRC in SMART, link resets in kernel logssmartctl -A UDMA_CRC_Error_Count; reseat cables
Marginal PSU or power deliveryIntermittent multi-device errors under load, devices dropping and reappearingKernel logs for link downs correlated with load
Software edge caseErrors only during scrub or after rebuild, hardware replacement changes nothingOpenZFS version and release notes for checksum fixes

Two notes on the table. First, a single bad HBA producing CKSUM errors on every drive behind it is a well-documented failure mode: operators have replaced drives, cables, and firmware without improvement, and only replacing the HBA stopped the errors. Second, there are rare OpenZFS-side causes. If CKSUM errors appear on all members of a raidz vdev during scrubs and persist across complete hardware replacement, you may be hitting a parity reconstruction edge case rather than a hardware fault. Check that you are on a current OpenZFS release; OpenZFS 2.3.7 and 2.4.2 included fixes for rare checksum errors after rebuilds. On those branches, go to 2.3.7+ or 2.4.2+ before attributing the pattern to hardware.

Quick checks

All read-only. Run these before touching hardware.

# 1. Full error picture: which devices, which counters, how many
zpool status -v

# 2. Machine-parseable counters for trending
zpool status -p

# 3. Kernel view of the storage bus: resets, aborts, timeouts
dmesg | grep -i -E "ata|sas|reset|timeout" | tail -50

# 4. SMART per device: reallocated sectors, pending sectors, UDMA_CRC
smartctl -A /dev/sdX

# 5. Which devices share a controller (PCI topology)
ls -l /dev/disk/by-path/

# 6. ECC/EDAC error logs, if the platform has ECC
dmesg | grep -i edac
grep -i edac /var/log/messages 2>/dev/null | tail -20

What to look for in the output:

  • CKSUM counters roughly equal across mirror members (for example, identical counts on both legs of a two-way mirror) is a widely reported indicator of corruption outside the disks, typically RAM. The diagnostic value is the shared-source pattern rather than the raw counter arithmetic; mirror read placement is load- and implementation-dependent. SMART on both devices will usually be spotless.
  • Errors clustered on devices behind one PCI device in by-path output points at the HBA or the cabling/backplane downstream of it, not the disks.
  • Non-zero UDMA_CRC_Error_Count in SMART means the drive itself detected corruption on the SATA link. That is a cable, connector, or backplane signal, not a RAM signal. Zero UDMA_CRC across all devices with rising CKSUM pushes suspicion toward RAM or the controller.
  • SATA link resets or SAS task aborts in dmesg correlate transport instability with the ZFS error window. READ/WRITE counters rising alongside CKSUM on the same controller’s devices reinforces a controller or cabling cause.
  • EDAC reports of corrected or uncorrected memory errors settle the RAM question immediately on ECC systems. Uncorrected ECC errors plus multi-device CKSUM means replace the DIMM, not the disks.

Note on counters: READ/WRITE/CKSUM counters in zpool status are cumulative since the last zpool clear, and there is a reported OpenZFS issue (#11545) where scrub-repaired checksum errors may not increment the counter. PR #11609 improved counting and was in 2.1.0, but #11545 remains open, so do not assume the behavior is identical in every release. A zero counter is not proof of no corruption; trend the counters rather than reading them once.

How to diagnose it

  1. Confirm the pattern is multi-device. Run zpool status -v. If only one device has non-zero CKSUM and it is growing, this is a dying-disk problem, not this article. Replace the disk. If two or more unrelated devices have errors in the same window, continue.

  2. Map the topology. Use ls -l /dev/disk/by-path/ to see which PCI device each affected disk sits behind. Errors confined to one HBA or one expander point at the controller path. Errors spanning controllers point at RAM.

  3. Check SMART on every affected device. smartctl -A per disk. Rising Reallocated_Sector_Ct or Current_Pending_Sector on one device plus CKSUM means that device is genuinely dying. Clean SMART with zero UDMA_CRC everywhere means the disks are bystanders.

  4. Check the kernel log for transport errors. dmesg resets, aborts, and timeouts on a shared port implicate cabling, backplane, or HBA. A completely quiet kernel log with climbing CKSUM on multiple devices implicates RAM.

  5. Check ECC logs if available. Any EDAC activity, corrected or not, makes the DIMM a replacement candidate. Absence of EDAC data usually means non-ECC RAM, which is itself the finding: you have no way to detect memory corruption except the pattern you are already looking at.

  6. Test memory offline. Schedule a maintenance window and boot memtest86 (or memtest86+). Run several full passes; four is a practical starting point, and marginal DIMMs can pass a single pass. Also remove any memory or CPU overclocking, including XMP profiles, before concluding the DIMMs are bad. Unstable memory clocks produce exactly this symptom on otherwise good hardware.

  7. Rule out software before replacing everything. If errors appear only during scrubs or only after a resilver, and hardware tests clean, check your OpenZFS version against release notes for checksum-related fixes. Do not skip this step: there are documented cases of operators replacing HBA, PSU, motherboard, CPU, and RAM while the actual cause was software.

  8. Fix the root cause, then repair the pool. After the faulty component is replaced, run a full scrub to let ZFS repair corrupted blocks from redundancy. Only after the scrub completes clean should you consider zpool clear to reset the counters. Clearing counters before the root cause is fixed destroys your trend data.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Per-device CKSUM counters, trendedThe single-versus-multiple-device split is only visible over timeCKSUM rising on 2+ devices in the same window
Per-device READ/WRITE error countersTransport failures alongside CKSUM point at controller or cablingREAD/WRITE growing on all devices behind one HBA
SMART UDMA_CRC_Error_CountDiscriminates link-layer corruption from in-memory corruptionNon-zero and climbing on multiple devices
SMART Reallocated/Pending sectorsProves or clears individual disksRising on a device you suspected was a bystander
EDAC corrected/uncorrected error countsDirect evidence of DIMM failure on ECC systemsAny uncorrected error; rising corrected counts
Kernel storage messages (resets, aborts, timeouts)Hardware-side confirmation of bus instabilityResets correlating with CKSUM increments
Scrub repair counts and permanent error listMeasures how much damage was written while the fault was activezpool status -v listing files with permanent errors

Fixes

Bad RAM

Replace the failing DIMM. If the system runs non-ECC memory and memtest86 passes but suspicion remains high (symmetric mirror errors, clean hardware everywhere else), swap the DIMMs anyway or test on known-good ECC-capable hardware. Marginal non-ECC memory can pass testers and still corrupt under production load patterns.

After replacement, run a full scrub. Blocks corrupted between the device and memory will be repaired from redundancy. Blocks corrupted in memory before checksumming cannot be detected by the scrub, because their checksums match. If the corruption window was long and the data matters, validate critical datasets against backups or replicas. Check zpool status -v for permanent errors; that list is your recovery inventory.

Going forward, use ECC RAM on any host running ZFS. Every write passes through memory before checksumming, so memory integrity is the root of the trust chain.

Failing HBA or controller

Update or reflash controller firmware to a known-good level, reseat the card, and watch counters. If errors continue, replace the HBA. Do not replace the drives behind it first: they will all test clean individually, and replacing them wastes money and triggers pointless resilvers, which are themselves risk windows. After the controller is replaced, scrub, then clear counters.

Cabling or backplane

Reseat or replace cables for affected devices, starting with any device showing UDMA_CRC. On backplane-based chassis, check for expander firmware updates and reseat the affected slots. These errors are often load- or temperature-dependent, so a quiet hour after reseating is not proof of a fix. Trend UDMA_CRC for days.

After any fix

Do not clear counters until the root cause is confirmed fixed and a full scrub has completed. zpool clear resets the evidence. If permanent errors were recorded, work through the file list from zpool status -v and restore affected files from backup before clearing. See ZFS permanent errors have been detected for that recovery process.

Prevention

  • Run ECC memory on ZFS hosts. The single highest-leverage prevention for this failure class. Non-ECC systems have no way to detect the corruption source that produces this exact symptom.
  • Alert on error counter growth, not just pool state. A pool stays ONLINE the entire time this is happening. zpool status -x will say “all pools are healthy” while RAM poisons every vdev. Any non-zero CKSUM on a production pool warrants investigation; growth on multiple devices warrants urgency.
  • Scrub on a schedule, weekly to biweekly for production. Multi-device CKSUM patterns are usually discovered by a scrub. Longer gaps mean longer corruption windows.
  • Trend, do not snapshot. Export per-device counters to a time-series system. The diagnostic value is in the pattern across devices over time, which a single zpool status cannot show.
  • Keep a hardware map. Knowing which drives sit behind which HBA, expander, and backplane turns a multi-hour triage into minutes. Record it before the incident.
  • Do not overclock memory on storage hosts. XMP and memory overclocks trade exactly the stability ZFS depends on for bandwidth a fileserver rarely needs.

How Netdata helps

  • Per-device error counter trends: Netdata tracks READ, WRITE, and CKSUM counters per vdev over time, which is the view needed to distinguish “one disk dying” from “everything behind one component being corrupted”.
  • Correlation with system-level signals: multi-device CKSUM growth can be viewed alongside SMART attributes and kernel error activity on the same node, so the RAM-versus-controller discrimination happens on one dashboard instead of across several terminals.
  • Scrub visibility: tracking scrub completion and repair counts closes the loop after the faulty component is replaced, confirming the pool is repaired rather than merely quiet.
  • Alerting on growth, not state: alerts fire when error counters increment on any device, so the pattern surfaces even while the pool reports ONLINE.
  • Memory and hardware error context: on ECC systems, hardware-level memory error signals on the host can be correlated with the first CKSUM increments to pinpoint when the DIMM started failing.