The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / zfs / zfs-slog-endurance-wearout ▌

Operations Guides

ZFS SLOG endurance: the SSD that wears out from concentrated sync writes

A SLOG (Separate Intent Log) device is usually the smallest, fastest SSD in the box, and it is almost always the first one to die. Every synchronous write in the pool, from every dataset and every application, lands on this one device before the application gets its acknowledgment. That concentration is the point of a SLOG, and it is also why the device burns through write endurance far faster than the data vdevs around it.

ZFS will not warn you about remaining endurance. It exposes SLOG activity through zil_itx_metaslab_slog_bytes and zil_itx_metaslab_slog_write, but not device wear or remaining life. The pool stays ONLINE, zpool status -x reports healthy, and the SSD’s wear counter climbs silently until the device reaches its endurance limit. Many devices then fail abruptly; some enterprise drives instead enter a vendor-defined protected mode. Your sync-heavy workloads (NFS, databases, anything calling fsync) still find out at the worst possible time.

This article covers why SLOG wear is concentrated, why ZFS cannot show it to you, how to monitor it at the device level, what actually happens when the device fails, and a replacement policy that keeps you ahead of the wear curve.

Why the SLOG absorbs so much write traffic

Every synchronous write (O_SYNC, fsync, NFS, database commit) must be recorded in the ZFS Intent Log before ZFS acknowledges it. Without a SLOG, the ZIL lives on the pool’s data vdevs and sync write latency is pool-speed. Adding a SLOG moves the ZIL onto a dedicated fast device, and sync write latency becomes SLOG-speed. That is the performance win.

The cost is write concentration:

  • Async writes accumulate in RAM and flush to the pool in transaction groups, spread across all data vdevs.
  • Sync writes bypass that path. Each one is written to the ZIL immediately, which means each one is a write to the SLOG device.
  • On a pool serving NFS or database workloads, effectively every write is a sync write, so the SLOG sees a write volume no single data vdev ever experiences.

Two properties of the ZIL make this worse than it looks. First, ZIL data is transient: entries are freed once their transactions commit during TXG sync, so the SLOG holds only a few seconds of data at any time while absorbing an enormous lifetime write volume. Second, the SLOG is write-only in normal operation. It is only ever read during pool import after an unclean shutdown, to replay acknowledged sync writes. You get no read-side warning signs; the device just keeps absorbing writes.

flowchart LR
  A[Sync writes: fsync, O_SYNC, NFS] --> B[ZIL]
  B --> C[SLOG device]
  D[Async writes] --> E[TXG in RAM]
  E --> F[Pool data vdevs]
  C -. read only after crash .-> G[ZIL replay at import]

What happens when the SLOG wears out

Endurance exhaustion on an SSD is often more abrupt than a failing HDD with growing latency and reallocated sectors, but the exact behavior is firmware-specific. Some enterprise NVMe drives implement vendor-defined protection such as reduced performance, asynchronous warning events, or read-only mode before data is lost. Do not assume a universal cliff; trend SMART and consult the drive vendor’s endurance behavior.

What a SLOG failure does to the pool:

  • SLOG failure during normal operation is not data loss. ZFS falls back to writing the ZIL on the pool’s data vdevs. Data already acknowledged is safe.
  • The failure is a performance event, not an availability event. Sync write latency jumps from SLOG-speed to pool-speed, which can be a 10x to 100x regression for NFS and database workloads. The pool state does not change to DEGRADED, because the SLOG is not a data device. zpool status -x may still report all pools healthy while applications are timing out.
  • The one losing combination is an unmirrored SLOG plus a crash or power loss before the next TXG commits. If the SLOG fails and the system goes down in the same window, the ZIL entries for recently acknowledged sync writes are gone. This is why production SLOG devices should be mirrored: a single SLOG failure must never coincide with a sync-write durability gap.

Because a failed SLOG does not degrade pool state, the failure is invisible to the most common health check. You catch it by checking the log vdev explicitly in zpool status, by watching sync write latency, and, long before that, by watching the wear counters.

Why ZFS cannot show you SLOG wear

ZFS does not expose SLOG wear or remaining endurance. It does expose SLOG activity: zil_itx_metaslab_slog_bytes, zil_itx_metaslab_slog_write, and zil_itx_metaslab_slog_alloc in the global ZIL kstat. By contrast, l2_write_bytes in /proc/spl/kstat/zfs/arcstats counts bytes written to the L2ARC, not to the SLOG; alerting on it as a SLOG proxy monitors the wrong device.

The same kstat also tracks commit activity with zil_commit_count, zil_commit_writer_count, zil_commit_stall_count, and zil_commit_error_count. Those counters show whether the ZIL is busy, stalling, or erroring, but none of them estimates consumed NAND endurance; SMART and NVMe health data remain the wear source.

So wear monitoring must come from the device itself, via SMART and NVMe health data:

  • NVMe drives report Percentage_Used, the controller’s estimate of consumed endurance as a percentage of the rated budget.
  • SATA/SAS SSDs report attributes such as Media_Wearout_Indicator or equivalent vendor wear counters, plus total host writes.
  • Total bytes written against the rated TBW (terabytes written) from the manufacturer’s spec sheet gives you an absolute endurance budget independent of the SMART percentage.
# Check SLOG device wear (NVMe example)
smartctl -a /dev/nvme0 | grep -iE "percentage_used|data_units_written"

# SATA/SAS SSD wear attributes
smartctl -A /dev/sdX | grep -iE "wear|reallocat|pending"

# Confirm which device is the SLOG and its state
zpool status <pool>

The logs section of zpool status shows the SLOG vdev, its state, and its error counters. Any non-ONLINE state or non-zero error counts on the log vdev is a ticket.

Estimating runway and setting a replacement threshold

SMART wear indicators are counters, not alerts. You have to trend them. The useful calculation is:

Days remaining = (Rated endurance - Lifetime writes) / Daily write rate

For Percentage_Used, the equivalent is: sample the value weekly, compute the percent consumed per week, and project the date you cross your replacement threshold. A SLOG on a busy NFS pool can consume several percent per month, which looks harmless until you realize the device will cross 70% inside a year.

Replacement thresholds:

  • Replace the SLOG at roughly 70% wear for production workloads. SLOG and L2ARC devices get a more aggressive threshold than data drives (commonly 80%) because their failure is sudden, their workload is concentrated, and the blast radius is every sync-writing application on the pool.
  • Trend the rate, not just the level. A device at 40% with a flat rate is fine. A device at 40% that added 15 points last quarter has a date with your maintenance window.
  • Mirror the SLOG. With a mirrored log vdev, one device hitting its wear limit is a routine replacement, not a sync-write outage risk. Create the mirror with zpool add <pool> log mirror <devA> <devB>.

Replacing a worn SLOG

Replacement is orderly when you do it before failure. SLOG removal through ZFS (zpool remove on the log device) is a clean, controlled operation; an endurance-exhaustion failure is not.

# 1. Confirm current log vdev layout
zpool status <pool>

# 2. Replace in place (preferred; keeps redundancy if mirrored)
zpool replace <pool> <old-slog-device> <new-slog-device>

# 3. Or remove the worn device and add the new one
zpool remove <pool> <old-slog-device>
zpool add <pool> log <new-slog-device>

Operational notes:

  • Do this during a normal maintenance window. During the swap, sync writes temporarily use the remaining mirror leg or fall back to the pool ZIL, so sync latency may rise briefly.
  • After replacement, reset your wear baseline for the new device and remove the old device’s wear data from your trending.
  • If wear accumulated faster than expected, reconsider the device class: the SLOG should be chosen for endurance, not just latency. A device that wore out in 18 months under your sync workload will do it again.

To reduce SLOG wear at the source, audit which datasets actually need synchronous semantics. sync=standard (the default) only involves the ZIL when the application requests it; sync=always forces every write through the ZIL and multiplies SLOG write volume. sync=disabled bypasses the ZIL entirely, but it sacrifices crash durability for acknowledged writes and is not appropriate for databases or anything that relies on fsync. Treat it as a workload decision, not a wear workaround.

Signals to watch in production

SignalWhy it mattersWarning sign
SMART Percentage_Used / wear attributeDirect measure of consumed endurance; the authoritative SLOG wear signalCrossing 70%, or climbing several points per month
Total bytes written vs rated TBWAbsolute endurance budget independent of SMART estimationTrending toward the manufacturer rating
Wear rate (percent per week/month)Drives runway estimation and replacement schedulingImplies replacement date inside the next maintenance window
Log vdev state in zpool statusSLOG failure does not degrade pool state; this is the only place it showsAny non-ONLINE state or non-zero errors on the log vdev
zil_commit_stall_count / zil_commit_error_countCommit stalls and errors point at ZIL/SLOG troubleIncrementing counters
Sync write latency (zpool iostat -l, syncq_wait)A worn or failed SLOG shows up as sync latency regression before anything elseSustained rise against baseline with reads unaffected

One correlation is worth memorizing: write latency rising while read latency stays normal, with the pool fully ONLINE, is the signature of a ZIL/SLOG problem, not a data disk problem.

How Netdata helps

  • Device-level SMART collection tracks Percentage_Used and wear attributes for the SLOG device continuously, so the wear trend exists as a time series instead of a quarterly surprise.
  • ZFS kstat collection covers the ZIL commit counters (zil_commit_stall_count, zil_commit_error_count) and the arcstats family, so you can watch ZIL health alongside device health. It also makes the L2ARC/SLOG distinction explicit, since l2_write_bytes lives in arcstats and belongs to the cache device.
  • Pool latency and queue metrics from zpool iostat let you correlate a sync-latency regression with log vdev state and confirm the fallback-to-pool-ZIL pattern.
  • The correlation that shortens diagnosis: sync write latency rising + log vdev errors + SMART wear near threshold means schedule a replacement. Sync latency rising + log vdev healthy + TXG sync times extending points at pool-side write pressure instead. Same symptom, different root cause, and the signal set separates them in minutes.