The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / zfs / zfs-l2arc-ineffective ▌

Operations Guides

ZFS L2ARC ineffective: a cache that burns SSD endurance for nothing

You added an SSD as an L2ARC device expecting faster reads. Months later, read latency has not moved, the SSD’s wear indicator is climbing, and the ARC is smaller than it should be. The L2ARC is being written to constantly and read from almost never. You are paying for the cache in RAM and drive endurance and getting nothing back.

This is one of the most common ZFS mistakes: L2ARC is assumed to be beneficial by default, and almost nobody validates it after adding it. Many workloads get zero benefit from it, and for those workloads the device is a net negative: it consumes ARC memory for its index headers and burns SSD write endurance caching blocks that are never read again.

This article covers how L2ARC actually works, how to measure whether yours is earning its keep, and how to remove it cleanly when it is not.

What this means

L2ARC is not a general-purpose read cache. Three design properties explain most “ineffective L2ARC” situations:

It focuses on ARC candidates, not first reads. L2ARC does not cache data on first read. In the traditional mode (total cache capacity below twice arc_c_max), a block must be read into the ARC and evicted before L2ARC writes it. In current OpenZFS, a sufficiently large L2ARC (at least twice arc_c_max) can use persistent markers to cache content before eviction, but still follows ARC recency and feed-rate limits. If your working set fits in the ARC, there are few useful candidates, and the L2ARC sits mostly empty. If your workload is streaming (backups, media, analytics scans), blocks are read once and never again, so caching them is pure waste.

It only helps on the second access. The first read of any block always goes to disk. The second read only hits L2ARC if the block was evicted from the ARC in between. Workloads without that revisit pattern see no hits at all.

Its index lives in your RAM. Every block cached on L2ARC costs ARC memory for an L2-only header; the exact size depends on OpenZFS release and platform ABI, with operator estimates historically ranging from about 70 to 170 bytes. Current OpenZFS allocates an L2-only prefix of the ARC header and reports the actual total in l2_hdr_size, so measure it rather than assuming the historical constant. A large L2ARC device can consume gigabytes of ARC just for the index, and that RAM comes directly out of what the ARC could use for actual cached data. A low-hit L2ARC makes the ARC smaller and less effective while providing nothing in return: a double loss.

flowchart TD
  A[Application read] --> B{In ARC?}
  B -->|yes| C[Served from RAM]
  B -->|no| D{In L2ARC?}
  D -->|yes| E[Served from SSD]
  D -->|no| F[Read from pool disks]
  F --> G[Block enters ARC]
  G --> H[ARC eviction under pressure]
  H --> I[L2ARC feed writes block to SSD]
  I --> J[Header costs ARC RAM per block]
  J -.->|never re-read| K[SSD endurance spent, zero hits]

Without persistent L2ARC (OpenZFS 2.0+, enabled by default via l2arc_rebuild_enabled=1), the L2ARC is also completely cold after every reboot. Even with persistence, the index is rebuilt on pool import and the cache takes time to become useful again. If you reboot frequently, you may pay the fill cost over and over without ever reaching steady state.

Common causes

CauseWhat it looks likeFirst thing to check
Working set fits in ARCL2ARC nearly empty, near-zero hits, ARC hit ratio already highARC hit ratio and eviction rate in arcstats
Streaming or scan-heavy workloadHigh l2_write_bytes, near-zero l2_read_bytes, blocks read once and never againCompare l2_read_bytes vs l2_write_bytes over time
Working set far larger than ARC plus L2ARCHigh miss rates everywhere, L2ARC churns constantlyl2_hits vs l2_misses ratio
Frequent reboots without persistenceL2ARC cold after every boot, hit ratio never climbsUptime vs L2ARC hit ratio trend
L2ARC headers eating ARC RAMl2_hdr_size is a large fraction of ARC size, ARC data hit ratio degradedl2_hdr_size vs ARC size in arcstats
Cache device too small or too slow to matterL2ARC fills and wraps without capturing the hot setl2_size and l2_asize vs workload working set

Quick checks

All of these are read-only and safe to run on a production system.

# Pull the L2ARC-relevant counters from arcstats
awk '/^l2_hits|^l2_misses|^l2_size|^l2_asize|^l2_hdr_size|^l2_read_bytes|^l2_write_bytes|^l2_io_error|^l2_cksum_bad/ {print $1, $3}' /proc/spl/kstat/zfs/arcstats
# Compute the L2ARC hit ratio
awk '/^l2_hits/ {h=$3} /^l2_misses/ {m=$3} END {if (h+m>0) printf "L2ARC hit ratio: %.2f%%\n", h/(h+m)*100; else print "No L2ARC activity"}' /proc/spl/kstat/zfs/arcstats
# Compare header RAM cost against total ARC size
awk '/^l2_hdr_size/ {hdr=$3} /^size/ {s=$3} END {printf "l2_hdr_size: %.1f MiB, ARC size: %.1f GiB, headers = %.2f%% of ARC\n", hdr/1048576, s/1073741824, hdr/s*100}' /proc/spl/kstat/zfs/arcstats
# Check whether L2ARC writes vastly exceed reads (burning endurance for nothing)
awk '/^l2_read_bytes/ {r=$3} /^l2_write_bytes/ {w=$3} END {printf "l2_read_bytes: %.1f GiB, l2_write_bytes: %.1f GiB, read:write ratio 1:%.1f\n", r/1073741824, w/1073741824, (r>0?w/r:w)}' /proc/spl/kstat/zfs/arcstats
# Confirm the cache device state and any device-level errors
zpool status -v
# Check SSD wear on the cache device (adjust device name)
smartctl -A /dev/sdX

The arcstats counters are cumulative since boot, so a hit ratio computed from them hides recent behavior. Take two samples minutes or hours apart and compute the ratio over the delta to see what the L2ARC is doing now, not what it did last month.

How to diagnose it

  1. Establish the interval hit ratio. Sample l2_hits and l2_misses twice, at least an hour apart during representative load, and compute delta_hits / (delta_hits + delta_misses). Sustained ratios below roughly 10 to 30 percent mean the L2ARC is not earning its overhead. The decision threshold is workload-dependent, but if you are consistently at the low end of that range, the cache is decorative.

  2. Check the read-to-write balance. Compare l2_read_bytes to l2_write_bytes over the same interval. Writes massively exceeding reads is the signature of the L2ARC burnout pattern: the feed thread is streaming data onto the SSD that nobody reads back, and every one of those bytes consumes drive endurance.

  3. Quantify the RAM cost. Divide l2_hdr_size by the ARC size. Header overhead should stay well under a tenth of the ARC. If headers are consuming 5 to 10 percent of the ARC while the hit ratio is low, the L2ARC is a net negative: removing it gives that RAM back to the ARC, which is a strictly better cache.

  4. Check whether the workload can ever benefit. Ask whether the same blocks are re-read after ARC eviction. Backups, log archival, video streaming, and analytics scans are read-once workloads. Random re-reads of a working set modestly larger than RAM are where L2ARC helps. If your workload is the former, no tuning will fix it.

  5. Factor in reboot frequency. If the system reboots often and you are not on OpenZFS 2.0 or later (or persistence is disabled), the cache restarts cold every boot and may never warm up before the next reboot. Check uptime against your hit ratio trend.

  6. Rule out the opposite problem first. Before blaming L2ARC for slow reads, confirm the ARC itself is healthy. A low ARC hit ratio caused by a shrunken or undersized ARC is a different problem with a different fix. See the related guides below.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
l2_hits / (l2_hits + l2_misses)Direct measure of L2ARC effectivenessSustained below 10-30% on interval deltas
l2_write_bytes vs l2_read_bytesEndurance spend vs benefit deliveredWrites far exceed reads over long intervals
l2_hdr_size vs ARC sizeRAM tax the cache imposes on the ARCHeaders approaching a tenth of ARC with low hits
l2_size / l2_asizeHow much of the device is actually usedDevice full and churning with low hit ratio
l2_io_error, l2_cksum_badCache device health (tolerated, but degrades value)Any sustained increment
SSD SMART wear indicatorsEndurance consumed by L2ARC feed writesWear climbing faster than the rest of your fleet
ARC hit ratioThe cache that actually mattersDropping while l2_hdr_size grows

Fixes

Remove the L2ARC device

If the interval hit ratio is sustained below roughly 10 to 30 percent and headers are consuming meaningful ARC, the correct fix is removal:

# Identify the cache device first with zpool status, then remove it
zpool remove <pool> <cache-device>

This is safe: L2ARC contains only cached copies, never unique data. Removal is orderly, and the header RAM returns to the ARC immediately. Reads that would have hit L2ARC fall back to the pool disks, which for a low-hit cache is barely measurable. It is not disruptive, but a maintenance window costs you nothing if you want to be conservative.

After removal, watch the ARC hit ratio for a few days. It should improve as the freed header RAM is put to use.

Keep it, but only if the numbers justify it

If the interval hit ratio is comfortably above 30 percent and l2_hdr_size is a small fraction of the ARC, the L2ARC is working. Leave it alone. Keep trending the hit ratio and SSD wear, because workload drift can turn a good L2ARC into a bad one over time.

Reduce the RAM tax and write burn if you keep it

  • Right-size the device. Header cost scales with cached blocks, so a smaller cache device means fewer headers. Estimate the header RAM from device size and your dataset recordsize before you ever attach the device.
  • Restrict what gets cached. The secondarycache dataset property controls what L2ARC stores per dataset: all, metadata, or none. Setting secondarycache=none on backup, archive, or other scan-heavy datasets stops one-time-read data from ever entering the feed, cutting both SSD wear and header overhead while keeping the cache for workloads that benefit.
  • Throttle the feed. L2ARC write speed is bounded by the l2arc_write_max and l2arc_write_boost tunables. Lowering the write rate reduces endurance consumption at the cost of slower cache warmup. The default rose from 8 MiB to 32 MiB per second in OpenZFS 2.2.7; check the installed release before assuming a value.
  • Enable persistence if you reboot. On OpenZFS 2.0 and later, l2arc_rebuild_enabled=1 (the default) rebuilds the L2ARC index on pool import so the cache survives reboots. On older versions, every reboot is a cold start.

Do not confuse this with an ARC problem

If your real symptom is a low ARC hit ratio, adding or keeping L2ARC is usually the wrong response. The better levers are raising zfs_arc_max if the ARC is artificially capped, adding RAM, or reducing memory competition from applications. L2ARC is a second-tier cache with a RAM tax; the first tier is almost always where the win is.

Prevention

  • Validate before and after adding L2ARC. Estimate header RAM cost up front, and set a review checkpoint after deployment: if the interval hit ratio is below your threshold after a few weeks of representative load, remove it.
  • Trend, do not snapshot. Point-in-time zpool status and cumulative arcstats counters hide drift. Export l2_hits, l2_misses, l2_hdr_size, l2_read_bytes, and l2_write_bytes to a time-series system so you can compute interval ratios and spot workload changes.
  • Monitor cache device wear like a SLOG. L2ARC and SLOG devices take concentrated write load. Track SSD SMART or NVMe wear indicators and plan replacement before exhaustion, same as you would for a log device.
  • Set secondarycache deliberately on scan-heavy datasets rather than letting one-time-read data flow into the cache by default.
  • Cap the ARC explicitly. An uncapped ARC plus L2ARC headers plus applications is how you end up in memory-pressure incidents. See the related guides on zfs_arc_max and ARC memory behavior.

How Netdata helps

  • Netdata collects the ZFS ARC statistics from /proc/spl/kstat/zfs/arcstats per second, including l2_hits, l2_misses, l2_hdr_size, l2_read_bytes, and l2_write_bytes, so you see the L2ARC hit ratio and read/write balance as trends rather than cumulative snapshots.
  • Correlating L2ARC hits against ARC hits and pool disk read IOPS shows whether the L2ARC is actually absorbing reads that would otherwise hit the data vdevs, or whether disk reads are unchanged while the SSD fills.
  • Charting l2_hdr_size next to ARC size makes the RAM tax visible: you can see header growth coinciding with ARC data hit ratio decline.
  • Per-second L2ARC write rates combined with disk-level metrics and device wear tracking let you quantify endurance spend per day and project SSD lifetime under the current feed rate.
  • Anomaly detection on the L2ARC hit ratio catches workload drift, such as a new batch job that turned a healthy cache into a write-only endurance sink, without manual threshold babysitting.