The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / lvm / lvm-snapshot-origin-latency-cow-overhead ▌

Operations Guides

LVM snapshot slowing the origin: copy-on-write write amplification

Write latency on a logical volume doubles or worse with no obvious cause. The disks are healthy, dmesg shows no I/O errors, no RAID resync is running, and the workload has not changed. The one thing that did change: someone created a snapshot, often hours or days ago, usually for a backup that has long since finished.

This is a traditional (thick) LVM snapshot working as designed. While the snapshot exists, the first write to every chunk of the origin triggers a copy-on-write (COW) cycle: read the original chunk, write it to the snapshot’s exception store, then write the new data. One application write becomes a minimum of three physical I/Os. Under write-heavy load, origin latency doubles or worse, and it degrades further as the exception store fills.

The tell is the timeline: latency spikes the moment a snapshot appears and recovers the instant it is removed or invalidates. If latency steps up at snapshot creation and steps back down at removal, you have your answer.

What this means

A traditional LVM snapshot is not a copy. It is a small pre-allocated exception store plus a rule in the kernel’s device-mapper layer: “before you overwrite any chunk of the origin for the first time since I was created, save the old chunk to me first.” That is what lets the snapshot present a point-in-time view of the origin. The cost is paid by the origin, not the snapshot:

  • First write to a chunk: read old data, write old data to the exception store, write new data. Three I/Os instead of one.
  • Repeat writes to an already-copied chunk: no COW, normal write path. The penalty concentrates on first-touch writes; reads on the origin are largely unaffected.
  • Every additional snapshot on the same origin adds its own COW copy. Each first-touch write triggers COW on all snapshots, so stacked snapshots compound the overhead.
  • The exception store receives scattered writes. If it sits on the same PV as the origin (the default), snapshot and origin traffic compete for the same disk.

Only LVs with active snapshots are affected. If one LV is slow and its neighbors on the same disks are fine, that asymmetry is itself a strong hint.

Thin snapshots are a different mechanism: they share the thin pool and have no fixed per-snapshot exception store, so they do not impose this specific COW penalty. Their failure mode is pool exhaustion, which freezes every thin LV in the pool at once. That is covered in LVM thin pool out of data space. For the full layering picture, see How LVM actually works in production.

flowchart LR
  A["App write to origin LV"] --> B{"Active thick snapshot?"}
  B -->|"no"| C["Write new data: 1 IO"]
  B -->|"yes, first write to this chunk"| D["Read old chunk"]
  D --> E["Write old chunk to exception store"]
  E --> F["Write new data"]
  F --> G["Total: 3 IOs minimum"]
  B -->|"yes, chunk already copied"| C
  H["Each extra snapshot on the origin"] -. "adds one more COW copy" .-> E

Common causes

CauseWhat it looks likeFirst thing to check
Backup snapshot left behind after the backup finishedsnap_percent creeping up for days, latency elevated the whole timelvs -o lv_name,origin,snap_percent,lv_time against backup job logs
Multiple snapshots on one originLatency far worse than 2x, stepping up with each snapshot addedCount snapshots per origin in lvs output
High-churn origin with an undersized exception storesnap_percent climbing fast, latency worsening as it fillssnap_percent sampled twice, minutes apart
Exception store on the same PV as the originOne PV saturated while sibling PVs are idlelvs -o lv_name,seg_pe_ranges plus per-disk utilization
Unexpected write burst on the origin (deploy, batch job, reindex)snap_percent jumps, latency spike starts with the burstApplication write throughput versus its baseline

Quick checks

All of these are read-only and safe to run during an incident.

# 1. List all snapshots, their origins, and how full they are
lvs -o lv_name,vg_name,origin,snap_percent,lv_attr -S 'seg_type=snapshot'

Any LV listed here is an active thick snapshot. Empty output means no thick snapshots exist and this article does not apply.

# 2. Check for already-invalidated snapshots
# An invalid snapshot shows 'I' (uppercase) in its lv_attr state field
# (position 5, per lvs(8)).
lvs -o lv_name,lv_attr -S 'seg_type=snapshot' --noheadings

An invalid snapshot has already lost its data; the kernel skips invalid snapshots, so origin overhead disappears when it invalidates. An invalid snapshot plus a latency graph that already recovered means the performance incident is over but the restore point is gone.

# 3. Map LV names to dm devices
dmsetup ls
# 4. Measure write latency on the origin's dm device (two samples, 10s apart)
grep ' dm-' /proc/diskstats > /tmp/dm.a
sleep 10
grep ' dm-' /proc/diskstats > /tmp/dm.b
awk 'NR==FNR {w[$3]=$8; m[$3]=$11; next}
     ($3 in w) && ($8 > w[$3]) {
       printf "%-14s writes=%6d  avg_write_ms=%.2f\n", $3, $8-w[$3], ($11-m[$3])/($8-w[$3])
     }' /tmp/dm.a /tmp/dm.b

/proc/diskstats fields are cumulative counters, so you need a delta between samples: average write latency is milliseconds-spent-writing divided by writes completed. If sysstat is installed, iostat -x 10 2 on the dm device gives the same answer with less awk.

# 5. Check queueing on the origin right now (ios_in_progress)
awk '$3 ~ /^dm-/ {print $3, "inflight="$12}' /proc/diskstats

Sustained inflight counts well above baseline mean writes are queueing behind the COW cycle.

# 6. Find which PVs hold the snapshot's exception store
lvs -o lv_name,vg_name,seg_pe_ranges -S 'seg_type=snapshot'

If those ranges sit on the same PV as the origin, COW traffic is hammering the same disk the application is writing to.

# 7. Look for invalidation events in the kernel log
dmesg | grep -i 'invalidating snapshot'

“Invalidating snapshot: Unable to allocate exception” means the exception store hit 100% and the snapshot is dead.

# 8. Infer the amplification factor: PV write volume versus origin write volume
iostat -x 10 2

Device-mapper exposes no direct amplification metric; you infer it by comparing I/O at the origin dm device against the underlying PV. A PV doing roughly 3x the origin’s write throughput while a snapshot is active is the mechanism showing up in numbers.

One caveat: lvs and friends take LVM metadata locks and can hang on a badly stuck system. If they do, fall back to dmsetup status, which reads from kernel memory.

How to diagnose it

  1. Establish the symptom. From check 4, write latency on the origin’s dm device is 2x or more above its normal baseline, sustained for minutes. The penalty lands on writes; reads stay close to normal.
  2. List the snapshots on that origin (check 1). Exactly one LV with origin pointing at your slow LV is the common case.
  3. Correlate the timeline. Compare the snapshot’s creation time with the moment latency stepped up. This correlation is the distinguishing feature of the failure. If latency started before the snapshot existed, keep looking: check thin pool usage (lvs -o lv_name,data_percent,metadata_percent), mirror resync progress (copy_percent), and dmesg for device errors.
  4. Check for compounding. More than one snapshot on the origin, or an exception store sharing a PV with the origin, makes the same mechanism worse.
  5. Check the snap_percent trajectory. Fast growth means the snapshot is also racing toward invalidation, which is a second incident (a lost restore point) stacked on the performance one.
  6. Confirm by removal, if operations allow it. Removing the snapshot is both the fix and the proof: origin latency returns to baseline immediately.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
snap_percent per snapshotCountdown to invalidation; growth rate reflects origin churnAbove 70% and rising; jumps of several percent between samples are normal because allocation happens in chunks
Origin dm write latencyThe user-facing symptom2x or more over baseline, sustained longer than 5 minutes
Snapshot count and age per originEach extra snapshot adds a COW copy; old snapshots are usually forgotten onesMore than one snapshot per origin; any snapshot older than 24 hours on a high-churn LV
Inflight I/O on the origin dm deviceShows queueing from serialized COW workSustained elevation versus baseline
Utilization of the PV holding the exception storeSame-disk contention between COW writes and origin writesOne PV near saturation while siblings are idle
Kernel invalidation messagesThe snapshot died; latency recovers but the restore point is gone“Invalidating snapshot” in dmesg

Fixes

If the backup is done: remove the snapshot

# WARNING: irreversible. The point-in-time view is gone.
lvremove <vg>/<snapshot>

This is the definitive fix. Origin latency returns to baseline immediately and the exception store’s space is freed at once. Two checks first: confirm the backup that used the snapshot actually completed and is restorable, and confirm nothing is still reading the snapshot (a mounted snapshot, an in-flight copy job).

If the snapshot must live longer: extend it

# Grow the exception store to avoid invalidation
lvextend -L +<size>G <vg>/<snapshot>

This buys protection against overflow and nothing more. It does not reduce COW overhead; latency stays elevated for as long as the snapshot exists. Use it when a backup is still running and the snapshot is trending toward 100%. Requires free extents in the VG.

If the snapshot exists for rollback: merge it

# WARNING: destructive. The origin reverts to the snapshot's point-in-time state.
# All data written to the origin since the snapshot was taken is lost.
lvconvert --merge <vg>/<snapshot>

Merging restores the origin to the snapshot’s point-in-time state and eliminates the COW overhead once the merge completes. It discards everything written to the origin since the snapshot was taken, and it requires the snapshot to still be valid. Only choose this when rollback was the intent all along.

If several snapshots are stacked on the origin

Remove all but the one you actually need. There is no way to make stacked thick snapshots cheap: every first write pays the COW penalty once per snapshot. If your backup tooling creates a new snapshot without deleting the previous one, fix the tooling.

Recreate with better placement and chunk size

If you snapshot write-heavy origins regularly:

  • Place the exception store on a different PV than the origin at creation time (pass the target PV as a positional argument to lvcreate) so COW traffic does not fight the origin for the same disk.
  • Create with a larger chunk size (lvcreate -s -c). The default is small, 4 KiB on many distributions, which produces many small COW operations under random write loads. Larger chunks mean fewer, bigger copies; the tradeoff is that each first-touch write copies more data if the application writes in small blocks. Chunk size is fixed at creation, so changing it means removing and recreating the snapshot.

Long-term: move snapshot-heavy workloads to thin provisioning

Thin snapshots share the pool and carry no fixed exception store, so they avoid this origin-side COW penalty and the cliff-edge invalidation. The risk moves to the pool: a full thin pool freezes every thin LV in it, so pool data and metadata monitoring become the critical signals instead. Migration means recreating the volumes, so plan it as maintenance, not as an incident fix.

Prevention

  • Treat snapshots as short-lived objects: create at backup start, remove at backup end, and alert on anything older than 24 hours on a busy origin.
  • Alert on snap_percent at 80% and treat 95% as urgent. Overflow is instant and irreversible.
  • Keep one snapshot per origin. Make stacking a deliberate, documented exception.
  • Size the exception store at least 2x the total write volume expected during the snapshot’s lifetime.
  • Keep COW traffic off the origin’s PV where the layout allows it.
  • Baseline per-LV write latency so a 2x step-change is an alert rather than a user complaint. Alert on deviation from baseline, not absolute thresholds.
  • If snapshots are part of the daily workflow, evaluate thin provisioning and monitor the pool instead.

How Netdata helps

  • Netdata charts per-dm-device write latency and throughput from /proc/diskstats at per-second resolution, computing the deltas for you. The step-change at snapshot creation and the instant recovery at removal both show up on the same graph.
  • Overlaying the origin LV’s write throughput against the underlying PV’s makes the roughly 3x amplification visible, which is as close as you can get to a COW metric since device-mapper does not expose one.
  • Per-PV disk utilization surfaces the same-disk contention case without manual iostat work.
  • Snapshot usage and age alerting catches snap_percent climbing toward invalidation while you are focused on the latency incident, so the restore point is not silently lost in the background.
  • Retained history lets you line the latency step up against change records (snapshot cron entries, backup jobs) during the postmortem. LVM itself keeps no history; without time-series data this correlation is guesswork.