The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / lvm / lvm-thin-pool-queue-if-no-space-hang ▌

Operations Guides

LVM thin pool queue_if_no_space: why a full pool hangs instead of erroring

The server looks dead. Processes pile up in uninterruptible sleep, shells that touch certain filesystems never return, and even lvs hangs. There are no I/O errors in the application logs, no kernel oops, nothing in dmesg that screams hardware. Teams burn the first hour of this incident investigating a kernel or storage fault, when the actual cause is an LVM thin pool that ran out of data space.

The reason there are no errors is the pool’s when-full policy. The default for LVM thin pools is queue_if_no_space: when the pool fills, the kernel queues write I/O instead of failing it. The alternative, error_if_no_space, returns write errors immediately, which applications can see, log, and handle. Most operators have never checked which policy their pools use, because the default is silent until the pool fills.

This article covers how to confirm a hang is a full thin pool, how to recover, and how to audit and change the policy before it happens again.

What this means

A thin pool is a pair of internal LVs (data and metadata) that backs overcommitted thin volumes. When the data LV fills to 100%, the device-mapper thin-pool target cannot allocate blocks for new writes. What happens next is a policy decision baked into the dm table:

  • queue_if_no_space (the default): writes are queued in the kernel. Every process doing write I/O to any thin LV in the pool goes into D-state. Because all thin LVs share the pool, they all freeze at once. If the root filesystem, a journal, or swap sits on the pool, the whole system progressively locks up.
  • error_if_no_space: writes fail immediately with ENOSPC. Databases, VMs, and applications see real errors and can retry, abort, or alert.

There is a kernel-level safety valve: the no_space_timeout module parameter for dm_thin_pool, which defaults to 60 seconds. Queued writes that wait longer than the timeout are failed with errors. Even when a timeout is configured and fires, it does not save you: tens of seconds of frozen I/O already trips application timeouts, new writes keep re-queueing behind the failed ones, and the lvmthin documentation warns that tuning the timeout can lead to memory exhaustion, hung tasks, and deadlocks. Operationally, queue_if_no_space means “the system hangs.”

flowchart TD
  A[Thin pool data reaches 100%] --> B{When-full policy?}
  B -->|queue_if_no_space - default| C[Writes queue in kernel]
  C --> D[Processes enter D-state]
  D --> E[lvm tools hang on pool I/O]
  E --> F[Looks like kernel or hardware hang]
  B -->|error_if_no_space| G[Writes fail with ENOSPC]
  G --> H[Applications see errors and react]

The cruelest part is the observability gap. lvs, vgs, and pvs read metadata from the PVs and take LVM locks. When the pool is wedged, those commands hang too, so the standard diagnostic toolkit goes blind at exactly the wrong moment. dmsetup talks directly to the kernel and still works. That asymmetry is the key to fast diagnosis.

Common causes

CauseWhat it looks likeFirst thing to check
Thin pool data space at 100% with queue policyAll thin LV I/O frozen, D-state processes accumulating, no errors returneddmsetup status --target thin-pool used/total data blocks
VG has no free space, so the pool cannot be extendedPool at or near 100%, no headroom to grow intovgs -o vg_name,vg_free
Auto-extend expected but not actually activePool fills with no extension attempt; threshold 100 means disabledthin_pool_autoextend_threshold in /etc/lvm/lvm.conf
no_space_timeout disabled or set very highHang never resolves into errors, even after minutes/sys/module/dm_thin_pool/parameters/no_space_timeout
Not LVM at all: failing PV, multipath, or SAN lossSimilar D-state pileup but with I/O errors in dmesg and pool not fulldmesg, pvs for missing PVs
dm device stuck suspended from a resize or reloaddmsetup info shows suspended with no operation in progressdmsetup info -c -o name,suspended

Quick checks

All of these are read-only. Lead with dmsetup; it does not take LVM locks or read from disk.

# 1. Thin pool fullness and policy, straight from the kernel
dmsetup status --target thin-pool
# Fields: ... <used_meta>/<total_meta> <used_data>/<total_data> ... queue_if_no_space|error_if_no_space
# used_data/total_data at or near 1:1 confirms a full pool.

# 2. D-state processes, the visible symptom of queued I/O
ps -eo pid,stat,wchan:30,comm | awk '$2 ~ /D/'

# 3. Confirm a stuck process is waiting on dm/thin
cat /proc/<pid>/stack
# Look for device-mapper or thin-pool functions in the stack.

# 4. Any suspended dm devices
dmsetup info -c --noheadings -o name,suspended

# 5. Pool usage via LVM (may hang during an active incident)
lvs -o lv_name,vg_name,lv_attr,data_percent,metadata_percent
# 'D' in position 9 of lv_attr means the pool is out of data space.

# 6. Which when-full policy each pool uses
lvs -o lv_name,vg_name,lv_when_full

# 7. Recovery headroom: can the pool be extended at all
vgs -o vg_name,vg_size,vg_free

# 8. Rule out hardware as the cause of the hang
dmesg | grep -i 'I/O error' | tail -20

How to diagnose it

  1. Establish that the system is I/O-hung, not dead. If you still have a working shell, run ps and count D-state processes. A growing count of processes stuck for more than 60 seconds each is an active I/O blockage, not a CPU or memory event.

  2. Run dmsetup status --target thin-pool. If this returns immediately while lvs hangs, you are already looking at the answer: the LVM management plane is blocked on the same storage that is hung. Compare used_data to total_data. At or near 1:1, the pool is full.

  3. Read the policy flag in the same output. queue_if_no_space at the end of the status line explains the absence of errors. This is the moment the “server hung” ticket becomes an LVM capacity incident.

  4. Confirm the wait channel. For one or two D-state processes, read /proc/<pid>/stack and check for device-mapper or thin-pool frames. This distinguishes a thin pool freeze from an NFS hang or a dying disk, which can look identical from the ps output alone.

  5. Check dmesg for competing explanations. Block I/O errors, device resets, or path failures point at hardware or multipath instead. A full thin pool is quiet in the kernel log; a failing disk usually is not.

  6. Check recovery headroom before touching anything. vgs -o vg_free tells you whether lvextend on the pool can succeed. This determines which recovery path you take.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Thin pool data_percentThe countdown to the freezeAbove 85%, or growth projecting 100% within hours
dmsetup status used/total data blocksWorks when LVM tools hang; the incident-time source of truthRatio approaching 1:1
LV health (lv_attr position 9)D means the pool is already out of data spaceAny D on a thin pool
When-full policy per poolDetermines hang versus error at exhaustionqueue_if_no_space on production pools that could handle errors
D-state process count and durationThe user-visible hang symptomProcesses stuck over 60 seconds, count growing
Thin pool metadata_percentSecond, independent exhaustion pathAbove 75%
VG free spaceWhether recovery by extension is even possibleBelow 10%

Thin pool usage is updated by the kernel periodically and can lag reality by tens of seconds. Treat these as trend signals, not tripwires.

Fixes

Recover the hung pool

The only clean recovery is giving the pool more space. Everything else is a workaround.

# Extend the pool if the VG has free space
lvextend -L +50G vg0/thinpool

As soon as the pool has free data blocks, queued writes complete and D-state processes unblock on their own. D-state processes cannot be killed, not even with SIGKILL; freeing pool space is what releases them.

Caveats that matter during the incident:

  • LVM commands may hang because they read metadata from PVs involved in the freeze. Run recovery commands from a shell that has not touched the hung filesystems. If lvextend blocks, your remaining options narrow quickly.
  • If the VG is full, you must add capacity first: pvcreate a new device and vgextend the VG, then extend the pool. If no spare device exists, deleting an unneeded thin LV or old thin snapshots frees pool blocks directly.
  • Last resort: forced reboot (for example via echo b > /proc/sysrq-trigger or out-of-band power control). This is disruptive and risks filesystem corruption on every thin LV that had queued writes, and the pool will still be full after boot. Extend the pool before workloads restart and refill it.
  • After recovery, reclaim space: run fstrim on the thin LV filesystems so deleted data returns blocks to the pool, and review snapshots sharing the pool.

Switch the pool to error_if_no_space

For workloads that handle write errors sanely (most databases, queue systems, and anything with its own retry logic), failing fast is far better than hanging:

# Change the policy on an existing pool
lvchange --errorwhenfull y vg0/thinpool

# Verify
lvs -o lv_name,lv_when_full vg0/thinpool
dmsetup status --target thin-pool

To make new pools default to erroring, set activation/error_when_full = 1 in /etc/lvm/lvm.conf. The lvchange --errorwhenfull option exists in LVM since lvm2 2.02.115 (the RHEL 7.1 era); on older releases the policy can only be set through dmsetup table manipulation. The policy flag is part of the dm table, so lvchange --errorwhenfull applies it by reloading the pool table (a brief suspend/resume), not by deactivating the pool.

Understand the tradeoff before flipping the policy globally. With error_if_no_space, applications that do not handle ENOSPC well will crash or corrupt instead of hanging. A hang preserves the option of extending the pool and continuing as if nothing happened; an error forces the application to cope. Choose per pool, per workload.

There is no runtime dmsetup message switch for this policy on current kernels: the thin-pool target accepts messages only for operations like create_thin, delete, and set_transaction_id, so changing the policy requires a table reload (lvchange --errorwhenfull, or manual dmsetup table manipulation).

Consider the timeout, cautiously

no_space_timeout bounds how long queued writes wait before failing. Raising it gives you more time to extend the pool before errors hit applications; setting it to 0 disables the timeout and the queue becomes truly indefinite. The lvmthin documentation warns that disabling timeouts can cause memory exhaustion, hung tasks, and deadlocks. Treat this as a deliberate tuning decision, not a safety feature, and leave the default unless you have a specific reason.

Prevention

  • Audit the policy now, not during the incident. Run lvs -o lv_name,vg_name,lv_when_full on every host with thin pools and record the answer. This is a Level 4 maturity item in the LVM monitoring maturity model for a reason: most teams discover their default the hard way.
  • Alert on pool fill, not on the hang. data_percent above 85% is a same-shift ticket; above 95% is urgent. By the time D-state processes accumulate, you are in the incident.
  • Do not trust auto-extend blindly. The default thin_pool_autoextend_threshold is 100, which means disabled. See LVM thin pool auto-extend not working: threshold 100 means disabled. Even when configured, auto-extend fails silently if dmeventd is down or the VG is full.
  • Keep VG headroom. The pool can only be extended into free VG space. Track vg_free and its runway; see LVM volume group running low on free space.
  • Monitor D-state processes as a corroborating signal. A rising count of processes stuck on dm devices turns “pool at 97%” into “pool freezing workloads now,” which is the difference between a ticket and a page.
  • Use dmsetup status in your monitoring path. LVM tools take locks and perform I/O; they hang during the exact incidents you most need to observe. Kernel-side status does not.

How Netdata helps

  • Netdata collects per-dm-device I/O from the kernel block layer, so inflight I/O and latency on thin pool devices are visible even when LVM user-space tools are hung.
  • Process state tracking surfaces D-state accumulation, letting you correlate “processes stuck on I/O” with the pool filling rather than chasing a phantom kernel bug.
  • Disk space and block device trends give you the data_percent trajectory and growth rate, so you can alert on runway (hours to full) instead of a static threshold.
  • Because collection is per-second and local, the moments right before the freeze (latency spikes from reclaim activity, inflight I/O climbing) are captured rather than averaged away.
  • Correlating dm device saturation, D-state count, and pool fill on one dashboard compresses the diagnosis from “why is the server hung” to “pool full, queue policy, extend now” in one screen.