The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / lvm / lvm-thin-pool-needs-repair ▌

Operations Guides

LVM thin pool check needed: thin_check and lvconvert --repair

A thin pool that reports “check needed” is telling you the kernel no longer trusts the pool’s metadata B-tree. The pool may still be serving I/O, or it may already be erroring writes. Either way, the wrong move (resizing metadata, forcing activation, running the wrong repair path on an oversized metadata LV) can turn a recoverable state into permanent data loss.

This state usually appears after metadata exhaustion, an unclean shutdown during a metadata update, or an underlying storage error. It shows up in lvs output as c (check needed) or C (check needed, suspended) in position 5 of lv_attr, often alongside a health flag in position 9. If you have not read the background on how thin pools are structured, see how LVM actually works in production first. This page covers only detection, validation, and repair.

What this means

Every thin pool is backed by two internal LVs: a data LV and a much smaller metadata LV. The metadata LV holds the block mapping table (a B-tree) for every thin volume and snapshot in the pool. When the device-mapper thin target hits a metadata problem it cannot resolve online, it sets the needs_check flag in the pool superblock and surfaces it to LVM. The kernel’s thin-provisioning documentation is blunt about the consequence: if metadata space is exhausted or a metadata operation fails, the pool errors I/O until the pool is taken offline and repair is performed.

Two things follow from that:

  • Repair is an offline operation. You cannot fix a flagged pool while it is active. lvconvert --repair requires the pool to be deactivated first.
  • Repair is not guaranteed. The lvmthin(7) man page warns that data may be unrecoverable. If the device details tree or the data mapping tree itself is damaged, thin_repair may produce incomplete output. Treat backups as the real recovery path and repair as the fast path.

“Check needed” is frequently what comes after metadata exhaustion, which is why thin pool metadata full is the saturation state worth alerting on first.

Common causes

CauseWhat it looks likeFirst thing to check
Metadata exhaustionmetadata_percent hit 100%, health flag set, data% may be moderatelvs -o lv_name,data_percent,metadata_percent,lv_attr
Unclean shutdown or crash during metadata updateFlag appears after power loss or kernel panicjournalctl -k around the crash window
Underlying device errors on the metadata LVI/O errors in dmesg for the PV hosting _tmetadmesg | grep -i 'I/O error' and pvs
Failed metadata operation under loadRandom-write-heavy workload, many snapshots, metadata climbing for weeksMetadata growth trend; snapshot count in lvs -o lv_name,origin

The insidious case is metadata exhaustion with low data usage. Operators watching only data_percent never see it coming. The monitoring checklist covers why both percentages need separate alerting.

Quick checks

All of these are read-only and safe to run during an incident. Prefer dmsetup over lvs if the system is under I/O stress: lvs takes LVM locks and reads metadata from disk, and can hang on an unhealthy pool. dmsetup reads kernel state directly.

# Find pools flagged check needed (position 5 = c or C) and their health (position 9)
lvs -a -o lv_name,vg_name,lv_attr,data_percent,metadata_percent

# Kernel view of the pool: used/total metadata and data blocks, ro/rw, queue vs error policy
dmsetup status --target thin-pool

# Confirm which devices are actually active and whether any are suspended
dmsetup info -c -o name,attr,suspended | grep -i <vg>

# Kernel log evidence: dm-thin metadata errors, I/O errors on backing devices
dmesg | grep -iE 'dm-thin|thin|metadata|I/O error' | tail -50

Interpreting what you see:

  • c in position 5 with the pool still active: the flag is set but the pool has not yet failed. Plan a repair window before it gets worse.
  • C (check needed, suspended): the pool is already suspended and I/O is blocked.
  • M in position 9: metadata has gone read-only. The pool is in emergency mode.
  • F in position 9: the pool has failed. Writes are erroring.

Note: lv_attr position 9 uses uppercase for these thin pool health flags: (F)ailed, out of (D)ata space, and (M)etadata read only, per lvs(8).

How to diagnose it

  1. Identify the pool and its metadata LV. Run lvs -a -o lv_name,vg_name,lv_attr,lv_size,data_percent,metadata_percent and note the pool name and the size of its hidden [pool_tmeta] LV. You need the metadata LV size later; it determines which repair path is safe.

  2. Check kernel status without LVM locks. dmsetup status <vg>-<pool>-tpool (or dmsetup status --target thin-pool) shows used_metadata/total_metadata and used_data/total_data, plus whether the pool is read-only and whether it queues or errors on no space.

  3. Validate the metadata. If the pool is active and you cannot take it down yet, run thin_check --metadata-snap /dev/mapper/<vg>-<pool>_tmeta to check a metadata snapshot of the live pool. Running thin_check directly against the live metadata device fails with “Device or resource busy”. thin_check comes from the device-mapper-persistent-data package (also shipped as thin-provisioning-tools on some distributions). For a faster but less thorough pass, --skip-mappings skips the per-device mapping trees.

  4. Assess blast radius. List the thin volumes and snapshots in the pool, check which are mounted or in use, and confirm which ones have current backups. This decides how aggressive you can be.

flowchart TD
  A[Pool shows check needed] --> B{Pool active?}
  B -->|yes| C[thin_check --metadata-snap to validate]
  B -->|no| D[Confirm backups exist]
  C --> D
  D --> E[Unmount thin LVs, lvchange -an VG/pool]
  E --> F{Metadata LV over 15.81 GiB?}
  F -->|yes| G[Manual path: thin_dump/restore to smaller LV]
  F -->|no| H[lvconvert --repair VG/pool]
  H -->|success| I[Extend metadata, activate, fsck thin LVs]
  H -->|failure| G
  G -->|failure| J[Restore from backup]

Metrics and signals to monitor

SignalWhy it mattersWarning sign
metadata_percentThe leading indicator; exhaustion is the most common path to check neededAbove 75%, or any sustained upward trend
lv_attr position 5Direct surface of the needs_check flagc or C on any pool
lv_attr position 9 (health)Distinguishes degraded from failedAny value other than -
dmsetup status metadata ratioWorks when lvs hangs; kernel-side truthused/total metadata approaching 1:1
D-state processes on dm devicesUser-visible symptom of a pool erroring or queuing I/OCount growing over 60 seconds
Kernel log dm-thin messagesFirst place metadata operation failures appearAny dm-thin error or read-only transition

Fixes

Standard repair: lvconvert –repair

This is the supported path on a pool whose metadata LV is within the kernel’s size limit.

  1. Stop everything using the pool. Unmount filesystems on thin LVs, stop VMs or containers backed by them, and deactivate the thin LVs, then the pool:
# Deactivate thin LVs first, then the pool itself
lvchange -an <vg>/<thin_lv>
lvchange -an <vg>/<pool>

lvconvert --repair refuses to run on an active pool. Verify deactivation with dmsetup ls; a pool that still shows active mappings was not actually deactivated. If lvchange -an appears to succeed but mappings remain, re-run with -v and check for leftover device-mapper entries or processes still holding the devices open.

  1. Run the repair:
# Repair thin pool metadata (offline, destructive class of operation)
lvconvert --repair <vg>/<pool>

This runs thin_repair against the damaged metadata LV and writes the repaired copy to the VG’s pmspare LV. On success, the pmspare becomes the new metadata LV and the old damaged one is renamed to <pool>_meta<N>. Two prerequisites bite in practice: the pmspare LV must be at least as large as the damaged metadata LV (extend it or create a replacement first if not), and the VG needs free space for the swap.

  1. Extend metadata before reactivating. Whatever filled the metadata area will fill it again. Grow it now:
# Grow pool metadata while the pool is still inactive
lvextend --poolmetadatasize +<size>G <vg>/<pool>

Do not attempt this on a pool that still carries the needs_check flag. The kernel refuses to resize metadata until repair is done, and on a corrupted pool the resize attempt can hang. Repair first, extend second, in that order.

  1. Activate and check the upper layers. The kernel documentation strongly recommends running consistency checks (fsck) on the thin volumes after any pool repair. Metadata repair can leave individual block mappings inconsistent even when the pool as a whole validates.

The oversized-metadata trap

There is a hard kernel limit on thin pool metadata size: 4145152 4K blocks, about 15.81 GiB. On older toolchains (RHEL 7 era lvm2 and device-mapper-persistent-data, and Ubuntu 16.04’s lvm2 2.02.133), lvconvert --repair on a metadata LV larger than that writes a repaired superblock claiming the larger size (4161600 blocks has been observed). The kernel then refuses to activate the pool with an error along the lines of “metadata device (4145152 blocks) too small: expected 4161600”. Red Hat Bug 2028639 tracks this and was closed WONTFIX for RHEL 7.

Modern LVM2 (2.03.x) added the allocation/thin_pool_crop_metadata setting in lvm.conf to crop metadata to 15.81 GiB for backward compatibility, confirming that the kernel limit persists; however, whether lvconvert --repair itself now handles oversized metadata correctly in all 2.03.x releases is not confirmed from available sources.

If your metadata LV is at or above 16 GiB, or you are on an affected toolchain, skip lvconvert --repair and take the manual path below. Newer LVM versions also offer thin_pool_crop_metadata in lvm.conf, which crops metadata size to 15.81 GiB for compatibility; enable it only when you actually need to interoperate with older toolchains.

Manual repair path

When lvconvert --repair fails or is unsafe:

  1. Create a new, empty metadata LV, sized under the 15.81 GiB kernel limit.
  2. Run thin_repair -i /dev/<vg>/<pool>_tmeta -o /dev/<vg>/<new_meta> to rebuild the mapping tree into the new LV.
  3. Verify the result: thin_check /dev/<vg>/<new_meta>.
  4. Swap it in: lvconvert --thinpool <vg>/<pool> --poolmetadata <vg>/<new_meta>.

On very old toolchains with the size bug, the equivalent workaround is thin_dump from the damaged metadata, thin_restore into a freshly created smaller metadata LV, then the same swap. thin_check --clear-needs-check-flag exists to clear the flag after a successful check, but only use it when you have verified the metadata is actually clean; clearing the flag on broken metadata just hides the warning. The --auto-repair option fixes only trivial issues such as metadata leaks; it is not a substitute for thin_repair on structural damage.

If all repair paths fail, restore from backup. That is the outcome the man page warns about, and it is why the pool’s health flags deserve paging-level attention before they ever reach this state.

Prevention

  • Alert on metadata_percent early. Ticket at 75%, urgent at 90%. Metadata is cheap; over-provision it aggressively at pool creation time rather than relying on extension later. See the metadata exhaustion guide for sizing rationale.
  • Never let the pool get here through data exhaustion either. A pool hitting out of data space under queue_if_no_space hangs I/O for every thin volume at once, and the failed-write churn also consumes metadata.
  • Verify auto-extend actually works. The default thin_pool_autoextend_threshold of 100 means disabled. See thin pool auto-extend not working and the low water mark warning that fires before the freeze.
  • Keep VG free space for the repair path. lvconvert --repair needs a usable pmspare and room to swap metadata LVs. A VG at zero free extents removes your repair options at the worst moment; see volume group free space low and insufficient free extents.
  • Track the maturity basics. The LVM monitoring maturity model puts metadata percent, health flags, and dmeventd state at Level 2. If you are not collecting those, check needed will be your first warning, and it arrives late.

How Netdata helps

  • Netdata collects thin pool data_percent and metadata_percent per pool, so the slow metadata climb that precedes a check-needed event is visible as a trend, not a surprise.
  • LV health state changes (read-only metadata, failed pools, check-needed flags) surface as alerts, so you learn about the flag from monitoring rather than from failed writes.
  • Correlating pool saturation with VG free space on the same dashboard answers the critical repair-time question instantly: is there room to extend metadata or rebuild pmspare?
  • D-state process counts and block-device I/O error charts sit next to LVM signals, providing the corroboration that distinguishes a pool erroring I/O from one that is merely flagged.
  • Per-second collection catches the rapid metadata jumps that follow snapshot creation bursts or random-write storms, which minute-interval polling routinely misses.