A thin pool that reports “check needed” is telling you the kernel no longer trusts the pool’s metadata B-tree. The pool may still be serving I/O, or it may already be erroring writes. Either way, the wrong move (resizing metadata, forcing activation, running the wrong repair path on an oversized metadata LV) can turn a recoverable state into permanent data loss.
This state usually appears after metadata exhaustion, an unclean shutdown during a metadata update, or an underlying storage error. It shows up in lvs output as c (check needed) or C (check needed, suspended) in position 5 of lv_attr, often alongside a health flag in position 9. If you have not read the background on how thin pools are structured, see how LVM actually works in production first. This page covers only detection, validation, and repair.
What this means
Every thin pool is backed by two internal LVs: a data LV and a much smaller metadata LV. The metadata LV holds the block mapping table (a B-tree) for every thin volume and snapshot in the pool. When the device-mapper thin target hits a metadata problem it cannot resolve online, it sets the needs_check flag in the pool superblock and surfaces it to LVM. The kernel’s thin-provisioning documentation is blunt about the consequence: if metadata space is exhausted or a metadata operation fails, the pool errors I/O until the pool is taken offline and repair is performed.
Two things follow from that:
- Repair is an offline operation. You cannot fix a flagged pool while it is active.
lvconvert --repairrequires the pool to be deactivated first. - Repair is not guaranteed. The lvmthin(7) man page warns that data may be unrecoverable. If the device details tree or the data mapping tree itself is damaged,
thin_repairmay produce incomplete output. Treat backups as the real recovery path and repair as the fast path.
“Check needed” is frequently what comes after metadata exhaustion, which is why thin pool metadata full is the saturation state worth alerting on first.
Common causes
| Cause | What it looks like | First thing to check |
|---|---|---|
| Metadata exhaustion | metadata_percent hit 100%, health flag set, data% may be moderate | lvs -o lv_name,data_percent,metadata_percent,lv_attr |
| Unclean shutdown or crash during metadata update | Flag appears after power loss or kernel panic | journalctl -k around the crash window |
| Underlying device errors on the metadata LV | I/O errors in dmesg for the PV hosting _tmeta | dmesg | grep -i 'I/O error' and pvs |
| Failed metadata operation under load | Random-write-heavy workload, many snapshots, metadata climbing for weeks | Metadata growth trend; snapshot count in lvs -o lv_name,origin |
The insidious case is metadata exhaustion with low data usage. Operators watching only data_percent never see it coming. The monitoring checklist covers why both percentages need separate alerting.
Quick checks
All of these are read-only and safe to run during an incident. Prefer dmsetup over lvs if the system is under I/O stress: lvs takes LVM locks and reads metadata from disk, and can hang on an unhealthy pool. dmsetup reads kernel state directly.
# Find pools flagged check needed (position 5 = c or C) and their health (position 9)
lvs -a -o lv_name,vg_name,lv_attr,data_percent,metadata_percent
# Kernel view of the pool: used/total metadata and data blocks, ro/rw, queue vs error policy
dmsetup status --target thin-pool
# Confirm which devices are actually active and whether any are suspended
dmsetup info -c -o name,attr,suspended | grep -i <vg>
# Kernel log evidence: dm-thin metadata errors, I/O errors on backing devices
dmesg | grep -iE 'dm-thin|thin|metadata|I/O error' | tail -50
Interpreting what you see:
cin position 5 with the pool still active: the flag is set but the pool has not yet failed. Plan a repair window before it gets worse.C(check needed, suspended): the pool is already suspended and I/O is blocked.Min position 9: metadata has gone read-only. The pool is in emergency mode.Fin position 9: the pool has failed. Writes are erroring.
How to diagnose it
Identify the pool and its metadata LV. Run
lvs -a -o lv_name,vg_name,lv_attr,lv_size,data_percent,metadata_percentand note the pool name and the size of its hidden[pool_tmeta]LV. You need the metadata LV size later; it determines which repair path is safe.Check kernel status without LVM locks.
dmsetup status <vg>-<pool>-tpool(ordmsetup status --target thin-pool) showsused_metadata/total_metadataandused_data/total_data, plus whether the pool is read-only and whether it queues or errors on no space.Validate the metadata. If the pool is active and you cannot take it down yet, run
thin_check --metadata-snap /dev/mapper/<vg>-<pool>_tmetato check a metadata snapshot of the live pool. Runningthin_checkdirectly against the live metadata device fails with “Device or resource busy”.thin_checkcomes from thedevice-mapper-persistent-datapackage (also shipped asthin-provisioning-toolson some distributions). For a faster but less thorough pass,--skip-mappingsskips the per-device mapping trees.Assess blast radius. List the thin volumes and snapshots in the pool, check which are mounted or in use, and confirm which ones have current backups. This decides how aggressive you can be.
flowchart TD
A[Pool shows check needed] --> B{Pool active?}
B -->|yes| C[thin_check --metadata-snap to validate]
B -->|no| D[Confirm backups exist]
C --> D
D --> E[Unmount thin LVs, lvchange -an VG/pool]
E --> F{Metadata LV over 15.81 GiB?}
F -->|yes| G[Manual path: thin_dump/restore to smaller LV]
F -->|no| H[lvconvert --repair VG/pool]
H -->|success| I[Extend metadata, activate, fsck thin LVs]
H -->|failure| G
G -->|failure| J[Restore from backup]Metrics and signals to monitor
| Signal | Why it matters | Warning sign |
|---|---|---|
metadata_percent | The leading indicator; exhaustion is the most common path to check needed | Above 75%, or any sustained upward trend |
lv_attr position 5 | Direct surface of the needs_check flag | c or C on any pool |
lv_attr position 9 (health) | Distinguishes degraded from failed | Any value other than - |
dmsetup status metadata ratio | Works when lvs hangs; kernel-side truth | used/total metadata approaching 1:1 |
| D-state processes on dm devices | User-visible symptom of a pool erroring or queuing I/O | Count growing over 60 seconds |
| Kernel log dm-thin messages | First place metadata operation failures appear | Any dm-thin error or read-only transition |
Fixes
Standard repair: lvconvert –repair
This is the supported path on a pool whose metadata LV is within the kernel’s size limit.
- Stop everything using the pool. Unmount filesystems on thin LVs, stop VMs or containers backed by them, and deactivate the thin LVs, then the pool:
# Deactivate thin LVs first, then the pool itself
lvchange -an <vg>/<thin_lv>
lvchange -an <vg>/<pool>
lvconvert --repair refuses to run on an active pool. Verify deactivation with dmsetup ls; a pool that still shows active mappings was not actually deactivated. If lvchange -an appears to succeed but mappings remain, re-run with -v and check for leftover device-mapper entries or processes still holding the devices open.
- Run the repair:
# Repair thin pool metadata (offline, destructive class of operation)
lvconvert --repair <vg>/<pool>
This runs thin_repair against the damaged metadata LV and writes the repaired copy to the VG’s pmspare LV. On success, the pmspare becomes the new metadata LV and the old damaged one is renamed to <pool>_meta<N>. Two prerequisites bite in practice: the pmspare LV must be at least as large as the damaged metadata LV (extend it or create a replacement first if not), and the VG needs free space for the swap.
- Extend metadata before reactivating. Whatever filled the metadata area will fill it again. Grow it now:
# Grow pool metadata while the pool is still inactive
lvextend --poolmetadatasize +<size>G <vg>/<pool>
Do not attempt this on a pool that still carries the needs_check flag. The kernel refuses to resize metadata until repair is done, and on a corrupted pool the resize attempt can hang. Repair first, extend second, in that order.
- Activate and check the upper layers. The kernel documentation strongly recommends running consistency checks (fsck) on the thin volumes after any pool repair. Metadata repair can leave individual block mappings inconsistent even when the pool as a whole validates.
The oversized-metadata trap
There is a hard kernel limit on thin pool metadata size: 4145152 4K blocks, about 15.81 GiB. On older toolchains (RHEL 7 era lvm2 and device-mapper-persistent-data, and Ubuntu 16.04’s lvm2 2.02.133), lvconvert --repair on a metadata LV larger than that writes a repaired superblock claiming the larger size (4161600 blocks has been observed). The kernel then refuses to activate the pool with an error along the lines of “metadata device (4145152 blocks) too small: expected 4161600”. Red Hat Bug 2028639 tracks this and was closed WONTFIX for RHEL 7.
If your metadata LV is at or above 16 GiB, or you are on an affected toolchain, skip lvconvert --repair and take the manual path below. Newer LVM versions also offer thin_pool_crop_metadata in lvm.conf, which crops metadata size to 15.81 GiB for compatibility; enable it only when you actually need to interoperate with older toolchains.
Manual repair path
When lvconvert --repair fails or is unsafe:
- Create a new, empty metadata LV, sized under the 15.81 GiB kernel limit.
- Run
thin_repair -i /dev/<vg>/<pool>_tmeta -o /dev/<vg>/<new_meta>to rebuild the mapping tree into the new LV. - Verify the result:
thin_check /dev/<vg>/<new_meta>. - Swap it in:
lvconvert --thinpool <vg>/<pool> --poolmetadata <vg>/<new_meta>.
On very old toolchains with the size bug, the equivalent workaround is thin_dump from the damaged metadata, thin_restore into a freshly created smaller metadata LV, then the same swap. thin_check --clear-needs-check-flag exists to clear the flag after a successful check, but only use it when you have verified the metadata is actually clean; clearing the flag on broken metadata just hides the warning. The --auto-repair option fixes only trivial issues such as metadata leaks; it is not a substitute for thin_repair on structural damage.
If all repair paths fail, restore from backup. That is the outcome the man page warns about, and it is why the pool’s health flags deserve paging-level attention before they ever reach this state.
Prevention
- Alert on metadata_percent early. Ticket at 75%, urgent at 90%. Metadata is cheap; over-provision it aggressively at pool creation time rather than relying on extension later. See the metadata exhaustion guide for sizing rationale.
- Never let the pool get here through data exhaustion either. A pool hitting out of data space under
queue_if_no_spacehangs I/O for every thin volume at once, and the failed-write churn also consumes metadata. - Verify auto-extend actually works. The default
thin_pool_autoextend_thresholdof 100 means disabled. See thin pool auto-extend not working and the low water mark warning that fires before the freeze. - Keep VG free space for the repair path.
lvconvert --repairneeds a usable pmspare and room to swap metadata LVs. A VG at zero free extents removes your repair options at the worst moment; see volume group free space low and insufficient free extents. - Track the maturity basics. The LVM monitoring maturity model puts metadata percent, health flags, and dmeventd state at Level 2. If you are not collecting those, check needed will be your first warning, and it arrives late.
How Netdata helps
- Netdata collects thin pool
data_percentandmetadata_percentper pool, so the slow metadata climb that precedes a check-needed event is visible as a trend, not a surprise. - LV health state changes (read-only metadata, failed pools, check-needed flags) surface as alerts, so you learn about the flag from monitoring rather than from failed writes.
- Correlating pool saturation with VG free space on the same dashboard answers the critical repair-time question instantly: is there room to extend metadata or rebuild pmspare?
- D-state process counts and block-device I/O error charts sit next to LVM signals, providing the corroboration that distinguishes a pool erroring I/O from one that is merely flagged.
- Per-second collection catches the rapid metadata jumps that follow snapshot creation bursts or random-write storms, which minute-interval polling routinely misses.
Related guides
- LVM thin pool metadata full: the exhaustion that can corrupt the pool
- LVM thin pool out of data space: every thin volume freezes at once
- LVM reached low water mark for data device: the thin pool warning before the freeze
- LVM thin pool auto-extend not working: threshold 100 means disabled
- LVM volume group running low on free space: vg_free and runway estimation
- LVM Insufficient free extents: the volume group is out of space
- LVM cannot extend a logical volume: adding a PV when the VG is full
- LVM filesystem full while the volume group has space: the resize step everyone forgets
- LVM has free space but striped or mirrored allocation still fails
- How LVM actually works in production: a mental model for operators
- LVM monitoring checklist: the signals every production volume manager needs
- LVM monitoring maturity model: from survival to expert






