The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / lvm / lvm-raid-mirror-degraded ▌

Operations Guides

LVM RAID or mirror degraded: a leg is dead and you are one failure from data loss

Your LVM RAID1 or mirrored logical volume is still serving reads and writes. Applications see no errors. Filesystems are mounted. But one leg is dead and you are running on a single copy with zero redundancy. The next disk failure, cable disconnect, or SAN path loss is total data loss.

The default raid_fault_policy is warn. When a leg fails, dmeventd logs a warning but takes no repair action. No resync is triggered, no alert fires unless your monitoring explicitly checks for degraded state. The system runs indefinitely on the surviving leg until the second failure removes all copies.

What this means

When a leg fails, device-mapper removes it from the active set and routes I/O to the surviving leg or legs. Data is still accessible, but the redundancy is gone.

The failure surfaces through two interfaces:

  • dmsetup status <vg>-<lv> shows per-leg health characters. A means alive and in-sync. D means dead or failed. a means alive but not in-sync (normal during rebuild).
  • lvs shows aggregate LV health. The lv_health_status field reports “partial” when a device is permanently gone, or “refresh needed” when a device was transiently missing and returned. The lv_attr field shows p in position 9 for partial, or r for refresh needed.

The critical distinction is between D and a in dmsetup status output. D means a device has failed and is no longer participating in the array. a means alive but not yet synchronized, which is expected during initial sync or rebuild. Alert on D, or on a copy_percent that was previously 100 and has dropped below 100.

Fixed in LVM 2.02.155 (WHATS_NEW: “Correcting value in copy_percent() for 100%”). On older LVM versions, copy_percent may report 100.00 due to rounding even when the device is not fully in-sync. For those versions, rely on dmsetup status health characters to determine true sync state.

flowchart TD
    A["copy_percent dropped below 100
or lv_attr shows p or r"] --> B{"Check dmsetup status
per-leg health chars"} B -->|D present| C["Leg is dead or failed"] B -->|a present, no D| D["Resync in progress
Check if progressing"] B -->|All A| E["Check pv_attr for
missing PV"] C --> F{"Is the failed device
visible to the OS?"} F -->|No, gone| G["lvconvert --repair
with replacement PV"] F -->|Yes, returned| H["lvchange --refresh
to trigger resync"] F -->|Yes, still present| I["lvconvert --replace
with new PV"] D -->|Progressing| J["Normal rebuild
Monitor to completion"] D -->|Stalled over 1 hour| K["Investigate device
errors blocking rebuild"]

Common causes

CauseWhat it looks likeFirst thing to check
Disk hardware failuredmsetup status shows D for one leg, dmesg shows I/O errors for a specific block device`dmesg
SAN LUN unpresented or zoned awayPV shows as [unknown] in pvs, VG attr shows partialSAN management console, verify LUN masking and zoning
Multipath all paths downPV missing, underlying paths show errorsmultipath -ll to check path status
Cable or connector failureSimilar to disk failure, may be intermittent with resets in dmesgCheck /sys/block/<device>/device/state
Controller or HBA failureMultiple PVs on same controller affected simultaneouslyIdentify which PVs share a controller
Accidental device removalPV missing after maintenance window or hot-unplugVerify physical device presence and udev events
Transient link loss with recoverylv_health_status shows “refresh needed” instead of “partial”Device returned but needs manual refresh

Quick checks

# Check copy_percent and health for all RAID/mirror LVs
lvs -o lv_name,vg_name,lv_attr,copy_percent,lv_health_status -S 'seg_type=~raid|seg_type=~mirror'

# Get per-leg health characters from device-mapper
dmsetup status <vg>-<lv>
# Look for D (dead), a (alive not in-sync), A (alive in-sync)

# Check which PVs are missing
pvs -o pv_name,vg_name,pv_attr,pv_size,pv_free
# Missing PVs show as [unknown]

# Check VG partial status
vgs -o vg_name,vg_attr,vg_missing_pv_count
# 'p' in attr position 4 means partial (missing PV)

# Verify each PV device exists
for pv in $(pvs --noheadings -o pv_name); do
  [ -b "$pv" ] && echo "$pv: OK" || echo "$pv: MISSING"
done

# Check for I/O errors on underlying devices
dmesg | grep -i 'I/O error\|offline\|not ready' | tail -30

# Check RAID-specific fields
lvs -o lv_name,raid_sync_action,raid_mismatch_count,copy_percent -S 'seg_type=~raid'

# Check multipath if SAN-attached
multipath -ll

# Check if resync is progressing (run twice, compare copy_percent)
lvs -o lv_name,copy_percent && sleep 30 && lvs -o lv_name,copy_percent

# Check dmeventd is running (relevant for raid_fault_policy)
systemctl is-active lvm2-monitor.service
pgrep -x dmeventd

How to diagnose it

  1. Identify which LVs are degraded. Run lvs -o lv_name,vg_name,lv_attr,copy_percent,lv_health_status -S 'seg_type=~raid|seg_type=~mirror'. Any LV with copy_percent below 100, a non-empty lv_health_status, or lv_attr position 9 not - needs investigation.

  2. Get per-leg health from device-mapper. Run dmsetup status <vg>-<lv> for each suspect LV. D confirms a dead leg. a without D means resync in progress (normal). All A with copy_percent below 100 may indicate the rounding bug on older LVM or the very start of a resync.

  3. Determine which PV backed the failed leg. Check pvs for missing devices. A PV showing as [unknown] means the device disappeared from the system entirely. Cross-reference with pvs --segments -o pv_name,lv_name,seg_start_pe,seg_size_pe to map PVs to LV segments.

  4. Check the kernel log for the cause. Run dmesg | grep -i 'I/O error\|offline\|not ready\|device not found' to find when and why the device disappeared. SCSI timeouts, link resets, and medium errors each produce distinct messages.

  5. Check the physical layer. For SAN-attached storage, run multipath -ll to verify path status. For local disks, check /sys/block/<device>/device/state. For NVMe, check /sys/class/nvme/nvme*/state.

  6. Distinguish permanent loss from transient failure. If lv_health_status shows “refresh needed”, the device was temporarily absent and has returned. This requires lvchange --refresh to clear, not a full repair. If the status shows “partial”, the device is gone and needs replacement.

  7. Verify whether resync is progressing. If rebuilding (lowercase a in health chars, no D), check that copy_percent is advancing. Run lvs -o lv_name,copy_percent twice with a 30-second gap. If the value is unchanged after an hour on a reasonably sized LV, the rebuild is stalled and you need to investigate the underlying device.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
copy_percent per RAID/mirror LVWas 100, now below 100 means redundancy lost or resync startedAny value below 100 that was previously 100
dmsetup status health charactersPer-leg D is definitive proof of a dead deviceAny D in the health character string
lv_health_status fieldAggregate LV health in human-readable form“partial” or “refresh needed”
lv_attr position 9Single-character health flag for quick filteringp (partial) or r (refresh needed)
PV accessibilityIdentifies which physical device is goneAny PV showing as [unknown]
vgs VG attr position 4VG-level partial flagp means the VG has missing PVs
Resync progress rateDetects stalled rebuildscopy_percent unchanged for over 1 hour
raid_mismatch_countData inconsistency between legs detected by scrubAny nonzero value after a scrub completes

Fixes

The failed PV is gone (not visible to the OS)

You need a replacement device with enough free extents to hold the failed leg’s data.

# DESTRUCTIVE: This destroys any existing data on the device
pvcreate /dev/new_device

# Add it to the VG if not already a member
vgextend <vg> /dev/new_device

# Repair the RAID LV, allocating the new leg from the replacement PV
lvconvert --repair <vg>/<lv> /dev/new_device

lvconvert --repair allocates a new leg from the specified PV (or from VG free space if no PV is named), copies data from the surviving leg, and brings the new leg into sync. The failed leg’s metadata is cleaned up.

Confirmed by lvmraid(7): integrity must be disabled before repair/replace commands can be used. If raidintegrity is enabled on the LV, you may need to disable it before repair: lvconvert --raidintegrity n <vg>/<lv>. Re-enable after repair completes.

The repair command requires free extents in the VG. If insufficient space exists, it fails with an error indicating how many extents are needed versus available.

The failed PV returned (transient failure)

If the device was temporarily disconnected (SAN path blip, cable reseat) and is now visible again, the LV shows “refresh needed” in lv_health_status. The array has not automatically re-synced the returned device.

# Non-destructive: trigger resync of the returned device
lvchange --refresh <vg>/<lv>

After refresh, monitor copy_percent to confirm the resync completes. The device rejoins as a (alive, not in-sync) and transitions to A (alive, in-sync) when the copy reaches 100 percent.

The failing PV is still visible but unreliable

When a device is present but showing hardware errors, bad sectors, or SMART warnings, use lvconvert --replace to swap it out without waiting for total failure.

# Prepare the replacement as a PV first
pvcreate /dev/new_device
vgextend <vg> /dev/new_device

# Replace the failing device
lvconvert --replace /dev/failing_device <vg>/<lv>

Unlike --repair, --replace works on devices that are still visible and participating in the array. Data is copied to the new device, then the old leg is removed.

Repairing an active system volume

For LVs that must stay active (root, swap, mounted application volumes), run lvconvert --repair with the LV active. The repair does not require deactivation. Monitor I/O latency during the resync, as rebuild traffic competes with production workload for disk bandwidth.

After any repair: verify full redundancy

# Confirm all legs show A (alive, in-sync)
dmsetup status <vg>-<lv>

# Confirm copy_percent is 100
lvs -o lv_name,copy_percent -S 'seg_type=~raid|seg_type=~mirror'

# Confirm lv_health_status is empty
lvs -o lv_name,lv_health_status

Do not consider the incident resolved until every leg shows A and copy_percent is 100 with no health flags.

Prevention

  • Alert on any RAID/mirror LV with copy_percent below 100 that was previously at 100. This catches both new failures and stalled resyncs. Gate new LVs (which legitimately start below 100 during initial sync) by tracking which LVs have ever reached 100.
  • Alert on any D health character in dmsetup status output. This is the definitive failure signal.
  • Alert on lv_health_status showing any non-empty value. Both “partial” and “refresh needed” require operator action.
  • Monitor resync progress rate. A stalled resync (copy_percent unchanged for over an hour on a reasonably sized LV) often indicates an underlying device problem.
  • Verify dmeventd is running. With the default raid_fault_policy of warn, dmeventd logs the failure but does not auto-repair. Confirm the daemon is at least logging so you can correlate events.
  • Document LV-to-PV-to-physical-device mappings. When a leg fails, you need to know immediately which physical device to replace. Maintain this mapping outside LVM metadata so it survives PV loss.
  • Run periodic RAID scrubs. lvchange --syncaction check <vg>/<lv> verifies data consistency between legs and populates raid_mismatch_count. Investigate any nonzero mismatch count.
  • Verify physical independence of legs. Two PVs on the same controller, disk shelf, or SAN fabric provide no real redundancy. Verify when the array is initially created.

How Netdata helps

  • Tracks copy_percent trends across all RAID and mirror LVs, alerting when a previously synced LV drops below 100.
  • Surfaces lv_health_status and lv_attr position 9 as discrete metrics, enabling alerts on any non-healthy state without parsing command output in scripts.
  • Correlates per-device disk I/O latency and error rates with LVM degradation events, so you can see whether the surviving leg is also showing wear.
  • Collects kernel log entries (dmesg) alongside LVM metrics, providing device-level context for LVM-level degradation without running multiple commands manually.
  • Monitors dmeventd process presence and lvm2-monitor service state, so you know whether the safety net is actually running.