The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nvme / nvme-endurance-runway-planning ▌

Operations Guides

NVMe endurance runway: projecting time-to-replacement from wear signals

NVMe wear-out is one of the few storage failures that usually announces itself months in advance. The announcement is quiet: a monotonically rising percentage_used, a slow drift in available_spare, maybe the first few media_errors. If you only look at current values, you find out when the drive crosses a threshold. If you trend the rates, you can procure before the drive becomes an incident.

This guide is a planning procedure for local PCIe-attached NVMe devices. It turns three wear signals into a replacement date, then cross-checks that date against spare-block consumption and early media failure. It does not apply to NVMe-oF, ZNS, SPDK, or virtualized cloud “NVMe” devices where SMART data is absent, synthetic, or owned by the provider.

The core rule: treat percentage_used as the procurement clock, available_spare as the physical runway, and media_errors as the signal that can invalidate both.

What this projection gives you

A useful endurance runway answers four operational questions:

  • When to buy. Start procurement when percentage_used crosses 80%, or when the projected date to 80% is inside your procurement lead time.
  • When to schedule. Schedule replacement when percentage_used crosses 90%, when available_spare approaches its vendor threshold, or when spare consumption accelerates.
  • When to stop trusting the estimate. New media_errors, critical warning bit 2, or a falling available_spare out of proportion to percentage_used means the smooth wear model is no longer the main risk.
  • When the drive class is wrong. A sudden increase in wear rate usually means the workload changed, the drive is too full, write amplification increased, or a consumer-class drive is doing an enterprise-class job.

Do not page on percentage_used alone. High endurance consumed is a planning signal, not an emergency. The emergency signals are read-only mode, reliability degraded with rising media errors, controller resets, and device disappearance.

Prerequisites

  • Device access. You need read access to the NVMe character device, for example /dev/nvme0. Namespace devices such as /dev/nvme0n1 report block I/O, but SMART wear data is controller-level.
  • A time window. SMART wear counters are coarse and polled. Use at least 14 days for a first trend and 30 days for procurement decisions. Hourly deltas are noise.
  • Drive identity. Record model, firmware, capacity, rated TBW or DWPD, and whether the drive is consumer QLC, consumer TLC, or enterprise TLC. The endurance scale varies enormously: consumer QLC is often in the 100-300 TBW range, while enterprise TLC is commonly specified around 1-10 DWPD.
  • Fill level context. A nearly full drive has higher garbage collection pressure and can wear faster than the same workload on a mostly empty drive. Keep capacity utilization in the same trend view.

Signals to use

SignalSource fieldWhat it contributesMain trap
Endurance consumedpercentage_usedVendor estimate of rated life consumed; monotonic fuseVendor-specific estimate; can exceed 100 and still run
Spare blocksavailable_spare, available_spare_thresholdPhysical margin for bad-block replacement; bit 0 trips at thresholdLast portion depletes non-linearly; threshold varies by vendor
Media failuremedia_errorsEarly proof that ECC and retries are no longer enoughLifetime counter; only rate of increase matters
Host write volumedata_units_written, power_on_hoursCross-check against TBW or DWPD ratingHost-visible writes only; not NAND writes
Reliability assessmentcritical_warning bits 0, 2, 3Drive’s own urgent state: spare low, reliability degraded, read-onlyA single bitmask hides different severities

Collect the raw values with read-only commands:

# Read SMART wear and health fields for one controller
nvme smart-log /dev/nvme0

# Machine-readable form for trending
nvme smart-log /dev/nvme0 -o json | jq '{percent_used, avail_spare, spare_thresh, media_errors, data_units_written, power_on_hours, critical_warning}'

Do not average these values across unlike drive models; trend per model and per workload.

Procedure

  1. Baseline the drive class. Record model, firmware, capacity, rated TBW or DWPD, and current fill level. A projection without the drive class is not actionable, because 5% per month means very different things on a 100 TBW QLC part and a 3 DWPD enterprise part.

  2. Compute the daily endurance rate. Use at least a 14-day window, preferably 30 days:

    daily_pct = (percentage_used_now - percentage_used_then) / days_between_samples
    

    If daily_pct <= 0, discard the window. The counter should not go down; a flat or negative delta means bad samples, stale SMART data, a replaced drive, or a vendor reporting quirk.

  3. Project days to rated end of life:

    days_to_100 = (100 - percentage_used_now) / daily_pct
    

    Also project days to the action thresholds:

    days_to_80 = max(0, (80 - percentage_used_now) / daily_pct)
    days_to_90 = max(0, (90 - percentage_used_now) / daily_pct)
    
  4. Apply the final-10% safety factor. Spare depletion accelerates near end of life, so the last 10% is not linear. Once percentage_used_now >= 90, or once any projection enters the 90-100 band, halve the remaining runway:

    effective_days = days_to_100 / 2
    

    This is deliberately conservative. The goal is to replace before spare exhaustion and read-only mode, not to prove the datasheet exact.

  5. Compute the spare-block runway. Convert available_spare decline into a daily drop over the same window:

    spare_daily_drop = (available_spare_then - available_spare_now) / days_between_samples
    spare_days = (available_spare_now - available_spare_threshold) / spare_daily_drop
    

    Treat spare_days as invalid if spare_daily_drop <= 0. A long flat period followed by a step down is common on some controllers; use a longer window before trusting it.

  6. Take the minimum credible runway. Use the smaller of the endurance runway and spare runway, then downgrade confidence if media_errors is increasing. A smooth percentage_used projection does not protect you from a NAND batch problem that shows up first as media errors.

  7. Cross-check with host write volume. Convert data_units_written to bytes using 512,000 bytes per unit (thousands of 512-byte units):

    tb_written = data_units_written * 512000 / 1e12
    

    Compare lifetime TB written with rated TBW where the vendor publishes one. For DWPD-rated drives, estimate recent DWPD from the delta in data_units_written, drive capacity, and elapsed days. This is host-visible write volume: the FTL can write more to NAND than the host wrote, so standard SMART cannot give you true write amplification.

  8. Classify the action.

    StateTriggerAction
    Normalpercentage_used < 80, spare above 2x threshold, no new media errorsKeep trending monthly
    Procurepercentage_used >= 80, or projected 80% date inside lead timeBuy replacement, validate spares, check fleet siblings
    Schedulepercentage_used >= 90, spare near 2x threshold, or spare runway shorter than endurance runwayPlan migration before failure signals appear
    Replace nowspare at or below threshold, critical warning bit 0, bit 3 read-only, or bit 2 with rising media errorsProtect data first, then replace
    Wrong classwear rate spikes after workload change, or rate exceeds class expectationFix workload or move to higher-endurance drive
flowchart TD
  A[Collect SMART wear fields] --> B[Trend percentage_used and available_spare]
  B --> C[Project endurance and spare runway]
  C --> D{Media errors rising or critical_warning set?}
  D -- yes --> E[Protect data and replace]
  D -- no --> F{Runway inside lead time?}
  F -- no --> G[Keep monthly trend]
  F -- yes --> H[Procure at 80 and schedule by 90]

Reading the result

A worked example with synthetic numbers: a drive goes from 61% to 64% percentage_used in 30 days. daily_pct = 3 / 30 = 0.1. At 64%, days_to_100 = 36 / 0.1 = 360 days, days_to_80 = 160 days, and days_to_90 = 260 days. If procurement takes 90 days, procurement starts when projected days-to-80 falls below 90, not when the drive is already at 80.

Now change one input: available_spare falls from 41 to 38 over the same 30 days with a threshold of 10. spare_daily_drop = 3 / 30 = 0.1, so spare_days = 28 / 0.1 = 280. The spare runway is shorter than the endurance runway to 100 and close to the 90% schedule window. This drive should be scheduled earlier than the percentage_used-only estimate suggests.

If media_errors increases during the window, stop presenting the result as a date. Present it as a risk state: verify redundancy, rewrite or back up cold data if your architecture requires it, and replace during the next maintenance window. If critical warning bit 2 is also set, treat it as active degradation rather than planning.

Common pitfalls

  • Using one sample. A single SMART read tells you state, not runway. Without rate of change, 30% used can be safe for years or five months from replacement.
  • Trusting 100% as a cliff. The NVMe model allows percentage_used above 100, and many drives keep running past it. Operational risk rises, but the exact failure point is vendor- and workload-dependent.
  • Ignoring the spare curve. available_spare can remain high for a long time and then fall faster near the end. The final 10% needs a safety factor, not a straight line.
  • Confusing host writes with NAND writes. data_units_written excludes FTL write amplification. It is useful for workload cross-checks, not for proving exact cell wear.
  • Averaging unlike drives. Consumer QLC, consumer TLC, and enterprise TLC have different ratings, reporting behavior, over-provisioning, and failure tolerance.
  • Alerting on the raw bitmask. Critical warning bit 0 is a replacement signal, bit 2 plus rising media errors is active degradation, and bit 3 is write refusal. One alert for critical_warning != 0 hides the response.
  • Forgetting fill level. Drives under heavy garbage collection pressure can wear faster than the same host write rate on a drive with more free space. Keep logical free space in the same review.

Signals to monitor

SignalWhy it mattersWarning sign
percentage_used rateMain procurement clockSustained increase above baseline, or roughly >1% per week without a planned workload change
available_spare vs thresholdPhysical bad-block marginAt or below 2x threshold, accelerating downward trend, or at/below threshold
media_errors rateEarliest media failure signalAny sustained increase; page-level concern when paired with critical warning bit 2
data_units_written deltaWorkload cross-check against TBW or DWPDWrite rate no longer matches drive class or approved workload
critical_warning bitsDrive’s own urgent assessmentbit 0 spare low, bit 2 reliability degraded, bit 3 read-only
Drive fill levelGC pressure and wear amplification contextSustained operation with little logical free space

How Netdata helps

  • Track nvme.device_estimated_endurance_perc as a trend, not a one-off gauge, so percentage_used rate of change is visible before the 80% and 90% thresholds.
  • Compare nvme.device_available_spare_perc against the vendor spare threshold and watch for acceleration, not only the crossing.
  • Alert on nvme.device_media_errors_rate as a rate, because the lifetime counter never decreases and isolated old errors can be benign.
  • Use nvme.device_io_transferred_count to cross-check host write volume against the endurance trend and expose workload changes that invalidate the projection.
  • Correlate endurance, spare, media errors, temperature, and per-device I/O on one timeline so a wear-rate spike can be tied to a workload, fill-level, or thermal change instead of being treated as a mysterious SMART jump.