The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nvme / nvme-how-it-works-in-production ▌

Operations Guides

How NVMe actually works in production: a mental model for operators

Most storage incidents on NVMe are misdiagnosed for the same reason: operators debug them as if NVMe were a faster SATA disk. It is not. NVMe is a host-to-controller communication protocol that exposes flash storage over PCIe, and the failure modes live in places generic disk monitoring never looks: the PCIe link, the controller firmware, and the flash translation layer sitting between your filesystem and the NAND.

The mental model that makes NVMe behavior predictable has three layers: the PCIe transport, the controller, and the flash media. Once you hold this model, most “mystery slowness” and “healthy SMART, dead drive” incidents stop being mysterious. Scope here is local PCIe-attached NVMe. NVMe-oF, ZNS, and SPDK change the monitoring model fundamentally and are not covered.

What it is and why it matters

A SATA disk is mostly a passive device: the OS sends commands, the disk obeys, and SMART is a thin reporting layer. An NVMe device is a small embedded computer. It has its own firmware, its own DRAM (or borrowed host RAM), its own scheduler, and an internal mapping layer that constantly relocates your data without telling you. It manages wear, heat, and error correction autonomously, and it makes decisions, like throttling or going read-only, that your OS only learns about after the fact.

This is why NVMe failure archetypes look the way they do:

  1. It gets slow before it dies. Latency rises as FTL overhead increases, long before outright failure.
  2. It runs out of spare blocks. Write amplification exhausts over-provisioned space and performance collapses.
  3. It thermal-throttles silently. No error, just gradually decreasing performance until thermal equilibrium.
  4. The controller hangs. A firmware bug or internal fault stops all I/O until the kernel resets it.
  5. It wears out predictably but silently. SMART shows declining life but nothing alerts until critical.
  6. The PCIe link degrades. Signal integrity issues cause retransmissions that look like high latency.

Every one of these maps to a specific layer. Debugging NVMe well means knowing which layer you are looking at.

How it works: the three layers

flowchart TD
  APP[Application / filesystem] --> BLK[Linux block layer]
  BLK -->|writes commands into SQ| SQ[Submission queues in host RAM]
  SQ -->|DMA fetch| CTRL[NVMe controller]
  CTRL -->|completion entries| CQ[Completion queues]
  CQ -->|interrupt / MSI-X| BLK
  CTRL -->|FTL: LBA to NAND page map| FTL[Flash translation layer]
  FTL --> NAND[NAND dies across channels]
  PCIE[PCIe link: Gen x width, AER error reporting] -.transport.- CTRL

Layer 1: the PCIe transport

The NVMe device is a PCIe endpoint. It negotiates a link with a specific generation (Gen3/Gen4/Gen5) and width (x1/x2/x4), and that negotiation is not guaranteed to hold.

The critical operational fact: the link can degrade silently. A Gen4 x4 device running at Gen3 x2 delivers one-quarter of its rated bandwidth with zero errors visible to the filesystem, zero SMART changes, and nothing in dmesg. The drive looks healthy. Everything is just slower.

PCIe also has its own error reporting layer, AER (Advanced Error Reporting), that operates below NVMe entirely. Correctable errors are retransmitted transparently, so there is no data loss, but each retransmission is latency. A slightly loose M.2 connector can cause thousands of retransmissions: the drive reports healthy, and latency is 10x worse. This is the diagnostic signal teams most consistently miss.

Layer 2: the controller

The controller is where almost all interesting behavior lives. Its key components:

  • Submission and completion queues. The host writes commands into submission queue (SQ) entries in host memory; the controller fetches them via DMA, processes them, and writes completion entries to the completion queue (CQ). Each CPU core typically gets its own queue pair, and queues can hold up to 64K entries each. (The Linux driver defaults to 1024 entries per I/O queue, settable via the io_queue_depth module parameter.) Queue depth and queue count determine how much parallelism the drive can exploit. A dedicated admin queue (always queue 0) handles management commands like identify, log pages, and firmware operations.
  • DRAM or HMB. Enterprise drives carry their own DRAM for mapping tables and write buffering. Consumer drives are often DRAM-less and borrow host RAM via HMB (Host Memory Buffer). HMB is used mainly as an L2P table cache, so the performance penalty shows up mainly on random I/O.
  • The flash translation layer (FTL). Maps logical block addresses to physical NAND pages and handles wear leveling, garbage collection, and bad block management. The FTL is entirely opaque to the host, and its behavior under pressure is the single largest source of latency variance in NVMe.
  • Garbage collection. NAND must erase entire blocks before rewriting. GC moves valid pages out of partially used blocks to free them. When the drive is full or worn, GC competes directly with host I/O for controller resources, causing latency spikes.
  • Wear leveling. Distributes writes across blocks so hot blocks do not die early. Consumes background bandwidth.
  • Write buffer / SLC cache. Most controllers program part of the NAND in fast pseudo-SLC mode to absorb write bursts. When sustained writes exhaust it, throughput drops off a cliff, 50 to 80 percent in a step function, with no errors. This is by design, not a defect.
  • Thermal management. The controller throttles itself at vendor-defined thresholds (WCTEMP for warning, CCTEMP for critical). Throttling is invisible to the OS unless you watch temperature and throughput at the same time.

Layer 3: the flash media

Modern devices have multiple NAND dies operating in parallel across channels. Write latency depends on how well the controller distributes writes across that parallelism, which is an FTL problem again.

NAND cells have finite program/erase cycles, varying enormously by cell type (on the order of 100K cycles for SLC down to roughly 1K for QLC). As blocks wear out, the controller remaps them to spare blocks. When spares run out, the error rate accelerates and the drive approaches read-only mode. Two more media behaviors matter operationally: read disturb (repeated reads to a block can flip bits in adjacent cells, forcing the controller to relocate data in the background) and retention (stored charge leaks over time, faster at high temperature and high P/E counts, so cold data can rot).

Where it shows up in production

Map each symptom to a layer first, then confirm with the layer-specific signal:

SymptomLikely layerMechanism
Bandwidth capped, zero errorsPCIe transportLink trained down below max speed/width
Latency 10x worse, SMART cleanPCIe transportCorrectable AER errors being silently retransmitted
Gradual slowdown under load, recovers when idleControllerThermal throttling at WCTEMP
Sudden 50-90% write throughput drop, normal tempController / mediaSLC cache exhausted, or GC stall on a nearly full drive
Periodic write latency spikes, drive over 80% fullController (FTL)GC thrashing under space pressure
Read latency sporadically high, cold dataMediaRetention loss, reads requiring retry passes
5-30 second total I/O stalls, then recoveryControllerFirmware hang, kernel driver resets the controller
Drive refuses writesMediaSpare blocks exhausted, controller set read-only mode

A few deployment facts that change which layer dominates:

  • Consumer vs enterprise. Consumer drives lack DRAM (HMB instead), lack power-loss protection, throttle aggressively, and have lower endurance. Enterprise drives have full DRAM, PLP capacitors, higher DWPD, and more consistent latency. The same thresholds cannot be applied to both.
  • M.2 vs U.2/U.3/EDSFF. M.2 drives lack active cooling and thermal-throttle much faster. An M.2 drive without a heatsink under sustained write load will throttle while the CPU sits idle.
  • Drive fullness is a performance variable. Below roughly 80% fill the FTL has room to work; past that, GC pressure rises and write performance degrades. Maintaining 15-20% free space is operational headroom, not waste. TRIM matters for the same reason: without it, the FTL thinks the drive is fuller than it logically is.

Why SMART is not ground truth

Every SMART value is self-reported by controller firmware. That has three consequences.

First, firmware can be wrong or silent. Controllers can stop updating SMART counters during internal error states, and cheap drives may never assert critical warning bits even while failing. If SMART values have not changed in 24+ hours on an active drive, the data is stale, not stable.

Second, the fields are vendor estimates, not measurements. percentage_used is the vendor’s prediction of consumed endurance; the spec allows values above 100 and drives routinely run at 150% or more. The critical_warning field is a single bitmask where bit 2 (“reliability degraded”) is asserted at the vendor’s discretion about what counts as significant media or internal errors.

Third, SMART only covers the media and controller layers. The PCIe transport is invisible to SMART entirely. A drive at Gen3 x2 with accumulating AER retransmissions reports perfect health.

The fix is cross-referencing: treat SMART as the drive’s opinion, and check it against OS-level observations (block layer I/O counters, kernel logs, sysfs link status, AER counters) before trusting it.

Signals to watch in production

SignalWhy it mattersWarning sign
Controller state (/sys/class/nvme/nvmeX/state)Most direct availability signal: live, resetting, deadAny non-live state sustained over 30 seconds
critical_warning bits (SMART)The drive’s own emergency flags, per bitBit 3 (read-only) is an outage; bit 0 (spare low) is a replacement signal, not a page
media_errors rateUncorrectable NAND errors; lifetime counterAny sustained increment over zero
percentage_used + rateEndurance consumed; fuse that only goes upOver 80% plan replacement; over 1% per week suggests write amplification or workload mismatch
available_spare vs spare_threshRunway before read-only modeBelow 2x vendor threshold; accelerating consumption rate
Composite temperature + warning_temp_timeThermal throttling is silentTemperature near WCTEMP; warning time counter increasing
unsafe_shutdowns rateEach one risks data loss on drives without PLPAny increment during normal operation
PCIe link: current_link_speed/width vs maxSilent bandwidth capCurrent below max, either field
PCIe AER counters (aer_dev_correctable, aer_dev_fatal, aer_dev_nonfatal)Transport-layer signal integrityAny sustained non-zero correctable rate; any uncorrectable error
Kernel log: nvme timeout / reset messagesController hangs are not in SMARTOne reset is a ticket; repeated resets are a page

Note the split: SMART fields come from nvme smart-log, but controller state, link status, and AER counters are sysfs and kernel-log only. A monitoring setup that only polls smart-log sees roughly two of the three layers.

How Netdata helps

  • Netdata’s NVMe collector breaks critical_warning into per-bit dimensions (read_only, nvm_subsystem_reliability, available_spare, temp_threshold, volatile_mem_backup_failed), so you can alert per bit with the right severity instead of one blanket alert for the whole byte.
  • Lifetime SMART counters are exposed as rates where the rate is the signal: media errors, error log entries, and unsafe shutdowns only matter when they increment.
  • Endurance signals (percentage_used, available spare) are charted over time, which makes the consumption trajectory visible. The rate of spare depletion is a stronger leading indicator than the current value.
  • Composite temperature is charted alongside throughput, so the telltale thermal-throttling correlation (temperature up, throughput down, zero errors) is visible on one dashboard instead of requiring manual cross-referencing.
  • Warning and critical composite temperature time counters are collected, so thermal stress that happened between polls is not lost.

Two gaps to cover yourself: the PCIe transport layer (link speed/width, AER counters) is not collected by the NVMe collector and requires sysfs access, and controller resets live only in kernel logs.