The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nvme / nvme-controller-reset-loop ▌

Operations Guides

NVMe controller reset loop: repeated resets from a firmware hang

Your kernel log shows the same cycle over and over: an I/O command stalls until the 30-second io_timeout fires, the NVMe driver resets the controller, I/O briefly recovers, then the whole thing hangs again. Each cycle costs your applications 5 to 30 seconds of stalled I/O. Databases time out, requests fail, clusters rebalance, and minutes later it happens again.

This is the controller firmware hang loop. The firmware hits a bug triggered by a specific command sequence, queue depth, I/O size, or power state transition. The kernel detects the timeout, resets the controller, and replays the in-flight commands. The same trigger recurs, so the controller hangs again. The loop can continue indefinitely.

The defining characteristic is what is absent: temperature is normal, PCIe AER counters are clean, and SMART shows no media errors. Nothing in the environment explains the resets. That absence is the fingerprint of a firmware defect rather than a hardware or thermal problem.

What this means

A single controller reset is the kernel’s brute-force recovery: a command timed out, the driver gave up waiting, and it reset the controller to replay in-flight I/O. One reset can be a transient glitch. Repeated resets with the same shape (latency spike to the timeout value, reset, brief recovery, another hang) mean the controller keeps walking into the same faulty internal state.

The trigger is usually one of:

  • A specific command sequence or opcode pattern the firmware mishandles.
  • A queue depth or I/O size that pushes the firmware into an unhandled path.
  • An APST (Autonomous Power State Transition) event, where the controller enters a low-power state and fails to wake cleanly.

Because the trigger is workload-dependent, the loop often correlates with a particular application, a time of day, or a recent change in I/O pattern. Because it is firmware-internal, no host-side counter will show you the cause. SMART stays clean. The drive reports healthy right up until it stops answering.

flowchart LR
  A[Normal I/O] --> B[Trigger: command seq, queue depth, I/O size, or APST]
  B --> C[Firmware hangs, controller stops responding]
  C --> D[Latency spikes to 30s io_timeout]
  D --> E[Kernel resets controller, replays I/O]
  E --> F[Brief recovery]
  F --> B
  E -.->|recovery fails| G[Device dead or removed, PAGE]

Common causes

CauseWhat it looks likeFirst thing to check
Firmware bug hit by workload patternResets recur under a specific application or I/O shape, environment cleanFirmware version via nvme id-ctrl, vendor advisories for that version
APST power state transition fails to wakeResets correlate with idle-to-active transitions, often on consumer drivesDisable APST and watch whether resets stop
Severe PCIe bus faultResets plus AER errors, link retrains, possible device disappearanceAER counters and dmesg for pcieport errors
Fried or dying controllerResets escalate into failed recovery, controller state goes dead/sys/class/nvme/nvmeX/state, error log entries
Corrupted firmwareLoop starts after an interrupted or failed firmware updatenvme fw-log history and slot state

The lookalikes matter because their fixes differ completely. A PCIe fault is a reseat or retimer problem. A firmware hang is a firmware update or an RMA. Do not skip the rule-out steps.

Quick checks

All read-only and safe to run during an incident.

# Confirm the reset loop pattern in the kernel log
journalctl -k --no-pager | grep -i "nvme" | grep -i "timeout\|reset\|aborting"

# Count resets per device over the current boot
journalctl -k --no-pager | grep "nvme0" | grep -c "reset controller"

# Check current controller state (live, resetting, dead, deleting)
cat /sys/class/nvme/nvme0/state

# Capture the firmware revision before anything else
nvme id-ctrl /dev/nvme0 | grep -E "^(fr|mn|sn)"

# Check firmware slot history for recent updates
nvme fw-log /dev/nvme0

# Rule out media degradation: these should be flat and near zero
nvme smart-log /dev/nvme0 | grep -E "media_errors|num_err_log_entries|critical_warning|temperature"

# Rule out PCIe transport faults: all counters should be zero
cat /sys/class/nvme/nvme0/device/aer_dev_correctable
cat /sys/class/nvme/nvme0/device/aer_dev_fatal
cat /sys/class/nvme/nvme0/device/aer_dev_nonfatal

# Check the error log for what the controller recorded before each hang
nvme error-log /dev/nvme0

Note the ordering. Capture the firmware version first, before any reset, workaround, or swap. Once you replace the drive, the evidence of which firmware was running is gone, and you need it for the vendor case and for fleet-wide exposure assessment.

How to diagnose it

  1. Confirm the loop shape. In journalctl -k, you should see I/O <N> QID <N> timeout, reset controller messages, repeating with a period of minutes. Between resets, I/O completes normally. If resets appear without preceding I/O timeouts, or the device disappears entirely instead of resetting, you are looking at a different failure (see the lookalikes above).

  2. Rule out thermals. Check composite temperature against the drive’s WCTEMP. A thermal reset comes with temperature near or above the warning threshold and TMT transition counters climbing. In a firmware hang loop, temperature is normal and stays normal through the resets.

  3. Rule out the PCIe transport. AER correctable and uncorrectable counters must be zero or flat. Any uncorrectable error, or a rising correctable rate, points at the physical layer: connector, riser, retimer, slot. Also confirm current_link_speed and current_link_width still match max. A degrading link produces retransmissions and latency, not a clean timeout-then-reset cycle.

  4. Rule out media failure. media_errors flat, critical_warning zero, available spare stable. Media failure produces errors and read retries, not a silent controller that stops answering commands.

  5. Correlate the trigger. Look at what was running when each hang started. The triggers are specific: a command sequence, a queue depth, an I/O size, or an APST transition. If resets cluster around idle periods or wake-from-idle, suspect APST. If they cluster under one application’s load, suspect a workload-pattern trigger.

  6. Check the firmware version against known issues. Firmware hang loops are typically version-specific. Search the vendor’s release notes and advisories for your exact fr revision. Also check nvme fw-log for a recent update: a loop that started right after a firmware commit points at the new firmware, not the workload.

  7. Assess recovery quality. After each reset, does the controller return to live and stay there until the next trigger? If recovery starts failing (state stuck in resetting beyond about 30 seconds, or transitioning to dead), escalate immediately. The threshold: PAGE at two or more resets in an hour with failed recovery. At that point you are one bad cycle from losing the device mid-write.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Reset count from kernel log (timeout, reset controller per hour)The primary signal; not tracked in SMART>= 2 in an hour, or any failed recovery
Controller state (/sys/class/nvme/nvmeX/state)Ground truth on whether recovery succeededAnything but live sustained > 30s
Pre-reset I/O latencyThe latency spike to the timeout value precedes every hangLatency climbing toward 30s on a normally sub-millisecond device
Error log entries rateCaptures non-media errors a firmware bug generatesRising rate with zero media errors
Media errors rateRule-out signal for the loop patternAny increase means this is not a clean firmware hang
PCIe AER countersRule-out signal for transport faultsAny uncorrectable error or rising correctable rate
Composite temperatureRule-out signal for thermal resetsNear or above WCTEMP during reset events
Firmware version (fleet-wide)Exposure tracking once a bad version is identifiedAny host still running the affected revision

Fixes

Disable APST (workaround, not a fix)

If resets correlate with idle transitions, disable autonomous power state transitions:

# Disable APST on the controller
nvme set-feature /dev/nvme0 -f 0x0c -v 0

This is disruptive-adjacent rather than destructive, but test it on one host first. If the loop stops, APST was the trigger. The underlying firmware bug is still there; you have only removed the stimulus. Treat this as a bridge to a firmware update, and if you keep it, make it persistent across reboots via the kernel parameter nvme_core.default_ps_max_latency_us=0 (kernel cmdline or /etc/modprobe.d/); a runtime set-feature alone is re-programmed by the kernel at controller initialization.

Modify the I/O pattern (workaround)

If the trigger is a specific queue depth or I/O size, changing the workload shape can stop the loop: different I/O scheduler settings, reduced queue depth, or batching changes at the application layer. This is only practical if you have identified the trigger, and it carries the same caveat as APST: you are avoiding the bug, not fixing it.

Firmware update (the actual fix)

If a newer firmware revision addresses the hang, schedule a maintenance window and update. Firmware updates on NVMe are controller-level operations: brief I/O interruption is expected, and a controller reset or activation step is part of the process. Verify the new revision afterward with nvme id-ctrl and record the change in nvme fw-log. Keep the old version string in your records; if the new firmware introduces its own regression, you need the before-and-after.

Hardware replacement

If no fixed firmware exists, or the vendor confirms the revision you are on cannot be patched, replace the drive. Before swapping, capture the firmware version, serial number (nvme id-ctrl | grep sn), SMART log, and error log from the old device. You need these for the RMA and for confirming whether the replacement ships with the same affected firmware.

What not to do

Do not keep rebooting or reloading the driver as remediation. A host-initiated reset clears the current hang but the trigger recurs, and every cycle is another 5 to 30 second I/O stall against your applications. Do not raise io_timeout to mask the symptom either: it stretches each stall window without addressing why the controller stopped answering.

Prevention

  • Track firmware versions fleet-wide. Once one drive hits a firmware hang loop, every drive on the same revision is exposed. Firmware version tracking is operational hygiene, and this failure mode is exactly why. Unexpected firmware changes should also alert.
  • Alert on resets, not just device loss. A single reset is a TICKET. Two or more in an hour, or any reset with failed recovery, is a PAGE. Resets are only visible in kernel logs, not SMART, so you need log-pattern monitoring, not just device polling.
  • Baseline controller state. Alert on /sys/class/nvme/nvmeX/state leaving live for more than 30 seconds. This catches stuck resets and failed recoveries faster than waiting for filesystem errors.
  • Watch pre-reset latency. A latency trend climbing toward the timeout value, on a device that normally answers in microseconds, is your earliest warning that the next hang is coming.
  • Stage firmware updates. Roll new firmware to a small cohort first and watch reset counts and error log rates before fleet-wide deployment. A firmware hang loop that starts right after an update is a rollout you want to halt, not propagate.
  • Record workload context at each incident. The trigger is workload-specific. Noting which application, I/O size, or queue depth was active when each hang occurred is what lets you, or the vendor, reproduce it.

How Netdata helps

  • Netdata’s NVMe collector tracks media errors rate, error log entries rate, and critical warning bits continuously, which gives you the rule-out evidence in one view: a clean SMART picture alongside active resets is the firmware hang fingerprint.
  • Composite temperature and thermal management transition charts let you confirm or exclude a thermal cause in seconds instead of correlating logs by hand.
  • Per-device I/O latency and throughput charts show the pre-reset spike and post-reset recovery pattern, so you can see the loop’s rhythm and measure how long each stall lasts.
  • Unsafe shutdown and power cycle counters stay flat through a reset loop, which helps distinguish software recovery (controller reset) from power events.
  • Alerting on state and error-rate signals means the second reset of the hour pages you before the third one takes the device offline mid-write.