The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nvme / nvme-io-queue-depth-saturation ▌

Operations Guides

NVMe queue depth saturation: command slots, io_timeout, and deep queues

You are looking at a host where NVMe latency has climbed, iostat shows the device busy, and nothing is erroring. No media errors, no kernel I/O error lines, SMART looks clean. The question is whether the drive is simply working as hard as it can, or whether something is wrong inside it. Queue depth is the signal that separates those two cases.

NVMe was designed for deep parallelism: each CPU core typically gets its own submission queue (SQ) and completion queue (CQ) pair, and a single queue can hold up to 64K command entries. Most enterprise drives reach peak throughput somewhere between QD 64 and QD 256 across all queues. Past that point, more outstanding commands buy nothing but latency: every extra command sits in a queue slot waiting for the controller to drain the ones ahead of it, and completion time grows linearly with how deep you stack.

When the controller cannot drain queues fast enough for long enough, the situation escalates from slow to broken: individual commands age past the kernel’s I/O timeout (default 30 seconds), the driver logs a timeout, aborts the command, and eventually resets the controller. That escalation has a specific signature in the kernel log, and the queue depth telemetry before it tells you why it happened.

What this means

An NVMe command occupies a slot in a submission queue from the moment the host writes it until the controller posts a completion entry. “Command slots” is just that capacity: queue depth per queue times the number of queues. Saturation means outstanding commands are piling up faster than the controller completes them, so the queue stays full and new commands wait.

There are two fundamentally different reasons a queue can stay full:

  1. The device is at its legitimate ceiling. The workload is issuing more parallel I/O than the drive’s rated parallelism. Throughput is at or near spec, latency rises smoothly with depth. Nothing is wrong with the drive; the fix is at the workload or capacity layer.
  2. The drive is draining slowly. Garbage collection, thermal throttling, SLC cache exhaustion, or a firmware problem has cut the controller’s internal completion rate. The queue backs up as a symptom. The giveaway: host-visible throughput drops while the queue stays deep.

The diagnostic value of queue depth is the relationship between depth, latency, and throughput, not any one of them alone:

flowchart TD
  A[QD rising] --> B{Latency behavior}
  B -->|Latency stable, throughput rising| C[Healthy: exploiting more parallelism]
  B -->|Latency rising, throughput at ceiling| D[Saturated by legitimate load]
  B -->|Latency rising, throughput LOW| E[Sick drive: GC, thermal, SLC cliff, firmware]
  D --> F{Sustained beyond io_timeout?}
  E --> F
  F -->|Yes, default 30s| G[Command timeout, abort, controller reset]
  F -->|No| H[Steady-state degradation, no errors logged]

Note the branch on the right: a deep queue by itself raises latency without producing a single error anywhere. If you only alert on errors, you can run for months in state D or E before the first timeout fires.

Common causes

CauseWhat it looks likeFirst thing to check
Legitimate load exceeds device parallelismQD high, latency elevated, throughput at spec ceiling, no errorsCompare observed IOPS/throughput to the drive’s rating for your I/O size and mix
GC stall / write cliffWrite throughput drops 50-90% abruptly, latency spikes, temperature normal, no media errorsDrive fill level; whether TRIM/discard is enabled
Thermal throttlingLatency climbs as composite temperature climbs; throughput declines gradually; QD backs upnvme smart-log temperature and thermal management transition counters
Sick drive (low QD, high latency)Latency high even when QD is low; controller busy time high relative to host IOPScontroller_busy_time in SMART vs. actual host I/O rate
Host-side queue misconfigurationPer-queue imbalance across hctx, one core saturated, deep queue with mediocre throughputScheduler set to something other than none; IRQ/NUMA affinity
Sustained slot exhaustion past io_timeoutnvme nvmeX: I/O <N> QID <N> timeout in dmesg, then resetsdmesg timeout and reset pattern

Quick checks

All read-only, all safe to run during an incident.

# Instantaneous in-flight I/O count (field 9 of stat, ios_in_progress)
awk '{print $9}' /sys/block/nvme0n1/stat

# In-flight reads and writes separately
cat /sys/block/nvme0n1/inflight

# Per-hardware-context queue occupancy (one hctx per CPU queue set)
for f in /sys/kernel/debug/block/nvme0n1/hctx*/busy; do echo "$f: $(cat $f)"; done

# Current I/O timeout the driver is enforcing
cat /sys/module/nvme_core/parameters/io_timeout

# Look for timeout and reset escalation in the kernel log
dmesg | grep -i "nvme.*timeout\|Resetting controller"

# Average latency and I/O rate from block counters (two samples)
iostat -xp nvme0n1 1 3

# Is the controller itself busy when the host is not?
nvme smart-log /dev/nvme0 | grep -i "controller_busy_time\|temperature\|media_errors"

# Device-reported maximum queue capabilities
nvme id-ctrl /dev/nvme0 | grep -i "sqes\|cqes\|nn"

Two caveats on tools you may already be reaching for. iostat’s %util is not a saturation signal for NVMe: it measures the fraction of time the device had at least one request outstanding, and a massively parallel device can pin at 100% util while still having deep headroom. Likewise avgqu-sz often reads implausibly low on multi-queue devices. Use ios_in_progress and inflight instead.

How to diagnose it

  1. Establish the QD-latency-throughput triple. Sample ios_in_progress frequently (sub-second; it is an instantaneous snapshot, so single samples mislead) alongside computed average latency from /sys/block/nvme0n1/stat deltas and throughput from iostat. The rule: QD rising with stable latency is healthy parallelism; QD rising with rising latency is saturation; QD rising with rising latency and falling throughput is a sick drive.

  2. Do the load-shed test. If you can pause or throttle the workload briefly, watch latency as QD falls toward single digits. Latency that collapses back to baseline at low QD means the drive itself is fine and you were over-driving it. Latency that stays high at QD 1-4 means the problem is internal to the device (GC, thermal, firmware, media). This single test resolves most of the ambiguity.

  3. Check per-queue balance. Compare hctx*/busy values across hardware contexts. One queue doing all the work while others idle points at IRQ affinity or NUMA placement problems on the host, not the drive. Every queue equally deep points at device-level limits.

  4. Check the sick-drive candidates. Composite temperature climbing alongside latency (thermal throttle), drive over 80-90% full with discard disabled (GC thrash), a step-function write throughput cliff (SLC cache exhaustion), controller_busy_time high while host IOPS is low (internal contention). Any of these explains a slow drain rate.

  5. Check for timeout escalation. I/O <N> QID <N> timeout lines mean commands aged past io_timeout (default 30 seconds): the queue was effectively stuck, not just deep. Count occurrences and look at what preceded them: a reset loop, thermal events, or a load spike.

  6. Compare against device limits. nvme id-ctrl reports the controller’s queue entry capabilities (SQES/CQES). The kernel-side per-queue depth is set by the driver (the io_queue_depth module parameter; the default has been 1024 since the parameter was introduced in kernel 4.13, and the driver’s queue depth was 1024 before that). Saturation at the host queue layer looks different from saturation at the device layer: host-side slots filling while the device reports low utilization means the bottleneck is above the drive.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
ios_in_progress (stat field 9), time-weightedThe actual depth of outstanding I/OPersistently above your device’s optimal QD band (64-256) with rising latency
/sys/block/.../inflight (reads/writes split)Tells you which direction is backing upWrites pinned deep while reads flow: GC or SLC pressure
Average I/O latency from stat deltasThe cost of the depthLatency growing linearly with QD; any mean above 5x baseline
Throughput vs. rated specSeparates saturated from sickHigh QD + throughput well below spec = internal drain problem
controller_busy_time rate vs. host IOPSController-internal saturationController busy near 100% while host IOPS is low
dmesg timeout/reset linesSlot exhaustion escalationAny I/O timeout on a QID; more than one reset per week
Per-hctx busy distributionHost-side balanceOne queue hot, others idle
Composite temperature + TMT transitionsThermal as the drain-rate causeLatency tracking temperature upward

Fixes

Workload is legitimately over-driving the device

Reduce application-side concurrency to land in the device’s optimal band. For a database, that means tuning I/O thread counts and async I/O depth; for fio-based validation, cap iodepth. The target is the QD where throughput stops improving: past roughly QD 64-256 on most drives you are paying latency for zero throughput. If the workload genuinely needs more, the fix is capacity (more devices, a higher-class drive), not deeper queues.

You can also trim host-side queueing so backpressure lands on the application instead of inflating device latency. nr_requests in /sys/block/nvme0n1/queue/ controls how many requests the block layer stages per direction (on current kernels it defaults to the driver’s queue depth, i.e. 1024 for NVMe; the classic block-layer default used to be 128). Lowering it makes overload fail fast and visibly at the application layer instead of hiding as device-side latency.

Drive is draining slowly (GC, thermal, SLC cliff)

Attack the drain rate, not the queue:

  • GC thrash: enable continuous discard or schedule fstrim, and get the drive below roughly 80% fill. TRIM is what tells the FTL which blocks are free; without it the drive treats itself as perpetually full.
  • Thermal throttle: fix airflow or heatsinking. Throttling is self-protecting, so performance returns when temperature drops, but a drive that throttles daily is degrading.
  • SLC write cliff: this is by design. If sustained write throughput matters, either reduce burst duration or use a drive whose native TLC/QLC write rate meets your floor.

Host-side queue misconfiguration

Confirm the scheduler is none (cat /sys/block/nvme0n1/queue/scheduler); any other scheduler adds latency in front of a device that manages its own queues. Fix IRQ affinity so MSI-X vectors spread across cores, and keep I/O submission on the NUMA node local to the device. A single saturated hctx with idle siblings is a configuration problem, not a hardware one.

io_timeout is firing

Raising io_timeout (module parameter on nvme_core; 30s default) stops the resets but fixes nothing: the commands are still waiting 30+ seconds for completion. Treat a raised timeout as blast-radius control for workloads that would rather stall than see a reset (some cloud block devices recommend the maximum value for exactly this reason), never as a fix. If timeouts fire under normal load, the drain rate is the problem; work the sections above. If timeouts fire in a repeating reset cycle with no thermal or PCIe correlate, suspect firmware and check the vendor’s advisories for your firmware version.

Prevention

  • Baseline the QD-latency curve per drive model at deploy time, so “latency at QD 128” has a known-good reference. Deviation from your own curve is a far better alert than any absolute threshold.
  • Alert on the combination, not the depth. QD alone is noisy; QD persistently high with latency rising and throughput flat or falling is the actionable condition. Reserve paging for QD pinned near device maximum with completions stalled (pre-timeout state) or actual io_timeout firings.
  • Track the sick-drive precursors independently: drive fill level, discard configuration, composite temperature trend, and controller_busy_time ratio. These move before the queue backs up.
  • Size for the optimal band. If steady-state load needs more than QD 64-256 of parallelism to hit its throughput target, the device class is wrong for the workload and every future growth step will buy latency instead of IOPS.

How Netdata helps

  • Per-second block-layer I/O rates and derived average latency per NVMe namespace, so the QD-latency relationship is visible continuously rather than reconstructed from two iostat samples during an incident.
  • Device SMART signals (composite temperature, thermal management transitions, controller busy time, media errors) on the same timeline as throughput, which is exactly the correlation that separates a saturated drive from a sick one.
  • Temperature-to-throughput correlation that makes thermal throttling obvious: throughput declining in lockstep with rising composite temperature, with no errors.
  • controller_busy_time collected alongside host-visible IOPS, exposing the “controller busy, host idle” internal-contention state that block stats alone cannot show.
  • Kernel log monitoring context for io_timeout and controller reset events, so escalation from deep queues to timeouts is captured and alertable rather than buried in dmesg.