The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / smartctl-disk-monitoring / smartctl-blind-spot-vm-usb ▌

Operations Guides

SMART blind spots: VMs, USB bridges, and drives you think you're watching

SMART monitoring is deployed, the dashboard is green, but some drives are invisible to smartctl. No alert fired. The pattern: expected physical drive count exceeds drives returning SMART data. The drives may be healthy or failing. You cannot tell because no telemetry is collected, and the monitoring system treats “smartctl returned no data” as “no problem.”

Four configurations create this gap: virtual machines, USB-attached drives, cloud block devices, and hardware RAID controllers. In each case, smartctl cannot reach the physical drive’s firmware. The fix is to move monitoring to where SMART data is accessible and use the correct device-type flags.

What a SMART blind spot is

This is a TICKET-level instrumentation gap, not a drive failure. No individual drive has necessarily failed. But drives can fail silently inside the blind spot, and you will not know until the filesystem reports I/O errors or the application loses data.

Detection requires two inputs: an inventory of expected physical drives per host (from the hardware manifest, RAID controller enumeration, lsscsi, lspci, or lsblk), and the set of drives where smartctl successfully returns SMART data. Any expected drive absent from the successful set is a blind spot.

The trap: most monitoring tools report absence of data as absence of problems. The dashboard stays green. No alert fires. A drive could accumulate hundreds of reallocated sectors while the monitoring system shows nothing wrong because it never collected data from that drive.

flowchart TD
    A[SMART data missing for a physical drive] --> B{How is the drive connected?}
    B -->|Inside a VM| C[Monitor on the hypervisor host]
    B -->|USB enclosure| D[Try -d sat, then vendor flags]
    B -->|Cloud block storage| E[No physical SMART to read]
    B -->|RAID controller| F[Use -d megaraid,N or cciss,N]

Virtual machines: the hypervisor wall

Hypervisors do not pass SMART commands to guest VMs. The VM sees a virtual disk (VMware Virtual disk, VBOX HARDDISK, Hyper-V virtual disk) that presents no SMART capability. Running smartctl -i inside the guest returns:

SMART support is: Unavailable - device lacks SMART capability.

The hypervisor presents a block device abstraction, not a physical drive. SMART commands operate below the virtualization layer and are not forwarded to the guest. Even when a physical disk is passed through via virtio or SCSI emulation, SMART commands may fail with “unsupported scsi opcode.” The virtualization layer translates block I/O but does not forward the ATA PASS-THROUGH or NVMe passthrough commands that smartctl relies on.

Run monitoring on the hypervisor host. For Proxmox, ESXi, or KVM, run smartctl on the host OS, not inside any guest. The hypervisor host has direct access to the physical SATA, SAS, or NVMe controllers.

PCIe controller passthrough is the only reliable VM-level option. Pass through the entire SATA or SAS controller to the guest via PCIe passthrough. The guest then talks directly to the hardware, and SMART commands reach the physical drives. Disk-level passthrough (passing individual drives via virtio or SCSI emulation) does not work for SMART.

USB bridges: the pass-through lottery

USB-to-SATA and USB-to-NVMe bridges are the most common source of SMART blind spots in backup servers and external storage. The bridge chip sits between the host and the drive, and not all bridges forward the ATA or NVMe commands that smartctl sends.

When smartctl cannot identify the USB bridge, it prints:

/dev/sdX: Unknown USB bridge [0xVENDOR:0xPRODUCT (0xREV)]

The bridge’s USB vendor/product ID pair is not in smartmontools’ drivedb.h database. The default auto-detection (-d auto) only works for USB IDs already catalogued there.

SATA drives behind USB

Try these device-type flags in order:

PriorityFlagWhen to use
1-d satStandard SCSI-ATA Translation. Works for most modern bridges.
2-d sat,12Older kernels or bridges that require 12-byte CDBs.
3-d sat,16Explicitly force the 16-byte CDB; SAT already defaults to 16-byte unless 12 is selected.
4-d usbcypressCypress bridge chips.
5-d usbjmicronJMicron bridges.
6-d usbjmicron,xJMicron bridges when you need 48-bit ATA commands (required for -l xerror). The ,x suffix is disabled by default.
7-d usbsunplusSunplus bridges.

The -d sat,auto variant only uses SAT if the SCSI INQUIRY data reports a SATL (the vendor string is “ATA”). Otherwise it falls back to the SCSI device type. This can help avoid false matches on bridges that partially implement SAT.

NVMe drives behind USB

NVMe drives behind USB bridges require completely different flags. The -d sat flag does not work for NVMe and returns “unsupported scsi opcode.” The correct device type depends on the bridge chipset:

Bridge chipsetFlag
JMicron JMS583-d sntjmicron
ASMedia ASM2362-d sntasmedia
Realtek RTL9210/1-d sntrealtek

For dual-protocol enclosures that support both NVMe and SATA, smartctl 7.5 introduced combined flags like -d sntjmicron/sat and -d sntasmedia/sat. If the NVMe Identify command fails, the device type falls back to -d sat. This is useful when you are not sure whether the enclosure holds an NVMe or SATA drive.

Updating the drive database

If your bridge is unknown, run update-smart-drivedb to fetch the latest drivedb.h from the smartmontools repository. Newly added bridge IDs may be recognized without upgrading smartmontools itself. This is the first thing to try when you see “Unknown USB bridge” on a chipset that should be supported.

When nothing works

Some older bridge chips do not support SMART pass-through at all. No -d flag will help. The only option is to replace the enclosure with one that uses a known-compatible bridge, or connect the drive directly via SATA or SAS.

Cloud block devices: virtualized by design

AWS EBS, GCP Persistent Disk, and Azure Managed Disks are virtualized block devices. They do not expose physical drive SMART data because there is no single physical drive to report against. Running smartctl against these devices returns nothing meaningful or fails silently.

If your workload runs on cloud block storage, you cannot monitor disk health with SMART. Rely on the cloud provider’s volume health metrics, I/O error rates from the instance OS, and your application’s own error handling.

Exception: instance-store NVMe. AWS documents NVMe SSD instance store volumes on some instance types. They may expose standard SMART/Health data when the instance OS can issue NVMe admin commands; verify each instance type and AMI with smartctl -a /dev/nvme0n1. If it succeeds, monitor fields such as Available Spare, Percentage Used, and Media and Data Integrity Errors.

Hardware RAID controllers: the abstraction layer

RAID controllers (LSI/Broadcom MegaRAID, HP Smart Array, Adaptec) present a virtual SCSI device to the OS. Standard smartctl /dev/sda queries this virtual device, not the individual physical drives behind the controller. SMART is either unavailable or returns controller-level data that masks per-drive health.

Use controller-specific passthrough (requires root):

# MegaRAID / LSI / Broadcom: access physical drive N
smartctl -a -d megaraid,0 /dev/sda
smartctl -a -d megaraid,1 /dev/sda

# HP Smart Array (cciss-based)
smartctl -a -d cciss,0 /dev/sda

# SAT passthrough through MegaRAID for SATA drives behind the controller
smartctl -a -d sat+megaraid,0 /dev/sda

The number N is the physical drive slot as enumerated by the controller. Find the mapping with the controller’s management utility (storcli, megacli, or ssacli).

Even with correct passthrough flags, some controllers filter or cache SMART data. Results may be stale or incomplete compared to direct SATA or SAS access. Cross-check with the controller’s own health reporting, which tracks drive status, rebuild progress, and predictive failure at the controller level.

Detecting blind spots at fleet scale

  1. Build a hardware inventory. Enumerate every expected physical drive per host using lsscsi, lspci (for NVMe), lsblk, or the RAID controller’s drive enumeration.
  2. Query each expected device. Run smartctl with the correct -d flag for every drive, including RAID controller passthrough where needed.
  3. Compare counts. Expected drive count vs. drives returning valid SMART data. Any gap is a blind spot.
  4. Track changes over time. A new blind spot after infrastructure changes (new RAID controller, VM migration, enclosure swap) means the monitoring configuration needs updating.
# Enumerate block devices and their transport
lsblk -d -o NAME,TYPE,MODEL,TRAN

# First-pass SMART accessibility check for direct-attached devices.
# Does NOT detect drives behind RAID controllers; those need -d megaraid,N etc.
# Requires root. Globs match up to 26 SATA/SAS and 10 NVMe controllers.
for dev in /dev/sd? /dev/nvme?n?; do
    if [ -b "$dev" ]; then
        echo "=== $dev ==="
        sudo smartctl -i "$dev" 2>&1 | grep -iE "SMART support|Device Model|Product|Unavailable|Unknown USB"
    fi
done

Common blind spot triggers after infrastructure changes:

  • New RAID controller installed. Existing smartctl configurations stop returning data unless updated with -d megaraid,N or equivalent. The controller presents virtual devices that respond to basic queries but return no per-drive SMART.
  • Migration from bare metal to VM. SMART monitoring that worked on the physical host returns empty data inside the guest. The hypervisor wall blocks all SMART commands.
  • New USB enclosures deployed. Different bridge chips require different -d flags. A configuration that worked with the previous enclosure may silently fail with the new one.

A monitoring system that silently collects no data for a drive is worse than no monitoring at all, because it creates a false sense of coverage. The dashboard says everything is fine, but only because it has no data to say otherwise.

How Netdata helps

  • Device presence tracking. Netdata’s SMART collector runs smartctl on a schedule and tracks which devices return valid data. A device that previously reported and stops responding is visible as a gap.
  • Telemetry blind spot detection. The SMART Telemetry Blind Spot signal compares expected physical drive count against successfully monitored drives. When the monitored count drops below expected, Netdata raises a TICKET-level alert flagging the instrumentation gap before it masks a real failure.
  • Host-level I/O correlation. When a drive is in a blind spot, Netdata still collects host-level I/O metrics from the kernel. Latency spikes, I/O errors in dmesg, and queue depth changes may surface the problem even when SMART data is unavailable.
  • Per-device flag configuration. Netdata’s collector supports custom smartctl arguments per device, so drives behind RAID controllers or USB bridges can be queried with the correct -d passthrough flag.
  • Hypervisor host deployment. Netdata runs on the host OS. When installed on the hypervisor host (Proxmox, ESXi, KVM), it has direct access to physical drive SMART data, bypassing the VM blind spot entirely.