The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nvidia-gpu / nvidia-gpu-mig-monitoring ▌

Operations Guides

Monitoring MIG-partitioned NVIDIA GPUs: per-instance metrics

On a MIG-enabled GPU (A100, A30, H100 and later), the physical card is partitioned into isolated GPU instances, each with its own SMs, memory partition, and failure domain. Every monitoring habit built around whole-GPU metrics breaks here: aggregate utilization, aggregate memory, and even some health counters describe the card, not the tenant. An instance can be OOM while nvidia-smi shows healthy free memory on the card.

The second surprise: nvidia-smi and NVML do not attribute utilization to MIG devices at all. Per-instance utilization reads as N/A by design. The supported path for per-instance metrics is DCGM, queried with MIG-aware entity types, and not every DCGM field is available at instance granularity.

This guide covers what MIG changes about the metric surface, which DCGM entities and fields work per instance, how to enable MIG-aware collection, and the version-dependent behavior you need to know before trusting the numbers.

Why per-GPU metrics mislead on MIG

MIG partitions compute and memory but not everything:

flowchart TD
    GPU["Physical GPU (A100 / H100)"]
    GPU --> GI1["GPU instance 0
SMs + dedicated HBM partition"] GPU --> GI2["GPU instance 1
SMs + dedicated HBM partition"] GPU --> GI3["GPU instance 2
SMs + dedicated HBM partition"] GI1 --> CI1["Compute instance(s)"] GI2 --> CI2["Compute instance(s)"] GPU -.-> SHARED["Shared: thermal domain, power budget,
ECC sources, PCIe/NVLink path, driver"]

Isolated per instance:

  • SM allocation (compute capacity)
  • HBM framebuffer partition (with hardware memory isolation between tenants)
  • L2 cache slices and memory controllers assigned to the partition

Shared across all instances on the card:

  • Thermal domain: one instance running hot heats the whole die
  • Power budget: all instances draw from one enforced power limit
  • The PCIe/NVLink path and the driver itself

This split dictates the monitoring model. Utilization and framebuffer usage must be read per instance, because each tenant only sees their partition. Thermal, power, and several reliability signals remain card-level, because the hardware domain is shared. You need both views: per-instance for capacity and workload health, aggregate for thermal/power and hardware degradation. For the underlying failure-model background, see How an NVIDIA GPU actually works in production.

The DCGM entity model

DCGM exposes the same metrics at different entity granularities. The entity type prefix determines what you are reading:

Entity prefixGranularityUse for
GPUPhysical device aggregateThermal, power, ECC, PCIe, XID context
GPU-IPer GPU instancePer-tenant utilization, framebuffer
CIPer compute instanceFiner-grained attribution within a GPU instance

The profiling field IDs are the same across entity types. What changes is whether the value is the card aggregate or the instance slice. A useful consequence: GPU-level graphics-engine activity is scaled by the instance’s SM allocation ratio. With two instances each holding 42 of 98 SMs, the card-level activity attributed to one instance is roughly (42/98) times the instance-level reading. Do not compare instance-level and card-level activity numbers directly without that scaling in mind.

One enumeration gotcha: MIG-enabled GPUs enumerate differently depending on DCGM configuration. The parent GPU may be invisible and only MIG instances appear, or vice versa. If your DCGM GPU count does not match nvidia-smi -L, check MIG mode and group configuration before assuming a hardware fault.

Which fields work per MIG instance

Field availability at MIG granularity depends on GPU generation and DCGM version. Treat this table as a starting point and verify against your DCGM version’s field documentation.

FieldPer-instance supportNotes
DCGM_FI_DEV_FB_FREE / DCGM_FI_DEV_FB_USED / DCGM_FI_DEV_FB_RESERVEDYesFramebuffer per instance partition. The core per-tenant memory signals.
DCGM_FI_PROF_GR_ENGINE_ACTIVEYes (profiling)Graphics engine active. Enabled by default in dcgm-exporter’s default counters CSV.
DCGM_FI_PROF_PIPE_TENSOR_ACTIVEYes (profiling)Tensor pipe activity. Default in dcgm-exporter CSV.
DCGM_FI_PROF_DRAM_ACTIVEYes (profiling)HBM activity for the instance’s memory partition. Default in dcgm-exporter CSV.
DCGM_FI_PROF_PCIE_TX_BYTES / DCGM_FI_PROF_PCIE_RX_BYTESYes (profiling)Default in dcgm-exporter CSV.
DCGM_FI_PROF_SM_ACTIVE / DCGM_FI_PROF_SM_OCCUPANCYAvailable but commented out by default in dcgm-exporterUncomment in the counters CSV if you need occupancy.
Other DCGM_FI_PROF_* (1000-series) profiling fieldsVariesNot all profiling fields are available at MIG granularity; support depends on GPU generation and DCGM version. Verify on your stack.
ECC error fields (DCGM_FI_DEV_ECC_SBE_VOLATILE / DCGM_FI_DEV_ECC_DBE_VOLATILE)Card-level in practiceECC sources are shared hardware; DCGM documents these fields at device granularity only. Correlate per-card.
NVLink bandwidth countersGPU device levelNVLink is a card-level resource on MIG-enabled GPUs; DCGM documents NVLink fields at device granularity. Treat per-instance NVLink bandwidth as not available.

Two field-ID cautions. First, field IDs and names have shifted across DCGM releases; confirm names against dcgmi dmon -l or your version’s field reference rather than hardcoding IDs from documentation written for another release. Second, requesting an unsupported field typically returns a blank or zero value without an error, which is worse than an explicit failure: your dashboard shows a plausible-looking flatline.

Enabling MIG-aware collection

Prerequisites:

  • MIG mode enabled and instances created (nvidia-smi mig -lgi lists GPU instances, nvidia-smi mig -lci lists compute instances).
  • DCGM installed. The MIG user guide recommends DCGM v2.0.13 or later for GPU utilization metrics and v3 or later for MIG device monitoring; current releases are 4.x.
  • nv-hostengine running as root. Profiling metrics require administrator privileges; an unprivileged host engine will silently lack the DCGM_FI_PROF_* fields.
  • On Ampere and earlier, profiling metrics require the datacenter-gpu-manager-4-proprietary package. It is bundled in the dcgm-exporter container but needs separate installation for host deployments.

Manual verification with dcgmi dmon, using entity-aware queries against a group containing your MIG devices:

# Confirm MIG instances exist and note their IDs
nvidia-smi mig -lgi
nvidia-smi mig -lci

# Watch profiling fields per MIG entity (group must include the MIG devices)
dcgmi dmon -e 1001,1002,1003,1004 -g <mig-group>

For dcgm-exporter (the Prometheus path), the default default-counters.csv already enables the core profiling fields listed in the table above. If you need SM_ACTIVE or SM_OCCUPANCY, uncomment them in your counters file and restart the exporter. XID error and clock-event cumulative counters are also opt-in and commented out by default.

One tool to avoid for this job: dcgmi stats does not support profiling metrics for MIG devices. SM utilization and memory utilization come back as N/A or Not Found for MIG instances. This is a confirmed limitation (NVIDIA DCGM GitHub issue #58). Use dcgmi dmon with entity prefixes instead.

Version and generation differences

These differences bite during fleet upgrades and mixed-generation clusters:

  • MIG mode persistence. On Ampere (A100, A30), MIG mode persists across reboots via InfoROM. On Hopper and later (H100, H200, B200), MIG mode is not persistent: it resets when the driver reloads. After any driver reload or reboot on Hopper+, instances are gone until you re-enable MIG and recreate them. Most teams handle this with a systemd unit that reapplies the MIG configuration at boot. This also means your per-instance time series will have gaps and re-created instance identities after every reload; do not alert on the gap as a node failure.
  • GPU reset requirement. Enabling MIG mode requires a GPU reset on Ampere, which is disruptive: drain all workloads from the GPU first. On Hopper+, no reset is required.
  • Minimum driver versions for MIG (per the NVIDIA MIG user guide): A100/A30 need R525 (>=525.53), H100/H200 need R450 (>=450.80.02), B200 needs R570 (>=570.133.20).
  • On DGX systems, stop nvsm and DCGM services before enabling MIG mode, or the mode change fails with “In use by another client” errors.

Known gotchas

  • N/A utilization is by design, not a fault. NVML does not support utilization attribution to MIG instances. Any pipeline scraping nvidia-smi --query-gpu=utilization.gpu for a MIG device will produce N/A forever. Migrate the collection to DCGM entities.
  • DCGM_FI_PROF_SM_ACTIVE above 100% on MIG devices has been reported on NVIDIA forums (February 2024) and was not documented as fixed in DCGM 4.x at review time. Treat >100% readings as a known anomaly, not proof of a broken collector.
  • Utilization not reaching 100% under burn tests on partitioned A100s has been reported against dcgm-exporter (issue #639). If your instance shows a plateau below 100% under a known-saturating workload, check the issue before distrusting the workload.
  • Instance metrics lie by omission across tenants. One instance can be thermally throttled because a neighbor instance saturated the shared power budget. Per-instance metrics alone will not explain this; you need card-level power draw and throttle reasons alongside.
  • Placement failures are invisible in utilization data. A MIG profile that cannot be placed (wrong granularity, fragmented free slices) produces no metric at all, just a failed creation command. Alert on configuration drift from the expected instance set, not on the absence of metrics.

Signals to watch in production

SignalGranularityWhy it mattersWarning sign
Framebuffer used/free (FB_USED, FB_FREE)Per instanceInstance OOM is a cliff; the card can look healthy while one tenant is out of memory>90% of the instance partition sustained, or steady growth (leak)
PROF_GR_ENGINE_ACTIVEPer instanceActual compute activity per tenantZero during scheduled work; unexpected 100% on a tenant that should be idle
PROF_DRAM_ACTIVEPer instanceMemory-bandwidth saturation of the instance partitionSustained near-saturation with low SM activity (memory-bound)
PROF_PIPE_TENSOR_ACTIVEPer instanceTensor-core engagement for ML workloadsNear zero on a training tenant (possible wrong-precision path)
Card-level temperature and throttle reasonsPer GPUShared thermal domain; one tenant’s heat throttles everyonehw_thermal_slowdown active while no single instance looks busy
Card-level power draw vs enforced limitPer GPUShared power budget; cross-tenant contentionsw_power_cap active plus tenant throughput complaints
ECC corrected/uncorrected countsPer GPUHardware degradation under all tenantsNew uncorrected event (page); accelerating corrected rate (ticket). See NVIDIA GPU ECC errors
MIG configuration vs expected instance setPer GPUPlacement failures and drift are silentInstance count or profiles differ from the declared baseline

How Netdata helps

  • Netdata’s NVIDIA GPU collector surfaces per-device utilization, framebuffer, temperature, power, and throttle reasons at per-second resolution, which is what you need to catch the cross-tenant thermal and power contention that per-instance tools cannot see.
  • On MIG systems, the card-level view and the per-instance view answer different questions; correlating a tenant’s throughput drop against card-level sw_power_cap or thermal throttle events is usually the fastest path to “my instance is fine, the card is contended.”
  • Framebuffer tracking per device catches the slow-growth leaks that precede instance OOMs, and pairs naturally with the fragmentation failure mode covered in CUDA out of memory with free memory available.
  • ECC counter and XID-aware alerting at the card level catches the shared-hardware degradation that affects every tenant on the GPU regardless of partition boundaries.
  • Anomaly detection on utilization and DRAM activity per device helps flag tenants behaving outside their historical envelope, which on multi-tenant MIG nodes is often the first sign of a runaway job.