The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nvidia-gpu / nvidia-gpu-monitoring-checklist ▌

Operations Guides

NVIDIA GPU monitoring checklist: the signals every production GPU fleet needs

Most GPU monitoring setups fail in one of two ways. Either they collect a handful of nvidia-smi counters and page on the wrong things (raw temperature, memory percentage, idle PCIe downgrade), or they collect everything and alert on nothing meaningful. Both failure modes come from the same root cause: no shared vocabulary for which signals matter, at what fidelity, and with what alert semantics.

This checklist organizes the signals a production GPU fleet needs into four tiers: survival, operational, mature, and expert. Each tier builds on the previous one. The intent is not that every fleet reaches expert. The intent is that you can place your current setup on the ladder and know which gaps to close next, in what order.

One thing before the tiers: many GPU signals are persistent state, not events. Retired page counts, aggregate ECC counters, and row-remap failure flags are stored in the GPU’s InfoROM and stay nonzero forever. Alerting on the raw value means paging forever on a GPU that was remediated months ago. Where a signal is state, this checklist says so, and the correct pattern is alerting on deltas or transitions. This distinction is the single biggest source of alert fatigue in GPU fleets.

flowchart TD
  L4["Level 4: Expert - DCGM depth, NCCL, fabric, predictive"]
  L3["Level 3: Mature - retired pages, row remap, NVLink, drift"]
  L2["Level 2: Operational - power, throttle reasons, ECC, PCIe"]
  L1["Level 1: Survival - reachability, temp, VRAM, fatal XIDs"]
  L4 --> L3 --> L2 --> L1

Level 1: Survival

The minimum viable set. A fleet at Level 1 can answer two questions: is the GPU there, and is it about to die or corrupt data.

SignalHow to checkWarning sign
GPU reachabilitynvidia-smi -L exit code, per-GPU with -i NNon-zero exit, timeout, or hang on one GPU
Driver responsivenessTime a lightweight query: timeout 5 nvidia-smi --query-gpu=gpu_name --format=csv,noheader -i 0Sustained latency above 2s; healthy is under 100ms
GPU die temperaturenvidia-smi --query-gpu=temperature.gpu --format=csv,noheader,nounitsSustained within 5 degrees C of the GPU’s max operating temp
Framebuffer used / totalnvidia-smi --query-gpu=memory.used,memory.total --format=csv,noheader,nounitsSustained linear growth over hours (leak), not the level itself
Fatal XIDs in kernel logdmesg -T | grep -i "NVRM: Xid" or journalctl -k | grep -i "NVRM: Xid"New XID 48 (double-bit ECC) or XID 79 (fallen off the bus)

Two Level 1 rules that prevent bad pages:

  • Gate reachability alerts on duration. Brief nvidia-smi failures during boot or driver reload are normal. Page only on sustained unreachability (for example, more than 60 seconds) and after an uptime gate. One stuck GPU can hang a whole multi-GPU nvidia-smi query, so always probe per-GPU with -i N.
  • Do not page on VRAM percentage. PyTorch and TensorFlow caching allocators deliberately hold 95 to 99 percent of framebuffer memory. High usage is normal. OOM detection belongs in application logs (CUDA out of memory), not in a memory threshold. Track the rate of change for leaks instead.

Level 2: Operational

Professional production monitoring. This tier adds the signals that explain why performance degraded, which is what Level 1 cannot do. Most fleets should be here.

SignalHow to checkWarning sign
Power draw vs limitnvidia-smi --query-gpu=power.draw,power.limit,power.default_limit,enforced.power.limit --format=csv,noheader,nounitsDraw sustained above 95 percent of enforced.power.limit and sw_power_cap active
Clock throttle reasonsnvidia-smi --query-gpu=clocks_event_reasons.hw_thermal_slowdown,clocks_event_reasons.hw_power_brake_slowdown,clocks_event_reasons.sw_power_cap,clocks_event_reasons.sw_thermal_slowdown --format=csv,noheaderhw_thermal_slowdown active during production compute, sustained over 60s
Current vs max clocksnvidia-smi --query-gpu=clocks.current.sm,clocks.max.sm,clocks.current.memory,clocks.max.memory --format=csv,noheader,nounitsSM clock below 80 percent of max under heavy load
ECC corrected + uncorrected countsnvidia-smi --query-gpu=ecc.errors.corrected.volatile.total,ecc.errors.uncorrected.volatile.total --format=csv,noheader,nounitsAny new uncorrected error (delta); accelerating corrected rate
PCIe link gen and widthnvidia-smi --query-gpu=pcie.link.gen.gpucurrent,pcie.link.gen.gpumax,pcie.link.width.current,pcie.link.width.max --format=csv,noheaderCurrent below max during active workloads
Fan speed (where applicable)nvidia-smi --query-gpu=fan.speed --format=csv,noheader,nounits0 percent with temperature rising, or sustained 100 percent with high temp
Process listnvidia-smi --query-compute-apps=pid,process_name,used_gpu_memory --format=csv,noheaderUnexpected processes; zombies holding memory after job exit
Persistence modenvidia-smi --query-gpu=persistence_mode --format=csv,noheaderAnything other than Enabled on a production node

Operational-tier rules that matter:

  • Alert on throttle reasons, not raw temperature. 80 degrees C under full load is normal for many GPUs. hw_thermal_slowdown active is the real signal; temperature tells you severity and direction. One hot GPU means a local cooling problem; all GPUs hot means HVAC.
  • Power at the limit is not automatically a fault. Datacenter operators often cap power.limit below power.default_limit intentionally for rack density. Only treat it as throttling when sw_power_cap is active or clocks and throughput actually degrade. Compare against enforced.power.limit, which is the real ceiling.
  • Idle PCIe downgrade is by design. Drivers drop the link generation at idle to save power (a Gen5 GPU showing Gen2 at idle is normal). Only alert when the link is degraded under load. Note the field name: pcie.link.gen.current is deprecated; use pcie.link.gen.gpucurrent.
  • ECC counters need event semantics. Page on a new uncorrected error (delta in the volatile counter, or a new XID 48/95 in the log) on a GPU with active or recent production work. A historical nonzero aggregate on a drained, quarantined card is an urgent ticket, not a 3 a.m. page. Volatile counters reset on driver reload, so track deltas, not absolutes.
  • Persistence mode is a checkbox with real consequences. Without it (nvidia-smi -pm 1, requires root, persist it via a systemd unit), the driver unloads between jobs: seconds of first-call latency, monitoring gaps, and scheduler races. It resets on reboot, so verify it rather than setting it once.

Level 3: Mature

Full-spectrum instrumentation. This tier is where you stop reacting to failures and start predicting them, and where alert semantics get subtle.

SignalHow to checkWarning sign
Retired pagesnvidia-smi --query-retired-pages=retired_pages.single_bit_ecc.count,retired_pages.double_bit_ecc.count,retired_pages.pending --format=csv,noheaderAny increase (delta); pending = Yes; approaching the 64-page hard limit
Row remapping (Ampere and later)nvidia-smi --query-remapped-rows=remapped_rows.correctable,remapped_rows.uncorrectable,remapped_rows.pending,remapped_rows.failure --format=csv,noheaderUncorrectable rows, pending remaps needing reboot, transition to failure = true
HBM memory temperaturenvidia-smi --query-gpu=temperature.memory --format=csv,noheader,nounitsApproaching the GPU’s memory thermal limit (returns N/A on GDDR cards)
NVLink status and errorsnvidia-smi nvlink -s, nvidia-smi nvlink -eAny CRC/replay errors; an expected link down during active multi-GPU work
PCIe throughput and replay errorsnvidia-smi dmon -s t -d 1; nvidia-smi -q -d PCIE | grep -i replaySustained above 80 percent of practical link bandwidth; rising replay rate
BAR1 utilizationnvidia-smi --query-gpu=memory.bar1.total,memory.bar1.used,memory.bar1.free --format=csv,noheader,nounitsAbove 90 percent in GPUDirect RDMA environments
ECC mode statusnvidia-smi --query-gpu=ecc.mode.current --format=csv,noheaderDisabled on a datacenter GPU processing production data
Full XID classificationKernel log monitoring with per-XID routingFatal: 48, 64, 79, 95. App bugs: 13, 31, 43. Informational: 45, 63, 92
Configuration driftQuery persistence mode, compute mode, ECC mode, MIG mode, power limits, driver and VBIOS versionsAny deviation from the per-node-class baseline

Note: nvidia-smi -q -d PCIE exposes PCIe replay counters on driver versions that support it (added in nvidia-smi v340 per the changelog). DCGM exposes PCIe replays as field DCGM_FI_DEV_PCIE_REPLAY_TOTAL (field 202). Verify field availability against your installed driver version with nvidia-smi -q -d PCIE | grep -i replay before wiring this into alerting.

Mature-tier rules:

  • Retired pages and row remap are state, not events. Counts are monotonic and persist across reboots in InfoROM. Alert on increases, never on the raw value. remapped_rows.failure is latched: once true it stays true, so alert on the transition (or on a new XID 64), not the boolean. A GPU with failure = true is RMA-eligible but keeps running.
  • 64 retired pages is the cliff. At the driver’s 64-page retirement limit, no further retirement is possible and the next double-bit error is unrecoverable. Start planning replacement well before that.
  • XID routing must be per-code. XID 79 (fallen off the bus) and new XID 48/95 events on a busy GPU are pages. XID 63 (row remap event) is ECC self-healing working correctly and should never page. XID 13/31/43 are application bugs, not hardware. Getting this mapping wrong sends operators chasing hardware when the code is at fault, or ignoring real hardware faults as noise.
  • Verify field names against your driver. The correct flags are --query-retired-pages and --query-remapped-rows (display flag -d ROW_REMAPPER), and the ECC fields are ecc.errors.corrected.* / ecc.errors.uncorrected.*. Field names from older docs or generated configs (ecc.errors.single_bit_total, -d REMAPPED_ROWS) do not exist and fail silently. Validate with nvidia-smi --help-query-gpu.
  • Zero ECC errors is not proof of health. It can mean ECC is disabled, which turns every memory error into silent data corruption. Always check ecc.mode.current before trusting error counters. Some cloud providers ship with ECC off.

Level 4: Expert

Deep specialization, mostly relevant at scale: multi-node training clusters, NVSwitch systems, MIG multi-tenancy, and predictive maintenance.

  • DCGM daemon liveness and enumeration. If you run DCGM, the daemon (nv-hostengine) is itself a signal: pgrep -x nv-hostengine for presence, dcgmi diag -r 1 for responsiveness. Compare dcgmi dmon -e 1001,1002 -c 1 against nvidia-smi -L and lspci -d 10de:; any enumeration mismatch means a container misconfiguration, MIG grouping issue, or a GPU on its way off the bus. A live daemon can still serve stale cached data when NVML calls block, so treat query latency as its own signal.
  • NCCL collective latency. Baseline-relative only: absolute thresholds are meaningless across topologies and message sizes. Latency above 2x baseline, or NCCL timeouts, ticket. NCCL can silently fall back from NVLink to PCIe; training still runs, just much slower. Zero NVLink utilization during multi-GPU training is a classic silent-catastrophe pattern.
  • IB/RoCE fabric health. perfquery and ibdiagnet for InfiniBand, ethtool -S error counters for RoCE. Fabric faults are invisible from nvidia-smi and surface only as slow collectives or timeouts, so they are routinely misattributed to GPUs.
  • Fabric Manager on NVSwitch systems. On DGX/HGX, nv-fabricmanager down means NVLink connectivity through the switch is lost even though every GPU looks individually healthy. systemctl status nvidia-fabricmanager; page only with an uptime gate and active NVSwitch-dependent jobs failing.
  • MIG per-instance health. On MIG-enabled A100/H100, aggregate nvidia-smi numbers lie: one instance can be OOM while the card-level view looks fine, and instances share thermal and power domains. Monitor per-instance via DCGM (dcgmi dmon -g <mig-group>), not the physical card.
  • Reset recurrence and host-starvation correlation. Track GPU reset frequency per device; a GPU that resets more than twice in 24 hours is on a path to permanent failure. Correlate low SM utilization with host CPU, disk, and network saturation to catch host-starved GPUs before someone buys more GPUs to fix a data-loading problem.
  • Predictive wear modeling. Trend ECC corrected-error rate acceleration (not counts), thermal baseline drift month over month, and retired-page runway toward the 64-page limit. This is what turns GPU replacement from an incident into a scheduled task.

The alert-semantics cheat sheet

The same signal often needs different handling depending on whether it is an event or latched state. Keep this mapping next to your alert config:

PatternExamplesAlert on
Binary eventXID 79, new XID 48/95, NVLink down during active workNew occurrence, with uptime and duration gates
Latched stateremapped_rows.failure, aggregate ECC counters, retired page countsTransitions and deltas only
Level with contextTemperature, power draw, VRAM usedNever alone; corroborate with throttle reasons or application errors
Workload-gatedPCIe gen/width, straggler detectionOnly during active workloads
ConfigurationPersistence mode, ECC mode, compute mode, power limitsDrift from baseline, ticket severity

Common mistakes this checklist prevents

  • Paging on VRAM percentage when framework caching allocators make 95 percent usage normal.
  • Paging on raw temperature instead of throttle reasons.
  • Ignoring accelerating corrected ECC rates because “corrected means fine.” Acceleration is the leading indicator of an uncorrectable error.
  • Treating idle PCIe downgrade as a hardware fault.
  • Aggregating MIG metrics at the card level and missing per-instance OOM.
  • Alerting on raw retired-page and row-remap values, producing permanent re-alerts on remediated hardware.
  • Sampling once per minute. Temperature spikes, XID events, and PCIe errors live and die in seconds; sampling at 10 seconds or faster is the floor.

How Netdata helps

  • Netdata’s NVIDIA GPU collector queries NVML directly at per-second resolution, which is the sampling rate short-lived events like throttle transitions, power excursions, and temperature spikes actually require.
  • Die temperature, HBM temperature, power draw versus limit, clock speeds, and throttle reasons are charted together per GPU, so a thermal or power throttling cascade reads as one correlated view instead of five separate graphs.
  • ECC corrected and uncorrected counters are collected as time series, making rate-of-change and delta alerting (the correct semantics for these signals) straightforward rather than a custom scripting exercise.
  • Per-process GPU memory and utilization alongside host CPU, disk I/O, and network metrics on the same node makes host-starved-GPU and straggler diagnosis a correlation exercise instead of a guess.
  • Because Netdata also monitors the host, kernel logs, and systemd units, XID events, nv-fabricmanager state, and driver reloads land on the same timeline as the GPU metrics they explain.