The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nvidia-gpu / nvidia-gpu-dcgm-daemon-down ▌

Operations Guides

DCGM nv-hostengine down or hung: GPU telemetry goes dark or stale

Your GPU dashboards flatline, or worse, they do not. The graphs keep rendering smooth, plausible values for temperature, utilization, and power, but nothing on the node has actually changed in twenty minutes. Both symptoms point at the same component: nv-hostengine, the DCGM daemon that sits between NVML and every monitoring tool you run.

nv-hostengine fails in two distinct ways, and only one of them is obvious. The obvious failure is a dead process: no data, gaps in every time series, alerts firing on absent metrics. The dangerous failure is an alive-but-hung daemon: the process exists, the socket accepts connections, and DCGM serves stale values from its in-memory field value cache because an underlying NVML call is blocked on the driver. Dashboards look normal. Threshold alerts stay quiet because the cached values are inside bounds. You find out when a human notices the numbers have not moved.

This guide covers detecting both states, telling a hung DCGM apart from a hung driver, and recovering without making things worse.

What this means

nv-hostengine is the core DCGM daemon. It runs as root, polls GPU state through NVML on background health threads, and stores recent samples in an in-memory field value cache. Clients, including dcgmi, dcgm-exporter, and any custom collectors, query the daemon, not the driver directly.

That layering is why a hang is so deceptive. When an NVML call blocks, for example because the driver is deadlocked or a GPU has wedged mid-query, the health thread stalls but the daemon process keeps running. Client queries are answered from the cache with the last good sample. Every downstream consumer sees fresh-looking data that is actually frozen in time.

Two consequences follow:

  • Process liveness is not proof of health. pgrep nv-hostengine succeeding, a listening socket, or an HTTP 200 from dcgm-exporter tells you nothing about whether real NVML queries are completing.
  • A DCGM hang can be a symptom, not the disease. A hung NVML call is often the first visible sign of a driver deadlock or a GPU hardware fault stalling the bus. Restarting the daemon without checking the driver can leave the new process hanging on initialization.
flowchart TD
  A[GPU telemetry looks wrong] --> B{nv-hostengine process present?}
  B -- No --> C[Dead daemon: check logs, restart, find crash cause]
  B -- Yes --> D{Active probe returns fresh values?}
  D -- Yes --> E[Daemon healthy: check exporter and scrape config]
  D -- No or unchanged values --> F{nvidia-smi per-GPU query responds?}
  F -- Yes --> G[Wedged daemon: restart nv-hostengine]
  F -- No --> H[Driver/GPU fault: do NOT restart DCGM alone]
  H --> I[Check dmesg for Xid, plan driver reload or reboot]

Common causes

CauseWhat it looks likeFirst thing to check
Daemon crashed or was killedNo process, gaps in all DCGM-derived metricspgrep -x nv-hostengine, systemd/journal logs
Alive-but-hung on a blocked NVML callMetrics update on schedule but values never change; dcgmi commands hangRun a field query with a timeout and compare two samples
Driver deadlock or GPU wedge underneath DCGMAll GPUs stale simultaneously, nvidia-smi also hangs, Xid errors in dmesgdmesg -T | grep -i "NVRM: Xid", per-GPU nvidia-smi with timeout
Driver updated without restarting nv-hostengineDaemon holds a stale NVML reference; queries fail or misbehave after a driver upgradeCompare daemon start time against driver install time
Systemd restart loop masking a crashProcess exists but start time is minutes old; metrics flappingProcess uptime vs node uptime, restart counter
Kubernetes probe blind spotdcgm-exporter pod Ready, /health returns 200, but metrics staleQuery actual DCGM field values, not the HTTP endpoint
Version mismatch between exporter and hostengineInstability, in reported cases a segfault of the hostengineMatch dcgm-exporter and DCGM package versions

Quick checks

All of these are read-only and safe.

# 1. Is the process there at all, and how long has it been running?
pgrep -x nv-hostengine
ps -o pid,etime,cmd -C nv-hostengine

A very short etime on a node that has been up for weeks means the daemon recently restarted. Find out why before treating its return as recovery.

# 2. Active liveness probe: run a real query with a hard timeout
timeout 10 dcgmi dmon -e 1001,1002 -c 1
echo "exit: $?"

This is the check that matters. It forces a field query through the daemon. Exit code 124 (timeout) or a hang means the daemon is not completing queries even though the process is alive. A healthy daemon returns in well under 2 seconds.

# 3. Freshness check: take two samples and compare
dcgmi dmon -e 1002 -c 2 -d 5000

If the value is identical across samples on a node under load, treat the data as stale. A value unchanged across intervals, or older than roughly 2x the collection interval, is suspect.

# 4. Cross-check with the driver directly, per GPU, with a timeout
timeout 5 nvidia-smi --query-gpu=gpu_name --format=csv,noheader -i 0

If nvidia-smi answers quickly while DCGM hangs, the wedge is in the daemon. If nvidia-smi also hangs, the problem is the driver or the GPU, and DCGM is a victim. Always query per-GPU with -i N: one wedged GPU can hang a whole-node query.

# 5. Look for driver-level faults
dmesg -T | grep -i "NVRM: Xid" | tail -20
# 6. Rapid DCGM diagnostic as a deeper responsiveness check
timeout 30 dcgmi diag -r 1
# 7. On Kubernetes: what is the probe actually testing?
kubectl describe pod -n gpu-operator -l app=nvidia-dcgm-exporter | grep -A4 -i "readiness\|liveness"

If the probe is an HTTP GET on port 9400 /health, it proves the web server is up, not that DCGM is answering.

How to diagnose it

  1. Classify the failure: dead or hung. Run check 1. No process means dead: go to step 5. Process present means you must distinguish hung from healthy with an active probe, not with liveness signals.

  2. Probe with a real field query under a timeout (check 2). A timeout or hang means the daemon is wedged. Capture the output of check 3 to confirm values are frozen rather than merely quiet.

  3. Determine whether the driver is also stuck (checks 4 and 5). This decides your recovery path. A wedged daemon over a healthy driver is a daemon restart. A wedged daemon over a hung driver is a driver or hardware incident. Do not restart DCGM alone against a hung driver: the new process will hang on initialization.

  4. Check timing correlations. Did a driver update land recently without a daemon restart? Did the daemon’s start time coincide with a crash loop? Did metrics go stale at the same moment across all GPUs (driver-level) or just some (possibly one GPU wedging a shared query path)?

  5. For a dead daemon, find the cause before restarting. Review the service logs and journal for a segfault, OOM kill, or fatal NVML error. Restarting without understanding the crash invites a repeat. After restart, verify GPU enumeration (dcgmi dmon -e 1001 -c 1 against nvidia-smi -L): a daemon that comes back seeing zero GPUs has a driver or permissions problem, not a monitoring gap.

  6. In Kubernetes, verify what “healthy” means. A dcgm-exporter pod can be Ready while the engine inside is wedged, because the default probes only exercise the HTTP server. Confirm staleness by comparing a metric value across two scrape intervals before concluding anything about the node.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
nv-hostengine process presence and uptimeThe floor: no daemon, no GPU telemetryProcess missing, or restart without an operational trigger
Active probe latency (dcgmi dmon/diag -r 1 under timeout)Proves real NVML queries complete, not just that the process existsResponse over 2s, or any timeout
Field freshness (value delta across samples)Catches the alive-but-hung case that liveness checks missIdentical values across intervals under load, or data older than 2x collection interval
Per-GPU nvidia-smi latency with timeoutSeparates daemon wedge from driver hangOver 2s sustained, or timeout on any GPU
Xid events in dmesgA DCGM hang is often the first symptom of a driver or GPU faultNew fatal Xids (48, 79, 95) or repeated timeouts
GPU enumeration count vs inventoryA restarted daemon may come back blind to some devicesFewer GPUs in DCGM than in nvidia-smi -L
Scrape success rate for dcgm-exporterDistinguishes exporter/network failure from engine failureScrapes succeeding with frozen values

Fixes

Dead daemon

Restart the service after capturing logs. Then immediately verify two things: the daemon answers an active probe (check 2), and it enumerates the expected GPUs. If the crash was a segfault or OOM kill, treat it as a bug or sizing issue to track, not a one-off. Keep DCGM and dcgm-exporter versions matched; mismatched versions have been reported to crash the hostengine itself.

Hung daemon over a healthy driver

Restart nv-hostengine. This is disruptive to monitoring but not to running GPU work: clients lose telemetry during the restart, and the field value cache is rebuilt, so expect 1-2 polling intervals of empty or zero data before values stabilize. Do not alarm on that warmup window. After restart, confirm freshness with a two-sample comparison.

Hung daemon over a hung driver

Do not just bounce DCGM. A driver deadlock or wedged GPU requires a coordinated driver reload or node reboot, and that means draining or checkpointing GPU workloads first. Collect evidence before you clear the state: hung process list, dmesg Xid history, and per-GPU query results. If a single GPU is implicated (one device hangs nvidia-smi -i N while others respond), Xid 79 and friends in dmesg will usually name it. Follow the driver-incident path in NVIDIA-SMI has failed because it couldn’t communicate with the NVIDIA driver.

After any driver update

Restart nv-hostengine as part of the driver upgrade procedure. The daemon can hold a stale NVML reference across a driver reload and misbehave in ways that look like a new bug but are just version skew between the daemon and the kernel module.

Prevention

  • Monitor liveness two ways, always. Process presence plus an active probe that executes a real DCGM field query under a timeout. Either alone is insufficient: the first misses hangs, the second is meaningless if the process is gone.
  • Add a freshness check to your pipeline. Alert when a value that should move under load (utilization, power, temperature) is unchanged across multiple collection intervals, or when the newest sample is older than 2x the collection interval. This is the only reliable catch for stale-cache failure.
  • Fix Kubernetes probes. A readiness probe that hits the exporter’s HTTP health port will keep a wedged pod in rotation. The probe must exercise DCGM itself. The upstream dcgm-exporter chart confirms the blind spot: its default liveness and readiness probes use an HTTP /health request on the service port.
  • Use DCGM’s built-in hang detection. Since DCGM 4.4.2, nv-hostengine has hang detection enabled by default; it logs when a hang is detected. The DCGM_HANGDETECT_TERMINATE environment variable escalates that to daemon termination, which converts a silent hang into a visible, alertable death. DCGM_HANGDETECT_EXPIRY_SEC tunes the timeout (minimum 120s, in 60s steps). Check these settings against the installed DCGM release, because the feature is new in 4.4.2 and module/diagnostic behavior can differ.
  • Pin and align versions. Deploy dcgm-exporter and the DCGM hostengine packages together, and restart the daemon on every driver update.
  • Alert on unexplained restarts. Any nv-hostengine restart outside a deployment or maintenance window is a signal, not a recovery.

How Netdata helps

  • Netdata’s NVIDIA GPU collector queries through NVML directly at per-second resolution, so a frozen DCGM layer does not silently freeze your only view of the node: you can compare Netdata’s live values against DCGM-derived dashboards to spot staleness.
  • Per-second temperature, power, and utilization series make freshness anomalies obvious: a healthy GPU under load never prints identical values for minutes.
  • Process monitoring of nv-hostengine (presence, uptime, restarts) turns unexplained daemon restarts into alertable events instead of silent flapping.
  • Xid and NVRM error visibility from kernel logs, correlated on the same timeline as GPU metrics, lets you see whether a telemetry freeze preceded or followed a driver-level fault.
  • Anomaly detection on GPU utilization and power flags the flatline pattern itself, which is the signature of the alive-but-hung failure that threshold alerts cannot catch.