The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nvidia-gpu / nvidia-gpu-thermal-cascade-dense-servers ▌

Operations Guides

NVIDIA GPU thermal cascade in dense servers: one hot GPU heats its neighbours

Three of the eight GPUs in a node are thermal throttling. Training step time has doubled. The instinct is to suspect the cooling system as a whole, or the workload. But line up per-GPU temperatures and one unit stands out: GPU 5 is running 12 degrees hotter than everything else, and the GPUs downstream of it in the airflow path are the ones throttling.

This is the thermal cascade pattern, specific to dense GPU servers (DGX, HGX, and similar trays) where GPUs share a cooling path. One GPU’s exhaust air is another GPU’s inlet air. A single unit with a local cooling problem raises the intake temperature of its neighbours, pushing them past their thermal limits. A throttled GPU still draws substantial power, so it keeps generating heat while producing less useful work. The system settles into a stable, badly degraded equilibrium: several GPUs throttled, all because of one.

The fix is almost never “reduce the workload.” Find the originating unit and its local cooling fault, then break the cascade there.

What this means

GPU thermal protection works in stages. As die temperature rises past the GPU Max Operating Temp, SW thermal slowdown reduces clocks. If temperature keeps climbing to the GPU Slowdown Temp, HW thermal slowdown applies an aggressive 2x or greater clock reduction. Past the Shutdown Temp, the GPU powers off. Exact thresholds are model-specific; read them from the hardware with nvidia-smi -q.

In a standalone server, thermal throttling is a single-GPU story. In a dense tray, the physics change:

  1. A GPU develops a local cooling problem: degraded thermal paste, a poorly seated heatsink, an airflow blockage, or a tray position that receives pre-heated air.
  2. Its die temperature climbs, and it begins to throttle.
  3. A throttled GPU does not stop consuming power. It typically continues drawing most of its power budget while delivering a fraction of its former throughput. Heat output stays high; useful work drops.
  4. That heat leaves the card as exhaust and becomes inlet air for the GPUs downstream in the chassis airflow path.
  5. Those GPUs start warmer, hit their own thermal limits earlier, and throttle too.
  6. The system reaches equilibrium at a degraded performance level. Every downstream GPU looks symptomatic, but only one is the cause.

The distinguishing feature is the shape of the temperature map: one GPU significantly hotter than the rest, usually in a specific physical position (middle or end of the tray), with the throttling GPUs clustered downstream of it. An all-GPUs-hot pattern is a different problem: it points to the environment (HVAC, ambient, blocked rack airflow), not a single unit.

flowchart TD
    A[GPU 5: local cooling fault] --> B[GPU 5 throttles but still draws power]
    B --> C[Exhaust heat raises inlet temp of downstream GPUs]
    C --> D[GPU 6 hits SW thermal slowdown]
    C --> E[GPU 7 hits SW thermal slowdown]
    D --> F[GPU 6 continues heating shared airflow]
    E --> F
    F --> G[GPU 7 escalates to HW thermal slowdown]
    G --> H[Equilibrium: multiple GPUs throttled, one root cause]

Common causes

CauseWhat it looks likeFirst thing to check
Degraded thermal paste or poorly seated heatsink on one GPUOne GPU much hotter than peers at the same workload; gap grows over monthsPer-GPU temperature spread; physical inspection and reseat
Airflow blockage inside the chassis (cabling, debris, missing baffle)One or two adjacent GPUs hot; fans at max but temps still climbingPhysical inspection of the tray; IPMI fan and inlet sensors
Unfavourable physical position (middle or end of tray receives pre-heated air)Same position runs hot across multiple nodes of the same modelCompare temperature maps across identical nodes
Chassis fan failure or degraded fan (air-cooled systems)Rising temps with fans reporting max speed, or a fan reporting 0Chassis fan speeds via IPMI/BMC sensors
Ambient or HVAC problemAll GPUs in the node (and likely the rack) hot togetherInlet air temperature via BMC/IPMI; room and CRAC status
Recirculation at rack level (missing blanking panels, hot air looping back)Inlet temps above spec on several nodes; worse at specific rack positionsRack inlet temperatures across rows

The critical split: one hot GPU means local cooling, all hot GPUs means environmental. Everything in diagnosis flows from that.

Quick checks

All of these are read-only and safe on a production node.

# Per-GPU die temperature, one line per GPU
nvidia-smi --query-gpu=index,temperature.gpu --format=csv,noheader,nounits

# Which throttle reasons are active right now, per GPU
nvidia-smi --query-gpu=index,clocks_event_reasons.active --format=csv,noheader

# Individual thermal and power throttle flags per GPU
nvidia-smi --query-gpu=index,clocks_event_reasons.sw_thermal_slowdown,clocks_event_reasons.hw_thermal_slowdown,clocks_event_reasons.sw_power_cap --format=csv,noheader

# Current vs maximum SM clocks: the gap shows throttle severity
nvidia-smi --query-gpu=index,clocks.current.sm,clocks.max.sm --format=csv,noheader,nounits

# Power draw vs the enforced limit: throttled GPUs still draw power
nvidia-smi --query-gpu=index,power.draw,enforced.power.limit --format=csv,noheader,nounits

# Detailed temperature view, including thresholds, for one suspect GPU
nvidia-smi -i 5 -q -d TEMPERATURE

# HBM temperature (datacenter GPUs only; returns N/A on GDDR cards)
nvidia-smi --query-gpu=index,temperature.memory --format=csv,noheader,nounits

On the chassis side, check inlet temperature and fan speeds through the BMC:

# Chassis inlet temperature and fan speeds (datacenter GPUs report fan N/A; the chassis fans matter)
ipmitool sensor list | grep -i -E "inlet|fan|temp"

And check whether the driver has logged anything relevant:

# GPU error events in the kernel log
dmesg -T | grep -i "NVRM: Xid"

How to diagnose it

  1. Build the temperature map. Pull per-GPU die temperatures and lay them out in physical slot order, not just index order. (Index-to-slot mapping is system-specific; check your platform documentation.) You are looking for the shape: one outlier versus a uniform rise.

  2. Read the throttle reasons, not just temperatures. A GPU at 83C with sw_thermal_slowdown active is throttling; a GPU at 80C with no throttle reasons active is not. hw_thermal_slowdown active for a sustained period during production compute is a paging condition. Note which GPUs are throttling and which are merely warm.

  3. Trace the cascade upstream. The throttling GPUs are victims. The originating unit is the hottest GPU upstream of them in the airflow path, typically middle or end of the tray in front-to-back designs. If the hottest GPU is also throttling hardest and its downstream neighbours are next worst, you have the cascade. If temperatures are uniform and high across all GPUs, skip to step 6: this is environmental.

  4. Check whether the origin GPU is still drawing power. Compare power.draw against enforced.power.limit. A throttled GPU still drawing most of its limit confirms the amplifier mechanism: it is heating the shared airflow without delivering throughput.

  5. Correlate with chassis data. Pull inlet temperature and fan speeds from IPMI/BMC. Inlet temp above spec across the node points to rack or room issues. Normal inlet temp with one GPU 10C or more above peers points to local cooling on that unit: paste, heatsink seating, or a blockage.

  6. Rule out the environmental case. If all GPUs are hot, check inlet air temperature against spec, HVAC status, and rack-level recirculation (missing blanking panels, hot-aisle air looping back to intakes). No amount of per-GPU work fixes a hot room.

  7. Physically inspect the originating unit. With the node drained and powered down per your platform’s procedure, check heatsink seating, thermal paste condition, airflow baffles and shrouds, and cable routing that might block airflow. If the same physical position runs hot across several identical nodes, suspect the tray position or a shared assembly fault rather than one bad card.

  8. Confirm the cascade is broken after the fix. After remediation, put the node back under representative load and re-map temperatures and throttle reasons. Downstream GPUs should return to full clocks within minutes once inlet air normalizes. The repaired unit should hold temperature within a few degrees of its peers.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Per-GPU die temperatureLeading signal: rises before clocks dropOne GPU more than roughly 10C above peers under equal load
clocks_event_reasons.sw_thermal_slowdownFirst-stage thermal throttle confirmationActive during production compute
clocks_event_reasons.hw_thermal_slowdownAggressive 2x+ clock reduction; cooling is failingActive and sustained beyond 60 seconds during compute
clocks.current.sm vs clocks.max.smQuantifies throttle severityBelow 80% of max under heavy load
power.draw vs enforced.power.limitShows the throttled-but-still-heating amplifierHigh draw with reduced clocks and low throughput
Per-GPU temperature spreadSeparates local faults from environmental onesSpread widening over weeks or months
Chassis inlet temperature (BMC/IPMI)Tells you what the GPUs are breathingAbove platform spec, or rising trend
Chassis fan speeds (BMC/IPMI)Cooling capacity response100% fans with temperatures still climbing
Temperature baseline trend over monthsCatches paste degradation and dust before throttling startsShrinking gap between peak-load temp and throttle threshold

Do not alert on raw temperature alone. 80C under full load is normal for many GPUs. Alert on throttle reasons, and use temperature for severity and triage context.

Fixes

Local cooling fault on one GPU

Reseat the heatsink and replace thermal paste. The most common root cause of the one-hot-GPU pattern, especially on hardware older than a couple of years. It requires draining the node and following your platform’s service procedure. Tradeoff: node downtime, but it is the actual fix rather than a workaround.

Clear airflow obstructions. Reroute cables, reinstall missing baffles or shrouds, clear dust. Often zero-cost and immediately effective.

Replace the chassis fan or repair the cooling loop if IPMI data shows a fan at 0 or a fan at 100% that is not moving air. On liquid-cooled systems, pump degradation can present identically: one unit hot while others are fine.

Environmental causes

Restore rack airflow discipline. Install blanking panels in unoccupied rack units and seal obvious recirculation paths. Hot air looping from the exhaust side back to intakes raises inlet temperature for every node in the rack and can start cascades on nodes that were previously fine.

Fix the cooling plant issue. If inlet temperature is above spec across rows, escalate to facilities. No GPU-level action compensates for hot intake air.

Buying time while you fix the root cause

Migrate or shed workload on the affected node. If the platform supports it, move jobs off the node while it is serviced. This is the cleanest stopgap.

Reduce the power limit on the originating GPU as a temporary measure. Capping power reduces heat output at the cost of performance on that unit, which can relieve the downstream neighbours while you schedule repair. Treat it strictly as a bridge: a permanently power-capped GPU masking a paste problem will resurface later, and power capping is easy to forget. If you do this, track it as configuration drift.

Do not just restart the workload or reboot the node. The cascade is a physical equilibrium; it will re-form as soon as load returns.

Prevention

  • Monitor the temperature spread, not just the maximum. The spread between the hottest and coolest GPU in a node is the earliest cascade indicator. Alert on the spread and on throttle reasons, not on absolute temperature thresholds.
  • Trend the thermal baseline per GPU over months. A slowly rising baseline under constant workload indicates paste degradation or dust accumulation while there is still time to schedule maintenance calmly. Rule of thumb: keep peak workload temperature at least 10C below the throttle threshold.
  • Watch inlet temperature and fan duty per node. Rising inlet temps or fans working harder for the same workload are rack-level leading indicators.
  • Sample fast enough. Thermal excursions and throttle events can start and resolve in seconds. Minute-resolution sampling will miss the onset of a cascade; collect temperature, power, clocks, and throttle reasons at 10 seconds or faster.
  • After any service that involves reseating cards or heatsinks, re-baseline. Compare the repaired unit’s temperature against its peers under load before returning the node to the production pool.
  • In distributed training, treat asymmetric thermals as a straggler risk. One throttled GPU in a data-parallel job gates every collective. Cross-GPU comparison within the job catches the cascade from the workload side too.

How Netdata helps

Netdata shortens this diagnosis mostly by making the per-GPU comparison and the time correlation trivial:

  • Per-second per-GPU temperature, power draw, and SM clocks on one screen, so the one-hot-GPU outlier and its downstream victims are visible at a glance rather than after a dozen nvidia-smi -i N queries.
  • Clock throttle reasons collected alongside temperature and clocks, so you can confirm a hot GPU is actually throttling (sw_thermal_slowdown / hw_thermal_slowdown) rather than just warm.
  • The throttled-but-still-drawing-power amplifier is directly visible: power draw staying near the enforced limit while SM clocks collapse.
  • Long retention at high resolution, which is what makes the months-long thermal baseline drift (paste degradation, dust) visible before the first throttle event.
  • Correlation with host-level metrics and hardware sensor data, so chassis inlet temperature and fan behaviour sit next to GPU temperatures when you separate local faults from environmental ones.
  • In multi-GPU training, per-GPU side-by-side views expose the asymmetric pattern that turns one throttling GPU into a job-wide straggler.