The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nvidia-gpu / nvidia-gpu-retired-pages ▌

Operations Guides

NVIDIA GPU retired pages: page retirement, pending retirements, and end of life

Retired pages are the permanent record of a GPU’s memory hardware degrading. Every retirement means the driver found a framebuffer page it could no longer trust and removed it from the allocatable pool for good. The counts persist in the GPU’s InfoROM across reboots and driver reloads, which makes them one of the few genuinely cumulative hardware health signals nvidia-smi exposes.

Operators usually hit this signal one of two ways: monitoring flags retired_pages.pending = Yes and nobody knows whether to panic, or someone asks why a GPU with 40 GB of HBM shows slightly less memory than its siblings. Both require understanding the difference between SBE-retired and DBE-retired pages, what a pending retirement means for running workloads, and where the end-of-life line sits.

For the broader ECC picture that feeds into retirement, see NVIDIA GPU ECC errors: corrected, uncorrected, volatile, and aggregate.

What page retirement is and why it matters

Dynamic page retirement is the NVIDIA driver’s mechanism for quarantining framebuffer pages that have proven unreliable. When the ECC subsystem determines a page is bad, the driver records the page’s address and marks it so no future allocation can touch it. The GPU keeps running, workloads keep allocating, and the lost capacity per page is small (on the order of 64 KB per page). The mechanism exists so that memory degradation is survivable.

It matters for three reasons:

  1. It is one-directional. A retired page never comes back. The counts only go up, they survive reboots, and they are stored in the GPU’s InfoROM, not on the host. Retired page counts are the closest thing a GPU has to a wear gauge.
  2. The retirement table has a hard limit. The driver can track a maximum of 64 retired pages. Once the table is full, no further pages can be retired, and a subsequent uncorrectable error on a new address has no mitigation left. A GPU at or near 64 retired pages is at the end of its useful life.
  3. Retirement is the evidence trail for replacement decisions. Fleet operators, and NVIDIA’s own RMA process, use retired page counts and causes to decide whether a board gets swapped.

Retired pages are not an acute fault. The GPU continues operating with reduced capacity, and ECC continues to protect active workloads. The job here is to read the signal correctly and make a good replacement decision, not to respond to an incident.

How page retirement works

There are two retirement causes, and they mean very different things about the hardware.

SBE retirement is preventive. When a page accumulates repeated single-bit (correctable) ECC errors at the same address, the driver concludes the page is marginal and retires it before it produces something worse. Background soft errors are normal at low rates, but repeated errors at one address indicate a physically weakening cell.

DBE retirement is reactive. A double-bit (uncorrectable) ECC error means data corruption has already occurred. The page that produced it is retired immediately, and the driver logs XID 48 to the kernel log. Any DBE-retired page is a serious event: whatever was in that page may have been corrupt, and any workload running at the time should have its outputs validated.

The full lifecycle:

stateDiagram-v2
  [*] --> Healthy: page in allocatable pool
  Healthy --> Suspect: repeated SBEs at same address
  Healthy --> BadPage: DBE (XID 48)
  Suspect --> PendingRetire: retirement recorded in InfoROM
  BadPage --> PendingRetire: retirement recorded in InfoROM
  PendingRetire --> Retired: driver reload / GPU reset / reboot (XID 63)
  Retired --> [*]: permanently blacklisted
  PendingRetire --> RetirementFailed: XID 64 (reset + escalate)

The intermediate state in that diagram is where most operational confusion lives.

Reading the counters

Retired page data is not available through --query-gpu. It has its own query path:

# Counts and pending-blacklist status
nvidia-smi -q -d PAGE_RETIREMENT

# Per-page CSV dump with addresses and cause per page
nvidia-smi --query-retired-pages=gpu_uuid,retired_pages.address,retired_pages.cause --format=csv

The -d PAGE_RETIREMENT output reports the SBE retirement count, the DBE retirement count, and a pending flag (Yes or No), plus per-page addresses and causes in the detailed view. Current --query-retired-pages fields expose the per-page address, timestamp, and cause (Single Bit ECC or Double Bit ECC); derive totals and pending state from the detailed display or DCGM counters. Older documentation refers to the pending field as “Pending Page Blacklist”; newer nvidia-smi output labels it “Pending”. Both mean the same thing.

If you collect through DCGM instead, the equivalent field IDs are:

DCGM fieldMeaning
390Total pages retired due to SBE
391Total pages retired due to DBE
392Pending retirement status

Two interpretation rules save you from common mistakes:

  • Alert on deltas, not on raw counts. These counters are monotonic and persistent. A GPU that retired two pages eight months ago and has been stable since is not an emergency; alerting on “count > 0” will page you forever on hardware that is fine. Track increases between samples and alert on new retirements.
  • Volatile versus aggregate does not apply here. Unlike ECC error counters, retired page counts are already lifetime-persistent. There is no reset-on-driver-load behavior to account for. The count you read is the count since the board shipped.

Pending retirements: the dangerous state

A pending retirement means the GPU knows a page is bad, has written the address to InfoROM, but is still capable of allocating that page. The retirement does not take effect until the driver is reloaded. Until then, the bad page remains in the allocatable pool and can produce further errors, including another DBE, while workloads are running on it.

NVIDIA’s own documentation is blunt about this: pages that are retired but not yet blacklisted can still be allocated and may cause further reliability issues.

Applying a pending retirement requires one of:

# Option 1: GPU reset (disruptive - kills all CUDA contexts on the GPU)
nvidia-smi -i <gpu> -r

# Option 2: reboot the node
# If XID 63 is missing after XID 48, also exit persistence mode before
# reset/reboot to force the documented InfoROM race workaround:
nvidia-smi -i <gpu> -pm 0

These actions are disruptive to anything running on the GPU. nvidia-smi -r will fail while processes hold contexts, so the practical sequence is: drain the GPU or the node, confirm no compute apps remain (nvidia-smi --query-compute-apps=pid --format=csv,noheader), then reset, then verify. Remember to re-enable persistence mode afterward if you disabled it (nvidia-smi -i <gpu> -pm 1).

One more gotcha: there is a documented race in persistence mode between logging errors to InfoROM and ending a CUDA job. The symptom is seeing XID 48 in dmesg but never seeing the follow-up XID 63 (page retired). If you hit that pattern, exiting persistence mode forces the pending writes to flush. See NVIDIA persistence mode: why the GPU keeps re-initializing and P-state flaps for the broader persistence tradeoffs, and NVIDIA’s dynamic page retirement documentation for the mechanism.

After applying the retirement, verify with the query from the previous section: the pending-blacklist status should return to No and the SBE or DBE count should have incremented by the number of pages that were pending.

End of life: when retired pages mean replacement

The retirement table holds a maximum of 64 pages. That number drives the entire end-of-life model:

StateInterpretationAction
Any new DBE-retired pageUncorrectable error occurred; corruption possiblePut the GPU on the replacement queue regardless of total count; validate recent workload outputs
Total approaching ~48 pages (~75% of limit)Majority of retirement budget consumedSchedule replacement at the next maintenance window
Total >= 60 pagesNearly at the 64-page hard limitUrgent replacement; NVIDIA’s RMA eligibility guidance starts here
Total = 64 pagesRetirement table fullNo further mitigation exists; any new bad page is unrecoverable. Remove from production

Two nuances on this table.

Any DBE retirement puts the GPU on the replacement queue, even at a low total count. A single uncorrectable error is a qualitatively different event from a pile of preventive SBE retirements: the hardware has already produced corruption once. The SBE count can grow to a dozen pages on a GPU that serves reliably for years; a single DBE-retired page is a replacement conversation.

Retirement rate matters as much as the absolute count. A GPU that retired 10 pages in its first week and none since is a different risk from one that retires a page a month on a steady slope. Extrapolate the slope against the 64-page limit to estimate runway, the same way you would for any consumable resource. Accelerating SBE retirement is the classic leading edge of the silent memory degradation pattern.

Ampere and later: row remapping changes the picture

On A100 and later datacenter GPUs, row remapping supplements page retirement as the memory error mitigation mechanism. Instead of blacklisting a whole page, the hardware remaps the faulty DRAM row to a spare row, which is finer-grained. Page retirement still exists, but on these GPUs you should also watch the remapper:

# Row remapping status (Ampere+ only; older GPUs return N/A)
nvidia-smi --query-remapped-rows=remapped_rows.correctable,remapped_rows.uncorrectable,remapped_rows.pending,remapped_rows.failure --format=csv,noheader

# Human-readable form - note the flag is ROW_REMAPPER, not REMAPPED_ROWS
nvidia-smi -q -d ROW_REMAPPER

The field to fear here is remapped_rows.failure = true, which means a row-remapping failure has occurred and the GPU can no longer self-heal the affected region. Like the pending flag, remapped_rows.pending > 0 requires a reset or reboot to apply. The same delta-tracking discipline applies: failure is latched state in InfoROM, so alert on the transition to true, not on the boolean, or you will re-alert on the same GPU forever. For row-remapped GPUs, NVIDIA’s DRAM RMA policy is different: the failure flag, validated by field diagnostics, is the DRAM RMA criterion. It can be set by a ninth uncorrectable-row attempt on a bank with eight already remapped, a retry of a row already remapped, or the 512-total uncorrectable-remap limit; Blackwell adds HBM channel repair when a spare channel is available. The 64-page retirement thresholds apply to the page-retirement table, not row remapping.

Signals to watch in production

SignalWhy it mattersWarning sign
SBE retired-page total (delta)Preventive retirements; slope predicts runwayAny increase; acceleration over weeks
DBE retired-page total (delta)Reactive retirements after corruptionAny increase; replacement queue regardless of total
PAGE_RETIREMENT pending status / DCGM field 392Bad page still allocatable until reattachmentYes / 1; schedule a drain and reset promptly
Total retired pages vs 64Remaining mitigation budget~48 (schedule replacement), >= 60 (urgent)
remapped_rows.pending / failure (Ampere+)Pending remap needs reset; failure means a repair limit/rule was hitpending > 0, transition of failure to true
XID 48 / 63 / 64 in dmesgThe event trail behind the counters48 without a following 63 (persistence-mode race); any 64
Corrected ECC error rateFeeds future SBE retirementsRate acceleration, per the ECC guide

For how these fit into a full monitoring stack, see NVIDIA GPU monitoring checklist: the signals every production GPU fleet needs.

How Netdata helps

  • Netdata’s NVIDIA GPU collector tracks retired page counts and pending status per GPU over time, which is exactly the delta-tracking view this signal needs: trend lines on monotonic counters instead of raw values.
  • Because retirement is the downstream product of ECC errors, having SBE rate, DBE events, and retired page counts on the same dashboard lets you watch a degradation story develop: corrected error rate accelerates, first SBE retirement lands, pending flag sets, retirement applies after reset.
  • Per-second sampling catches the pending flag as soon as it sets, so the window where a known-bad page is still allocatable is measured in the time it takes you to schedule a drain, not the time it takes someone to notice.
  • Long retention on these slow-moving counters is what makes runway extrapolation possible: a retirement every six weeks is invisible on a 24-hour graph and obvious on a 6-month one.
  • Correlating XID events from system logs with counter transitions closes the loop on the persistence-mode race (XID 48 without XID 63) that pure counter polling would miss.