The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nvidia-gpu / nvidia-gpu-xid-13-graphics-engine-exception ▌

Operations Guides

NVIDIA Xid 13: Graphics Engine Exception

You found a line like this in the kernel log:

NVRM: Xid (PCI:0000:65:00.0): 13, pid='<unknown>', name=<unknown>, Graphics Engine Exception.

Or, more likely, you found the application-side symptom first: a training job died with a generic CUDA error, a run produced NaN loss, or an inference service started returning garbage, and only after digging into dmesg did Xid 13 show up.

Xid 13 is the most ambiguous Xid code NVIDIA emits. It means the graphics engine raised an exception while executing work. The three candidate causes are a bad CUDA kernel (out-of-bounds access, illegal instruction), a driver bug, or genuine hardware degradation. Most isolated occurrences are software. The operational problem is telling which case you are in, because the correct response ranges from “file a bug against the application” to “RMA the GPU”.

This guide is about making that call quickly and defensibly.

What this means

The graphics engine is the unit that dispatches work to the SMs. When it hits a condition it cannot execute (an instruction it does not recognize, a descriptor that fails sanity checks, a memory access the fault handler escalates), it raises an exception, the faulting CUDA context is typically destroyed, and NVRM logs Xid 13.

Two properties matter for operations:

  1. The GPU survives. Unlike Xid 48 (double-bit ECC) or Xid 79 (fallen off the bus), Xid 13 does not, by itself, indicate that the GPU is damaged. The application that owned the faulting context dies or errors out; other contexts and the device usually keep running.
  2. The log line alone cannot classify the cause. NVIDIA’s own Xid documentation treats Xid 13 as “typically an application bug, rarely hardware”. The classification work happens in the correlation, not in the message.

That second point is why Xid 13 deserves a real triage procedure rather than a reflexive “ignore it, it’s an app bug”.

Common causes

CauseWhat it looks likeFirst thing to check
Bad CUDA kernel (out-of-bounds access, illegal instruction)Xid 13 fires when one specific application or job runs; the same job reproduces itRun the app under Compute Sanitizer memcheck or cuda-gdb
Driver or firmware bugXid 13 appears after a driver upgrade, on a specific driver branch, or under a known-problematic driver variant (open vs proprietary kernel module)Check whether the driver version changed recently and whether the issue tracks that version across nodes
Genuine hardware degradationSame GPU throws Xid 13 across different applications; frequency increases over weeks; ECC counts or retired pages are also movingECC error counts, retired pages, row remapping status on that GPU
Environmental stress masquerading as a faultXid 13 correlates with thermal throttling or power events on the same GPUTemperature and clock throttle reasons around the event timestamp

The single most useful discriminator: isolated occurrences are usually software; recurring occurrences, especially the same GPU across different applications, point to hardware.

Quick checks

All of these are read-only and safe to run during an incident.

# 1. Find all Xid events, with timestamps, for the affected GPU
dmesg -T | grep -i "NVRM: Xid"

# 2. Same via journalctl if dmesg has wrapped
journalctl -k | grep -i "NVRM: Xid"

# 3. ECC error counters on the GPU that logged Xid 13 (volatile = since driver load)
nvidia-smi -i <N> --query-gpu=ecc.errors.corrected.volatile.total,ecc.errors.uncorrected.volatile.total --format=csv,noheader,nounits

# 4. Retired pages (persistent hardware degradation record)
nvidia-smi -i <N> -q -d PAGE_RETIREMENT

# 5. Row remapping status (Ampere and later)
nvidia-smi -i <N> -q -d ROW_REMAPPER

# 6. Thermal and throttle state
nvidia-smi -i <N> --query-gpu=temperature.gpu,clocks_event_reasons.active --format=csv,noheader

# 7. What was running on the GPU when it faulted
nvidia-smi -i <N> --query-compute-apps=pid,process_name,used_gpu_memory --format=csv,noheader

# 8. Driver version, to correlate with recent changes
nvidia-smi --query-gpu=driver_version --format=csv,noheader

Two things to note while collecting:

  • Count Xid 13 occurrences per GPU per time window. The playbook heuristic is that more than roughly one per hour, or recurrence across different applications on the same GPU, moves you into hardware-suspect territory. This is an operational heuristic, not an NVIDIA-published threshold; NVIDIA’s Xid catalog classifies Xid 13 as “typically an application bug” and does not prescribe a rate threshold.
  • Check whether the faulting application is always the same binary. If three unrelated frameworks all trigger Xid 13 on GPU 4 and none of them trigger it on GPUs 0-3, the application-bug hypothesis is effectively dead.

How to diagnose it

The decision you are making is: software, driver, or hardware. Work it in this order because the cheap, high-signal checks come first.

  1. Establish the recurrence pattern. Grep the kernel log history for every Xid 13 on this node. Group by GPU PCI address and by time. One event six months ago is noise. A rising cadence on one GPU is a degradation signature.

  2. Check whether the fault follows the application or the GPU. If your scheduler can place the same job on a different GPU, do it. If the same job triggers Xid 13 on any GPU it lands on, you are looking at an application bug. If any job triggers Xid 13 only on this specific GPU, suspect the GPU.

  3. Correlate with hardware health signals on the faulting GPU. Pull ECC counters, retired pages, row remapping status, temperature, and throttle reasons around the event timestamps. Xid 13 plus accelerating single-bit ECC errors, new retired pages, or pending row remaps on the same GPU is a hardware story. Xid 13 with clean ECC history and no thermal events is a software story.

  4. Rule out the driver. Did the driver version change shortly before the first occurrence? Are other nodes on the same driver branch and GPU model reporting Xid 13? There are known cases where a specific driver variant (for example, the open kernel module on certain consumer/workstation cards) produces Xid 13/31 under specific workloads while the proprietary branch does not. A driver rollback or upgrade that makes the errors stop is diagnostic in itself.

  5. If it reproduces with one application, debug the application. Run the workload under Compute Sanitizer (compute-sanitizer --tool memcheck) or attach cuda-gdb. The older cuda-memcheck was deprecated in CUDA Toolkit 11.5 and removed in later toolkit releases; Compute Sanitizer (included since CUDA 11.0) is the current tool. Out-of-bounds accesses and illegal instructions show up directly, and this closes the case fastest for the software branch. NVIDIA’s Xid catalog for Xid 13 specifically recommends: “Run the application in cuda-gdb or the Compute Sanitizer memcheck tool, or run the application with CUDA_DEVICE_WAITS_ON_EXCEPTION=1 and then attach later with cuda-gdb.”

  6. If hardware is suspected, run diagnostics. dcgmi diag -r 3 (or -r 4 where supported) exercises the GPU under load and surfaces faults that intermittent Xid 13s hint at. Run it on a drained node; the heavy diagnostic is disruptive to anything sharing the GPU. A clean diagnostic result plus recurring Xid 13 in production is still evidence. Intermittent faults often pass short diagnostics.

flowchart TD
  A[Xid 13 in kernel log] --> B{Recurring on this GPU?}
  B -->|No, isolated| C[Likely application bug - track and move on]
  B -->|Yes| D{Same application each time?}
  D -->|Yes| E[Run app under Compute Sanitizer or cuda-gdb]
  D -->|No, different apps| F{ECC, retired pages, remaps, thermal anomalies on this GPU?}
  F -->|Yes| G[Hardware degradation - drain, diagnose, plan RMA]
  F -->|No| H{Driver changed recently or known bad branch?}
  H -->|Yes| I[Driver rollback or upgrade, then observe]
  H -->|No| J[dcgmi diag -r 3 on drained node, monitor recurrence]

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Xid 13 event frequency per GPUThe primary recurrence signal; cadence separates noise from degradationMore than ~1/hour, or an accelerating trend
Xid 13 across applicationsSame GPU faulting under different workloads kills the app-bug hypothesisMultiple unrelated jobs triggering it on one GPU
ECC corrected error rate (volatile)Accelerating single-bit errors indicate degrading memory that can also surface as engine exceptionsRate doubling over days, or above peer GPUs of the same model
ECC uncorrected errorsData corruption has occurred; recent outputs may be invalidAny new event (delta in volatile counter)
Retired pages / row remappingPersistent record of memory hardware degradationAny increase, pending retirements, or a row remapping failure flag
GPU die temperature and throttle reasonsThermal stress can contribute to intermittent faultsThrottle reasons active around Xid 13 timestamps
Application CUDA errors in job logsThe user-visible manifestation of Xid 13Generic CUDA error, unspecified launch failure, NaN loss correlated with Xid timestamps

The correlation is the classification. None of these signals alone answers “software or hardware”; together they usually do.

Fixes

Application bug confirmed (Compute Sanitizer or cuda-gdb finds the fault)

Fix the kernel. Common root causes are out-of-bounds memory access, illegal instruction from a miscompiled or wrong-arch binary, and misuse of CUDA APIs. This is the most common outcome and the cheapest. After the fix, watch the node for a few days to confirm Xid 13 does not recur.

Driver-suspect

If the errors track a driver version or variant:

  • Upgrade or roll back the driver on an affected node and observe. This is disruptive to running workloads, so drain first.
  • If you are on the open kernel module and hitting reproducible Xid 13/31 faults that the proprietary branch does not produce, switching branches is a legitimate diagnostic step, not just a workaround.
  • File a bug with NVIDIA with the full Xid line, driver version, GPU model, and a reproducer if you have one. Inconclusive software debugging is a stated reason to escalate in NVIDIA’s Xid guidance.

Hardware-suspect

  • Drain the GPU from production scheduling. Recurring Xid 13 under varied workloads means you cannot trust the outputs, and silent corruption is a worse failure mode than a crash.
  • Run dcgmi diag -r 3 on the drained node to gather evidence for the RMA conversation.
  • Validate or discard recent work products from that GPU. If the fault window overlaps a training run, treat checkpoints from that window as suspect.
  • Replace the GPU if diagnostics fail, ECC counters are moving, or Xid 13 keeps recurring after driver changes are ruled out. A GPU that throws recurring Xid 13s is on a trajectory; it rarely gets better.

Prevention

  • Log Xid events centrally with per-GPU, per-code counts. Xid 13 is only classifiable over time and across applications. Node-local dmesg wraps and loses the history you need.
  • Track ECC counters and retired pages as trends, not snapshots. The acceleration of corrected errors is the leading indicator; the absolute count is nearly useless.
  • Keep driver rollouts staged. A driver branch that introduces Xid 13 should be caught on a canary node, not across the fleet.
  • Run Compute Sanitizer on new or changed kernels in CI before production deployment. Most Xid 13s are application bugs, and catching them pre-production keeps them out of your kernel log entirely.
  • Baseline per-GPU health so that “this GPU behaves differently from its peers” is a measurement, not a hunch.

How Netdata helps

  • Xid events in kernel logs are collected and classified per code, so a recurring Xid 13 on one GPU shows up as a trend rather than a line you grep for after the incident.
  • ECC corrected and uncorrected error counts, retired pages, and row remapping status are charted per GPU, which makes the “is this GPU degrading” half of the classification answerable in seconds.
  • Temperature, power draw, and clock throttle reasons on the same dashboard let you check whether thermal or power events line up with Xid 13 timestamps.
  • Per-second GPU metrics matter here because the useful correlation window around a fault is short; minute-resolution data smears the event.
  • Per-GPU comparison across the node and fleet supports the strongest discriminator: does this GPU behave differently from identical peers under the same workloads.