The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nvidia-gpu / nvidia-gpu-bar1-memory-exhaustion ▌

Operations Guides

NVIDIA BAR1 memory exhaustion: mapping failures with free framebuffer

Your job dies with a CUDA allocation or mapping error, but nvidia-smi shows gigabytes of framebuffer free. You check for fragmentation, restart the job, and it fails again at the same place. The resource that ran out is not VRAM. It is the BAR1 aperture, the PCIe-mapped window the CPU uses to reach GPU memory directly.

BAR1 exhaustion is a distinct failure mode from framebuffer OOM. A GPU can have most of its framebuffer free and still refuse new mappings because every process, IPC handle, and GPUDirect RDMA registration consumes space in a small, shared aperture. It is most common in multi-process, MPS, containerized, and multi-tenant environments where many processes map GPU memory at once.

This article covers how to confirm BAR1 is the bottleneck, what consumes it, and how to recover without blindly rebooting the node.

What this means

BAR1 (Base Address Register 1) is a PCIe region that maps GPU memory into the host CPU’s address space. The driver uses it for CPU-visible allocations, CUDA IPC between processes, GPUDirect RDMA registrations, and various internal mappings. Every concurrent consumer takes a slice.

The aperture is finite and often surprisingly small. On older and consumer GPUs without Resizable BAR, BAR1 can be as small as 256 MiB regardless of how much VRAM the card has. On datacenter GPUs with Resizable BAR enabled (which requires both BIOS and OS support), BAR1 is typically sized to match the framebuffer — for example, an A100 80 GB reports 32768 MiB of BAR1 when Resizable BAR is enabled — which makes exhaustion unlikely.

When BAR1 fills up, new mapping requests fail. Depending on the call path, that surfaces as a CUDA allocation error, a CUDA IPC failure, or a GPUDirect RDMA registration failure that forces a fallback to slower bounce buffers. The framebuffer itself is fine. The failure is in the mapping layer.

flowchart TD
  A["CUDA allocation or IPC fails"] --> B{"Framebuffer free?"}
  B -->|"No"| C["Real VRAM exhaustion
see CUDA OOM guide"] B -->|"Yes, GB free"| D["Check BAR1 used vs total"] D -->|"BAR1 > 80%"| E["BAR1 exhaustion:
too many mappings,
leaked mappings, or tiny BAR1"] D -->|"BAR1 low"| F["Different cause:
FB fragmentation,
reserved overhead, app bug"]

Common causes

CauseWhat it looks likeFirst thing to check
Too many concurrent processes mapping one GPUBAR1 grows with process count; failures start when a new job lands on an already shared GPUCount compute apps per GPU and compare against BAR1 usage
Leaked or unreleased mappingsBAR1 stays high after jobs exit; climbs monotonically over daysCompare BAR1 before and after killing all GPU processes
Small BAR1 aperture (no Resizable BAR)BAR1 total is 256 MiB on a card with tens of GB of VRAMnvidia-smi -q -d MEMORY BAR1 total
CUDA IPC heavy workloadsFailures correlate with inter-process tensor sharing (for example, inference servers passing tensors between workers)Correlate failures with IPC usage in application logs
GPUDirect RDMA registrationsBAR1 high on nodes doing InfiniBand/RoCE direct transfers; throughput degrades before outright failureBAR1 usage on RDMA-enabled nodes; registration errors in fabric logs
Driver-branch mapping bugsBAR1 climbs under sustained mapping churn with no workload growth; may end in GPU lockupDriver version against known issues; dmesg for NVRM mapping errors

Quick checks

All of these are read-only and safe on a production node.

# Human-readable BAR1 usage per GPU (Total / Used / Free)
nvidia-smi -q -d MEMORY

# Machine-readable BAR1 usage
nvidia-smi --query-gpu=memory.bar1.total,memory.bar1.used,memory.bar1.free --format=csv,noheader,nounits

The exact BAR1 field names vary by driver branch and were not verified against every branch. If the query returns an error, list the names your driver supports and adjust:

# Confirm the BAR1 field names on your driver
nvidia-smi --help-query-gpu | grep -i bar1
# Live view: fb and bar1 columns (in MB), refreshing
nvidia-smi dmon -s m

# Which processes hold GPU contexts, and how much FB each uses
nvidia-smi --query-compute-apps=pid,process_name,used_gpu_memory --format=csv,noheader

# Driver messages: mapping failures and Xid events
dmesg -T | grep -i "NVRM"

Two things to note when reading output. First, nvidia-smi does not attribute BAR1 per process, so attribution is indirect: count processes and watch BAR1 as they come and go. Second, in containers the PID shown is the host PID, not the container PID, so you may need namespace translation to find the real owner.

In dmesg, some driver branches log explicit mapping failures such as NVRM: dmaAllocMapping... can't alloc VA space for mapping when BAR1 virtual address space runs out. The exact message text is driver-version dependent; the absence of a message does not rule BAR1 out.

How to diagnose it

  1. Confirm the symptom shape. Collect the application error and nvidia-smi output at failure time. If the error is an allocation or mapping failure while memory.free shows gigabytes available, BAR1 or fragmentation are the candidates. If free memory is genuinely near zero, you have a normal framebuffer OOM; see CUDA out of memory: diagnosing NVIDIA GPU framebuffer exhaustion.

  2. Read BAR1 directly. Run nvidia-smi -q -d MEMORY on the affected GPU and compute used/total. Above 80% in a multi-process environment is the investigate line; above 90% in a GPUDirect RDMA environment is a ticket-level condition.

  3. Check the aperture size. If BAR1 total is 256 MiB on a card with far more VRAM, Resizable BAR is off or unsupported on this host. That alone makes exhaustion plausible under modest process counts.

  4. Correlate with process count. List compute apps. If BAR1 tracks the number of concurrent processes, you have a capacity problem: too many mappers for the aperture. If BAR1 is high with few or no processes, you have leaked mappings.

  5. Test the leak hypothesis. Drain the GPU of workloads (or pick a quiet window) and watch BAR1. If it does not fall when processes exit, mappings are being held by stale contexts or a driver-level leak. Check dmesg for NVRM errors and note the driver version.

  6. Correlate with Xid events. Application-side Xids (13, 31, 43) point at application bugs; mapping failures from aperture exhaustion are a driver-resource problem. Keeping the two apart changes who fixes it. See NVIDIA Xid 31: GPU memory page fault (invalid address) if Xid 31 appears.

  7. Check for the wedged-GPU end state. Sustained mapping churn against a full aperture has, on some driver branches, ended in a GPU that stops responding. If nvidia-smi starts hanging or one GPU stops answering queries, treat it as a separate, more severe incident: see nvidia-smi hangs or is unresponsive.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
bar1.used / bar1.total ratioThe direct measure of aperture pressure>80% in multi-process environments; >90% in GPUDirect RDMA environments
BAR1 trend over daysLeaked mappings show up as a floor that never resetsBaseline ratcheting upward across job boundaries
Compute app count per GPUProxy for mapping concurrencyCount climbing toward the point where BAR1 historically saturates
memory.free at failure timeDistinguishes BAR1 exhaustion from real VRAM OOMAllocation errors with large free framebuffer
Application CUDA IPC / mapping errorsThe user-visible symptomRecurring mapping failures at roughly constant BAR1 levels
NVRM messages in dmesgDriver-side confirmation of mapping failurecan't alloc VA space or NV_ERR_NO_MEMORY style messages; exact wording varies by driver branch
GPUDirect RDMA throughputRegistration failure forces bounce-buffer fallbackThroughput drop on RDMA nodes with high BAR1

Fixes

Reduce concurrent mappings on the GPU

If BAR1 pressure tracks process count, cap how many processes share the device. Co-schedule fewer tenants per GPU, or stagger job starts so mappings do not peak simultaneously. On multi-tenant nodes this is a scheduling policy fix, not a GPU fix.

If you use CUDA IPC heavily (for example, worker processes sharing tensors), reduce the number of IPC peers per GPU or restructure so tensors move through fewer mapped handles.

Clear leaked mappings

Find and stop the processes holding stale contexts. nvidia-smi --query-compute-apps shows what the driver thinks is attached; zombie contexts from crashed processes may need a kill -9 on the host PID. Watch BAR1 as each process exits to identify the leaker.

A known historical variant: PyTorch DataLoader setups with num_workers > 0 leaked BAR1 when the main process was killed with SIGKILL/SIGTERM on certain driver branches (reported fixed in 515.65.01 on some affected branches; the PyTorch issue tracking the problem remains open, so verify against your driver branch). If your driver is older than that and you kill training jobs routinely, upgrading the driver is the durable fix.

If processes are gone and BAR1 stays allocated, the mappings are held at the driver level. A GPU reset releases them:

# Destructive: kills all work on the GPU and fails if any process is attached
nvidia-smi --gpu-reset

Reset only after draining the device. If the reset fails or the GPU does not come back clean, the remaining option is a node reboot. See Resetting a wedged NVIDIA GPU for the full decision tree.

Enlarge the aperture: Resizable BAR

If BAR1 is 256 MiB, enabling Resizable BAR lets the aperture grow to match the framebuffer on supported GPUs. This requires both BIOS support (an above-4G decoding / Resizable BAR option) and OS/driver support; on the open kernel modules, the NVreg_EnableResizableBar=1 module parameter exists since driver 530.41.03 (the option is checked in the open kernel modules’ nv-pci.c source).

Caveats before you schedule the reboot:

  • This is a BIOS change plus a reboot, so it needs a maintenance window.
  • On multi-GPU hosts, PCIe address space is finite. There are reports of systems where most GPUs get a large BAR1 but one GPU gets a smaller one because the host ran out of address space. Verify BAR1 size on every GPU after the change, not just the first.
  • Consumer platforms vary widely in whether the option exists and works.

Upgrade the driver for mapping-churn bugs

Recent open-kernel-module branches have had bugs where sustained mapping reuse churn exhausts BAR1 virtual address space and locks the GPU, even at low framebuffer usage. If your BAR1 climbs under a stable workload and dmesg shows NVRM mapping errors, check the open-gpu-kernel-modules issue tracker for your branch and plan a driver upgrade. Treat a GPU that has locked up once from this as suspect until the driver is changed.

Prevention

  • Alert on the ratio, not the absolute. Track bar1.used / bar1.total per GPU. Investigate at 80% in multi-process environments; ticket at 90% where GPUDirect RDMA is in play.
  • Baseline per workload class. A single-process training job should sit near zero BAR1. Multi-tenant inference nodes will have a higher normal. Alert on deviation from the workload’s own baseline.
  • Enforce per-GPU tenancy limits. The number of processes that can safely map one GPU is a property of the aperture size, not the framebuffer size. Encode it in the scheduler.
  • Track BAR1 across job boundaries. A floor that rises week over week is a leak; catch it before the cliff.
  • Verify BAR1 size at provisioning. Include BAR1 total in node acceptance checks, especially after BIOS changes or hardware swaps. A node that silently comes up with a 256 MiB aperture will fail later under multi-process load.
  • For the broader GPU signal picture, see NVIDIA GPU monitoring checklist: the signals every production GPU fleet needs and the GPU subsystem model in How an NVIDIA GPU actually works in production.

How Netdata helps

  • Netdata collects per-GPU memory metrics at per-second resolution, so BAR1 pressure is visible as a trend rather than a postmortem surprise; minute-resolution polling misses the ramp before failures.
  • Charting BAR1 used next to framebuffer used makes the signature obvious: flat or falling FB with climbing BAR1 means a mapping problem, not an OOM problem.
  • Per-GPU process counts alongside BAR1 usage let you see whether pressure tracks tenancy (capacity problem) or grows independently (leak).
  • Anomaly detection flags a BAR1 baseline that is ratcheting upward across job boundaries, the early leak signal that static thresholds miss.
  • Correlating BAR1 with application error timing and system log events in one view shortens the path from “CUDA error” to “aperture exhausted.”