The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / activemq / activemq-systemusage-memory-vs-heap ▌

Operations Guides

ActiveMQ memoryUsage vs JVM heap: the two memory budgets teams confuse

An ActiveMQ Classic broker can hit 100% MemoryPercentUsage and silently block every producer while the JVM heap sits at 40%. The same broker, on a different day, can die from OutOfMemoryError while MemoryPercentUsage shows 50% and every queue-depth dashboard looks calm. Both incidents look contradictory until you understand that ActiveMQ tracks two separate memory budgets, and only one of them is the JVM’s.

Teams monitor one budget, assume it represents the other, and get surprised by whichever one they ignored. This article covers what each budget counts, how they drift apart, how to size them relative to each other, and which signals detect the drift.

Scope: ActiveMQ Classic 5.x/6.x. Artemis has a fundamentally different memory model (paging and global-max-size rather than usage-manager flow control) and nothing below applies to it.

What each budget actually is

JVM heap is the real memory. It is the Java heap bounded by -Xmx, holding everything the broker’s JVM allocates: message bodies sitting in cursors, destination metadata, connection and session state, JMX MBean trees, pending acknowledgment records, KahaDB index structures, and every non-ActiveMQ object the runtime creates. When this budget is exhausted, the JVM throws OutOfMemoryError or dies in a GC thrash spiral. Nothing in ActiveMQ configuration can save you from it.

ActiveMQ memoryUsage is internal accounting on top of the heap. It is a counter, configured in activemq.xml inside the <systemUsage> block, that tracks the bytes charged for pending (undispatched and unacknowledged) messages across all destinations. It is not a reservation and it is not a measurement of heap. It is the broker’s own bookkeeping of “how much message data am I holding,” maintained separately from what the JVM has actually allocated.

The distinction matters because the two budgets enforce completely different failure behaviors:

  • memoryUsage hitting 100% triggers producer flow control. The broker stops reading from producer sockets. Producer send() calls block silently, with no exception and no log entry on the producer side by default. The broker stays alive. Upstream services hang.
  • JVM heap hitting its ceiling triggers GC collapse and OOM. Long GC pauses freeze all threads, clients hit their wireFormat.maxInactivityDuration heartbeat timeout (default 30000ms) and disconnect, reconnection storms pile more objects onto the heap, and the cycle accelerates until the process dies or is OOM-killed. The broker does not stay alive.
flowchart TD
  HEAP["JVM heap (-Xmx)
real memory"] MU["memoryUsage counter
message bytes only
limit: percentOfJvmHeap or absolute"] OTHER["Non-message heap
MBeans, connection state,
destination metadata, index, sessions"] FC["memoryUsage = 100%
producer flow control
send() blocks silently"] OOM["heap exhausted
GC spiral, OOM kill,
broker down"] HEAP --> MU HEAP --> OTHER MU -->|"limit reached"| FC OTHER -->|"grows unchecked"| OOM MU -.->|"does not count"| OTHER

memoryUsage is a subset of heap in spirit but not in accounting: the bytes it counts live in the heap, but the heap also holds everything the counter ignores. Either side can run out first.

How the two budgets drift apart

The drift is not a bug. It falls out of what each budget counts. There are two failure directions, and both are common in production.

Direction 1: heap exhausted while memoryUsage shows headroom

The memoryUsage counter only charges message bytes. It does not count:

  • JMX MBeans and their metadata. Every destination creates several MBeans. A broker with thousands of dynamically created destinations carries tens of thousands of MBeans, all in heap, none charged to memoryUsage. This is the destination-explosion failure pattern: heap and GC pressure climb while message memory stays low.
  • Connection and session state. Thousands of connections, each with sessions, consumers, producers, and prefetch bookkeeping, consume heap outside the counter.
  • Advisory destinations. Each advisory topic is a real destination with real MBeans and state, created silently on top of your application destinations.
  • KahaDB index structures and page cache effects in the JVM.
  • Redelivery state, transaction state, scheduler state, network bridge bookkeeping in Network of Brokers deployments.

This is the silently catastrophic pattern: low MemoryPercentUsage with high heap usage means the two budgets are misaligned and OOM is possible. Your message-flow dashboards look green while the JVM is dying of metadata bloat. It is also why a broker can enter the GC pause death spiral (long pauses, heartbeat timeouts, reconnection storms) with MemoryPercentUsage nowhere near 100%.

Direction 2: memoryUsage full while the heap has room

This is the reverse misconfiguration, and it is self-inflicted. If memoryUsage is set far below what the heap could carry, or if per-destination memory limits on <policyEntry> elements are set too small, the broker flow-controls producers while gigabytes of heap sit idle. Symptoms:

  • Producer send() calls hang with no exception. Upstream services stall and their owners blame their own code, connection pools, or the network.
  • Broker-level MemoryPercentUsage is at 100%, or one destination’s MemoryPercentUsage is at 100% while the broker-level counter is fine.
  • JVM heap usage is moderate. GC is calm. Nothing looks “broken” except that throughput has collapsed.

Because producer flow control is silent by default, the first team to notice is usually the one whose HTTP API started timing out, three services upstream of the actual producer. The related guide on producer flow control covers the diagnosis of the blocked-send side in detail.

Sizing: why memoryUsage must be 60-70% of heap, never 100%

The rule is absolute: never set memoryUsage equal to max heap. The broker needs heap for everything the counter does not track. If the message counter is allowed to consume the entire heap, the first burst of messages leaves zero room for MBeans, connection state, and dispatch bookkeeping, and the broker OOMs at exactly the moment it is under the most load.

The shipped default in current 5.x configurations is:

<systemUsage>
  <systemUsage>
    <memoryUsage>
      <memoryUsage percentOfJvmHeap="70"/>
    </memoryUsage>
    ...
  </systemUsage>
</systemUsage>

Seventy percent is a rule of thumb, not a guarantee. The right value depends on how much non-message heap your workload carries:

  • Many destinations, many connections, NoB, heavy advisory traffic: non-message heap is large. Keep memoryUsage at 50-60% of heap and treat the default 70% as too generous.
  • Few destinations, few connections, large message bodies: message bytes dominate. 70% is reasonable.
  • Very large heaps: the absolute headroom (heap minus memoryUsage limit) is what matters. On a 32GB heap, 70% leaves ~9.6GB for non-message state, comfortable for almost any deployment. On a 2GB heap, 70% leaves ~600MB, which a destination explosion can eat in an afternoon.

You can also set an absolute limit (limit="2 gb") instead of percentOfJvmHeap. Absolute limits are safer when the same activemq.xml is deployed to hosts or containers with different -Xmx values, because a percentage silently rescales with heap and can become wrong without anyone editing the file. If you run the broker in a container, remember that the container memory limit and the JVM heap interact independently of all of this: if -Xmx approaches the container limit, the OOM killer terminates the process before the JVM throws OutOfMemoryError, and neither budget’s metrics will tell you why.

Two adjacent limits deserve the same skepticism. storeUsage is a configured limit compared against KahaDB consumption, not against physical disk; if the store limit exceeds real disk capacity, the filesystem fills first. tempUsage caps non-persistent overflow spill. All three systemUsage values are accounting limits layered on real resources, and each can be misaligned with the resource it proxies.

Signals to watch in production

Monitor both budgets independently, plus the ratio between them:

SignalWhy it mattersWarning sign
MemoryPercentUsage (Broker MBean)The broker’s message-memory budget. At 100%, producer flow control blocks all sends silently.>80% sustained, or climbing 5%/min. 100% with active producers is a page.
Per-destination MemoryPercentUsageIsolates which destination is consuming the message budget. One runaway queue can starve all producers.100% on any destination, or one destination dominating broker-level usage.
JVM heap used after major GC (java.lang:type=Memory, HeapMemoryUsage)The real memory budget, including everything the counter ignores. Post-GC value is the honest floor.>85% after GC, or the post-GC floor trending upward over days. >95% after GC with GC distress is a page.
The gap between the twoMisalignment detector. This is the signal almost nobody plots.Low MemoryPercentUsage with high heap usage: non-message heap is growing, OOM possible while message dashboards look green. The reverse (memoryUsage pinned at 100%, heap calm) means your limit is too small or flow control is misconfigured.
GC pause duration and frequencyHeap pressure’s early symptom. Pauses over the 30s default inactivity timeout disconnect every client.Pauses >2s, full GC more than once per 5 minutes, or pause duration trending up.
Total destination countEach destination costs heap in MBeans and metadata regardless of message volume.Count growing unbounded, or thousands of destinations with zero producers and consumers.
TempPercentUsageNon-persistent overflow spilling to disk, an early sign the memory budget is undersized for the workload.Sustained non-zero values.

All of these are JMX attributes, reachable via Jolokia on the web console (default port 8161) or any JMX client. HeapMemoryUsage comes from java.lang:type=Memory; MemoryPercentUsage and MemoryLimit come from the broker MBean (org.apache.activemq:type=Broker,brokerName=<name>). Pull both on the same polling interval so the gap between them is a real comparison, not two unrelated time series.

Checking alignment by hand

A quick manual audit takes three queries. Compare what the broker thinks it is using for messages against what the JVM is actually holding:

# Broker message-memory budget
curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/org.apache.activemq:type=Broker,brokerName=localhost/MemoryPercentUsage'

curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/org.apache.activemq:type=Broker,brokerName=localhost/MemoryLimit'

# Real JVM heap
curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/java.lang:type=Memory/HeapMemoryUsage'

From the three results, derive:

  1. Effective message budget = MemoryLimit (or percentOfJvmHeap x -Xmx if configured as a percentage).
  2. Non-message heap pressure = heap used after a major GC, minus roughly the message bytes in flight. If heap-after-GC is high while MemoryPercentUsage is low, the difference is MBeans, connection state, and metadata, and it will not be freed by draining queues.
  3. Headroom check = is MemoryLimit between 60% and 70% of heap max? If it equals heap max, fix it before the next traffic burst does it for you.

If non-message heap is the growing half, check the usual suspects in order: destination count (the Queues and Topics arrays on the broker MBean), connection count (CurrentConnectionsCount), and orphaned durable subscriptions accumulating pending messages and state.

Common misuses, condensed

  • Setting memoryUsage equal to -Xmx. Leaves no heap for anything that is not a message body. The broker OOMs under load. Keep it at 60-70%.
  • Monitoring only MemoryPercentUsage. Catches flow control, misses every non-message heap failure: destination explosion, MBean bloat, connection-state growth.
  • Monitoring only JVM heap. Catches OOM risk, misses flow control entirely. Producers block at 100% memoryUsage with the heap half empty.
  • Trusting the 70% default blindly on small heaps. On a 1-2GB heap the remaining 30% is small in absolute terms and easily consumed by a few thousand destinations.
  • Percentage limits with heterogeneous -Xmx across environments. The same XML silently means different byte budgets in staging and production.
  • Assuming Artemis semantics. Artemis pages to disk and uses global-max-size; the flow-control-at-100% model is Classic-specific.

How Netdata helps

The two-budget problem is a correlation problem, and correlation is where per-second monitoring earns its place:

  • Both budgets on one timeline. Netdata collects ActiveMQ broker metrics alongside JVM heap metrics from the same host, so MemoryPercentUsage and heap-after-GC are directly comparable instead of living in two tools.
  • The gap as an anomaly. When heap usage climbs while broker memory usage stays flat, ML anomaly scoring flags the divergence before the OOM, which is exactly the misalignment pattern described above.
  • GC pauses next to connection drops. Long GC pauses and the resulting client disconnect/reconnect sawtooth appear on adjacent charts, making the heap-side failure mode recognizable in seconds.
  • Destination and connection counts next to heap. The usual causes of non-message heap growth (destination explosion, connection leaks) are visible on the same dashboard as the heap curve they produce.
  • Per-destination memory usage. When broker-level usage rises, per-destination breakdowns identify the noisy queue without JMX archaeology.