The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / activemq / activemq-per-destination-memory-limit ▌

Operations Guides

ActiveMQ per-destination memory usage: one noisy queue blocking every producer

Every producer on the broker is stuck in send(). Broker-level MemoryPercentUsage reads 100, the broker log shows the usage manager hitting its memory limit, and upstream services are timing out. But when you list queue depths, almost every queue is empty. One queue holds nearly all the pending messages.

That one queue has consumed the broker’s shared memory pool. Every pending message is charged against its destination’s memory accounting and against the broker-wide system memory limit. If either limit is reached, flow control activates. With no per-destination memory limit set, ActiveMQ Classic does not throttle just that queue’s producers: it throttles everyone’s. This is the default behavior, not a bug.

This article covers how to confirm which destination is responsible, why the blocking is silent, and how per-destination limits contain the blast radius. For the broader failure model, see How ActiveMQ Classic actually works in production.

What this means

During the memory accounting step of the message flow, each message’s footprint is charged against two budgets: the destination’s and the broker’s system memory limit (configured via <systemUsage><memoryUsage> in activemq.xml). When either budget is exhausted, the broker stops reading from the producer’s socket, creating TCP backpressure. From the producer’s perspective, send() simply blocks. There is no exception, no log entry on the producer side, and no timeout unless you configure one.

There are two levels of flow control, and they have very different blast radii:

  • Per-destination flow control. A destination with its own memoryLimit blocks only producers sending to that destination.
  • Broker-level flow control. When the shared pool is exhausted, every producer on every destination blocks, regardless of which destination caused it.

Per-destination MemoryPercentUsage is the signal that tells you which level you are dealing with and who is responsible. On a multi-tenant broker, where unrelated applications share one JVM, this is the difference between “tenant A is throttled” and “the whole platform is down.”

flowchart TD
  A[Consumer on queue A stalls] --> B[Messages pile up in queue A]
  B --> C[Queue A memory usage climbs toward 100 percent]
  C --> D[Shared broker memory pool fills]
  D --> E[Broker-level flow control engages]
  E --> F[Producers on all destinations block in send]
  C -.->|with per-destination memoryLimit| G[Flow control scoped to queue A only]
  G --> H[Other destinations keep flowing]

One mechanism worth knowing: because flow control works by the broker not reading from the producer’s socket, it operates at connection granularity. If several producer sessions share one JMS connection and any of them hits flow control, sends on that connection stall together.

Common causes

CauseWhat it looks likeFirst thing to check
Stalled or slow consumer on one queueDequeue rate collapsed on one destination; inflight pinned at prefetchConsumerCount and InFlightCount on that destination
No per-destination memoryLimit configuredOne destination’s usage roughly equals broker usage; nothing isolates it<destinationPolicy> in activemq.xml
Producer burst or replay storm on one destinationEnqueue spike on one destination; memory climbs in minutesEnqueueCount delta on that destination
Non-persistent flood with VM cursorsMemory climbs with little or no store growthDelivery mode on the producer; TempPercentUsage
Offline durable subscriber on a topicTopic memory rising; a durable subscription with growing pending count and zero consumersPendingQueueSize on subscription MBeans
Broker memory limit too small for the workloadFlow control fires at normal traffic levelsMemoryLimit attribute versus your baseline backlog

The first two causes usually appear together: the slow consumer is the trigger, the missing per-destination limit is the structural weakness that turns one team’s problem into everyone’s incident.

Quick checks

These are read-only. They assume the embedded web console is reachable on port 8161 (bound to 127.0.0.1 by default since 5.16) and use the default credentials from the stock config. Adjust brokerName and credentials for your deployment.

# 1. Is broker-level flow control active?
curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/org.apache.activemq:type=Broker,brokerName=localhost/MemoryPercentUsage'

# 2. Which queues are consuming the shared pool?
curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/org.apache.activemq:type=Broker,brokerName=localhost,destinationType=Queue,destinationName=*/MemoryPercentUsage'

# 3. Check topics as well (offline durable subscribers accumulate silently)
curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/org.apache.activemq:type=Broker,brokerName=localhost,destinationType=Topic,destinationName=*/MemoryPercentUsage'

# 4. Depth on the suspect queue
curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/org.apache.activemq:type=Broker,brokerName=localhost,destinationType=Queue,destinationName=MY.QUEUE/QueueSize'

# 5. Are consumers attached to it?
curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/org.apache.activemq:type=Broker,brokerName=localhost,destinationType=Queue,destinationName=MY.QUEUE/ConsumerCount'

# 6. Are consumers holding messages without acking?
curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/org.apache.activemq:type=Broker,brokerName=localhost,destinationType=Queue,destinationName=MY.QUEUE/InFlightCount'

# 7. Broker log: flow control and usage manager messages (path varies by install)
grep -i "memory limit" /opt/activemq/data/activemq.log | tail -20

# 8. Rule out a different saturation path
curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/org.apache.activemq:type=Broker,brokerName=localhost/StorePercentUsage'
curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/org.apache.activemq:type=Broker,brokerName=localhost/TempPercentUsage'

If check 1 returns 100 while exactly one destination in checks 2-3 is near its limit and everything else is idle, you have the noisy-neighbor pattern this article is about.

How to diagnose it

  1. Confirm broker-level flow control. Broker MemoryPercentUsage at 100 with active producers, plus Usage Manager ... reached memory limit lines in the broker log, means producers are being blocked right now. When an individual destination fills, the broker also publishes to ActiveMQ.Advisory.FULL.Queue.<name> (or the topic equivalent), which tells you which destination tripped first.
  2. Rank destinations by MemoryPercentUsage. The noisy destination stands out: near 100 while peers are idle. Note that per-destination usage can briefly exceed 100 during bursts before flow control engages. Treat a value over 100 as a burst in progress, not a broken metric.
  3. Explain the accumulation. ConsumerCount of zero means nothing is draining the queue. Consumers connected but InFlightCount equal to consumer count times prefetch (default 1000 for queues) means consumers are holding messages without acknowledging them: stuck or overloaded, not absent.
  4. Check the topic case separately. A durable subscriber that went offline without unsubscribing accumulates every published message indefinitely. Look for a subscription with growing PendingQueueSize and no active consumer.
  5. Check JVM heap independently. java.lang:type=Memory HeapMemoryUsage is not the same thing as ActiveMQ’s memory accounting. Either can fill first, and they fail differently: one triggers flow control, the other triggers GC stalls or an OOM kill.
  6. Decide the immediate action. Drain (fix or restart the stalled consumer), shed (purge the queue, destructive: messages are discarded, so confirm with the owning team first), or cap (apply a per-destination limit so it cannot recur this way).

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Per-destination MemoryPercentUsageIsolates which destination is consuming the shared poolAny destination climbing toward 100 while others are idle
Broker MemoryPercentUsageAt 100, every producer on the broker blocksAbove 80 sustained; 100 with active producers is an incident
QueueSize per destinationThe backlog behind the memory usageGrowing while consumers are connected
ConsumerCount per destinationZero consumers means nothing drains the queueBelow the expected count for that destination
InFlightCount versus prefetchConsumers holding messages without ackingInflight equal to consumer count times prefetch, sustained
Dequeue rateWhether the queue is draining at allZero or collapsed with pending messages present
Enqueue/dequeue ratioAccumulation rate; gives you runway to the cliffSustained above 1.0 on the noisy destination

The degradation curve here is a cliff, not a slope. At 99 percent everything works; at 100 producers block instantly. Alerts that fire at 100 arrive after the incident has already started.

Fixes

Set per-destination memory limits (the structural fix)

Give destinations their own budget with the memoryLimit attribute on a <policyEntry>:

<destinationPolicy>
  <policyMap>
    <policyEntries>
      <policyEntry queue=">" producerFlowControl="true" memoryLimit="100mb"/>
    </policyEntries>
  </policyMap>
</destinationPolicy>

With memoryLimit set, a destination that fills its own budget flow-controls only its own producers. The noisy queue is contained; every other destination keeps flowing. When a destination has its own limit, its cursor high-water mark (cursorMemoryHighWaterMark) is evaluated against that per-destination limit rather than the broker-wide pool, so capped destinations also start paging toward the store earlier. See the per-destination policies documentation for the full attribute list. Add stricter entries for known-hot queues and looser wildcard entries for the rest.

Sizing tradeoff: too small and you throttle a healthy destination during normal peaks; too large and the cap never protects the shared pool. Size from each destination’s observed baseline backlog plus headroom for its worst legitimate burst. On multi-tenant brokers, every destination should have a cap; an uncapped wildcard is the hole the next incident walks through.

Unblock the stalled consumer (the root cause fix)

Per-destination limits contain the blast radius, but the stalled consumer is why memory filled at all. Find the consumer whose inflight count is pinned at its prefetch and restart or rebalance it. Disconnecting a slow consumer forces its inflight messages to be redelivered to the remaining consumers, which often restores drain immediately. That is disruptive (messages redeliver, ordering within the affected stream can change), so prefer it over a broker restart, not over a clean consumer fix. Once the queue drains, memory accounting releases and flow control lifts on its own.

Make blocking visible to producers

Silent blocking is what turns a broker problem into a mystery across five upstream services. Configure sendFailIfNoSpaceAfterTimeout so a producer that would block instead receives a ResourceAllocationException after the timeout, which it can log, alert on, and retry. It can be set on the connection factory and, since 5.16.0, per destination via the policy entry. Tradeoff: your producers must actually handle the exception; an unhandled exception in a fire-and-forget send path is its own kind of silent failure.

Turning flow control off is a different failure mode, not a fix

Setting producerFlowControl="false" on a destination changes blocking into spooling to the temp store or, depending on sendFailIfNoSpace, dropping messages. If the temp store fills with flow control disabled, non-persistent messages can be discarded silently. Use this only on destinations where loss is acceptable, and pair it with sendFailIfNoSpace="true" so producers at least get an error.

What not to do first

  • Do not restart the broker. On restart the store replays the backlog and cursors page messages back into memory, so the noisy queue refills the pool and you are back at 100 percent, minus the time you lost. Nothing structural has changed.
  • Do not just raise the broker memoryUsage. A bigger shared pool with no per-destination caps is the same incident on a longer fuse. Also keep the broker memory limit around 60 to 70 percent of JVM max heap; push it higher and you trade flow control for GC stalls and OOM risk, which are worse.

Prevention

  • Per-destination memory limits. Set memoryLimit on every policyEntry, with tighter caps on shared and multi-tenant brokers, so no single destination can drain the shared pool.
  • Alerting below the cliff. Ticket above 80 percent on both per-destination and broker-level memory. Flow control is a cliff-edge at 100, so threshold alerts at 100 only confirm an outage already in progress.
  • Loud producers. Configure sendFailIfNoSpaceAfterTimeout so blocked sends surface as exceptions in application logs instead of silent hangs.
  • Leading indicators. Track the enqueue/dequeue ratio and queue depth growth on busy destinations. A ratio sustained above 1.0 is your runway warning long before memory is involved.
  • Durable subscription hygiene. Unsubscribe decommissioned durable subscribers; an offline durable subscription accumulates every published message forever and shows up later as unexplained topic memory pressure.
  • Heap alignment. Keep broker memoryUsage near 60 to 70 percent of JVM max heap and monitor both numbers independently, because either one can fill first.

How Netdata helps

  • Netdata’s ActiveMQ monitoring charts per-destination MemoryPercentUsage alongside broker-level memory, so ranking destinations during an incident is a glance at a dashboard instead of a sequence of JMX queries.
  • Correlating destination memory with QueueSize, ConsumerCount, and dequeue rate on the same view is what lets you distinguish a stalled consumer from a producer burst in minutes.
  • Alerts at 80 percent catch the climb before the cliff-edge at 100 where flow control engages, which is the difference between a ticket and a page.
  • JVM heap is charted next to ActiveMQ’s internal memory accounting, which matters because the two fail differently and the most common misdiagnosis here is confusing them.
  • Per-second history shows whether the growth was a spike (replay or traffic event) or a slow burn (consumer gradually falling behind), and that shape determines which fix applies.