The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / activemq / activemq-prefetch-tuning ▌

Operations Guides

ActiveMQ prefetch limit: unacked hoarding versus idle consumers

You have four consumers on a queue, producers are sending steadily, and yet throughput is a quarter of what you expect. ConsumerCount says 4. QueueSize looks small. DequeueCount is crawling. The broker looks healthy and the consumers look connected, but three of them are doing nothing.

This is the classic prefetch mismatch. The ActiveMQ prefetch limit controls how many messages the broker pushes to a consumer before it requires acknowledgments back. It is a client-side buffer the broker fills eagerly. With the default of 1000 for queues, the first consumer to connect can absorb the entire visible backlog into its prefetch buffer, where those messages are invisible to load balancing, no longer read as pending depth, and pinned against broker memory until they are acked.

What the prefetch limit actually controls

The prefetch limit is the maximum number of messages the broker will dispatch to a consumer without waiting for acknowledgments. It exists to keep consumers busy: without prefetching, every message would require a round trip, and consumer throughput would collapse to network latency per message.

Defaults in ActiveMQ Classic 5.x:

Destination typeDefault prefetch
Queue (persistent and non-persistent)1000
Topic (persistent)100
Topic (non-persistent)32766 per the official docs; some sources and versions report 32767 (Short.MAX_VALUE)

The key operational facts:

  • Prefetched messages are dispatched but not acknowledged. They appear in InFlightCount, and QueueSize includes inflight messages. A queue with QueueSize=1000 and InFlightCount=1000 is, from the broker’s dispatch perspective, drained. Everything is sitting in consumer buffers.
  • Prefetched messages still count against broker memory accounting until acknowledged. A large prefetch across many consumers is a real memory commitment, not a free performance knob.
  • Once a consumer’s prefetch buffer is full, the broker stops dispatching to it until acks free up slots. Per the official documentation, dispatch resumes in top-up batches once the consumer has acked roughly half of the prefetched messages.
  • Setting prefetch to 0 switches the consumer to polling mode: it pulls one message at a time instead of receiving pushed messages. This changes the performance profile completely and, as a side effect reported by operators, can inflate the console’s Enqueues counter.

For the broader picture of how dispatch, cursors, and memory accounting fit together, see how ActiveMQ Classic actually works in production.

How dispatch and prefetch interact

Queue dispatch in ActiveMQ Classic is competing-consumer: each message goes to exactly one consumer. The dispatch thread is nominally round-robin, but there is a subtlety that bites operators constantly: the broker fills one consumer’s prefetch buffer before moving on to the next. The official dispatch policies documentation is explicit that with many consumers, a high prefetch value, and a small number of messages, load balancing is not fair.

flowchart LR
  P[Producers] --> Q[Queue pending cursor]
  Q --> D[Dispatch thread]
  D -->|fills prefetch 1000| CA[Consumer A buffer FULL]
  D -.->|nothing left to push| CB[Consumer B idle]
  D -.->|nothing left to push| CC[Consumer C idle]
  CA -->|acks free slots| D

Consumer A is not greedy or buggy. It connected first, and the broker did exactly what the prefetch setting told it to do. The messages in A’s buffer are now exclusive to A: they cannot be redispatched to B or C unless A disconnects or its acks free slots and new messages arrive.

The two failure modes

Unacked hoarding

One consumer holds a large number of dispatched-but-unacked messages while processing them slowly. Consequences:

  • Hidden backlog. Queue depth looks modest because the messages are inflight, not pending. Operators watching only QueueSize conclude the queue is fine while hundreds of messages sit unprocessed in one consumer’s buffer. The extreme version is the zombie consumer pattern described in InFlightCount high: QueueSize near zero, InFlightCount high, dequeue rate near zero.
  • Pinned broker memory. Unacked messages remain charged against memory accounting. A consumer with prefetch 1000 holding large messages can push a destination toward its memory limit on its own.
  • Pinned store. For persistent messages, journal files are only reclaimable when every message in the file is acknowledged. A hoarded message that never gets acked pins its journal file, which is how prefetch misconfiguration feeds the journal files not deleted problem.
  • Slow failure detection. The consumer is connected and the broker sees no error. Only the inflight-versus-dequeue correlation reveals it.

Idle consumers

The mirror image: prefetch too large relative to message volume, so the dispatch thread drains everything into whichever consumer it visits first and the rest starve. You pay for four consumer instances and get the throughput of one. This is worst when message volume is low or bursty and processing time is long, because the hoarding consumer never drains its buffer fast enough for the imbalance to self-correct.

The failure mode in the other direction is prefetch too small: consumers finish each message, then wait for the ack round trip and the next dispatch before they have more work. Throughput drops even though CPU on the consumer side is idle. You see this as a low, steady dequeue rate with healthy consumers and a queue that never drains as fast as it should.

Acknowledgment mode changes the math

The prefetch limit interacts with ack mode, and you cannot reason about one without the other.

  • AUTO_ACKNOWLEDGE. The client acks automatically, so InFlightCount stays lower and prefetch slots free quickly. The official performance tuning documentation and the client source state that with optimizeAcknowledge the ack batch is sent once delivered plus acked messages reach 65% of the prefetch limit, or every 300 ms (optimizeAcknowledgeTimeOut) when consumption is slow (an older documentation version cites 50%). The trap: acking is decoupled from how fast your downstream processing actually drains, so the broker keeps a consumer’s buffer topped up even while that consumer’s work is backing up and a peer sits idle. Prefetch hoarding can still happen; it just does not always show up as high broker inflight.
  • CLIENT_ACKNOWLEDGE and transacted sessions. Inflight runs high by design, because acks (or commits) lag dispatch deliberately. High inflight here is not a stuck consumer by itself; the warning sign is inflight pinned at exactly consumer_count x prefetch with a collapsing dequeue rate.
  • optimizeAcknowledge. Changes when batch acks are sent, which changes how quickly prefetch slots free: the client sends a batch when delivered plus acked messages reach 65% of the prefetch limit, or when optimizeAcknowledgeTimeOut (default 300 ms) expires. If you enable it, verify the interaction with your prefetch size empirically.

Sizing prefetch against your workload

There is no universally correct number. Size against message processing time, message size, and how evenly work must fan out.

WorkloadSuggested starting pointWhy
Fast, uniform processing (sub-100ms), high volume100-1000 (default is fine)Batching amortizes round trips; imbalance self-corrects because buffers drain quickly
Slow, uneven processing (seconds per message), must fan out evenly1-10Each consumer holds almost nothing, so dispatch stays fair and a stuck consumer pins almost nothing
Large messagesLower than defaultPrefetched bytes are charged to broker memory; 1000 large messages is a serious memory commitment
STOMP or scripting-language consumers that cannot buffer1The official docs recommend prefetch 1 for consumers that cannot cache prefetched messages
Pooled/cached consumers (e.g. caching connection factories)0 (polling) or disable consumer cachingPooled consumers defer close, so prefetched messages sit unconsumed in a buffer nobody is draining

The prefetch=1 recommendation for slow, uneven workloads deserves emphasis. Operators often raise prefetch chasing throughput, when the actual problem is that one slow message blocks 999 others in the same buffer. With prefetch 1, a slow message delays exactly one message, and the remaining work flows to the other consumers.

Where this shows up in production

Concrete checks that distinguish hoarding from starving on a live broker:

# Queue depth vs inflight: are messages pending or sitting in consumer buffers?
curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/org.apache.activemq:type=Broker,brokerName=localhost,destinationType=Queue,destinationName=MY.QUEUE/QueueSize'

curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/org.apache.activemq:type=Broker,brokerName=localhost,destinationType=Queue,destinationName=MY.QUEUE/InFlightCount'

# How many consumers are supposed to be sharing the load?
curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/org.apache.activemq:type=Broker,brokerName=localhost,destinationType=Queue,destinationName=MY.QUEUE/ConsumerCount'

Reading the results:

  • QueueSize roughly equal to InFlightCount, with ConsumerCount > 1 and a low dequeue rate: one or few consumers are hoarding. Check per-consumer DispatchedQueueSize or MessageCountAwaitingAcknowledge on the subscription MBeans to find which one.
  • InFlightCount near zero, dequeue rate low, consumers healthy and idle: prefetch may be too small, or the bottleneck is downstream of the broker entirely.
  • InFlightCount pinned at ConsumerCount x prefetch sustained: every consumer’s buffer is full. If dequeue is not keeping pace, the consumers are saturated or stuck, which is the pattern covered in InFlightCount high.

Broker-side memory pressure is the other tell. If hoarded messages are large, watch MemoryPercentUsage on the destination and the broker. Prefetch-driven hoarding is one of the less obvious paths to the flow-control cliff described in MemoryPercentUsage climbing.

Signals to watch in production

SignalWhy it mattersWarning sign
InFlightCount per destinationMessages sitting in consumer prefetch buffersEquals consumer_count x prefetch sustained, or high while QueueSize is near zero
InFlightCount / QueueSize ratioShows how much of the visible depth is actually hoarded vs pendingRatio near 1.0 with multiple consumers: load balancing is defeated
Per-consumer DispatchedQueueSizeIdentifies which specific consumer is hoardingOne consumer at its prefetch limit while peers are at zero
Dequeue rate vs consumer countReveals whether all consumers are actually workingThroughput of one consumer while N are connected
MemoryPercentUsage (destination and broker)Hoarded unacked messages still cost broker memoryClimbing memory with modest queue depth
Message age of oldest pending messageThe latency symptom hoarding hidesOld messages despite shallow queue depth; see oldest message age

A useful composite alert: page only when inflight is pinned near total prefetch AND dequeue has collapsed AND message age is rising, sustained on a critical queue. Any one of those alone false-pages on transacted or delayed-ack consumers.

How Netdata helps

  • Netdata’s ActiveMQ collector pulls JMX destination metrics including QueueSize, InFlightCount, ConsumerCount, and enqueue/dequeue counts, so the hoarding signature (depth equal to inflight, dequeue collapsed, consumers connected) is visible on one dashboard instead of three curl commands.
  • Per-second collection catches the transient version of this problem: a consumer that hoards for 30 seconds during a GC pause and then recovers, which minute-scrape monitoring smooths into invisibility.
  • Correlating InFlightCount against MemoryPercentUsage on the same timeline shows when a prefetch buffer, not a true backlog, is what is pushing the broker toward flow control.
  • Alerting on the ratio of inflight to queue size, and on dequeue rate per connected consumer, turns the idle-consumer failure mode into a detectable condition rather than a quiet waste of consumer capacity.
  • Historical retention lets you verify a prefetch change actually worked: compare inflight distribution and dequeue rate before and after, rather than trusting the first five minutes.