The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / activemq / activemq-message-age-oldest-pending ▌

Operations Guides

ActiveMQ oldest message age: the queue latency depth alone cannot show

Every ActiveMQ dashboard has queue depth on it. Almost none show the age of the oldest pending message, which is the number that actually maps to your SLA. A queue holding 50 messages whose head has been waiting 45 minutes is a much bigger problem than a queue holding 50,000 messages whose head is 2 seconds old and draining fast. Depth alone cannot tell you which of those two situations you are looking at.

ActiveMQ Classic does not expose oldest message age as a direct JMX attribute on most versions in the field. You have to derive it, and the derivation has two traps that produce confidently wrong numbers if you do not know about them. The timestamp you read belongs to the producer’s clock, not the broker’s. And the cheapest way to read it, browsing the queue, gets expensive exactly when the queue is deep enough that you need the number most.

This article covers how the age signal is derived on ActiveMQ Classic, the clock-skew and browse-cost traps, what a rising age is actually telling you, and how to alert on age instead of raw depth. For the broader broker signal taxonomy, see the monitoring checklist.

Why queue depth is not a latency signal

QueueSize is a backlog gauge, not a latency gauge. Three properties make it ambiguous as an SLA signal:

  • It includes inflight messages. QueueSize counts messages dispatched to consumers but not yet acknowledged. A queue showing 1,000 with InFlightCount at 1,000 is drained from the broker’s perspective; the messages are sitting in consumer prefetch buffers. High prefetch makes this worse: messages live in client memory and the broker looks emptier than the pipeline really is.
  • Depth has no universal threshold. A batch queue may sit at tens of thousands of messages by design. A synchronous order queue is in trouble at a few hundred. Raw counts force you to maintain per-queue magic numbers that still say nothing about wait time.
  • Time-to-clear is a proxy, not a measurement. depth / dequeue_rate is a useful estimate, but it breaks down exactly when you need it: when dequeue rate collapses to zero the estimate is undefined, bursty rates make it swing wildly, and it assumes strict FIFO dispatch, which selectors and message groups violate.

The age of the message at the head of the queue answers the SLA question directly: how long has the longest-waiting message been waiting. That is why a shallow queue with old messages is worse than a deep queue with fresh ones, and why age belongs on the dashboard next to depth, not inferred from it.

How oldest message age is derived

There is no OldestMessageAge attribute on the destination MBean in Classic 5.x releases prior to 5.19.0. Some versions expose MinEnqueueTime, but it is not reliable across versions, so do not build alerting on it. The practical methods, in order of how often you will use them:

1. Browse the queue and read JMSTimestamp of the head

The destination MBean has a browse() operation. The first message in the result is the head of the queue, and its JMSTimestamp header is epoch milliseconds set at send time. Age is now - JMSTimestamp.

# Browse the head of a queue via Jolokia. WARNING: expensive on deep queues.
curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/exec/org.apache.activemq:type=Broker,brokerName=localhost,destinationType=Queue,destinationName=MY.QUEUE/browse'
# Parse JMSTimestamp from the first element of the returned array.
# age_seconds = (current_epoch_ms - JMSTimestamp) / 1000

The command-line equivalent shows headers without writing a client:

# Show message headers including JMSTimestamp for a queue
activemq-admin browse --amqurl tcp://localhost:61616 -Vheader MY.QUEUE

The official command-line tools documentation warns that browse may not return all messages due to broker configuration and resource limits. For age purposes that is fine, because you only need the head, but do not use browse output as an inventory of the queue.

2. The StatisticsPlugin’s firstMessageTimestamp

The StatisticsPlugin (available since 5.3) can return firstMessageTimestamp for a destination when the statistics query message requests it. The reply message’s own JMSTimestamp is the broker’s current time, so the client can compute age = reply_JMSTimestamp - firstMessageTimestamp using only broker-side clocks. Internally this performs a browse of one message from the head, so the per-call cost is small, but it is still a browse. The upstream motivation for adding a tracked attribute instead was precisely that a browse-based statistic is difficult to scale when you poll hundreds or thousands of queues.

3. Tracked enqueue timestamps on 5.19.0 and 6.2.0+

AMQ-8463 (first released in 5.19.0 and 6.2.0) added advanced message statistics to destination MBeans: EnqueuedMessageTimestamp reports the producer’s JMSTimestamp and EnqueuedMessageBrokerInTime the broker arrival time of the most recently enqueued message, tracked as messages flow through the destination, with no browse required. The feature is disabled by default and is enabled per destination with advancedMessageStatisticsEnabled="true" on the policy entry. Because the value is overwritten on every enqueue it reflects the newest message, not the oldest pending one; treat it as a flow-activity signal for continuous alerting, and use browse or the StatisticsPlugin when you need the head message’s age.

If you are pinned to an older release, you are back to methods 1 and 2, used sparingly.

Two traps: producer clocks and browse cost

Trap 1: JMSTimestamp is the producer’s clock

JMSTimestamp is set by the sending client from its local clock. The broker does not rewrite it. Consequences:

  • Clock skew between producers and the broker directly distorts computed age. A producer 10 minutes ahead makes every message look 10 minutes old on arrival. A producer behind the broker can yield negative ages.
  • In mixed fleets (different hosts, containers, cloud regions), skew varies per producer, so the distortion is not even consistent.
  • If your hosts are not disciplined by NTP, treat browse-derived age as approximate. Fixing time sync is the real fix.

Two broker-side alternatives exist. JMSActiveMQBrokerInTime is an ActiveMQ-specific pseudo-property backed by the OpenWire Message.brokerInTime field, stamped with the broker’s clock when the message arrives; it is more reliable for age math, but it is not part of the JMS spec, reads back as 0 on messages the broker never stamped, and is not readable over non-OpenWire protocols (STOMP, AMQP, MQTT). The broker’s timestamp plugin can overwrite JMSTimestamp with broker time, but the official documentation warns this breaks JMS compliance: the timestamp the producer sees on the message after send() will differ from what the consumer observes. Choose deliberately.

Trap 2: browsing is expensive on deep queues

A browse pages messages out of the store into memory. Page size is bounded by the destination policy’s maxBrowsePageSize (default 400). On a queue with 100K or more pending messages, aggressive or full browses spike broker memory and CPU and can slow checkpointing, which is the last thing you want on a queue that is already backing up.

Practical rules:

  • Browse only critical queues, and only the head. Never run full-queue browses in a monitoring loop.
  • Keep the polling cadence slow (tens of seconds to minutes). Age is an SLA signal; it does not need per-second resolution, and high-frequency JMX polling on a broker with many destinations is itself a measurable load.
  • On 5.19.0+/6.2.0+, prefer the tracked enqueue timestamp attributes and stop browsing entirely.
  • For ad-hoc human inspection, the web console’s first-page view is cheaper than a scripted browse.

What rising age is telling you

Age is a symptom, not a cause. The same rising number maps to several distinct failure modes, and the cheap JMX attributes you already collect disambiguate them before you touch the queue contents:

flowchart TD
  A[Oldest message age rising past SLO] --> B{ConsumerCount above zero?}
  B -->|no| C[Consumers down or partitioned]
  B -->|yes| D{Dequeue rate above zero?}
  D -->|yes but age rising| E[Consumers lagging: too slow or too few]
  D -->|no| F{InFlightCount pinned at prefetch?}
  F -->|yes| G[Stuck or zombie consumers]
  F -->|no| H[Selector mismatch or pinned message group]
  A -.->|age stays young but ExpiredCount rising| I[Silent expiry at the head]

Working through the branches:

  • ConsumerCount is zero. Nobody is draining the queue. This is the cleanest case and the one most dashboards already catch, but age tells you how bad the backlog is in time units while consumer count only tells you that it exists.
  • Consumers connected, dequeue zero, inflight pinned at prefetch. The zombie consumer pattern: messages were dispatched into prefetch buffers and are never acknowledged. The broker thinks it did its job. If inflight equals consumer_count x prefetch_size for more than a couple of minutes, the consumers are stuck, not slow. Restarting or disconnecting them requeues the inflight messages to healthy consumers.
  • Consumers connected, dequeue zero, inflight low. The broker is not even dispatching. Suspect a selector mismatch (no connected consumer’s selector matches the head message) or JMS message groups, where all messages for a group are pinned to one consumer and a slow group owner holds the whole group hostage. Aggregate depth hides this completely.
  • Dequeue positive but age still rising. Consumers are working but losing ground. This is a capacity problem: scale consumers out or speed up their downstream dependency.
  • Age stays young while ExpiredCount climbs. The subtle one. Messages are expiring at the head as fast as new ones arrive, so the head is always fresh and depth can look stable with zero dequeue. The broker looks healthy while silently losing business events. Expired messages route to the DLQ by default unless processExpired="false", so DLQ growth is your corroborating signal.
  • Dequeue looks healthy but age was the complaint. Check whether the “dequeues” are actually DLQ transfers. Messages moved to the DLQ count as dequeued from the source queue, so a poison-message loop drains age and depth while the redelivery rate climbs. Redelivery is the leading indicator; DLQ growth is the lagging one.

Alerting on age instead of depth

Raw depth is not 3AM-safe: it pages for batch queues doing their job and stays silent for shallow queues full of stale work. Structure the alerting like this:

  • Define the threshold in time, per queue. Base it on the business SLO for end-to-end processing. Where messages carry a TTL, express the threshold as a fraction of that TTL. There is no global number.
  • Ticket on age alone. Oldest message past the per-queue SLO means the SLA is already breached; someone should look during working hours even if consumers are running.
  • Page on the composite. On critical queues: oldest message age past SLO AND (consumer count at zero OR dequeue rate collapsed), sustained, with broker uptime above 600 seconds to exclude restart recovery. The composite kills the false pages from transacted consumers, batch drains, and post-restart catch-up, all of which transiently push age up.
  • Keep time-to-clear as a companion, not a replacement. When dequeue rate is healthy and roughly FIFO holds, depth / dequeue_rate is a good runway estimate. When the rate is zero or selectors are in play, only the head timestamp tells the truth.

Signals to watch in production

SignalWhy it mattersWarning sign
Oldest message age (derived)The direct SLO/latency measurement for the queueAge past per-queue SLO, sustained
QueueSizeBacklog gauge, includes inflightAbove 2x baseline for >15 min, or stable with zero dequeue
Dequeue rate (from DequeueCount deltas)Consumption velocityZero while consumers are connected; healthy-looking but actually DLQ transfers
InFlightCount vs prefetchDetects stuck consumers that depth missesInflight equals consumer count x prefetch for >2 min
ConsumerCountWhether anyone is draining at allZero on a critical queue with pending messages
ExpiredCountSilent loss that keeps the head deceptively youngRising with stable depth and low dequeue
Redelivery rateLeading indicator for poison messages before DLQ growthSustained above baseline
DLQ QueueSizeLagging confirmation of poison/expiry loopsAny unexpected growth

How Netdata helps

  • Netdata’s ActiveMQ monitoring charts the per-destination signals this diagnosis depends on (depth, consumer count, unacked messages, cumulative enqueue/dequeue counters) at high resolution, so consumption velocity and backlog are visible on the same timeline without running browse operations against the broker.
  • Derived rates from the cumulative DequeueCount make the “consumers connected but dequeue collapsed” condition visible, which is the branch that turns rising age into a page.
  • Charting expired counts and DLQ depth next to queue depth catches the two cases where head age stays deceptively young: silent expiry and DLQ transfers counting as dequeues.
  • Alerts on composite conditions (consumer count at zero with nonzero depth, unacked pinned against expected prefetch) implement most of the page logic from cheap attributes, so browse-based age checks can be reserved for critical queues on a slow cadence.
  • ML anomaly detection on dequeue rate and depth flags the consumption collapse that usually precedes an age SLO breach, giving you the ticket before the page.