The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / activemq / activemq-monitoring-checklist ▌

Operations Guides

ActiveMQ monitoring checklist: the signals every production broker needs

Most ActiveMQ monitoring setups fail in one of two ways. Either they only watch process liveness and queue depth, and discover producer flow control from angry users. Or they export every JMX attribute into a dashboard nobody reads, and the one signal that mattered is buried on page four.

This checklist organizes the signals that matter into four maturity levels, from “is the broker alive” to “can you prove end-to-end message flow works.” Each level lists the signal, where it comes from, and the threshold that makes it actionable. Scope is ActiveMQ Classic (5.x and 6.x Classic stream). Artemis has a different store, memory model, and MBean tree; do not apply these thresholds to it.

Use it two ways: as an audit of what you have today, and as a build order for what to add next.

Why queue depth alone is not enough

The four signals teams most often monitor are process up/down, total connections, queue depth, and CPU. None of them catch ActiveMQ’s two most common production failures: producer flow control and store exhaustion.

Flow control is silent by design. When a destination or broker memory limit is reached, the broker stops reading from the producer’s socket. The producer’s send() blocks with no exception, no log line, no timeout. Upstream services hang and operators see a “slow application,” not a full broker. Store exhaustion is the mirror image: KahaDB journal files accumulate for weeks because one unacknowledged message pins an entire 32MB journal file, until persistent messaging halts.

Both failures are visible days in advance if you watch the right counters. The checklist below is built around catching them early.

flowchart TD
  L1[Level 1: Survival - is the broker alive and accepting work]
  L2[Level 2: Operational - throughput, backlog, consumer health]
  L3[Level 3: Mature - per-destination detail, latency, topology]
  L4[Level 4: Expert - synthetic checks, saturation ratios, internals]
  L1 --> L2 --> L3 --> L4

Level 1: survival

The absolute minimum. If you monitor nothing else, monitor these. Each one maps directly to a total-service-loss failure mode.

SignalSourceWarning sign
Broker process accepting connectionsTCP connect to the OpenWire port (default 61616)Port closed or connect timeout on the active broker
Broker memory usage percentJMX MemoryPercentUsage on the Broker MBean100% means producer flow control is active; ticket above 80% sustained
Store usage percentJMX StorePercentUsage100% means persistent messaging is halted; ticket above 80%
JVM heap used after GCJMX java.lang:type=Memory, HeapMemoryUsageAbove 95% after a major GC, repeated across GC windows
Disk free on the KahaDB partitionOS df on the data partitionAbove 90% used; ticket at 80%
Consumer count on critical queuesPer-destination ConsumerCountZero on a queue that should be draining, with enqueue rate above zero
DLQ depthQueueSize on ActiveMQ.DLQAny unexpected growth; every message is a failed business transaction

Two of these need extra care:

Broker memory is not JVM heap. MemoryPercentUsage is ActiveMQ’s internal accounting against its configured memoryUsage limit. JVM heap can be exhausted by MBeans, connection state, and cursor metadata while broker memory shows headroom. Monitor both, and keep the ActiveMQ memory limit at roughly 60-70% of JVM max heap.

Store usage is not disk usage. StorePercentUsage measures against the limit configured in activemq.xml. If that limit is set higher than the physical partition, the OS fills the disk before ActiveMQ’s counter reaches 100%. Watch both independently. This is one of the most common ways teams get surprised.

HA caveat: in a shared-storage pair, the standby broker’s process runs but its transport connectors are not started. Do not alert on the standby’s closed ports.

Level 2: operational

Everything in Level 1, plus the signals that tell you whether the broker is keeping up and where pressure is building. This is the level a competent production team should reach.

SignalSourceWarning sign
Per-queue depthPer-destination QueueSizeSustained above 2x baseline for more than 15 minutes
Enqueue and dequeue ratesEnqueueCount / DequeueCount (cumulative; derive rates)enqueue - dequeue sustained positive; zero dequeue with nonzero depth
Per-destination consumer countConsumerCountBelow the expected minimum for that queue
GC pause duration and frequencyjava.lang:type=GarbageCollector MBeansPauses above 1s, or full GC more than once per 5 minutes
Connection countCurrentConnectionsCount on the Broker MBeanDeviation over 50% from baseline; sawtooth pattern
Inflight message countPer-destination InFlightCountInflight pinned at prefetch size for more than 2 minutes
Temp store usageTempPercentUsageAny sustained nonzero value
File descriptor usageOpenFileDescriptorCount vs MaxFileDescriptorCountAbove 70% of limit
KahaDB journal file countFilesystem count of db-*.logCount above 2x baseline or growing steadily
Expired message countPer-destination ExpiredCountAny unexpected sustained expiry
Redelivery rateDestination/consumer stats, JMSXDeliveryCountSustained increase above baseline

Three interpretive rules prevent the most common misreads at this level:

  1. QueueSize includes inflight messages. A queue showing QueueSize=1000 with InFlightCount=1000 is drained from the broker’s perspective; all the messages are sitting in consumer prefetch buffers. The inverse is the zombie consumer pattern: QueueSize=0 with high InFlightCount and no dequeue rate means consumers received messages and never acked. Always read depth and inflight together.

  2. Dequeue counts DLQ transfers. Messages moved to the dead letter queue increment the source queue’s dequeue counter. A “healthy” dequeue rate can actually be the broker discarding poison messages. Cross-check dequeue against DLQ enqueue rate before trusting it.

  3. GC pauses cascade into connection storms. The default OpenWire wireFormat.maxInactivityDuration is 30000ms. A GC pause longer than that disconnects every client, which then reconnects simultaneously, allocating more objects and making the next GC worse. If you see a sawtooth in connection count, look at GC logs first.

On file descriptors: the default Linux ulimit of 1024 is the single most common production misconfiguration. Each client connection, journal file, and log file consumes one. Raise the limit to at least 65536 and alert on open FDs above 70% of it.

Level 3: mature

Everything above, plus per-destination isolation, latency, and topology signals. This level is where you stop reacting to saturation and start seeing it days out.

  • Per-destination memory usage. Isolates the noisy destination before it trips broker-wide flow control. Without per-destination limits set via <policyEntry>, one runaway queue can starve every other producer.
  • Message age on critical queues. The most business-relevant latency signal: a shallow queue with a two-hour-old oldest message is worse than a deep queue of fresh messages. There is no direct JMX attribute; derive it by browsing the first message and reading JMSTimestamp. Queue browses are expensive on deep queues, so sample, do not poll aggressively. The timestamp comes from the producer’s clock, so NTP skew distorts it.
  • Per-consumer dispatch metrics. DispatchedQueueSize and MessageCountAwaitingAcknowledge on subscription MBeans identify which specific consumer is stuck, not just that a consumer is stuck.
  • Durable subscriber pending count. Offline durable subscriptions accumulate messages forever. A decommissioned test subscriber that was never unsubscribed is a permanent storage leak. Alert on any offline subscriber with growing pending messages.
  • Network bridge status and throughput (Network of Brokers only). Bridge down is a partition; bridge connected with zero throughput is demand-forwarding broken. Both brokers look healthy in isolation while messages pile up on one and consumers idle on the other. Watch store-and-forward replay on reconnect; the burst can push the receiving broker into flow control.
  • Total destination count. Each destination creates at least four MBeans. Unbounded growth (usually dynamic destination creation without cleanup) produces JMX sluggishness, then GC pressure, then outage. Alert above 2x expected count. Advisory topics inflate this silently; set advisorySupport="false" where you do not need advisories.
  • Temporary destination count. Temp destinations should cycle with request-reply traffic. Monotonic growth is a leak, usually connection pooling keeping connections (and their temp destinations) alive.
  • Store write latency. iostat -x on the KahaDB device. Journal fsync latency is the hard ceiling on persistent throughput. Under 2ms is healthy on SSD; sustained writes above 10ms mean investigate.
  • KahaDB index file size. Watch db.data. Under 100MB is healthy; above 1GB degrades lookup performance and stretches crash-recovery startup to tens of minutes.
  • Thread count. Roughly 50 plus 1-2 per connection with the default TCP transport. Growth without connection growth is a leak. NIO transport decouples threads from connections, so this signal is less useful there.
  • Authentication failure rate. From broker logs. Sporadic failures are usually a misconfigured client; broad multi-source failures after a credential change are urgent.
  • HA role and lock state (shared-storage HA only). Split-brain, both brokers active, risks store corruption and is a page. A standby that cannot acquire the lock after the active dies means failover is broken.

Level 4: expert

Signals teams add after the third or fourth major incident, when the aggregate metrics looked green but the system was still wrong.

  • Canary message round-trip. Continuously send a test message to a canary queue and consume it back, measuring latency and success. This is the single best health signal because a broker can be metric-green and functionally dead (dispatch bug, selector mismatch, authorization change blocking consumers). A TCP connect proves a listener exists; a canary proves the broker works.
  • Prefetch buffer saturation. The inflight/prefetch ratio per consumer, approaching 1.0. This is the earliest form of the slow-consumer signal, before memory usage moves.
  • Connection create/destroy rate. Churn is invisible in the connection count. A broker can hold a steady 500 connections while 100 clients reconnect every second, burning threads and GC.
  • Per-message-group depth. With JMS message groups, one stuck group owner is invisible in aggregate depth.
  • Advisory topic resource consumption. Monitoring the monitoring overhead.
  • Producer count per destination. Catches anomalous producers (a misconfigured retry loop) before their enqueue rate shows up in aggregates.
  • Scheduled message count. Messages in the scheduler store are invisible to normal queue depth.
  • MBean count and JMX query latency. JMX polling itself becomes a bottleneck at high destination counts; bulk queries and longer scrape intervals matter.

The thresholds worth committing to memory

These come up in nearly every ActiveMQ incident review:

  • MemoryPercentUsage 100% equals producers silently blocked. Page when producers are active.
  • StorePercentUsage 100% equals persistent messaging halted. Page on the active role.
  • Disk above 90% on the KahaDB partition equals imminent failure, and may fire before StorePercentUsage does.
  • Heap above 95% after major GC, repeated, equals OOM imminent.
  • GC pause above 30s equals every OpenWire client disconnects (default inactivity timeout).
  • Zero consumers on a critical queue with pending messages equals page.
  • Any DLQ growth equals a processing failure per message. Ticket, always.
  • FDs above 90% of limit with accept failures equals page.

How Netdata helps

The value here is correlation, not collection. The signals above only work in combination:

  • Netdata charts JVM heap, GC pause time, and thread count alongside per-broker and per-destination JMX metrics on the same node, so the GC-to-connection-storm cascade shows up as one timeline instead of three tools.
  • Broker memory usage, store usage, and temp store usage are collected as distinct signals, which keeps the “ActiveMQ memory is not JVM heap” and “store limit is not disk space” distinctions visible instead of averaged away.
  • Per-second collection catches the sawtooth connection pattern and inflight spikes that minute-interval scrapes flatten into normal-looking averages.
  • Disk space and block device latency on the KahaDB partition sit on the same dashboard as StorePercentUsage, so you can see the store limit and the physical disk racing each other before either hits the wall.
  • Alerting on ConsumerCount per destination, DLQ depth, and ExpiredCount catches the silent-correctness-loss patterns that throughput dashboards never surface.