The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / activemq / activemq-disk-full-kahadb ▌

Operations Guides

ActiveMQ disk full on the KahaDB partition: write failures and store corruption risk

The KahaDB partition is at 100%. The broker log shows journal write failures, persistent producers have stopped, and you are in the worst ActiveMQ failure mode: the one where freeing space may not be enough, because an in-flight write at the moment the disk filled can leave the journal or index corrupt.

This is not the same incident as StorePercentUsage hitting 100%. StorePercentUsage measures KahaDB against the configured storeUsage limit in activemq.xml. Disk full measures the partition against physical capacity with df. They are independent limits, and either can fire first. If storeUsage is larger than the partition, the OS runs out of space while the broker still thinks it has headroom, and the failure arrives with no warning from any ActiveMQ metric. This guide covers the OS-level case; for the configured-limit case, see ActiveMQ store is full.

The corruption risk is specific: the broker acks persistent producers only after the journal fsync completes. A disk that fills mid-write can leave a partially written journal record or a torn index update. On the next restart, KahaDB recovery may fail, take hours replaying journals, or require a forced index rebuild.

What this means

KahaDB keeps three things on this partition, typically under data/kahadb/ (or activemq-data/ depending on packaging):

  • db-*.log: sequential journal files, 32MB each by default. This is the write-ahead log every persistent message hits first.
  • db.data: the B-tree index mapping message IDs to journal locations. Page-file I/O, and it grows with pending message count.
  • db.redo: the redo log used during recovery.

The same partition usually also holds the temp store (data/tmp_storage) for non-persistent overflow and often the broker log files. They all compete for the same bytes.

When the partition fills, the sequence is:

flowchart TD
  A[Partition reaches 100%] --> B[Journal write fails: ENOSPC]
  B --> C[Persistent sends cannot be journaled]
  C --> D[Persistent messaging halts]
  B --> E[Write interrupted mid-record]
  E --> F[Journal or index corruption risk]
  F --> G[Next restart: recovery fails or runs long]
  D --> H[Producers block or error]

A journal file is only reclaimable when every message in it has been acknowledged, so disk consumption is pinned by the oldest unacked messages, not by average throughput. That is why this incident usually has a slow, visible runway (days of journal growth) followed by a cliff.

Common causes

CauseWhat it looks likeFirst thing to check
Message backlog (consumers lagging)Journal file count grows steadily over days; queue depths highQueueSize and consumer count on the big queues
DLQ accumulationActiveMQ.DLQ grows; DLQ messages pin journal filesQueueSize on ActiveMQ.DLQ
Journal file pinningMany db-*.log files remain after backlog drains; one unacked message pins a whole 32MB fileJournal file count vs. actual pending messages
Temp store growthtmp_storage consuming space; TempPercentUsage elevateddu -sh on the temp store directory
Broker logs on same partitionactivemq.log and rotated logs consuming gigabytesdu -sh on the log directory
Non-ActiveMQ processesSomething else on the host writing to the same mountdu -x on the mount’s top-level directories
Store limit larger than diskStorePercentUsage well under 100 while df shows fullCompare storeUsage config to partition size

Quick checks

# 1. Confirm the partition and its usage
df -h /opt/activemq/data/kahadb/

# 2. See what is consuming space (stay on this filesystem with -x)
du -xh --max-depth=1 /opt/activemq/data/ | sort -rh | head -20

# 3. Count and size the journal files
ls /opt/activemq/data/kahadb/db-*.log | wc -l
du -sh /opt/activemq/data/kahadb/

# 4. Check the index size (large index also means slow recovery later)
ls -lh /opt/activemq/data/kahadb/db.data

# 5. Compare broker-side store accounting vs physical disk
curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/org.apache.activemq:type=Broker,brokerName=localhost/StorePercentUsage'

# 6. Check DLQ depth, the most common silent space consumer
curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/org.apache.activemq:type=Broker,brokerName=localhost,destinationType=Queue,destinationName=ActiveMQ.DLQ/QueueSize'

# 7. Find what else is on this mount (non-ActiveMQ consumers)
du -xh --max-depth=1 "$(df --output=target /opt/activemq/data/kahadb/ | tail -1)" | sort -rh | head -20

# 8. Look for write failures in the broker log
grep -i "no space\|IOException\|store.*error" /opt/activemq/data/activemq.log | tail -20

All read-only. Adjust paths to your layout; the KahaDB directory default is activemq-data or data/kahadb under the install, but packaging varies.

How to diagnose it

  1. Confirm which limit fired. Compare df output with StorePercentUsage from JMX. If df is at 100% and StorePercentUsage is at 60%, the configured storeUsage limit exceeds the physical partition and the OS filled first. If both are at 100%, you have both problems; read the store is full guide alongside this one.

  2. Identify the space consumer. The du output from check 2 names the responsible subdirectory. Journal files dominating means message accumulation. tmp_storage dominating means non-persistent overflow. Log directory dominating means log retention, a much easier fix.

  3. If journal files dominate, find what pins them. Journal files are deleted only when every message they contain is acked. Check DLQ depth first (no TTL by default, pins files forever), then queue depths and offline durable subscribers. A queue with a handful of ancient unacked messages can pin many journal files.

  4. Assess corruption exposure before restarting anything. If the broker is still running, check the broker log for journal write errors. If the broker already crashed or was killed during the full-disk window, assume the index may be inconsistent and plan recovery time into the incident. Index recovery time scales with index size and journal count; a multi-GB db.data can mean 30+ minutes of startup.

  5. Check the ext4 reserved blocks factor. ext4 reserves 5% of blocks for root by default. The broker typically runs as a non-root user, so it can hit ENOSPC while df still shows a few percent “free”. On a 1TB partition that hidden reserve is about 50GB. Confirm with tune2fs -l /dev/<device> | grep -i reserved.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Disk free on KahaDB partition (df)The actual failure trigger; independent of broker accounting>80% used; PAGE at >90% on the active broker role
StorePercentUsage (JMX)Broker-side store limit; fires before or after disk full depending on configClimbing steadily; diverging sharply from disk usage
Journal file countDirect measure of KahaDB growth and pinningCount >2x baseline or monotonic growth
DLQ depthDLQ messages have no TTL and pin journal filesAny sustained non-zero growth
TempPercentUsage (JMX)Temp store shares the partitionSustained non-zero
db.data sizeIndex growth; also determines recovery time after a crash>500MB; >1GB means slow recovery and degraded lookups
Queue depth on critical queuesThe upstream driver of store growthEnqueue rate exceeding dequeue rate sustained

Keep at least 20% of the partition free for journal rotation, cleanup, and filesystem overhead; 30% if you need to survive a worst-case consumer outage.

Fixes

Free space immediately (broker still running)

If the partition is full but the broker is alive, free space from something that is not KahaDB first: rotate or delete old broker log files, clear application logs or dumps on the same mount, and remove any non-ActiveMQ files identified in diagnosis. This buys room without touching the store.

Do not manually delete db-*.log files. Journal files are the store. Deleting them by hand is data loss and near-certain corruption.

Drain the backlog

If journal growth is from a live backlog, the clean fix is consumption: restart or scale consumers, then let KahaDB’s cleanup task reclaim journal files. Reclamation is not instant; the checkpoint/cleanup cycle runs periodically (default every 30s), so disk frees in steps after the backlog drains.

Purge or export the DLQ

If the DLQ is the pinner, export messages you need for forensics (they carry JMSDestination and failure context), then purge. Purging is destructive: purged messages are gone, so if your business requires reprocessing, export first. After purging, set a TTL on DLQ messages so this cannot silently recur.

Expand the partition

If the workload legitimately needs more store, grow the filesystem or move KahaDB to a larger volume. When you do, reconcile the configured storeUsage limit with physical capacity: the limit should sit below partition size with at least 20% headroom, so the broker’s own accounting trips before the OS does.

Reclaim the ext4 reserve (with care)

tune2fs -m reduces the reserved block percentage. On large dedicated data partitions, lowering it from 5% to 1% recovers meaningful space. It is a live-safe operation on mounted ext4, but treat it as a one-time capacity correction, not a substitute for headroom, and never do it on the root filesystem.

Recover from actual corruption

If the broker crashed during the full-disk window and fails to start, the recovery path is a forced index rebuild: stop the broker, delete db.data and db.redo (never the journal files), and restart. KahaDB rebuilds the index by replaying all journal files, which takes time proportional to store size. Two cautions: startup options such as ignoreMissingJournalfiles and checkForCorruptJournalFiles can get a broker past corrupt journal entries but may lose messages, and if cleanup already deleted journal files that held state for inactive durable subscribers, those subscriptions may not survive the rebuild (the surviving subscription state depends on what the cleanup task already removed from the store). Test the procedure in staging before you need it in production.

Prevention

  • Monitor disk free independently of StorePercentUsage. The two limits are decoupled; alert on both. Page at >90% disk used on the active broker role only, ticket at >80%.
  • Size storeUsage below physical capacity. Leave at least 20-30% of the partition for rotation, cleanup, and the filesystem itself. If storeUsage exceeds the partition, you have configured this incident to happen.
  • Give KahaDB a dedicated partition or volume. Sharing with broker logs, temp store, and OS files turns unrelated growth into a store-corruption event.
  • Bound the DLQ. Per-destination DLQs plus a TTL, plus an alert on any non-zero depth. Unbounded DLQ is the most common root cause of slow store exhaustion.
  • Enable KahaDB integrity checks proactively. checksumJournalFiles (default true since 5.9.0, false before) and checkForCorruptJournalFiles (default false) let the broker detect, rather than silently propagate, journal damage.
  • Account for the ext4 reserve in capacity math. Either lower it on the dedicated data volume or subtract it from usable capacity in your alerting thresholds.
  • On 5.12+, consider schedulePeriodForDiskUsageCheck. The broker can periodically recheck actual disk space and shrink store/temp limits when other processes consume the disk. It mitigates the “external consumer fills the partition” case; it does not fix a store limit that was oversized from the start. The attribute is on the broker element (added in 5.12.0, default 0 = disabled).
  • Alert on journal file count growth. It is the earliest leading indicator, days before disk full. Pair it with DLQ depth and store usage trend for the full picture of the store exhaustion spiral.

How Netdata helps

  • Per-mount disk usage at per-second granularity, so you see the KahaDB partition filling as a trend, not a surprise, with alerting thresholds you can set below the ext4 reserve boundary.
  • Correlation between OS disk metrics and broker JMX metrics (StorePercentUsage, TempPercentUsage, queue depth, DLQ depth) on one dashboard: the comparison step in this guide, done continuously instead of at 3 a.m.
  • Journal file count and directory sizes via file/directory collection, catching the pinning pattern (files accumulating while queues look drained) before disk full.
  • Disk I/O latency on the store device alongside enqueue rate, distinguishing “disk full” from “disk slow” when persistent throughput degrades.
  • Anomaly detection on disk growth rate, which flags the slow exhaustion spiral days ahead of the cliff even when absolute usage is still below static thresholds.