The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-pulsar / apache-pulsar-unacked-messages-dispatch-freeze ▌

Operations Guides

Apache Pulsar unacked messages at the limit: the silent dispatch freeze

Consumers are connected. Producers are publishing. Backlog looks flat or zero. But no messages are being processed.

Pulsar brokers enforce a per-subscription limit on unacknowledged messages (maxUnackedMessagesPerSubscription, default 200,000) and a per-consumer limit (maxUnackedMessagesPerConsumer, default 50,000). When unacked messages hit either ceiling, the broker stops dispatching new messages to that subscription or consumer. No error is returned to the client. No exception is thrown. Dispatch pauses silently.

Most monitoring focuses on backlog and connection count. Both can look healthy during a freeze. Messages were already dispatched to consumers, so backlog may read zero. Consumers remain connected at the TCP level. The definitive indicators are the unacked message count itself and the blockedSubscriptionOnUnackedMsgs flag in topic stats.

Mechanism

When a consumer receives a message, the broker increments the unacked counter for that subscription. When the consumer acknowledges, the counter decrements. If consumers process slower than the dispatch rate, or stop acknowledging entirely, the counter climbs. At the limit, the broker sets blockedSubscriptionOnUnackedMsgs to true and halts further dispatch.

The freeze persists until the unacked count drops below the limit. This happens when consumers acknowledge messages, when ack timeouts trigger redelivery (if configured), or when messages are otherwise cleared from the unacked set. If consumers are genuinely stuck and no ack timeout is configured, the freeze persists indefinitely.

flowchart TD
    A[Consumer connected, receiving messages] --> B[Consumer stops or slows acknowledging]
    B --> C[Unacked count rises toward limit]
    C --> D{Unacked at limit?}
    D -->|No| C
    D -->|Yes| E[Broker sets blockedSubscriptionOnUnackedMsgs = true]
    E --> F[Dispatch stops for this subscription]
    F --> G[Consumers stay connected, no errors]
    G --> H[Backlog may read stable or zero]
    H --> I[Silent freeze: no forward progress]

There is also a per-consumer limit. When maxUnackedMessagesPerConsumer (default 50,000) is exceeded for an individual consumer, the broker sets blockedConsumerOnUnackedMsgs to true for that consumer. In a shared subscription with multiple consumers, this can affect individual consumers independently.

Version-specific gotchas

Two upstream bugs complicate diagnosis on certain Pulsar versions:

  • The blocked metric was unreliable in older 2.x builds. The pulsar_subscription_blocked_on_unacked_messages Prometheus metric could report 0 even when the subscription was actively blocked. The fix (PR #18621, merged Nov 2022) is in 2.10.x/2.11.x patch releases and 3.0.0. On earlier 2.x builds, do not rely on this metric alone; use the Admin API field blockedSubscriptionOnUnackedMsgs instead.
  • A regression caused permanent consumer freeze on some 3.0.x through 4.0.x versions. Consumers could stop receiving messages even when pulsar_subscription_unacked_messages was 0 and pulsar_subscription_blocked_on_unacked_messages was 0, because the per-consumer blockedConsumerOnUnackedMsgs flag got stuck at true after acknowledgements completed (apache/pulsar#22657; regression introduced by PR #20990). If you see freeze symptoms but the subscription-level metrics look clean, check the per-consumer blocked flag in topic stats.

There is also a long-standing issue where pulsar_subscription_unacked_messages can report negative values due to batch acknowledgement counting problems. Monitor the absolute value or cross-check with the Admin API for authoritative counts.

Common causes

CauseWhat it looks likeFirst thing to check
Consumer processing too slowUnacked count rising steadily, low ack rateConsumer application logs for processing latency
Consumer bug: never acknowledgesUnacked rising linearly, ack rate zeroConsumer code path for acknowledge() calls
Downstream dependency failureUnacked rising, consumer logs show DB or API timeoutsConsumer’s downstream service health
Poison messageHigh redelivery rate, unacked rising then plateauingConsumer error logs for repeated failures on the same message
Ack timeout misconfiguredUnacked elevated, redelivery cycling without progressConsumer ackTimeout and ackTimeoutTickDuration settings
Stuck blocked flag (bug)Freeze symptoms but metrics show 0 unacked and not blockedPer-consumer blockedConsumerOnUnackedMsgs in topic stats

Quick checks

Run these read-only checks on the affected topic and subscription. None of them modify state.

# Check unacked message count per subscription from Prometheus metrics
curl -s http://<broker-host>:8080/metrics | grep pulsar_subscription_unacked_messages

# Check whether the subscription is flagged as blocked
curl -s http://<broker-host>:8080/metrics | grep pulsar_subscription_blocked_on_unacked_messages

# Get detailed subscription stats via Admin API
pulsar-admin topics stats persistent://tenant/namespace/topic
# Look for:
#   subscriptions.<name>.unackedMessages
#   subscriptions.<name>.blockedSubscriptionOnUnackedMsgs
#   subscriptions.<name>.consumers[].blockedConsumerOnUnackedMsgs
#   subscriptions.<name>.consumers[].availablePermits

# Check dispatch rate vs publish rate for the affected topic
curl -s http://<broker-host>:8080/metrics | grep -E "pulsar_rate_(in|out)"

# Check redelivery rate for the subscription
curl -s http://<broker-host>:8080/metrics | grep pulsar_subscription_msg_rate_redeliver

# Verify consumer connections are still active
curl -s http://<broker-host>:8080/admin/v2/persistent/tenant/namespace/topic/stats | \
  jq '.subscriptions | to_entries[] | {sub: .key, consumers: (.value.consumers | length)}'

The availablePermits field in consumer stats deserves attention. A value of 0 means the client library’s internal receive queue is full and receive() is not being called. There is no documented -1 sentinel, but negative availablePermits values have been observed on stuck consumers (apache/pulsar#24418), so treat 0 or negative as stalled. This is a separate mechanism from the unacked limit, but it produces the same symptom: the dispatcher stops sending messages. If availablePermits is 0 and blockedSubscriptionOnUnackedMsgs is false, the bottleneck is client-side, not broker-side.

How to diagnose it

  1. Identify the affected subscription. Compare pulsar_rate_in and pulsar_rate_out per topic. A topic with non-zero publish rate but zero dispatch rate is your target. If multiple subscriptions exist on the topic, check each independently.

  2. Check the unacked count. Pull pulsar_subscription_unacked_messages for the subscription. Compare against your configured maxUnackedMessagesPerSubscription. If you have not changed the default, the limit is 200,000. Alerting at 50% of the limit (100,000 for the default) gives a leading indicator before the freeze triggers.

  3. Check the blocked flag. Query the Admin API for blockedSubscriptionOnUnackedMsgs on the subscription and blockedConsumerOnUnackedMsgs on each consumer. The Admin API is more authoritative than the Prometheus metric, especially on versions before 3.0.0.

  4. Check availablePermits. If the subscription is not blocked but dispatch has stopped, check availablePermits on each consumer. A value of 0 means the client is not requesting more messages. This points to a client-side issue, not a broker-side block.

  5. Check redelivery rate. High redelivery (pulsar_subscription_msg_rate_redeliver) relative to dispatch rate indicates poison messages or processing failures. If every message is being redelivered, the consumer is receiving but failing to process, which means it will never ack and the unacked count will stay elevated.

  6. Check the backlog. If pulsar_subscription_back_log is zero or stable while pulsar_rate_out is zero, messages were dispatched but not acknowledged. This is the signature of the unacked freeze. A growing backlog with zero dispatch rate means no consumers are connected at all, which is a different problem.

  7. Rule out the version-specific freeze bug. If unacked is 0, the subscription is not blocked, dispatch is zero, and consumers are connected, check blockedConsumerOnUnackedMsgs per consumer. If it is stuck at true, you may be hitting the stuck-blocked-flag regression (apache/pulsar#22657, present in 3.0.2+ until fixed). The temporary workaround is to restart the affected consumer.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
pulsar_subscription_unacked_messagesCount of dispatched but unacknowledged messages. Approaching the limit means dispatch freeze is imminent.Sustained value above 50% of maxUnackedMessagesPerSubscription
pulsar_subscription_blocked_on_unacked_messagesBinary flag indicating the subscription is dispatch-frozen.Value of 1 (blocked)
blockedSubscriptionOnUnackedMsgs (Admin API)Authoritative blocked flag, more reliable than the Prometheus metric on older versions.true
blockedConsumerOnUnackedMsgs (Admin API, per consumer)Per-consumer block flag. Can be stuck at true due to known bugs.true while unacked is 0
availablePermits (per consumer)Client-side flow control. 0 means the client is not pulling messages.Any value that does not increase over time
pulsar_subscription_msg_rate_redeliverMessages being redelivered indicate processing failures.Above 10% of dispatch rate
pulsar_rate_out vs pulsar_rate_inDispatch rate vs publish rate. Divergence means backlog or freeze.rate_out at zero with non-zero rate_in
pulsar_subscription_back_logMessages not yet dispatched. Can be zero during a freeze.Zero or stable while dispatch is stopped (confirms freeze, not consumer absence)

Fixes

Consumer is processing too slowly

The consumer receives messages but its processing pipeline cannot keep up with the dispatch rate.

  • Scale out consumers. If the subscription type is Shared or Failover, adding consumer instances increases parallelism. For Key_Shared, ensure keys are well-distributed across consumers.
  • Increase the unacked limit. Raising maxUnackedMessagesPerSubscription (broker-level config) gives consumers more headroom. This is a bandage, not a fix. If consumers cannot keep up, a higher limit just delays the freeze.
  • Tune consumer receiver queue size. A larger receiver queue means more messages buffered client-side, but it also means more memory per consumer and more unacked messages in flight.

Consumer is not acknowledging at all

The consumer receives messages but never calls acknowledge(). This is a code defect.

  • Review the consumer’s processing pipeline. Ensure acknowledge() is called on every successfully processed message, including in error paths that handle and recover from failures.
  • Check for swallowed exceptions. If the consumer catches an exception, logs it, but does not nack or ack the message, the message stays unacked indefinitely (or until ack timeout triggers redelivery).

Downstream dependency failure

The consumer is blocked waiting on a slow database, API, or external service. Messages pile up in the unacked set because processing cannot complete.

  • Fix the downstream dependency. This is the root cause.
  • Set appropriate timeouts. If the consumer is stuck on a downstream call, it should time out and either nack the message or move on. Without a timeout, one hung downstream call can freeze the entire subscription.
  • Configure ackTimeout on the consumer. This causes the broker to redeliver messages that have been unacked for too long. Without ackTimeout configured (the default is no timeout), unacked messages stay unacked indefinitely.

Poison message

A specific message always fails processing. The consumer receives it, fails, and the message gets redelivered. Redelivery rate spikes.

  • Configure a dead letter topic. After maxRedeliveryCount redeliveries, the message routes to a DLQ and stops clogging the subscription.
  • Skip the message manually. If you identify the poison message by ID, you can acknowledge it explicitly to clear it from the unacked set. This loses the message, so coordinate with the application team.
  • Fix the consumer. If the consumer should handle this message type, fix the deserialization or processing logic.

Stuck blocked flag (version bug)

On affected Pulsar versions, blockedConsumerOnUnackedMsgs can remain true after all messages are acknowledged. The consumer stops receiving permanently.

  • Restart the affected consumer. This clears the stuck flag temporarily. The bug will recur.
  • Upgrade. Upgrade to a version containing PR #23796 (merged Feb 2025) to resolve permanently. Do not rely on a specific patch number without checking your release’s changelog for that PR.

Prevention

  • Alert on unacked count at 50% of the limit. By the time you hit 100%, the freeze is already active and consumers are stuck. Alert when pulsar_subscription_unacked_messages exceeds 50% of maxUnackedMessagesPerSubscription for more than 10 minutes.
  • Configure ackTimeout on consumers. Without it, unacked messages never expire. A reasonable ackTimeout (matching your expected processing time) ensures stuck messages eventually get redelivered rather than sitting in the unacked set forever.
  • Use dead letter topics. Poison messages cause repeated redelivery, consuming broker resources and keeping unacked counts elevated. A DLQ with an appropriate maxRedeliveryCount prevents infinite retry loops.
  • Monitor the blocked flag directly. Even with unacked alerting, a separate alert on pulsar_subscription_blocked_on_unacked_messages == 1 (or the Admin API equivalent) provides a definitive “this is frozen right now” signal.
  • Track redelivery rate. High redelivery relative to dispatch rate is a leading indicator of poison messages and processing failures that will eventually push unacked counts up.
  • Upgrade past known bugs. If you are on an affected version range, the stuck-blocked-flag regression is a live risk. Plan an upgrade to a patched release.

How Netdata helps

  • Per-second resolution on pulsar_subscription_unacked_messages. The difference between a consumer slowly falling behind and one that has hit the wall is visible in the slope of the unacked count. Per-second granularity makes that slope legible before the freeze triggers.
  • Correlate unacked count with dispatch rate. When pulsar_rate_out drops to zero while pulsar_subscription_unacked_messages is at the limit, the diagnosis is immediate. Netdata’s correlated dashboards put these signals on the same timeline without manual cross-referencing.
  • ML anomaly detection on unacked trends. A slowly climbing unacked count that has not yet crossed the 50% threshold is easy to miss in threshold-based alerting. Anomaly detection flags the deviation from baseline earlier.
  • Alert at 50% of the limit. Configurable alerts on the unacked-to-limit ratio give you the leading indicator before the freeze happens, not after.
  • Track the blocked flag as a dedicated signal. The pulsar_subscription_blocked_on_unacked_messages metric is surfaced directly as a binary state change rather than requiring inference from rate drops.