The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-pulsar / apache-pulsar-subscription-cursor-leak ▌

Operations Guides

Apache Pulsar abandoned subscriptions: cursor leaks that pin storage forever

Bookie disk usage climbs steadily across the cluster, but publish rates are flat and there is no traffic spike. Per-subscription backlogs look fine for the subscriptions you know about. When you try to delete an old topic to reclaim space, the operation fails with a message about active subscriptions. The topic has not had a consumer in weeks.

This is the signature of a silent subscription cursor leak. Applications create dynamically-named durable subscriptions and then disconnect without unsubscribing. Each abandoned subscription leaves behind a cursor pinned at its last acknowledged position. Pulsar cannot delete any message after that cursor’s position until the subscription acknowledges past it, which never happens because no consumer is connected. Storage grows monotonically, and the growth is invisible in per-subscription backlog dashboards because those dashboards only track subscriptions the team knows about.

By the time a bookie approaches its read-only threshold, dozens or hundreds of abandoned cursors may be pinning terabytes of data that cannot be garbage collected.

What this means

A Pulsar subscription is a durable cursor into a topic’s managed ledger. When a consumer creates a subscription, the broker records a mark-delete position and begins tracking acknowledged messages. Even after the consumer disconnects, the cursor persists. Pulsar’s design guarantee is that messages are retained until all subscriptions have acknowledged them.

This guarantee becomes a liability when subscriptions are abandoned. The managed ledger’s garbage collection cannot reclaim ledger entries that sit after the oldest cursor position. If a cursor was last advanced three months ago, every message published since then is pinned in storage, regardless of whether every other subscription on the topic has long since acknowledged those same messages.

The cascade looks like this:

flowchart TD
    A[App creates dynamic subscription] --> B[Consumer disconnects without unsubscribe]
    B --> C[Cursor frozen at old mark-delete position]
    C --> D[Managed ledger GC cannot reclaim entries after cursor]
    D --> E[Storage grows despite flat traffic]
    E --> F[Bookie disk usage climbs toward threshold]
    F --> G[Bookie hits diskUsageThreshold and goes read-only]
    G --> H[New writes fail if enough bookies are read-only]

The storage cost scales with publish rate multiplied by time since abandonment for each abandoned cursor. A topic producing 1,000 messages per second with an abandoned cursor from 90 days ago has roughly 7.8 billion messages pinned.

Common causes

CauseWhat it looks likeFirst thing to check
Microservice per-instance subscriptionsSubscription names contain hostnames, pod names, or instance IDs; consumer count is zero after deployment rolloutTopic stats, look for subscriptions matching pod or host naming patterns
Batch processing framework leftoversFrameworks like Spark or Flink create durable subscriptions and do not clean them up on job terminationSubscription names matching job or application names with zero consumers
Test or dev subscriptions left behindSubscriptions with names like “test”, “debug”, or developer usernames; zero consumers on production topicsSubscription list on production topics for non-production naming
Failed deployment cleanupApplication creates subscriptions on startup but crash or shutdown path does not call unsubscribeCompare subscription creation timestamps against deployment events

Quick checks

Run these read-only commands to identify abandoned subscriptions and assess storage impact.

# List all topics in a namespace
curl -s http://<broker-host>:8080/admin/v2/persistent/tenant/namespace | jq -r '.[]'

# Get subscription details for a topic: consumer count and backlog
curl -s http://<broker-host>:8080/admin/v2/persistent/tenant/namespace/topic/stats \
  | jq '.subscriptions | to_entries[] | {sub: .key, consumers: (.value.consumers | length), backlog: .value.msgBacklog}'

# Find subscriptions with zero consumers across all topics in a namespace
# NOTE: topic names from the list endpoint are local names; URL-encode if needed
curl -s http://<broker-host>:8080/admin/v2/persistent/tenant/namespace | jq -r '.[]' | while read topic; do
  curl -s "http://<broker-host>:8080/admin/v2/persistent/tenant/namespace/${topic}/stats" \
    | jq -r --arg t "$topic" \
      '.subscriptions | to_entries[] | select(.value.consumers | length == 0) | "\($t)\t\(.key)\t\(.value.msgBacklog)"'
done

# Check cursor positions relative to last confirmed entry
curl -s http://<broker-host>:8080/admin/v2/persistent/tenant/namespace/topic/internalStats \
  | jq '{lastConfirmedEntry: .lastConfirmedEntry, cursors: [.cursors | to_entries[] | {name: .key, markDeletePosition: .value.markDeletePosition}]}'

# Check bookie disk usage
curl -s http://<bookie-host>:8000/metrics | grep bookie_ledger_dir

# Check per-subscription backlog from Prometheus metrics
curl -s http://<broker-host>:8080/metrics | grep pulsar_subscription_back_log

# Count subscriptions on a topic
curl -s http://<broker-host>:8080/admin/v2/persistent/tenant/namespace/topic/stats \
  | jq '.subscriptions | length'

How to diagnose it

  1. Identify topics with growing storage but flat traffic. Pull bookie disk usage trends over weeks using the bookie_ledger_dir_{path}_usage metric. Cross-reference with publish rate (pulsar_rate_in). If disk grows while publish rate is flat, something is preventing garbage collection.

  2. List all subscriptions on affected topics. Use the stats endpoint to enumerate every subscription. Look at the consumer count for each one. Subscriptions with zero connected consumers are candidates for abandonment.

  3. Verify abandonment. Before deleting anything, confirm with application teams that the subscription is truly unused. Some subscriptions may have consumers that reconnect intermittently, such as batch consumers that connect once per day. Query internal stats for markDeletePosition and compare it across checks spaced hours or days apart. A cursor whose position does not advance is abandoned.

  4. Check backlog age, not just size. A subscription with a small backlog might still be abandoned if that backlog has not advanced in weeks. Compare cursor mark-delete positions over time. The internal stats endpoint shows markDeletePosition per cursor, which tells you where each cursor sits relative to lastConfirmedEntry.

  5. Assess storage impact. For each abandoned subscription, the pinned storage is approximately the total data written to the topic since the cursor’s last advancement. Prioritize cleanup of subscriptions with the oldest cursors on the highest-traffic topics.

  6. Confirm the pattern at scale. If one topic has abandoned subscriptions, check the entire namespace or tenant. The application behavior that caused the leak likely affected multiple topics. Run the namespace-wide scan from the quick checks section to enumerate all zero-consumer subscriptions.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Subscription count per topicGrowing count without corresponding consumer growth indicates leaksCount increasing over days or weeks without deployment changes
Subscriptions with zero consumersDirect indicator of abandoned cursorsAny sustained zero-consumer subscription on an active topic
Bookie disk usage (bookie_ledger_dir_{path}_usage)Pinned storage eventually fills disksSustained growth despite flat publish rate
Subscription backlog (pulsar_subscription_back_log)Abandoned subscriptions may show large or frozen backlogsNon-zero backlog on a subscription with zero consumers
bookie_SERVER_STATUSTerminal state if disk fills to thresholdValue drops to 0, meaning read-only
Cursor mark-delete position (internal stats)Shows exactly where each cursor is frozen in the ledgerPosition unchanged over days or weeks
Topic deletion failuresAttempting cleanup surfaces the problemError: “Topic has active subscriptions or producers”

Fixes

Unsubscribe confirmed abandoned subscriptions

Once you have confirmed a subscription is truly abandoned, remove it:

# Unsubscribe a single abandoned subscription
pulsar-admin topics unsubscribe \
  --subscription <subscription-name> \
  persistent://tenant/namespace/topic

The cursor is removed, and the managed ledger can immediately begin garbage collecting entries that were pinned by that cursor, subject to retention policy and other cursors on the same topic.

After unsubscribing, monitor bookie disk usage. BookKeeper’s entry log garbage collection runs asynchronously, so reclaimed space may not appear immediately. The bookie_ACTIVE_ENTRY_LOG_SPACE_BYTES metric should begin trending down as GC reclaims entry logs that contained the pinned data.

Batch unsubscribe for multiple abandoned subscriptions

There is no built-in batch unsubscribe command. Script the operation carefully:

# WARNING: Destructive. Verify each subscription is truly abandoned before running.
# Review the echo output first by removing the pulsar-admin line.
curl -s http://<broker-host>:8080/admin/v2/persistent/tenant/namespace | jq -r '.[]' | while read topic; do
  for sub in $(curl -s "http://<broker-host>:8080/admin/v2/persistent/tenant/namespace/${topic}/stats" \
    | jq -r '.subscriptions | to_entries[] | select(.value.consumers | length == 0) | .key'); do
    echo "Unsubscribing: persistent://tenant/namespace/${topic} / ${sub}"
    pulsar-admin topics unsubscribe \
      --subscription "$sub" \
      "persistent://tenant/namespace/${topic}"
  done
done

Always verify the subscription list with application owners before running this against production.

Truncate as a last resort

If a topic has so many abandoned subscriptions that manual cleanup is impractical, pulsar-admin topics truncate moves all cursors to the end of the topic and deletes all messages up to that point. This is destructive: it acknowledges all existing messages for every subscription, including active ones. Use this only when you have confirmed no active consumer needs the historical data.

# DESTRUCTIVE: Moves all cursors to end, deletes all existing messages
pulsar-admin topics truncate persistent://tenant/namespace/topic

Prevention

Set subscriptionExpirationTimeMinutes at the namespace level. This tells the broker to automatically delete subscriptions that have had no connected consumers for the configured duration. Set it per namespace rather than relying on the global broker.conf default, since different applications have different lifecycle needs.

# Set subscription expiration to 7 days (10080 minutes) for a namespace
pulsar-admin namespaces set-subscription-expiration-time \
  -t 10080 \
  tenant/namespace

The default value is 0, meaning inactive subscriptions are never deleted automatically. The broker checks for expired subscriptions periodically based on subscriptionExpiryCheckIntervalInMinutes, which defaults to 5 minutes.

Caveat: expiration may not work on completely inactive topics. If a topic has no connected producers and no consumers on any subscription, the broker may not evaluate subscription expiration for it. This is a known limitation. For topics that are purely accumulating data with no active traffic, you may need to manually unsubscribe abandoned subscriptions.

Subscription expiry is evaluated only for topics currently loaded on the broker. Since Pulsar 3.0.0 (PR #17573) the cursor’s last-active timestamp is not persisted across topic unload/reload, so on frequently-rebalanced topics the idle clock resets and expiry may never fire (PR #25244, still open). This matters even when the previous caveat’s conditions are not met: a topic with active producers but an idle subscription can still evade expiry.

Caveat for versions before Pulsar 2.8.1 and 2.9.0. Before PR #11253 (fixed in 2.8.1 and 2.9.0), the namespace-level field was a plain integer: setting -t 0 did not disable namespace expiry (0 silently fell back to the broker-level subscriptionExpirationTimeMinutes). The fix made the field nullable (0 = disabled, null = inherit the broker setting). After upgrading, existing namespaces may still carry a stored 0; clear it with pulsar-admin namespaces remove-subscription-expiration-time tenant/namespace, then set the desired value.

Ensure applications call unsubscribe on shutdown. The root cause is almost always application-side. Microservices that create per-instance subscriptions must unsubscribe during graceful shutdown. Batch processing frameworks that create durable subscriptions must clean them up when jobs terminate. Review the application lifecycle code for the topics where you found leaks.

Monitor subscription count as a time series. Track the number of subscriptions per topic over weeks, not just the current value. Any monotonically increasing trend without corresponding consumer growth is a leak. Alert on sustained subscription count growth.

Track backlog age, not just backlog size. A subscription with a backlog of 100 messages that has not advanced in 30 days is a leak. A subscription with a backlog of 100,000 messages that is actively draining is healthy. Age reveals abandonment; size alone does not. Compare cursor mark-delete positions to the managed ledger’s lastConfirmedEntry to compute how far behind each cursor is.

How Netdata helps

Netdata’s per-second metrics collection surfaces several signals relevant to cursor leak diagnosis:

  • Bookie disk usage trends at per-second resolution make slow monotonic growth from pinned storage visible long before disk thresholds trigger. Anomaly detection flags unusual growth patterns that deviate from the established baseline.
  • Subscription backlog metrics (pulsar_subscription_back_log) collected per-topic and per-subscription let you correlate specific subscriptions with storage growth, even when those subscriptions are not in your actively monitored set.
  • Bookie server status (bookie_SERVER_STATUS) transitions to read-only are immediately visible, giving you the terminal-state alert if cursor leaks have already pushed a bookie past its disk threshold.
  • Publish and dispatch rates alongside disk usage trends let you confirm the “flat traffic but growing storage” signature that distinguishes cursor leaks from legitimate backlog accumulation.
  • Active connection counts per broker, correlated with subscription data from the stats API, help identify subscriptions that have zero connected consumers.