The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-pulsar / apache-pulsar-backlog-age-vs-size ▌

Operations Guides

Apache Pulsar backlog age vs size: the latency depth alone cannot show

Backlog size tells you how much data is waiting. Backlog age tells you how long it has been waiting. Monitoring only one leaves you blind to entire classes of consumer health failures.

A stable backlog of 1,000 entries is unremarkable if those entries are 5 seconds old. The same 1,000 entries from 5 hours ago means a cursor is stuck, a consumer is down, or an application has silently stopped acknowledging. Size alone cannot distinguish these scenarios. Age can.

What it is and why it matters

Backlog in Pulsar is the gap between the latest published message position and the furthest-behind subscription cursor’s mark-delete position. Every persistent subscription has its own backlog. A topic with five subscriptions has five independent backlogs.

Two metrics measure this gap from different angles:

  • Backlog size (pulsar_subscription_back_log, entries): How many entries sit between the current write position and the cursor’s mark-delete position. This is a volume metric.
  • Backlog age (pulsar_storage_backlog_age_seconds, seconds): How old the oldest unacknowledged message is. This is a latency metric.

Size tells you magnitude. Age tells you severity.

A backlog of 50,000 entries that is 2 seconds old is healthy throughput lag. The same 50,000 entries at 3 hours old is a stuck cursor. A backlog of 200 entries at 4 hours old is small in volume but ancient in age, typically a cursor leak or an abandoned subscription holding a position near the tip.

Patterns age reveals that size hides:

  • Cursor leaks: Abandoned subscriptions hold cursors that never advance. Backlog size may be small if the subscription was created near the tip, but age will be very old for whatever backlog does exist.
  • Abandoned subscriptions: A subscription with zero consumers but non-zero backlog. Size might be stable and small. Age grows monotonically.
  • Slow-consumer SLA breaches: A consumer processing at 95% of publish rate shows stable backlog size but continuously growing age. Size looks fine. Age tells the truth.

How it works

Pulsar’s managed ledger model creates one cursor per subscription. The cursor tracks a mark-delete position that advances as messages are acknowledged. The backlog is the set of entries between the mark-delete position and the current write position.

flowchart LR
    A[Producer writes entry] --> B[Backlog: not yet acknowledged]
    B --> C[Dispatched to consumer]
    C --> D{Consumer acks?}
    D -->|Yes| E[Cursor mark-delete advances]
    D -->|No / timeout| F[Redelivery loop]
    B -.->|Size metric| G["pulsar_subscription_back_log
(entries, not messages)"] B -.->|Age metric| H["pulsar_storage_backlog_age_seconds
(seconds, oldest unacked)"] C -.->|In-flight metric| I["pulsar_subscription_unacked_messages
(dispatched, not acked)"] F -.->|Retry metric| J["pulsar_subscription_msg_rate_redeliver"]

Backlog size: estimated, not exact

pulsar_subscription_back_log counts entries, not individual messages. If a producer batches 1,000 messages into a single entry, that batch counts as 1 entry in the backlog. A backlog of 1,000 entries might represent 1,000 messages or 1,000,000 messages depending on batch size. For message-level granularity, the admin API’s analyzeSubscriptionBacklog endpoint (PIP-187) inspects individual messages rather than entries.

For byte-level backlog at the subscription level, there is no dedicated pulsar_subscription_back_log_size Prometheus metric. The exposeSubscriptionBacklogSizeInPrometheus setting (default false) controls whether the broker computes each subscription’s backlog size in bytes when building stats; that size is surfaced per subscription through the admin API (pulsar-admin topics stats -sbs / --get-subscription-backlog-size), while Prometheus continues to expose only the entry counts (pulsar_subscription_back_log) and the topic-level byte metric pulsar_storage_backlog_size. Without it, you only have entry counts.

At the topic level, pulsar_storage_backlog_size reports backlog size in bytes, but it aggregates across all subscriptions on the topic. A single slow subscription can hide behind fast ones in the aggregate.

Backlog size is also estimated, not exact. It is calculated by summarizing ledger sizes from the current active ledger up to the ledger containing the oldest unacknowledged message. For partitioned topics, updates are asynchronous.

Backlog age: cached and update-gated

pulsar_storage_backlog_age_seconds was introduced in PIP-323, first available in Pulsar 3.2.0. It reports the age of the oldest unacknowledged message at the topic level. Two admin API fields accompany it in topic stats: oldestBacklogMessageAgeSeconds and oldestBacklogMessageSubscriptionName, which identify the age and the responsible subscription.

The metric is cached. It updates only every backlogQuotaCheckIntervalInSeconds (default 60 seconds). Between checks, the value reflects the last check, not the current state. A 60-second staleness window is built in.

A related subscription-level field, earliestMsgPublishTimeInBacklogs (in milliseconds), persists across broker restarts. The topic-level age metric does not persist and resets on broker unload or restart.

For per-subscription age detail, inspect earliestMsgPublishTimeInBacklogs in each subscription’s stats via the admin API. The Prometheus metric is topic-level only and cannot provide per-subscription granularity.

Version bugs that break the age metric

Two bugs in the PIP-323 implementation affected pulsar_storage_backlog_age_seconds:

  1. The metric only updated when a message_age backlog quota policy was set on the topic. Topics without that policy showed stale or zero age regardless of actual backlog age. Fixed in Pulsar 3.0.8, 3.3.3, and 4.0.1 (PR #23619). On affected versions, the age metric is unreliable unless you configure time-based backlog quotas on every topic you want to monitor.

  2. oldestBacklogMessageAgeSeconds kept increasing on open ledgers even when the subscription was caught up. When the cursor’s mark-delete position pointed to an open ledger, the age would not reset. Fixed in Pulsar 3.0.15, 3.3.10, 4.0.8, 4.1.2, and 4.2.0 (PR #24915).

If the age metric seems stuck or nonsensical on your deployment, check whether these bugs apply to your version.

The precise time check tradeoff

By default, preciseTimeBasedBacklogQuotaCheck=false. When false, the broker uses the ledger’s close time (in-memory) to estimate message age. When true, the broker reads the oldest message from BookKeeper to obtain its exact publish timestamp. The precise check is accurate but I/O expensive because it forces a bookie read on every check interval. For most deployments, the in-memory estimate is sufficient for alerting. Reserve the precise check for forensic investigation.

Where it shows up in production

Stable size, growing age: the silent SLA breach

A consumer processing at 95% of publish rate maintains a stable backlog size. Messages accumulate at 5% but the backlog also drains. Size plateaus. Without an age metric, the dashboard looks healthy.

Age tells a different story. The oldest message gets progressively older. If your SLA requires consumption within 60 seconds, age climbing past that threshold is the breach, regardless of what size says.

Zero backlog, consumers not processing

A near-zero pulsar_subscription_back_log does not guarantee consumers are making forward progress. Per the official metric definition it counts entries in the unacknowledged state, which includes entries already dispatched and held by a consumer; a consumer holding many unacked messages therefore does not register as backlog. Pair backlog with metrics that capture different failure modes.

pulsar_subscription_unacked_messages tracks messages dispatched to consumers but not yet acknowledged. When this count hits maxUnackedMessagesPerSubscription (default 200,000), the broker stops dispatching. No error surfaces. The consumer stays connected, the backlog stays flat, and no forward progress is made. The blockedSubscriptionOnUnackedMsgs flag in topic stats turns true when this limit is hit, which is the definitive dispatch-freeze indicator.

Pair backlog metrics with:

  • pulsar_subscription_unacked_messages to detect dispatched-but-unacked pressure
  • pulsar_subscription_msg_rate_redeliver to detect consumers that receive but cannot process (poison messages, downstream failures)

Small backlog, very old age: cursor leak

An abandoned subscription created near the topic tip has a small backlog but very old backlog age. No consumer is connected, so the cursor never advances. Every check interval, age increases. Size stays low because the subscription started late.

This pattern is invisible with size-only monitoring. Over weeks, multiple abandoned subscriptions accumulate, each holding back garbage collection. The oldestBacklogMessageSubscriptionName field in topic stats identifies which subscription is the oldest, making it easier to find the culprit.

Stable size, reasonable age, high redelivery: poison message

A consumer stuck on a single message that always fails processing shows a small backlog (possibly 1 entry), a reasonable age, but high redelivery rate. The message is dispatched, fails, gets nack’d or times out, and is redelivered. No forward progress, but the backlog metrics look fine.

This is why pulsar_subscription_msg_rate_redeliver must be paired with backlog metrics. Redelivery rate above approximately 10% of dispatch rate warrants investigation. At 100% redelivery, zero forward progress is being made.

What backlog age alone cannot show

Age tells you the oldest unacked message is old. It does not tell you why. The same age reading could mean:

  • A single abandoned subscription on a topic with otherwise healthy consumers
  • A poison message blocking one consumer in a shared subscription
  • A consumer that crashed hours ago with no auto-restart
  • A downstream dependency (database, API) timing out on every message

Age also does not distinguish between many slightly-old messages and one very-old message. A backlog of 10,000 entries at 5 minutes average age might contain one entry that is 4 hours old, buried behind 9,999 entries that are seconds old. The topic-level metric reports the oldest, which is useful for detecting the worst case, but it does not give you the distribution.

For per-subscription detail, the admin API is required. pulsar_storage_backlog_age_seconds is topic-level only. It is not exposed when exposeTopicLevelMetricsInPrometheus=false because it cannot be meaningfully aggregated to namespace level.

The full signal set

SignalWhat it measuresWhat it misses
pulsar_subscription_back_log (entries)Cursor position gap per subscriptionWhether messages were dispatched; batched vs individual
pulsar_subscription_back_log_no_delayedBacklog excluding delayed messagesDelayed message backlog hidden
pulsar_storage_backlog_age_seconds (seconds)Age of oldest unacked message (topic-level)Per-subscription detail; updates every 60s
pulsar_subscription_unacked_messagesDispatched but unacked countWhether processing is succeeding
pulsar_subscription_msg_rate_redeliverMessages redelivered after failureWhy processing failed
pulsar_subscription_msg_rate_expiredTTL deleting unacked messagesWhether TTL is correctly configured
oldestBacklogMessageSubscriptionNameWhich subscription is oldestWhether it is abandoned or just slow

No single metric covers all failure modes. Size plus age plus unacked plus redelivery gives you the full picture.

Tradeoffs and when to use each

Backlog size alone is sufficient for tailing workloads where consumers must keep up with the tip. If the backlog is zero or near-zero and stable, the consumer is keeping up. Growth in size is the primary detection signal.

Backlog age is essential when you have SLA requirements on message processing latency, subscriptions that might be abandoned (dynamic subscription names, ephemeral microservice deployments), slow consumers that process most messages but occasionally stall, or a need to distinguish “caught up with a small lag” from “stuck at the tip.”

Both are insufficient for determining whether consumers are actually processing successfully. Both measure the broker’s view of cursor positions, not consumer application health. pulsar_subscription_unacked_messages and pulsar_subscription_msg_rate_redeliver are required for that layer of visibility.

Retention and backlog quota interaction. Retention settings must exceed backlog quota settings. If retention is configured too aggressively relative to backlog quota, operators hit “Please increase retention quota and retry” errors. Additionally, consumer_backlog_eviction quota policy silently drops the oldest unacknowledged messages when backlog exceeds the quota, which is data loss presented as capacity management. If this policy is active, backlog size will appear stable while messages are being silently deleted. Check pulsar_subscription_msg_rate_expired to detect this.

How Netdata helps

  • Per-second collection of pulsar_subscription_back_log captures backlog growth rates at finer granularity than the 60-second cached age metric.
  • Correlating pulsar_subscription_back_log with pulsar_subscription_unacked_messages on a single dashboard surfaces the “zero backlog but not processing” pattern.
  • Anomaly detection on backlog age flags monotonic growth on stable-size backlogs, indicating cursor leaks or abandoned subscriptions.
  • Pairing pulsar_subscription_msg_rate_redeliver with backlog metrics provides the “received but not processed” signal that neither size nor age surfaces.
  • Per-broker collection across the full cluster provides per-subscription backlog trends without manual aggregation across broker boundaries.