The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-pulsar / apache-pulsar-geo-replication-backlog ▌

Operations Guides

Apache Pulsar geo-replication backlog: replication lag and your real RPO window

Geo-replication in Pulsar is asynchronous by design. Messages are persisted locally and acknowledged to producers before they reach the remote cluster. Every message sitting in the replication backlog is a message that would be lost if you failed over right now. The replication backlog is your recovery point objective (RPO) window, measured in real time.

The common operational mistake is checking whether replication is “working” (connected, producing to the remote cluster) without checking how far behind it is. A replicator can be fully connected and actively shipping messages while being hours behind. The dashboard shows green. The failover plan assumes near-zero data loss. The reality is data loss measured in hours.

What it is and why it matters

Pulsar geo-replication uses internal replication producers per topic per remote cluster. Each replicated topic maintains a replicator cursor that tracks the local position up to which messages have been successfully sent to the remote cluster. The gap between the topic’s current write position and the replicator cursor position is the replication backlog.

Two metrics capture this gap (names confirmed in current Pulsar broker metrics; label sets vary by version):

  • pulsar_replication_backlog: message count pending replication to each remote cluster. Labels include cluster, namespace, topic, and remoteCluster.
  • pulsar_replication_delay_in_seconds: time elapsed since the oldest unreplicated message was published. This is the direct measure of replication lag in time units.

Message count tells you volume, not urgency. A backlog of 50,000 messages at 10,000 messages/second is 5 seconds of RPO exposure. The same 50,000 messages at 1 message/second is nearly 14 hours. Always convert backlog to time using pulsar_replication_delay_in_seconds, or estimate it from the backlog size and replication rate.

The lag is not a bug or a degradation signal by itself. It is an inherent property of asynchronous replication. Your job as an operator is to know whether the current lag is within your declared RPO, not whether it is zero.

How it works

The replication data path

When a producer publishes a message to a replicated topic, the local broker:

  1. Appends the message to the local managed ledger.
  2. Writes to the local BookKeeper ensemble and waits for ack quorum (Qa) acknowledgments.
  3. Acknowledges the producer. The message is now durably persisted locally.
  4. The replicator reads the message from the managed ledger and sends it to the remote cluster via an internal producer over the WAN.
  5. The remote broker receives the message, persists it to its own BookKeeper ensemble, and acknowledges.
  6. The replicator cursor advances only after the remote broker acknowledges.

Steps 4 through 6 happen asynchronously. The producer does not wait for them. This is the window of potential data loss during failover.

flowchart LR
    P[Producer] -->|1 publish| LB[Local Broker]
    LB -->|2 write + Qa ack| LDB[Local Bookies]
    LB -->|3 producer ack| P
    LB -->|4 async read| REP[Replicator]
    REP -->|5 WAN produce| RB[Remote Broker]
    RB -->|6 write + Qa ack| RDB[Remote Bookies]
    RB -->|ack| REP
    REP -->|cursor advance| LB

The replicator cursor advances only on step 6. Until then, every message between the cursor and the write head is in the replication backlog and is at risk during failover.

What limits replication throughput

Cross-region WAN bandwidth and latency are the fundamental constraints. TCP windowing means throughput is bounded by window size divided by round-trip time. High inter-region latency directly caps the maximum replication rate regardless of available bandwidth. If local publish rate exceeds this ceiling, backlog grows continuously.

When pulsar_replication_connected_count > 0 and backlog is growing continuously, the replicator is hitting this cross-region throughput limit. The connection is healthy, the producer is active, but the pipe is too narrow. This is a capacity problem, not a failure.

Where it shows up in production

Normal: transient backlog during spikes

Non-zero replication backlog is expected during traffic bursts. If the local publish rate temporarily exceeds replication throughput, backlog accumulates. The signal that this is benign: the backlog recovers to near-zero after the spike passes, and pulsar_replication_delay_in_seconds returns to baseline.

Problem: continuous growth with connected replicator

When backlog grows continuously for more than 30 minutes and pulsar_replication_connected_count > 0, the replicator is connected but cannot keep up. This is the most commonly misdiagnosed pattern. Teams see “connected” and assume replication is healthy, missing that the RPO window is growing unbounded.

Check the replication rate metrics. If pulsar_replication_rate_out and pulsar_replication_throughput_out are flat while backlog grows, the replicator has hit the WAN throughput ceiling. Remediation options are limited: increase cross-region bandwidth, reduce publish rate, or add more replication producers by partitioning.

Problem: silently stopped replication

This pattern is tracked in Pulsar issue #25097 and was fixed in Pulsar 4.0.12 / 4.2.3.

A more insidious pattern: replication for some topics randomly stops during normal operation, often triggered by bursts of high publish rate or BookKeeper iowait. The replicator shows connected: true with msgRateOut: 0.0 and a growing backlog. No error messages appear in logs.

If you see this signature (connected replicator, zero output rate, growing backlog on specific topics), check per-topic replication stats from pulsar-admin topics stats, since aggregate metrics can mask individual topic stalls.

The reported recovery, confirmed by operators in issue #25097, is disabling and re-enabling replication for the affected namespace. This is disruptive: it interrupts replication for all topics in the namespace. Apply it during a maintenance window or only when the RPO exposure from continued stall exceeds the disruption cost.

Problem: backlog quota stalling replication

Geo-replication producers on the remote cluster are subject to the same backlog quotas as regular producers. If the remote cluster has a producer_request_hold or producer_exception backlog quota policy and the remote-side subscription falls behind, the replication producer can be blocked. The replicator cursor may enter a “NoLedger” state with active: false. Backlog grows on the local side even though the network and the remote cluster are otherwise healthy.

Check the remote cluster’s backlog quota configuration with pulsar-admin namespaces get-backlog-quotas <namespace> if you suspect this.

Partial replication hides aggregate problems

Aggregate replication metrics can hide topic-level failures. If 99 of 100 topics are replicating fine and one is stuck, the aggregate backlog may look stable. Always check per-topic replication stats for critical topics, not just namespace or cluster aggregates.

Common misuses and risks

Treating “replication working” as “DR-ready”

Teams verify that replication is connected and producing, then assume they can fail over with near-zero data loss. Without checking pulsar_replication_delay_in_seconds against the declared RPO, the actual data loss on failover could be hours.

Your failover runbook must include a pre-failover gate: check pulsar_replication_delay_in_seconds for every remote cluster. If the delay exceeds the declared RPO, you either wait for replication to catch up or accept the data loss explicitly.

Modifying cluster configuration triggers cascading deletions

Modifying the clusters list at the namespace or topic policy level can trigger automatic topic deletions on excluded clusters. This is a data-loss risk. Always verify the impact of cluster configuration changes before applying them, and maintain independent backups.

Replicated subscriptions have limitations

Replicated subscriptions do not replicate individual acknowledgments. Only the mark-delete position (baseline cursor) is replicated via periodic snapshots (default every 1000ms, replicatedSubscriptionsSnapshotFrequencyMillis). Messages acknowledged out of order may be redelivered after failover. For deployments with more than 2 clusters, the snapshot timeout (default 30 seconds, replicatedSubscriptionsSnapshotTimeoutSeconds) may need to be increased to 60 seconds to avoid replication stalls.

Signals to watch in production

SignalWhy it mattersWarning sign
pulsar_replication_delay_in_secondsDirect measure of RPO exposure in time units. This is your failover data loss window.Sustained value exceeding declared RPO.
pulsar_replication_backlogMessage count pending replication. Must be converted to time to assess urgency.Continuous growth rather than recovery after spikes.
pulsar_replication_connected_countWhether the replicator to each remote cluster is connected. Value > 0 means connected.Drops to 0 with growing backlog indicates a disconnected replicator.
pulsar_replication_disconnected_countCount of disconnected replication sessions.Any non-zero value indicates connection problems to a remote cluster.
pulsar_replication_rate_outMessages per second being replicated to the remote cluster.Flat or declining while backlog grows means throughput ceiling hit.
pulsar_replication_throughput_outBytes per second being replicated. More useful than message rate for bandwidth assessment.Saturating near the WAN bandwidth limit.
Local pulsar_rate_inLocal publish rate. Compare against replication rate to predict backlog trajectory.Publish rate consistently exceeding replication rate_out.
pulsar_replication_rate_expiredRate of messages TTL-evicted during replication before reaching the remote cluster.Non-zero means messages are being deleted before replication completes.

How Netdata helps

  • Per-second collection of pulsar_replication_backlog and pulsar_replication_delay_in_seconds with the remoteCluster label lets you see the RPO window for each remote cluster individually, not just in aggregate.
  • Correlating pulsar_replication_delay_in_seconds against local pulsar_rate_in in the same dashboard reveals whether growing lag is caused by a publish spike (transient) or a structural throughput mismatch (continuous growth).
  • ML anomaly detection on pulsar_replication_backlog distinguishes normal burst-driven accumulation from unusual sustained growth that may indicate a stuck replicator or WAN degradation.
  • Alerting on pulsar_replication_delay_in_seconds with per-remoteCluster granularity lets you set RPO-based thresholds per DR target rather than relying on generic message-count thresholds.
  • Correlating replication metrics with cross-region network bandwidth and remote broker health in a single view shortens root-cause isolation.