The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-pulsar / apache-pulsar-zookeeper-watch-explosion ▌

Operations Guides

Apache Pulsar ZooKeeper watch explosion: thousands of watches killing metadata latency

Topic lookups slow down. Bundle ownership transfers stall. Broker sessions flicker between connected and disconnected. ZooKeeper reports thousands, sometimes tens of thousands, of registered watches. This is a watch explosion, and the root cause is rarely in ZooKeeper itself. It is in client connection churn.

Watches accumulate when consumers disconnect and reconnect in waves. Each reconnection registers metadata watches for topic ownership, subscription state, and policy changes. During a mass reconnection event (broker restart, network blip, load balancer cycle), thousands of ephemeral znodes are created and deleted in rapid succession. Each creation and deletion triggers watch registration and notification. ZK processes requests sequentially, so as notification load grows, request latency climbs. Once latency exceeds session timeout, brokers lose sessions, triggering bundle unloads and another round of client reconnections. The feedback loop tightens until the cluster thrashes.

What this means

A ZK watch is a one-time trigger registered against a znode. When the znode changes, ZK notifies the watching client, which must re-register for continued notifications. In Pulsar, brokers, bookies, and clients maintain watches on metadata paths: bundle ownership, topic policies, schema registry, cluster configuration.

A watch explosion occurs when the rate of znode creation, deletion, and modification overwhelms ZK’s request processing pipeline. The ensemble spends more time dispatching watch notifications than servicing metadata operations. Request latency rises because of ZK’s sequential processing model: as the outstanding request queue grows, each new request waits longer, causing client timeouts, which trigger reconnections, which generate more watch registrations.

The degradation curve is not linear. ZK latency stays flat until a tipping point, then spikes hard. By the time dashboards show elevated latency, the cascade has often already begun.

flowchart TD
    A[Consumer disconnect storm] --> B[Mass ephemeral znode churn]
    B --> C[Watch registrations spike]
    C --> D[ZK request latency rises]
    D --> E{Exceeds session timeout?}
    E -->|No| F[Metadata ops slow
lookups fail] E -->|Yes| G[Broker sessions expire] G --> H[Bundles unloaded
to other brokers] H --> I[Clients reconnect
to new owners] I --> C

Common causes

CauseWhat it looks likeFirst thing to check
Consumer disconnect/reconnect stormWatch count spikes alongside pulsar_active_connections fluctuations and bundle unload eventsecho wchs | nc <zk-host> 2181 for current count
Excessive topic countHigh baseline watch count proportional to topic count, slow degradation over time as topics accumulatepulsar_topics_count per broker vs cluster average
Load balancer thrashingBundle unload rate (pulsar_lb_unload_bundle_total) sustained above baseline; each unload triggers client reconnectionsBroker logs for rapid ownership cycling
ZK transaction log disk I/O bottleneckZK latency spikes correlate with disk I/O on the ZK transaction log volumeiostat -x 1 on the ZK transaction log disk
ZK ensemble member failureQuorum degraded, latency rises as remaining members absorb loadecho stat | nc <zk-host> 2181 on each node

Quick checks

All read-only and safe to run during an active incident. Note: ZK 4-letter words require 4lw.commands.whitelist to include the relevant commands in zoo.cfg.

# Total watch count on each ZK node
echo wchs | nc <zk-host> 2181

# Watches by session (shows which sessions hold the most watches)
# WARNING: can be expensive on large ensembles; use judiciously during incidents
echo wchc | nc <zk-host> 2181

# Detailed ZK monitoring output: watch count, latency, node count, followers
echo mntr | nc <zk-host> 2181

# ZK server stats: latency, connections, outstanding requests
echo stat | nc <zk-host> 2181

# Broker metadata-store session events (Pulsar exposes no zookeeper_connected Prometheus metric; see the log grep below)

# Bundle unload rate (reconnection source)
curl -s http://<broker-host>:8080/metrics | grep pulsar_lb_unload_bundle_total

# Active connections for reconnection storm signature
curl -s http://<broker-host>:8080/metrics | grep pulsar_active_connections

# Lookup failures (early failure indicator)
curl -s http://<broker-host>:8080/metrics | grep pulsar_broker_lookup

# Broker logs for session expiry events
grep -i "Session expired\|Connection loss" /var/log/pulsar/broker.log | tail -20

How to diagnose it

  1. Confirm the watch count is abnormal. Run echo wchs | nc <zk-host> 2181 on each ZK ensemble member. Compare against your established baseline. Thousands of watches on a cluster that should have hundreds is the signature. Sustained ZK request latency above 10ms is an early warning; above 100ms means failure is likely imminent. (These thresholds come from operational incident reports; Pulsar’s own documentation does not define them.)

  2. Identify the source of watches. Run echo wchc | nc <zk-host> 2181 to see which sessions hold the most watches. If a single broker session holds a disproportionate share, that broker is the source. If watches are distributed across many sessions, the issue is broad client churn.

  3. Correlate with connection churn. Check pulsar_active_connections across brokers. Fluctuating counts (drops followed by spikes) indicate a reconnection storm. If available, check pulsar_connection_created_total_count and pulsar_connection_closed_total_count; the rate of change between them indicates churn velocity.

  4. Correlate with bundle unload rate. Check pulsar_lb_unload_bundle_total. Each bundle unload drops client connections for affected topics, causing reconnections that generate new watch registrations. A sustained unload rate above 1 per minute outside maintenance indicates load balancer thrashing.

  5. Check ZK transaction log disk. ZK writes every transaction to its log disk synchronously before acknowledging. If that disk is slow or shared with another workload, request latency rises independently of watch count. Run iostat -x 1 on the ZK node and examine %util and await on the transaction log device.

  6. Verify ensemble quorum health. Run echo stat | nc <zk-host> 2181 on each member. Confirm all members are present, latency is consistent across nodes, and the outstanding request count is near zero. A degraded member (slow disk, GC pause, network issue) drags the entire ensemble because writes require quorum acknowledgment.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
ZK watch count (wchs or mntr)Direct measure of watch pressure on the ensembleCount growing without corresponding traffic growth
ZK request latencyLeading indicator for session timeout cascadeSustained average > 50ms; > 100ms is critical
pulsar_lb_unload_bundle_totalEach unload triggers client reconnections and new watchesSustained > 1/min outside maintenance
pulsar_active_connectionsReconnection storms generate watch churnFluctuating counts or unexplained growth
pulsar_broker_lookup_failuresLookup failures indicate ZK cannot serve metadataFailure rate > 1% of total lookups sustained
Broker metadata-store session events (log Session expired / Connection loss)Session loss is the cascade triggerSession events on any broker
ZK transaction log disk I/OZK writes are synchronous to this diskawait elevated or %util consistently high
ZK client connection countMore sessions mean more watchesGrowing trend correlating with latency

Fixes

Stop the reconnection storm

The immediate goal is to break the feedback loop between watch churn and ZK latency.

If a consumer fleet is in a reconnection loop (client bug, expired credential, broker repeatedly dropping connections), identify the applications with the highest reconnection rate from broker logs. Tight retry loops with no backoff are the most common trigger. Ensure client libraries use exponential backoff with jitter.

If the storm was triggered by a broker restart or bundle unload, wait for the system to settle. Do not restart brokers to “fix” the issue. Restarts generate additional reconnection load and worsen the cascade.

Reduce ZK pressure from topic count

If the watch baseline is high because of topic count rather than churn, the cluster has outgrown its ZK capacity.

  • Audit and reduce topic count. Consolidate topics where possible. Partitioned topics with many partitions multiply metadata load because each partition is a separate managed ledger with its own znodes.
  • Enable metadata batching (introduced in Pulsar 2.10; the default is true — verify it is not disabled in broker.conf). This batches metadata operations to reduce ZK request count. Relevant parameters: metadataStoreBatchingMaxDelayMillis=5, metadataStoreBatchingMaxOperations=1000, metadataStoreBatchingMaxSizeKb=128. Batching introduces up to 5ms of artificial latency per operation. For clusters with few topics, disabling batching may yield better latency.

Batching applies to metadata read/write operations (get/put/delete/children); watch registrations are not batched (AbstractBatchedMetadataStore).

Tune session timeout

If ZK latency is elevated but not catastrophic, increasing zooKeeperSessionTimeoutMillis buys time by preventing premature session expiry. Pulsar’s default is 30 seconds (zooKeeperSessionTimeoutMillis=-1 falls back to metadataStoreSessionTimeoutMillis=30_000; earlier configs set the literal 30000). Raising it to 60 seconds gives brokers more headroom during latency spikes.

The tradeoff: a higher session timeout means slower detection of genuinely dead brokers. The load balancer will not reassign bundles from a dead broker until its session expires.

Configure session expiry policy (Pulsar 2.10+)

Set zookeeperSessionExpiredPolicy=reconnect to prevent broker shutdown on ZK session expiry. Instead of shutting down, the broker reconnects and re-owns bundles from its in-memory cache. This prevents the simultaneous broker shutdown that turns a ZK latency spike into a cluster-wide outage. The default is reconnect since Pulsar 2.10 (2.8/2.9 shipped shutdown).

Add ZK resources (buys time, does not fix the root cause)

Scaling the ZK ensemble (more members, more CPU, faster transaction log disks) increases throughput and reduces latency. This is a valid stopgap but does not address the underlying watch pressure. If topic count or client churn continues to grow, you will hit the wall again.

The transaction log disk is the most impactful upgrade. ZK writes are synchronous and sequential. A dedicated NVMe device for the transaction log eliminates disk I/O as a bottleneck. Do not share this disk with snapshots, ZK data, or any other workload.

Consider Oxia migration

Current Pulsar docs recommend Oxia for new clusters and document a live migration framework from ZooKeeper (Oxia support landed experimentally in Pulsar 3.3.0, PIP-335). For clusters hitting persistent ZK capacity limits, Oxia eliminates the watch-based coordination model that causes watch explosions. This is a strategic infrastructure decision, not an incident response action.

Prevention

  • Monitor ZK watch count proactively. Track wchs or mntr output over time, establish a baseline, and alert on sustained growth.
  • Correlate watch growth with bundle unload rate and client reconnection rate. The three signals together tell the story: reconnections drive watch churn, which drives ZK latency, which drives session expiry, which drives more reconnections.
  • Treat sustained ZK latency above 10ms as an early warning. Do not dismiss intermittent latency spikes as transient.
  • Audit topic count growth. Topics accumulate over time from testing, abandoned features, and dynamic topic creation. High topic count is the multiplier that turns normal reconnection churn into a watch explosion.
  • Configure client reconnection with backoff and jitter. Tight retry loops are the most common trigger for reconnection storms. Ensure all client libraries use exponential backoff.
  • Set zookeeperSessionExpiredPolicy=reconnect on all brokers. This is the single most effective configuration change to prevent a ZK latency spike from cascading into a full cluster outage.

How Netdata helps

  • Per-second metric collection captures ZK latency spikes and watch count changes that 15-second scrape intervals miss. The feedback loop between watch churn and latency can develop in under a minute.
  • Correlating ZK watch count with pulsar_lb_unload_bundle_total, pulsar_active_connections, and pulsar_broker_lookup_failures on a single timeline makes it immediately visible whether a latency spike is caused by watch pressure, bundle thrashing, or disk I/O.
  • Anomaly detection on ZK request latency and connection count flags early deviation from baseline before latency crosses critical thresholds.
  • JVM metrics (heap usage, GC pause times) for ZK servers reveal whether latency spikes are caused by ZK-side GC pauses rather than watch pressure.
  • Disk I/O metrics on the ZK transaction log volume distinguish disk-bottlenecked latency from CPU-bottlenecked watch processing.