The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-pulsar / apache-pulsar-publish-latency-high ▌

Operations Guides

Apache Pulsar publish latency high: reading pulsar_broker_publish_latency P99

Your alert fired on sustained P99 elevation of pulsar_broker_publish_latency above 2x the rolling baseline. Producers are seeing slow acknowledgements.

The metric pulsar_broker_publish_latency is a Summary metric exposed on the broker Prometheus endpoint. It measures the time from when the broker receives a message from a producer through the BookKeeper write path (write quorum Qw, ack quorum Qa) and back to the client callback. It is broker-side only: it excludes client-to-broker network time and producer-side batching delay. The Summary type provides quantiles at 0.5, 0.95, 0.99, 0.999, 0.9999, and 1.0. Alert on P99. P50 can look healthy while P99 is spiking, and it is the tail that triggers producer timeouts.

Healthy clusters typically see P99 below 10ms on SSD-backed bookies. The absolute threshold is workload-dependent; relative degradation (2x rolling baseline) is the more reliable alerting signal.

What this means

When P99 publish latency elevates, the broker is taking too long to persist messages through BookKeeper and acknowledge the producer. The bottleneck is in the write path: broker processing, BookKeeper journal fsync, quorum ack wait, GC pauses, or network contention between broker and bookies.

Publish latency is a downstream symptom. The root cause is almost always journal disk saturation, GC pressure, or network degradation, and it has been building for seconds or minutes before publish latency spikes. One slow bookie in the write quorum is enough to inflate latency for every topic whose ensemble includes that bookie.

Common causes

CauseWhat it looks likeFirst thing to check
Journal sync saturationOne or more bookies show elevated JOURNAL_SYNC latency; force write queue growing; ADD_ENTRY_IN_PROGRESS not drainingbookie_journal_JOURNAL_SYNC P99 on each bookie
Broker GC death spiralPublish latency spikes correlate with GC pauses; ZK sessions dropping; bundle unloads increasingBroker GC logs or jstat -gc
Batching distortionP99 elevated but throughput (bytes/sec) stable; message rate low relative to entry rateCorrelate with pulsar_rate_in and producer batch settings
Slow bookie in ensembleOnly topics using specific bookies affected; other topics normal; per-bookie ADD_ENTRY latency shows outlierPer-bookie bookkeeper_server_ADD_ENTRY_REQUEST latency
Broker-bookie network contentionJournal sync latency normal on bookies but publish latency still high; TCP retransmits elevatedNetwork interface utilization and latency between broker and bookie hosts

Quick checks

These are safe, read-only checks. Run them in order.

# Check publish latency P99 on the broker
curl -s http://<broker-host>:8080/metrics | grep pulsar_broker_publish_latency

# Check publish rate to rule out batching distortion
curl -s http://<broker-host>:8080/metrics | grep -E "pulsar_(rate|throughput)_in"

# Check journal sync latency on each bookie
# Pulsar's bookkeeper.conf serves metrics on 8000 (httpServerPort=8000); upstream BookKeeper defaults to 8080
curl -s http://<bookie-host>:8000/metrics | grep bookie_journal_JOURNAL_SYNC

# Check journal force write queue depth (earliest saturation signal)
curl -s http://<bookie-host>:8000/metrics | grep JOURNAL_FORCE_WRITE_QUEUE_SIZE

# Check bookie write queue pressure
curl -s http://<bookie-host>:8000/metrics | grep bookkeeper_server_ADD_ENTRY_IN_PROGRESS

# Check raw journal disk I/O stats
iostat -x 1

# Check broker GC behavior
jstat -gc <broker-pid> 1000

# Check broker heap state
jcmd <broker-pid> GC.heap_info

# Check active connections for leak or storm patterns
curl -s http://<broker-host>:8080/metrics | grep pulsar_active_connections

How to diagnose it

The diagnostic flow narrows from “publish latency is high” to a specific root cause by checking the write path in order of likelihood.

flowchart TD
    A["P99 > 2x baseline"] --> B{"rate_in stable?"}
    B -- "No" --> C["Check producer errors
or throttling"] B -- "Yes" --> D{"P50 also elevated?"} D -- "Yes, large batches" --> E["Correlate with
producer batch config"] D -- "No, tail only" --> F{"Journal sync
P99 elevated?"} F -- "Yes" --> G["Check force write queue
and disk iostat"] F -- "No" --> H{"Broker GC
pauses > 1s?"} H -- "Yes" --> I["Investigate heap or
direct memory"] H -- "No" --> J["Check broker-bookie
network latency"]
  1. Confirm the degradation is real, not a batching artifact. Pull pulsar_rate_in alongside the latency spike. If message rate dropped but throughput in bytes per second stayed stable, producers may have increased batch size. Higher batch sizes inflate per-batch publish latency because the broker waits for the batch to fill. If both rate and throughput dropped, the spike is real.

  2. Isolate broker-side from bookie-side. Check bookie_journal_JOURNAL_SYNC on each bookie. If journal sync P99 is elevated on one or more bookies, the problem is in the storage layer. If all bookie journal sync latencies are normal, the bottleneck is on the broker (GC, thread contention, direct memory) or the network between broker and bookies.

  3. If bookie-side: trace the write queue. Check bookie_journal_JOURNAL_FORCE_WRITE_QUEUE_SIZE first. This is the earliest saturation signal in the stack. It rises before JOURNAL_SYNC latency spikes and before ADD_ENTRY_IN_PROGRESS grows. If the queue is sustained above zero for more than a few seconds, the journal disk cannot drain fsync operations fast enough. Run iostat -x 1 on the journal disk to confirm high %util and elevated await.

  4. If broker-side: check GC and memory. Use jstat -gc <broker-pid> 1000 to check GC frequency and duration. Any full GC pause above 1 second will stall the broker event loop, causing publish latency spikes and potentially ZK session timeouts. Check direct memory separately: Pulsar does not expose it as a Prometheus metric. Use JMX (java.nio:type=BufferPool,name=direct) or compare process RSS against heap usage. A large gap between RSS and heap indicates direct memory consumption from Netty buffers.

  5. If neither: check the network. If journal sync is normal and GC is clean, investigate broker-to-bookie network latency. Check TCP retransmit rates and NIC utilization on both broker and bookie hosts. The broker writes to Qw bookies in parallel and waits for Qa acks. Network latency directly adds to the quorum ack wait.

  6. Check for the slow-bookie-in-ensemble pattern. Publish latency is dominated by the slowest bookie in the write quorum. You do not need all bookies to be slow. One degraded disk or one bookie with elevated journal latency is enough. Compare bookkeeper_server_ADD_ENTRY_REQUEST latency across all bookies in the ensemble. A single outlier bookie with 5x the latency of others will bottleneck every topic whose ensemble includes it.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
pulsar_broker_publish_latency (P99)Primary producer-experience SLISustained above 2x rolling 1-hour baseline
pulsar_rate_inDistinguishes batching artifacts from real degradationRate drops while throughput stays stable
bookie_journal_JOURNAL_SYNC (P99)Physical limit of write throughputAbove 2x baseline; SSD P99 above 5ms indicates disk stress
bookie_journal_JOURNAL_FORCE_WRITE_QUEUE_SIZEEarliest write-path saturation signalSustained above 0 for more than 10 seconds
bookkeeper_server_ADD_ENTRY_IN_PROGRESSBookie write queue pressureQueue not draining within 30 seconds after traffic bursts
Broker GC pause timesGC stalls block the event loop and spike latencyAny full GC above 1 second
pulsar_active_connectionsConnection storms or leaks stress direct memoryUnexplained growth over days or weeks

Fixes

Journal disk saturation

If iostat shows the journal disk approaching bandwidth limits and the force write queue is not draining:

  • Short term: Reduce publish rate to give the journal disk room to drain. If a single bookie is the bottleneck, consider decommissioning it to trigger ledger recovery to other bookies. This is disruptive and generates heavy replication I/O: run bin/bookkeeper shell decommissionbookie -bookieid <bookie-address:port> only if the bookie is clearly degraded.
  • Medium term: Add bookies to the cluster to distribute write load across more journal disks. New ledgers will be placed on the new bookies automatically.
  • Long term: Dedicate an NVMe device exclusively for the bookie journal. Journal and entry log storage on the same physical disk is the most common architecture mistake in Pulsar deployments. See Apache Pulsar bookie journal and ledger storage on one disk.

Broker GC death spiral

If GC pauses correlate with publish latency spikes:

  • Short term: Reduce managedLedgerCacheSizeMB to free heap. This may increase bookie read load as cache evictions push reads to BookKeeper, so monitor pulsar_ml_cache_evictions after the change.
  • Medium term: Increase JVM heap (-Xmx) if the working set has outgrown the current allocation. Check for memory leaks in custom Pulsar Functions or IO connectors.
  • Consider the GC algorithm: If still running G1GC, switching to ZGC dramatically reduces pause times. Pulsar’s shipped pulsar_env.sh has selected ZGC as the default GC since 2.11, so this mainly applies to deployments that override PULSAR_GC. ZGC pauses are sub-millisecond versus G1GC’s potential multi-second stops under heap pressure. See Apache Pulsar broker GC death spiral.

Network contention

If journal sync is clean and GC is healthy but publish latency is still elevated:

  • Check NIC utilization on broker and bookie hosts. Sustained utilization above 70% leaves no headroom for bursts.
  • Check TCP retransmit rates. Elevated retransmits indicate packet loss or congestion.
  • Verify that broker-to-bookie traffic is not competing with other workloads on shared network infrastructure.

Batching distortion

If the latency spike is a measurement artifact from producer batching:

  • This is not a real degradation. The broker is processing batches efficiently.
  • If the elevated latency violates producer-side SLAs, reduce the batch size or batch delay in the producer configuration to reduce batch fill time.

Slow bookie in ensemble

If one bookie shows consistently higher ADD_ENTRY latency than others:

  • Investigate the bookie’s journal disk health. Run SMART checks on the disk. Check for noisy neighbors on shared cloud storage (EBS latency spikes are common).
  • If the bookie is permanently degraded, decommission it gracefully and let auto-recovery redistribute its data. Monitor auditor_NUM_UNDER_REPLICATED_LEDGERS during recovery to ensure it trends to zero.

Prevention

  • Alert on P99, never P50. P50 can look healthy while P99 is spiking. Set the alert threshold relative to a rolling baseline (2x the 1-hour rolling average), not an absolute number.
  • Monitor the force write queue as a leading indicator. It rises before journal sync latency spikes and before publish latency degrades. See Apache Pulsar journal force write queue growing.
  • Keep journal and entry log on separate physical disks. This is the single highest-impact architecture decision for write latency.
  • Use a 15-second or shorter scrape interval for latency metrics. Sub-minute spikes are invisible at longer intervals.
  • Monitor direct memory alongside heap. Direct memory exhaustion crashes the broker with no heap-level warning. See Apache Pulsar OutOfDirectMemoryError.
  • Track per-bookie write latency distribution. A single slow bookie in the ensemble is the most common partial-degradation pattern. Standard deviation of ADD_ENTRY latency across bookies is the signal.

How Netdata helps

  • Netdata collects pulsar_broker_publish_latency at per-second resolution, so sub-minute P99 spikes are visible rather than averaged away.
  • Publish latency correlated with bookie_journal_JOURNAL_SYNC in a single view makes it immediately clear whether the bottleneck is storage or broker-side.
  • The force write queue depth appears alongside publish latency, exposing the leading indicator before latency degrades.
  • JVM GC pause metrics next to publish latency make GC-induced spikes obvious within seconds.
  • pulsar_rate_in plotted against publish latency distinguishes batching artifacts from real degradation at a glance.