The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-pulsar
APACHE PULSAR · OPERATIONS PLAYBOOK

Apache Pulsar's two-layer split: stateless brokers, a storage layer where every write waits on one journal fsync, and a metadata store that can fence the whole fleet

A messaging system that separates serving from storage: brokers own nothing durably and can be moved or fenced at any moment, BookKeeper bookies persist every message behind a journal fsync that gates the entire write path, and ZooKeeper coordinates all of it. We trace how that design behaves under load, where a problem on one layer surfaces as a symptom on another, and what to do when it does.

"

Pulsar's separation of serving from storage is its great strength and the reason its incidents are so hard to place: the symptom you see on the broker almost always started one layer down.

The defaults work. Until one bookie's journal disk cannot fsync fast enough and every producer whose topic touches that bookie slows down at once — because a write is not acknowledged until the ack quorum has it on disk. Until a broker's off-heap memory fills and Netty throws OutOfDirectMemoryError, killing the broker while the heap dashboard still reads green. Until a GC pause outlasts the ZooKeeper session timeout, the broker is fenced, loses its bundles, and every client on it reconnects at once. Until a consumer stalls, backlog crosses its backlog quota, and the policy you never set turns a consumer problem into a producer outage. Until a bookie fills to diskUsageThreshold, flips to read-only, and new ledgers fail with NotEnoughBookiesException.

These guides are written for engineers who already run Pulsar, not for people learning what a topic is. The goal is the mental model of how brokers, BookKeeper, and the metadata store actually behave under load, the failure patterns that keep recurring across the two layers, the monitoring story that catches them before they page anyone, and the runbooks you wish someone had handed you before your last incident.

How Apache Pulsar actually runs in production

Pulsar is not one system but three tiers that fail independently: a stateless broker fleet that owns topics only by lease, a BookKeeper storage layer where durability and the write path actually live, and a metadata store that coordinates everything. A message travels down through all of them and back up to consumers — and most production failures live in the seams between tiers, not inside any one of them.

01
producers / clients
Every producer and consumer first does a topic lookup to find which broker owns the topic, then holds a TCP connection. Each connection costs a file descriptor and Netty direct memory for buffers. Client libraries reconnect automatically — which is a feature until a bundle unload or fenced broker makes thousands of them reconnect at once.
CLIENT
02
brokers + bundle ownership
Brokers are stateless. Each owns a set of namespace bundles — hash-range slices of the topic space — assigned by the load balancer and recorded in the metadata store. A broker owns a topic only by lease; when it dies, is overloaded, or loses its session, its bundles move and its topics are fenced and reopened elsewhere.
BROKER
03
managed ledger + entry cache
Each topic is a managed ledger: a continuous append stream wrapping a sequence of BookKeeper ledgers. Recently written entries sit in an off-heap entry cache so connected consumers read from broker memory. A cache miss sends the read down to the bookies, adding disk I/O on the storage layer.
LEDGER
04
BookKeeper bookies
The storage layer, and where durability lives. Bookies know nothing about topics — they store opaque ledger entries. Each managed ledger writes to an ensemble of bookies with a write quorum (Qw) and ack quorum (Qa); the producer is acknowledged only after Qa bookies confirm. The slowest bookie in the write set sets the publish latency.
BOOKIE
05
journal + entry log disks
Each bookie writes a sequential, per-entry-fsync'd write-ahead journal on one dedicated disk, and interleaved entry logs (the actual payloads) on another. The journal fsync is the physical heartbeat of write throughput; sharing it with any other workload — including the entry-log reads — is the single most common architecture mistake.
STORAGE
06
subscriptions + cursors
Consumers read through subscriptions, each holding a cursor (mark-delete position) in the metadata store. Delivered-but-unacknowledged messages count against a per-subscription limit; hit it and dispatch silently freezes. An abandoned subscription's cursor pins every message after it, so storage cannot be reclaimed.
CONSUME
07
metadata store (ZooKeeper)
ZooKeeper (still the default in 3.x; Oxia and etcd exist) is the single source of truth: bundle-to-broker ownership, ledger metadata, cursors, schemas, policies. Its latency is a leading indicator for the whole cluster — when it crosses the session timeout, brokers are fenced and the cluster destabilises.
METADATA
08
JVM runtime (heap + direct memory)
Brokers and bookies are JVMs. Heap holds connection and cursor state; direct (off-heap) memory holds Netty buffers and the entry cache and is invisible to heap metrics. A full GC pause that outlasts the ZK session timeout is what turns memory pressure into a fenced broker and the GC death spiral.
JVM

Why this matters: 'Pulsar is slow' or 'publishes are timing out' can come from one saturated journal disk, an off-heap memory crash the heap dashboard hides, a ZooKeeper latency storm fencing brokers, a single hot broker, a stalled consumer tripping a backlog quota, a bookie gone read-only, or ownership oscillating between brokers. The symptom rhymes but each tier has a different signal — and a different fix.

The failures you'll actually see

Most Pulsar incidents fall into a small set of recurring patterns that cross the broker/BookKeeper/metadata seams. Recognise the shape, and triage gets dramatically faster.

CRITICAL

The write stall

A bookie's journal disk cannot fsync fast enough. The journal force-write queue grows, add-entry operations queue behind it, and because a write is not acked until the ack quorum has it on disk, every producer whose topic touches that bookie slows down. Publish latency climbs, throughput collapses, and the whole write path waits on one disk. This is the single most common performance failure in Pulsar.

  • bookie_journal_JOURNAL_SYNC P99 well above the device baseline
  • JOURNAL_FORCE_WRITE_QUEUE_SIZE sustained above zero
  • bookkeeper_server_ADD_ENTRY_IN_PROGRESS not draining after bursts
  • pulsar_broker_publish_latency rising on topics using that bookie
Investigate
CRITICAL

The off-heap broker crash

Netty uses direct (off-heap) memory for every network buffer and the managed ledger cache. Under a connection storm, large messages, or an oversized cache, direct memory fills and Netty throws OutOfDirectMemoryError — the broker can no longer allocate buffers for any I/O and is dead in all but PID. Heap monitoring shows plenty of room, so standard JVM dashboards miss it entirely.

  • io.netty.util.internal.OutOfDirectMemoryError in broker logs
  • Process RSS far larger than configured heap
  • Broker health check failing while the process still exists
  • All topics on the broker unavailable at once
Investigate
CRITICAL

The ZooKeeper session cascade

ZooKeeper latency rises — watch explosion, transaction-log disk I/O, a GC pause on a ZK server. Brokers miss heartbeats, their sessions expire, and they are fenced and lose bundle ownership. Every client reconnects at once, the reconnect metadata storm pushes ZK latency higher, and more sessions expire. The cluster thrashes even though every broker process is running fine.

  • Session expired events on multiple brokers simultaneously
  • Metadata store latency elevated before any broker symptom
  • Bundle reassignment and client reconnections spiking together
  • Broker lookup failures and client-side LookupException rising
Investigate
ACTIVE

The backlog avalanche

A consumer falls behind. Backlog grows, so reads miss the entry cache and go to the bookies, whose disks now serve reads and writes; write latency rises, dispatch slows, more consumers fall behind, more reads hit the bookies. Left alone, backlog fills bookie disk or crosses its quota — at which point the policy either holds producers, errors them, or silently evicts the oldest messages.

  • pulsar_subscription_back_log growing continuously with consumers connected
  • pulsar_rate_out flat or falling while pulsar_rate_in is normal
  • Managed ledger cache miss rate climbing, bookie read latency rising
  • Publish latency rising once bookie disks serve both reads and writes
Investigate
CRITICAL

Not enough bookies to write

The broker rolls to a new ledger periodically and on recovery, but cannot satisfy the ensemble and quorum requirement — often after a bookie filled to diskUsageThreshold and went read-only, or after too many bookies were restarted at once. Ledger creation fails with NotEnoughBookiesException and new messages cannot be written durably. Rack-awareness makes it subtler: enough bookies numerically, not enough distinct failure domains.

  • NotEnoughBookiesException / BKException on ledger create in broker logs
  • One or more bookies at bookie_SERVER_STATUS == 0 (read-only)
  • bookie_ledger_writable_dirs == 0 on affected bookies
  • Publish errors on topics whose ledgers need to roll
Investigate
IMMINENT

Ownership oscillation and fencing loops

Two brokers keep trading a bundle: A opens the managed ledgers, the balancer unloads it, B fences the ledgers with LedgerFencedException and opens new ones, the balancer moves it back. Each cycle seals a ledger, opens another, and reconnects every client on those topics. Normal rebalancing converges in minutes; this can run indefinitely, and if a new ledger cannot be created after fencing, the topic enters a broken state.

  • LedgerFencedException with no matching planned bundle transfer
  • Rapid acquired / released ownership log lines on two brokers
  • Bundle unload rate high with constant client reconnections
  • Spiky publish latency on the contested topics
Investigate
Choosing a tool

Best Apache Pulsar Monitoring Tools Ranked (2026)

A ranked review of the tools teams actually shortlist here, what each one is genuinely good at, and how the pricing behaves as you scale.

Apache Pulsar monitoring maturity levels

Pulsar observability works in four practical levels. Each is a complete operation, not a stepping stone. Pick the level that matches how much your cluster matters. Most production deployments should land at the second level — and monitor BookKeeper as a first-class concern, not an afterthought.

Level 1: Survival

Know that something is wrong

Survival monitoring is the floor. With these signals you can answer one question: are all three tiers alive and is data flowing? You will not learn what broke, but you will learn that something broke before users do. Survival is enough for dev clusters and non-critical pipelines.

  • Broker process up HTTP admin endpoint on :8080 answering, not just the JVM alive?
  • Bookie process up + writable bookie_SERVER_STATUS == 1; read-only means it stopped accepting writes.
  • Metadata store (ZooKeeper) up ruok / quorum healthy — the coordinator every tier depends on.
  • Message rates in and out pulsar_rate_in / pulsar_rate_out — is data actually flowing?
  • Bookie ledger disk usage Approaching diskUsageThreshold flips the bookie to read-only.

Level 2: Operational

Diagnose most incidents on your own

Operational monitoring is what most production clusters should target. Survival tells you something is wrong; operational tells you what. With this coverage your team can usually diagnose an incident on its own: write stalls, backlogs, ZooKeeper degradation, connection pressure, grey lookup failures.

  • Per-namespace publish / dispatch rate Per-topic rates catch partial failures cluster averages hide.
  • Subscription backlog on critical topics Alert on rate-of-change, not absolute size.
  • Broker publish latency P99 The producer SLI; P50 looks fine while P99 spikes.
  • Bookie journal sync latency P99 The physical write-throughput limit; the cluster heartbeat.
  • Active connections + FD ratio Each connection costs an FD and direct memory; watch the trend.
  • Metadata store latency The #1 leading indicator; > 50ms sustained is a warning.
  • Broker lookup failures New clients can't find their topic while old ones work — grey failure.
  • Bookie disk usage trend Runway to read-only, not just the current value.

Level 3: Mature

Catch problems before they become incidents

Mature monitoring catches problems before they wake anyone up. A subscription quietly not acking, the entry cache thrashing, ledgers left under-replicated, one broker owning far too many topics, a bookie write queue that no longer drains. None of these page you on day one. They become page-out incidents on day thirty.

  • Per-subscription backlog + redelivery Redelivery storms mean receiving without progressing.
  • Unacked messages vs the freeze limit At maxUnackedMessagesPerSubscription, dispatch freezes silently.
  • Managed ledger cache hit / miss Misses send consumer reads down to the bookies.
  • Under-replicated ledger count Data at risk after a bookie failure; should trend to zero.
  • Add-entry in-progress + force-write queue Write-path saturation before latency spikes.
  • Bundle unload rate Steady-state unloads should be rare; high means thrashing.
  • Topic count per broker One hot broker while cluster averages look healthy.
  • TLS certificate expiry Not a broker metric; an expired cert is a total, silent outage.

Level 4: Expert

Reactive instrumentation after real incidents

Expert signals enter your stack the day after a specific incident proved you needed them. Direct memory, GC pauses, fencing churn, entry-log GC lag, dead-letter arrival. Most teams never need every signal here. Add the ones your incident history says you do — and remember Pulsar exposes several of these only via JMX or the logs, not Prometheus.

  • Broker direct (off-heap) memory Via JMX; the invisible killer heap dashboards never show.
  • GC pause times A pause past the ZK session timeout fences the broker.
  • Journal add-entry latency percentiles Finer-grained than sync latency for disk diagnosis.
  • Ledger fencing events / ownership churn Fencing with no bundle transfer means split ownership.
  • Entry-log active vs total space BookKeeper GC falling behind on space reclamation.
  • Dead-letter queue arrival rate A ledger of application processing failures.
  • Message expiration rate TTL deleting unread messages — silent consumer-side loss.
  • ZooKeeper watch count Watch explosions crush metadata latency (wchs).

Operating mistakes worth avoiding

The traps Pulsar teams keep falling into. Each has a clear, well-known fix. Most teams only learn it after an incident.

Watching rates but not subscription backlog

Teams monitor publish and consume rates but not backlog directly. Backlog is the accumulation signal — it grows silently until it fills bookie disk, hits retention, or trips a <code>backlog quota</code> that throttles producers. By the time it fires an alert, the damage has already cascaded. Monitor per-subscription backlog and its growth rate; alert on sustained growth, not an absolute threshold.

Ignoring broker direct (off-heap) memory

Standard JVM monitoring shows heap. Pulsar brokers use significant direct memory for Netty buffers and the managed ledger cache, and Pulsar does NOT expose it as a Prometheus metric. Teams see 'heap is fine at 60%' and cannot explain why the broker hung or crashed with <code>OutOfDirectMemoryError</code>. Set <code>MaxDirectMemorySize</code> explicitly and watch it via JMX, or track process RSS minus heap.

Not separating journal and storage disk metrics

Bookies use separate disks — they must — for the write-ahead journal and the entry logs. Teams monitor aggregate disk I/O and miss that the journal disk is saturated. Journal saturation directly blocks the entire write path; it is the most impactful failure point in the stack. Monitor per-device I/O and use the <code>journalIndex</code> label to isolate the journal.

Missing ZooKeeper degradation until the cascade

ZK latency creeps up over days. Brokers occasionally lose ownership and recover, and it gets dismissed as transient — until latency crosses the session timeout and the whole cluster destabilises in a session cascade. Treat sustained metadata latency above 10ms as an early warning; it is the number-one leading indicator for cluster-wide failures.

No visibility into consumer acknowledgment patterns

Teams see consumers connected and messages dispatched and assume progress. They miss that consumers are not acking — driving redelivery storms, or hitting <code>maxUnackedMessagesPerSubscription</code> and freezing dispatch entirely. Compare acknowledgment rate to dispatch rate; the gap, plus the redelivery rate, is the real indicator.

Reactive capacity planning on bookie disk

Teams wait until a bookie disk is 90% full to expand. Bookie disk growth is predictable from retention policy and publish rate, and a full bookie flips to read-only — which can cascade into cluster-wide write failure. Plan expansion at 70%, estimate runway to 90%, and factor in that a consumer outage accelerates the fill.

Monitoring backlog size but not backlog age

A backlog of 1,000 messages from five seconds ago is nothing; 1,000 messages from five hours ago is an SLA breach. Size is volume, age is latency. Age is also what reveals cursor leaks and abandoned subscriptions that pin storage. Without it, teams react to normal batch spikes and miss the degradation that actually matters.

Alerting on average latency instead of P99

Publish latency P50 can look beautiful while P99 spikes — and P99 is what producers experience during bursts and what triggers their timeouts. Teams that alert on mean or P50 miss the degradation entirely. Alert on P99 for publish and journal-sync latency, and remember broker-side latency excludes client-to-broker network and batching.

Apache Pulsar runbooks in this section

Each guide is a focused runbook for one symptom or topic. Pick one when you have an incident, or use the categories to learn the area.

WHERE TO GO NEXT

Setting up Apache Pulsar monitoring, or putting out a fire?

If you're starting from scratch, the monitoring checklist is the path of least regret. If you're mid-incident, jump straight to the symptom that matches what you're seeing.