The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-pulsar / apache-pulsar-write-ack-quorum-explained ▌

Operations Guides

Apache Pulsar ensemble, write quorum, and ack quorum: what E, Qw, and Qa actually guarantee

Every persistent message in Apache Pulsar passes through BookKeeper’s quorum system, governed by three parameters: ensemble size (E), write quorum (Qw), and ack quorum (Qa). Configured per namespace and inherited by topics, they determine how many bookies receive each entry, how many must confirm a durable write before the producer is acknowledged, and how the cluster behaves when bookies fail, slow down, or are taken offline.

The nuance is in what each parameter does not guarantee. Write quorum is not a durability floor for acknowledged entries. Ack quorum is not the number of copies that will exist after the write completes. Ensemble size constrains ledger creation in ways that surprise teams during maintenance windows: you can have plenty of bookies but still be unable to create new ledgers.

These three parameters sit behind most write-path incidents: publish latency spikes, NotEnoughBookiesException during rolling restarts, broker OOM under degraded bookie conditions, and gaps between expected and actual replication depth.

What it is and why it matters

BookKeeper persists each managed ledger’s data across a set of bookies called an ensemble. Three values define the persistence policy.

Ensemble size (E) is the total number of bookies assigned to a single ledger, fixed at ledger creation time. Every entry in the ledger is distributed across a subset of these E bookies.

Write quorum (Qw) is the number of bookies that receive each individual entry. This is the intended replication factor for each entry. When Qw is less than E, entries are striped across different subsets of the ensemble.

Ack quorum (Qa) is the number of bookies that must acknowledge a durable write (journal fsync complete) before the broker acknowledges the producer. This is the minimum replication factor guaranteed to the producer. The system tolerates up to Qa - 1 bookie failures without data loss for acknowledged entries.

The invariant is strict: E >= Qw >= Qa. Ledger creation fails if this is violated.

Set the policy per namespace:

# Set persistence policy for a namespace
pulsar-admin namespaces set-persistence <tenant/namespace> -e <E> -w <Qw> -a <Qa>

# Check the current policy
pulsar-admin namespaces get-persistence <tenant/namespace>

Changes apply to new ledgers only. Existing ledgers keep their current ensemble until they roll over.

When get-persistence returns all zeros, the namespace inherits broker-level defaults — managedLedgerDefaultEnsembleSize=2, managedLedgerDefaultWriteQuorum=2, managedLedgerDefaultAckQuorum=2 in broker.conf.

How it works

The striping algorithm

When Qw is less than E, entries are striped across the ensemble. For a given entry ID, the write quorum is the subsequence of the ensemble starting at bookie index (entry_id mod E), with length Qw. Wrapping is modular.

For example, with E=5 and Qw=3:

Entry IDStart indexBookies written to
00B1, B2, B3
11B2, B3, B4
22B3, B4, B5
33B4, B5, B1
44B5, B1, B2
50B1, B2, B3

There are exactly E distinct write quorums in any ensemble. The pattern repeats every E entries. When Qw equals E, every entry goes to all bookies and no striping occurs.

The write path

When a producer sends a message, the broker appends it to the managed ledger, which writes the entry to Qw bookies in parallel. Each bookie writes the entry to its journal and fsyncs. Once Qa of those Qw bookies confirm the fsync, the broker acknowledges the producer.

flowchart TD
    P[Producer sends message] --> B[Broker: managed ledger append]
    B --> S[Select Qw bookies from ensemble of E]
    S --> W1[Bookie A: journal fsync]
    S --> W2[Bookie B: journal fsync]
    S --> W3[Bookie C: journal fsync]
    W1 -->|ack| GATE{Qa of Qw acks received?}
    W2 -->|ack| GATE
    W3 -->|ack| GATE
    GATE -->|Yes| ACK[Broker acks producer]

Publish latency is bounded by the Qa-th fastest fsync among the Qw bookies in the write set. One slow bookie that happens to be the Qa-th responder is enough to drag latency for every entry that includes it.

Where it shows up in production

NotEnoughBookiesException requires E bookies, not Qw

BookKeeper requires at least E available bookies to create a new ledger, because the ensemble is selected at ledger creation time, before any entries are written.

If you set E=5 and Qw=3, you might expect that losing 2 bookies is safe since you only need 3 for writes. But the broker cannot create new ledgers until at least 5 bookies are available. Existing ledgers continue to accept writes (their ensemble is already assigned), but ledger rollover fails. During rolling bookie restarts, this causes NotEnoughBookiesException in broker logs and publish errors for affected topics.

Setting E higher than Qw reduces write availability without improving durability for any individual entry. It spreads entries across more bookies (load distribution) at the cost of stricter availability requirements for ledger creation.

The slowest bookie in the write set drags publish latency

The broker waits for Qa acks out of Qw writes. Publish latency for each entry depends on which specific bookies are in its write set and how fast they fsync. If one bookie’s journal disk degrades, every entry that includes that bookie sees elevated latency.

The impact is partial: only topics whose ledgers include the degraded bookie are affected. Other topics continue normally. This partial impact makes the problem harder to detect in aggregate metrics. Per-bookie journal sync latency is the signal that isolates the culprit — the Prometheus metric is bookie_journal_JOURNAL_SYNC (there is no ..._LATENCY variant).

Entries can remain at Qa copies, never reaching Qw

When Qa is less than Qw, an entry is acknowledged after Qa bookies confirm. The remaining (Qw - Qa) bookies may still be processing the write. Under normal conditions, they complete and the entry reaches Qw copies. But several scenarios can leave entries permanently at Qa replication depth:

  • Ensemble changes during ledger recovery replace bookies, and pending writes to the removed bookies are abandoned.
  • Ledger closure races with in-flight writes.
  • A bookie fails after Qa acks but before all Qw writes complete.

This is a legal state of the BookKeeper protocol. Entries that reach Qa but not Qw remain at Qa copies indefinitely. The “Guaranteed Write Quorum” protocol, which would enforce Qw copies for all acknowledged entries, has been formally verified in TLA+ (Jack Vanlightly’s BookKeeper TLA+ specifications) but has not been adopted by any released BookKeeper as of 2026.

Qw is an intended replication factor, not a guaranteed one. Qa is the only replication depth you can rely on for acknowledged entries.

Broker memory pressure when Qa < Qw and a bookie is slow

When Qa is less than Qw and a bookie responds slowly to add-entry requests, the broker holds pending operations in memory waiting for all Qw writes to complete or time out. Entries that have already reached Qa but have not received all Qw acks keep their pending add operations queued. Under sustained slow-bookie conditions, these pending operations accumulate in the broker’s direct memory, eventually causing OOM.

This is documented in Pulsar issue #14861 (“Broker OOM when the bookie is slowly to respond and AQ is smaller than WQ”, reproduced with E=2, W=2, A=1). The root cause is that the gap between Qa and Qw creates a window where operations are acknowledged to the producer but cannot be fully retired on the broker side. The wider the gap (Qw minus Qa), the more pending state can accumulate.

Rack-awareness constrains placement

When rack-aware placement policy is enabled, the BookKeeper client selects bookies from different failure domains. The relevant bookkeeper.conf settings are enforceMinNumRacksPerWriteQuorum (default false) and minNumRacksPerWriteQuorum (default 2). With E=3, Qw=3 and minNumRacksPerWriteQuorum=3, the ensemble must include bookies from at least 3 distinct racks. If the required number of racks is not available, ledger creation fails.

Before decommissioning a bookie, verify that the E >= Qw >= Qa invariant still holds with one fewer bookie in the cluster. If decommissioning drops available bookies below E, all new ledger creation for that namespace fails.

Tradeoffs and when to use it

Common configurations

ConfigurationUse caseTradeoff
E=2, Qw=2, Qa=2Maximum safety, minimal bookie countNo striping. Cannot tolerate any bookie failure for writes. Requires exactly 2 bookies available for ledger creation.
E=3, Qw=3, Qa=2Balanced durability and availabilityTolerates 1 bookie failure without data loss. Requires 3 available bookies for new ledgers. Most common production setting.
E=3, Qw=2, Qa=1Maximum throughput, minimal latencyCan lose data on single bookie failure. Qa=1 blocks ledger recovery if that bookie is down. Unsafe for most workloads.
E=5, Qw=3, Qa=2Spread load across more bookiesStriping degrades read performance. Requires 5 available bookies for ledger creation. Entries may remain at Qa=2 copies permanently.

Why E > Qw usually hurts more than it helps

Striping (E > Qw) distributes entries across more bookies, which can smooth write load. But it carries two costs that usually outweigh the benefit.

Read performance degrades. BookKeeper optimizes for sequential reads from a single bookie. When entries are striped across different bookie subsets, a consumer reading sequentially must hit multiple bookies instead of reading from one. This increases read fan-out and latency, particularly for catch-up reads that bypass the broker cache.

Availability decreases. Ledger creation requires E available bookies. Setting E higher than necessary means more bookies must be up to create new ledgers, which makes rolling restarts and maintenance windows riskier.

BookKeeper’s sticky read optimization (bookkeeperEnableStickyReads=true) only takes effect when E equals Qw; with striping enabled, sticky reads provide no benefit. The default is true in current Pulsar broker.conf, which states the caveat explicitly.

Multiple BookKeeper committers recommend setting E = Qw unless you have a specific, measured reason to stripe.

Why Qa = 1 is dangerous

Setting ack quorum to 1 means the producer is acknowledged after a single bookie confirms. If that bookie fails before the remaining Qw - 1 writes complete, the entry exists on only one node. Recovery cannot always determine whether that bookie holds the entry, which can block ledger recovery entirely.

Qa = 1 should be reserved for non-critical, ephemeral data where data loss is acceptable and recovery speed does not matter.

The metadata persistence parameters are dead code

The ManagedLedgerConfig fields metadataEnsembleSize, metadataWriteQuorumSize, and metadataAckQuorumSize have had zero effect since Pulsar 2.2 (2018): PR #2535 refactored ledger creation so cursor ledgers use the data-ledger settings, and the fields’ getters are never called (see Pulsar issue #26212). The dead fields were removed on the master/5.0 development line by PR #26293 (August 2026), but still exist, unused, in the 3.x and 4.x lines. If you are setting them expecting them to control cursor ledger replication, they do nothing.

Signals to watch in production

SignalWhy it mattersWarning sign
bookie_journal_JOURNAL_SYNC (per bookie, P99)Journal fsync is on the write critical path. The Qa-th slowest bookie in each write set sets publish latency.One bookie’s P99 is 2x or more above others in the same ensemble.
pulsar_broker_publish_latency (P99)End-to-end write latency as seen by the broker. Reflects quorum ack wait time.Sustained elevation over baseline, especially when only some topics are affected.
bookkeeper_server_ADD_ENTRY_IN_PROGRESSWrite queue depth on each bookie. Indicates the bookie cannot keep up with incoming writes.Queue not draining within 30 seconds after traffic bursts.
bookie_journal_JOURNAL_FORCE_WRITE_QUEUE_SIZEEarliest signal for journal disk saturation. Rises before sync latency spikes.Sustained non-zero depth between write batches.
bookie_SERVER_STATUSWhether each bookie is writable. Read-only bookies reduce the pool available for new ledgers.Transitions to 0 (read-only). Check against configured E for ledger creation risk.
auditor_NUM_UNDER_REPLICATED_LEDGERSLedgers with fewer copies than configured. Indicates recovery is needed or failing.Non-zero count that does not trend to zero after a bookie event.
NotEnoughBookiesException in broker logsLedger creation failure. Available bookies below E, or rack-awareness constraints not met.Any occurrence during non-maintenance periods.

Metric names in this table were checked against BookKeeper’s Prometheus output and the Pulsar 3.0.x/4.0.x references (bookie_SERVER_STATUS, bookie_journal_JOURNAL_SYNC, bookie_journal_JOURNAL_FORCE_WRITE_QUEUE_SIZE, bookkeeper_server_ADD_ENTRY_IN_PROGRESS, auditor_NUM_UNDER_REPLICATED_LEDGERS, pulsar_broker_publish_latency).

How Netdata helps

  • Per-bookie journal sync latency at per-second resolution isolates which bookie in an ensemble is the Qa-th slow responder. Correlate journal sync spikes with broker publish latency to identify a degraded journal disk.
  • Add-entry in-progress and force write queue depth surface write-path saturation before journal sync latency spikes. These are the earliest indicators that a bookie cannot keep up.
  • Bookie server status changes fire immediately when a bookie goes read-only. Compare available bookie count against your configured E to assess ledger creation risk before producers see errors.
  • Under-replicated ledger count tracked over time shows whether auto-recovery is keeping up after a bookie failure. A growing count means durability risk is accumulating.
  • Anomaly detection on publish latency flags the partial, topic-specific latency spikes caused by a single degraded bookie, even when aggregate cluster metrics look normal.