The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / activemq / activemq-ha-split-brain ▌

Operations Guides

ActiveMQ shared-storage HA split-brain: two brokers holding the store lock

In a shared-storage HA pair, exactly one ActiveMQ Classic broker is supposed to hold the KahaDB store lock and run active. The second broker sits in a polling loop, waiting for the lock, with its transport connectors down. Split-brain is the state where that invariant breaks: both brokers believe they are active, both accept client connections, and both write to the same KahaDB store on shared storage.

Two writers on one store means message duplication, ordering violations, and real risk of journal and index corruption. Clients connected to the “wrong” broker get messages the other broker also delivers. Recovery is not just “restart one node”: you have to decide which broker’s view of the store is authoritative, and the store itself may already be damaged.

The trigger is almost always the storage layer, not ActiveMQ. NFS lock-manager edge cases, NFSv4 lease and grace-period races, a filesystem that does not implement POSIX locking correctly (OCFS2 is the classic offender), or a misconfigured SAN can all let the standby acquire a lock the active broker still believes it holds. NFS-based locking in particular is notoriously unreliable for this purpose.

This article covers how to confirm split-brain fast, how to contain it without making the corruption worse, the root causes worth checking, and how to keep it from recurring.

What this means

The shared file locker works like this: the first broker to start takes an exclusive file lock on the lock file in the KahaDB directory (using Java’s FileLock, which maps to OS-level advisory locking). That broker becomes active, starts its transport connectors, and serves clients. The second broker blocks in a retry loop, periodically attempting to acquire the same lock. When the active broker shuts down cleanly or crashes and the OS releases the lock, the standby acquires it and promotes.

Split-brain happens when the lock mechanism lies. The standby’s lock acquisition succeeds even though the active broker is still running and writing. From that point:

  • Both brokers accept producer and consumer connections.
  • Both write journal files (db-*.log) and update the index (db.data) in the same directory.
  • Producers may be load-balanced or failed over onto either broker, so the same logical message stream is written twice with interleaved journal state.
  • KahaDB’s index, which assumes a single writer, can end up pointing at journal entries written by either broker, or at overwritten state.
flowchart TD
  A[Storage disruption: NFS outage, partition, lock daemon failure] --> B[Standby retries lock acquisition]
  B --> C{Does the lock actually exclude the active broker?}
  C -->|Yes: normal failover| D[Standby promotes, old active is down]
  C -->|No: lock mechanism lies| E[Both brokers hold the lock]
  E --> F[Both accept connections and write KahaDB]
  F --> G[Duplicate delivery and ordering violations]
  F --> H[Journal and index corruption]
  H --> I[Broker may fail to restart after containment]

The dangerous property of this failure is that both brokers can look healthy in isolation. Each is accepting connections, enqueuing, and dequeuing. The failure is only visible when you look at the pair, at the lock state, or at client-level symptoms like duplicate deliveries.

Common causes

CauseWhat it looks likeFirst thing to check
NFSv4 outage and grace-period raceAfter an NFS server outage or network partition heals, both brokers become active (documented as AMQ-5549)Storage event timeline vs. lock acquisition timestamps in both broker logs
lockKeepAlivePeriod disabled (0)Active broker never notices it lost the lock; standby acquires silentlyLocker configuration in activemq.xml
Filesystem without proper POSIX locking (OCFS2, SMB/CIFS)Split-brain from first standby start, or after any blipWhat filesystem hosts the KahaDB directory (mount, df -T)
NFSv3 stale lock after abnormal terminationInverse failure: standby can never promote because a dead client’s lock is never releasedNFS server lock state; whether failover works at all
SAN misconfigurationBoth nodes see the LUN but fencing or locking is not enforcedStorage vendor fencing and multipath configuration
Keep-alive and acquire intervals misalignedKeep-alive period longer than the acquire sleep interval defeats detectionRelative values of lockKeepAlivePeriod and lockAcquireSleepInterval

Quick checks

All of these are read-only. Run them on both brokers before changing anything.

# 1. Which brokers are actually serving client traffic?
#    In a healthy pair, only one should have the transport port listening.
ss -tlnp | grep 61616

# 2. Does the broker MBean exist? Standby brokers typically do not
#    register the full MBean tree while waiting for the lock.
curl -s -u admin:admin \
  'http://localhost:8161/api/jolokia/read/org.apache.activemq:type=Broker,brokerName=localhost/BrokerId'

# 3. Search both broker logs for lock acquisition and release events.
grep -i "lock" /opt/activemq/data/activemq.log | tail -50

# 4. What filesystem is the KahaDB directory on?
df -T /opt/activemq/data/kahadb/
mount | grep -E "nfs|ocfs|gfs|cifs"

# 5. Who owns the lock file right now, and when was it touched?
ls -l /opt/activemq/data/kahadb/lock

# 6. Is the store already showing damage (journal growth from two writers)?
ls -lh /opt/activemq/data/kahadb/

# 7. Storage latency on the store device. For NFS mounts use nfsiostat
#    (ships with nfs-utils); iostat only covers block devices.
nfsiostat 1 5

The locker’s log lines (from the SharedFileLocker and LockFile sources): when a broker starts waiting for the lock it logs INFO Database <lockfile> is locked by another server. This broker is now in slave mode waiting a lock to be acquired; when keep-alive detects the lock was externally modified or removed, the master logs INFO Lock file <path>, locked at <date>, has been modified at <date> or Lock file <path>, does not exist.

A few notes on interpretation:

  • If both hosts answer on the transport port and both return a BrokerId, you have split-brain (or someone started two independent brokers against one store, which is the same problem).
  • df -T telling you the store is on OCFS2 or CIFS is itself a finding: the shared file locker does not work correctly on those filesystems. OCFS2 only supports fcntl-style locking, which is incompatible with Java’s FileLock, so both brokers can believe they hold the lock. CIFS/SMB is not supported for this use.
  • Compare lock acquisition timestamps in the two broker logs against your storage or network event timeline. The classic NFSv4 signature is: storage outage, standby promotes during the outage, storage heals, old active resumes writing without realizing it lost the lock.

How to diagnose it

  1. Confirm both brokers are active. Use checks 1 and 2 above on both hosts. One active and one waiting is normal; two active is the incident. If only one is active but clients report duplicates, check for a second broker process on the same host or an orphaned process from a failed restart.

  2. Establish the timeline. Pull lock acquisition and release events from both broker logs and line them up with NFS server logs, network device logs, or hypervisor events. You are looking for the window where the standby acquired the lock while the active was still running. This tells you the mechanism (outage race vs. bad filesystem vs. bad config) and which broker has been writing longer.

  3. Check the locker configuration. In activemq.xml, find the persistence adapter’s locker settings. lockKeepAlivePeriod defaults to 0 (disabled) and, per the official locker documentation, is not applicable before 5.9.0. If it is 0, the active broker never revalidates its lock and will never demote itself. The keep-alive mechanism exists precisely to close the “I still think I am master” window.

  4. Verify the filesystem actually enforces exclusive locks. If the store is on NFS, confirm it is NFSv4 (not v3, which has the opposite problem: locks survive abnormal client death and can block failover forever). If it is a cluster filesystem, confirm it is GFS2 or another POSIX-locking-correct option, not OCFS2. If it is a SAN LUN, confirm with your storage team that fencing is enforced.

  5. Assess store damage. Look at journal file count and modification times in the KahaDB directory. Interleaved modification times across recent journal files are consistent with two writers. The definitive test comes later: whether the surviving broker starts cleanly and passes KahaDB recovery without errors.

  6. Assess client impact. Ask the application teams whether they saw duplicate deliveries or out-of-order processing during the window. This determines whether you need application-level deduplication or replay cleanup after the broker side is fixed.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
HA role / lock state per brokerThe direct split-brain detectorBoth brokers reporting active, or full MBean tree present on both
Transport connector listeners on both nodesStandby should not accept client connectionsPort 61616 listening on both hosts simultaneously
Lock acquisition/release events in broker logThe forensic timeline and an alertable eventAcquisition on standby without a corresponding clean release on active
NFS/SAN health and store write latencyStorage disruption precedes the lock raceLatency spikes, mount hangs, retransmissions before failover events
Connection count per brokerClients spread across two brokers confirms both-activeNonzero client connections on the host that should be standby
Enqueue rate per brokerTwo writers on one storeBoth brokers showing concurrent enqueue on the same destinations
Store usage and journal file countCorruption and dual-writer growthAbnormal journal growth rate during the both-active window

Both-active detection is a PAGE. Store corruption risk makes this one of the few conditions where waking someone up at 3 a.m. is always right. Conversely, do not page on the standby broker’s transport connectors being down: that is the expected state.

Fixes

Immediate containment: stop one broker

This is disruptive and there is no way around it. While both brokers run, the store keeps getting worse. Pick a winner and stop the loser.

  • Prefer keeping the broker that acquired the lock first in the timeline you built in diagnosis, unless it is clearly unhealthy. Its journal history is more likely to be the coherent one.
  • Stop the losing broker with a normal shutdown first. Reserve kill -9 for a broker that will not stop; an unclean shutdown adds recovery work on the next start.
  • Before restarting anything, take a copy of the entire KahaDB directory if disk allows. If recovery later destroys evidence or loses messages, you will want the pre-incident state.

Do not let the stopped broker start again automatically until the root cause is fixed. Disable the service or systemd auto-restart on that node temporarily, or it may re-acquire the “lock” and recreate the split-brain.

Repair and restart the survivor

Start the surviving broker alone and watch the log through KahaDB recovery. Journal replay and index rebuild after an unclean or contested state can take minutes to hours on a large store, and the transport ports may be open before the broker accepts client connections. Do not declare recovery until a canary send and receive succeeds.

If the broker refuses to start or reports journal corruption, the ignoreMissingJournalfiles=true and checkForCorruptJournalFiles=true options can get it up, but they can lose messages. Treat that as a data-loss decision, not a routine flag.

Fix the locking configuration

If you must stay on shared-file locking, enable lock keep-alive by setting lockKeepAlivePeriod on the <kahaDB> element (it is a persistence-adapter attribute, not a locker attribute) so the active broker periodically revalidates its lock file and demotes itself if the file was lost or modified:

<persistenceAdapter>
  <kahaDB directory="/shared/kahadb" lockKeepAlivePeriod="5000">
    <locker>
      <shared-file-locker lockAcquireSleepInterval="10000"/>
    </locker>
  </kahaDB>
</persistenceAdapter>

Two rules from the AMQ-5549 discussion: lockKeepAlivePeriod of 0 disables the protection entirely, and the keep-alive period must be shorter than the acquire sleep interval (at most half is the recommended relationship), or the detection logic cannot work as intended.

Fix the storage layer

  • NFS mount options matter. The AMQ-5549 testing suggests aggressive timeouts so the client notices storage failure quickly: timeo=100,retrans=1,soft,noac. Be aware that soft mounts trade lock reliability for failure detection and can surface I/O errors to the broker mid-write; understand that tradeoff before adopting it.
  • Move off NFS if you can. NFS-based locking for this purpose is notoriously unreliable, and the JIRA record shows no mount-option combination that fully eliminated the dual-active window. A SAN LUN with proper fencing, or GFS2 for a cluster filesystem, is the supported-grade alternative. Do not use OCFS2 or CIFS/SMB.
  • Fix NFSv3 specifically if that is what you have: locks are not released on abnormal client termination, which produces the mirror-image failure where failover never happens. Recovery from a stuck NFSv3 lock can require restarting the affected ActiveMQ instances.

Consider a different locker or topology

The Lease Database Locker has been usable with the KahaDB persistence adapter since 5.9.0 (the official locker documentation notes the combination requires a <statements/> child element in the locker). Because the lease lives in a database instead of on the filesystem, it is a common community workaround for NFS-backed stores where file locking is unreliable. It has its own failure modes: it adds a database dependency to the failover path, and lease-renewal timing must be tuned so a slow database does not cause spurious failover. Validate it in a test broker before adopting it in production.

If you are on ActiveMQ Artemis rather than Classic, the internals differ: split-brain fixes such as ARTEMIS-4143 land in specific versions (2.29.0 and later for that issue), Artemis has a built-in network health check that can stop a broker during a partition, shared storage needs a filesystem with real exclusive-lock support, and duplicate node IDs are surfaced by the client log code AMQ212034 (“There are more than one servers on the network broadcasting the same node id”) — that check runs only in discovery-group topologies, not with static connectors.

Prevention

  • Alert on both-active, not just broker-down. The single most important check: periodically verify that exactly one broker in the pair has its transport connectors up and its full MBean tree registered. Two is a page.
  • Alert on lock events. Treat lock acquisition on the standby without a preceding clean shutdown of the active as a page-level event. That ordering is the split-brain signature.
  • Monitor the storage layer as a first-class dependency. NFS/SAN latency, mount health, and storage error logs belong on the same dashboard as broker health, because the storage event always precedes the lock event.
  • Test failover regularly. A planned failover in a maintenance window exercises the lock path and tells you whether your filesystem honors exclusive locking before you find out at 3 a.m.
  • Keep clocks synchronized on both broker hosts and the storage server. Your forensic timeline depends on comparable timestamps.
  • Know your filesystem’s locking semantics and document them in the runbook. “KahaDB is on NFS” should immediately raise the question “which NFS version, with which mount options, and why not SAN?”
  • Plan the application side. Assume that any both-active window produces duplicates. Idempotent consumers or deduplication keys turn a split-brain from a correctness incident into an availability incident.

How Netdata helps

  • Pair-level role visibility: tracking process, port, and JMX availability for both brokers on one dashboard makes “two actives” obvious instead of requiring someone to check each node separately.
  • Storage correlation: disk latency and I/O error signals on the KahaDB device, graphed next to broker enqueue rates, surface the NFS/SAN event that precedes the lock race.
  • Connection and enqueue asymmetry: nonzero client connections or enqueue rate on the node that should be standby is an early both-active indicator, often visible before clients report duplicates.
  • Log-based lock events: alerting on lock acquisition and release patterns in the broker log catches the failover-ordering anomaly at the moment it happens rather than after corruption.
  • Store growth signals: journal file count and disk usage on the shared partition show the abnormal dual-writer growth rate during and after an incident, which helps size the recovery effort.