<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>High Availability on Netdata</title><link>https://www.netdata.cloud/tags/high-availability/</link><description>Recent content in High Availability on Netdata</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sat, 22 Aug 2026 13:58:56 +0300</lastBuildDate><atom:link href="https://www.netdata.cloud/tags/high-availability/index.xml" rel="self" type="application/rss+xml"/><item><title>Netdata Parents For Intelligent Observability</title><link>https://www.netdata.cloud/product/netdata-parents/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/product/netdata-parents/</guid><description>Streaming aggregators that centralize data from thousands of nodes while keeping intelligence distributed at the edge. Handle million metrics per second with minimal resources, active-active clustering for high availability, and ML-powered insights without centralized bottlenecks.</description></item><item><title>Zero-Downtime Monitoring For Real-Time Deployments</title><link>https://www.netdata.cloud/features/architecture/zero-downtime-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/architecture/zero-downtime-monitoring/</guid><description>Netdata&amp;rsquo;s distributed architecture eliminates single points of failure in monitoring, delivering sub-2-second visibility with automatic failover, zero data loss, and production-safe resource overhead—ensuring complete observability during the moments that matter most.</description></item><item><title>Elasticsearch Yellow Cluster: Unassigned Shards Fix</title><link>https://www.netdata.cloud/academy/elasticsearch-yellow-cluster-access/</link><pubDate>Sun, 07 Sep 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/elasticsearch-yellow-cluster-access/</guid><description>&lt;p&gt;You run a health check on your production cluster and the result comes back: &lt;code&gt;status: yellow&lt;/code&gt;. It&amp;rsquo;s not the dreaded red status, so your application is likely still serving requests, but this is a critical warning sign. An Elasticsearch yellow cluster status is a direct indication that your data&amp;rsquo;s high availability is compromised. While all your primary shards are active, one or more replica shards have failed to be assigned to a node. Ignoring this warning can lead to data loss if another node fails, or worse, it could be a symptom of a network partition risking an Elasticsearch split-brain.&lt;/p&gt;</description></item><item><title>Consul Service Discovery Failures: Causes &amp; Fixes</title><link>https://www.netdata.cloud/academy/consul-service-discovery-failures/</link><pubDate>Wed, 03 Sep 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/consul-service-discovery-failures/</guid><description>&lt;p&gt;It’s a scenario that keeps DevOps and SRE teams up at night: your application logs fill with connection errors, services start failing, and you realize a critical component can&amp;rsquo;t find the database it depends on. The culprit? A breakdown in your service discovery mechanism. For many, that mechanism is HashiCorp Consul, the backbone of modern microservice architectures. When Consul falters, your entire ecosystem can become unstable.&lt;/p&gt;&#10;&lt;p&gt;Understanding how to diagnose these failures is crucial. The problem often lies deep within the operational layers—agent communication issues, misconfigured health checks, or disruptions in the gossip protocol that maintains cluster state. In this guide, we&amp;rsquo;ll dissect the most common causes of Consul service discovery failures, providing you with the tools to troubleshoot and resolve them. More importantly, we&amp;rsquo;ll show you how to shift from a reactive, fire-fighting mode to a proactive one, using comprehensive monitoring to build a truly resilient Consul deployment.&lt;/p&gt;</description></item><item><title>Redis Sentinel Failover &amp; Split-Brain Recovery Guide</title><link>https://www.netdata.cloud/academy/redis-cluster-split/</link><pubDate>Sun, 24 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/redis-cluster-split/</guid><description>&lt;p&gt;You&amp;rsquo;re on call. An alert fires—your Redis master node is unreachable. Your heart rate quickens. You&amp;rsquo;ve set up Redis Sentinel for high availability, but is it working? Did the failover succeed? Or worse, are you now in a Redis cluster split-brain situation where two nodes think they&amp;rsquo;re the master, leading to data inconsistency and eventual loss? In these critical moments, blindly trusting the automation isn&amp;rsquo;t enough; you need to verify what&amp;rsquo;s happening.&lt;/p&gt;</description></item><item><title>Ecommerce Infrastructure: Components &amp; Benefits</title><link>https://www.netdata.cloud/academy/ecommerce-infrastructure/</link><pubDate>Tue, 29 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/ecommerce-infrastructure/</guid><description>&lt;p&gt;The world of online shopping has exploded, offering convenience and accessibility like never before. Businesses can reach global audiences, and consumers can browse and buy from anywhere. But behind every successful online store, from small boutiques to giants like Amazon, lies a complex system working tirelessly: the &lt;strong&gt;ecommerce infrastructure&lt;/strong&gt;. Without this foundation, websites crash during peak traffic, customer data gets compromised, and orders get lost – leading to frustrated customers and lost revenue.&lt;/p&gt;</description></item><item><title>What Is Database Clustering? Types &amp; Benefits</title><link>https://www.netdata.cloud/academy/whatisdatabaseclusteringtypesbenefits/</link><pubDate>Thu, 06 Mar 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/whatisdatabaseclusteringtypesbenefits/</guid><description>&lt;p&gt;&lt;strong&gt;Database clustering&lt;/strong&gt; is a robust strategy employed to enhance the &lt;strong&gt;performance&lt;/strong&gt;, &lt;strong&gt;scalability&lt;/strong&gt;, and &lt;strong&gt;availability&lt;/strong&gt; of databases by orchestrating the distribution of data across multiple servers. This setup doesn&amp;rsquo;t just streamline handling big data; it also keeps things running smoothly, even if some nodes go down. That way, your service stays up and available no matter what happens.&lt;/p&gt;&#10;&lt;p&gt;By diving into the various &lt;strong&gt;types of database clustering architectures&lt;/strong&gt;, in this article we explain how each setup addresses specific needs and challenges within IT environments. We&amp;rsquo;ll also examine the tangible benefits that database clustering brings to businesses, from improved data redundancy to enhanced query response times.&lt;/p&gt;</description></item><item><title>6 + 1 Effective Strategies to Reduce Unplanned Downtime</title><link>https://www.netdata.cloud/academy/6-+-1-effective-strategies-to-reduce-unplanned-downtime/</link><pubDate>Wed, 30 Oct 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/6-+-1-effective-strategies-to-reduce-unplanned-downtime/</guid><description>&lt;p&gt;For DevOps and SRE teams, unplanned downtime may be a nightmare since it can ruin everything, from customer satisfaction to corporate operations. With more and more systems relying on constant availability, any unplanned downtime must be eliminated to ensure reliability of service. In the article below, we will discuss how to utilize &lt;a href="https://www.netdata.cloud/academy/what-is-infrastructure-monitoring-and-why-you-need-it/"&gt;infrastructure monitoring&lt;/a&gt;, monitoring tools, development and operations best practices in order to &lt;a href="https://www.netdata.cloud/academy/what-is-uptime-monitoring/"&gt;minimize unplanned downtime&lt;/a&gt;.&lt;/p&gt;&#10;&lt;h2 id="what-is-unplanned-downtime"&gt;What is Unplanned Downtime&lt;/h2&gt;&#10;&lt;p&gt;Unplanned downtime occurs when an application or system or infrastructure component fails without notice, causing an interruption. These disruptions could be due to a network issue, or a problem in the software or hardware or even human error. The business expenses are often substantial in terms of revenue loss and negative publicity. So, putting a plan in place to reduce downtime is critical.&lt;/p&gt;</description></item><item><title>How To Achieve High Availability In CI/CD With Observability</title><link>https://www.netdata.cloud/academy/ci-cd-high-availability/</link><pubDate>Sun, 09 Jun 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/ci-cd-high-availability/</guid><description>&lt;p&gt;Your CI/CD pipeline is the backbone of your software delivery process. When it works, code flows smoothly from commit to production. But what happens when it breaks? A failed pipeline means stalled feature releases, delayed bug fixes, and frustrated developers unable to ship their work. To prevent this, you need to treat your CI/CD infrastructure with the same rigor as your production applications, and that starts with making it highly available.&lt;/p&gt;</description></item><item><title>Netdata Parents (Streaming and Replication)</title><link>https://www.netdata.cloud/blog/netdata-parents-streaming-replication/</link><pubDate>Fri, 30 Jun 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-parents-streaming-replication/</guid><description>&lt;h2 id="what-are-they-and-why-do-we-need-them"&gt;What are they and why do we need them?&lt;/h2&gt;&#10;&lt;p&gt;A “Parent” is a Netdata Agent, like the ones we install on all our systems, but is configured as a central node that receives, stores and processes metrics data from other Netdata “Child” nodes in our infrastructure.&lt;/p&gt;&#10;&lt;p&gt;Netdata Parents are flexible. You can have one big active-active cluster of Netdata Parents, or you can spread a lot of independent Parents across the infrastructure.&lt;/p&gt;</description></item><item><title>Infinite Scalability: Monitoring Without Limits</title><link>https://www.netdata.cloud/blog/netdata-inifinite-scalability/</link><pubDate>Thu, 04 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-inifinite-scalability/</guid><description>&lt;p&gt;Scalability is crucial for monitoring systems as it ensures that they can accommodate growth, maintain performance, provide flexibility, optimize costs, enhance fault tolerance, and support informed decision-making, all of which are critical for effective infrastructure management.&lt;/p&gt;&#10;&lt;!--truncate--&gt;&#10;&lt;p&gt;Most monitoring solutions struggle with scalability, mainly because of:&lt;/p&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;&lt;strong&gt;High data volume and velocity&lt;/strong&gt;: Monitoring systems generate vast amounts of data and as the infrastructure grows, so does the volume and velocity of these data.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Resource constraints&lt;/strong&gt;: Scalability requires efficient resource utilization, leading to bottlenecks and performance issues as the monitored environment grows.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Architectural limitations&lt;/strong&gt;: Monitoring systems are usually designed with certain architectural constraints that limit their scalability. Most open source solutions rely on monolithic or centralized architectures that can become overwhelmed at scale.&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;For open source solutions scalability has always been a challenge, increasing their complexity significantly (check for example the scalability issues of Prometheus), while for commercial solutions it usually results in increased data collection to visualization latency and cost.&lt;/p&gt;</description></item><item><title>Server Uptime Monitoring: Core Benefits For High Performance</title><link>https://www.netdata.cloud/blog/server-uptime-monitoring-why-do-we-need-it/</link><pubDate>Tue, 02 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/server-uptime-monitoring-why-do-we-need-it/</guid><description>&lt;p&gt;&lt;img src="../2023-05-02-server-uptime-monitoring-why-do-we-need-it/img/stacked-netdata.png" alt="Server Uptime Monitoring: Core Benefits For High Performance"&gt;&lt;/p&gt;&#10;&lt;p&gt;Server uptime monitoring tracks the availability and reliability of servers within your infrastructure.&lt;/p&gt;&#10;&lt;!-- truncate --&gt;&#10;&lt;h2 id="what-is-server-uptime-monitoring"&gt;What Is Server Uptime Monitoring?&lt;/h2&gt;&#10;&lt;p&gt;Server uptime monitoring is the process of continuously tracking the operational status of your servers to ensure optimal performance and availability for users.&lt;/p&gt;&#10;&lt;p&gt;With Netdata, you gain access to real-time, high-resolution monitoring that goes beyond basic checks, providing a detailed overview of your entire infrastructure.&lt;/p&gt;</description></item><item><title>Why is data replication important?</title><link>https://www.netdata.cloud/blog/why-is-data-replication-important/</link><pubDate>Wed, 12 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/why-is-data-replication-important/</guid><description>&lt;p&gt;High availability. This is what every monitoring tool needs to ensure that you never compromise on IT infrastructure visibility.&lt;!--truncate--&gt; On top of high availability, do you really want to enable all available features on your production system? It is important for the monitoring tool to have a low footprint on your CPU consumption and memory usage. Let’s dive deeper into the recommended way of configuring Netdata to ensure high availability and a low resource footprint through data replication.&lt;/p&gt;</description></item><item><title>Ceph FS_DEGRADED: standby MDS failed to take over a rank</title><link>https://www.netdata.cloud/guides/ceph/ceph-fs-degraded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-fs-degraded/</guid><description>&lt;p&gt;&lt;code&gt;FS_DEGRADED&lt;/code&gt; fires when at least one CephFS rank is &lt;code&gt;failed&lt;/code&gt; or &lt;code&gt;damaged&lt;/code&gt; and a standby did not promote. Clients can usually still reach the filesystem through surviving ranks, but you are running without the failover reserve the MDS cluster was sized to provide.&lt;/p&gt;&#10;&lt;p&gt;&lt;code&gt;FS_DEGRADED&lt;/code&gt; is the precursor to &lt;code&gt;MDS_ALL_DOWN&lt;/code&gt;. If the last active rank fails before you restore a healthy standby, CephFS becomes fully unavailable. The playbook classifies &lt;code&gt;FS_DEGRADED&lt;/code&gt; active for more than 120 seconds as a TICKET. If CephFS is a primary storage interface, treat 120 seconds as the upper bound on response time, not a soft target.&lt;/p&gt;</description></item><item><title>CockroachDB clock skew cascade: how shared NTP drift causes quorum loss</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-clock-skew-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-clock-skew-cascade/</guid><description>&lt;p&gt;Multiple CockroachDB nodes crashed overnight. The logs show &amp;ldquo;clock synchronization error: this node is more than 500ms away from at least half of the known nodes.&amp;rdquo; You restart them, and they crash again. Some ranges are now unavailable. The cluster is losing quorum.&lt;/p&gt;&#10;&lt;p&gt;A shared NTP failure caused multiple nodes to drift past CockroachDB&amp;rsquo;s self-termination threshold in quick succession. Single-node clock skew is bad but recoverable. Multi-node skew from a shared NTP source can take down quorum faster than the cluster can heal.&lt;/p&gt;</description></item><item><title>CockroachDB connection storm after failover: reconnect stampedes and surviving-node overload</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-connection-storm-after-failover/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-connection-storm-after-failover/</guid><description>&lt;p&gt;When a CockroachDB node dies, every client connected to it reconnects at the same time. Without jittered backoff in client connection pools, hundreds or thousands of new connections land on the surviving nodes within seconds. The survivors are not overwhelmed by additional query load. They are overwhelmed by connection overhead: per-connection goroutines, session memory allocations, TLS handshakes, and SQL planner initialization.&lt;/p&gt;&#10;&lt;p&gt;The signature is a 2-3x spike in &lt;code&gt;sql_conns&lt;/code&gt; on surviving nodes, with a simultaneous jump in &lt;code&gt;sys_goroutines&lt;/code&gt; and &lt;code&gt;sql_mem_root_current&lt;/code&gt;. Latency rises across all nodes, not just those receiving the flood. The storm usually self-resolves as pools stabilize. But if memory pressure reaches OOM, one node failure can cascade into a multi-node outage.&lt;/p&gt;</description></item><item><title>CockroachDB node liveness failure: heartbeats, lease redistribution, and flapping</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-node-liveness-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-node-liveness-failure/</guid><description>&lt;p&gt;Lease transfer spikes, briefly unavailable ranges, and client errors such as ambiguous results or connection resets indicate node liveness failure. In the logs, nodes transition to not-live and back within seconds. When the cluster decides a node cannot renew its liveness heartbeat, it redistributes leases. If the node recovers fast enough to renew but not fast enough to stay healthy, it flaps: an oscillating state more destructive than a clean outage because it repeatedly interrupts in-flight work and prevents stable failover.&lt;/p&gt;</description></item><item><title>CockroachDB replica unavailable: lost quorum and stuck Raft groups</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-replica-unavailable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-replica-unavailable/</guid><description>&lt;p&gt;&lt;code&gt;ranges_unavailable&lt;/code&gt; has gone nonzero. Clients are seeing &amp;ldquo;replica unavailable&amp;rdquo; errors. Some portion of your keyspace cannot be read or written. This is an active availability incident.&lt;/p&gt;&#10;&lt;p&gt;In CockroachDB, every range (a ~512 MiB slice of the keyspace) is replicated across multiple nodes. A write requires quorum acknowledgment from a majority of replicas before it commits. When too few replicas are reachable, the range loses quorum and cannot serve reads or writes. The &lt;code&gt;ranges_unavailable&lt;/code&gt; metric tracks exactly this condition: ranges with no leaseholder or with lost Raft quorum.&lt;/p&gt;</description></item><item><title>CockroachDB under-replicated ranges: ranges_underreplicated and the healing margin</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-under-replicated-ranges/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-under-replicated-ranges/</guid><description>&lt;p&gt;When &lt;code&gt;ranges_underreplicated&lt;/code&gt; rises above zero, CockroachDB has ranges with fewer live replicas than the configured replication factor. The replicate queue is already working to heal them by transferring Raft snapshots to new target nodes, but until healing completes, those ranges have reduced or zero fault tolerance margin.&lt;/p&gt;&#10;&lt;p&gt;The key operational question is not whether under-replicated ranges exist (they will, transiently, after almost any node event), but whether the cluster is healing fast enough to stay ahead of the next failure. This is the healing margin: the buffer between the current replication state and the point where a range loses quorum and becomes unavailable.&lt;/p&gt;</description></item><item><title>Elasticsearch master instability: frequent elections and metadata overload</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-master-instability-flapping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-master-instability-flapping/</guid><description>&lt;p&gt;Index creation requests time out. &lt;code&gt;_cluster/health&lt;/code&gt; hangs or returns timeouts. The node listed by &lt;code&gt;_cat/master&lt;/code&gt; changes every few minutes outside planned maintenance. Shard allocation stalls, and new indices stay red or unassigned even though all data nodes are reachable. These symptoms indicate a master node that cannot keep up with cluster state updates, triggering repeated elections and leaving the cluster without stable coordination.&lt;/p&gt;&#10;&lt;p&gt;This is metadata overload. The elected master maintains the cluster state: a heap-resident data structure describing every index, shard, mapping, alias, pipeline, and node. On every change, the master serializes and publishes the state to all nodes. Updates are processed serially, so any delay in serialization, heap allocation, or node acknowledgment blocks subsequent metadata operations. When metadata churn is high or the state is oversized, the master falls behind, pending tasks accumulate, and if the master misses enough heartbeat checks, remaining master-eligible nodes trigger a new election. Until a stable master converges, writes, allocations, and administrative operations stall.&lt;/p&gt;</description></item><item><title>Elasticsearch master_not_discovered_exception: no elected master and stalled writes</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-no-master-not-discovered/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-no-master-not-discovered/</guid><description>&lt;p&gt;HTTP 503 and &lt;code&gt;master_not_discovered_exception&lt;/code&gt; mean bulk indexing, index creation, mapping updates, and shard allocation checks are being rejected. The cluster cannot process writes or administrative work. The data nodes may be healthy, but without an elected master, the cluster cannot update shard routing, publish state changes, or acknowledge document writes. The root cause is usually one of four problems inside the master-eligible cohort.&lt;/p&gt;&#10;&lt;p&gt;Elasticsearch 7.0 and later use a consensus protocol called Zen2 for master election. Only master-eligible nodes vote. The elected master maintains the cluster state, which describes every index, shard, mapping, alias, and node. That state is serialized and published to all nodes on every change. Without a master, state updates halt and the default &lt;code&gt;cluster.no_master_block&lt;/code&gt; rejects all operations until election completes.&lt;/p&gt;</description></item><item><title>Elasticsearch node left the cluster: fault detection, reallocation, and recovery</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-node-left-cluster/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-node-left-cluster/</guid><description>&lt;p&gt;Your cluster health turned yellow and &lt;code&gt;number_of_nodes&lt;/code&gt; dropped by one. The master logs a &lt;code&gt;NODE_LEFT&lt;/code&gt; event, shards are unassigned, and the remaining nodes absorb extra load. In the next minute, the allocator decides whether to move data. Misread the cause and a transient restart becomes an expensive reallocation storm, or a genuine hardware failure goes unaddressed while replicas rebalance.&lt;/p&gt;&#10;&lt;p&gt;This guide covers how Elasticsearch decides a node is gone, what happens to its shards, and how to recover without deepening the incident.&lt;/p&gt;</description></item><item><title>HAProxy backend losing servers: active server count and cascade risk</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-backend-losing-servers-capacity/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-backend-losing-servers-capacity/</guid><description>&lt;p&gt;A backend can lose half its servers and HAProxy will still report the BACKEND aggregate row as UP. As long as one server is available, the rollup status says traffic is being served, and status-based dashboards stay green. Meanwhile the surviving servers are absorbing redistributed load, their session counts are climbing, and you are one health check failure away from a full 503 outage for that backend.&lt;/p&gt;&#10;&lt;p&gt;This is the leading edge of the backend collapse cascade: initial server failure, traffic redistribution, survivor overload, survivor health check failures, more redistribution, total collapse. The operators who catch it early are watching the active server count and the load on survivors, not the aggregate status.&lt;/p&gt;</description></item><item><title>Kafka ISR shrinking: IsrShrinksPerSec, flapping, and the cascade to offline</title><link>https://www.netdata.cloud/guides/kafka/kafka-isr-shrink-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-isr-shrink-storm/</guid><description>&lt;p&gt;&lt;code&gt;IsrShrinksPerSec&lt;/code&gt; is climbing on your leaders and &lt;code&gt;UnderReplicatedPartitions&lt;/code&gt; is no longer zero. If it is flapping &amp;ndash; shrinks followed by expands every few minutes &amp;ndash; the path ends with &lt;code&gt;OfflinePartitionsCount&lt;/code&gt; rising and &lt;code&gt;acks=all&lt;/code&gt; producers throwing &lt;code&gt;NotEnoughReplicasException&lt;/code&gt;. This guide covers that path: how a lagging follower becomes a cluster-wide problem, how to separate flapping from one-way degradation, and how to stop the cascade before partitions go offline.&lt;/p&gt;&#10;&lt;h2 id="what-this-means"&gt;What this means&lt;/h2&gt;&#10;&lt;p&gt;Kafka leaders maintain an In-Sync Replica set (ISR): followers that have fetched within &lt;code&gt;replica.lag.time.max.ms&lt;/code&gt; (default 30 seconds since Kafka 2.5.0; 10 seconds before 2.5.0). When a follower stops fetching or falls behind beyond that window, the leader removes it. &lt;code&gt;IsrShrinksPerSec&lt;/code&gt; measures the velocity of these removals. &lt;code&gt;IsrExpandsPerSec&lt;/code&gt; measures replicas catching up and rejoining.&lt;/p&gt;</description></item><item><title>Kafka min.insync.replicas and acks: configuring durability you actually have</title><link>https://www.netdata.cloud/guides/kafka/kafka-min-insync-replicas-misconfigured/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-min-insync-replicas-misconfigured/</guid><description>&lt;p&gt;Most operators set producers to &lt;code&gt;acks=all&lt;/code&gt; and assume the cluster acks only when every replica has the message. It does not. With &lt;code&gt;acks=all&lt;/code&gt;, the broker waits only for the current in-sync replica set (ISR). Because the ISR shrinks dynamically when followers lag, a partition with replication factor three can have an ISR of one &amp;ndash; the leader itself. Without raising &lt;code&gt;min.insync.replicas&lt;/code&gt; from its default, the leader acks with zero followers caught up. Your durability guarantee collapses to leader-only persistence, and you only find out when the leader dies and data is missing.&lt;/p&gt;</description></item><item><title>Kafka NotEnoughReplicasException: acks=all writes rejected below min.insync.replicas</title><link>https://www.netdata.cloud/guides/kafka/kafka-not-enough-replicas-exception/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-not-enough-replicas-exception/</guid><description>&lt;p&gt;Producers are throwing &lt;code&gt;org.apache.kafka.common.errors.NotEnoughReplicasException&lt;/code&gt; or &lt;code&gt;NotEnoughReplicasAfterAppendException&lt;/code&gt;, and &lt;code&gt;acks=all&lt;/code&gt; writes are failing while &lt;code&gt;acks=1&lt;/code&gt; or &lt;code&gt;acks=0&lt;/code&gt; writes may still succeed. The affected partitions no longer have enough in-sync replicas to satisfy &lt;code&gt;min.insync.replicas&lt;/code&gt;. The immediate operational question is whether the ISR shrink is a transient recovery blip or a sustained degradation that will block writes until you fix the follower.&lt;/p&gt;&#10;&lt;h2 id="what-this-means"&gt;What this means&lt;/h2&gt;&#10;&lt;p&gt;The leader tracks followers caught up within &lt;code&gt;replica.lag.time.max.ms&lt;/code&gt; (default 30s, changed from 10s in Kafka 2.5.0) in the In-Sync Replica set (ISR). For &lt;code&gt;acks=all&lt;/code&gt;, the leader waits for all current ISR members before acknowledging the producer. If the ISR size drops below &lt;code&gt;min.insync.replicas&lt;/code&gt;, the leader rejects the produce request. The broker-level default for &lt;code&gt;min.insync.replicas&lt;/code&gt; is 1, so a lone leader can acknowledge alone. In practice, with &lt;code&gt;replication.factor=3&lt;/code&gt; and &lt;code&gt;acks=all&lt;/code&gt;, set &lt;code&gt;min.insync.replicas=2&lt;/code&gt; so a single follower loss blocks writes instead of silently weakening durability.&lt;/p&gt;</description></item><item><title>Keepalived Monitoring</title><link>https://www.netdata.cloud/monitoring-101/keepalived-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/keepalived-monitoring/</guid><description>&lt;h2 id="keepalived-monitoring"&gt;Keepalived Monitoring&lt;/h2&gt;&#10;&lt;h3 id="what-is-keepalived"&gt;What Is Keepalived?&lt;/h3&gt;&#10;&lt;p&gt;Keepalived is a robust software used primarily for high-availability and load-balancing on Linux systems. It leverages VRRP (Virtual Router Redundancy Protocol) to increase network uptime by offering failover protocols. By monitoring network interfaces, Keepalived helps in ensuring that servers are capable of handling traffic efficiently.&lt;/p&gt;&#10;&lt;h3 id="monitoring-keepalived-with-netdata"&gt;Monitoring Keepalived With Netdata&lt;/h3&gt;&#10;&lt;p&gt;Monitoring Keepalived with Netdata is a seamless process that ensures you have real-time insights into the performance and availability of your network infrastructure. Netdata uses an &lt;strong&gt;openmetrics (Prometheus) exporter&lt;/strong&gt; to retrieve Keepalived metrics. This means that, to monitor Keepalived, Netdata can accept data from any Prometheus exporter, allowing you to benefit from automated dashboards and alerts without needing a dedicated Prometheus server or Grafana. This integration allows for comprehensive monitoring of Keepalived metrics, ensuring your high-availability setup runs smoothly.&lt;/p&gt;</description></item><item><title>Kubernetes Controller-Manager Leader Election Failures</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-controller-manager-leader-election/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-controller-manager-leader-election/</guid><description>&lt;p&gt;Your Deployment has stopped scaling. Nodes cordoned hours ago are still draining. Garbage collection is paused, and orphaned volumes are not being cleaned up. The kube-controller-manager runs these reconciliation loops, and in an HA cluster only the leader performs work. When leader election fails, the controller-manager exits, and the control plane stops acting on desired state. Existing workloads keep running, but nothing new is managed.&lt;/p&gt;&#10;&lt;p&gt;The kube-controller-manager coordinates through a Lease object in the &lt;code&gt;coordination.k8s.io&lt;/code&gt; API group. The leader must renew the lease before &lt;code&gt;--leader-elect-renew-deadline&lt;/code&gt; (default 10 seconds) elapses. The lease itself expires after &lt;code&gt;--leader-elect-lease-duration&lt;/code&gt; (default 15 seconds). Renewal is attempted every &lt;code&gt;--leader-elect-retry-period&lt;/code&gt; (default 2 seconds). If a write to etcd is too slow, if the API server is saturated, if RBAC is stripped, or if the election timing is misconfigured, the leader loses the lock, logs &lt;code&gt;leaderelection lost&lt;/code&gt;, and exits. During the gap, no instance holds a valid lease, so controllers stop reconciling.&lt;/p&gt;</description></item><item><title>MongoDB no primary / election storm: repeated elections and write outages</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-no-primary-election-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-no-primary-election-storm/</guid><description>&lt;p&gt;Applications log &amp;ldquo;not primary&amp;rdquo; errors. &lt;code&gt;rs.status()&lt;/code&gt; shows a different &lt;code&gt;PRIMARY&lt;/code&gt; than thirty seconds ago. MongoDB logs repeat &lt;code&gt;&amp;quot;Starting an election&amp;quot;&lt;/code&gt; and &lt;code&gt;&amp;quot;Stepping down&amp;quot;&lt;/code&gt;. Each election costs 2-12 seconds of write unavailability. More than two in ten minutes is an election storm.&lt;/p&gt;&#10;&lt;p&gt;This pattern is more dangerous than a single failover because it creates rolling write outages that do not self-stabilize. Drivers reconnect, retry buffers fill, and application latency degrades even when a primary exists. Root causes usually fall into three categories: the primary is too slow to answer heartbeats, the network is dropping or delaying packets between members, or a misconfigured priority is forcing a healthy primary to step down.&lt;/p&gt;</description></item><item><title>MongoDB not master error: writes hitting a non-primary node after failover</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-not-master-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-not-master-error/</guid><description>&lt;p&gt;A node restart, network partition, or planned stepdown triggers a MongoDB election. Seconds later, application logs show &lt;code&gt;NotWritablePrimary&lt;/code&gt; (code 10107) or the legacy string &lt;code&gt;not master and slaveOk=false&lt;/code&gt;. Writes fail against a node that used to be PRIMARY, even though the cluster has elected a new one.&lt;/p&gt;&#10;&lt;p&gt;This guide covers how to find the root cause and stop it from recurring.&lt;/p&gt;&#10;&lt;h2 id="what-this-means"&gt;What this means&lt;/h2&gt;&#10;&lt;p&gt;MongoDB replica sets elect exactly one PRIMARY at a time. When a failover occurs, the old primary steps down and a secondary is promoted. Application drivers discover the new topology through the replica set seed list and refresh their connection pools automatically. Between stepdown and election completion, there is a brief window with no writable primary. After the new primary is elected, drivers should route writes there.&lt;/p&gt;</description></item><item><title>MongoDB not primary and secondaryOk=false: reading from a secondary and how to fix it</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-not-primary-and-secondaryok-false/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-not-primary-and-secondaryok-false/</guid><description>&lt;p&gt;Your application logs show &lt;code&gt;NotPrimaryNoSecondaryOk&lt;/code&gt; (code 13435) with the message &lt;code&gt;&amp;quot;not master and slaveOk=false&amp;quot;&lt;/code&gt;. Metrics show read failures against a specific host. The &lt;code&gt;mongod&lt;/code&gt; process is running, replica set heartbeats are clean, and replication lag looks normal. The cluster is not down. The error is a routing decision: a client sent a read to a replica set member that is not the primary, without declaring that reading from a non-primary is acceptable.&lt;/p&gt;</description></item><item><title>MongoDB rollback after failover: silent data loss and the rollback directory</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-rollback-after-failover/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-rollback-after-failover/</guid><description>&lt;p&gt;A replica set member in &lt;code&gt;ROLLBACK&lt;/code&gt; state, or an application reporting vanished documents after failover, means a former primary held writes that never reached a majority. When that node rejoins, MongoDB erases the divergent history and writes the removed data to files under &lt;code&gt;&amp;lt;dbPath&amp;gt;/rollback/&lt;/code&gt;. The application may have received acknowledgment for those writes. With &lt;code&gt;w:1&lt;/code&gt;, acknowledgment meant only that the primary applied the write. It did not guarantee replication to a majority or survival through failover. That is silent data loss.&lt;/p&gt;</description></item><item><title>MySQL GTID errant transactions: detecting replication divergence</title><link>https://www.netdata.cloud/guides/mysql/mysql-gtid-errant-transactions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-gtid-errant-transactions/</guid><description>&lt;p&gt;Before a planned failover, or after an orchestrator aborts with a GTID consistency error, replication can appear healthy: &lt;code&gt;SHOW REPLICA STATUS&lt;/code&gt; reports both threads running, &lt;code&gt;Seconds_Behind_Source&lt;/code&gt; is near zero, and the error log is quiet. Yet comparing GTID sets between source and replica reveals mismatching numbers. Errant transactions are GTIDs present in a replica&amp;rsquo;s &lt;code&gt;gtid_executed&lt;/code&gt; set that the source never generated.&lt;/p&gt;&#10;&lt;p&gt;Ordinary lag resolves as the replica catches up. Errant transactions represent true divergence. Promoting a replica with extra GTIDs propagates those transactions into the new source and downstream replicas, creating split-brain that is expensive to reverse. This guide covers distinguishing benign lag from dangerous divergence, pinpointing offending GTIDs, and deciding between empty-transaction injection and a full rebuild.&lt;/p&gt;</description></item><item><title>Patroni Monitoring</title><link>https://www.netdata.cloud/monitoring-101/patroni-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/patroni-monitoring/</guid><description>&lt;h2 id="patroni-monitoring"&gt;Patroni Monitoring&lt;/h2&gt;&#10;&lt;h3 id="what-is-patroni"&gt;What Is Patroni?&lt;/h3&gt;&#10;&lt;p&gt;Patroni is an open-source high-availability solution for PostgreSQL databases, providing automatic failover and reliable replication solutions. It is designed to be easy to configure and highly reliable, making it a popular choice for database administrators who require robust failover management.&lt;/p&gt;&#10;&lt;h3 id="monitoring-patroni-with-netdata"&gt;Monitoring Patroni With Netdata&lt;/h3&gt;&#10;&lt;p&gt;Monitoring Patroni with Netdata gives users real-time insights into their Patroni clusters. By utilizing an openmetrics (Prometheus) exporter like &lt;a href="https://github.com/gopaytech/patroni_exporter"&gt;Patroni Exporter&lt;/a&gt;, Netdata can seamlessly ingest metrics from any Prometheus exporter, allowing for automated dashboards, alerts, and more—all without the need for a Prometheus server or Grafana setup. This approach ensures that you have everything you need to monitor Patroni efficiently with minimal setup.&lt;/p&gt;</description></item><item><title>PostgreSQL replication lag: detection, diagnosis, and fixes</title><link>https://www.netdata.cloud/guides/postgres/postgres-replication-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-replication-lag/</guid><description>&lt;p&gt;Replication lag is the distance between the last WAL record generated on the primary and the last record applied on a replica. In asynchronous streaming replication, a few seconds of lag is normal. When lag grows without bound, your recovery point objective becomes fiction.&lt;/p&gt;&#10;&lt;p&gt;Lag often grows silently. Replication processes stay connected, WAL streams flow, and uptime checks stay green while the byte gap creeps from megabytes to gigabytes. Promoting a replica that is hours behind destroys the consistency your application assumes.&lt;/p&gt;</description></item><item><title>RabbitMQ network partition detected: split-brain in cluster_status</title><link>https://www.netdata.cloud/guides/rabbitmq/rabbitmq-network-partition-detected/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/rabbitmq/rabbitmq-network-partition-detected/</guid><description>&lt;p&gt;You ran &lt;code&gt;rabbitmqctl cluster_status&lt;/code&gt; (or your monitoring scraped &lt;code&gt;/api/nodes&lt;/code&gt;) and saw the &amp;ldquo;Network Partitions&amp;rdquo; section listing node names, or a non-empty &lt;code&gt;partitions&lt;/code&gt; array on one or more nodes. The cluster has split: some nodes can no longer see each other over the Erlang distribution link, and each side now has its own view of the world.&lt;/p&gt;&#10;&lt;p&gt;This is a paging condition. What happens next depends on &lt;code&gt;cluster_partition_handling&lt;/code&gt;: the two sides may be accepting writes independently and diverging (the &lt;code&gt;ignore&lt;/code&gt; default), the minority side may have frozen and stopped serving clients (&lt;code&gt;pause_minority&lt;/code&gt;), or nodes may be about to restart themselves (&lt;code&gt;autoheal&lt;/code&gt;). Each outcome has a different blast radius, and the wrong response makes it worse.&lt;/p&gt;</description></item><item><title>RabbitMQ node down: telling a dead broker apart from a partitioned one</title><link>https://www.netdata.cloud/guides/rabbitmq/rabbitmq-node-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/rabbitmq/rabbitmq-node-down/</guid><description>&lt;p&gt;A node has disappeared from your cluster view. The first question is not &amp;ldquo;how do I bring it back&amp;rdquo; but &amp;ldquo;what actually happened to it.&amp;rdquo; A dead broker (VM crashed, process killed, host gone) and a partitioned broker (Erlang node still running but cut off from its peers) look almost identical from one vantage point and completely different from another. The recovery steps are opposite: you restart a dead node, but restarting a partitioned node mid-partition can make split-brain worse.&lt;/p&gt;</description></item><item><title>RabbitMQ node is quorum critical: checking before a rolling restart</title><link>https://www.netdata.cloud/guides/rabbitmq/rabbitmq-node-quorum-critical/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/rabbitmq/rabbitmq-node-quorum-critical/</guid><description>&lt;p&gt;You are about to restart a RabbitMQ node: a rolling upgrade, an OS patch, an instance resize. Before you stop it, one question decides whether this is routine maintenance or a queue outage: if this node goes down right now, does any quorum queue or stream lose its online majority?&lt;/p&gt;&#10;&lt;p&gt;RabbitMQ ships a purpose-built check for exactly this. &lt;code&gt;rabbitmq-diagnostics check_if_node_is_quorum_critical&lt;/code&gt; returns unhealthy when stopping the target node would drop a quorum queue below the number of online members it needs to accept writes. The same logic is exposed over HTTP at &lt;code&gt;GET /api/health/checks/node-is-quorum-critical&lt;/code&gt;, which makes it usable from automation, load balancer health gates, and CI-driven upgrade pipelines.&lt;/p&gt;</description></item><item><title>Redis cluster_slots_pfail &gt; 0: impending node failure in a cluster</title><link>https://www.netdata.cloud/guides/redis/redis-cluster-slots-pfail/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-cluster-slots-pfail/</guid><description>&lt;p&gt;&lt;code&gt;cluster_slots_pfail &amp;gt; 0&lt;/code&gt; means at least one hash slot is mapped to a node that a peer suspects is down. In Redis Cluster, PFAIL is unilateral: any node raises it when another stops answering gossip PINGs for longer than &lt;code&gt;cluster-node-timeout&lt;/code&gt;. Slots continue to serve traffic; the cluster has not yet agreed the node is dead.&lt;/p&gt;&#10;&lt;p&gt;Brief spikes are expected during background saves, AOF rewrites, or any main-thread freeze. Sustained non-zero values indicate a real problem: network partition, node crash, or overload. If the majority of masters confirm the suspicion within twice &lt;code&gt;cluster-node-timeout&lt;/code&gt;, PFAIL escalates to FAIL. The affected slots become unavailable until a replica wins election. In a three-master cluster, losing two primaries leaves the survivor without quorum. The cluster enters a zombie state where no failover can proceed. Investigate PFAIL while you still have quorum and before automatic escalation.&lt;/p&gt;</description></item><item><title>Redis CLUSTERDOWN / cluster_state:fail: slot coverage and recovery</title><link>https://www.netdata.cloud/guides/redis/redis-cluster-state-fail/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-cluster-state-fail/</guid><description>&lt;p&gt;&lt;code&gt;CLUSTERDOWN The cluster is down&lt;/code&gt; means at least one of the 16384 hash slots lacks a healthy master. With &lt;code&gt;cluster-require-full-coverage yes&lt;/code&gt; (the default), a single missing slot blocks all writes. This guide covers diagnosing the root cause, recovering safely, and preventing recurrence.&lt;/p&gt;&#10;&lt;h2 id="what-this-means"&gt;What this means&lt;/h2&gt;&#10;&lt;p&gt;Redis Cluster shards the keyspace across 16384 hash slots. Each slot must be assigned to a master node that is reachable and healthy to count toward &lt;code&gt;cluster_slots_ok&lt;/code&gt;. When &lt;code&gt;cluster_slots_assigned&lt;/code&gt; drops below 16384, or &lt;code&gt;cluster_slots_fail&lt;/code&gt; becomes non-zero because a node has been marked FAIL by quorum, the cluster transitions to &lt;code&gt;cluster_state:fail&lt;/code&gt;. Clients receive &lt;code&gt;CLUSTERDOWN&lt;/code&gt; for operations hashing to affected slots.&lt;/p&gt;</description></item><item><title>Redis READONLY You can't write against a read only replica - causes and fixes</title><link>https://www.netdata.cloud/guides/redis/redis-readonly-cant-write-against-replica/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-readonly-cant-write-against-replica/</guid><description>&lt;p&gt;Your application hits &lt;code&gt;(error) READONLY You can't write against a read only replica&lt;/code&gt;. Writes fail; reads work. The connected Redis instance thinks it is a replica, so it rejects mutating commands. The replica is behaving correctly. The problem is a write-capable client routed to a node that is not the current primary. This typically happens in three situations: a routing bug that sends writes to a replica endpoint, stale client topology after a failover or upgrade, or an instance that was accidentally demoted at runtime.&lt;/p&gt;</description></item><item><title>SQL Server AG send and redo queues growing: replication lag and failover RTO</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-ag-send-redo-queue-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-ag-send-redo-queue-growing/</guid><description>&lt;p&gt;Two queues decide whether your Always On Availability Group can actually fail over: the send queue (log generated on the primary but not yet shipped to the secondary) and the redo queue (log received by the secondary but not yet replayed). When either grows without bound, replication lag is the visible symptom, but the hidden cost is failover RTO. On forced or automatic failover, the new primary must drain the entire redo queue before it accepts writes, so a queue that looks tolerable during steady state can turn a 30-second failover into a 30-minute one.&lt;/p&gt;</description></item><item><title>SQL Server Availability Group not synchronizing: NOT_HEALTHY replicas and failover risk</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-ag-not-synchronizing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-ag-not-synchronizing/</guid><description>&lt;p&gt;The symptom arrives as an alert or a dashboard color change: a synchronous-commit secondary replica is reporting &lt;code&gt;synchronization_health_desc = NOT_HEALTHY&lt;/code&gt; or &lt;code&gt;connected_state_desc = DISCONNECTED&lt;/code&gt; in &lt;code&gt;sys.dm_hadr_availability_replica_states&lt;/code&gt;. The primary is still accepting writes, but the protection you assumed is degraded or gone.&lt;/p&gt;&#10;&lt;p&gt;In synchronous-commit mode, the primary waits for the secondary to harden log records before acknowledging commits. When the secondary drops or stops keeping up, the primary either continues unprotected or stops accepting writes entirely, depending on &lt;code&gt;required_synchronized_secondaries_to_commit&lt;/code&gt;. Either way, your recovery point objective and your recovery time objective are both at risk.&lt;/p&gt;</description></item><item><title>SQL Server database in SUSPECT or RECOVERY_PENDING: an offline database and how to recover it</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-database-suspect-recovery-pending/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-database-suspect-recovery-pending/</guid><description>&lt;p&gt;A production database is showing &lt;code&gt;state_desc = SUSPECT&lt;/code&gt; or &lt;code&gt;RECOVERY_PENDING&lt;/code&gt; in &lt;code&gt;sys.databases&lt;/code&gt;. Applications cannot open connections to that database. Users are seeing login failures, query timeouts, or generic &amp;ldquo;database cannot be opened&amp;rdquo; errors. The SQL Server instance itself is up, and every other database on it may be fine.&lt;/p&gt;&#10;&lt;p&gt;&lt;code&gt;RECOVERY_PENDING&lt;/code&gt; rarely means corruption. It usually means SQL Server could not get the resources it needed during recovery: a missing file, a full log volume, a permissions change, or a transient I/O failure at startup. &lt;code&gt;SUSPECT&lt;/code&gt; is more serious because recovery actually ran and failed, but it still does not automatically mean data loss. The wrong move is to jump straight to &lt;code&gt;DBCC CHECKDB&lt;/code&gt; with &lt;code&gt;REPAIR_ALLOW_DATA_LOSS&lt;/code&gt;. The right move is to fix the underlying resource, re-run recovery, and only fall back to repair or restore when that fails.&lt;/p&gt;</description></item><item><title>SQL Server instance down: no response on port 1433 and where to look first</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-instance-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-instance-down/</guid><description>&lt;p&gt;Your availability probe just fired: TCP connect plus &lt;code&gt;SELECT 1&lt;/code&gt; against port 1433 has failed three or more times over at least 60 seconds. Before you restart anything, separate the two failure modes that get lumped together as &amp;ldquo;SQL Server is down&amp;rdquo;. They have different causes, different fixes, and different blast radii.&lt;/p&gt;&#10;&lt;p&gt;Mode one: no TCP connect at all. The listener is not accepting connections on 1433 (or the named instance&amp;rsquo;s dynamic port). The service is stopped, the host is down, the network path is broken, or the listener is misconfigured, commonly after an AlwaysOn failover. The engine is not there to talk to.&lt;/p&gt;</description></item><item><title>vSphere HA 'Insufficient resources to satisfy configured failover level': admission control</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-ha-insufficient-resources/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-ha-insufficient-resources/</guid><description>&lt;p&gt;The error &amp;ldquo;Insufficient resources to satisfy configured failover level for vSphere HA&amp;rdquo; is admission control refusing a VM power-on, vMotion, or reservation change because granting it would leave the cluster without enough spare capacity to honor the configured HA failover policy. Admission control is doing its job: protecting the restart guarantee after a host failure.&lt;/p&gt;&#10;&lt;p&gt;The cluster may physically hold more capacity than admission control lets you commit. A cluster with 500 GHz of CPU and 2 TB of RAM may only let you deploy against roughly 70% of that, with the rest held in reserve so HA can restart protected VMs after a host failure. Operators who bought hardware expecting to use all of it hit this wall during provisioning and reach for the disable switch.&lt;/p&gt;</description></item><item><title>ZooKeeper quorum loss: no leader elected and every write is failing</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-quorum-loss-no-writes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-quorum-loss-no-writes/</guid><description>&lt;p&gt;Every write to your ZooKeeper ensemble is timing out. Clients report &lt;code&gt;ConnectionLoss&lt;/code&gt; and &lt;code&gt;SessionExpired&lt;/code&gt;. Downstream systems that depend on ZK for coordination, such as Kafka controller elections or HBase region assignment, are cascading into failure. On the surviving ZK nodes, &lt;code&gt;ruok&lt;/code&gt; still returns &lt;code&gt;imok&lt;/code&gt;. The process is alive; the ensemble is not.&lt;/p&gt;&#10;&lt;p&gt;Quorum loss is ZooKeeper&amp;rsquo;s worst-case availability scenario. When fewer than &lt;code&gt;floor(N/2)+1&lt;/code&gt; voting members can communicate, no leader can be elected and every write fails. Surviving nodes sit in &lt;code&gt;LOOKING&lt;/code&gt; state, unable to make progress through ZAB.&lt;/p&gt;</description></item></channel></rss>