<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Netdata: Monitoring and troubleshooting transformed on Netdata</title><link>https://www.netdata.cloud/</link><description>Recent content in Netdata: Monitoring and troubleshooting transformed on Netdata</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Wed, 15 Jul 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://www.netdata.cloud/index.xml" rel="self" type="application/rss+xml"/><item><title>Native macOS Monitoring: Logs, Sensors, GPU &amp; Hardware Health</title><link>https://www.netdata.cloud/blog/macos-monitoring/</link><pubDate>Wed, 15 Jul 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/macos-monitoring/</guid><description>&lt;p>&lt;img src="../images/macos-monitoring.svg" alt="Native macOS monitoring with Netdata: unified logs, power, sensors, GPU, per-app metrics, storage, and network">&lt;/p>
&lt;p>We&amp;rsquo;ve overhauled macOS monitoring in the latest Netdata release. Netdata already collects system metrics on Macs at per-second resolution; this release completes the picture with logs and hardware telemetry, areas that previously required users to run CLI tools like &lt;code>log show&lt;/code> and &lt;code>powermetrics&lt;/code>. The new collectors read this data through Apple&amp;rsquo;s own frameworks, allowing users to trace application and OS errors and catch hardware issues early.&lt;/p></description></item><item><title>9 Best Container Monitoring Tools &amp; How To Choose</title><link>https://www.netdata.cloud/resources/best-container-monitoring-tools/</link><pubDate>Thu, 09 Jul 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/resources/best-container-monitoring-tools/</guid><description/></item><item><title>SolarWinds NPM Alternative: With Per-Second Metrics</title><link>https://www.netdata.cloud/solarwinds-npm-alternative/</link><pubDate>Mon, 29 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solarwinds-npm-alternative/</guid><description>Replace 5-minute SNMP polling with per-second network monitoring, ML anomaly detection, and zero-template device discovery.</description></item><item><title>SolarWinds NTA Alternative: NetFlow Traffic Analyzer</title><link>https://www.netdata.cloud/solarwinds-nta-alternative/</link><pubDate>Mon, 29 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solarwinds-nta-alternative/</guid><description>Replace SolarWinds NTA with Netdata for modern, per-second NetFlow analysis, interactive Sankey diagrams, street-level maps, and zero-config deployment.</description></item><item><title>SolarWinds Observability Alternative: Real-Time &amp; AI</title><link>https://www.netdata.cloud/solarwinds-observability-alternative/</link><pubDate>Mon, 29 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solarwinds-observability-alternative/</guid><description>Netdata is the modern alternative to SolarWinds Hybrid Cloud Observability: per-second resolution, edge-native ML on every metric, zero-config deployment, and flat per-node pricing with full data sovereignty.</description></item><item><title>SolarWinds Orion Alternative: One Agent, Full Stack</title><link>https://www.netdata.cloud/solarwinds-orion-alternative/</link><pubDate>Mon, 29 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solarwinds-orion-alternative/</guid><description>Replace SolarWinds Orion Platform with Netdata: one agent, per-second visibility, ML anomaly detection, AI troubleshooting, and network flow analysis without stacked module licenses.</description></item><item><title>SolarWinds SAM Alternative: Server &amp; App Monitoring</title><link>https://www.netdata.cloud/solarwinds-sam-alternative/</link><pubDate>Mon, 29 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solarwinds-sam-alternative/</guid><description/></item><item><title>Fleet Observability: Linux Edge Device Monitoring</title><link>https://www.netdata.cloud/blog/fleet-observability/</link><pubDate>Sun, 28 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/fleet-observability/</guid><description>&lt;p>It feels less like managing devices and more like remote babysitting. You check the dashboard, everything is green, and then a customer in the field tells you a device has been down for two days. At a handful of servers, the rare failure is an event. Across thousands of distributed Linux endpoints — robots in warehouses, EV chargers across a city, kiosks in retail, IoT gateways in the field — the rare failure becomes a daily occurrence, and the tools built for a datacenter quietly stop telling you the truth.&lt;/p></description></item><item><title>Distributed &amp; Edge Fleet Monitoring</title><link>https://www.netdata.cloud/solutions/use-cases/edge-fleet-monitoring/</link><pubDate>Thu, 25 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/edge-fleet-monitoring/</guid><description>Netdata brings real-time, AI-powered observability to distributed fleets of Linux endpoints — robots, kiosks, POS terminals, IoT gateways, and remote servers — keeping data and intelligence at the edge where it belongs.</description></item><item><title>EV Charging Network Monitoring</title><link>https://www.netdata.cloud/solutions/industries/ev-charging/</link><pubDate>Thu, 25 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/industries/ev-charging/</guid><description>Netdata gives charge point operators per-second visibility into every charger controller, with edge-resident storage and streaming that survives flaky site connectivity.</description></item><item><title>Retail POS and Kiosk Fleet Monitoring</title><link>https://www.netdata.cloud/solutions/industries/pos-kiosk-monitoring/</link><pubDate>Thu, 25 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/industries/pos-kiosk-monitoring/</guid><description>Netdata gives per-second visibility into every POS terminal and self-service kiosk across thousands of stores, with edge-resident monitoring that survives flaky in-store networks.</description></item><item><title>Robot Fleet Monitoring with Per-Second Edge Intelligence</title><link>https://www.netdata.cloud/solutions/industries/robotics/</link><pubDate>Thu, 25 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/industries/robotics/</guid><description>Netdata runs on each robot, collecting per-second metrics for compute, memory, storage, thermals, and network — keeping thousands of devices observable even on intermittent site networks, without centralized SaaS.</description></item><item><title>School Safety Device Fleet Monitoring</title><link>https://www.netdata.cloud/solutions/industries/education-device-monitoring/</link><pubDate>Thu, 25 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/industries/education-device-monitoring/</guid><description>Netdata runs on every campus device, giving IT teams per-second visibility into compute, connectivity, and OS health across thousands of school safety devices and endpoints.</description></item><item><title>What Is Edge Monitoring? Edge Observability Explained</title><link>https://www.netdata.cloud/academy/what-is-edge-monitoring/</link><pubDate>Thu, 25 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-edge-monitoring/</guid><description>&lt;p>Edge monitoring means running the full observability pipeline - collection, storage, anomaly detection, alerting, and dashboards - on or near the monitored device itself, instead of shipping all raw telemetry to a central system first. Each node in the fleet collects and processes its own metrics locally, and only forwards what is needed upstream. This model is essential for distributed fleets of thousands of devices where centralizing all telemetry is economically and technically impractical.&lt;/p></description></item><item><title>NetFlow vs sFlow vs IPFIX: Differences &amp; When To Use</title><link>https://www.netdata.cloud/academy/netflow-sflow-ipfix/</link><pubDate>Wed, 24 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/netflow-sflow-ipfix/</guid><description>&lt;p>NetFlow, sFlow, and IPFIX are three network traffic export protocols that give visibility into traffic flows without requiring full packet capture. NetFlow and IPFIX aggregate packets into flow records on the device before exporting them, providing exact (if unsampled) accounting. sFlow takes a fundamentally different approach: it exports every Nth packet header plus interface counters, making it stateless and scalable but statistical rather than exact. IPFIX is the IETF-standardized evolution of NetFlow v9, adding vendor-neutral extensibility through template-based information elements.&lt;/p></description></item><item><title>Network Monitoring Dashboard</title><link>https://www.netdata.cloud/features/network/dashboard/</link><pubDate>Wed, 24 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/network/dashboard/</guid><description>Netdata&amp;rsquo;s Network Monitoring Dashboard unifies SNMP device metrics, NetFlow/sFlow traffic analysis, SNMP traps, and live topology into a single real-time, interactive view.</description></item><item><title>Network Topology Mapping Explained</title><link>https://www.netdata.cloud/academy/network-topology-mapping/</link><pubDate>Wed, 24 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/network-topology-mapping/</guid><description>&lt;p>Network topology mapping is the process of discovering and visualizing how devices, endpoints, and services connect across a network. It captures both physical (Layer 2) relationships - which switch port links to which device - and logical (Layer 3) relationships - how subnets, routes, and autonomous systems reach each other. A topology map can range from a hand-drawn diagram to a live, continuously updated graph built from protocol data such as LLDP, CDP, ARP, FDB, OSPF, and BGP.&lt;/p></description></item><item><title>Real Time Network Monitoring: Topology, NetFlow, SNMP</title><link>https://www.netdata.cloud/blog/network-monitoring/</link><pubDate>Wed, 24 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/network-monitoring/</guid><description>&lt;p>Interface counters tell you a port is busy. Bytes in, bytes out, errors, drops. That&amp;rsquo;s enough to know a link is saturated, but not enough to know which conversations are saturating it, which devices are involved, or how a problem propagates across your network. For that you&amp;rsquo;ve traditionally needed dedicated network performance monitoring tools, usually expensive, usually a separate console from the rest of your monitoring.&lt;/p>
&lt;p>Today we&amp;rsquo;re closing that gap. Netdata has had solid network interface monitoring for a long time through its native collectors and SNMP support. We&amp;rsquo;ve now built out the rest of the picture, and it adds up to NPM-class network monitoring: live network topology, NetFlow and sFlow traffic analysis, SNMP device monitoring across 200+ vendor profiles, SNMP trap handling, and a dedicated network monitoring dashboard. All of it runs alongside the infrastructure, application, and container metrics Netdata already collects, on the same timeline, in the same platform.&lt;/p></description></item><item><title>SNMP Monitoring Guide: How To Monitor Network Devices</title><link>https://www.netdata.cloud/academy/snmp-monitoring-guide/</link><pubDate>Wed, 24 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/snmp-monitoring-guide/</guid><description>&lt;p>To monitor SNMP devices you enable an SNMP agent on the device, open UDP/161 from the monitoring host, point your monitoring tool at the device with the correct credentials, and poll interface and health counters on a fixed interval. SNMP works by having a central manager query an agent that exposes numeric objects (OIDs) defined in MIBs, with optional push notifications (traps) for event-driven signals. The bulk of day-to-day network monitoring (interface utilization, errors, port state, CPU, memory) comes from periodic polling of those OIDs.&lt;/p></description></item><item><title>SNMP Traps vs Polling: What Is The Difference?</title><link>https://www.netdata.cloud/academy/snmp-traps-vs-polling/</link><pubDate>Wed, 24 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/snmp-traps-vs-polling/</guid><description>&lt;p>SNMP polling is a pull model where a monitoring server periodically queries a device over UDP port 161 to read metrics and state. SNMP traps are a push model where the device itself sends an unsolicited notification over UDP port 162 the instant an event occurs. The two are not competing choices: mature network monitoring uses polling for continuous metrics and traps for immediate event alerts.&lt;/p>
&lt;h2 id="what-is-snmp-polling">What is SNMP polling?&lt;/h2>
&lt;p>Polling is the classic request-response model at the heart of SNMP-based monitoring. A central manager (your monitoring system) sends a GET, GETNEXT, or GETBULK request to an SNMP agent running on a router, switch, UPS, or other networked device. The agent responds with the requested values.&lt;/p></description></item><item><title>What Is NetFlow? Network Flow Monitoring Explained</title><link>https://www.netdata.cloud/academy/what-is-netflow/</link><pubDate>Wed, 24 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-netflow/</guid><description>&lt;p>NetFlow is a network protocol, originally developed by Cisco, for collecting metadata about IP traffic flows as they pass through a router, switch, or firewall. A flow is a unidirectional sequence of packets that share the same key fields - classically the 5-tuple of source IP, destination IP, source port, destination port, and protocol. NetFlow does not capture packet payloads; it exports compact metadata records that a collector stores and analyzes, giving you visibility into who is talking to whom, how much bandwidth they are using, and when.&lt;/p></description></item><item><title>5 Best SolarWinds Alternatives for 2026</title><link>https://www.netdata.cloud/blog/solarwinds-alternatives-2026/</link><pubDate>Tue, 23 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/solarwinds-alternatives-2026/</guid><description>&lt;p>As organizations modernize their infrastructure and embrace cloud-native architectures, traditional monitoring solutions are showing their age. SolarWinds, while a long-established player in IT management, was designed for an era of static, on-premise infrastructure. In 2026, teams are seeking alternatives that can keep pace with dynamic, distributed systems—and they&amp;rsquo;re finding better options.&lt;/p>
&lt;h2 id="what-is-solarwinds-is-it-still-the-right-choice">What Is SolarWinds? Is It Still The Right Choice?&lt;/h2>
&lt;p>SolarWinds Platform (formerly Orion) has been a cornerstone of IT monitoring for decades, offering comprehensive coverage of network devices, servers, applications, and databases. It&amp;rsquo;s particularly strong in traditional enterprise environments with extensive hardware monitoring needs.&lt;/p></description></item><item><title>NetFlow Traffic Analyzer</title><link>https://www.netdata.cloud/features/network/netflow-traffic-analyzer/</link><pubDate>Tue, 23 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/network/netflow-traffic-analyzer/</guid><description>Real-time NetFlow, IPFIX, and sFlow analysis with top talkers, Sankey diagrams, geographic traffic maps, and automatic GeoIP/ASN/NetBox enrichment — built into the Netdata Agent.</description></item><item><title>Network Topology Viewer</title><link>https://www.netdata.cloud/features/network/topology-viewer/</link><pubDate>Tue, 23 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/network/topology-viewer/</guid><description>Netdata Network Topology Viewer maps live TCP/UDP connections between processes, containers, and endpoints, plus SNMP device fabric via LLDP, CDP, BGP, and OSPF.</description></item><item><title>SNMP &amp; Network Device Monitoring</title><link>https://www.netdata.cloud/features/network/network-device-monitoring/</link><pubDate>Tue, 23 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/network/network-device-monitoring/</guid><description>Real-time SNMP device monitoring with auto-matched vendor profiles, ML anomaly detection, and zero-config discovery.</description></item><item><title>SNMP Trap Monitoring</title><link>https://www.netdata.cloud/features/network/snmp-traps/</link><pubDate>Tue, 23 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/network/snmp-traps/</guid><description>Native SNMP trap receiver with MIB-resolved decoding, storm controls, health alerts, and SIEM forwarding — built into Netdata.</description></item><item><title>SolarWinds Price Increases 2026: What Customers Need to Know</title><link>https://www.netdata.cloud/blog/solarwinds-price-increases-2026/</link><pubDate>Tue, 23 Jun 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/solarwinds-price-increases-2026/</guid><description>&lt;p>If you&amp;rsquo;re a SolarWinds customer facing renewal, you&amp;rsquo;ve likely noticed significant changes to pricing and licensing terms in 2024-2025. You&amp;rsquo;re not alone. At Netdata, we&amp;rsquo;ve been speaking with dozens of SolarWinds customers who are reassessing their monitoring strategies in light of these changes. This post provides factual information about what&amp;rsquo;s changed, the real impact on organizations, and a practical framework for evaluating your path forward.&lt;/p>
&lt;h2 id="whats-happening-with-solarwinds-pricing">What&amp;rsquo;s Happening with SolarWinds Pricing?&lt;/h2>
&lt;p>In February 2025, SolarWinds was acquired by private equity firm Turn/River Capital in a $4.4 billion transaction. As is common with PE-backed acquisitions, this has led to significant changes in pricing and business terms. Based on customer reports and public information, renewal prices have increased by 100-300% for many customers. One customer on the SolarWinds community forum reported their renewal more than doubled, a 225% increase from the previous year.&lt;/p></description></item><item><title>Grafana vs Datadog: The honest comparison in 2026</title><link>https://www.netdata.cloud/comparisons/grafana-vs-datadog/</link><pubDate>Sat, 30 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/grafana-vs-datadog/</guid><description/></item><item><title>High Cardinality Metrics At Scale: A Better Playbook</title><link>https://www.netdata.cloud/blog/high-cardinality-metrics-observability-scale/</link><pubDate>Sat, 30 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/high-cardinality-metrics-observability-scale/</guid><description>&lt;p>The &amp;ldquo;high cardinality is expensive&amp;rdquo; sentence has become observability&amp;rsquo;s version of &amp;ldquo;in this economy&amp;rdquo;: said so often that nobody questions whether it&amp;rsquo;s true. Every vendor pricing page invokes it. Every glossary article repeats it. Every architecture diagram shows aggregation buffers placed &lt;em>before&lt;/em> the storage layer.&lt;/p>
&lt;!--truncate-->
&lt;p>&amp;ldquo;High cardinality is expensive&amp;rdquo; is not a fact about the universe; it&amp;rsquo;s a fact about one architectural choice: centralizing time-series storage and querying it through an index that scales with unique series. Once you accept that choice, everything follows. You pay per metric, you drop labels you wish you&amp;rsquo;d kept, you pre-aggregate before storage, and you discover that the bug you were debugging only existed at the full resolution you already threw away.&lt;/p></description></item><item><title>How To Reduce Alert Fatigue With Anomaly Detection</title><link>https://www.netdata.cloud/solutions/use-cases/alert-fatigue/</link><pubDate>Sat, 30 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/alert-fatigue/</guid><description>Most alert-fatigue tools manage noise after the fact. Netdata prevents it by running 18 ML models per metric with consensus voting — anomalies only fire when multiple models agree, suppressing the false positives that drive on-call burnout.</description></item><item><title>The best AI-powered observability platforms in 2026</title><link>https://www.netdata.cloud/resources/best-ai-powered-observability-platforms/</link><pubDate>Sat, 30 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/resources/best-ai-powered-observability-platforms/</guid><description/></item><item><title>The best DevOps monitoring tools in 2026</title><link>https://www.netdata.cloud/resources/best-devops-monitoring-tools/</link><pubDate>Sat, 30 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/resources/best-devops-monitoring-tools/</guid><description/></item><item><title>The Best Infrastructure Monitoring Tools In 2026</title><link>https://www.netdata.cloud/resources/best-infrastructure-monitoring-tools/</link><pubDate>Sat, 30 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/resources/best-infrastructure-monitoring-tools/</guid><description/></item><item><title>The best open-source observability tools in 2026</title><link>https://www.netdata.cloud/resources/best-open-source-observability-tools/</link><pubDate>Sat, 30 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/resources/best-open-source-observability-tools/</guid><description/></item><item><title>Netdata Skills: Teach Your AI Coding Agent To Monitor</title><link>https://www.netdata.cloud/blog/netdata-skills/</link><pubDate>Wed, 20 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-skills/</guid><description>&lt;p>There&amp;rsquo;s a growing ecosystem of AI coding agents: Claude Code, Cursor, Copilot, Codex, Gemini CLI, Windsurf, and others. They&amp;rsquo;re good at writing code, but they don&amp;rsquo;t inherently know how to instrument that code for observability, configure monitoring infrastructure, or troubleshoot production systems using real telemetry data. That knowledge lives in documentation, runbooks, and the heads of your senior SREs.&lt;/p>
&lt;p>We&amp;rsquo;ve open-sourced a repository that encodes this knowledge into a format AI agents can use directly. &lt;a href="https://github.com/netdata/skills">netdata/skills&lt;/a> is a collection of agent skills, published in the open &lt;a href="https://agentskills.io">agentskills.io&lt;/a> format, that teach AI coding agents how to set up Netdata, instrument applications with OpenTelemetry, build collector pipelines, troubleshoot 49 specific technologies, and verify everything against live data via MCP.&lt;/p></description></item><item><title>OpenTelemetry Backend</title><link>https://www.netdata.cloud/opentelemetry/</link><pubDate>Wed, 20 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/opentelemetry/</guid><description>An open OTEL backend with native OTLP ingestion, per-second granularity, ML anomaly detection, and infrastructure correlation. Traces coming soon. No per-metric pricing. No vendor lock-in.</description></item><item><title>OpenTelemetry and Netdata, Today</title><link>https://www.netdata.cloud/blog/opentelemetry-metrics-and-logs-ingestion/</link><pubDate>Fri, 15 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/opentelemetry-metrics-and-logs-ingestion/</guid><description>&lt;p>OpenTelemetry has become the default way to instrument applications and ship telemetry. The hard part has never been the data model. It&amp;rsquo;s been picking a backend that handles OTLP without quietly turning into a per-metric bill or a black box that swallows your data.&lt;/p>
&lt;p>Netdata is a native OTLP backend. Stand up an OpenTelemetry Collector with any of its hundreds of receivers, point its OTLP exporter at Netdata, and you get per-second charts, ML anomaly detection on every signal, AI-assisted troubleshooting, and infrastructure correlation, with no per-metric, per-series, or per-host charges. Metrics and logs work today. Trace support is coming soon.&lt;/p></description></item><item><title>Dashboard Playlists: Cycle Through Dashboards in TV Mode</title><link>https://www.netdata.cloud/blog/dashboard-playlists/</link><pubDate>Tue, 12 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/dashboard-playlists/</guid><description>&lt;p>When we shipped TV mode, we heard almost immediately: &amp;ldquo;Great, but I have five dashboards and one screen.&amp;rdquo; A single dashboard on a wall display covers one view of your infrastructure. If you want to rotate between your network overview, database health, application metrics, and infrastructure summary, someone has to walk over and click, or you&amp;rsquo;re buying more screens.&lt;/p>
&lt;p>Dashboard playlists solve this. You can now select a sequence of dashboards to cycle through in TV mode, with a configurable rotation interval. Set it up once, open the TV mode URL on your display, and the screen rotates through your chosen dashboards on its own.&lt;/p></description></item><item><title>Azure Local Migration: Monitor Both Sides In One View</title><link>https://www.netdata.cloud/blog/azure-local-migration/</link><pubDate>Mon, 11 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/azure-local-migration/</guid><description>&lt;p>&lt;img src="../images/azure-local-migration.svg" alt="Monitoring Your Azure to Azure Local Migration: One Dashboard for Both Sides">&lt;/p>
&lt;p>More organizations are moving workloads from Azure public cloud to Azure Local (formerly Azure Stack HCI) than most people realize. The reasons vary: data sovereignty requirements, latency-sensitive workloads that need to be closer to the edge, cost optimization for predictable workloads where reserved cloud capacity doesn&amp;rsquo;t make financial sense, or regulatory constraints that require data to stay on-premises.&lt;/p></description></item><item><title>Geo Maps: See Where Your Infrastructure Lives</title><link>https://www.netdata.cloud/blog/geo-maps/</link><pubDate>Sat, 09 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/geo-maps/</guid><description>&lt;p>When your infrastructure is spread across regions, data centers, branch offices, or edge locations, knowing where a node is physically located matters more than people usually admit. During an incident, &amp;ldquo;the node in the Singapore POP&amp;rdquo; communicates faster than a hostname. When you&amp;rsquo;re planning capacity, seeing geographic clustering tells you something that a flat list of nodes doesn&amp;rsquo;t. When a subset of your fleet starts misbehaving, the first question is often &amp;ldquo;is this regional?&amp;rdquo;&lt;/p></description></item><item><title>NVIDIA DCGM Collector: Deep GPU Monitoring For AI</title><link>https://www.netdata.cloud/blog/nvidia-dcgm-monitoring/</link><pubDate>Mon, 04 May 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/nvidia-dcgm-monitoring/</guid><description>&lt;p>&lt;img src="../images/dcgm-collector.svg" alt="NVIDIA DCGM Collector: Deep GPU Monitoring for Data Center and AI Infrastructure">&lt;/p>
&lt;p>GPU infrastructure is expensive and increasingly central to production workloads. Whether you&amp;rsquo;re running ML training jobs, inference serving, video transcoding, or HPC workloads, understanding what your GPUs are actually doing, and what&amp;rsquo;s going wrong when performance degrades, is not optional. The problem is that NVIDIA&amp;rsquo;s Data Center GPU Manager (DCGM) exposes an enormous amount of telemetry, but getting that data into a monitoring system in a useful, organized way has traditionally required significant setup and custom dashboarding work.&lt;/p></description></item><item><title>Misconfigured Alert Detection: Tuning Made Easy</title><link>https://www.netdata.cloud/blog/identifying-misconfigured-alerts/</link><pubDate>Tue, 28 Apr 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/identifying-misconfigured-alerts/</guid><description>&lt;p>&lt;img src="../images/miscon-alerts-1.png" alt="Misconfigured Alert Detection: Find the Alerts That Need Tuning">&lt;/p>
&lt;p>Netdata ships with hundreds of stock alerts. They cover a wide range of infrastructure conditions and they&amp;rsquo;re designed with sensible defaults. But &amp;ldquo;sensible defaults&amp;rdquo; and &amp;ldquo;correct for your environment&amp;rdquo; are not the same thing. A CPU threshold that&amp;rsquo;s perfectly reasonable for a build server might generate constant noise on a machine running batch jobs. An alert that&amp;rsquo;s critical for production might be irrelevant in staging, where it fires daily and everyone ignores it.&lt;/p></description></item><item><title>Azure Monitor Collector: Monitor Azure Infrastructure</title><link>https://www.netdata.cloud/blog/azure-monitor/</link><pubDate>Mon, 27 Apr 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/azure-monitor/</guid><description>&lt;p>&lt;img src="../images/azure-monitor-collector.svg" alt="Azure Monitor Collector: Monitor Your Entire Azure Infrastructure From Netdata">&lt;/p>
&lt;p>If you&amp;rsquo;re running infrastructure on Azure, you&amp;rsquo;ve probably dealt with the split between your Azure-native monitoring and the rest of your stack. Your VMs, databases, and Kubernetes clusters generate platform metrics through Azure Monitor, but those metrics live in a separate world from the OS-level, application, and on-prem metrics you&amp;rsquo;re already watching in Netdata. You end up checking two (or more) places during incidents, building mental bridges between dashboards that don&amp;rsquo;t talk to each other.&lt;/p></description></item><item><title>Database Performance Monitoring: 14+ DBs Supported</title><link>https://www.netdata.cloud/blog/dbm/</link><pubDate>Fri, 24 Apr 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/dbm/</guid><description>&lt;p>&lt;img src="../images/dbm-hero.svg" alt="Database Performance Monitoring: Query-Level Visibility Across 14+ Databases">&lt;/p>
&lt;p>Netdata has always collected database metrics: connections, throughput, replication lag, buffer cache hit ratios, and so on. These tell you that something is wrong, but they don&amp;rsquo;t tell you why. When your PostgreSQL response time spikes, the metric alone doesn&amp;rsquo;t tell you which query is responsible. For that, you&amp;rsquo;ve traditionally needed to SSH into the box, connect to the database, and run diagnostic queries manually. Or set up a separate database monitoring tool entirely.&lt;/p></description></item><item><title>Nagios Plugins: Run Existing Checks &amp; Custom Scripts</title><link>https://www.netdata.cloud/blog/nagios-plugins/</link><pubDate>Wed, 22 Apr 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/nagios-plugins/</guid><description>&lt;p>A lot of teams have a collection of Nagios plugins and custom monitoring scripts that have been running reliably for years. Some are standard community plugins for checking disk health or SSL certificate expiry. Others are homegrown Bash or Python scripts that check something very specific to the business: whether an API endpoint returns the right payload, whether a batch job completed on time, whether a queue depth is within bounds. These scripts work, they&amp;rsquo;re battle-tested, and nobody wants to rewrite them.&lt;/p></description></item><item><title>Secrets Management: Remove Credentials From Configs</title><link>https://www.netdata.cloud/blog/secrets-management/</link><pubDate>Mon, 20 Apr 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/secrets-management/</guid><description>&lt;p>If you&amp;rsquo;re running Netdata collectors that connect to databases, APIs, or other authenticated services, there&amp;rsquo;s a good chance you have passwords sitting in plain-text configuration files right now. It works, but it&amp;rsquo;s the kind of thing that makes security teams nervous and makes credential rotation painful. Every password change means editing config files and restarting collectors.&lt;/p>
&lt;p>Netdata now supports secrets management natively. Instead of putting credentials directly in your collector configurations, you reference them using a resolver syntax, and Netdata resolves the actual values at runtime from whatever source you choose: environment variables, files on disk, the output of a command, or a centralized secret store like HashiCorp Vault or AWS Secrets Manager.&lt;/p></description></item><item><title>Netdata Complete Tutorial by LearnLinux.tv</title><link>https://www.netdata.cloud/resources/netdata-tutorial/</link><pubDate>Thu, 16 Apr 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/resources/netdata-tutorial/</guid><description>LearnLinux.tv&amp;rsquo;s Jay LaCroix built a complete video course on Netdata. Eight episodes taking you from first install to production best practices.</description></item><item><title>Smarter Alerts: Test, Review &amp; Preview Schedules</title><link>https://www.netdata.cloud/blog/smarter-alerts/</link><pubDate>Thu, 16 Apr 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/smarter-alerts/</guid><description>&lt;p>Alert fatigue usually isn&amp;rsquo;t caused by one thing. It&amp;rsquo;s the accumulation of thresholds that are slightly too sensitive, alerts that fire during known maintenance windows, and historical patterns that nobody has the tools to review easily. Fixing it requires better visibility into how alerts actually behave over time, and a way to test changes before they hit production.&lt;/p>
&lt;p>We&amp;rsquo;ve shipped three improvements to alerting in Netdata that address different parts of this problem: the ability to evaluate alert definitions against historical data before deploying them, a timeline view of alert transitions, and a schedule preview for recurring silencing rules.&lt;/p></description></item><item><title>TV Mode: Put Your Dashboards on the Big Screen</title><link>https://www.netdata.cloud/blog/tv-mode/</link><pubDate>Tue, 14 Apr 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/tv-mode/</guid><description>&lt;p>One of the most common requests we&amp;rsquo;ve gotten since launching custom dashboards is deceptively simple: &amp;ldquo;How do I put this on a TV?&amp;rdquo; Teams want their dashboards on wall-mounted screens in NOCs, war rooms, and open office spaces. The dashboard is already built. The data is already there. They just need a way to display it on a screen that nobody is logged into, without exposing the full Netdata Cloud interface.&lt;/p></description></item><item><title>New Custom Dashboards: Metrics, Logs &amp; Live Commands</title><link>https://www.netdata.cloud/blog/new-custom-dashboards/</link><pubDate>Sun, 12 Apr 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/new-custom-dashboards/</guid><description>&lt;p>Custom dashboards in Netdata have always let you pull charts together on-the-fly into a single view. That&amp;rsquo;s useful, but it&amp;rsquo;s also limited. In practice, when you&amp;rsquo;re running an incident or reviewing a service, you don&amp;rsquo;t just want charts. You want to see the output of &lt;code>top&lt;/code> alongside your CPU metrics. You want slow query logs next to your database latency charts. You want an infrastructure summary card that tells you how many nodes in a room are healthy without having to click through to find out.&lt;/p></description></item><item><title>Alert Acknowledgement: Mark It as Seen, Keep Working</title><link>https://www.netdata.cloud/blog/alert-acknowledge/</link><pubDate>Fri, 10 Apr 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/alert-acknowledge/</guid><description>&lt;p>If you&amp;rsquo;ve ever opened the alerts tab during a busy period, you know the problem. There are alerts you&amp;rsquo;ve already looked at, alerts someone on your team is handling, and alerts that fired on a known issue that&amp;rsquo;s being worked on. They all sit together in the same list alongside the new ones you haven&amp;rsquo;t seen yet. There&amp;rsquo;s no way to say &amp;ldquo;I&amp;rsquo;ve seen this, move on&amp;rdquo; without silencing or disabling the alert entirely, which is a much heavier action than the situation calls for.&lt;/p></description></item><item><title>Expanded Chart View: Investigate Without Leaving the Chart</title><link>https://www.netdata.cloud/blog/charts-expanded-view/</link><pubDate>Wed, 08 Apr 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/charts-expanded-view/</guid><description>&lt;p>Charts in Netdata have always been interactive. You can zoom, pan, select time ranges, and see per-second granularity across thousands of metrics. But when you spotted something interesting, the next steps usually meant leaving the chart: opening another tab to check a related metric, navigating to the correlation tool, or pulling up a different time range for comparison. The investigation workflow lived outside the chart, even though the chart was where the investigation started.&lt;/p></description></item><item><title>Monitor Your Azure to Azure Local Migration</title><link>https://www.netdata.cloud/solutions/use-cases/azure-local-migration/</link><pubDate>Wed, 08 Apr 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/azure-local-migration/</guid><description>Netdata gives migration teams a single pane of glass across Azure cloud resources and Azure Local cluster nodes. Per-second metrics, edge-based ML, and 800+ collectors mean you can baseline source workloads, validate destination hardware, run parallel cutovers, and operate the new environment without changing tools.</description></item><item><title>Conversations: Ask Netdata About Anything You're Looking At</title><link>https://www.netdata.cloud/blog/converse-with-everything-in-netdata/</link><pubDate>Thu, 02 Apr 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/converse-with-everything-in-netdata/</guid><description>&lt;p>Netdata AI can already troubleshoot your alerts and generate Insights reports. What it couldn&amp;rsquo;t do, until now, was have a back-and-forth conversation. You could get a one-shot analysis, but you couldn&amp;rsquo;t ask follow-up questions, pull in additional context, or go from a quick question to a full investigation without starting over.&lt;/p>
&lt;p>We&amp;rsquo;ve added a conversational layer to Netdata AI. You&amp;rsquo;ll notice a new blue chat icon throughout Netdata Cloud, on charts, in the alerts table, on Insights reports, and in the reports list. Click it, and you&amp;rsquo;re in a conversation where the thing you clicked on is already the context. No copy-pasting metric names, no explaining what you&amp;rsquo;re looking at.&lt;/p></description></item><item><title>Node Groups: Organize Infrastructure Into Views</title><link>https://www.netdata.cloud/blog/node-groups/</link><pubDate>Wed, 01 Apr 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/node-groups/</guid><description>&lt;p>When you&amp;rsquo;re managing a handful of nodes, the flat list in the nodes tab works fine. When you&amp;rsquo;re managing hundreds or thousands, it becomes a wall of hostnames. You end up applying the same filters repeatedly: all the production database servers, all the nodes in eu-west, all the Kubernetes workers in the staging cluster. The filters work, but they don&amp;rsquo;t persist, and there&amp;rsquo;s no way to share them with the rest of your team.&lt;/p></description></item><item><title>Netdata Product Roadmap</title><link>https://www.netdata.cloud/roadmap/</link><pubDate>Fri, 13 Mar 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/roadmap/</guid><description>Netdata&amp;rsquo;s product roadmap and strategic investment areas. From AI-native observability to full-stack signal coverage, see where Netdata is headed.</description></item><item><title>Introducing the Netdata Cloud MCP Server</title><link>https://www.netdata.cloud/blog/netdata-cloud-mcp-server/</link><pubDate>Fri, 27 Feb 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-cloud-mcp-server/</guid><description>&lt;p>The Netdata Cloud MCP Server is now available — giving AI agents and assistants direct access to your Netdata through a single endpoint at &lt;code>app.netdata.cloud/api/v1/mcp&lt;/code>.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="ai-is-changing-how-we-monitor-infrastructure">AI Is Changing How We Monitor Infrastructure&lt;/h2>
&lt;p>If you&amp;rsquo;re an engineer in 2026, chances are AI is already part of your daily workflow, whether that&amp;rsquo;s a general-purpose assistant like ChatGPT, Claude, or Gemini that you bounce questions off, or a coding agent like &lt;a href="https://docs.anthropic.com/en/docs/claude-code">Claude Code&lt;/a>, &lt;a href="https://openai.com/index/codex/">Codex&lt;/a>, &lt;a href="https://www.cursor.com/">Cursor&lt;/a>, or &lt;a href="https://windsurf.com/">Windsurf&lt;/a> that writes and debugs code alongside you. These tools are incredibly powerful, but until now, they&amp;rsquo;ve been blind to what&amp;rsquo;s actually happening on your infrastructure.&lt;/p></description></item><item><title>Netdata for Homelabs</title><link>https://www.netdata.cloud/homelab/</link><pubDate>Wed, 18 Feb 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/homelab/</guid><description>Netdata&amp;rsquo;s Homelab plan gives you enterprise-grade monitoring at a price that respects your hobby. Per-second metrics, ML anomaly detection, unlimited custom dashboards, and zero configuration — for $90/year.</description></item><item><title>Howard Conference &amp; Expo 2026: Smarter Observability</title><link>https://www.netdata.cloud/blog/howard-expo-2026/</link><pubDate>Tue, 03 Feb 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/howard-expo-2026/</guid><description>&lt;p>The Netdata team will be at the &lt;strong>Howard Conference and Expo &amp;ldquo;Game On&amp;rdquo;&lt;/strong> event, &lt;strong>February 24-26, 2026 at the Grand Hotel Marriott Resort in Fairhope, Alabama&lt;/strong>. We&amp;rsquo;re looking forward to meeting IT leaders and practitioners to talk about real-time observability—what it actually looks like in practice, and where traditional monitoring falls short.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="what-well-be-showing">What We&amp;rsquo;ll Be Showing&lt;/h2>
&lt;p>Stop by our booth to see Netdata in action and chat with our team about what you&amp;rsquo;re dealing with in your own environment.&lt;/p></description></item><item><title>India DevOps Show 2026: Modern Observability Recap</title><link>https://www.netdata.cloud/blog/india-devops-show-2026/</link><pubDate>Tue, 03 Feb 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/india-devops-show-2026/</guid><description>&lt;p>DevOps has fundamentally transformed how organizations build and deliver software. But as deployment velocity increases and infrastructure becomes more dynamic, the gap between shipping code and truly understanding system behavior continues to widen. Teams need observability that keeps pace with their pipelines, not tools that slow them down or break the budget.&lt;/p>
&lt;p>Netdata is proud to participate as a &lt;strong>Silver Partner&lt;/strong> at the &lt;strong>10th Edition India DevOps Show 2026&lt;/strong>, taking place on &lt;strong>February 13, 2026 at Aloft ORR Hotel, Bengaluru&lt;/strong>. We&amp;rsquo;re excited to engage with India&amp;rsquo;s vibrant DevOps community and share our vision for efficient, intelligent observability.&lt;/p></description></item><item><title>Tech Show London 2026: Cloud &amp; AI Observability Recap</title><link>https://www.netdata.cloud/blog/techshow-london-2026/</link><pubDate>Tue, 03 Feb 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/techshow-london-2026/</guid><description>&lt;p>The intersection of cloud and AI is creating unprecedented infrastructure complexity. As organizations race to deploy AI workloads alongside traditional cloud services, the demand for intelligent, high-fidelity observability has never been greater. Understanding what&amp;rsquo;s happening across your entire stack, in real time, is no longer a luxury, it&amp;rsquo;s a necessity.&lt;/p>
&lt;p>That&amp;rsquo;s why the Netdata team is excited to be part of &lt;strong>Tech Show London 2025&lt;/strong>, taking place &lt;strong>March 4-5 at ExCeL London&lt;/strong>. We&amp;rsquo;ll be in the &lt;strong>Cloud &amp;amp; AI Infrastructure&lt;/strong> zone, ready to show you how modern observability should work.&lt;/p></description></item><item><title>Database Monitoring Software Without Query Languages</title><link>https://www.netdata.cloud/solutions/built-for/dbas/</link><pubDate>Tue, 27 Jan 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/built-for/dbas/</guid><description>Netdata gives DBAs complete database visibility with per-second query performance, replication lag tracking, lock analysis, and connection pool monitoring across MySQL, PostgreSQL, SQL Server, Oracle, and MongoDB - all from one dashboard with zero configuration.</description></item><item><title>Netdata vs ScienceLogic | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/sciencelogic/</link><pubDate>Wed, 14 Jan 2026 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/sciencelogic/</guid><description>Comprehensive comparison of Netdata and ScienceLogic Skylar One platforms, highlighting real-time per-second visibility, transparent pricing, and zero-configuration deployment advantages over enterprise-focused 5-minute polling.</description></item><item><title>Netdata vs SolarWinds | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/solarwinds/</link><pubDate>Tue, 30 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/solarwinds/</guid><description/></item><item><title>Introducing Real-Time Conversations with Netdata AI</title><link>https://www.netdata.cloud/blog/ai-conversations/</link><pubDate>Tue, 23 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/ai-conversations/</guid><description>&lt;p>Over the past few months, we&amp;rsquo;ve seen incredible adoption of our AI Investigations and Insights reports. Teams are using them to automate the deep, thoughtful analysis required for complex post-mortems, capacity planning, and performance optimization. These comprehensive reports are fantastic when you need a well-researched, shareable document.&lt;/p>
&lt;p>But what about the moments &lt;em>during&lt;/em> an investigation? What about the rapid-fire &amp;ldquo;what if&amp;rdquo; questions and the quick exploration of hypotheses that happen in the heat of the moment? For that, you need speed and interactivity. You need a partner you can have a real-time dialogue with.&lt;/p></description></item><item><title>Update Your Netdata Agent</title><link>https://www.netdata.cloud/please-update-your-agent/</link><pubDate>Fri, 19 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/please-update-your-agent/</guid><description/></item><item><title>About Us</title><link>https://www.netdata.cloud/about/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/about/</guid><description>Meet the team behind Netdata - engineers who transformed observability by distributing intelligence to the edge, making infrastructure transparent for teams worldwide.</description></item><item><title>Access Control: Secure Observability Without Risk</title><link>https://www.netdata.cloud/features/enterprise/access-control/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/enterprise/access-control/</guid><description>Secure your observability infrastructure with Netdata&amp;rsquo;s edge-native access control: granular RBAC, automated SSO provisioning, true multi-tenancy through Spaces and Rooms, comprehensive audit logging, and SOC 2 Type 2 certification - all while keeping your data on your infrastructure.</description></item><item><title>AI Co-Engineer For Instant Root Cause Insights</title><link>https://www.netdata.cloud/features/aiml/ai-co-engineer/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/aiml/ai-co-engineer/</guid><description>Netdata&amp;rsquo;s AI Co-Engineer combines edge-native machine learning with flexible AI integration, providing instant expert-level insights while keeping your data sovereign and secure.</description></item><item><title>AI Reporting Software For Observability Data</title><link>https://www.netdata.cloud/features/aiml/reporting/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/aiml/reporting/</guid><description>Generate professional infrastructure reports in minutes, not hours. Netdata&amp;rsquo;s AI reporting delivers automated insights across metrics, logs, and alerts - from capacity planning to root cause analysis - with zero configuration and complete data sovereignty.</description></item><item><title>AIOps Platform: Edge-Native ML, Zero Configuration</title><link>https://www.netdata.cloud/features/aiml/aiops/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/aiml/aiops/</guid><description>Enterprise AIOps intelligence without enterprise complexity. Edge-native ML, automated insights, and transparent pricing deliver operational excellence from day one.</description></item><item><title>Algorithmic Dashboards For Every Metric</title><link>https://www.netdata.cloud/features/architecture/algorithmic-dashboards/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/architecture/algorithmic-dashboards/</guid><description>Netdata&amp;rsquo;s algorithmic dashboards transform observability from a months-long configuration project into instant, comprehensive visibility. Every metric visualized automatically, every relationship revealed through intelligent point-and-click analysis, every engineer productive from day one.</description></item><item><title>Anomaly Advisor: Root Cause In Seconds, Not Hours</title><link>https://www.netdata.cloud/features/aiml/anomaly-advisor/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/aiml/anomaly-advisor/</guid><description>Netdata Anomaly Advisor uses edge-native machine learning with 18-model consensus to eliminate 99% of false positives while surfacing root causes in the top 30-50 metrics from thousands collected. Get sub-2-second correlation analysis at any scale without configuration, training delays, or specialist expertise.</description></item><item><title>Anomaly Detection Software: 99% Fewer False Positives</title><link>https://www.netdata.cloud/features/aiml/anomaly-detection/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/aiml/anomaly-detection/</guid><description>Production-ready ML anomaly detection from installation. 18 consensus models per metric achieve 99% false positive reduction in anomaly detection while catching issues competitors miss. Zero configuration. Zero false promises.</description></item><item><title>Anomaly Detection With 99% Fewer False Positives</title><link>https://www.netdata.cloud/features/aiml/machine-learning/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/aiml/machine-learning/</guid><description>Netdata trains 18 independent ML models per metric at the edge, achieving 99% false positive reduction in anomaly detection through unanimous consensus - all included at no additional cost.</description></item><item><title>Application Performance Monitoring Software</title><link>https://www.netdata.cloud/solutions/use-cases/application-performance/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/application-performance/</guid><description>Netdata revolutionizes application performance monitoring with distributed edge intelligence, zero-configuration deployment, and AI-powered troubleshooting that delivers enterprise-grade observability at a fraction of traditional APM costs.</description></item><item><title>AWS Monitoring Software With 1 Second Visibility</title><link>https://www.netdata.cloud/solutions/technologies/aws-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/technologies/aws-monitoring/</guid><description>Netdata delivers true real-time AWS monitoring with 1-second granularity, ML-powered anomaly detection, and predictable per-node pricing - eliminating CloudWatch&amp;rsquo;s complexity and cost unpredictability.</description></item><item><title>Azure Monitoring Tool With Real-Time Visibility</title><link>https://www.netdata.cloud/solutions/technologies/azure-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/technologies/azure-monitoring/</guid><description>Netdata delivers true real-time Azure monitoring with per-second granularity, 90% lower costs, and zero learning curve. Monitor Azure VMs, containers, databases, and applications with automated dashboards, ML anomaly detection, and AI-powered root cause analysis.</description></item><item><title>Blast Radius Detection For Faster Incident Response</title><link>https://www.netdata.cloud/features/aiml/blast-radius-detection/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/aiml/blast-radius-detection/</guid><description>Netdata reveals blast radius dynamically through real-time anomaly correlation and ML-powered pattern recognition, showing the complete story from first failure to full impact in seconds.</description></item><item><title>Built-In MCP For AI Assistants To Query Metrics &amp; Logs</title><link>https://www.netdata.cloud/features/aiml/mcp/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/aiml/mcp/</guid><description>Transform infrastructure troubleshooting with AI assistants that understand your systems. Netdata&amp;rsquo;s built-in MCP server provides real-time metrics, ML anomaly detection, logs, and live system state - all accessible through natural language queries.</description></item><item><title>Cloud Monitoring Software With Real-Time Metrics &amp; AI</title><link>https://www.netdata.cloud/solutions/use-cases/cloud-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/cloud-monitoring/</guid><description>Transform cloud operations with Netdata&amp;rsquo;s distributed observability platform. Get per-second insights, AI-powered troubleshooting, and predictable costs - all while keeping your data sovereign.</description></item><item><title>Cloud On-Premises Solution With Complete Observability</title><link>https://www.netdata.cloud/product/cloud-on-premises/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/product/cloud-on-premises/</guid><description>Self-hosted Netdata Cloud for air-gapped environments and regulated industries. Complete control plane within your datacenter, zero external dependencies, full feature parity with SaaS. SOC 2 Type 2 certified observability that meets DORA, NIS2, HIPAA, and FedRAMP requirements by design.</description></item><item><title>Container Monitoring Software With Zero Configuration</title><link>https://www.netdata.cloud/solutions/use-cases/container-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/container-monitoring/</guid><description>Transform container monitoring with Netdata&amp;rsquo;s edge-native platform. Get per-second metrics, automated dashboards, ML anomaly detection, and AI troubleshooting - all without complex setup or unpredictable costs.</description></item><item><title>Custom Dashboards With Real-Time Precision</title><link>https://www.netdata.cloud/features/visualization/custom-dashboards/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/visualization/custom-dashboards/</guid><description>Netdata&amp;rsquo;s algorithmic dashboards eliminate the false choice between pre-built rigidity and custom complexity. Each chart is a complete analytical tool equivalent to 25+ Grafana charts, providing 360° views with simple point-and-click. No query languages. No dashboard building. No maintenance burden.</description></item><item><title>Data Center Monitoring Software With Real-Time Metrics</title><link>https://www.netdata.cloud/solutions/use-cases/datacenters/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/datacenters/</guid><description>Transform data center operations with Netdata&amp;rsquo;s real-time monitoring platform. Per-second metrics, automated ML anomaly detection, and AI-powered troubleshooting deliver 80% faster MTTR at 90% lower cost than traditional solutions.</description></item><item><title>Data Sovereignty With Metadata-Only Cloud</title><link>https://www.netdata.cloud/features/enterprise/data-sovereignty/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/enterprise/data-sovereignty/</guid><description>Netdata delivers complete observability while keeping every metric and log exactly where you need it: on your infrastructure, in your jurisdiction, under your control. Meet GDPR, NIS2, DORA, and HIPAA requirements through architecture, not configuration.</description></item><item><title>Database Monitoring Software With Real-Time Visibility</title><link>https://www.netdata.cloud/solutions/use-cases/database-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/database-monitoring/</guid><description>Real-time database monitoring with AI-powered troubleshooting, zero configuration, and predictable costs. Monitor 15+ database platforms with per-second granularity and ML-based anomaly detection.</description></item><item><title>DevOps Monitoring Software With ML Anomaly Detection</title><link>https://www.netdata.cloud/solutions/built-for/devops/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/built-for/devops/</guid><description>Netdata delivers complete observability for DevOps teams with per-second metrics, zero-configuration deployment, and AI-powered root cause analysis. Monitor everything from bare metal to Kubernetes with a single platform that replaces 7 tools while reducing costs by 90%.</description></item><item><title>Distributed Observability For Scale, Speed &amp; Control</title><link>https://www.netdata.cloud/features/dataplatform/metrics-management/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/dataplatform/metrics-management/</guid><description>Collect unlimited metrics at per-second granularity with zero configuration. Netdata&amp;rsquo;s edge-native architecture eliminates cardinality cost traps while delivering ML-based anomaly detection on every metric.</description></item><item><title>Distributed Observability Without Centralized Pipelines</title><link>https://www.netdata.cloud/features/architecture/distributed-observability-data-pipeline/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/architecture/distributed-observability-data-pipeline/</guid><description>Netdata&amp;rsquo;s distributed architecture processes observability data at the edge, eliminating centralized bottlenecks while delivering per-second insights, automatic ML anomaly detection, and 90% cost reduction compared to traditional pipeline solutions.</description></item><item><title>eBPF Network Monitoring With Real-Time Insights</title><link>https://www.netdata.cloud/features/dataplatform/ebpf-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/dataplatform/ebpf-monitoring/</guid><description>Get kernel-level network insights - bandwidth, connections, retransmissions - embedded in your infrastructure observability platform. Zero eBPF expertise required. Zero configuration. Zero surprise bills.</description></item><item><title>Edge Observability Without Cloud Dependency</title><link>https://www.netdata.cloud/features/architecture/edge-computing/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/architecture/edge-computing/</guid><description>Monitor thousands of edge nodes with per-second precision, &amp;lt;5% CPU overhead, and 90% cost reduction. Netdata&amp;rsquo;s edge-native architecture delivers complete visibility with autonomous operation, local ML, and zero configuration.</description></item><item><title>Eliminate Log Pipelines: Native Log Query At Scale</title><link>https://www.netdata.cloud/features/dataplatform/zero-pipeline-logs/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/dataplatform/zero-pipeline-logs/</guid><description>Netdata&amp;rsquo;s Zero Pipeline Logs architecture eliminates traditional log aggregation pipelines by leveraging native system formats. Query logs directly where they live - no shipping, parsing, or indexing required - achieving 90% cost reduction and sub-2-second latency at any scale.</description></item><item><title>Enterprise Monitoring At Scale Without Bottlenecks</title><link>https://www.netdata.cloud/features/architecture/infinite-scalability/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/architecture/infinite-scalability/</guid><description>Experience observability that grows naturally with your infrastructure - no rewrites, no bottlenecks, no compromise. Monitor everything, everywhere, at per-second resolution with predictable costs that scale with nodes, not data volume.</description></item><item><title>Financial Transaction Monitoring Software</title><link>https://www.netdata.cloud/solutions/industries/finance/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/industries/finance/</guid><description>Transform financial services operations with per-second monitoring, AI-powered troubleshooting, and zero-pipeline logs. Built for compliance, optimized for performance, designed for lean teams.</description></item><item><title>Gaming Infrastructure Monitoring &amp; Observability</title><link>https://www.netdata.cloud/solutions/industries/gaming/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/industries/gaming/</guid><description>Netdata delivers true real-time monitoring for gaming infrastructure with per-second metrics, edge-based ML anomaly detection, and AI-powered troubleshooting - enabling lean teams to maintain 99.9% uptime while reducing monitoring costs by 90%.</description></item><item><title>GCP Monitoring Tool With Per-Second Visibility</title><link>https://www.netdata.cloud/solutions/technologies/gcp-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/technologies/gcp-monitoring/</guid><description>Netdata delivers true real-time GCP monitoring with per-second granularity, zero-configuration deployment, and predictable per-node pricing - solving the cost unpredictability and complexity that plague organizations.</description></item><item><title>Hetzner Monitoring With Real-Time Observability</title><link>https://www.netdata.cloud/solutions/technologies/hetzner-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/technologies/hetzner-monitoring/</guid><description>Complete monitoring solution for Hetzner infrastructure with 1-second granularity, automated hardware health tracking, and AI-powered troubleshooting. Eliminates native monitoring gaps while reducing costs by 90%.</description></item><item><title>High Cardinality Protection For Unlimited Metrics</title><link>https://www.netdata.cloud/features/architecture/extreme-cardinality/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/architecture/extreme-cardinality/</guid><description>Automated multi-layer protection handles extreme cardinality at the edge, enabling unlimited observability without manual tuning or cost explosions.</description></item><item><title>HPC Monitoring Software With Per-Second Metrics</title><link>https://www.netdata.cloud/solutions/use-cases/hpc/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/hpc/</guid><description>Transform HPC operations with distributed edge-native monitoring that delivers sub-2-second insights, 90% cost reduction, and linear scalability to 100,000+ nodes - without the complexity.</description></item><item><title>Hybrid Cloud Monitoring Solution At 90% Lower Cost</title><link>https://www.netdata.cloud/solutions/use-cases/hybrid-cloud-observability/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/hybrid-cloud-observability/</guid><description>Transform hybrid cloud monitoring with Netdata&amp;rsquo;s distributed edge-native architecture. Get per-second visibility, ML-powered insights, and 90% cost savings without pipelines, sampling, or vendor lock-in.</description></item><item><title>Infrastructure Health Monitoring Software</title><link>https://www.netdata.cloud/solutions/industries/healthcare/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/industries/healthcare/</guid><description>Netdata delivers HIPAA-compliant, real-time observability for healthcare with 90% cost savings, per-second monitoring, and zero-configuration deployment. Monitor EHRs, medical devices, and critical infrastructure with edge-based ML and complete data sovereignty.</description></item><item><title>Infrastructure Monitoring Software With AI &amp; ML</title><link>https://www.netdata.cloud/solutions/use-cases/infrastructure-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/infrastructure-monitoring/</guid><description>Transform infrastructure monitoring with Netdata&amp;rsquo;s edge-native platform. Get per-second visibility, ML anomaly detection on every metric, and AI-powered troubleshooting - all while keeping your data sovereign and reducing costs by 90%.</description></item><item><title>Infrastructure Observability Without Instrumentation</title><link>https://www.netdata.cloud/features/enterprise/zero-code-instrumentation/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/enterprise/zero-code-instrumentation/</guid><description>Complete infrastructure visibility without touching code. Netdata monitors processes, containers, networks, and logs automatically through kernel-level instrumentation and intelligent auto-discovery.</description></item><item><title>Instant Observability UI For Every Engineer</title><link>https://www.netdata.cloud/product/netdata-ui/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/product/netdata-ui/</guid><description>Algorithmic dashboards, ML-powered insights, and point-and-click analysis that turns 100,000+ metrics into actionable intelligence in seconds.</description></item><item><title>IoT Monitoring Solution With Full Data Sovereignty</title><link>https://www.netdata.cloud/solutions/use-cases/iot-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/iot-monitoring/</guid><description>Real-time IoT monitoring with per-device metrics, ML-powered anomaly detection, and transparent pricing. Monitor thousands of sensors, gateways, and edge devices with &amp;lt;2s latency.</description></item><item><title>Join Us</title><link>https://www.netdata.cloud/join-us/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/join-us/</guid><description>Build infrastructure observability at massive scale with Netdata. Remote-first culture, open source core, 76K+ GitHub stars. Own the full stack, ship code used by millions, work with world-class engineers.</description></item><item><title>Kubernetes Monitoring Tool Without Centralized Data</title><link>https://www.netdata.cloud/solutions/technologies/kubernetes-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/technologies/kubernetes-monitoring/</guid><description>Transform Kubernetes observability with Netdata&amp;rsquo;s edge-native architecture. Get per-second visibility, ML anomaly detection on every metric, and AI-powered root cause analysis - all at 90% lower cost than traditional solutions.</description></item><item><title>Linux Monitoring Software With Per-Second Metrics</title><link>https://www.netdata.cloud/solutions/technologies/linux-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/technologies/linux-monitoring/</guid><description>Netdata delivers complete Linux monitoring with per-second granularity, edge-based ML anomaly detection, and zero-pipeline logs - all with zero configuration and predictable per-node pricing.</description></item><item><title>Live Tab - Real-Time Netdata Functions</title><link>https://www.netdata.cloud/features/visualization/live/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/visualization/live/</guid><description>The Live tab provides real-time access to Netdata Functions - specialized routines from collectors that deliver live information from your monitored nodes, including database monitoring, network topology maps, process explorers, and diagnostics.</description></item><item><title>LLM Monitoring Platform With Per-Second Visibility</title><link>https://www.netdata.cloud/solutions/use-cases/llm-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/llm-monitoring/</guid><description>Real-time infrastructure monitoring for LLM deployments. Track GPU utilization, container resources, database performance, and system metrics with per-second granularity and ML-powered anomaly detection.</description></item><item><title>Log Management Software With 90% Cost Reduction</title><link>https://www.netdata.cloud/features/dataplatform/logs-management/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/dataplatform/logs-management/</guid><description>Eliminate expensive log pipelines and centralized clusters. Netdata queries systemd-journal and Windows Event Logs directly at the edge, delivering 90% cost savings, sub-second queries, and complete data sovereignty - all while maintaining superior analysis accuracy than traditional platforms.</description></item><item><title>Manufacturing Monitoring Software Without Downtime</title><link>https://www.netdata.cloud/solutions/industries/manufacturing/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/industries/manufacturing/</guid><description>Transform manufacturing operations with distributed observability that processes data at the edge. Monitor production metrics, predict equipment failures, and troubleshoot issues 80% faster - all while keeping costs predictable and data sovereign.</description></item><item><title>Microservices Monitoring Tool Without Sampling</title><link>https://www.netdata.cloud/solutions/use-cases/microservices-observability/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/microservices-observability/</guid><description>Netdata delivers comprehensive microservices observability with per-second metrics, ML-based anomaly detection, and AI-powered troubleshooting - all without the complexity, cost overruns, or tool sprawl that plague traditional solutions.</description></item><item><title>Monitoring Agents For Accurate Production Insights</title><link>https://www.netdata.cloud/product/netdata-agents/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/product/netdata-agents/</guid><description>Discover how Netdata Agents revolutionize infrastructure monitoring with edge-native intelligence, per-second granularity, and zero-configuration deployment. Complete observability in 60 seconds.</description></item><item><title>MSP Monitoring Software With Multi-Tenant Isolation</title><link>https://www.netdata.cloud/solutions/built-for/msp/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/built-for/msp/</guid><description>Netdata delivers the most cost-effective, scalable observability platform for MSPs - solving tool sprawl, alert fatigue, and unpredictable costs with true real-time monitoring, physical multi-tenancy, and edge-based ML.</description></item><item><title>Multi-Cloud Monitoring Tools With Predictable Pricing</title><link>https://www.netdata.cloud/solutions/use-cases/multi-cloud-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/multi-cloud-monitoring/</guid><description>Break free from fragmented multi-cloud monitoring. Netdata delivers unified per-second visibility across AWS, Azure, GCP, and on-premises with edge-native ML, zero configuration, and transparent per-node pricing - eliminating tool sprawl and surprise bills.</description></item><item><title>Multi-Tenant Observability Access Without Data Egress</title><link>https://www.netdata.cloud/features/dataplatform/alerts-notifications/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/dataplatform/alerts-notifications/</guid><description>Edge-intelligent alerting that delivers component-level precision through 400+ pre-configured templates, zero-configuration deployment, and sub-2-second evaluation latency - all while maintaining visibility during network partitions.</description></item><item><title>Native iOS &amp; Android Apps With AI Troubleshooting</title><link>https://www.netdata.cloud/product/mobile-apps/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/product/mobile-apps/</guid><description>Transform on-call monitoring with Netdata&amp;rsquo;s mobile apps for iOS and Android. Get push notifications for infrastructure alerts, AI-powered root cause analysis, and real-time dashboard access - all at a fixed, predictable price that includes mobile access for your entire team.</description></item><item><title>Netdata Parents For Intelligent Observability</title><link>https://www.netdata.cloud/product/netdata-parents/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/product/netdata-parents/</guid><description>Streaming aggregators that centralize data from thousands of nodes while keeping intelligence distributed at the edge. Handle million metrics per second with minimal resources, active-active clustering for high availability, and ML-powered insights without centralized bottlenecks.</description></item><item><title>Netdata vs Amazon CloudWatch | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/cloudwatch/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/cloudwatch/</guid><description>Netdata provides 90% cost reduction vs CloudWatch with superior real-time monitoring (1-second vs 1-60 minutes), comprehensive ML on all metrics, and zero-configuration simplicity. Deploy in 60 seconds and eliminate CloudWatch&amp;rsquo;s unpredictable costs while gaining better visibility.</description></item><item><title>Netdata vs Azure Monitor | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/azuremonitor/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/azuremonitor/</guid><description/></item><item><title>Netdata vs BetterStack | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/betterstack/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/betterstack/</guid><description>Netdata and BetterStack serve different observability needs. Netdata excels at deep infrastructure monitoring with per-second metrics and ML intelligence, while BetterStack focuses on incident coordination and uptime monitoring. Learn which solution fits your technical requirements.</description></item><item><title>Netdata vs CardinalHQ | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/cardinalhq/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/cardinalhq/</guid><description>Comprehensive comparison of Netdata&amp;rsquo;s edge-native monitoring platform versus CardinalHQ&amp;rsquo;s AI-powered observability optimization layer. Learn which solution fits your infrastructure needs.</description></item><item><title>Netdata vs Catchpoint | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/catchpoint/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/catchpoint/</guid><description>Netdata delivers real-time infrastructure visibility with per-second metrics, ML-based anomaly detection, and AI-powered troubleshooting - the essential internal monitoring layer that external testing platforms cannot provide.</description></item><item><title>Netdata vs Checkmk | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/checkmk/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/checkmk/</guid><description>Netdata vs Checkmk comparison: Real-time monitoring (1-second vs 60-second), zero configuration (vs 700+ rule sets), AI-powered troubleshooting, 90% cost reduction, and operational simplicity. See why cloud-native teams choose Netdata for modern infrastructure monitoring.</description></item><item><title>Netdata vs Coralogix | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/coralogix/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/coralogix/</guid><description>Netdata and Coralogix serve complementary observability needs. While Coralogix excels at centralized log analytics and SIEM, Netdata provides unmatched real-time infrastructure monitoring with per-second granularity, zero-configuration deployment, and transparent per-node pricing. Learn how Netdata addresses common Coralogix pain points including steep learning curves, slow production rollouts, unpredictable costs, and UI performance issues.</description></item><item><title>Netdata vs Coroot | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/coroot/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/coroot/</guid><description>Netdata vs Coroot comparison: production-ready observability platform vs emerging Kubernetes tool. Learn why Netdata provides superior maturity, universal coverage, and operational simplicity for real-world infrastructure.</description></item><item><title>Netdata vs Cribl | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/cribl/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/cribl/</guid><description>Netdata provides real-time infrastructure monitoring with ML-based anomaly detection and built-in dashboards. Cribl optimizes log pipelines and routes data for cost efficiency. Learn how these complementary solutions work together for complete observability.</description></item><item><title>Netdata vs Dash0 | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/dash0/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/dash0/</guid><description>Netdata vs Dash0: Real-time edge intelligence with predictable pricing versus OpenTelemetry-native centralized monitoring. Compare features, costs, and capabilities to choose the right observability platform.</description></item><item><title>Netdata vs Datadog | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/datadog/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/datadog/</guid><description/></item><item><title>Netdata vs Dynatrace | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/dynatrace/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/dynatrace/</guid><description>Netdata provides superior infrastructure monitoring at a fraction of Dynatrace&amp;rsquo;s cost, with true per-second granularity and complete data sovereignty. Learn when to use each platform and how a hybrid approach delivers complete observability at significantly lower total cost.</description></item><item><title>Netdata vs EdgeDelta | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/edgedelta/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/edgedelta/</guid><description>Complete edge-native observability vs telemetry pipelines. See how Netdata provides real-time monitoring, ML-powered insights, and automated dashboards without the complexity of data routing infrastructure.</description></item><item><title>Netdata vs ELK Stack | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/elk/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/elk/</guid><description>A comprehensive comparison between Netdata and the ELK Stack for infrastructure monitoring and observability. Learn which solution fits your real-time monitoring, log management, and troubleshooting needs.</description></item><item><title>Netdata vs Fluentd | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/fluentd/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/fluentd/</guid><description>Netdata provides complete real-time observability with metrics, logs, ML anomaly detection, and AI troubleshooting in a single platform. Fluentd excels at log collection and routing. Learn how they complement each other and when Netdata replaces entire monitoring stacks.</description></item><item><title>Netdata vs Grafana | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/grafana/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/grafana/</guid><description>Comprehensive comparison of Netdata and Grafana monitoring platforms, highlighting real-time performance, cost efficiency, and operational simplicity.</description></item><item><title>Netdata vs Graylog | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/graylog/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/graylog/</guid><description>Netdata and Graylog serve different but complementary purposes in modern observability stacks. Netdata provides real-time infrastructure metrics monitoring with ML-based anomaly detection, while Graylog offers centralized log management for security analysis and forensic investigations. Together, they deliver complete observability.</description></item><item><title>Netdata vs Grepr | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/grepr/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/grepr/</guid><description/></item><item><title>Netdata vs Guance | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/guance/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/guance/</guid><description/></item><item><title>Netdata vs HetrixTools | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/hetrixtools/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/hetrixtools/</guid><description>When HetrixTools alerts you&amp;rsquo;re down, Netdata shows you why. Compare external uptime monitoring with internal infrastructure observability powered by ML and AI.</description></item><item><title>Netdata vs Honeycomb | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/honeycomb/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/honeycomb/</guid><description>Compare Netdata and Honeycomb for infrastructure monitoring. Learn how Netdata provides comprehensive system visibility, predictable per node pricing, and zero-configuration deployment—solving the gaps in Honeycomb&amp;rsquo;s application-centric observability.</description></item><item><title>Netdata vs IBM Instana | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/instana/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/instana/</guid><description/></item><item><title>Netdata vs Icinga | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/icinga/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/icinga/</guid><description>Comprehensive comparison of Netdata and Icinga monitoring platforms, highlighting real-time performance analysis, automated deployment, ML-based anomaly detection, and modern observability capabilities.</description></item><item><title>Netdata vs Lansweeper | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/lansweeper/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/lansweeper/</guid><description/></item><item><title>Netdata vs Last9 | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/last9/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/last9/</guid><description/></item><item><title>Netdata vs LibreNMS | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/librenms/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/librenms/</guid><description>Compare Netdata and LibreNMS monitoring platforms. Learn how Netdata delivers per-second real-time monitoring, AI-powered anomaly detection, and native Windows support while LibreNMS excels at SNMP network device monitoring.</description></item><item><title>Netdata vs LogicMonitor | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/logicmonitor/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/logicmonitor/</guid><description>Comprehensive comparison of Netdata and LogicMonitor monitoring platforms, highlighting real-time performance, cost efficiency, and deployment simplicity advantages.</description></item><item><title>Netdata vs Logz.io | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/logz/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/logz/</guid><description>Netdata provides real-time, edge-native observability with predictable per-node pricing and zero-configuration deployment. Unlike Logz.io&amp;rsquo;s centralized log aggregation model with volume-based billing, Netdata distributes intelligence to your systems - delivering 90% cost reduction, 80% faster MTTR, and complete data sovereignty.</description></item><item><title>Netdata vs ManageEngine | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/manageengine/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/manageengine/</guid><description>Netdata vs ManageEngine: Real-time observability comparison for cloud-native infrastructure. See how per-second monitoring, distributed architecture, and zero-config deployment transform operations.</description></item><item><title>Netdata vs Mezmo | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/mezmo/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/mezmo/</guid><description>Comprehensive comparison of Netdata and Mezmo observability platforms, highlighting real-time monitoring capabilities, cost efficiency, and deployment simplicity.</description></item><item><title>Netdata vs Monit | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/monit/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/monit/</guid><description>Compare Netdata&amp;rsquo;s real-time observability platform with Monit&amp;rsquo;s lightweight process supervision. Learn how Netdata provides per-second monitoring, ML anomaly detection, AI troubleshooting, and unlimited scalability versus Monit&amp;rsquo;s basic service checks and 1,000-host ceiling.</description></item><item><title>Netdata vs N-able: Real-Time Monitoring Comparison</title><link>https://www.netdata.cloud/comparisons/nable/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/nable/</guid><description>Netdata provides per-second infrastructure monitoring with ML-based anomaly detection and AI troubleshooting - capabilities N-able&amp;rsquo;s 5-10 minute intervals and basic dashboards can&amp;rsquo;t match. See how Netdata solves the monitoring gaps N-able customers experience daily.</description></item><item><title>Netdata vs Nagios | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/nagios/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/nagios/</guid><description/></item><item><title>Netdata vs Netmon: Real-Time Observability Compared</title><link>https://www.netdata.cloud/comparisons/netmon/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/netmon/</guid><description>Netdata transforms infrastructure monitoring with real-time per-second data collection, ML-based anomaly detection, and AI-powered troubleshooting - replacing traditional appliance-based solutions with modern cloud-native observability.</description></item><item><title>Netdata vs New Relic | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/newrelic/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/newrelic/</guid><description/></item><item><title>Netdata vs Observium | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/observium/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/observium/</guid><description>Netdata vs Observium comparison: Real-time full-stack observability with ML and AI versus SNMP-based network device monitoring. See how Netdata&amp;rsquo;s distributed architecture, per-second granularity, and automated intelligence deliver comprehensive infrastructure visibility that network-only monitoring cannot provide.</description></item><item><title>Netdata vs Prometheus: No PromQL Complexity</title><link>https://www.netdata.cloud/comparisons/prometheus/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/prometheus/</guid><description>Netdata vs Prometheus comparison: Real-time monitoring with zero configuration, built-in ML, and 90% lower costs. See how Netdata solves Prometheus pain points while maintaining enterprise-grade capabilities.</description></item><item><title>Netdata vs PRTG | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/prtg/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/prtg/</guid><description/></item><item><title>Netdata vs Rakuten SixthSense | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/rakuten/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/rakuten/</guid><description>Netdata provides per-second real-time infrastructure monitoring with AI-powered troubleshooting for 90% less than traditional APM tools. Compare features, pricing, and capabilities to see why infrastructure teams choose Netdata over Rakuten SixthSense.</description></item><item><title>Netdata vs Sensu | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/sensu/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/sensu/</guid><description>Netdata vs Sensu: Real-time monitoring that just works vs a platform in decline. See why thousands of teams are switching to Netdata for instant deployment, beautiful dashboards, and 90% cost savings.</description></item><item><title>Netdata vs Sentry | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/sentry/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/sentry/</guid><description>Netdata and Sentry serve different but complementary purposes in modern observability. Netdata provides real-time infrastructure monitoring with ML-based anomaly detection, while Sentry excels at application error tracking. Learn how using both tools together delivers complete visibility at lower total cost.</description></item><item><title>Netdata vs SigLens | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/siglens/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/siglens/</guid><description>Comprehensive comparison of Netdata and SigLens observability platforms. Learn how these complementary tools work together to provide complete infrastructure monitoring and log analysis at a fraction of traditional enterprise costs.</description></item><item><title>Netdata vs SigNoz | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/signoz/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/signoz/</guid><description>Comprehensive comparison of Netdata and SigNoz observability platforms, highlighting operational simplicity, real-time performance, and cost efficiency for infrastructure monitoring teams.</description></item><item><title>Netdata vs Splunk | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/splunk/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/splunk/</guid><description/></item><item><title>Netdata vs StackState | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/stackstate/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/stackstate/</guid><description>Real-time monitoring in 60 seconds vs months of topology setup. Netdata provides per-second observability, ML anomaly detection, and AI troubleshooting at 90% lower cost - perfect for lean teams who need comprehensive monitoring without enterprise complexity.</description></item><item><title>Netdata vs Sumo Logic | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/sumologic/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/sumologic/</guid><description>Comprehensive comparison of Netdata and Sumo Logic: real-time monitoring, pricing models, deployment complexity, and use cases. Learn which platform fits your DevOps and observability needs.</description></item><item><title>Netdata vs Uptime Kuma | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/uptimekuma/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/uptimekuma/</guid><description/></item><item><title>Netdata vs UptimeRobot | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/uptimerobot/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/uptimerobot/</guid><description>Netdata and UptimeRobot serve different purposes: UptimeRobot monitors external availability while Netdata provides deep infrastructure diagnostics. Learn when to use each tool and how Netdata&amp;rsquo;s component-level alerts and ML-based anomaly detection provide deeper visibility.</description></item><item><title>Netdata vs VictoriaMetrics | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/victoriametrics/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/victoriametrics/</guid><description>Comprehensive comparison of Netdata vs VictoriaMetrics covering real-time monitoring, ML anomaly detection, visualization, scalability, and total cost of ownership.</description></item><item><title>Netdata vs Virtana: Real-Time Operations Monitoring</title><link>https://www.netdata.cloud/comparisons/virtana/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/virtana/</guid><description>Netdata delivers per-second operational monitoring for engineering teams, while Virtana provides strategic cost optimization for IT executives. Learn how these complementary platforms serve different needs within modern organizations.</description></item><item><title>Netdata vs WhatsUp Gold | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/whatsup-gold/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/whatsup-gold/</guid><description>Netdata vs WhatsUp Gold: Real-time observability with per-second granularity, zero configuration, and built-in ML/AI vs traditional Windows-centric network monitoring. See how Netdata delivers 90% cost reduction and 80% faster MTTR for modern infrastructure.</description></item><item><title>Netdata vs Zabbix | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/zabbix/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/zabbix/</guid><description/></item><item><title>Netdata vs Zenoss | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/zenoss/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/zenoss/</guid><description>Comprehensive comparison of Netdata and Zenoss monitoring platforms, highlighting Netdata&amp;rsquo;s advantages in real-time performance, ease of use, resource efficiency, and transparent pricing.</description></item><item><title>Network Monitoring Software For Education</title><link>https://www.netdata.cloud/solutions/industries/education/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/industries/education/</guid><description>Netdata delivers comprehensive observability for educational institutions with zero-configuration deployment, complete data sovereignty, and 90% cost savings. Monitor everything from learning management systems to research clusters with per-second precision.</description></item><item><title>Observability Cost Control With Per-Node Pricing</title><link>https://www.netdata.cloud/features/enterprise/cost-efficiency/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/enterprise/cost-efficiency/</guid><description>Transform observability from unpredictable cost center to fixed operational expense. Netdata&amp;rsquo;s distributed architecture and transparent pricing deliver 90% cost savings while providing unlimited metrics, logs, and per-second visibility.</description></item><item><title>Open Source</title><link>https://www.netdata.cloud/open-source/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/open-source/</guid><description>The world&amp;rsquo;s most popular open source monitoring platform. Deploy complete infrastructure observability in 60 seconds with zero configuration. Trusted by millions of engineers worldwide.</description></item><item><title>Operations Center Monitoring &amp; 24/7 NOC Teams</title><link>https://www.netdata.cloud/solutions/built-for/ops/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/built-for/ops/</guid><description>Real-time observability platform designed for 24/7 operations centers. Per-second monitoring, zero-configuration deployment, and AI-powered troubleshooting reduce MTTR by 80%.</description></item><item><title>Our Values and Promises</title><link>https://www.netdata.cloud/values/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/values/</guid><description>Discover the 12 principles that guide every decision at Netdata. From radical transparency to edge-native intelligence, learn how we invest engineering effort to eliminate complexity and align our success with your operational needs.</description></item><item><title>Per-Second Observability Without Cardinality Limits</title><link>https://www.netdata.cloud/features/architecture/real-time-at-scale/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/architecture/real-time-at-scale/</guid><description>Netdata delivers true real-time monitoring at planetary scale through distributed edge intelligence - maintaining per-second granularity and sub-2-second latency whether you&amp;rsquo;re monitoring 10 nodes or 100,000, with 90% cost reduction and 80% MTTR improvement.</description></item><item><title>Point-and-Click Analysis Without Query Languages</title><link>https://www.netdata.cloud/features/visualization/no-query-language/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/visualization/no-query-language/</guid><description>Experience genuine queryless observability with Netdata&amp;rsquo;s algorithmic dashboards and NIDL framework. Each chart delivers the equivalent of 25+ Grafana visualizations through simple point-and-click navigation - making expert-level analysis accessible to every engineer from day one.</description></item><item><title>Proxmox Monitoring With Real-Time Observability</title><link>https://www.netdata.cloud/solutions/technologies/proxmox-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/technologies/proxmox-monitoring/</guid><description>Netdata delivers enterprise-grade Proxmox monitoring with per-second granularity, automated dashboards, and ML anomaly detection - all without the complexity or cost of traditional solutions.</description></item><item><title>Query Systemd-Journal Logs Without Pipelines</title><link>https://www.netdata.cloud/solutions/use-cases/systemd-journal-logs/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/systemd-journal-logs/</guid><description>Netdata revolutionizes systemd-journal log management by querying logs directly where they live, eliminating costly pipelines and delivering faster queries with zero configuration. Get complete visibility into Linux system logs with AI-powered troubleshooting, unified metrics correlation, and 90% cost reduction compared to traditional log management platforms.</description></item><item><title>Real-Time GPU Monitoring For AI Infrastructure</title><link>https://www.netdata.cloud/solutions/industries/ai/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/industries/ai/</guid><description>Monitor AI training clusters, inference APIs, and ML workloads with Netdata&amp;rsquo;s edge-native platform. Per-second visibility, ML-powered insights, predictable pricing.</description></item><item><title>Real-Time Infrastructure Monitoring For Freelancers</title><link>https://www.netdata.cloud/solutions/built-for/freelancers/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/built-for/freelancers/</guid><description>Real-time infrastructure monitoring built for technical freelancers managing multiple client environments. Zero configuration, AI-powered insights, and predictable costs let you focus on delivering value, not managing monitoring tools.</description></item><item><title>Real-Time Infrastructure Monitoring For Tech Industry</title><link>https://www.netdata.cloud/solutions/industries/technology/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/industries/technology/</guid><description>Netdata delivers per-second observability for technology infrastructure with ML-based anomaly detection, zero-configuration deployment, and predictable per-node pricing - solving the cost, complexity, and visibility challenges facing modern tech teams.</description></item><item><title>Real-Time Infrastructure Observability For CISOs</title><link>https://www.netdata.cloud/solutions/built-for/cisos/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/built-for/cisos/</guid><description>Transform infrastructure security monitoring with real-time visibility, edge-based ML anomaly detection, and zero-configuration deployment. Netdata empowers CISOs to detect threats faster, reduce operational burden, and maintain data sovereignty - all at 90% lower cost than traditional SIEMs.</description></item><item><title>Real-Time Infrastructure Troubleshooting With AI</title><link>https://www.netdata.cloud/solutions/use-cases/troubleshooting/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/troubleshooting/</guid><description>Transform troubleshooting from hours to minutes with Netdata&amp;rsquo;s real-time observability platform. Get per-second metrics, automatic anomaly detection, and AI-guided investigations—all with zero configuration and 90% cost savings.</description></item><item><title>Real-Time Monitoring &amp; ML Detection For Sysadmins</title><link>https://www.netdata.cloud/solutions/built-for/sysadmins/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/built-for/sysadmins/</guid><description>Purpose-built observability for sysadmins: per-second metrics, automatic dashboards, ML-powered anomaly detection, and predictable pricing. Monitor everything from bare metal to Kubernetes without learning query languages or building dashboards.</description></item><item><title>Real-Time Observability For 24/7 Operations</title><link>https://www.netdata.cloud/solutions/use-cases/continuous-operations/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/continuous-operations/</guid><description>Real-time observability built for continuous operations. Per-second metrics, ML anomaly detection, and AI root cause analysis keep your infrastructure running around the clock.</description></item><item><title>Real-Time Observability For Developers Who Ship Fast</title><link>https://www.netdata.cloud/solutions/built-for/developers/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/built-for/developers/</guid><description>Stop context switching between tools. Netdata provides developers with real-time infrastructure visibility, AI-assisted debugging, and zero-configuration monitoring—all in one unified platform that integrates directly into your IDE and workflow.</description></item><item><title>Real-Time Observability For Government &amp; Public Sector</title><link>https://www.netdata.cloud/solutions/industries/government/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/industries/government/</guid><description>Netdata delivers enterprise-grade observability for government agencies with per-second monitoring, configurable multi-year log retention, edge-based ML anomaly detection, and 90% cost savings vs. legacy monitoring - all while keeping data sovereign on your infrastructure.</description></item><item><title>Real-Time Observability For Platform Engineers</title><link>https://www.netdata.cloud/solutions/built-for/platform-engineers/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/built-for/platform-engineers/</guid><description>Empower platform engineering teams with Netdata&amp;rsquo;s distributed observability platform. Get per-second visibility, automated dashboards, ML anomaly detection, and predictable costs - all without query languages or complex pipelines.</description></item><item><title>Real-Time Observability For SRE Teams</title><link>https://www.netdata.cloud/solutions/built-for/sre/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/built-for/sre/</guid><description>Transform SRE operations with Netdata&amp;rsquo;s edge-native observability platform. Get per-second visibility, ML anomaly detection on every metric, and AI-powered troubleshooting at 90% lower cost than traditional solutions.</description></item><item><title>Real-Time Troubleshooting With Sub-2-Second Latency</title><link>https://www.netdata.cloud/features/visualization/troubleshooting/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/visualization/troubleshooting/</guid><description>Interactive debugging with per-second precision and full historical context. Netdata transforms troubleshooting through edge-native ML, automated correlation, and AI-powered analysis.</description></item><item><title>Red Hat OpenShift Monitoring Tool At Scale</title><link>https://www.netdata.cloud/solutions/technologies/redhat-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/technologies/redhat-monitoring/</guid><description>Netdata delivers real-time, AI-powered observability for Red Hat OpenShift with 90% cost savings, per-second metrics, ML anomaly detection on every metric, and zero-configuration deployment. Eliminate Prometheus memory exhaustion, Loki query timeouts, and RHACM complexity.</description></item><item><title>Retail Monitoring Software Without PromQL</title><link>https://www.netdata.cloud/solutions/industries/retail/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/industries/retail/</guid><description>Transform retail operations with Netdata&amp;rsquo;s edge-native observability platform. Monitor POS systems, payment processing, store networks, and inventory infrastructure in real-time with per-second granularity, ML-powered anomaly detection, and predictable per-node pricing.</description></item><item><title>Root Cause Analysis (RCA) With 80% MTTR Reduction</title><link>https://www.netdata.cloud/features/aiml/root-cause-analysis/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/aiml/root-cause-analysis/</guid><description>Netdata&amp;rsquo;s edge-native ML detects anomalies as they happen, automated correlation surfaces root causes in the top 30-50 results, and AI explains incidents in plain English - achieving 80% MTTR reduction at 90% lower cost.</description></item><item><title>Service Mesh Observability Without Overhead</title><link>https://www.netdata.cloud/solutions/use-cases/service-mesh-observability/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/service-mesh-observability/</guid><description>Netdata revolutionizes service mesh observability with true real-time monitoring, automatic anomaly detection, and transparent pricing - solving the cost explosion, performance overhead, and complexity that plague traditional solutions.</description></item><item><title>Synthetic Monitoring Solution For APIs &amp; Services</title><link>https://www.netdata.cloud/solutions/use-cases/synthetic-checks/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/synthetic-checks/</guid><description>Real-time synthetic monitoring for APIs, certificates, DNS, and network health - unified with infrastructure observability for faster troubleshooting and predictable costs.</description></item><item><title>Team Collaboration With Real-Time Observability</title><link>https://www.netdata.cloud/features/enterprise/team-collaboration/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/enterprise/team-collaboration/</guid><description>Unite teams with real-time visibility, intelligent collaboration features, and enterprise-grade access control - all while maintaining data sovereignty and predictable costs.</description></item><item><title>Telecom Network Monitoring Software At Scale</title><link>https://www.netdata.cloud/solutions/industries/telecom/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/industries/telecom/</guid><description>Transform telecom operations with distributed edge monitoring that delivers complete infrastructure visibility, ML-powered anomaly detection, and zero-configuration deployment in minutes.</description></item><item><title>Text-To-Alert: Create Alerts From Natural Language</title><link>https://www.netdata.cloud/blog/ai-alerts/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/ai-alerts/</guid><description>&lt;p>Netdata has an incredibly powerful alerting engine. But this can sometimes be a double-edged sword: the flexibility to build incredibly specific, intelligent alerts is immense, but mastering its syntax can feel like learning a new language. We’ve heard this from so many of you. You tell us that configuring alerts is often the steepest part of the learning curve, a task that falls to the one &amp;ldquo;Netdata expert&amp;rdquo; on the team who has spent the time digging through the documentation.&lt;/p></description></item><item><title>Tiered Retention With 0.6 Bytes/Sample Storage</title><link>https://www.netdata.cloud/features/dataplatform/tiered-retention/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/dataplatform/tiered-retention/</guid><description>Netdata redefines tiered retention through edge-native architecture that updates all storage tiers simultaneously during collection, eliminating batch jobs and maintenance overhead while providing 15× better efficiency than traditional systems.</description></item><item><title>Unified Observability Platform With Predictable Costs</title><link>https://www.netdata.cloud/solutions/use-cases/unified-observability-architecture/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/unified-observability-architecture/</guid><description>Netdata delivers unified observability through edge-native architecture - metrics, logs, and AI-powered insights at per-second granularity without the complexity, cost explosion, or tool sprawl of traditional platforms.</description></item><item><title>VMware Monitoring Software Without Centralized Data</title><link>https://www.netdata.cloud/solutions/technologies/vmware-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/technologies/vmware-monitoring/</guid><description>Netdata delivers per-second VMware monitoring with zero configuration, built-in ML anomaly detection, and 90% cost savings. Monitor vSphere, ESXi hosts, VMs, and hybrid clouds with true real-time visibility.</description></item><item><title>Web Server Monitoring Software For Complete Visibility</title><link>https://www.netdata.cloud/solutions/use-cases/webserver-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/webserver-monitoring/</guid><description>Transform web server monitoring with Netdata&amp;rsquo;s edge-native platform. Get per-second metrics, automated root cause analysis, and complete infrastructure visibility without complex setup or unpredictable costs.</description></item><item><title>Windows Event Log Monitoring Software</title><link>https://www.netdata.cloud/solutions/use-cases/windows-event-logs/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/windows-event-logs/</guid><description>Netdata revolutionizes Windows Event Log monitoring by combining infrastructure metrics with native log access, ML anomaly detection, and AI troubleshooting - delivering comprehensive observability at a fraction of traditional SIEM costs.</description></item><item><title>Windows Server Monitoring Software At Scale</title><link>https://www.netdata.cloud/solutions/technologies/windows-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/technologies/windows-monitoring/</guid><description>Transform Windows Server monitoring with Netdata&amp;rsquo;s edge-native platform. Get per-second visibility, automatic discovery, and AI-powered troubleshooting at a fraction of traditional costs.</description></item><item><title>Zero Configuration Observability In 60 Seconds</title><link>https://www.netdata.cloud/features/architecture/zero-configuration/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/architecture/zero-configuration/</guid><description>True zero-configuration observability: from installation to troubleshooting production issues in 60 seconds, without specialized skills or months of integration work.</description></item><item><title>Zero Data Storage In SaaS With Unified Visibility</title><link>https://www.netdata.cloud/product/netdata-cloud-saas/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/product/netdata-cloud-saas/</guid><description>Access your entire infrastructure from anywhere while all observability data stays on-premises. Unlimited horizontal scalability with predictable per-node pricing. Role-based access control, dynamic organization, infrastructure-level dashboards, centralized alerts, and managed AI troubleshooting.</description></item><item><title>Zero-Downtime Monitoring For Real-Time Deployments</title><link>https://www.netdata.cloud/features/architecture/zero-downtime-monitoring/</link><pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/architecture/zero-downtime-monitoring/</guid><description>Netdata&amp;rsquo;s distributed architecture eliminates single points of failure in monitoring, delivering sub-2-second visibility with automatic failover, zero data loss, and production-safe resource overhead—ensuring complete observability during the moments that matter most.</description></item><item><title>Dashboard Gaps During Outages: Research Report</title><link>https://www.netdata.cloud/resources/research/dashboard-gaps-during-outages/</link><pubDate>Mon, 01 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/resources/research/dashboard-gaps-during-outages/</guid><description>&lt;blockquote>
&lt;p>&lt;strong>Research Notice&lt;/strong>: This document was compiled through online research conducted on December 1, 2025.
It serves as reference material for our blog post: &lt;a href="https://www.netdata.cloud/blog/monitor-everything-is-an-anti-pattern/">Monitor Everything is an Anti-Pattern!&lt;/a>.
Sources are cited inline and summarized at the end.&lt;/p>
&lt;/blockquote>
&lt;h2 id="outage-incidents-with-monitoring-gaps-missing-dashboards--quantified-impact">Outage Incidents With Monitoring Gaps, Missing Dashboards &amp;amp; Quantified Impact&lt;/h2>
&lt;h2 id="executive-summary">Executive Summary&lt;/h2>
&lt;p>This report documents multiple real-world outage incidents where monitoring systems failed to collect critical data, engineers created dashboards ad-hoc during crises, and teams implemented systematic improvements post-incident. The research reveals consistent patterns across major technology companies including Cloudflare, Datadog, GitLab, AWS, Azure, PagerDuty, and Google SRE, with quantified financial impacts reaching &lt;strong>$2 million per hour&lt;/strong> for high-impact outages.&lt;/p></description></item><item><title>Monitor Everything is an Anti-Pattern!</title><link>https://www.netdata.cloud/blog/monitor-everything-is-an-anti-pattern/</link><pubDate>Mon, 01 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/monitor-everything-is-an-anti-pattern/</guid><description>&lt;p>&lt;strong>Bullshit and nonsense.&lt;/strong>&lt;/p>
&lt;p>But let&amp;rsquo;s take it from the beginning.&lt;/p>
&lt;p>The industry&amp;rsquo;s story goes something like this:&lt;/p>
&lt;blockquote>
&lt;p>&lt;em>&amp;ldquo;Monitor everything is universally recognized as an anti-pattern.&amp;rdquo;&lt;/em>&lt;br/>
&lt;em>&amp;ldquo;You&amp;rsquo;ll drown in metrics, burn out your engineers, and blow your budget.&amp;rdquo;&lt;/em>&lt;br/>
&lt;em>&amp;ldquo;Just focus on 3–10 signals — the Four Golden Signals, RED, USE — and ignore everything else.&amp;rdquo;&lt;/em>&lt;br/>
&lt;em>&amp;ldquo;Trust us, you don&amp;rsquo;t want that much telemetry.&amp;rdquo;&lt;/em>&lt;br/>
&lt;br/>
(&lt;a href="https://www.netdata.cloud/resources/research/monitor-everything-anti-pattern/">true, read the whole story here&lt;/a>)&lt;/p>
&lt;/blockquote>
&lt;p>Then, in the same breath:&lt;/p></description></item><item><title>Why 'Monitor Everything' Is An Anti-Pattern: Research</title><link>https://www.netdata.cloud/resources/research/monitor-everything-anti-pattern/</link><pubDate>Mon, 01 Dec 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/resources/research/monitor-everything-anti-pattern/</guid><description>&lt;blockquote>
&lt;p>&lt;strong>Research Notice&lt;/strong>: This document was compiled through online research conducted on December 1, 2025.
It serves as reference material for our blog post: &lt;a href="https://www.netdata.cloud/blog/monitor-everything-is-an-anti-pattern/">Monitor Everything is an Anti-Pattern!&lt;/a>.
Sources are cited inline and summarized at the end.&lt;/p>
&lt;/blockquote>
&lt;h2 id="tldr">TL;DR&lt;/h2>
&lt;p>&amp;ldquo;Monitor everything&amp;rdquo; is universally recognized as an anti-pattern by SRE experts, observability leaders, and major tech companies. The core reasons are:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Metric Fatigue&lt;/strong>: Teams become overwhelmed by excessive data, unable to identify critical signals&lt;/li>
&lt;li>&lt;strong>Alert Fatigue&lt;/strong>: 63% of organizations face 1,000+ daily alerts with 72-99% false positives, costing $300,000+/hour in missed incidents&lt;/li>
&lt;li>&lt;strong>Lack of Actionability&lt;/strong>: 97% of alerts are non-actionable noise rather than signals requiring response&lt;/li>
&lt;li>&lt;strong>High Costs&lt;/strong>: Organizations spend 20-40% of cloud budgets on observability (vs. optimal 10-15%), with cardinality explosions creating exponential cost increases&lt;/li>
&lt;li>&lt;strong>Employee Burnout&lt;/strong>: Costs $4,000-$21,000 per employee annually, totaling $5M+ for 1,000-person companies&lt;/li>
&lt;li>&lt;strong>System Complexity&lt;/strong>: Monitoring systems themselves become fragile, requiring constant maintenance&lt;/li>
&lt;li>&lt;strong>Monitoring Tools Creating Problems&lt;/strong>: Monitoring agents can cause the latency outliers they&amp;rsquo;re meant to detect&lt;/li>
&lt;/ol>
&lt;p>&lt;strong>Expert Consensus&lt;/strong>: Focus on 3-10 key metrics (Google&amp;rsquo;s Four Golden Signals, RED Method, USE Method) that indicate symptoms rather than attempting comprehensive monitoring of all possible metrics.&lt;/p></description></item><item><title>Gartner IOCS 2025: Tackling Observability Overspend</title><link>https://www.netdata.cloud/blog/gartner-iocs-2025/</link><pubDate>Fri, 07 Nov 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/gartner-iocs-2025/</guid><description>&lt;p>The observability market is facing a paradox. As organizations spend more than ever on monitoring tools, their infrastructure complexity continues to grow, and incident resolution times often remain stubbornly high. Teams are drowning in data, struggling with tool sprawl, and facing unpredictable, budget-breaking bills.&lt;/p>
&lt;p>This challenge, how to gain better visibility without spiraling costs, is one of the most critical conversations for IT leaders today.&lt;/p>
&lt;!--truncate-->
&lt;p>That&amp;rsquo;s why the Netdata team is heading to Las Vegas for the &lt;strong>Gartner IT Infrastructure, Operations &amp;amp; Cloud Strategies (IOCS) Conference&lt;/strong> from December 9-11, 2025. We&amp;rsquo;ll be there to discuss this challenge head-on and share our vision for a more efficient and intelligent future for observability.&lt;/p></description></item><item><title>ServiceNow Integration: Streamline Incident Response</title><link>https://www.netdata.cloud/blog/servicenow-integration/</link><pubDate>Fri, 07 Nov 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/servicenow-integration/</guid><description>&lt;p>When a critical alert fires at 2 AM, the last thing your on-call engineer should be doing is manual administrative work. Yet, for many teams, that&amp;rsquo;s exactly what happens. You see the alert in your monitoring tool, then you have to switch contexts, open a new browser tab, log into your ITSM platform, and manually create an incident—all while your systems are failing.&lt;/p>
&lt;p>This &amp;ldquo;swivel-chairing&amp;rdquo; between tools is slow, error-prone, and a significant drag on your Mean Time to Resolution (MTTR).&lt;/p></description></item><item><title>SOC 2 Type 2: Validated Security Controls Over Time</title><link>https://www.netdata.cloud/blog/soc2-type-2-compliance/</link><pubDate>Thu, 16 Oct 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/soc2-type-2-compliance/</guid><description>&lt;p>We&amp;rsquo;re excited to share that Netdata has successfully achieved SOC 2 Type 2 attestation.&lt;/p>
&lt;!--truncate-->
&lt;p>Following a five-month audit conducted by Sensiba LLP, we can now confirm that our security controls work consistently in practice. The audit covered the period from April 1 to August 31, 2025, and tested whether our controls operated effectively throughout that entire timeframe.&lt;/p>
&lt;p>Back in April, we announced our &lt;a href="https://www.netdata.cloud/blog/soc-2-type1/">SOC 2 Type 1 attestation&lt;/a>, which validated that our security controls were properly designed at a specific point in time. We also mentioned we were entering the monitoring period for Type 2. Today we can share the results.&lt;/p></description></item><item><title>Automate Infrastructure Analysis With AI Reports</title><link>https://www.netdata.cloud/blog/scheduled-reports-insights-investigations/</link><pubDate>Tue, 23 Sep 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/scheduled-reports-insights-investigations/</guid><description>&lt;p>The least exciting part of an operations or SRE role is often the manual, repetitive task of generating reports. It’s the Monday morning scramble to summarize weekly infrastructure health for the team, or the end-of-quarter push to build a capacity planning document. This is boilerplate work that pulls you away from critical engineering tasks.&lt;/p>
&lt;p>We believe that if a process is repeatable, it should be automated.&lt;/p>
&lt;!--truncate-->
&lt;p>That&amp;rsquo;s why we’re introducing &lt;strong>Scheduled AI Investigations and Insights&lt;/strong>. This new capability builds directly on our existing AI tools, allowing you to set your most important analyses on a recurring schedule. It’s like setting up a cron job for your infrastructure reporting, letting your Co-SRE do the heavy lifting for you.&lt;/p></description></item><item><title>Elasticsearch Yellow Cluster: Unassigned Shards Fix</title><link>https://www.netdata.cloud/academy/elasticsearch-yellow-cluster-access/</link><pubDate>Sun, 07 Sep 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/elasticsearch-yellow-cluster-access/</guid><description>&lt;p>You run a health check on your production cluster and the result comes back: &lt;code>status: yellow&lt;/code>. It&amp;rsquo;s not the dreaded red status, so your application is likely still serving requests, but this is a critical warning sign. An Elasticsearch yellow cluster status is a direct indication that your data&amp;rsquo;s high availability is compromised. While all your primary shards are active, one or more replica shards have failed to be assigned to a node. Ignoring this warning can lead to data loss if another node fails, or worse, it could be a symptom of a network partition risking an Elasticsearch split-brain.&lt;/p></description></item><item><title>Docker Layer Caching in CI Pipelines Cut Build Times by 70 %</title><link>https://www.netdata.cloud/academy/docker-layer-caching/</link><pubDate>Sat, 06 Sep 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/docker-layer-caching/</guid><description>&lt;p>You push a one-line code change, and your CI/CD pipeline kicks off. You grab a coffee, come back, and it&amp;rsquo;s still running, stuck on the &lt;code>docker build&lt;/code> step. Slow container builds are a silent productivity killer, delaying feedback, slowing down deployments, and frustrating developers. In an ephemeral CI environment where every job starts with a clean slate, Docker often has to rebuild your entire application image from scratch, every single time.&lt;/p></description></item><item><title>Linux Cgroups V2 Memory Throttling &amp; OOM Fix</title><link>https://www.netdata.cloud/academy/diagnosing-linux-cgroups/</link><pubDate>Fri, 05 Sep 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/diagnosing-linux-cgroups/</guid><description>&lt;p>Your critical service is lagging. Users are complaining about timeouts. You check your orchestration platform and see the dreaded &lt;code>OOMKilled&lt;/code> status on a container. You dive into the node&amp;rsquo;s logs (&lt;code>dmesg&lt;/code>) and confirm it: the kernel&amp;rsquo;s Out-of-Memory (OOM) killer has claimed another victim. The immediate fix is easy—restart the container, maybe give it more memory—but the real question remains unanswered: &lt;em>why&lt;/em> did it happen? Was it a sudden memory leak, a traffic spike, or something more subtle?&lt;/p></description></item><item><title>Designing Error Budget Policies For SLOs At Scale</title><link>https://www.netdata.cloud/academy/designing-error-budget-policies/</link><pubDate>Thu, 04 Sep 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/designing-error-budget-policies/</guid><description>&lt;p>In every engineering organization, there&amp;rsquo;s a constant, fundamental tension: the push to ship new features versus the need to maintain a stable, reliable service. Move too fast, and you risk outages that erode user trust. Move too slowly, and you risk being outpaced by the competition. For years, this balancing act was managed by intuition, late-night heroics, and tense priority meetings. Site Reliability Engineering (SRE) offers a better way: the error budget.&lt;/p></description></item><item><title>Consul Service Discovery Failures: Causes &amp; Fixes</title><link>https://www.netdata.cloud/academy/consul-service-discovery-failures/</link><pubDate>Wed, 03 Sep 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/consul-service-discovery-failures/</guid><description>&lt;p>It’s a scenario that keeps DevOps and SRE teams up at night: your application logs fill with connection errors, services start failing, and you realize a critical component can&amp;rsquo;t find the database it depends on. The culprit? A breakdown in your service discovery mechanism. For many, that mechanism is HashiCorp Consul, the backbone of modern microservice architectures. When Consul falters, your entire ecosystem can become unstable.&lt;/p>
&lt;p>Understanding how to diagnose these failures is crucial. The problem often lies deep within the operational layers—agent communication issues, misconfigured health checks, or disruptions in the gossip protocol that maintains cluster state. In this guide, we&amp;rsquo;ll dissect the most common causes of Consul service discovery failures, providing you with the tools to troubleshoot and resolve them. More importantly, we&amp;rsquo;ll show you how to shift from a reactive, fire-fighting mode to a proactive one, using comprehensive monitoring to build a truly resilient Consul deployment.&lt;/p></description></item><item><title>AI Troubleshooting GA With On-Demand Credits</title><link>https://www.netdata.cloud/blog/ai-credits/</link><pubDate>Tue, 02 Sep 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/ai-credits/</guid><description>&lt;p>Since launching our AI investigations and insights in a research preview, one thing has become clear: &lt;strong>automated root cause analysis delivers a significant return on investment.&lt;/strong> Teams have confirmed that instant insights don&amp;rsquo;t just save a few minutes; they fundamentally shorten incident response cycles, free up valuable engineering hours, and reduce the business impact of downtime.&lt;/p>
&lt;p>The preview successfully demonstrated this value, with 10 free AI sessions per month allowing teams to integrate AI into their workflows. Now, based on the success and maturity of the capabilities, we are proud to announce that &lt;strong>Netdata&amp;rsquo;s AI investigations and insights are graduating from research preview to General Availability.&lt;/strong>&lt;/p></description></item><item><title>Fix NGINX 502/504 Double-Proxy Errors With CloudFront</title><link>https://www.netdata.cloud/academy/cloudfront-nginx-origin-solving-proxy/</link><pubDate>Tue, 02 Sep 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/cloudfront-nginx-origin-solving-proxy/</guid><description>&lt;p>You’ve set up a modern, resilient architecture: NGINX as your robust origin server and Cloudflare as your global CDN and security layer. Then, it happens. A visitor reports seeing a dreaded white screen with &amp;ldquo;502 Bad Gateway&amp;rdquo; or &amp;ldquo;504 Gateway Timeout.&amp;rdquo; The immediate question is, who&amp;rsquo;s to blame? Is Cloudflare having an &lt;code>edge_error&lt;/code>, or is your NGINX origin server failing? This confusion is a common pitfall in a &lt;code>double_reverse_proxy&lt;/code> setup where requests pass through multiple layers.&lt;/p></description></item><item><title>PostgreSQL 17 Cardinality Estimation &amp; Tuning Tips</title><link>https://www.netdata.cloud/academy/cardinality-estimation-in-postgres/</link><pubDate>Mon, 01 Sep 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/cardinality-estimation-in-postgres/</guid><description>&lt;p>You&amp;rsquo;ve seen it before: a query that runs in milliseconds on your staging server takes minutes to execute in production. Or a seemingly simple &lt;code>JOIN&lt;/code> causes the &lt;code>query_planner&lt;/code> to choose a disastrously slow Nested Loop over a much faster Hash Join. In almost every case, the root cause of these performance mysteries is not a bug in PostgreSQL, but a flaw in its understanding of your data. This is the challenge of &lt;code>cardinality_estimation&lt;/code>. The planner is only as smart as the statistics it&amp;rsquo;s given, and when its &lt;code>row_estimation&lt;/code> is wrong, the consequences are severe.&lt;/p></description></item><item><title>Using FOR UPDATE SKIP LOCKED For Queue Workflows</title><link>https://www.netdata.cloud/academy/update-skip-locked/</link><pubDate>Fri, 29 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/update-skip-locked/</guid><description>&lt;p>One of the most common and powerful patterns in modern application development is the job queue. Whether you&amp;rsquo;re sending emails, processing images, or running complex calculations, offloading tasks to background workers is essential for building responsive and scalable systems. Many developers reach for dedicated queueing software like RabbitMQ or Redis, but for many use cases, your primary PostgreSQL database already has all the tools you need to build a robust, transactional, and incredibly performant &lt;code>job_queue_postgres&lt;/code>.&lt;/p></description></item><item><title>SSL/TLS Handshake Failures Causing NGINX 502 Errors</title><link>https://www.netdata.cloud/academy/ssl-tls-handshake-failures-nginx/</link><pubDate>Wed, 27 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/ssl-tls-handshake-failures-nginx/</guid><description>&lt;p>You&amp;rsquo;ve set up NGINX as a reverse proxy, pointing it to a secure upstream service. You refresh your browser, and instead of your application, you&amp;rsquo;re greeted by a stark &amp;ldquo;502 Bad Gateway&amp;rdquo; page. You check your upstream service; it&amp;rsquo;s running perfectly. Puzzled, you turn to the NGINX error logs and find a cryptic message that seems to complicate things even further: &lt;code>SSL_do_handshake() failed&lt;/code>.&lt;/p>
&lt;p>This scenario is frustratingly common. A 502 error suggests a problem with the upstream server, but in many cases, NGINX itself is failing to establish a secure connection with that server. The true culprit is an SSL/TLS handshake failure, often masked by this generic HTTP error code. Understanding the root cause is key to a quick resolution, preventing prolonged downtime and tedious debugging sessions.&lt;/p></description></item><item><title>How Autovacuum Causes PostgreSQL Deadlocks</title><link>https://www.netdata.cloud/academy/autovaccum-vs-deadlock/</link><pubDate>Tue, 26 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/autovaccum-vs-deadlock/</guid><description>&lt;p>You&amp;rsquo;ve meticulously optimized your application queries. Your transaction logic is sound. Yet, under heavy load, your system seizes up, logging the dreaded &amp;ldquo;deadlock detected&amp;rdquo; error. You dig into the logs, expecting to find two application transactions locked in a deadly embrace, but instead, you find a surprising culprit: one of the participants is the PostgreSQL &lt;code>autovacuum&lt;/code> process. How can a routine maintenance task, designed to keep the database healthy, be the cause of a production-stopping &lt;code>postgres_deadlock&lt;/code>?&lt;/p></description></item><item><title>Real-World PostgreSQL Deadlock Examples &amp; Fixes</title><link>https://www.netdata.cloud/academy/10-real-world-postgresql-deadlock/</link><pubDate>Mon, 25 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/10-real-world-postgresql-deadlock/</guid><description>&lt;p>It’s 3 AM. The pager screams. Your application is throwing a cascade of errors, and users are reporting that the system is completely frozen. You dive into the logs and see the same ominous message repeating over and over: &lt;code>ERROR: deadlock detected&lt;/code>. A PostgreSQL deadlock is one of the most abrupt and disruptive failures a database can experience. It&amp;rsquo;s not a performance degradation; it&amp;rsquo;s a hard stop where two or more transactions are locked in a fatal embrace, each waiting for a resource the other holds.&lt;/p></description></item><item><title>Redis Sentinel Failover &amp; Split-Brain Recovery Guide</title><link>https://www.netdata.cloud/academy/redis-cluster-split/</link><pubDate>Sun, 24 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/redis-cluster-split/</guid><description>&lt;p>You&amp;rsquo;re on call. An alert fires—your Redis master node is unreachable. Your heart rate quickens. You&amp;rsquo;ve set up Redis Sentinel for high availability, but is it working? Did the failover succeed? Or worse, are you now in a Redis cluster split-brain situation where two nodes think they&amp;rsquo;re the master, leading to data inconsistency and eventual loss? In these critical moments, blindly trusting the automation isn&amp;rsquo;t enough; you need to verify what&amp;rsquo;s happening.&lt;/p></description></item><item><title>Fix NGINX 503 Errors From Misconfigured limit_req</title><link>https://www.netdata.cloud/academy/rate-limiting-gone-wrong-nginx/</link><pubDate>Sat, 23 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/rate-limiting-gone-wrong-nginx/</guid><description>&lt;p>You’ve done the responsible thing. To protect your application from abusive bots and prevent any single user from overwhelming your services, you&amp;rsquo;ve implemented rate limiting in NGINX. You add the &lt;code>limit_req_zone&lt;/code> and &lt;code>limit_req&lt;/code> directives, push the configuration, and watch. But instead of seeing a drop in malicious traffic, your monitoring dashboards light up with a sea of red. A massive &lt;code>503 Service Unavailable&lt;/code> spike appears, and legitimate users are complaining they can&amp;rsquo;t access your site. Your shield has become a weapon turned against yourself.&lt;/p></description></item><item><title>Prometheus Alertmanager: Noise Reduction Rules</title><link>https://www.netdata.cloud/academy/prometheus-alert-manager/</link><pubDate>Fri, 22 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/prometheus-alert-manager/</guid><description>&lt;p>It’s 3 AM, and a torrent of PagerDuty notifications floods your phone. A single network partition has triggered a cascade, making every service instance that can&amp;rsquo;t reach the database fire its own individual alert. You&amp;rsquo;re drowning in hundreds of notifications, all pointing to the same root cause. This is alert fatigue, a critical problem for any SRE or on-call engineer. When you&amp;rsquo;re constantly bombarded with low-signal noise, you risk becoming desensitized, potentially overlooking the one critical alert that signals a major incident.&lt;/p></description></item><item><title>Tune NGINX Timeouts: Eliminate Upstream Errors</title><link>https://www.netdata.cloud/academy/nginx-eliminate-upstream-timeout/</link><pubDate>Thu, 21 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/nginx-eliminate-upstream-timeout/</guid><description>&lt;p>The &lt;code>504 Gateway Timeout&lt;/code> error is a familiar and frustrating sight for anyone managing web applications. It signifies a breakdown in communication, but not between the user and your server. Instead, it means NGINX, acting as your trusty reverse proxy, gave up waiting for a response from a backend, or &amp;ldquo;upstream,&amp;rdquo; service. This could be your Node.js application, a PHP-FPM process, or a Python microservice.&lt;/p>
&lt;p>While your first instinct might be to blame the application, the root cause is often a simple configuration mismatch. NGINX has built-in timers to protect itself from unresponsive backends. When a backend process takes too long—longer than NGINX is configured to wait—NGINX proactively closes the connection and serves the dreaded 504 error. Understanding and tuning directives like &lt;code>proxy_read_timeout&lt;/code> and &lt;code>fastcgi_read_timeout&lt;/code> is the key to resolving these issues and building a more resilient infrastructure.&lt;/p></description></item><item><title>Linux Load Average Spikes: IO Wait &amp; Bottlenecks</title><link>https://www.netdata.cloud/academy/linux-system-load-average-spike/</link><pubDate>Tue, 19 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/linux-system-load-average-spike/</guid><description>&lt;p>An alert fires: &amp;ldquo;High load average on production server.&amp;rdquo; Your heart rate quickens. You SSH into the machine and run a command like top, only to be confused. The CPU usage is hovering at 10%, but the load average is sky-high. What’s going on? If the CPU isn&amp;rsquo;t busy, what is the system &amp;ldquo;loaded&amp;rdquo; with? This common scenario highlights one of the most misunderstood metrics in &lt;a href="https://www.netdata.cloud/solutions/linux-monitoring/">Linux performance troubleshooting&lt;/a>: the system load average.&lt;/p></description></item><item><title>Kubernetes CrashLoopBackOff: Causes &amp; Quick Fixes</title><link>https://www.netdata.cloud/academy/kubernetes-crash-loop-backoff/</link><pubDate>Mon, 18 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/kubernetes-crash-loop-backoff/</guid><description>&lt;p>You’ve deployed your application to Kubernetes, but a quick check on your pods reveals the dreaded &lt;code>CrashLoopBackOff&lt;/code> status. Your pod is stuck in a restart loop, and your service is down. This isn&amp;rsquo;t an error itself, but a status indicating that Kubernetes is trying to start a container, it crashes, and Kubernetes waits an exponentially increasing amount of time before trying again. This back-off mechanism prevents a faulty pod from overwhelming the cluster with constant restart attempts.&lt;/p></description></item><item><title>Jenkins Declarative Pipeline: Parallelism Fixes</title><link>https://www.netdata.cloud/academy/jenkins-declarative-pipeline/</link><pubDate>Sun, 17 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/jenkins-declarative-pipeline/</guid><description>&lt;p>You&amp;rsquo;ve embraced the power of Jenkins Declarative Pipelines, and now you&amp;rsquo;re looking to take the next step in CI/CD optimization: running stages in parallel. The promise is alluring—slashing your build times by running tests, linting, and security scans simultaneously. You eagerly refactor your &lt;code>Jenkinsfile&lt;/code>, commit the change, and watch your pipeline trigger. But instead of blazing-fast execution, your build gets stuck, silently hanging with no obvious error.&lt;/p>
&lt;p>This frustrating scenario is often caused by two insidious problems: executor starvation and node label mismatch. While the &lt;code>parallel&lt;/code> step in Jenkins is incredibly powerful, using it naively can lead to deadlocked pipelines and misconfigured jobs that fail in subtle ways. Understanding how Jenkins allocates resources is the key to unlocking true parallelism without the headaches. This guide will walk you through these common pitfalls and provide the best practices to avoid them.&lt;/p></description></item><item><title>Buffering keepalive and Hidden 500-Series Errors in NGINX</title><link>https://www.netdata.cloud/academy/hidden-500-ngnix-series-errors/</link><pubDate>Sat, 16 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/hidden-500-ngnix-series-errors/</guid><description>&lt;p>Your NGINX instance is humming along, handling traffic flawlessly. Then, a sudden traffic spike or a deployment of a new feature that handles larger files occurs, and your monitoring dashboards light up with 500-series errors. You check the NGINX error logs, but find nothing conclusive—just generic &lt;code>502 Bad Gateway&lt;/code> or &lt;code>504 Gateway Timeout&lt;/code> messages that don&amp;rsquo;t point to a root cause. These are the hidden, frustrating errors that often stem not from your application code, but from poorly tuned NGINX buffering and &lt;code>keepalive&lt;/code> settings.&lt;/p></description></item><item><title>Fix Helm Chart Rollback &amp; Pending-Upgrade Failures</title><link>https://www.netdata.cloud/academy/helm-chart-rollback-failures/</link><pubDate>Fri, 15 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/helm-chart-rollback-failures/</guid><description>&lt;p>It’s a scenario that every Kubernetes operator dreads. A production deployment has gone wrong, you confidently initiate a Helm rollback, and then&amp;hellip; nothing. The command hangs, and a quick check reveals the dreaded &lt;code>pending-rollback&lt;/code> status. Your application is now in a broken state, and you&amp;rsquo;re blocked from deploying any new fixes. This is more than a minor inconvenience; it&amp;rsquo;s a critical failure that can leave your services unstable and your deployment pipeline paralyzed.&lt;/p></description></item><item><title>Fix NGINX 502 Bad Gateway In Kubernetes Ingress</title><link>https://www.netdata.cloud/academy/debugging-nginx-errors-inside-kubernetes/</link><pubDate>Wed, 13 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/debugging-nginx-errors-inside-kubernetes/</guid><description>&lt;p>You’ve deployed your application to Kubernetes, and everything seems fine until a user reports the dreaded &lt;code>502 Bad Gateway&lt;/code> or &lt;code>504 Gateway Timeout&lt;/code>. In a traditional VM setup, your first step would be to SSH into the server and check the NGINX configuration and logs. But in the ephemeral, abstracted world of Kubernetes, the problem is rarely that simple. A gateway error from the NGINX Ingress Controller is often a symptom of a deeper issue within the cluster’s networking or health management systems.&lt;/p></description></item><item><title>Blue-Green And Canary Deployments With NGINX</title><link>https://www.netdata.cloud/academy/blue-green-canary-deployments-nginx/</link><pubDate>Tue, 12 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/blue-green-canary-deployments-nginx/</guid><description>&lt;p>The moment of truth arrives. You&amp;rsquo;ve tested the new version of your application, the container image is pushed, and the deployment pipeline is ready. You click &amp;ldquo;deploy,&amp;rdquo; and a wave of anxiety hits. Will this be a smooth, zero-downtime rollout, or will your dashboards soon light up with &lt;code>502 Bad Gateway&lt;/code> and &lt;code>504 Gateway Timeout&lt;/code> errors? For many teams using advanced deployment strategies like Blue-Green or Canary, this fear is all too real.&lt;/p></description></item><item><title>Docker Compose Networking: Service Discovery &amp; Ports</title><link>https://www.netdata.cloud/academy/docker-compose-networking-mysteries/</link><pubDate>Sun, 10 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/docker-compose-networking-mysteries/</guid><description>&lt;p>You&amp;rsquo;ve meticulously crafted your docker-compose.yml file, your services build correctly, and the startup process finishes without a single error. Yet, when your application tries to connect to its database, you get a &amp;ldquo;Connection refused&amp;rdquo; or &amp;ldquo;Host not found&amp;rdquo; error. This is a classic and frustrating scenario in Docker Compose networking. Your containers are running, but they&amp;rsquo;re isolated in their own digital worlds, unable to communicate.&lt;/p>
&lt;p>These issues almost always boil down to two core mysteries: service discovery failed events, where containers can&amp;rsquo;t find each other, and docker compose port conflict problems, where port mappings are misunderstood or misconfigured. To build reliable containerized applications, you must look beyond the docker-compose.yml file and understand the virtual network that Docker creates. This guide will demystify that network, providing you with the tools and techniques to troubleshoot and resolve these common connectivity challenges.&lt;/p></description></item><item><title>Metric Cardinality In Observability: Strategies</title><link>https://www.netdata.cloud/academy/metric-cardinality-in-observability/</link><pubDate>Sun, 10 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/metric-cardinality-in-observability/</guid><description>&lt;p>It’s a story familiar to any SRE or DevOps engineer. You add a seemingly innocuous label to a key metric—&lt;code>user_id&lt;/code>, &lt;code>request_id&lt;/code>, &lt;code>container_id&lt;/code>—to gain deeper insight. Suddenly, your monitoring bill skyrockets, your &lt;code>Prometheus_TSDB&lt;/code> instance starts gasping for memory, and dashboards slow to a crawl. You have just triggered a &lt;code>label_explosion&lt;/code>, the single biggest challenge in modern metrics-based observability: &lt;code>metrics_cardinality&lt;/code>.&lt;/p>
&lt;p>High cardinality isn&amp;rsquo;t an edge case; it&amp;rsquo;s the new normal in a world of microservices, containers, and complex user interactions. As the number of unique time series grows into the millions or even billions, it places immense pressure on monitoring systems, impacting &lt;code>storage_costs&lt;/code>, query performance, and the fundamental ability to scale.&lt;/p></description></item><item><title>Early Deadlock Detection With pg_stat_kcache &amp; eBPF</title><link>https://www.netdata.cloud/academy/early-deadlock-detection/</link><pubDate>Sat, 09 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/early-deadlock-detection/</guid><description>&lt;p>The dreaded deadlock. For anyone managing a high-traffic PostgreSQL database, it&amp;rsquo;s a familiar nemesis. The database logs a &amp;ldquo;deadlock detected&amp;rdquo; message, a transaction is unceremoniously aborted, and your application has to handle the fallout. This reactive cycle is frustrating; by the time you&amp;rsquo;re alerted, the damage is already done. Traditional &lt;code>postgres_metrics&lt;/code> and monitoring tools are excellent at telling you &lt;em>that&lt;/em> a deadlock occurred, but they fall short of explaining the subtle conditions that led to it.&lt;/p></description></item><item><title>Save Hours on Troubleshooting with Automated Investigations</title><link>https://www.netdata.cloud/blog/automated-investigations/</link><pubDate>Mon, 04 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/automated-investigations/</guid><description>&lt;p>How many times has your team stared at a dashboard, pointed to a spike, and asked a question that charts alone can&amp;rsquo;t answer? &amp;ldquo;What was the real impact of that deployment?&amp;rdquo; &amp;ldquo;Why are our Kubernetes pods in the us-east-1 cluster suddenly crashing?&amp;rdquo; &amp;ldquo;Are we wasting money on overprovisioned servers?&amp;rdquo;&lt;/p>
&lt;!--truncate-->
&lt;p>Answering these questions is the real work of operations and SRE. It often kicks off a time-consuming scramble, sending engineers down rabbit holes for hours, days, or even weeks. You dig through logs, correlate metrics across services, and piece together clues from Slack conversations and Jira tickets.&lt;/p></description></item><item><title>Netdata Now Troubleshoots Your Alerts for You</title><link>https://www.netdata.cloud/blog/automated-alert-troubleshooting/</link><pubDate>Sun, 03 Aug 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/automated-alert-troubleshooting/</guid><description>&lt;p>The 2 AM pager alert. For anyone in Ops, SRE, or IT administration, those words trigger a familiar sense of dread. An alert has fired. Is it a real fire, or another false alarm waking you from a dead sleep? The pressure is on. Every minute of downtime costs money and reputation, but troubleshooting a complex system when you&amp;rsquo;re sleep-deprived is a Herculean task.&lt;/p>
&lt;!--truncate-->
&lt;p>This cycle is a massive drain on engineering resources. The daily grind of sifting through alerts, trying to distinguish signal from noise, and manually correlating metrics to find a root cause consumes countless hours. This constant firefighting leads to alert fatigue, where even critical notifications start to get ignored. The core questions are always the same: Is this a real problem? What is the potential impact? Why did this trigger? What do I do next? Answering them is a slow, manual, and often stressful process.&lt;/p></description></item><item><title>GitLab Runner Executor Failures: Docker &amp; Kubernetes</title><link>https://www.netdata.cloud/academy/gitlab-runner-executor/</link><pubDate>Thu, 24 Jul 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/gitlab-runner-executor/</guid><description>&lt;p>You’ve been there. You push your code, a new CI/CD pipeline kicks off in GitLab, and you wait for that satisfying green checkmark. Instead, you get a dreaded red &amp;lsquo;X&amp;rsquo;. The pipeline failed. Digging into the job logs, you find a cryptic message: &amp;ldquo;ERROR: Job failed (system failure): prepare environment: exit code 1&amp;rdquo;. The culprit is often the GitLab Runner executor—the very engine responsible for running your jobs—failing in its environment.&lt;/p></description></item><item><title>Apache Kafka Consumer Lag: Troubleshooting &amp; Fixes</title><link>https://www.netdata.cloud/academy/apache-kafka-consumer-lags/</link><pubDate>Wed, 23 Jul 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/apache-kafka-consumer-lags/</guid><description>&lt;p>You&amp;rsquo;ve built a powerful, real-time data pipeline with Apache Kafka, but suddenly, things grind to a halt. The dreaded consumer lag is exploding, alerts are firing, and your downstream applications are starved for data. This scenario is all too common for teams running Kafka at scale. Often, the culprit is a subtle interplay between consumer group rebalancing, partition assignment, and suboptimal consumer configurations. Understanding these mechanics is not just about fixing a problem; it&amp;rsquo;s about building resilient, high-throughput streaming systems from the ground up.&lt;/p></description></item><item><title>Apache mod_rewrite Debugging &amp; Log Analysis</title><link>https://www.netdata.cloud/academy/apache-http-server-mod/</link><pubDate>Wed, 23 Jul 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/apache-http-server-mod/</guid><description>&lt;p>You&amp;rsquo;ve crafted what seems like a perfect rewrite rule, uploaded it to your &lt;code>.htaccess&lt;/code> file, and refreshed your browser, only to be met with a &amp;ldquo;Too many redirects&amp;rdquo; error. You&amp;rsquo;re stuck in an Apache redirect loop, a frustrating and all-too-common problem when working with Apache mod_rewrite. This powerful module can manipulate URLs in almost any way imaginable, but its complexity means that a small mistake in a regular expression or a missing condition can lead to unexpected behavior, broken links, or infinite loops.&lt;/p></description></item><item><title>AWS ELB 5xx Surge Analysis: Health &amp; Draining</title><link>https://www.netdata.cloud/academy/aws-elb-surge-investigation/</link><pubDate>Wed, 23 Jul 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/aws-elb-surge-investigation/</guid><description>&lt;p>The alert notification chimes, and your dashboard lights up with an ELB 5xx error rate spike. Your application, running behind an AWS Application Load Balancer (ALB), is suddenly returning 502, 503, or 504 errors. This is a high-stakes scenario where every second of downtime impacts users. The load balancer is often just the messenger; the real culprit lies somewhere in your backend infrastructure. Is a target group unhealthy? Are your connections timing out? Is a misconfigured deployment process wreaking havoc?&lt;/p></description></item><item><title>25 NGINX Directives to Audit Before Opening a Support Ticket</title><link>https://www.netdata.cloud/academy/25-nginx-directives-to-audit/</link><pubDate>Wed, 09 Jul 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/25-nginx-directives-to-audit/</guid><description>&lt;p>You&amp;rsquo;ve built a robust application, but it&amp;rsquo;s sluggish or throwing intermittent errors. Your first instinct might be to blame the code, but often the culprit lies in the intricate web of your NGINX configuration. Before you spend hours debugging your application or drafting a support ticket, a thorough NGINX config check can save you time and reveal simple yet impactful misconfigurations.&lt;/p>
&lt;p>An NGINX configuration audit isn&amp;rsquo;t just for troubleshooting; it&amp;rsquo;s a proactive step toward better performance, tighter security, and greater stability. Many common issues related to performance bottlenecks, security vulnerabilities, or proxy errors stem from overlooked or poorly optimized directives. This checklist covers 25 essential directives you should review. Working through this list can help you pinpoint problems, implement best practices, and gain a deeper understanding of how NGINX operates.&lt;/p></description></item><item><title>Short vs Long Transactions In PostgreSQL Benchmarks</title><link>https://www.netdata.cloud/academy/benchmarking-short-long-transactions-in-postgresql/</link><pubDate>Wed, 09 Jul 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/benchmarking-short-long-transactions-in-postgresql/</guid><description>&lt;p>Every developer working with a database faces a fundamental design choice: how to structure their transactions. Should a complex business operation, involving multiple database writes, be wrapped in a single, all-or-nothing transaction? Or is it better to break it down into a series of smaller, faster commits? On the surface, the &lt;code>long_running_transaction&lt;/code> seems safer, guaranteeing perfect atomicity. But as traffic grows, a mysterious performance degradation can set in. Latency spikes, throughput drops, and users start seeing errors. The culprit is often not the queries themselves, but the invisible contention they create.&lt;/p></description></item><item><title>Netdata vs Auvik: Real-Time Auvik Alternative</title><link>https://www.netdata.cloud/comparisons/auvik/</link><pubDate>Tue, 24 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/auvik/</guid><description>Netdata vs Auvik comparison: real-time full-stack observability with ML vs network-centric MSP SaaS with NCM.</description></item><item><title>Netdata vs Cacti: Real-Time Cacti Alternative</title><link>https://www.netdata.cloud/comparisons/cacti/</link><pubDate>Tue, 24 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/cacti/</guid><description>Netdata vs Cacti comparison — real-time full-stack network observability vs legacy SNMP graphing.</description></item><item><title>Netdata vs ManageEngine OpManager: Which Is Better?</title><link>https://www.netdata.cloud/comparisons/opmanager/</link><pubDate>Tue, 24 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/opmanager/</guid><description>Netdata vs ManageEngine OpManager — real-time, unified, ML-powered observability vs traditional modular NMS.</description></item><item><title>Netdata vs NeDi: Network Discovery Tool Comparison</title><link>https://www.netdata.cloud/comparisons/nedi/</link><pubDate>Tue, 24 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/nedi/</guid><description>Compare Netdata and NeDi for network monitoring, topology, flow analysis, and full-stack observability.</description></item><item><title>Network Device Auto-Discovery</title><link>https://www.netdata.cloud/features/network/device-discovery/</link><pubDate>Tue, 24 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/features/network/device-discovery/</guid><description>Automatically discover, classify, and monitor every SNMP-reachable network device with zero manual MIB mapping.</description></item><item><title>Netdata Implements MCP Protocol</title><link>https://www.netdata.cloud/blog/netdata-mcp-server/</link><pubDate>Wed, 18 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-mcp-server/</guid><description>&lt;p>&lt;strong>Update (February 2026):&lt;/strong> Netdata now also provides MCP via Netdata Cloud at &lt;code>app.netdata.cloud/api/v1/mcp&lt;/code> for infrastructure-wide access (Business/Homelab plan). See &lt;a href="https://learn.netdata.cloud/docs/netdata-ai/mcp">MCP documentation&lt;/a> for details.&lt;/p>
&lt;hr>
&lt;p>We are excited to announce that Netdata has officially implemented the Model Context Protocol (MCP), joining the forefront of AI-powered infrastructure monitoring.&lt;/p>
&lt;p>By enabling direct connections between AI assistants and your data sources and tools, the Model Context Protocol is a new open standard that builds a crucial link between artificial intelligence and practical systems. Instead of the AI providing generic suggestions that don’t understand your environment, MCP enables it to communicate and then interact with your infrastructure data in real time.
Being among the first monitoring platforms to adopt this groundbreaking protocol, Netdata is at the forefront of intelligent observability. With the help of this integration, traditional monitoring becomes a dialogue that your entire technical team is able to participate in.
We encourage you to dive deeper into the technical foundations of MCP, explore &lt;a href="https://www.anthropic.com/news/model-context-protocol">Anthropic&amp;rsquo;s comprehensive&lt;/a> explanation, and discover how this protocol is reshaping the future of AI-data interaction.&lt;/p></description></item><item><title>How Can Generative AI (Gen AI) Be Used In Cybersecurity</title><link>https://www.netdata.cloud/academy/generative-ai-in-cybersecurity/</link><pubDate>Tue, 10 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/generative-ai-in-cybersecurity/</guid><description>&lt;p>Generative AI has exploded into the mainstream, changing how we create content, write code, and interact with technology. But beyond its creative applications, &lt;code>what is generative AI&lt;/code> doing in the high-stakes world of cybersecurity? The answer is complex. GenAI is rapidly becoming one of the most powerful tools for security professionals, but it&amp;rsquo;s also a formidable weapon in the hands of their adversaries.&lt;/p>
&lt;p>For security teams grappling with an ever-expanding attack surface and increasingly sophisticated threats, GenAI offers a way to automate, predict, and respond faster than ever before. However, understanding &lt;code>how generative AI can be used in cybersecurity&lt;/code> requires looking at both sides of the coin—its potential to fortify our defenses and its capacity to create entirely new kinds of attacks.&lt;/p></description></item><item><title>How To Find And Fix Memory Leaks in C or C++</title><link>https://www.netdata.cloud/academy/how-to-find-memory-leak-in-c/</link><pubDate>Tue, 10 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/how-to-find-memory-leak-in-c/</guid><description>&lt;p>Your application feels sluggish. It runs perfectly after a restart, but over hours or days, it slows to a crawl before eventually crashing. If you&amp;rsquo;re working with C or C++, this behavior is a classic symptom of a memory leak—a silent bug that can drain system resources and destabilize your services. Because these languages put memory management directly in your hands, understanding how to find and fix memory leaks is a critical skill for building robust, long-running applications.&lt;/p></description></item><item><title>How To Fix Packet Loss - A Step-by-Step Guide To Reduce It</title><link>https://www.netdata.cloud/academy/how-to-fix-packet-loss/</link><pubDate>Tue, 10 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/how-to-fix-packet-loss/</guid><description>&lt;p>You have a fast internet connection, but your online game is lagging, your video call is choppy, or your application feels unresponsive. You run a speed test, and everything looks fine—low ping, high bandwidth. So, what’s the problem? The likely culprit is &lt;code>packet loss&lt;/code>, a frustrating and often misunderstood network issue that can occur even on the best internet connections.&lt;/p>
&lt;p>Unlike high latency (ping), which is a measure of delay, &lt;code>packet loss meaning&lt;/code> refers to data that never arrives at its destination at all. These lost pieces of data, or &amp;ldquo;packets,&amp;rdquo; must be re-sent, creating hitches, stutters, and lag. This guide will walk you through what causes this issue and provide a clear, step-by-step process for &lt;code>how to fix packet loss&lt;/code>.&lt;/p></description></item><item><title>How To Secure Sensitive Data In Cloud Environments</title><link>https://www.netdata.cloud/academy/how-to-secure-sensitive-data-in-cloud/</link><pubDate>Tue, 10 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/how-to-secure-sensitive-data-in-cloud/</guid><description>&lt;p>Migrating to the cloud offers unparalleled benefits in scalability, flexibility, and collaboration. However, this shift also introduces new complexities and risks, especially when it comes to protecting your most valuable asset—your data. Securing sensitive data in the cloud is not a one-time task; it&amp;rsquo;s a continuous process that requires a deep understanding of your cloud environment, robust security practices, and a proactive mindset.&lt;/p>
&lt;p>Many organizations mistakenly assume their cloud service provider (CSP) handles all aspects of security. In reality, security is a shared responsibility. While the CSP is responsible for the security &lt;em>of&lt;/em> the cloud (i.e., the physical data centers and underlying infrastructure), you, the customer, are responsible for security &lt;em>in&lt;/em> the cloud. This includes securing your data, applications, and user access. Let&amp;rsquo;s explore the essential strategies for &lt;code>how to secure sensitive data in cloud environments&lt;/code>.&lt;/p></description></item><item><title>Kubernetes Security Posture Management (KSPM) Explained</title><link>https://www.netdata.cloud/academy/kubernetes-security-posture-management/</link><pubDate>Tue, 10 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/kubernetes-security-posture-management/</guid><description>&lt;p>Kubernetes has become the de facto standard for container orchestration, offering incredible power and scalability. But this complexity introduces a massive attack surface. A single misconfiguration in a YAML file, an overly permissive role, or a vulnerable container image can create a critical security gap. Traditional security tools, designed for static infrastructure, struggle to keep up with the dynamic, ephemeral nature of Kubernetes environments.&lt;/p>
&lt;p>This is where &lt;code>Kubernetes Security Posture Management&lt;/code> (KSPM) comes in. It’s not just another security tool; it&amp;rsquo;s a specialized practice designed to continuously monitor, assess, and enforce security policies across your entire Kubernetes ecosystem. For any team running mission-critical workloads on Kubernetes, understanding and implementing &lt;code>KSPM&lt;/code> is no longer a luxury—it’s a necessity.&lt;/p></description></item><item><title>OpenSearch vs Elasticsearch Which One Is Better In 2025?</title><link>https://www.netdata.cloud/academy/elasticsearch-vs-opensearch/</link><pubDate>Tue, 10 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/elasticsearch-vs-opensearch/</guid><description>&lt;p>If you&amp;rsquo;re building a system for log analytics, application search, or security monitoring, you&amp;rsquo;ve almost certainly faced a critical decision: OpenSearch vs Elasticsearch. For years, Elasticsearch was the undisputed king, but since 2021, its open-source fork, OpenSearch, has emerged as a powerful and popular alternative. What began as a dispute over licensing has evolved into two distinct projects with different philosophies, features, and performance characteristics.&lt;/p>
&lt;p>Choosing between them is no longer a simple matter. It&amp;rsquo;s a strategic decision that impacts everything from your budget and licensing compliance to your system&amp;rsquo;s performance and future scalability. As we move through 2025, the paths of these two search engines have diverged enough that making an informed choice is more important than ever. This guide breaks down the key differences to help you decide which platform is right for your needs.&lt;/p></description></item><item><title>What Is A Flaky Test How To Detect Fix &amp; Avoid Them</title><link>https://www.netdata.cloud/academy/flaky-tests/</link><pubDate>Tue, 10 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/flaky-tests/</guid><description>&lt;p>You push a new feature, all local tests pass, and you open a pull request. The continuous integration (CI) pipeline kicks off, but a few minutes later, you see a dreaded red &amp;lsquo;X&amp;rsquo;. A test failed. You scrutinize your code, find nothing wrong, and re-run the job. This time, it passes with a green checkmark. If this scenario feels familiar, you&amp;rsquo;ve encountered a flaky test.&lt;/p>
&lt;p>A flaky test is a test that exhibits non-deterministic behavior—it can both pass and fail across multiple runs without any changes to the code or its environment. While it might seem like a minor annoyance, test flakiness is a significant problem that can erode your team&amp;rsquo;s confidence in your test suite, slow down development velocity, and ultimately allow real bugs to slip into production. Understanding what causes these inconsistencies is the first step toward building a more reliable and trustworthy testing process.&lt;/p></description></item><item><title>What Is A Memory Leak In Java How To Detect And Fix Them</title><link>https://www.netdata.cloud/academy/java-memory-leak/</link><pubDate>Tue, 10 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/java-memory-leak/</guid><description>&lt;p>Your Java application runs smoothly after a fresh deploy, but over hours or days, its performance steadily degrades. Response times creep up, garbage collection pauses become longer and more frequent, and then, the inevitable happens: the application crashes, logging a fatal OutOfMemoryError. This classic scenario is often the calling card of a subtle but dangerous problem—a memory leak.&lt;/p>
&lt;p>Even though Java features automatic memory management via its garbage collector (GC), applications are not immune to leaks. A Java memory leak occurs when objects are no longer in use by the application, but the GC is unable to reclaim their memory because they are still being referenced. Over time, these orphaned objects accumulate, consuming the available heap space and leading to performance degradation and eventual failure. Understanding how to detect and fix these leaks is a critical skill for any Java developer.&lt;/p></description></item><item><title>What Is Distributed Tracing How It Works And Use Cases</title><link>https://www.netdata.cloud/academy/distributed-tracing/</link><pubDate>Tue, 10 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/distributed-tracing/</guid><description>&lt;p>Your team has deployed a new feature, but users are reporting that a specific action in your application is painfully slow. You look at the metrics for your frontend service, and they seem fine. You check the logs for the user authentication service, and there are no errors. The database CPU usage is normal. So where is the bottleneck? In a modern distributed architecture built with microservices, a single user request can trigger a complex chain reaction across dozens of independent services. Finding the root cause of a problem can feel like searching for a needle in a global haystack.&lt;/p></description></item><item><title>Agentless Network Monitoring A Complete Guide To Understand</title><link>https://www.netdata.cloud/academy/agentless-network-monitoring/</link><pubDate>Mon, 09 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/agentless-network-monitoring/</guid><description>&lt;p>As your infrastructure grows, so does the complexity of monitoring it. The traditional approach often involves deploying a software &amp;ldquo;agent&amp;rdquo; on every server, virtual machine, and container. While effective, this can quickly become a management nightmare. You have to install, configure, update, and secure hundreds or even thousands of agents, each &lt;a href="https://www.netdata.cloud/academy/how-to-fix-cpu-overload/">consuming precious CPU&lt;/a> and memory on the hosts you&amp;rsquo;re trying to monitor. What if there was a less intrusive way?&lt;/p></description></item><item><title>Garbage Collection In Java What It Is and How It Works</title><link>https://www.netdata.cloud/academy/java-garbage-collection/</link><pubDate>Mon, 09 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/java-garbage-collection/</guid><description>&lt;p>One of the most powerful features of the Java platform is its automatic memory management. Unlike languages like &lt;a href="https://www.netdata.cloud/academy/how-to-find-memory-leak-in-c/">C or C++&lt;/a>, where developers must manually allocate and deallocate memory, Java handles this process for you through a process called garbage collection (GC). This frees developers to focus on application logic rather than the complexities of memory management, which is a major reason for Java&amp;rsquo;s enduring popularity.&lt;/p>
&lt;p>But what exactly is garbage collection in Java, and how does it work under the hood? While it&amp;rsquo;s an automatic process, a solid understanding of the Java garbage collector is crucial for writing high-performance, stable applications and for troubleshooting &lt;a href="https://www.netdata.cloud/academy/java-memory-leak/">memory-related issues&lt;/a> like the dreaded &lt;code>OutOfMemoryError&lt;/code>.&lt;/p></description></item><item><title>How To View Docker Container Logs A Step-by-Step Guide</title><link>https://www.netdata.cloud/academy/docker-logs/</link><pubDate>Mon, 09 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/docker-logs/</guid><description>&lt;p>Your containerized application is misbehaving. The service is unresponsive, or worse, it&amp;rsquo;s &lt;a href="https://www.netdata.cloud/guides/docker/docker-container-exits-immediately/">crash-looping&lt;/a>. As a developer or SRE, your first instinct is to ask, &amp;ldquo;What do the logs say?&amp;rdquo; For applications running in &lt;a href="https://www.netdata.cloud/guides/docker/">Docker&lt;/a>, accessing and understanding container logs is the most fundamental troubleshooting skill you can possess. These logs are the raw, unfiltered story of what your application is doing, thinking, and feeling.&lt;/p>
&lt;p>But &lt;code>docker logging&lt;/code> is more than just a single action. It&amp;rsquo;s a comprehensive system with layers of functionality, from quickly tailing real-time output to implementing robust, production-grade log management strategies. This guide will walk you through everything you need to know to effectively &lt;code>view docker container logs&lt;/code>, starting with the basics and progressing to the best practices that will keep your applications observable and your on-call nights quiet.&lt;/p></description></item><item><title>Logging As A Service (LaaS): Simplify Log Management</title><link>https://www.netdata.cloud/academy/logging-as-a-service/</link><pubDate>Mon, 09 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/logging-as-a-service/</guid><description>&lt;p>In the world of modern software development, logs are the lifeblood of observability. They are the detailed, chronological record of every event, error, and transaction that occurs within your applications and infrastructure. When something goes wrong, logs are the first place developers and SREs turn to for answers. But with the rise of microservices, containers, and distributed cloud architectures, the sheer volume and complexity of log data have exploded.&lt;/p>
&lt;p>The old way of managing logs—SSHing into a server and using &lt;code>grep&lt;/code> to search through a text file—simply doesn&amp;rsquo;t scale. This approach becomes an impossible, time-consuming, and error-prone task when you have hundreds of services spread across dozens of servers. This is the problem that Logging as a Service (LaaS) was created to solve.&lt;/p></description></item><item><title>What Is AIOps (Artificial Intelligence For IT Operations)</title><link>https://www.netdata.cloud/academy/aiops/</link><pubDate>Mon, 09 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/aiops/</guid><description>&lt;p>Your team is drowning in data. Your monitoring systems generate thousands of alerts every day, most of which are just noise. When a real issue strikes your complex, distributed application, engineers scramble, jumping between a dozen different dashboards to piece together clues. This frantic, reactive firefighting is the reality for many IT Operations, DevOps, and SRE teams. It’s stressful, inefficient, and unsustainable.&lt;/p>
&lt;p>This is where AIOps (Artificial Intelligence for IT Operations) enters the picture. It&amp;rsquo;s not just another industry buzzword; it&amp;rsquo;s a fundamental shift in how we manage technology. AIOps is the practice of applying artificial intelligence—specifically machine learning and big data analytics—to automate and streamline every aspect of IT operations. By intelligently processing the massive volumes of data your systems generate, AIOps platforms help you move from a reactive state of constant crisis to a proactive, predictive model of management.&lt;/p></description></item><item><title>What Is Container Orchestration &amp; Why Do We Need It</title><link>https://www.netdata.cloud/academy/container-orchestaration/</link><pubDate>Mon, 09 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/container-orchestaration/</guid><description>&lt;p>If you&amp;rsquo;re building modern applications, you&amp;rsquo;re likely using containers. They&amp;rsquo;re lightweight, portable, and provide a consistent environment for your code. Deploying a single container is simple. Managing a handful is manageable. But what happens when your application grows from a few containers to hundreds or even thousands, all working together as a complex system of microservices?&lt;/p>
&lt;p>This is where manual management breaks down. You face a storm of questions: How do you deploy updates without downtime? How do you scale to handle a sudden traffic surge? What happens if a server fails in the middle of the night? Answering these questions manually is an operational nightmare. This is precisely the problem that container orchestration was created to solve.&lt;/p></description></item><item><title>What Is OpenTelemetry How It Works Benefits and Use Cases</title><link>https://www.netdata.cloud/academy/opentelemetry/</link><pubDate>Mon, 09 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/opentelemetry/</guid><description>&lt;p>If you manage modern applications, you&amp;rsquo;re dealing with complexity. Your system isn&amp;rsquo;t a single program running on one server; it&amp;rsquo;s a distributed web of microservices, serverless functions, and third-party APIs running across a hybrid cloud infrastructure. Understanding what&amp;rsquo;s happening inside this complex system is a monumental challenge. Each component and vendor has its own proprietary way of exporting performance data, leaving you to juggle dozens of agents and formats, leading to data silos and vendor lock-in.&lt;/p></description></item><item><title>What Is Remote Infrastructure Management (RIM)</title><link>https://www.netdata.cloud/academy/remote-infrastructure-management/</link><pubDate>Mon, 09 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/remote-infrastructure-management/</guid><description>&lt;p>The days of managing servers neatly racked in a single on-premise data center are over for most businesses. Today’s IT landscape is a sprawling mix of public cloud instances, private cloud infrastructure, edge devices, and SaaS applications. This distribution is powerful, but it creates a significant management challenge: How do you maintain control, ensure uptime, and secure an infrastructure that’s physically everywhere and nowhere at once?&lt;/p>
&lt;p>This is the problem that &lt;code>Remote Infrastructure Management&lt;/code> (RIM) solves. It’s a strategic approach that moves IT operations from a hands-on, on-site model to a centralized, remote-first paradigm. For DevOps and SRE teams, understanding RIM is no longer optional; it&amp;rsquo;s fundamental to building and maintaining resilient, scalable, and efficient systems in the modern era.&lt;/p></description></item><item><title>Whats The Difference Between PostgreSQL vs MySQL</title><link>https://www.netdata.cloud/academy/postgres-mysql-differences/</link><pubDate>Mon, 09 Jun 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/postgres-mysql-differences/</guid><description>&lt;p>When building an application, one of the most fundamental decisions you&amp;rsquo;ll make is choosing the right database. For many developers, this choice boils down to two of the most popular open-source relational databases in the world: PostgreSQL and MySQL. Both are powerful, reliable, and backed by strong communities, but they operate on different philosophies and excel in different areas.&lt;/p>
&lt;p>Making the right choice between &lt;code>PostgreSQL vs MySQL&lt;/code> isn&amp;rsquo;t about finding a single &amp;ldquo;best&amp;rdquo; database, but understanding which one aligns with your application&amp;rsquo;s specific needs. This guide breaks down the critical differences in their architecture, features, performance characteristics, and ideal use cases to help you make an informed decision for your next project.&lt;/p></description></item><item><title>Introducing Netdata Insights</title><link>https://www.netdata.cloud/blog/netdata-insights/</link><pubDate>Tue, 27 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-insights/</guid><description>&lt;p>We&amp;rsquo;ve been thinking a lot about synthesis lately.&lt;/p>
&lt;p>Netdata already samples every metric every second at the edge. Engineers told us the remaining pain point was synthesis, the ability to pull hours or days or months of high‑resolution time‑series into a concise explanation they could hand to a teammate (or use themselves to debug faster).&lt;/p>
&lt;!--truncate-->
&lt;p>You know the pattern. An incident happens, and suddenly you&amp;rsquo;re context-switching between dozens of dashboards, trying to reconstruct a timeline. Or you need to write a capacity planning report, and you&amp;rsquo;re copy-pasting screenshots into slides, manually correlating trends across different retention windows. The raw data is there, but the synthesis step (the part where you turn metrics into narrative) doesn&amp;rsquo;t scale.&lt;/p></description></item><item><title>What Is Canary Deployment? Benefits, Metrics &amp; Setup</title><link>https://www.netdata.cloud/academy/canary-deployment/</link><pubDate>Tue, 27 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/canary-deployment/</guid><description>&lt;p>Releasing new software versions can be a nerve-wracking experience. Even with rigorous testing, the real world of production traffic often uncovers unforeseen issues. A problematic deployment can lead to downtime, frustrated users, and a frantic scramble to roll back. This is where a canary deployment strategy shines, offering a more cautious and controlled approach to rolling out updates.&lt;/p>
&lt;p>Instead of a big-bang release, a canary release exposes the new version to a small subset of users first, allowing you to monitor its performance and gather feedback before a full-scale rollout. This technique significantly de-risks the deployment process, especially in complex environments like Kubernetes.&lt;/p></description></item><item><title>What Is Structured Logging And How To Utilize It Effectively</title><link>https://www.netdata.cloud/academy/what-is-structured-logging/</link><pubDate>Tue, 27 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-structured-logging/</guid><description>&lt;p>Dealing with gigabytes of plain text log files can feel like searching for a needle in a haystack, especially when critical systems are down. Unstructured logs make it incredibly difficult to quickly pinpoint issues, correlate events, or gain meaningful insights from the vast amount of data your applications generate. If you&amp;rsquo;ve ever struggled to filter logs by a specific customer ID or trace a transaction across multiple services, you understand the pain. This is precisely where structured logging steps in, transforming your log data from a chaotic stream of text into a valuable, queryable asset.&lt;/p></description></item><item><title>Nodejs Memory Leak How To Identify Debug And Avoid Them</title><link>https://www.netdata.cloud/academy/nodejs-memory-leak/</link><pubDate>Mon, 26 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/nodejs-memory-leak/</guid><description>&lt;p>In the fast-paced world of Node.js development, performance and reliability are non-negotiable. However, a silent saboteur often lurks in the shadows – the Node.js memory leak. These insidious issues can gradually degrade your application&amp;rsquo;s performance, leading to slowdowns, crashes, and frustrated users. Understanding how to effectively identify, debug, and &lt;a href="https://www.netdata.cloud/academy/how-to-find-memory-leak-in-c/">prevent memory leaks&lt;/a> is a critical skill for any developer, DevOps engineer, or SRE working with Node.js. This guide will walk you through the intricacies of Node.js memory management and equip you with the knowledge to tackle these challenging problems.&lt;/p></description></item><item><title>What Is A Bare Metal Server? Benefits &amp; Use Cases</title><link>https://www.netdata.cloud/academy/bare-metal-server/</link><pubDate>Mon, 26 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/bare-metal-server/</guid><description>&lt;p>In a world increasingly dominated by virtualized cloud environments, the concept of a bare metal server might seem like a throwback. However, these physical powerhouses continue to play a critical role in modern IT infrastructure, offering distinct advantages for specific workloads. Understanding what is a bare metal server and why you might choose one over virtualized alternatives is key to designing efficient and high-performing systems.&lt;/p>
&lt;h2 id="what-is-a-bare-metal-server-definition--how-it-works">What Is A Bare Metal Server? Definition &amp;amp; How It Works&lt;/h2>
&lt;p>At its core, a bare metal server is a physical computer server dedicated entirely to a single customer or tenant. Unlike virtual machines (VMs) that share the resources of a physical host through a hypervisor layer, a bare metal server gives the user direct, unmediated access to the underlying hardware – the &amp;ldquo;bare metal.&amp;rdquo; This means you have exclusive use of the server&amp;rsquo;s CPU, RAM, storage, and network interface cards.&lt;/p></description></item><item><title>Application Monitoring Best Practices In 2025</title><link>https://www.netdata.cloud/academy/application-monitoring-2025/</link><pubDate>Sun, 25 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/application-monitoring-2025/</guid><description>&lt;p>In today&amp;rsquo;s digitally driven world, application performance is not just a technical concern; it&amp;rsquo;s a critical business imperative. Users expect applications to be fast, reliable, and seamlessly available. Any deviation can lead to frustrated users, lost revenue, and damaged reputation. This makes application monitoring a cornerstone of successful software development and operations. It&amp;rsquo;s the continuous process of collecting, analyzing, and acting on data to ensure your applications meet performance, availability, and user experience expectations.&lt;/p></description></item><item><title>The Three Pillars Of Observability Logs Metrics And Traces</title><link>https://www.netdata.cloud/academy/pillars-of-observability/</link><pubDate>Sun, 25 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/pillars-of-observability/</guid><description>&lt;p>In the realm of modern, complex IT systems, especially those involving microservices, cloud-native architectures, and distributed environments, understanding what&amp;rsquo;s happening internally is paramount. This is where observability comes into play. Observability isn&amp;rsquo;t just about monitoring; it&amp;rsquo;s the ability to infer the internal state and health of a system by examining its outputs. To achieve this, we rely on what are commonly known as the three pillars of observability: logs, metrics, and traces.&lt;/p></description></item><item><title>What Is Event Correlation Benefits Use Cases And Techniques</title><link>https://www.netdata.cloud/academy/event-correlation/</link><pubDate>Sun, 25 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/event-correlation/</guid><description>&lt;p>In today&amp;rsquo;s complex and dynamic IT environments, organizations are inundated with a massive volume of events generated by countless sources – applications, servers, network devices, security systems, and more. This flood of data, while rich in potential insights, can easily become overwhelming. The core challenge lies in sifting through this &amp;ldquo;sea of data&amp;rdquo; to identify events that truly matter and understand their relationships. This is precisely where event correlation becomes indispensable. It&amp;rsquo;s the process of sensing and analyzing relationships between disparate events to uncover meaningful patterns, diagnose root causes, and enable proactive responses.&lt;/p></description></item><item><title>What Is Synthetic Transaction Monitoring</title><link>https://www.netdata.cloud/academy/synthetic-monitoring/</link><pubDate>Sun, 25 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/synthetic-monitoring/</guid><description>&lt;p>In today&amp;rsquo;s digital-first world, the performance and availability of web applications are paramount to business success. Users expect seamless, fast, and reliable interactions, and any hiccup can lead to frustration, lost revenue, and a tarnished brand reputation. This is where synthetic transaction monitoring (STM) emerges as a critical proactive strategy. Unlike monitoring that relies on actual user traffic, STM uses simulated user journeys to continuously test and validate application functionality, performance, and availability, even before real users encounter issues.&lt;/p></description></item><item><title>What Is Blue-Green Deployment? Benefits &amp; Process</title><link>https://www.netdata.cloud/academy/blue-green-deployment/</link><pubDate>Sat, 24 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/blue-green-deployment/</guid><description>&lt;p>Software teams today face the challenge of releasing updates quickly without disrupting users. Traditional deployment methods often cause downtime and raise the risk of bugs reaching production. Blue green deployment offers a solution: by running two identical environments and switching traffic seamlessly, teams can deliver new versions with minimal risk and zero downtime. In this guide, we’ll explain what blue green deployment is, how it works, its benefits and challenges, and the best practices for implementing it.&lt;/p></description></item><item><title>LLM Observability and Monitoring A Comprehensive Guide</title><link>https://www.netdata.cloud/academy/llm-observability/</link><pubDate>Wed, 21 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/llm-observability/</guid><description>&lt;p>The rapid proliferation of Large Language Models (LLMs) like GPT and LLaMA is transforming how businesses operate, from enhancing office productivity with tools like Microsoft Copilot to innovative fraud management at Stripe. However, deploying and managing these powerful LLM applications in production environments presents unique hurdles. The sheer size of these models, their complex architectures, and often non-deterministic outputs make them challenging to manage. If you&amp;rsquo;re a developer, DevOps engineer, or SRE, you understand that ensuring consistent performance, security, and accuracy in LLM-driven systems requires a new level of insight. This is where LLM observability and LLM monitoring become indispensable.&lt;/p></description></item><item><title>Understanding Digital Experience Monitoring DEM</title><link>https://www.netdata.cloud/academy/digital-experience-monitoring/</link><pubDate>Sun, 18 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/digital-experience-monitoring/</guid><description>&lt;p>In today&amp;rsquo;s digital-first world, the way users interact with your online services defines their perception of your brand. A slow-loading webpage, a confusing mobile app, or an unavailable critical feature can quickly lead to frustration and churn. This is where Digital Experience Monitoring (DEM) becomes indispensable, offering a critical lens into how users truly perceive and interact with your digital platforms. Understanding and optimizing these interactions is no longer a luxury but a necessity for businesses aiming to thrive.&lt;/p></description></item><item><title>Understanding Error Budgets And Their Importance In SRE</title><link>https://www.netdata.cloud/academy/error-budget/</link><pubDate>Sun, 18 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/error-budget/</guid><description>&lt;p>In the quest for flawless digital experiences, the reality is that 100% uptime is an elusive, if not impossible, goal. Systems inevitably encounter issues, and services can experience disruptions. This is where the concept of an &lt;strong>error budget&lt;/strong> becomes a cornerstone for modern &lt;a href="https://www.netdata.cloud/academy/sre-vs-devops-what-are-the-main-differences-between-them/">Site Reliability Engineering (SRE) and DevOps&lt;/a> practices. Understanding and effectively managing your &lt;strong>error budget&lt;/strong> can mean the difference between fostering innovation and constantly firefighting.&lt;/p>
&lt;p>So, &lt;strong>what is an error budget&lt;/strong>? Simply put, it&amp;rsquo;s the quantifiable amount of unreliability or downtime that a service can tolerate over a specific period without breaching its Service Level Objectives (SLOs) or upsetting users. It&amp;rsquo;s the acknowledged margin for error, a critical component in balancing the drive for new features with the imperative of maintaining a stable and reliable service. For SRE teams, the &lt;strong>sre error budget&lt;/strong> is not just a metric- it&amp;rsquo;s a vital tool for decision-making.&lt;/p></description></item><item><title>A Comprehensive Guide To Database Performance Optimization</title><link>https://www.netdata.cloud/academy/database-performance/</link><pubDate>Sat, 17 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/database-performance/</guid><description>&lt;p>Sluggish database performance can be a silent killer for applications, leading to frustrated users, missed opportunities, and a direct hit to your bottom line. In today&amp;rsquo;s data-intensive environments, ensuring your database operates at peak efficiency isn&amp;rsquo;t just a technical task- it&amp;rsquo;s a critical business imperative. If you&amp;rsquo;re grappling with slow queries, high resource consumption, or concerns about &lt;code>data scalability&lt;/code>, this guide will walk you through essential &lt;code>database optimization&lt;/code> strategies.&lt;/p>
&lt;h2 id="tldr-summary">TL;DR Summary&lt;/h2>
&lt;ul>
&lt;li>Database performance issues usually come down to a few repeat offenders: slow or poorly designed queries, CPU pressure, disk I/O bottlenecks, weak indexing, and locking or concurrency contention.&lt;/li>
&lt;li>To manage and improve performance, track core metrics like response time, throughput, scalability, and resource utilization (CPU, memory, disk I/O), then optimize iteratively rather than treating tuning as a one-time fix.&lt;/li>
&lt;li>The biggest wins typically come from query tuning (EXPLAIN/EXPLAIN ANALYZE, rewriting queries, caching) and smart indexing (right index types, best-practice column selection, and regular maintenance).&lt;/li>
&lt;li>For growth and resilience, use techniques like partitioning, sharding, connection pooling, read replicas, HA setups, and archiving via history tables, backed by continuous monitoring (for example, real-time per-second visibility and alerts) to catch issues before users feel them.&lt;/li>
&lt;/ul>
&lt;h2 id="understanding-the-roots-of-database-performance-issues">Understanding The Roots Of Database Performance Issues&lt;/h2>
&lt;p>Before diving into solutions, it&amp;rsquo;s crucial to understand what &lt;code>database performance&lt;/code> truly means and what typically causes it to degrade. Performance refers to the speed and efficiency with which your database handles queries, transactions, and data retrieval. When performance suffers, it’s often due to one or more common culprits.&lt;/p></description></item><item><title>OpenShift vs Kubernetes What Are The Differences</title><link>https://www.netdata.cloud/academy/openshift-vs-kubernetes/</link><pubDate>Fri, 16 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/openshift-vs-kubernetes/</guid><description>&lt;p>Container orchestration is a cornerstone of modern cloud-native application development, and when it comes to managing containers at scale, two names often dominate the conversation- OpenShift vs Kubernetes. While both platforms are designed to &lt;a href="https://www.netdata.cloud/academy/deployment-automation/">automate the deployment&lt;/a>, scaling, and management of containerized applications, they have distinct characteristics and cater to slightly different needs. Understanding the difference between OpenShift and Kubernetes is crucial for organizations making strategic decisions about their containerization strategy.&lt;/p></description></item><item><title>SOC 2 Type 1: Committed To Security &amp; Trust</title><link>https://www.netdata.cloud/blog/soc2-type1/</link><pubDate>Fri, 16 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/soc2-type1/</guid><description>&lt;p>We are pleased to announce that Netdata has successfully achieved SOC 2 Type 1 attestation!&lt;/p>
&lt;!--truncate-->
&lt;p>Following an independent examination performed by AssuranceLab CPAs LLC, the report confirms that—as of April 25, 2025—the design of Netdata’s controls meets the Security, Availability, and Confidentiality Trust Services Criteria defined by the AICPA.&lt;/p>
&lt;p>At Netdata, the security and integrity of the monitoring data our users entrust to us are paramount. This significant milestone, validated through a rigorous, independent third-party audit conducted by AssuranceLab, formally attests to the robustness of our security controls and practices as designed and implemented at a specific point in time.&lt;/p></description></item><item><title>What Is Apache Kafka Used For Everything You Need To Know</title><link>https://www.netdata.cloud/academy/what-is-apache-kafka/</link><pubDate>Fri, 16 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-apache-kafka/</guid><description>&lt;p>In today&amp;rsquo;s data-driven world, applications generate vast streams of information that need to be processed and acted upon in real time. Managing these high-velocity, high-volume data flows presents a significant challenge for developers, DevOps engineers, and Site Reliability Engineers (SREs). This is precisely where Apache Kafka shines, offering a powerful platform for building robust, scalable, and real-time data pipelines. Understanding what Apache Kafka is used for and how its underlying Kafka technology works is crucial for anyone looking to harness the power of event streaming.&lt;/p></description></item><item><title>What Is Incident Management Benefits Process Best Practices</title><link>https://www.netdata.cloud/academy/what-is-incident-management/</link><pubDate>Wed, 07 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-incident-management/</guid><description>&lt;p>When your critical services face unexpected disruptions, the clock starts ticking. For developers, DevOps engineers, and Site Reliability Engineers (SREs), understanding &lt;strong>what is incident management&lt;/strong> is paramount. A slow or disorganized response not only impacts users but can also strain resources and damage your organization&amp;rsquo;s reputation. Effectively managing these events is key to maintaining system stability and ensuring business continuity.&lt;/p>
&lt;p>&lt;strong>Incident management&lt;/strong> is the set of actions an organization takes to identify, analyze, correct, and prevent future occurrences of service disruptions or losses in operations. An &amp;ldquo;incident,&amp;rdquo; in ITIL terms, is any event that disrupts, or could disrupt, a service. This could range from a complete application outage to a web server running slowly, impacting productivity and posing a risk of total failure. The primary goal of &lt;strong>IT incident management&lt;/strong> is to restore normal service operation as quickly as possible and minimize the adverse impact on business operations.&lt;/p></description></item><item><title>What Is Database Concurrency? Problems &amp; Control Techniques</title><link>https://www.netdata.cloud/academy/what-is-database-concurrency/</link><pubDate>Sun, 04 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-database-concurrency/</guid><description>&lt;p>Imagine trying to book the very last seat on a popular flight online. At the exact same moment, another person clicks &amp;ldquo;confirm&amp;rdquo; for the same seat. How does the system ensure only one booking goes through and the database remains accurate? This scenario highlights the core challenge of &lt;strong>database concurrency&lt;/strong>.&lt;/p>
&lt;p>In today&amp;rsquo;s world, &lt;a href="https://www.netdata.cloud/academy/what-is-application-performance-monitoring-apm/">almost every application interacts with databases&lt;/a> accessed by multiple users or processes simultaneously. &lt;strong>Database concurrency&lt;/strong> is the ability of a Database Management System (DBMS) to handle these simultaneous operations efficiently while maintaining data integrity and consistency. For developers, DevOps engineers, and SREs, understanding concurrency is vital for building reliable and performant applications. Let&amp;rsquo;s explore what concurrency entails, the problems it can cause if unmanaged, and the techniques used to control it.&lt;/p></description></item><item><title>Normalized vs Denormalized - Choosing The Right Data Model</title><link>https://www.netdata.cloud/academy/normalized-vs-denormalized/</link><pubDate>Sat, 03 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/normalized-vs-denormalized/</guid><description>&lt;p>When designing databases, one of the fundamental decisions you&amp;rsquo;ll face is how to structure your data. Two primary approaches dominate this discussion: normalization and denormalization. Choosing between a &lt;strong>normalized vs denormalized&lt;/strong> model significantly impacts data integrity, storage efficiency, and query performance. Understanding this trade-off is crucial for developers, database administrators, and SREs responsible for building and maintaining reliable, efficient systems.&lt;/p>
&lt;p>Getting the data model right from the start can save considerable headaches down the line. Let&amp;rsquo;s explore what &lt;strong>normalized data&lt;/strong> and &lt;strong>denormalized data&lt;/strong> mean, their respective strengths and weaknesses, and how to decide which strategy best fits your specific needs.&lt;/p></description></item><item><title>Cloud Managed Services: Definition, Types &amp; Benefits</title><link>https://www.netdata.cloud/academy/what-are-cloud-managed-services/</link><pubDate>Fri, 02 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-are-cloud-managed-services/</guid><description>&lt;p>Moving to the cloud offers incredible benefits like scalability and flexibility, but managing cloud infrastructure isn&amp;rsquo;t always simple. Configuring networks, ensuring security, optimizing costs, performing maintenance, and staying compliant can quickly become overwhelming, especially for growing teams or those new to the cloud landscape. This is where &lt;strong>cloud managed services&lt;/strong> come into play.&lt;/p>
&lt;p>Understanding &lt;strong>managed cloud services&lt;/strong> is essential for developers, DevOps engineers, and SREs. It represents a strategic approach to handling cloud complexity, allowing technical teams to offload operational burdens and focus on innovation and core business objectives. Let&amp;rsquo;s explore what these services entail, how they work, and their potential advantages and disadvantages.&lt;/p></description></item><item><title>How To Check Your Firewall Logs On Windows</title><link>https://www.netdata.cloud/academy/how-to-check-firewall-logs-on-windows/</link><pubDate>Thu, 01 May 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/how-to-check-firewall-logs-on-windows/</guid><description>&lt;p>Every Windows system comes equipped with a built-in firewall, a critical component of its security posture. The &lt;strong>Microsoft Defender Firewall&lt;/strong> (previously Windows Firewall) acts as a gatekeeper, controlling incoming and outgoing network traffic based on predefined rules. While it diligently protects your system, its default configuration doesn&amp;rsquo;t tell you much about the traffic it&amp;rsquo;s allowing or blocking.&lt;/p>
&lt;p>This is where &lt;strong>Windows Firewall logs&lt;/strong> come in. These logs record detailed information about the firewall&amp;rsquo;s activity, providing invaluable insights for &lt;a href="https://www.netdata.cloud/academy/what-is-uptime-monitoring/">troubleshooting network connectivity problems&lt;/a>, identifying potential security threats, and ensuring compliance. For developers, DevOps engineers, and SREs, knowing &lt;strong>how to check firewall logs&lt;/strong> is a fundamental skill for &lt;a href="https://www.netdata.cloud/blog/windows-monitoring-improvements/">maintaining secure and reliable Windows environments&lt;/a>, whether on workstations or &lt;strong>Windows Server&lt;/strong> instances. This guide will walk you through enabling, locating, interpreting, and managing these essential logs.&lt;/p></description></item><item><title>Cloud Workload - Definition, Types &amp; Challenges</title><link>https://www.netdata.cloud/academy/cloud-workload/</link><pubDate>Wed, 30 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/cloud-workload/</guid><description>&lt;p>Cloud computing has transformed how businesses operate, offering unprecedented scalability, flexibility, and access to powerful resources. As organizations increasingly migrate applications and services to the cloud, you&amp;rsquo;ll frequently encounter the term &amp;ldquo;cloud workload.&amp;rdquo; But what is a cloud workload exactly? Understanding this concept is crucial, especially for DevOps engineers, SREs, and developers tasked with deploying, managing, and optimizing applications in cloud environments.&lt;/p>
&lt;p>A workload in cloud computing is essentially the specific amount of processing or computing task assigned to or running on cloud resources at any given time. It&amp;rsquo;s the fundamental unit of work – an application, a service, a set of processes – that consumes cloud resources like compute, storage, and networking. Misunderstanding or mismanaging these workloads can lead to performance issues, security vulnerabilities, and runaway costs. This guide dives into the cloud workload definition, explores the various types of workloads, and highlights the common challenges associated with their management.&lt;/p></description></item><item><title>Industrial Remote Monitoring: Process &amp; Examples</title><link>https://www.netdata.cloud/academy/industrial-remote-monitoring/</link><pubDate>Wed, 30 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/industrial-remote-monitoring/</guid><description>&lt;p>Imagine needing to know the exact operating temperature of a critical pump in a remote processing plant, or wanting to predict if a crucial piece of machinery on your factory floor is about to fail – all without sending a technician on-site. This is the power of &lt;strong>industrial remote monitoring&lt;/strong>. In today&amp;rsquo;s competitive landscape, industries from manufacturing to energy to logistics are under constant pressure to improve efficiency, reduce costs, and ensure safety. &lt;strong>Remote machine monitoring&lt;/strong> provides the critical visibility and data needed to achieve these goals.&lt;/p></description></item><item><title>Ecommerce Infrastructure: Components &amp; Benefits</title><link>https://www.netdata.cloud/academy/ecommerce-infrastructure/</link><pubDate>Tue, 29 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/ecommerce-infrastructure/</guid><description>&lt;p>The world of online shopping has exploded, offering convenience and accessibility like never before. Businesses can reach global audiences, and consumers can browse and buy from anywhere. But behind every successful online store, from small boutiques to giants like Amazon, lies a complex system working tirelessly: the &lt;strong>ecommerce infrastructure&lt;/strong>. Without this foundation, websites crash during peak traffic, customer data gets compromised, and orders get lost – leading to frustrated customers and lost revenue.&lt;/p></description></item><item><title>Database Backup - Types, Process and Benefits</title><link>https://www.netdata.cloud/academy/what-is-database-backup/</link><pubDate>Mon, 28 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-database-backup/</guid><description>&lt;p>Data is often called the lifeblood of modern organizations. From customer details and financial records to &lt;a href="https://www.netdata.cloud/solutions/webserver-monitoring/">application configurations and operational logs&lt;/a>, databases store the critical information that powers business operations. But what happens if that data is lost due to hardware failure, accidental deletion, software corruption, or a cyberattack? Without a safety net, the consequences can be catastrophic, leading to costly downtime, reputational damage, and potentially irreparable business harm. This is where &lt;strong>database backup&lt;/strong> becomes indispensable.&lt;/p></description></item><item><title>Server Security - What It Is &amp; Why It Is So Important</title><link>https://www.netdata.cloud/academy/server-security/</link><pubDate>Fri, 25 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/server-security/</guid><description>&lt;p>Servers are the workhorses of the modern digital world. They store critical data, run essential applications, host websites, and power the services we rely on every day. But with great power comes great responsibility – specifically, the responsibility of server security. An unsecured server is like leaving your front door wide open, inviting potential threats that can lead to data breaches, service disruptions, and significant financial or reputational damage.&lt;/p>
&lt;p>For developers, DevOps engineers, and Site Reliability Engineers (SREs), understanding and implementing robust server protection is not just an IT task; it&amp;rsquo;s fundamental to building reliable, trustworthy, and resilient systems. Whether you&amp;rsquo;re managing a single web server or a complex distributed infrastructure, knowing how to secure a server is paramount. This guide will walk you through what server security entails, why it&amp;rsquo;s critically important, and the essential server security best practices you need to implement.&lt;/p></description></item><item><title>What Is Web Server Capacity Planning &amp; How Does It Work?</title><link>https://www.netdata.cloud/academy/web-server-capacity-planning/</link><pubDate>Thu, 24 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/web-server-capacity-planning/</guid><description>&lt;p>Imagine launching a major marketing campaign or experiencing peak holiday traffic, only to have your website slow to a crawl or crash entirely. Users encounter frustrating error messages, abandon their carts, and your business suffers. This scenario often happens when web &lt;strong>servers are at capacity&lt;/strong>, unable to handle the incoming load. The solution? Proactive &lt;strong>web server capacity planning&lt;/strong>.&lt;/p>
&lt;p>For DevOps engineers, SREs, and system administrators, capacity planning isn&amp;rsquo;t just a &amp;ldquo;nice-to-have&amp;rdquo;; it&amp;rsquo;s a fundamental practice for ensuring the reliability, performance, and availability of web services. It&amp;rsquo;s about understanding your current &lt;strong>server capacity&lt;/strong>, anticipating future needs, and ensuring you have the right resources in place &lt;em>before&lt;/em> demand overwhelms your infrastructure. This guide explores what web server capacity planning involves, why it&amp;rsquo;s critical, and the process for implementing it effectively.&lt;/p></description></item><item><title>Container vs VM - Which Is Better Option For You</title><link>https://www.netdata.cloud/academy/container-vs-vm-which-is-better-for-you/</link><pubDate>Tue, 22 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/container-vs-vm-which-is-better-for-you/</guid><description>&lt;p>Deploying applications efficiently and reliably often requires isolating them from the underlying infrastructure and from each other. Virtualization makes this possible by creating virtual representations of computing resources. Two leading technologies dominate this space: containers and virtual machines (VMs). While both offer isolation and deployment benefits, they operate fundamentally differently. Choosing between a &lt;strong>container vs VM&lt;/strong> approach is a critical decision that impacts performance, resource usage, security, and deployment speed.&lt;/p></description></item><item><title>Infrastructure Monitoring vs Application Monitoring</title><link>https://www.netdata.cloud/academy/infrastructure-monitoring-vs-application-monitoring/</link><pubDate>Tue, 22 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/infrastructure-monitoring-vs-application-monitoring/</guid><description>&lt;p>Navigating the divide between infrastructure monitoring vs application monitoring requires a strategic approach. Infrastructure monitoring assesses system health, such as network bandwidth adequacy. &lt;em>Is your network performing optimally?&lt;/em> On the other hand, application monitoring focuses on software performance issues. &lt;em>What causes slow response times?&lt;/em>&lt;/p>
&lt;p>&lt;em>How can your organization merge these monitoring capabilities to proactively identify and address potential issues?&lt;/em>&lt;/p>
&lt;p>Let’s find out.&lt;/p>
&lt;h2 id="what-is-infrastructure-performance-monitoring-ipm">What Is Infrastructure Performance Monitoring (IPM)?&lt;/h2>
&lt;p>&lt;strong>Infrastructure Performance Monitoring&lt;/strong> (IPM) is a critical facet of IT operations. It focuses on the continuous oversight of essential system components such as hardware, cloud networks, and overall system performance. &lt;a href="https://www.netdata.cloud/academy/what-is-infrastructure-monitoring-and-why-you-need-it/">IPM involves real-time data analysis&lt;/a> to ensure system health and efficiency, with a focus on identifying and pinpointing issues that can affect system performance and stability.&lt;/p></description></item><item><title>What Is Real-Time Monitoring? 5 Benefits &amp; How It Works</title><link>https://www.netdata.cloud/academy/real-time-monitoring-benefits/</link><pubDate>Mon, 21 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/real-time-monitoring-benefits/</guid><description>&lt;p>&lt;strong>Real-time monitoring&lt;/strong> is an essential tool for any business aiming to safeguard its digital operations and boost network performance. If your goal is to maintain smooth operations and robust security, you shouldn’t overlook &lt;strong>real-time data monitoring&lt;/strong>, as it is crucial for detecting and addressing issues instantly, keeping your systems efficient and protected.&lt;/p>
&lt;p>In this article, we will explore how real-time monitoring works and its key benefits. Understanding these elements can significantly enhance your approach to network management and operational efficiency.&lt;/p></description></item><item><title>What Is Cloud Workload Protection</title><link>https://www.netdata.cloud/academy/cloud-workload-protection-and-challenges/</link><pubDate>Sun, 20 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/cloud-workload-protection-and-challenges/</guid><description>&lt;p>The shift to cloud computing offers incredible advantages in scalability, agility, and innovation. Businesses leverage cloud platforms like AWS, &lt;a href="https://www.netdata.cloud/solutions/technologies/azure-monitoring/">Azure&lt;/a>, and &lt;a href="https://www.netdata.cloud/solutions/technologies/gcp-monitoring/">GCP&lt;/a> to build and deploy applications faster than ever before. However, this migration introduces new complexities, particularly around securing the actual applications and processes running in these dynamic environments – the cloud workloads.&lt;/p>
&lt;p>Traditional security approaches focused on protecting the network perimeter are no longer sufficient. Cloud workload protection (CWP) has emerged as a critical security strategy designed specifically for the unique nature of cloud and hybrid environments. It focuses on securing the workload itself, regardless of where it runs. Understanding what CWP is, why it&amp;rsquo;s essential, and the role of a Cloud Workload Protection Platform (CWPP) is vital for anyone responsible for cloud workload security. This guide explores the core concepts, benefits, and challenges of CWP.&lt;/p></description></item><item><title>What Is Network Congestion &amp; How To Fix It</title><link>https://www.netdata.cloud/academy/what-is-network-congestion/</link><pubDate>Sun, 20 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-network-congestion/</guid><description>&lt;p>&lt;em>Site congestion found.&lt;/em> Let’s fix this before it gets worse! &lt;strong>Network congestion&lt;/strong> can cripple your digital operations, slowing down processes and frustrating users. What causes congestion in your network? How does it affect your system, and what can you do to resolve – or prevent – these traffic jams? Keep your network up and running. Here is how!&lt;/p>
&lt;h2 id="what-is-network-congestion">What Is Network Congestion?&lt;/h2>
&lt;p>&lt;strong>Network congestion&lt;/strong> is like a traffic jam on your data highway; it occurs when there&amp;rsquo;s more data trying to travel across a network than the available bandwidth can handle. This overload leads to annoying performance issues such as increased latency, jitter, packet loss, and reduced &lt;a href="https://aws.amazon.com/compare/the-difference-between-throughput-and-latency/" target="_blank">throughput&lt;/a>.&lt;/p></description></item><item><title>What Are Windows Event Logs? The Ultimate Guide</title><link>https://www.netdata.cloud/academy/what-are-windows-event-logs/</link><pubDate>Fri, 18 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-are-windows-event-logs/</guid><description>&lt;p>What could the dream of every IT professional be? Systems running flawlessly, no question about that. In reality, though, crashes, errors, and performance issues are inevitable. When problems arise, &lt;strong>Windows event logs&lt;/strong> provide you with a detailed record of what happened, helping you diagnose and resolve issues efficiently.&lt;/p>
&lt;p>&lt;a href="https://gs.statcounter.com/os-market-share/desktop/worldwide" target="_blank">Windows powers about 71.9% of desktops worldwide&lt;/a>, making it the most widely used operating system. Moreover, Windows servers generate a massive amount of event logs. Some can log up to &lt;a href="https://community.splunk.com/t5/Installation/Windows-Event-log-volume-is-extremely-high-on-one-server-where/m-p/135541" target="_blank">500 events per second&lt;/a>. With this kind of data flow, &lt;strong>managing and analyzing logs efficiently&lt;/strong> is crucial to keeping your systems running smoothly and securely.&lt;/p></description></item><item><title>All Types Of Databases: Advantages &amp; Examples</title><link>https://www.netdata.cloud/academy/types-of-databases/</link><pubDate>Thu, 17 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/types-of-databases/</guid><description>&lt;p>This is a tech-savvy exploration of &lt;strong>database types&lt;/strong>! From the structured precision of relational models to the dynamic versatility of NoSQL, we’re diving deep into each &lt;strong>DB type&lt;/strong>, unpacking their perks, and dishing out some examples to help you pick the perfect fit for your digital toolbox.&lt;/p>
&lt;h2 id="what-is-a-database">What Is A Database?&lt;/h2>
&lt;p>&lt;strong>A database is a structured collection of stored data&lt;/strong>, that allows easy access, management, and updating. They come in various types to suit different needs; from hierarchical databases that organize data in a tree-like format to relational databases that use tables and relationships among those data. Databases are essential to build any tech product from small mobile apps to large-scale enterprise systems. They help you handle, search, and efficiently utilize large data sets.&lt;/p></description></item><item><title>FreeBSD vs Linux - Which Is Better?</title><link>https://www.netdata.cloud/academy/freebsd-vs-linux/</link><pubDate>Tue, 15 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/freebsd-vs-linux/</guid><description>&lt;p>When choosing an operating system for servers, embedded systems, or even desktops, &lt;a href="https://www.netdata.cloud/solutions/built-for/developers/">developers&lt;/a> and system administrators often encounter two powerful, free, and open-source Unix-like options: FreeBSD and &lt;a href="https://www.netdata.cloud/solutions/technologies/linux-monitoring/">Linux&lt;/a>. Both share a common heritage tracing back to the original UNIX, but they have evolved along different paths, resulting in distinct philosophies, architectures, and strengths. The &lt;strong>FreeBSD vs Linux&lt;/strong> debate isn&amp;rsquo;t about one being definitively &amp;ldquo;better&amp;rdquo; but rather understanding which is better suited for &lt;em>your&lt;/em> specific needs.&lt;/p></description></item><item><title>Real-Time Data Visualization: Examples &amp; Use Cases</title><link>https://www.netdata.cloud/academy/real-time-data-visualization/</link><pubDate>Tue, 15 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/real-time-data-visualization/</guid><description>&lt;p>Data is increasingly generated at an unprecedented velocity. From application logs and system metrics to user interactions and IoT sensor readings, the sheer volume can be overwhelming. Simply collecting this data isn&amp;rsquo;t enough; the real value lies in understanding it &lt;em>as it happens&lt;/em>. This is where &lt;strong>real-time data visualization&lt;/strong> comes into play. It transforms relentless streams of raw data into clear, intuitive visuals, allowing you to grasp trends, spot anomalies, and make informed decisions instantly.&lt;/p></description></item><item><title>What Is Network Security Monitoring - A Comprehensive Guide</title><link>https://www.netdata.cloud/academy/network-security-monitoring/</link><pubDate>Mon, 14 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/network-security-monitoring/</guid><description>&lt;p>Imagine attackers silently infiltrating your network, hiding like the Greeks inside the Trojan Horse. They could remain undetected for weeks or even months, mapping your systems, &lt;a href="https://www.netdata.cloud/academy/how-to-secure-sensitive-data-in-cloud/">stealing sensitive data&lt;/a>, and preparing for a larger attack. By the time you realize they&amp;rsquo;re inside, significant damage might already be done. This scenario highlights a critical challenge in modern cybersecurity: perimeter defenses like firewalls are essential, but they aren&amp;rsquo;t foolproof. You need visibility &lt;em>inside&lt;/em> your network to detect threats that slip through.&lt;/p></description></item><item><title>Deployment Automation: Tools, Benefits &amp; Practices</title><link>https://www.netdata.cloud/academy/deployment-automation/</link><pubDate>Wed, 02 Apr 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/deployment-automation/</guid><description>&lt;p>Getting new features and bug fixes from a developer’s machine into the hands of users quickly and reliably is paramount when it comes to successful software development. Manual deployment processes, however, are often slow, error-prone, and stressful.&lt;/p>
&lt;p>Manual tasks in application deployments frequently lead to configuration errors and make the software deployment process a time consuming process, especially for complex deployments. This is where &lt;strong>deployment automation&lt;/strong> comes in – a crucial practice in &lt;a href="https://www.netdata.cloud/solutions/built-for/devops/">modern DevOps&lt;/a> and agile methodologies.&lt;/p></description></item><item><title>Monitoring Netdata Restarts: A Reliable Solution</title><link>https://www.netdata.cloud/blog/2025-03-06-monitoring-netdata-restarts/</link><pubDate>Thu, 06 Mar 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/2025-03-06-monitoring-netdata-restarts/</guid><description>&lt;p>For a tool like Netdata, monitoring crashes and abnormal events extends far beyond bug fixing—it&amp;rsquo;s essential for identifying edge cases, preventing regressions, and delivering the most dependable observability experience possible. With millions of daily downloads, each event provides a vital signal for maintaining the integrity of our systems.&lt;/p>
&lt;h2 id="the-challenge-with-traditional-solutions">The Challenge with Traditional Solutions&lt;/h2>
&lt;p>Over the years, we&amp;rsquo;ve evaluated many monitoring tools, each with significant limitations:&lt;/p>
&lt;table>
 &lt;thead>
 &lt;tr>
 &lt;th>Tool&lt;/th>
 &lt;th>Strengths&lt;/th>
 &lt;th>Limitations&lt;/th>
 &lt;/tr>
 &lt;/thead>
 &lt;tbody>
 &lt;tr>
 &lt;td>&lt;strong>Sentry&lt;/strong>&lt;/td>
 &lt;td>• Comprehensive error tracking features&lt;br/>• Detailed stack traces&lt;/td>
 &lt;td>• Per-event pricing model becomes prohibitive at scale&lt;br/>• Forces sampling which reduces visibility into critical issues&lt;br/>• Compromises complete error capture for cost control&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>&lt;strong>&lt;a href="https://www.netdata.cloud/solutions/technologies/gcp-monitoring/">GCP&lt;/a> BigQuery &amp;amp; Similar&lt;/strong>&lt;/td>
 &lt;td>• Powerful query capabilities&lt;br/>• Flexible data processing&lt;br/>• High scalability potential&lt;/td>
 &lt;td>• Complex reporting setup and maintenance&lt;br/>• Significant costs at high event volumes&lt;br/>• Requires specialized technical expertise&lt;/td>
 &lt;/tr>
 &lt;tr>
 &lt;td>&lt;strong>Other Solutions&lt;/strong>&lt;/td>
 &lt;td>• Various specialized features&lt;br/>• Some open-source flexibility&lt;/td>
 &lt;td>• Either too inflexible for custom requirements&lt;br/>• Or prohibitively expensive at full-capture scale&lt;br/>• Often require compromising between detail and cost&lt;/td>
 &lt;/tr>
 &lt;/tbody>
&lt;/table>
&lt;p>We consistently encountered these core challenges:&lt;/p></description></item><item><title>What Is Database Clustering? Types &amp; Benefits</title><link>https://www.netdata.cloud/academy/whatisdatabaseclusteringtypesbenefits/</link><pubDate>Thu, 06 Mar 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/whatisdatabaseclusteringtypesbenefits/</guid><description>&lt;p>&lt;strong>Database clustering&lt;/strong> is a robust strategy employed to enhance the &lt;strong>performance&lt;/strong>, &lt;strong>scalability&lt;/strong>, and &lt;strong>availability&lt;/strong> of databases by orchestrating the distribution of data across multiple servers. This setup doesn&amp;rsquo;t just streamline handling big data; it also keeps things running smoothly, even if some nodes go down. That way, your service stays up and available no matter what happens.&lt;/p>
&lt;p>By diving into the various &lt;strong>types of database clustering architectures&lt;/strong>, in this article we explain how each setup addresses specific needs and challenges within IT environments. We&amp;rsquo;ll also examine the tangible benefits that database clustering brings to businesses, from improved data redundancy to enhanced query response times.&lt;/p></description></item><item><title>Docker Monitoring Tool With Unlimited Containers</title><link>https://www.netdata.cloud/solutions/technologies/docker-monitoring/</link><pubDate>Mon, 27 Jan 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/technologies/docker-monitoring/</guid><description>Netdata brings the simplicity Docker promised to monitoring. One command install, instant visibility into every container&amp;rsquo;s CPU, memory, disk, and network. No agents per container, no complex pipelines, no surprise bills.</description></item><item><title>Infrastructure Monitoring For Network Engineers</title><link>https://www.netdata.cloud/solutions/built-for/network-engineers/</link><pubDate>Mon, 27 Jan 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/built-for/network-engineers/</guid><description>Netdata gives network engineers the full picture: live topology maps, NetFlow/sFlow/IPFIX flow analysis, SNMP monitoring with 100+ device profiles, a native SNMP trap receiver, per-second interface metrics, and ML that learns your traffic patterns. Deploy in minutes and stop hearing about outages from users.</description></item><item><title>Netdata vs Prometheus: A 2025 Performance Analysis</title><link>https://www.netdata.cloud/blog/netdata-vs-prometheus-2025/</link><pubDate>Thu, 23 Jan 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-vs-prometheus-2025/</guid><description>&lt;p>When it comes to infrastructure monitoring, performance, scalability, and efficiency are critical considerations. In this blog post, we revisit two widely adopted open-source monitoring solutions: &lt;strong>Netdata&lt;/strong> and &lt;strong>Prometheus&lt;/strong>. Both tools have introduced notable improvements in their latest versions, emphasizing scalability and enhanced efficiency.&lt;/p>
&lt;p>In our previous &lt;a href="https://www.netdata.cloud/blog/netdata-vs-prometheus-performance-analysis/">analysis&lt;/a>, we explored key differences between these systems, focusing on resource consumption and data retention. This follow-up expands on that comparison by subjecting both tools to a significantly larger workload. With the number of monitored nodes increased to 1000, containers to 80k, and metrics ingestion reaching 4.6 million metrics per second, we examine how each system performs under these demanding conditions, focusing on CPU utilization, memory requirements, disk I/O, network usage, and data retention during data ingestion.&lt;/p></description></item><item><title>Long-Term Data Storage and Retention in Netdata</title><link>https://www.netdata.cloud/blog/long-term-data-retention/</link><pubDate>Tue, 21 Jan 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/long-term-data-retention/</guid><description>&lt;p>Netdata&amp;rsquo;s database engine (dbengine) provides a sophisticated multi-tiered storage system designed for efficient long-term data retention while maintaining high granularity. This article explores the technical details of how Netdata handles metric storage, the advantages of its distributed architecture, and how to configure it for your specific needs.&lt;/p>
&lt;hr>
&lt;h2 id="database-engine-architecture">Database Engine Architecture&lt;/h2>
&lt;p>Netdata&amp;rsquo;s database engine provides a sophisticated, efficient solution for long-term metric storage through:&lt;/p>
&lt;ul>
&lt;li>Intelligent multi-tiered storage architecture&lt;/li>
&lt;li>Efficient compression and caching mechanisms&lt;/li>
&lt;li>Flexible retention strategies&lt;/li>
&lt;li>Distributed deployment options&lt;/li>
&lt;/ul>
&lt;p>This design enables organizations to maintain detailed historical data while optimizing storage use and maintaining query performance.&lt;/p></description></item><item><title>Network Monitoring Software With Real-Time Visibility</title><link>https://www.netdata.cloud/solutions/use-cases/network-monitoring/</link><pubDate>Wed, 15 Jan 2025 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/solutions/use-cases/network-monitoring/</guid><description>Real-time network monitoring with per-second visibility, ML-powered anomaly detection, and zero-configuration deployment.</description></item><item><title>Best Of Category Badges Earned In 2024: G2 &amp; Capterra</title><link>https://www.netdata.cloud/blog/netdata-best-of-category-badges-in-2024/</link><pubDate>Mon, 23 Dec 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-best-of-category-badges-in-2024/</guid><description>&lt;h2 id="netdata-featured-with-multiple-best-of-category-badges-in-2024">Netdata Featured with Multiple “Best Of” Category Badges in 2024&lt;/h2>
&lt;p>As we are close to the end of this year, we are thrilled to announce that &lt;a href="https://www.capterra.com/p/251845/Netdata/?utm_source=vp&amp;utm_medium=blog&amp;utm_campaign=ts-q4-2024" target="_blank">&lt;strong>Netdata&lt;/strong>&lt;/a> has been recognized with multiple “Best of” badges from Gartner Digital Markets brands: &lt;a href="https://www.capterra.com/?utm_source=vp&amp;utm_medium=blog&amp;utm_campaign=ts-q4-2024" target="_blank">&lt;strong>Capterra&lt;/strong>&lt;/a>, &lt;a href="https://www.softwareadvice.com/?utm_source=vp&amp;utm_medium=blog&amp;utm_campaign=ts-q4-2024" target="_blank">&lt;strong>Software Advice&lt;/strong>&lt;/a>, and &lt;a href="https://www.getapp.com/?utm_source=vp&amp;utm_medium=blog&amp;utm_campaign=ts-q4-2024" target="_blank">&lt;strong>GetApp&lt;/strong>&lt;/a>, leading software recommendation search engines.&lt;/p>
&lt;p>This “Best of” badges program is an independent assessment that evaluates user reviews to help buyers identify the highest-rated software companies in specific categories that offer the most popular solutions.&lt;/p></description></item><item><title>Getting Started With Netdata: Real-Time Monitoring</title><link>https://www.netdata.cloud/blog/getting-started-with-netdata-a-comprehensive-guide-to-real-time-monitoring/</link><pubDate>Thu, 19 Dec 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/getting-started-with-netdata-a-comprehensive-guide-to-real-time-monitoring/</guid><description>&lt;p>Now you can start monitoring thousands of metrics in real-time, detecting anomalies throughout your infra, and troubleshooting issues even mid-crisis!
Watch this &lt;a href="https://www.youtube.com/watch?v=z5m8JdwMOn8&amp;themeRefresh=1" target="_blank">1-minute video&lt;/a> for a quick intro! Get ready to be blown away!&lt;/p>
&lt;p>To fully utilize Netdata for monitoring your infrastructure, do these:&lt;/p>
&lt;h2 id="a-hrefhttpslearnnetdataclouddocsdeployment-guides-target_blank-deploy-netdataa">&lt;a href="https://learn.netdata.cloud/docs/deployment-guides/" target="_blank">① Deploy Netdata&lt;/a>&lt;/h2>
&lt;p>&lt;em>Install Netdata to your systems.&lt;/em>&lt;/p>
&lt;p>Netdata runs on &lt;a href="https://www.netdata.cloud/solutions/technologies/linux-monitoring/">Linux&lt;/a>, &lt;a href="https://www.netdata.cloud/solutions/technologies/windows-monitoring/">Windows&lt;/a>, FreeBSD and MacOS and can be installed on &lt;a href="https://www.netdata.cloud/academy/bare-metal-server/">bare-metal servers&lt;/a>, cloud VMs, Kubernetes, even weak &lt;a href="https://www.netdata.cloud/solutions/iot-monitoring/">IoT devices&lt;/a>.
We have carefully optimized it to be the fastest and most advanced monitoring you will ever need, and at the same time be extremely friendly, polite and respectful to your production systems and applications.&lt;/p></description></item><item><title>Real-Time Windows Server Monitoring Best Practices</title><link>https://www.netdata.cloud/academy/real-time-windows-server-monitoring-from-insights-to-action/</link><pubDate>Tue, 17 Dec 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/real-time-windows-server-monitoring-from-insights-to-action/</guid><description>&lt;iframe width="100%" height="450" src="https://www.youtube.com/embed/xc8PaCLdZDA?si=k5wcZHSP2Gyusdg4" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen>&lt;/iframe>
&lt;p>If you are looking to monitor your Windows Server Machines, this webinar will guide you through the latest strategies for real-time observability, system and infrastructure optimization. Featuring a hands-on live demo and insights from industry experts, this session will be full of actionable techniques to help you gain deeper visibility into your Windows infrastructure, troubleshoot faster, and improve performance with ease.&lt;/p></description></item><item><title>Customize Your Netdata Experience with Favorites</title><link>https://www.netdata.cloud/blog/customize-your-netdata-experience-with-favorites/</link><pubDate>Mon, 09 Dec 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/customize-your-netdata-experience-with-favorites/</guid><description>&lt;h3 id="make-netdata-yours-with-favorites">Make Netdata yours with Favorites&lt;/h3>
&lt;p>Monitoring is personal. Different systems, workloads, and use-cases mean different charts matter to different users. Netdata now makes it easier than ever to customize your experience by letting you &lt;strong>favorite charts&lt;/strong> you care about most.&lt;/p>
&lt;h3 id="how-it-works">How it works&lt;/h3>
&lt;p>You&amp;rsquo;ll now see a &lt;strong>heart icon&lt;/strong> next to any chart or section of charts. Click it, and it instantly gets added to your &lt;strong>favorites&lt;/strong>. favorites appear right above your system overview metrics, so you see what matters most—first.&lt;/p></description></item><item><title>20 Best DevOps, SRE &amp; Observability Conferences 2025</title><link>https://www.netdata.cloud/blog/20-devops-sre-observability-events-and-conferences-you-should-consider-in-2025/</link><pubDate>Wed, 04 Dec 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/20-devops-sre-observability-events-and-conferences-you-should-consider-in-2025/</guid><description>&lt;p>Every year, DevOps, SRE Sysadmin &amp;amp; IT leaders gather at conferences across the world to share not only knowledge but also the latest trends, tools, and strategies for achieving greater insight and control over complex systems. This guide brings you a comprehensive list of DevOps, SRE &amp;amp; observability events that can help you stay ahead in a rapidly evolving field. Whether you&amp;rsquo;re exploring the latest in tools, cloud observability, or AI-driven insights, this guide will lead you to the perfect event to expand your knowledge and network with industry leaders!&lt;/p></description></item><item><title>SolarWinds Migration Program | Netdata</title><link>https://www.netdata.cloud/migrate/solarwinds/</link><pubDate>Tue, 12 Nov 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/migrate/solarwinds/</guid><description/></item><item><title>KAUST Case Study: System Stability For Academia</title><link>https://www.netdata.cloud/case-studies/education/kaust/</link><pubDate>Mon, 11 Nov 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/education/kaust/</guid><description>&lt;h2 id="empowering-system-stability-in-a-complex-research-environment">Empowering System Stability in a Complex Research Environment&lt;/h2>
&lt;p>The Systems Research Group at King Abdullah University of Science and Technology (KAUST) manages a diverse and expansive computing environment essential for advancing scientific research. Given the critical nature of this environment, ensuring uptime and rapid troubleshooting is essential. However, without a robust monitoring solution, pinpointing performance issues or hardware failures was a slow, often reactive process that challenged the team’s ability to maintain a smooth operational flow.&lt;/p></description></item><item><title>Native Windows Agent: Real-Time Windows Monitoring</title><link>https://www.netdata.cloud/blog/netdata-native-windows-agent/</link><pubDate>Fri, 08 Nov 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-native-windows-agent/</guid><description>&lt;p>We are pleased to announce a significant advancement in system monitoring: the launch of &lt;a href="https://www.netdata.cloud/solutions/windows-monitoring/">Netdata&amp;rsquo;s first-ever Native Windows Agent&lt;/a>. This release represents a major step forward in our mission to provide comprehensive and efficient monitoring solutions across all platforms. With the introduction of the native Windows agent, we are extending our robust monitoring capabilities to Windows environments, enabling seamless and unified monitoring across diverse infrastructures. This development is a direct response to the needs of our user community, who have expressed a strong demand for a powerful and intuitive Windows monitoring solution.&lt;/p></description></item><item><title>Accurate Process Monitoring with Netdata</title><link>https://www.netdata.cloud/blog/accurate-process-monitoring-with-netdata/</link><pubDate>Mon, 04 Nov 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/accurate-process-monitoring-with-netdata/</guid><description>&lt;p>Understand why tracking cumulative resource consumption is crucial for accurate process monitoring.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="accurate-process-monitoring-with-netdata-why-tracking-cumulative-resource-consumption-matters">Accurate Process Monitoring with Netdata: Why Tracking Cumulative Resource Consumption Matters&lt;/h2>
&lt;p>Tracking the cumulative resource consumption of processes, including short-lived and exited children, is a rare feature in monitoring tools – and it’s one of the standout capabilities Netdata offers.&lt;/p>
&lt;p>Most mainstream solutions, like Datadog’s process monitoring and Prometheus’s Node Exporter, focus on active processes and only collect metrics per PID. Even specialized process monitors (&lt;code>top&lt;/code>, &lt;code>htop&lt;/code>, etc.) face the same limitations. They capture snapshots of currently running processes and often rely on fixed sampling intervals, which are too slow to catch very short-lived tasks. This approach falls short when trying to capture the resource footprint of dynamic applications, shell scripts, and complex process hierarchies.&lt;/p></description></item><item><title>Linux Load Average Myths and Realities</title><link>https://www.netdata.cloud/blog/linux-load-average-myths-and-realities/</link><pubDate>Sun, 03 Nov 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/linux-load-average-myths-and-realities/</guid><description>&lt;p>When it comes to monitoring system &lt;a href="https://www.netdata.cloud/solutions/linux-monitoring/">performance on Linux&lt;/a>, the load average is one of the most referenced metrics. Displayed prominently in tools like &lt;code>top&lt;/code>, &lt;code>uptime&lt;/code>, and &lt;code>htop&lt;/code>, it&amp;rsquo;s often used as a quick gauge of system load and capacity. But how reliable is it? For complex, multi-threaded applications, load average can paint a misleading picture of actual system performance.&lt;/p>
&lt;p>In this article, we&amp;rsquo;ll dive into the myths and realities of Linux load average, using insights from Netdata’s high-frequency, high-concurrency monitoring setup. Through this journey, we&amp;rsquo;ll uncover why load average spikes can occur even under steady workloads, and why a single metric is rarely enough to capture the true state of a system. Whether you&amp;rsquo;re a system administrator, developer, or performance enthusiast, this exploration of load average will help you interpret it more accurately and understand when it may—or may not—reflect reality.&lt;/p></description></item><item><title>Fix NGINX 500 Internal Server Error: Easy Steps</title><link>https://www.netdata.cloud/academy/how-to-fix-nginx-500-internal-server-error-simple-steps-for-troubleshooting/</link><pubDate>Thu, 31 Oct 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/how-to-fix-nginx-500-internal-server-error-simple-steps-for-troubleshooting/</guid><description>&lt;p>When managing a website or web application through NGINX, it&amp;rsquo;s quite likely you&amp;rsquo;ve stumbled upon the feared 500 Internal Server Error at least once. This error often causes a moment of panic as it typically indicates a problem lurking within the server&amp;rsquo;s operations. Regrettably, one finds that the message falls short in offering sufficient context to aid in uncovering the underlying reason.&lt;/p>
&lt;p>No need to fret! This guide is here to take you step by step through the usual culprits behind a 500 error when using NGINX and, even better, how to solve it. If you&amp;rsquo;re facing a problem like not having the right permissions, a setup mistake, or something different, don&amp;rsquo;t worry—we have all the steps you need to fix it.&lt;/p></description></item><item><title>6 + 1 Effective Strategies to Reduce Unplanned Downtime</title><link>https://www.netdata.cloud/academy/6-+-1-effective-strategies-to-reduce-unplanned-downtime/</link><pubDate>Wed, 30 Oct 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/6-+-1-effective-strategies-to-reduce-unplanned-downtime/</guid><description>&lt;p>For DevOps and SRE teams, unplanned downtime may be a nightmare since it can ruin everything, from customer satisfaction to corporate operations. With more and more systems relying on constant availability, any unplanned downtime must be eliminated to ensure reliability of service. In the article below, we will discuss how to utilize &lt;a href="https://www.netdata.cloud/academy/what-is-infrastructure-monitoring-and-why-you-need-it/">infrastructure monitoring&lt;/a>, monitoring tools, development and operations best practices in order to &lt;a href="https://www.netdata.cloud/academy/what-is-uptime-monitoring/">minimize unplanned downtime&lt;/a>.&lt;/p>
&lt;h2 id="what-is-unplanned-downtime">What is Unplanned Downtime&lt;/h2>
&lt;p>Unplanned downtime occurs when an application or system or infrastructure component fails without notice, causing an interruption. These disruptions could be due to a network issue, or a problem in the software or hardware or even human error. The business expenses are often substantial in terms of revenue loss and negative publicity. So, putting a plan in place to reduce downtime is critical.&lt;/p></description></item><item><title>12 Benefits You Get by Scaling with Netdata</title><link>https://www.netdata.cloud/blog/12-benefits-you-get-by-scaling-with-netdata/</link><pubDate>Wed, 23 Oct 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/12-benefits-you-get-by-scaling-with-netdata/</guid><description>&lt;p>&lt;a href="https://blogs.idc.com/2022/12/09/idc-futurescape-worldwide-future-of-digital-infrastructure-2023-predictions/" target="_blank">80% of decision-makers globally&lt;/a> acknowledge that digital infrastructure is essential for reaching business goals. However, IT infrastructure is becoming increasingly distributed and complex. Organizations are managing hundreds—even thousands—of nodes across cloud, on-premise, and edge environments.
This predicament makes effective monitoring across all systems more essential than ever. In turn, this drives the demand for valuable real-time insights through scalable &lt;a href="https://www.netdata.cloud/blog/understanding-monitoring-tools/" target="_blank">monitoring solutions&lt;/a> like Netdata​.&lt;/p>
&lt;p>If you’re managing a large infrastructure but haven’t fully embraced Netdata, now is the time to reconsider. Let&amp;rsquo;s take a look at the benefits you’ll get, below.&lt;/p></description></item><item><title>5 Best Datadog Alternatives For Monitoring &amp; Observability</title><link>https://www.netdata.cloud/blog/5-datadog-alternatives/</link><pubDate>Tue, 15 Oct 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/5-datadog-alternatives/</guid><description>&lt;p>As businesses rely more on &lt;a href="https://www.netdata.cloud/academy/what-is-infrastructure-monitoring-and-why-you-need-it/">infrastructure monitoring&lt;/a>, observability tools have become essential for keeping systems secure, responsive, and efficient. In this evolving monitoring landscape, Datadog is known as a leading analytics and monitoring tool, capable of collecting essential performance metrics from servers, databases, applications, and other IT infrastructures. But at what cost?&lt;/p>
&lt;h2 id="what-is-datadog-is-it-your-only-option">What Is Datadog? Is It Your Only Option?&lt;/h2>
&lt;p>Datadog is a cloud-based monitoring platform that helps organizations track the performance of their infrastructure, applications, and logs. It provides visibility across systems to ensure efficient operations and identify issues quickly.&lt;/p></description></item><item><title>ilert Integration: Streamline Monitoring &amp; Response</title><link>https://www.netdata.cloud/blog/netdata-integration-with-ilert-streamlining-monitoring-and-incident-response/</link><pubDate>Mon, 14 Oct 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-integration-with-ilert-streamlining-monitoring-and-incident-response/</guid><description>&lt;p>Netdata now integrates with ilert, a leading incident response platform. With this integration, the incident management features and alerting capabilities of ilert and the real-time systems monitoring provided by Netdata can be leveraged. By combining both systems, users can not only monitor their infrastructure with fine detail as never before, but also assure the responsiveness of critical alerts to the correct teams swiftly.&lt;/p>
&lt;h2 id="what-is-ilert">What is ilert?&lt;/h2>
&lt;p>&lt;a href="https://www.ilert.com/?utm_campaign=Netdata&amp;utm_source=integration&amp;utm_medium=organic" target="_blank">ilert&lt;/a> is an end-to-end platform for alerting, on-call management, and status pages, built for the new-age DevOps and SRE teams. It streamlines the entire incident management process by automating key aspects and providing powerful tools to improve efficiency, such as automated on-call duty, multi-level escalations, alert grouping, postmortem document creation, and much more. ilert integrates with monitoring systems, like Netdata, and sends alerts through multiple channels (SMS, phone, push, Slack, Microsoft Teams, etc.), ensuring that teams can promptly respond to issues before they impact service.&lt;/p></description></item><item><title>What Is Application Performance Monitoring (APM)?</title><link>https://www.netdata.cloud/academy/what-is-application-performance-monitoring-apm/</link><pubDate>Thu, 03 Oct 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-application-performance-monitoring-apm/</guid><description>&lt;p>Ensuring that applications are functioning as expected is essential in the software-driven world of today. One sub-standard performance and your app can fail to win people over, ultimately drive them away or be the reason for debilitating first impressions. This is where Application Performance Monitoring (APM) comes in and helps developers, DevOps teams, and especially an SRE (Site Reliability Engineer) to monitor their application running live on systems, allowing them to identify the issues faster before they impact users.&lt;/p></description></item><item><title>ChileAtiende Case Study: 25% Less AWS Downtime</title><link>https://www.netdata.cloud/case-studies/government/chileatiende/</link><pubDate>Tue, 24 Sep 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/government/chileatiende/</guid><description>&lt;h2 id="proactive-infrastructure-management-for-citizen-services">Proactive Infrastructure Management for Citizen Services&lt;/h2>
&lt;p>ChileAtiende, a government organization dedicated to delivering essential digital services, faced a growing challenge: ensuring the stability and reliability of its complex infrastructure. With an ever-expanding digital presence, the team needed an efficient way to monitor a hybrid environment that included containers, databases, and web services like nginx and Apache.&lt;/p>
&lt;p>The critical pain point for ChileAtiende was ensuring system reliability through real-time monitoring and quick response capabilities. They sought a solution that would allow them to anticipate potential problems and address them before they impacted the public.&lt;/p></description></item><item><title>TMB Case Study: Barcelona Transport Monitoring</title><link>https://www.netdata.cloud/case-studies/transportation/tmb/</link><pubDate>Fri, 20 Sep 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/transportation/tmb/</guid><description>&lt;h2 id="monitoring-a-complex-heterogeneous-infrastructure">Monitoring a Complex, Heterogeneous Infrastructure&lt;/h2>
&lt;p>Transports Metropolitans de Barcelona (TMB), the public transportation provider in Barcelona, manages a vast and varied infrastructure hosted on AWS and on bare metal. Monitoring this infrastructure presents significant challenges due to its heterogeneous nature—spanning Linux, Windows, Kubernetes, cloud, and on-premises environments.&lt;/p>
&lt;p>One of the main obstacles for TMB was finding a single tool capable of handling the breadth of this environment effectively. As the infrastructure evolved, it became essential to have a holistic view of all systems to avoid operational silos and improve response times to issues.&lt;/p></description></item><item><title>What is Alert Fatigue and How to Prevent It</title><link>https://www.netdata.cloud/academy/what-is-alert-fatigue-and-how-to-prevent-it/</link><pubDate>Fri, 20 Sep 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-alert-fatigue-and-how-to-prevent-it/</guid><description>&lt;h2 id="what-is-alert-fatigue">What is Alert Fatigue?&lt;/h2>
&lt;p>This is called alert fatigue where engineers especially &lt;a href="https://www.netdata.cloud/academy/sre-vs-devops-what-are-the-main-differences-between-them/">DevOps SREs&lt;/a> and on-call teams become numb to the many notifications from their monitoring tools. Instead of quickly responding to important issues, engineers may start ignoring or missing crucial alerts due to the overwhelming number of notifications they receive. That is no small issue because that causes slower incident response times, possible service outages, and less reliability.&lt;/p>
&lt;p>Robust monitoring is a must-have in the high-speed world of cloud infrastructure and microservices. However, when every little spike sets off an alarm, it is hard to separate the wheat from the chaff, so to speak. So in the end, alerts lose their meaning and it becomes a society of ignoring alarms, even when they shouldn&amp;rsquo;t be ignored.&lt;/p></description></item><item><title>What Is Observability? Definition, Benefits &amp; How It Works</title><link>https://www.netdata.cloud/academy/what-is-observability/</link><pubDate>Fri, 20 Sep 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-observability/</guid><description>&lt;h2 id="what-is-observability">What Is Observability?&lt;/h2>
&lt;p>Observability is the process of trying to figure out how a system works on the inside, by looking at what it does on the outside. It has even grown to be a central theme in today&amp;rsquo;s technology stacks where it is used to guarantee system availability and speed, particularly in intricate distributed architectures. But it is not just about the data, it is about the knowledge derived from the data that would help keep the system healthy, diagnose problems, and optimize performance.&lt;/p></description></item><item><title>What Is Uptime Monitoring? All SREs &amp; DevOps Teams Must Know</title><link>https://www.netdata.cloud/academy/what-is-uptime-monitoring/</link><pubDate>Thu, 19 Sep 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-uptime-monitoring/</guid><description>&lt;h2 id="uptime-monitoring-explained">Uptime Monitoring Explained&lt;/h2>
&lt;p>Keeping an eye on how well and how often servers, apps, services, and all the parts of your system are up and running is what &lt;a href="https://www.netdata.cloud/blog/server-uptime-monitoring-why-do-we-need-it/">uptime monitoring&lt;/a> is all about. For people in &lt;a href="https://www.netdata.cloud/academy/sre-vs-devops-what-are-the-main-differences-between-them/">Site Reliability Engineering (SRE) and DevOps teams&lt;/a>, making sure everything works almost all the time is super important. Keeping your services up and running means users run into less trouble and enjoy a more seamless connection without outages. This cuts down on the chance of expensive interruptions in business.&lt;/p></description></item><item><title>Synthetic Checks: Definition, Examples &amp; Benefits</title><link>https://www.netdata.cloud/academy/synthetic-checks-definition-and-everything-else-you-need-to-know/</link><pubDate>Wed, 18 Sep 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/synthetic-checks-definition-and-everything-else-you-need-to-know/</guid><description>&lt;h2 id="what-are-synthetic-checks">What Are Synthetic Checks?&lt;/h2>
&lt;p>Synthetic checks are proactive tests that simulate user interactions or network requests to monitor the availability and performance of services. Rather than waiting for an issue to be reported by a user, &lt;a href="https://www.netdata.cloud/integrations/data-collection/synthetic-checks/">synthetic checks&lt;/a> continuously run automated checks to detect problems before they affect real users. This makes them an essential tool for DevOps and Site Reliability Engineers (SREs) who aim to maintain a reliable and responsive infrastructure.&lt;/p></description></item><item><title>Infrastructure Monitoring: Key Benefits, Types &amp; Use Cases</title><link>https://www.netdata.cloud/academy/what-is-infrastructure-monitoring-and-why-you-need-it/</link><pubDate>Tue, 17 Sep 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-infrastructure-monitoring-and-why-you-need-it/</guid><description>&lt;h2 id="what-is-infrastructure-monitoring">What Is Infrastructure Monitoring?&lt;/h2>
&lt;p>&lt;strong>Infrastructure monitoring&lt;/strong> is the practice of continuously tracking the performance, availability, and overall health of IT systems such as servers, databases, networks, cloud services and so many more.
To put it simply, it&amp;rsquo;s about having clear visibility into how all the components of your applications and your entire infrastructure in general, are behaving. This is why we need and use the so-called monitoring / observability tools.&lt;/p></description></item><item><title>Important Changes to the Netdata Agent Dashboard</title><link>https://www.netdata.cloud/blog/important-changes-to-the-netdata-agent-dashboard/</link><pubDate>Mon, 12 Aug 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/important-changes-to-the-netdata-agent-dashboard/</guid><description>&lt;p>&lt;em>Important Notice: These changes ONLY impact users of the Netdata Agent Dashboard not connected to Netdata Cloud.&lt;/em>&lt;/p>
&lt;p>Dear Netdata Community,&lt;/p>
&lt;p>We are writing to inform you of upcoming changes to the Netdata Agent Dashboard, which will take effect in the coming weeks. This change impacts users from the soon to be released Netdata v2.0 onwards (and also on the Netdata v1.47 nightly releases).
Currently, the Open-Source Netdata Agents allow unauthorized and unlimited Agent dashboard access. From Netdata v2.0 onwards, all Netdata Dashboards (Agent and Cloud) will offer exactly the same functionality under the same policy. Netdata Agent Dashboard will use Netdata Cloud as an SSO provider, ensuring dashboard access is authenticated and validated by Netdata Cloud, users will have the option to proceed with an unauthorized local dashboard but this will no longer be the default.&lt;/p></description></item><item><title>Analyze &amp; Reduce Disk I/O Bottlenecks: Top Techniques</title><link>https://www.netdata.cloud/academy/reduce-disk-io-bottlenecks/</link><pubDate>Mon, 01 Jul 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/reduce-disk-io-bottlenecks/</guid><description>&lt;h2 id="understanding-disk-io-bottlenecks">Understanding Disk I/O Bottlenecks&lt;/h2>
&lt;p>Disk I/O (Input/Output) bottlenecks can severely impact the performance of your applications and services. These bottlenecks occur when the disk subsystem cannot keep up with the read/write requests from the CPU or memory, leading to slow response times and degraded performance. As a DevOps or SRE professional, understanding how to analyze and reduce these bottlenecks is crucial for maintaining optimal system health.&lt;/p>
&lt;h2 id="identifying-disk-io-bottlenecks">Identifying Disk I/O Bottlenecks&lt;/h2>
&lt;p>Before you can tackle disk I/O bottlenecks, you need to identify them accurately. Effective identification involves using a combination of monitoring tools and understanding key metrics.&lt;/p></description></item><item><title>Motohunt Case Study: Automotive Industry Monitoring</title><link>https://www.netdata.cloud/case-studies/automotive/motohunt/</link><pubDate>Thu, 27 Jun 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/automotive/motohunt/</guid><description>&lt;h2 id="navigating-system-complexity-with-netdata">Navigating System Complexity with Netdata&lt;/h2>
&lt;p>Motohunt is a big player in the automotive technology sector, delivering innovative products to the moto lovers. The organization&amp;rsquo;s dedication to maintaining a healthy production environment and understanding data evolution across systems underscores its commitment to excellence.
Netdata emerged as a key player in addressing Motohunt&amp;rsquo;s primary challenge of monitoring system health.&lt;/p>
&lt;blockquote>
&lt;p>&amp;ldquo;Netdata&amp;rsquo;s out-of-the-box functionality and real-time analytics have been instrumental in maintaining our system&amp;rsquo;s health. Its value efficiency and comprehensive coverage &amp;lsquo;just works,&amp;rsquo; providing us with immediate insights without a hefty price tag.&amp;rdquo;&lt;/p></description></item><item><title>Webinar: Maximize Uptime With Powerful Monitoring</title><link>https://www.netdata.cloud/academy/webinar-maximize-uptime-minimize-stress-powerful-monitoring-solutions/</link><pubDate>Thu, 27 Jun 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/webinar-maximize-uptime-minimize-stress-powerful-monitoring-solutions/</guid><description>&lt;iframe width="100%" height="450" src="https://www.youtube.com/embed/xOJwlU_cYAE?si=nidihOQR9VN72jqg" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen>&lt;/iframe>
&lt;p>Selecting the right observability solution is essential for maintaining the performance and reliability of your IT infrastructure. With countless options available, making an informed decision can be challenging. This webinar will provide you with a comprehensive guide to navigating this complex process.&lt;/p>
&lt;p>We&amp;rsquo;ll explore the critical factors to consider, such as scalability, ease of use, integration capabilities, and cost-effectiveness. You&amp;rsquo;ll gain insights into best practices for evaluating and implementing monitoring tools that align with your business goals and technical requirements.&lt;/p></description></item><item><title>High CPU Usage Detected How To Fix CPU Overload</title><link>https://www.netdata.cloud/academy/how-to-fix-cpu-overload/</link><pubDate>Mon, 10 Jun 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/how-to-fix-cpu-overload/</guid><description>&lt;p>You get an alert: &amp;ldquo;High CPU usage detected on server-db-01.&amp;rdquo; Your application feels sluggish, users are reporting timeouts, and the system is becoming unresponsive. A maxed-out CPU is a clear sign of trouble, grinding your operations to a halt and putting your service reliability at risk. But what does high CPU usage actually mean, and more importantly, how do you fix it?&lt;/p>
&lt;p>Sustained high CPU usage, often pegged at 100%, means your server&amp;rsquo;s processor is completely saturated. It&amp;rsquo;s trying to handle more tasks than it&amp;rsquo;s capable of, leading to performance degradation, increased latency, and potential crashes. For DevOps engineers, SREs, and developers, quickly diagnosing and resolving a CPU overload is a critical skill. This guide will walk you through identifying the culprits and restoring your system to optimal health.&lt;/p></description></item><item><title>How To Achieve High Availability In CI/CD With Observability</title><link>https://www.netdata.cloud/academy/ci-cd-high-availability/</link><pubDate>Sun, 09 Jun 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/ci-cd-high-availability/</guid><description>&lt;p>Your CI/CD pipeline is the backbone of your software delivery process. When it works, code flows smoothly from commit to production. But what happens when it breaks? A failed pipeline means stalled feature releases, delayed bug fixes, and frustrated developers unable to ship their work. To prevent this, you need to treat your CI/CD infrastructure with the same rigor as your production applications, and that starts with making it highly available.&lt;/p></description></item><item><title>Key Observability Metrics | Infrastructure &amp; APM Monitoring</title><link>https://www.netdata.cloud/academy/a-guide-to-the-most-important-observability-metrics/</link><pubDate>Thu, 06 Jun 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/a-guide-to-the-most-important-observability-metrics/</guid><description>&lt;h2 id="what-is-observability-the-fundamentals">What Is Observability? The Fundamentals&lt;/h2>
&lt;p>Noone can argue that observability is crucial for maintaining the health and performance of applications and infrastructure. Observability refers to the ability to measure and understand the state of a system based on the outputs it produces. This is extremely important for identifying, diagnosing, and resolving issues effectively and efficiently.&lt;/p>
&lt;p>Observability is essential for DevOps and &lt;a href="https://www.ibm.com/think/topics/site-reliability-engineering" target="_blank">SRE&lt;/a> teams as it provides a comprehensive, overall view of the infrastructure’s health, enabling proactive maintenance and quicker incident response. It involves collecting and analyzing a variety of data types, including logs, metrics, and traces, to gain insights into system behavior and it can help discover possible anomalies throughout the whole infrastructure.&lt;/p></description></item><item><title>Introducing Netdata's Dynamic Configuration Manager</title><link>https://www.netdata.cloud/blog/netdata-dynamic-configuration-manager/</link><pubDate>Wed, 05 Jun 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-dynamic-configuration-manager/</guid><description>&lt;p>We are thrilled to unveil the latest addition to the Netdata platform: the Dynamic Configuration Manager. This powerful new feature revolutionizes how you manage your monitoring and alerting configurations, making it easier and more efficient than ever before.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="key-features-of-the-dynamic-configuration-manager">&lt;strong>Key Features of the Dynamic Configuration Manager&lt;/strong>&lt;/h2>
&lt;ol>
&lt;li>Create and Modify Alerts from Every Chart
&lt;ul>
&lt;li>You can now create and modify alerts directly from any chart on your dashboard or from the dedicated Alerts tab. This streamlined process allows for quick adjustments and ensures your monitoring is always aligned with your current needs.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Configure Collectors from the Integrations Section on the Dashboard
&lt;ul>
&lt;li>Currently available for go.d collectors, this feature lets you configure collectors straight from the Integrations section. This means you can quickly identify what Netdata can monitor and set up your configurations in one go, without having to dig through multiple settings pages.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Submit Configurations to Multiple Nodes with One Click
&lt;ul>
&lt;li>Managing configurations across a large infrastructure can be time-consuming. With the Dynamic Configuration Manager, you can now submit configurations to multiple nodes simultaneously, saving you valuable time and effort.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Construct and Copy Configurations for IaC Solutions
&lt;ul>
&lt;li>For those using Infrastructure as Code (IaC) solutions, this feature allows you to construct and copy configurations easily, integrating them into your IaC workflows. This ensures your configurations are consistent and reproducible across different environments.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ol>
&lt;h2 id="how-to-use-the-dynamic-configuration-manager">&lt;strong>How to Use the Dynamic Configuration Manager&lt;/strong>&lt;/h2>
&lt;p>To help you get started with the Dynamic Configuration Manager, we’ve put together a quick guide using the Netdata demo environment.&lt;/p></description></item><item><title>Effective Strategies for Managing PostgreSQL Deadlocks</title><link>https://www.netdata.cloud/academy/effective-strategies-for-managing-postgresql-deadlocks/</link><pubDate>Mon, 03 Jun 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/effective-strategies-for-managing-postgresql-deadlocks/</guid><description>&lt;p>Managing PostgreSQL deadlocks is critical for maintaining database performance and ensuring smooth application operations. Deadlocks occur when two or more transactions hold locks that the other transactions need, creating a cycle of dependencies with no resolution. This guide explores effective strategies for preventing and resolving deadlocks in PostgreSQL, helping &lt;a href="https://www.netdata.cloud/solutions/built-for/devops/">DevOps&lt;/a> and &lt;a href="https://www.netdata.cloud/solutions/built-for/sre/">Site Reliability Engineers (SREs)&lt;/a> ensure optimal &lt;a href="https://www.netdata.cloud/academy/what-is-cardinality-in-databases-a-comprehensive-guide/">database performance&lt;/a>.&lt;/p>
&lt;h2 id="understanding-postgresql-deadlocks">Understanding PostgreSQL Deadlocks&lt;/h2>
&lt;h3 id="what-are-deadlocks">What Are Deadlocks?&lt;/h3>
&lt;p>A deadlock in PostgreSQL happens when two or more transactions block each other, each waiting for the other to release a lock. This mutual blocking results in a standstill where none of the transactions can proceed. PostgreSQL has mechanisms to detect and resolve deadlocks by aborting one of the transactions, but this can still lead to performance issues and data inconsistencies.&lt;/p></description></item><item><title>How to automate adding nodes to rooms in Netdata?</title><link>https://www.netdata.cloud/blog/how-can-netdata-agents-be-placed-in-different-rooms-in-an-automated-way/</link><pubDate>Tue, 28 May 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/how-can-netdata-agents-be-placed-in-different-rooms-in-an-automated-way/</guid><description>&lt;p>How we organize nodes (and the Netdata agents that are running on those nodes) across different rooms should reflect our architectural decision because the room is a logical container with its own user members and notification rules. So if we are monitoring large infrastructure we should be consistent with these rules and one way to achieve this is to choose automation. &lt;a href="https://registry.terraform.io/providers/netdata/netdata/latest">Netdata Cloud Terraform Provider&lt;/a> lets you automate this by provisioning all the cloud resources and giving you the credentials to spin up the Netdata Agents. In this article, we will concentrate on how in practice we can organize and assign nodes across different rooms in two scenarios, in each of them I&amp;rsquo;m using &lt;strong>non-production&lt;/strong> installation of the Netdata Agents:&lt;/p></description></item><item><title>How To Fix 504 Gateway Timeout Errors In NGINX</title><link>https://www.netdata.cloud/academy/how-to-diagnose-and-fix-504-gateway-timeout-errors-in-nginx/</link><pubDate>Mon, 27 May 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/how-to-diagnose-and-fix-504-gateway-timeout-errors-in-nginx/</guid><description>&lt;p>Encountering a &amp;lsquo;504 Gateway Timeout&amp;rsquo; error in Nginx can be frustrating, especially when it disrupts the availability of your web applications. This error indicates that the server, while acting as a gateway or proxy, did not receive a timely response from the upstream server. This guide will help DevOps and Site Reliability Engineers (SRE) diagnose and fix these errors efficiently.&lt;/p>
&lt;h2 id="tldr">TL;DR&lt;/h2>
&lt;ul>
&lt;li>A 504 Gateway Timeout means Nginx, acting as a proxy, didn&amp;rsquo;t get a response from the upstream server in time, so the problem sits in the backend, the network, or your timeout settings rather than in Nginx itself.&lt;/li>
&lt;li>Diagnose before you change anything: check the Nginx and upstream logs, measure backend response time with curl, and test connectivity with ping, traceroute, or mtr.&lt;/li>
&lt;li>The fastest stopgap is raising Nginx&amp;rsquo;s timeout directives, but the durable fix is making the upstream faster through resource scaling, query optimization, and load balancing.&lt;/li>
&lt;li>Adding Nginx caching reduces how often requests hit the upstream at all, which keeps a strained backend responsive and cuts down 504s.k.&lt;/li>
&lt;/ul>
&lt;h2 id="understanding-the-504-gateway-timeout-error">Understanding The &amp;lsquo;504 Gateway Timeout&amp;rsquo; Error&lt;/h2>
&lt;p>The &amp;lsquo;504 Gateway Timeout&amp;rsquo; error typically occurs in Nginx when:&lt;/p></description></item><item><title>How to Troubleshoot Slow Queries in MongoDB</title><link>https://www.netdata.cloud/academy/how-to-troubleshoot-slow-queries-in-mongodb/</link><pubDate>Mon, 27 May 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/how-to-troubleshoot-slow-queries-in-mongodb/</guid><description>&lt;p>MongoDB is a popular NoSQL database known for its flexibility and scalability. However, as with any database system, you might encounter slow queries that can impact the &lt;a href="https://www.netdata.cloud/academy/what-is-application-performance-monitoring-apm/">performance of your application&lt;/a>. In this guide, we’ll walk you through the steps to troubleshoot and optimize slow queries in MongoDB, ensuring your database runs efficiently.&lt;/p>
&lt;h2 id="understanding-the-basics">Understanding the Basics&lt;/h2>
&lt;p>Before diving into troubleshooting, it&amp;rsquo;s important to understand some basic concepts in MongoDB:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Collections and Documents:&lt;/strong> MongoDB stores data in collections, which are analogous to tables in relational databases. Each collection contains documents, which are JSON-like data structures.&lt;/li>
&lt;li>&lt;strong>Indexes:&lt;/strong> Indexes improve query performance by allowing the database to quickly locate the data without scanning every document.&lt;/li>
&lt;li>&lt;strong>Query Plans:&lt;/strong> MongoDB evaluates different ways to execute a query and chooses the most efficient plan.&lt;/li>
&lt;/ul>
&lt;h3 id="identifying-slow-queries">Identifying Slow Queries&lt;/h3>
&lt;p>Using the &lt;code>slowms&lt;/code> Parameter
MongoDB logs operations that take longer than a specified threshold. By default, this threshold (&lt;code>slowms&lt;/code>) is set to 100 milliseconds. You can adjust this setting to catch slower operations more effectively:&lt;/p></description></item><item><title>DevOps Best Practices Playbook For Performance</title><link>https://www.netdata.cloud/academy/devops-playbook-best-practices-for-success/</link><pubDate>Wed, 22 May 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/devops-playbook-best-practices-for-success/</guid><description>&lt;p>In software development and IT operations, DevOps has become a central practice, facilitating collaboration, increasing efficiency, and ensuring high-quality software delivery. However, the range of DevOps operations can be intimidating for beginners. This playbook aims to simplify DevOps by outlining best practices for developers and teams, enabling more efficient workflows and operational excellence.&lt;/p>
&lt;h2 id="devops-meaning-bridging-software-development--it-operations">Devops Meaning: Bridging Software Development &amp;amp; IT Operations&lt;/h2>
&lt;p>DevOps is a cultural and technical shift that bridges the historically isolated worlds of software development and IT operations. The goal is to construct a fully automated pipeline for developing, testing, delivering, and maintaining software. Think of it as running a race where each runner hands off the baton seamlessly, boosting overall speed and efficiency.&lt;/p></description></item><item><title>How To Speed Up Windows 10 &amp; 11: Performance Tips</title><link>https://www.netdata.cloud/academy/how-to-speed-up-windows/</link><pubDate>Fri, 17 May 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/how-to-speed-up-windows/</guid><description>&lt;p>Dealing with a slow computer is frustrating. Whether you&amp;rsquo;re compiling code, running &lt;a href="https://www.netdata.cloud/academy/container-vs-vm-which-is-better-for-you/">virtual machines&lt;/a>, analyzing data, or just trying to browse documentation, sluggish performance pc issues can significantly hinder your productivity and break your focus. Waiting for applications to load or the system to respond eats up valuable time and can turn simple tasks into chores. Fortunately, there are numerous ways to speed up computer performance on Windows.&lt;/p>
&lt;p>This guide provides actionable tips to optimize pc responsiveness, ranging from simple maintenance routines and software adjustments to more impactful hardware upgrades. We&amp;rsquo;ll cover how to improve computer performance and get your Windows system running smoothly again.&lt;/p></description></item><item><title>What Is An HPC Cluster - Key Components &amp; How It Works</title><link>https://www.netdata.cloud/academy/what-is-an-hpc-cluster/</link><pubDate>Wed, 01 May 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-an-hpc-cluster/</guid><description>&lt;p>Imagine needing to analyze petabytes of genomic data, simulate the airflow over a new aircraft wing, or predict complex financial market movements. A single computer, no matter how powerful, would struggle or take an impractically long time. This is where High-Performance Computing (HPC) clusters come in. They provide the immense computational power needed to tackle problems far beyond the reach of standard computing. If you&amp;rsquo;re stepping into roles involving large-scale data processing or complex simulations, understanding HPC clusters is essential.&lt;/p></description></item><item><title>Costa Tsaousis Interview: Homelab Show Insights</title><link>https://www.netdata.cloud/blog/interview-recap-with-costa-tsaousis-ceo-of-netdata-insights-from-the-homelab-show/</link><pubDate>Fri, 26 Apr 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/interview-recap-with-costa-tsaousis-ceo-of-netdata-insights-from-the-homelab-show/</guid><description>&lt;p>Costa Tsaousis, founder and chief visionary of Netdata, recently shared his expertise on real-time monitoring&amp;rsquo;s pivotal role in modern IT landscapes during an episode of ‘The Homelab Show’ on YouTube. If you missed the live stream, here’s an essential summary of the discussion’s key points.&lt;/p>
&lt;iframe width="560" height="315" src="https://www.youtube.com/embed/WvXab8MkRS4?si=QkJE_brYT3aZ1FZQ" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen>&lt;/iframe>
&lt;h2 id="the-necessity-of-real-time-monitoring">The Necessity of Real-Time Monitoring&lt;/h2>
&lt;p>Nowadays, data flows incessantly and operational demands are continuous, the importance of real-time monitoring cannot be overstated.&lt;/p></description></item><item><title>Decentralized Monitoring Explained</title><link>https://www.netdata.cloud/blog/decentralized-or-distributed-monitoring-explained/</link><pubDate>Thu, 11 Apr 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/decentralized-or-distributed-monitoring-explained/</guid><description>&lt;h2 id="introduction-to-distributed-observability">Introduction to Distributed Observability&lt;/h2>
&lt;p>Users often find themselves puzzled by the concepts of decentralized or distributed monitoring. This confusion is likely due to many monitoring systems claiming distributed capabilities, making it challenging to discern how Netdata stands out.&lt;/p>
&lt;p>To grasp the distinction, we must delve into the evolution of monitoring systems.&lt;/p>
&lt;p>When the first monitoring systems were created, about 20-25 years ago, they were built as SNMP collectors. The monitoring application was installed on a server, configured to discover network devices via SNMP, pulling data once every minute, per device. Simultaneously, the monitoring system was exposing a daemon to collect SNMP traps (key events generated by the network devices, pushed to the monitoring system).&lt;/p></description></item><item><title>Netdata is the only real-time monitoring solution: Justified</title><link>https://www.netdata.cloud/blog/netdata-is-the-only-real-time-monitoring-solution-justified/</link><pubDate>Wed, 10 Apr 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-is-the-only-real-time-monitoring-solution-justified/</guid><description>&lt;p>In the digital era, where data flows like a ceaseless river, real-time monitoring stands as a pivotal technology, allowing organizations to not only keep pace but also to deeply understand the intricate dance of their operational ecosystems. This technology is not just about keeping tabs; it&amp;rsquo;s about gaining a profound, almost intuitive sense of the micro-worlds within which systems, containers, services, and applications pulse and thrive.&lt;/p>
&lt;p>Real-time monitoring is the art and science of tracking system performance, activities, or transactions continuously and automatically, providing the ability to analyze and visualize data the moment it&amp;rsquo;s generated.&lt;/p></description></item><item><title>Understanding Monitoring Tools</title><link>https://www.netdata.cloud/blog/understanding-monitoring-tools/</link><pubDate>Wed, 10 Apr 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/understanding-monitoring-tools/</guid><description>&lt;p>If you care about operational excellence when it comes to your IT infrastructure, the role of monitoring systems is pivotal. As we navigate through the myriad of available monitoring tools, it becomes essential to understand the distinct architectures, styles, and focal points of various monitoring solutions, as well as the time-to-value they offer. This blog post aims to demystify the landscape of monitoring systems, providing a comprehensive overview that categorizes these tools into four primary architectural design principles.&lt;/p></description></item><item><title>SafetyDetectives: An Interview-With-Costa-Tsaousis</title><link>https://www.netdata.cloud/blog/safetydetectives/</link><pubDate>Mon, 01 Apr 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/safetydetectives/</guid><description>&lt;p>In a recent conversation with SafetyDetectives, Costa Tsaousis, CEO and founder of Netdata, shares insights into the inception and evolution of Netdata, a game-changing monitoring solution. With a background in fintech and a passion for real-time data processing, Tsaousis was driven to create Netdata in response to the significant gaps he identified in traditional monitoring tools. Emphasizing the importance of real-time data, comprehensive metrics collection, and the innovative use of machine learning, Tsaousis discusses how Netdata is setting new standards in the monitoring industry. His vision for Netdata not only challenges the status quo but also introduces a novel approach to cybersecurity, making it an essential tool for organizations worldwide.&lt;/p></description></item><item><title>Overcoming Monitoring Challenges in Shared VM Environments</title><link>https://www.netdata.cloud/case-studies/technology/id5/</link><pubDate>Thu, 28 Mar 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/technology/id5/</guid><description>&lt;h2 id="optimizing-resource-management-in-complex-vm-ecosystems">Optimizing Resource Management in Complex VM Ecosystems&lt;/h2>
&lt;p>ID5 Web Solutions, a web solutions provider in Brazil, faced significant challenges in managing its shared VM environment. With a diverse clientele depending on their infrastructure, the need for precise and efficient resource usage identification per user was paramount. The company sought to not only allocate workloads more effectively but also to ensure fair pricing strategies based on actual consumption.&lt;/p>
&lt;p>The complexity of monitoring several shared VMs and the necessity for detailed user-specific resource usage insights posed substantial hurdles. Additionally, ID5 Web Solutions aimed to incorporate a seamless alerting mechanism via Discord, enhancing their response capabilities to incidents and optimizing service availability.&lt;/p></description></item><item><title>Transforming Server Monitoring across Diverse Hosts</title><link>https://www.netdata.cloud/case-studies/fintech/urios/</link><pubDate>Thu, 28 Mar 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/fintech/urios/</guid><description>&lt;h2 id="mastering-monitoring-complexity">Mastering Monitoring Complexity&lt;/h2>
&lt;p>URIOS, operating in the financial services sector, faced a significant challenge in monitoring its diverse Linux server environments. Spread across various infrastructures and hosts, the complexity of managing these servers was growing. The organization required a solution that could centralize and unify its monitoring efforts, offering precise and numerous indicators to stay ahead of potential issues.&lt;/p>
&lt;p>The multitude of infrastructures and hosting providers made it difficult for them to maintain a consistent monitoring approach. They soon realized that they needed a tool that could bring everything together in a cohesive manner.&lt;/p></description></item><item><title>TTIC Case Study: Efficient System Monitoring</title><link>https://www.netdata.cloud/case-studies/education/ttic/</link><pubDate>Thu, 28 Mar 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/education/ttic/</guid><description>&lt;h2 id="enhancing-system-management-with-scarce-time-resources">Enhancing System Management with Scarce Time Resources&lt;/h2>
&lt;p>At the Toyota Technological Institute at Chicago, the challenge of system administration is uniquely compounded by the scarcity of time. As a one-man IT department, Adam Bohlander juggles various responsibilities, making efficient time management crucial. The institute&amp;rsquo;s evolving needs demand a monitoring solution that aligns with its mission of leading in computer science and information technology research and education.&lt;/p>
&lt;blockquote>
&lt;p>&amp;ldquo;I find the install quick and easy via ansible (my own playbooks) and defaults to be quite sane. This allows me to use it without investing that much time.&amp;rdquo;&lt;/p></description></item><item><title>Manage Netdata Cloud with Terraform</title><link>https://www.netdata.cloud/blog/netdata-terraform-provider/</link><pubDate>Wed, 27 Mar 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-terraform-provider/</guid><description>&lt;p>We proudly announce the release of the &lt;a href="https://registry.terraform.io/providers/netdata/netdata/latest">Netdata Cloud Terraform Provider&lt;/a>.&lt;/p>
&lt;p>It&amp;rsquo;s a step forward to make our platform more automated and compliant with the modern Infrastructure as Code approach. &lt;a href="https://www.terraform.io/">Terraform&lt;/a> is one of the leaders in the IaC tools with a rich ecosystem of providers and modules, now you can put a puzzle with Netdata Cloud to your stack.&lt;/p>
&lt;p>The initial iteration of the Netdata Cloud Terraform Provider supports the following resources:&lt;/p></description></item><item><title>University of Calgary: Netdata in Academia</title><link>https://www.netdata.cloud/case-studies/education/calgary-university/</link><pubDate>Tue, 26 Mar 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/education/calgary-university/</guid><description>&lt;h2 id="empowering-academic-research-with-real-time-monitoring">Empowering Academic Research with Real-Time Monitoring&lt;/h2>
&lt;p>At the &lt;strong>Machine Learning Lab&lt;/strong> at the &lt;strong>University of Calgary&lt;/strong>, managing a robust infrastructure of servers and workstations is critical for advancing their research in machine learning (ML). Assistant Professor &lt;strong>Yani Ioannou&lt;/strong> oversees this infrastructure, ensuring that these essential resources remain operational and are used to their fullest potential. The lab&amp;rsquo;s challenge lies not just in the maintenance of these resources but in minimizing downtime to keep the research moving forward.&lt;/p></description></item><item><title>Enhancing E-Learning Through Proactive Monitoring</title><link>https://www.netdata.cloud/case-studies/elearning/ingenium/</link><pubDate>Wed, 20 Mar 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/elearning/ingenium/</guid><description>&lt;h2 id="tackling-e-learning-monitoring-challenges">Tackling E-learning Monitoring Challenges&lt;/h2>
&lt;p>Ingenium Digital Learning, a major player in the e-learning sector, faces unique challenges in delivering high-quality educational services online. One primary concern they had was managing disk space usage and CPU load effectively, as users continuously upload data without a clear view of their capacity limits. Additionally, unexpected spikes in server load, often due to unannounced large-scale events, further complicate resource management.&lt;/p>
&lt;blockquote>
&lt;p>&amp;ldquo;Clients do not warn us of big events, which impacts the load on the servers. Users can upload data to our servers but don&amp;rsquo;t have a view on their total capacity.&amp;rdquo;&lt;/p></description></item><item><title>Dynatrace vs Datadog vs Instana vs Grafana vs Netdata!</title><link>https://www.netdata.cloud/blog/netdata-vs-datadog-dynatrace-instana-grafana/</link><pubDate>Sun, 17 Mar 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-vs-datadog-dynatrace-instana-grafana/</guid><description>&lt;p>In this post, we delve into the comparative analysis of the commercial offerings of five leading monitoring solutions—Dynatrace, &lt;a href="https://www.netdata.cloud/blog/5-datadog-alternatives/">Datadog&lt;/a>, Instana, Grafana, and Netdata. Our objective is to unravel the intrinsic value each of these services offers when applied to a real-world scenario. To accomplish this, we employed trial subscriptions of these services to monitor a set of Ubuntu servers and VMs, each hosting a pair of widely-used applications: NGINX and PostgreSQL, along with a couple of Docker and LXC containers. Additionally, we extended our monitoring to physical servers to evaluate the efficacy of these tools in capturing hardware and sensor data along with VMs monitored from the host.&lt;/p></description></item><item><title>Monitoring Industrial Equipment Manufacturing with Netdata</title><link>https://www.netdata.cloud/case-studies/manufacturing/autonoma/</link><pubDate>Thu, 14 Mar 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/manufacturing/autonoma/</guid><description>&lt;h2 id="empowering-precision-and-efficiency">Empowering Precision and Efficiency&lt;/h2>
&lt;p>Autonoma Technologies GmbH, a leader in the industrial equipment manufacturing sector, has been at the forefront of integrating digital solutions to address the complexities of modern industrial challenges. Monitoring many different services without clear insights into critical metrics for each, Autonoma turned to Netdata for a solution. The need for deep knowledge across various services to identify and monitor critical metrics was a significant hurdle.&lt;/p>
&lt;p>With a primary focus on monitoring vital systems like disk space, RabbitMQ, VerneMQ, NGINX, Redis, Fail2ban, and Kubernetes clusters, Autonoma required a robust and flexible tool to keep its production setup fast and reliable. Netdata became their solution of choice, offering out-of-the-box analysis for specific services that drastically cut down the time and effort needed for creating and adapting dashboards.&lt;/p></description></item><item><title>New Streamlined Plan Structure</title><link>https://www.netdata.cloud/blog/netdata-unified-plans/</link><pubDate>Wed, 06 Mar 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-unified-plans/</guid><description>&lt;blockquote>
&lt;p>&lt;strong>UPDATE:&lt;/strong> Netdata is introducing a streamlined plan structure, sunsetting Early Bird plans on 13-03-2024.&lt;/p>
&lt;/blockquote>
&lt;!--truncate-->
&lt;p>As the landscape of real-time monitoring evolves, so does the diversity and complexity of use cases that our community brings to Netdata. Our mission has always been to democratize monitoring by making it accessible, powerful, and scalable for everyone. With the rapid growth of our user base and their expanding needs, it&amp;rsquo;s become clear that our plan structure must evolve to maintain this mission sustainably.&lt;/p></description></item><item><title>ACPro Case Study: Effortless Manufacturing Monitoring</title><link>https://www.netdata.cloud/case-studies/manufacturing/acpro/</link><pubDate>Tue, 27 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/manufacturing/acpro/</guid><description>&lt;h2 id="unifying-monitoring-across-diverse-environments">Unifying Monitoring Across Diverse Environments&lt;/h2>
&lt;p>ACPro faced the challenge of overseeing a wide array of in-house applications scattered across numerous servers, all without a monitoring solution in place. This lack of visibility made it difficult to maintain the high uptime and availability standards that are critical in the manufacturing industry. The search for a monitoring tool that was both comprehensive and user-friendly led ACPro to Netdata, which stood out for its intuitive interface and extensive, yet easily digestible, metrics.&lt;/p></description></item><item><title>InRento Case Study: Fintech Cloud Monitoring Wins</title><link>https://www.netdata.cloud/case-studies/fintech/inrento/</link><pubDate>Tue, 27 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/fintech/inrento/</guid><description>&lt;h2 id="simplifying-complex-cloud-environments">Simplifying Complex Cloud Environments&lt;/h2>
&lt;p>The migration of Inrento&amp;rsquo;s infrastructure to the cloud introduced a new set of challenges, particularly in monitoring multiple servers with diverse roles. The transition from a single server environment to a distributed application on AWS required a solution that not only automated the monitoring process but also made it accessible and understandable to developers accustomed to a monolithic setup.&lt;/p>
&lt;blockquote>
&lt;p>&amp;ldquo;Netdata enables automated integration of new servers into our architecture. As we experiment with various configurations, we can leverage automated configuration that is prepared beforehand.&amp;rdquo;&lt;/p></description></item><item><title>Reduced Costs &amp; Kubernetes Efficiency</title><link>https://www.netdata.cloud/case-studies/technology/7-oaks-group/</link><pubDate>Tue, 27 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/technology/7-oaks-group/</guid><description>&lt;h2 id="simplifying-complexity-with-netdata">Simplifying Complexity with Netdata&lt;/h2>
&lt;p>At 7 Oaks Group, our journey with Netdata began out of curiosity and a desire to move away from the frustrations of complex monitoring setups. Our infrastructure, primarily composed of &lt;strong>Kubernetes clusters and a selection of EC2/VM instances&lt;/strong>, required a monitoring solution that could provide instant, meaningful insights without the hassle of intricate configuration. The immediate value we discovered in Netdata was nothing short of revolutionary, offering us a level of insight into our systems that was both effortlessly attainable and profoundly impactful.&lt;/p></description></item><item><title>Easy Doc: Revolutionizing IT Operations with Netdata</title><link>https://www.netdata.cloud/case-studies/technology/easydocs/</link><pubDate>Tue, 13 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/technology/easydocs/</guid><description>&lt;h2 id="overcoming-it-operational-challenges">Overcoming IT Operational Challenges&lt;/h2>
&lt;p>At Easy Doc Soluções Integradas, navigating the complexities of docker swarm cluster infrastructure presented significant challenges, particularly in monitoring the health and availability of containers and cluster nodes. The primary pain point was the difficulty in detecting and diagnosing the reasons behind API unavailability, especially since container restarts resulted in the loss of crucial logs. This obstacle hindered our ability to maintain high levels of service availability and performance, directly impacting our operational efficiency and customer satisfaction.&lt;/p></description></item><item><title>Streamlining IT Operations with Netdata</title><link>https://www.netdata.cloud/case-studies/technology/clew/</link><pubDate>Tue, 13 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/technology/clew/</guid><description>&lt;h2 id="addressing-monitoring-complexity">Addressing Monitoring Complexity&lt;/h2>
&lt;p>The main challenge for the organization was ensuring up-to-date, comprehensive monitoring across a diverse infrastructure, complicated by different operating systems, security requirements, and network segmentation. The varied nature of their environment made it difficult to maintain a unified view of system health and performance, often leading to delayed root cause analysis and inefficient troubleshooting processes.&lt;/p>
&lt;p>Netdata&amp;rsquo;s ease of deployment and out-of-the-box functionality have been game-changers, offering immediate visibility into their systems&amp;rsquo; health with minimal setup. Its comprehensive monitoring capabilities allowed them to quickly identify and address issues, reducing the time and effort previously required for these tasks.&lt;/p></description></item><item><title>Upcoming Homelab Plan</title><link>https://www.netdata.cloud/blog/netdata-homelab-plan/</link><pubDate>Wed, 07 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-homelab-plan/</guid><description>&lt;blockquote>
&lt;p>&lt;strong>UPDATE:&lt;/strong> On the 2024-02-08 A new Netdata Cloud Homelab plan will become available. This is aimed for home users and students.&lt;/p>
&lt;/blockquote>
&lt;h2 id="what-you-need-to-know">What you need to know?&lt;/h2>
&lt;ul>
&lt;li>
&lt;p>New Plan alert: We&amp;rsquo;re introducing a dedicated plan on Netdata Cloud—the Homelab plan. It&amp;rsquo;s tailored to meet the needs of home users and students, offering unrestricted access to Netdata features without the limitations seen in the Community plan.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Exclusively for Personal Use: The Homelab plan is designed for personal, non-commercial use only. To qualify, users will need to self-certify as a home user or student during the sign-up process.&lt;/p></description></item><item><title>Codyas: Global Server Management with Netdata</title><link>https://www.netdata.cloud/case-studies/technology/codyas/</link><pubDate>Thu, 01 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/technology/codyas/</guid><description>&lt;h2 id="overcoming-server-management-hurdles">Overcoming Server Management Hurdles&lt;/h2>
&lt;p>&lt;a href="https://www.codyas.com/">Codyas&lt;/a> faces the intricate challenge of managing a vast array of servers, equipped with different technologies for various clients around the globe. The primary obstacle has been the allocation of limited time and personnel towards monitoring tasks, which detracted from the organization&amp;rsquo;s core focus on development and customer service.&lt;/p>
&lt;blockquote>
&lt;p>&amp;ldquo;Netdata&amp;rsquo;s ease of installation and configuration has been a game-changer for us, allowing quick setup and efficient monitoring across our diverse server landscape,&amp;rdquo;&lt;/p></description></item><item><title>Enhancing Research Through Advanced Monitoring</title><link>https://www.netdata.cloud/case-studies/education/landcare/</link><pubDate>Thu, 01 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/education/landcare/</guid><description>&lt;h2 id="helping-scientists-focus-on-the-science">Helping Scientists focus on the Science&lt;/h2>
&lt;p>At &lt;a href="https://landcareresearch.co.nz/">Manaaki Whenua – Landcare Research&lt;/a>, managing multi-user Linux workstations for scientific computing presented a unique challenge. Tasked with overseeing system utilization without a dedicated role in the organization, the need for a solution that was intuitive and efficient became paramount. Netdata&amp;rsquo;s simplicity in setup and the ability to monitor systems remotely via the cloud addressed these challenges head-on, offering a seamless solution for workstation management.&lt;/p></description></item><item><title>From Bottlenecks to Breakthroughs</title><link>https://www.netdata.cloud/case-studies/technology/lancom-systems/</link><pubDate>Thu, 01 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/technology/lancom-systems/</guid><description>&lt;h2 id="simplifying-it-monitoring-challenges">Simplifying IT Monitoring Challenges&lt;/h2>
&lt;p>At LANCOM Systems, the transition to a more integrated and efficient IT monitoring system was driven by a desire to streamline operations and overcome the challenges associated with complex system monitoring. The company faced significant hurdles in integrating services within Docker containers and adapting to the programming language Go, despite a strong preference for Python. This challenge was further compounded by the need for a comprehensive monitoring solution that could offer extensive out-of-the-box features without necessitating extensive development work.&lt;/p></description></item><item><title>Leica Biosystems: Advancing Cancer Diagnostics</title><link>https://www.netdata.cloud/case-studies/healthcare/leica/</link><pubDate>Thu, 01 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/healthcare/leica/</guid><description>&lt;h2 id="navigating-offline-monitoring-challenges">Navigating Offline Monitoring Challenges&lt;/h2>
&lt;p>&lt;a href="https://www.leicabiosystems.com/">Leica Biosystems&lt;/a>, a leader in pathology and cancer diagnostics, faced a unique challenge: monitoring and observability in strictly offline environments. The need for a dashboard providing system insights in an offline capacity was paramount, especially given the constraints of remote access impossibility due to the offline nature of the devices used.&lt;/p>
&lt;blockquote>
&lt;p>&amp;ldquo;Monitoring complexity in offline environments was a significant hurdle. Netdata&amp;rsquo;s ability to export metrics allowed us to achieve the observability we needed without the need for online connectivity.&amp;rdquo;&lt;/p></description></item><item><title>Monitoring on Autopilot</title><link>https://www.netdata.cloud/case-studies/hosting/webnestify/</link><pubDate>Thu, 01 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/hosting/webnestify/</guid><description>&lt;h2 id="overcoming-scalability-issues">Overcoming Scalability Issues&lt;/h2>
&lt;p>Simon Gajdosik, the driving force behind &lt;a href="https://webnestify.cloud/">Webnestify&lt;/a>, identified scalability and maintainability as primary challenges in their monitoring strategy. The traditional approach with &lt;strong>Grafana and Prometheus, although effective, became untenable at scale&lt;/strong>. Netdata Cloud provided the much-needed solution, offering automated monitoring that effortlessly scales with the business, eliminating the complexity of manual configurations.&lt;/p>
&lt;h2 id="enhancing-operations-with-netdata">Enhancing Operations with Netdata&lt;/h2>
&lt;p>&lt;a href="https://webnestify.cloud/">Webnestify&lt;/a>&amp;rsquo;s adoption of Netdata Cloud provided significant &lt;strong>operational enhancements&lt;/strong>:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Time Savings:&lt;/strong> The automated service discovery and configuration capabilities of Netdata have saved Webnestify at least 10 hours weekly, allowing the team to focus more on business growth rather than infrastructure maintenance.&lt;/li>
&lt;li>&lt;strong>Productivity Improvements:&lt;/strong> With the efficiency gains from Netdata, Webnestify has increased its client onboarding capacity, directly contributing to business expansion.&lt;/li>
&lt;li>&lt;strong>Issue Resolution:&lt;/strong> The advanced alerting system of Netdata enabled proactive anomaly detection, preventing potential downtimes for critical ecommerce platforms and thereby ensuring uninterrupted service.&lt;/li>
&lt;li>&lt;strong>Cost Reduction:&lt;/strong> Shifting from a dedicated Grafana server to Netdata resulted in a 40% reduction in annual monitoring costs, translating to significant financial savings.&lt;/li>
&lt;li>&lt;strong>Enhanced Security:&lt;/strong> The ability to whitelist parent node IP addresses and use private IPs for child nodes improved the overall security posture.&lt;/li>
&lt;li>&lt;strong>Improved Troubleshooting and Metrics Accuracy:&lt;/strong> The intuitive dashboards and accurate metrics from Netdata have simplified the process of identifying and resolving issues, making troubleshooting 50% easier.&lt;/li>
&lt;li>&lt;strong>Downtime Reduction and SLA Improvements:&lt;/strong> With Netdata, Webnestify has seen a 20% reduction in downtime and significantly improved its SLA compliance by resolving issues more efficiently.&lt;/li>
&lt;/ul>
&lt;blockquote>
&lt;p>&amp;ldquo;A standout moment for us was when Netdata&amp;rsquo;s anomaly detection alerted us to a potential issue, allowing us to avert a crisis on a server hosting over 100 high-traffic ecommerce sites. This instance underscored the indispensable role of Netdata in our operations,&amp;rdquo;&lt;/p></description></item><item><title>Optimizing The FinTech Industry</title><link>https://www.netdata.cloud/case-studies/fintech/yellow-brick-road/</link><pubDate>Thu, 01 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/fintech/yellow-brick-road/</guid><description>&lt;h2 id="optimizing-the-fintech-industry-yellow-brick-roads-monitoring-journey-with-netdata">Optimizing The FinTech Industry: Yellow Brick Road&amp;rsquo;s Monitoring Journey with Netdata&lt;/h2>
&lt;p>&lt;strong>Yellow Brick Road&lt;/strong>, a leader in the financial services industry, embarked on a mission to fine-tune its system and kernel parameters to minimize network latency crucial for quick trading decisions. The challenge was not just in identifying what to measure but also in how to effectively use the data obtained.&lt;/p>
&lt;h3 id="the-netdata-solution">The Netdata Solution&lt;/h3>
&lt;p>Netdata Cloud came to the rescue, offering &lt;strong>comprehensive baseline monitoring&lt;/strong> that was both easy to deploy and cost-efficient. The adoption of Netdata, provided the team at Yellow Brick Road with the necessary tools to monitor their Nomad, Consul, bare metal, and systemd applications efficiently.&lt;/p></description></item><item><title>Real-Time Insights for Reliable Web Hosting</title><link>https://www.netdata.cloud/case-studies/hosting/trillium/</link><pubDate>Thu, 01 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/hosting/trillium/</guid><description>&lt;h2 id="simplifying-monitoring-for-enhanced-performance">Simplifying Monitoring for Enhanced Performance&lt;/h2>
&lt;p>&lt;a href="https://trillium.host/">Trillium&lt;/a> was confronted with the challenge of finding an effective monitoring solution that could be easily integrated with their systems, despite the non-technical background of some team members. The company aimed to uphold its commitment to outstanding service uptime and performance—a critical aspect in the hosting industry. The search for a monitoring tool that could provide comprehensive insights into the health of their infrastructure, particularly RAM and CPU usage, and uptime tracking, led them to Netdata.&lt;/p></description></item><item><title>Streamlining VM Monitoring with Netdata</title><link>https://www.netdata.cloud/case-studies/technology/comtegra/</link><pubDate>Thu, 01 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/technology/comtegra/</guid><description>&lt;h2 id="simplifying-virtual-machine-monitoring">Simplifying Virtual Machine Monitoring&lt;/h2>
&lt;p>Comtegra, a leader in IT solutions for storage and virtualization, faced a significant challenge in managing a dynamic environment of virtual machines (VMs). The frequent deployment and deletion of VMs, often within a single day, created a complex scenario that was both time-consuming and difficult to monitor. This challenge was particularly pronounced for Mariusz Mieszkowski, a storage specialist at Comtegra, who noted the inefficiency of performance testing and overall infrastructure management under these conditions.&lt;/p></description></item><item><title>Transforming Container and Server Health Insights</title><link>https://www.netdata.cloud/case-studies/gaming/nodecraft/</link><pubDate>Thu, 01 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/gaming/nodecraft/</guid><description>&lt;h2 id="revolutionizing-infrastructure-monitoring">Revolutionizing Infrastructure Monitoring&lt;/h2>
&lt;p>Nodecraft faced significant challenges in monitoring its ever-expanding fleet of dedicated servers and containers. The complexity of tracking system health and hardware performance became a bottleneck, impacting its ability to deliver the seamless gaming experiences Nodecraft is known for. With a small development team, the creation and maintenance of a custom solution were out of the question. Nodecraft&amp;rsquo;s journey towards finding a solution led it to Netdata, which has been transformative.&lt;/p></description></item><item><title>Transforming critical software development</title><link>https://www.netdata.cloud/case-studies/technology/adastec/</link><pubDate>Thu, 01 Feb 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/technology/adastec/</guid><description>&lt;h2 id="revolutionizing-vehicle-automation">Revolutionizing Vehicle Automation&lt;/h2>
&lt;p>&lt;a href="https://www.adastec.com/">ADASTEC&lt;/a>, a frontrunner in automated driving software for commercial vehicles, utilizes Netdata to ensure the highest standards of reliability and efficiency in their operations. Specializing in SAE Level-4 automation, ADASTEC empowers OEMs to craft modern, automated, and connected vehicles. The challenge of distinguishing crucial alerts within their software stack posed a significant hurdle, as Tayfun Yurdaer, a DevOps specialist at ADASTEC, explains. The essential task was identifying alerts pivotal for the mission-critical applications to respond appropriately, amidst the complexity of their software environment.&lt;/p></description></item><item><title>Empowering Digital Transformation</title><link>https://www.netdata.cloud/case-studies/government/falkland/</link><pubDate>Tue, 30 Jan 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/government/falkland/</guid><description>&lt;h2 id="maximizing-resources-elevating-services">Maximizing Resources, Elevating Services&lt;/h2>
&lt;p>The &lt;a href="https://www.falklands.gov.fk/">Falkland Islands Government&lt;/a> faced critical challenges in gaining visibility into the resource consumption of its digital assets. This lack of clarity hindered the ability to efficiently manage alerts and optimize resource usage across digital platforms. The integration of better monitoring tools was imperative to overcoming these obstacles and ensuring the seamless provision of government services.&lt;/p>
&lt;blockquote>
&lt;p>&amp;ldquo;Not having a clear picture on the resource consumption of our digital assets was a major challenge. With Netdata agent, I am now able to get a comprehensive set of data about each of my nodes, which has revolutionized how we manage resources and handle alerts.&amp;rdquo;&lt;/p></description></item><item><title>Revolutionizing Device Monitoring at SafeSize</title><link>https://www.netdata.cloud/case-studies/retail/safesize/</link><pubDate>Tue, 30 Jan 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/retail/safesize/</guid><description>&lt;h2 id="enhancing-efficiency--transparency-in-retail">Enhancing Efficiency &amp;amp; Transparency in Retail&lt;/h2>
&lt;p>&lt;a href="https://www.safesize.com/">SafeSize&lt;/a>, a trailblazer in the retail technology sector, leverages innovative 3D foot scanning and shoe recommendation solutions to enhance the shopping experience in both physical and online stores. However, managing and monitoring devices installed in diverse customer environments posed significant challenges. These devices, crucial for delivering SafeSize&amp;rsquo;s cutting-edge service, operate in varied network environments with limited access, necessitating a monitoring solution that is both efficient and resource-light.&lt;/p></description></item><item><title>Revolutionizing Server Monitoring at Tilkal</title><link>https://www.netdata.cloud/case-studies/logistics/tilkal/</link><pubDate>Tue, 30 Jan 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/logistics/tilkal/</guid><description>&lt;h2 id="tackling-server-monitoring-complexity">Tackling Server Monitoring Complexity&lt;/h2>
&lt;p>At &lt;a href="https://www.tilkal.com/">Tilkal SAS&lt;/a>, the challenge of aggregating infrastructure information from numerous standalone Linux servers into a singular platform was significant. The complexity was further compounded by the need to incorporate custom metrics emitted by proprietary services, alongside standard system statistics. This requirement for a holistic view of both standard and custom data sources necessitated a versatile and powerful monitoring tool.&lt;/p>
&lt;blockquote>
&lt;p>&amp;ldquo;The data we wanted to collect was not only standard monitoring stats but also custom stats emitted by our own services. Netdata&amp;rsquo;s ease of deployment and its cloud app&amp;rsquo;s capability to consolidate server stats in one place were exactly what we needed.&amp;rdquo;&lt;/p></description></item><item><title>Simplifying Complex Monitoring for Educational Institutions</title><link>https://www.netdata.cloud/case-studies/education/funiber/</link><pubDate>Tue, 30 Jan 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/education/funiber/</guid><description>&lt;h2 id="navigating-a-complex-monitoring-terrain">Navigating a Complex Monitoring Terrain&lt;/h2>
&lt;p>The Fundación Universitaria Iberoamericana (&lt;a href="https://www.funiber.us/">FUNIBER&lt;/a>), with its expansive network of institutions and students across Spain and beyond, faces the unique challenge of managing and configuring monitoring tools across a dynamically evolving IT infrastructure, compounded by the task of integrating data from disparate sources. The organization&amp;rsquo;s IT landscape is characterized by rapid changes, requiring flexible and adaptable monitoring solutions. Compatibility issues with legacy systems and proprietary data formats further complicate the seamless integration of monitoring data, presenting significant obstacles to maintaining a cohesive and efficient monitoring strategy. Jesús Ramos, a seasoned DevOps professional at FUNIBER, found himself at the crux of these challenges.&lt;/p></description></item><item><title>BPF vs eBPF: Key Differences Explained For DevOps &amp; SREs</title><link>https://www.netdata.cloud/academy/what-are-the-differences-between-bpf-and-ebpf-an-overview/</link><pubDate>Mon, 29 Jan 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-are-the-differences-between-bpf-and-ebpf-an-overview/</guid><description>&lt;h2 id="introduction-to-bpf--ebpf">Introduction To BPF &amp;amp; eBPF&lt;/h2>
&lt;p>For DevOps and Site Reliability Engineers (SREs), understanding and applying the right tools for monitoring and troubleshooting is crucial. Among the various technologies available, BPF (Berkeley Packet Filter) and its extended version, eBPF (extended Berkeley Packet Filter), have emerged as powerful tools. But what exactly are BPF and eBPF, and how do they differ? This article will delve into the basics of both, highlight their differences, and explore their applications in modern DevOps and SRE practices, with a special focus on observability.&lt;/p></description></item><item><title>Enhancing Digital Media Excellence with Netdata</title><link>https://www.netdata.cloud/case-studies/digital-media/shed/</link><pubDate>Mon, 29 Jan 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/case-studies/digital-media/shed/</guid><description>&lt;h2 id="navigating-network-complexity">Navigating Network Complexity&lt;/h2>
&lt;p>&lt;a href="https://shedmtl.com/">SHED&lt;/a> encountered a critical pain point in swiftly and accurately identifying bandwidth hogs within their network. This was a significant issue, as monitoring these elements was essential to pinpoint devices, applications, or users causing excessive bandwidth consumption. The lack of real-time visibility and comprehensive insights into data transfer rates and traffic patterns was a major hurdle in their ability to proactively manage and optimize network resources. This limitation directly impacted their network performance, leading to potential slowdowns or bottlenecks, affecting the efficiency of their operations. Moreover, this challenge posed risks in resource allocation and maintaining a smooth user experience. Additionally, this issue had implications for security, as unidentified bandwidth hogs could indicate security threats or unauthorized activities. Thus, SHED Montreal&amp;rsquo;s foremost challenge was the effective monitoring of file server and switch activities to swiftly identify and mitigate bandwidth hogs, crucial for ensuring optimal network performance, resource allocation, and maintaining a secure infrastructure.&lt;/p></description></item><item><title>SRE vs DevOps: Key Differences, Synergies &amp; Roles Explained</title><link>https://www.netdata.cloud/academy/sre-vs-devops-what-are-the-main-differences-between-them/</link><pubDate>Mon, 29 Jan 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/sre-vs-devops-what-are-the-main-differences-between-them/</guid><description>&lt;p>If you’re new to IT operations and software development, you&amp;rsquo;ve likely encountered the terms SRE (Site Reliability Engineering) and DevOps. Both play crucial roles in modern tech organizations, but understanding the differences between SRE and DevOps can be confusing. This article aims to clarify these concepts, explain their roles, and show how they work together to enhance software delivery and system reliability.&lt;/p>
&lt;h2 id="a-brief-history-of-devops--sre">A Brief History Of DevOps &amp;amp; SRE&lt;/h2>
&lt;p>To better understand the philosophies behind DevOps and Site Reliability Engineering (SRE), it helps to look at how each approach emerged.&lt;/p></description></item><item><title>What Is Cardinality In Databases: A Comprehensive Guide</title><link>https://www.netdata.cloud/academy/what-is-cardinality-in-databases-a-comprehensive-guide/</link><pubDate>Mon, 29 Jan 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-cardinality-in-databases-a-comprehensive-guide/</guid><description>&lt;h2 id="an-introduction-to-cardinality">An Introduction To Cardinality&lt;/h2>
&lt;p>Cardinality is a fundamental concept in databases that plays a crucial role in designing efficient databases and optimizing query performance. For those new to databases, understanding what is cardinality is essential for effective data management. This comprehensive guide will explain the meaning of cardinality, its types, and its &lt;a href="https://www.netdata.cloud/mongodb-monitoring/">impact on database performance&lt;/a>, making it accessible for both beginners and intermediate users.&lt;/p>
&lt;h2 id="what-is-cardinality-in-databases">What Is Cardinality In Databases?&lt;/h2>
&lt;p>Cardinality in databases refers to the uniqueness of data values contained in a column. It essentially measures how many distinct values exist in a column compared to the total number of rows in a table.&lt;/p></description></item><item><title>What Is Cloud Management? Ensure Security &amp; Optimize Costs</title><link>https://www.netdata.cloud/academy/what-is-cloud-management-how-to-maximize-efficiency/</link><pubDate>Mon, 29 Jan 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-cloud-management-how-to-maximize-efficiency/</guid><description>&lt;p>Cloud computing has become a cornerstone for businesses of all sizes. As more organizations migrate their workloads to the cloud, understanding cloud management and how to maximize efficiency is crucial. But what is cloud management exactly? This article will explore the concept, dive into the various aspects of cloud management, and provide practical tips on how to enhance efficiency, especially for beginners.&lt;/p>
&lt;h2 id="what-is-cloud-management">What Is Cloud Management?&lt;/h2>
&lt;p>Cloud management involves the control, orchestration, and administration of cloud computing resources and services. It encompasses a range of tasks, from deploying and &lt;a href="https://www.netdata.cloud/solutions/aws-monitoring/">monitoring applications&lt;/a> to ensuring security and compliance. Effective cloud management is essential for optimizing performance, controlling costs, and maintaining security across cloud environments.&lt;/p></description></item><item><title>What Is Continuous Profiling &amp; Why It Matters</title><link>https://www.netdata.cloud/academy/what-is-continuous-profiling-why-it-is-important-for-monitoring/</link><pubDate>Mon, 29 Jan 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/academy/what-is-continuous-profiling-why-it-is-important-for-monitoring/</guid><description>&lt;h2 id="introduction-to-profiling-a-key-tool-for-developers">Introduction to Profiling: A Key Tool for Developers&lt;/h2>
&lt;p>Profiling is like a health checkup for your software. Just as doctors use various tests to understand how your body is performing, developers use profiling to understand how their applications are running. This process helps to find out which parts of the application are working well and which parts need improvement.&lt;/p>
&lt;h2 id="what-is-profiling">What Is Profiling?&lt;/h2>
&lt;p>Profiling, in simple terms, is the process of measuring how your program uses resources like CPU time and memory. Think of it as a way to see inside your application and understand what&amp;rsquo;s going on.&lt;/p></description></item><item><title>IoT Monitoring Challenges: Key Issues &amp; How To Overcome Them</title><link>https://www.netdata.cloud/blog/iot-monitoring-challenges/</link><pubDate>Thu, 11 Jan 2024 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/iot-monitoring-challenges/</guid><description>&lt;p>With the increasing prevalence of IoT devices, which are being used in a wide range of applications, from smart homes and cities to industrial and agricultural systems, monitoring thei performance and health is extremely important. However, it’s essential to remember that monitoring IoT devices involves more than just tracking device-level data. In addition, monitoring data from the IoT platform or application layer is equally important.&lt;/p>
&lt;p>We’ll explore some of these topics in more detail and explain how Netdata can play an essential role in the monitoring of such devices, including some hints on how it can be set up for maximum performance in such scenarios.&lt;/p></description></item><item><title>Introducing Netdata’s Alerts Configuration Manager</title><link>https://www.netdata.cloud/blog/netdata-alerts-configuration-manager/</link><pubDate>Thu, 21 Dec 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-alerts-configuration-manager/</guid><description>&lt;p>Netdata introduces its latest feature, the Alerts Configuration Manager, transforming the way users configure and manage alerts in their Netdata environment. This powerful tool integrates directly into the Netdata Dashboard, offering a streamlined and intuitive interface for both novice and experienced users.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="what-is-the-alerts-configuration-manager">What is the Alerts Configuration Manager?&lt;/h2>
&lt;p>The Alerts Configuration Manager is an innovative feature available to users with Business subscriptions. It allows for the creation and customization of alerts directly from the Netdata Dashboard, employing a user-friendly UI wizard. This tool simplifies alert configuration, making it accessible and straightforward, even for those who are not deeply technical.&lt;/p></description></item><item><title>Cost Transparency: The True Cost Of Monitoring</title><link>https://www.netdata.cloud/blog/netdata-cost-transparency-unveiling-the-true-cost-of-monitoring/</link><pubDate>Fri, 01 Dec 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-cost-transparency-unveiling-the-true-cost-of-monitoring/</guid><description>&lt;p>Businesses are increasingly reliant on monitoring tools to ensure the seamless performance and reliability of their systems. However, the true cost of implementing and maintaining these tools is often obscured by hidden expenses. Our previous blog delved into the concealed costs associated with various monitoring solutions, such as Prometheus &amp;amp; Grafana (Open Source Monitoring) and commercial platforms like Datadog, Dynatrace, and NewRelic. These costs can manifest in various forms - from complex setups and maintenance to additional charges for advanced features.&lt;/p></description></item><item><title>Understanding the Netdata Methodology</title><link>https://www.netdata.cloud/blog/netdata-methodology/</link><pubDate>Sat, 04 Nov 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-methodology/</guid><description>&lt;p>In the dynamic landscape of modern infrastructure and multi cloud environments, observing and understanding system performance requires a new breed of tools—ones that keep pace with the &amp;rsquo;living&amp;rsquo; nature of modern infrastructure. This is the inflection point at which Netdata steps in, and aims to bring a fresh perspective to monitoring.&lt;/p>
&lt;h2 id="what-does-netdata-do-differently">What does Netdata do differently&lt;/h2>
&lt;p>There are key “cultural” faults that we believe hold the &lt;a href="https://www.netdata.cloud/blog/monitoring-vs-observability">observability&lt;/a> industry back. Let’s take a look through the prism of these faults and understand what Netdata is doing differently to address them.&lt;/p></description></item><item><title>Netdata Best Practices</title><link>https://www.netdata.cloud/blog/netdata-best-practices/</link><pubDate>Fri, 03 Nov 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-best-practices/</guid><description>&lt;p>Effective &lt;strong>system monitoring&lt;/strong> is non-negotiable in today&amp;rsquo;s complex IT environments. Netdata offers real-time performance and health monitoring with precision and granularity. But the key to harnessing its full potential lies in the optimization of your setup. Let’s ensure you are not just collecting data, but doing it in the most optimal way while gaining actionable insights from it.&lt;/p>
&lt;p>The starting point for optimization is a robust setup. Netdata is engineered for minimal footprint and can run on a wide range of hardware—from IoT devices to powerful servers. Time for a deep dive into each of these key areas and what the best practices you should follow, if you are serious about monitoring and optimizing your Netdata monitoring setup:&lt;/p></description></item><item><title>System Operators: The Role Of SysOps In IT Infrastructure</title><link>https://www.netdata.cloud/blog/system-operators-unlock-log-management-mastery-with-systemd-journal-and-netdata/</link><pubDate>Fri, 03 Nov 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/system-operators-unlock-log-management-mastery-with-systemd-journal-and-netdata/</guid><description>&lt;p>System operators know the drill: as the complexity of systems scales, so does the deluge of logs. Traditionally, taming this relentless tide demands a concoction of costly tools and laborious configurations—until now. The dynamic duo of &lt;code>systemd-journal&lt;/code> and Netdata is revolutionizing log management, turning what was once a Herculean task into a streamlined, powerful, and surprisingly straightforward process.&lt;/p>
&lt;h2 id="efficient-handling-of-volume-and-velocity">Efficient Handling of Volume and Velocity&lt;/h2>
&lt;p>&lt;code>systemd-journal&lt;/code> is built to manage the deluge of data that systems generate, &lt;strong>without buckling under the speed and volume of incoming logs&lt;/strong>. It captures logs at the source, facilitating direct and immediate processing. By utilizing the journal&amp;rsquo;s native mechanism to send logs to a central server, system operators can bypass the complexities of traditional &lt;strong>log centralization methods&lt;/strong>.&lt;/p></description></item><item><title>Upcoming Changes To Netdata Cloud Plans</title><link>https://www.netdata.cloud/blog/netdata-plan-changes/</link><pubDate>Thu, 02 Nov 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-plan-changes/</guid><description>&lt;blockquote>
&lt;p>&lt;strong>UPDATE&lt;/strong>: On the &lt;strong>2023-11-08&lt;/strong> Node and Dashboard limits will be applied on the Netdata Cloud &lt;strong>Community plan&lt;/strong>, while all current features of the Community plan will remain the same.&lt;/p>
&lt;/blockquote>
&lt;h2 id="what-you-need-to-know">What you need to know?&lt;/h2>
&lt;ul>
&lt;li>
&lt;p>On Netdata Cloud Free Community plan, the number of &lt;strong>active nodes&lt;/strong> that can be concurrently visualized on the Netdata dashboards, as well as the number of &lt;strong>active custom dashboards&lt;/strong> for accounts created after 2023-11-07 will be subject to limits.&lt;/p></description></item><item><title>Netdata vs Prometheus</title><link>https://www.netdata.cloud/blog/netdata-vs-prometheus-performance-analysis/</link><pubDate>Sat, 28 Oct 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-vs-prometheus-performance-analysis/</guid><description>&lt;p>In an era dominated by data-driven decision making, monitoring tools play an indispensable role in ensuring that our systems run efficiently and without interruption. When considering tools like &lt;strong>Netdata and Prometheus&lt;/strong>, performance isn&amp;rsquo;t just a number; it&amp;rsquo;s about empowering users with &lt;strong>real-time insights&lt;/strong> and enabling them to act with agility.&lt;/p>
&lt;p>There&amp;rsquo;s a genuine need in the community for tools that are not only comprehensive in their offerings but also &lt;strong>swift and scalable&lt;/strong>. This desire stems from our evolving digital landscape, where the ability to swiftly detect, diagnose, and &lt;a href="https://www.netdata.cloud/kubernetes-monitoring/">rectify anomalies&lt;/a> has direct implications on user experiences and business outcomes. Especially as infrastructure grows in complexity and scale, there&amp;rsquo;s an increasing demand for &lt;strong>monitoring tools&lt;/strong> to keep up and provide clear, timely insights.&lt;/p></description></item><item><title>Discover The New Netdata!</title><link>https://www.netdata.cloud/blog/discover-the-new-netdata/</link><pubDate>Fri, 27 Oct 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/discover-the-new-netdata/</guid><description>&lt;p>Missed the last &lt;strong>Netdata&lt;/strong> updates? Here is what is new:&lt;/p>
&lt;h2 id="explore-your-systemd-journal-logs-with-netdata">Explore your systemd-journal logs with Netdata&lt;/h2>
&lt;p>&lt;img src="https://github.com/netdata/blog/assets/139226121/7d2779c9-0efb-4491-8fe3-aedce1dc72fb" alt="systemd-journal-logs">&lt;/p>
&lt;p>Netdata &lt;a href="https://learn.netdata.cloud/docs/logs/systemd-journal/?utm_source=IL&amp;amp;utm_medium=internallinking&amp;amp;utm_campaign=new_netada">got a &lt;code>systemd&lt;/code>-journal logs explorer&lt;/a> to analyze your &lt;code>systemd&lt;/code>-journal logs, directly on their sources. By just installing &lt;strong>Netdata&lt;/strong> on any systemd based system, Netdata automatically finds all the &lt;strong>journal sources&lt;/strong> and presents a powerful dashboard to explore, search, filter and analyze your &lt;strong>logs&lt;/strong>. It works on both individual servers and journal centralization servers.&lt;/p></description></item><item><title>Improve Your Security With systemd-journal &amp; Netdata</title><link>https://www.netdata.cloud/blog/improve-your-security-with-systemd-and-netdata/</link><pubDate>Tue, 24 Oct 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/improve-your-security-with-systemd-and-netdata/</guid><description>&lt;p>&lt;strong>&lt;code>systemd&lt;/code> journals&lt;/strong> play a crucial role in the Linux system ecosystem, and understanding the importance of the logs contained within is essential for both system administrators and developers.&lt;/p>
&lt;p>For those unfamiliar, &lt;code>systemd&lt;/code> is an init system employed by Linux distributions, initiates the user space and oversees all ensuing processes. One of its key components, &lt;code>systemd&lt;/code> journal, assumes a central role in logging system activities and messages, delivering a host of benefits to both system administrators, developers and cyber security engineers. The &lt;code>systemd&lt;/code> journal functions as a logging system that gathers, archives, and oversees log messages and event data originating from a diverse array of system components, encompassing the kernel, system services, applications, and user activities.&lt;/p></description></item><item><title>Monitoring vs Observability: Key Differences</title><link>https://www.netdata.cloud/blog/monitoring-vs-observability/</link><pubDate>Tue, 24 Oct 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/monitoring-vs-observability/</guid><description>&lt;p>As systems increasingly shift towards distributed architectures to deliver application services, the roles of monitoring and observability have never been more crucial. Monitoring delivers the situational awareness you need to detect issues, while &lt;a href="https://www.netdata.cloud/academy/what-is-observability/">observability goes a step further&lt;/a>, offering the analytical depth to understand the root cause of those issues.&lt;/p>
&lt;p>Understanding the nuanced differences between monitoring and observability is crucial for anyone responsible for system health and performance. In dissecting these methodologies, we&amp;rsquo;ll explore their unique strengths, dive into practical applications, and illuminate how to strategically employ each to enhance operational outcomes.&lt;/p></description></item><item><title>Exploring systemd journal logs with Netdata</title><link>https://www.netdata.cloud/blog/exploring-systemd-journal-logs/</link><pubDate>Thu, 12 Oct 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/exploring-systemd-journal-logs/</guid><description>&lt;p>Today, we released our &lt;code>systemd&lt;/code> &lt;strong>journal plugin for Netdata&lt;/strong>, allowing you to explore, view, search, filter and analyze &lt;code>systemd&lt;/code> journal logs.&lt;/p>
&lt;p>Like most things about Netdata, this is a &lt;strong>zero-configuration plugin&lt;/strong>. You don’t have to do anything apart from &lt;strong>installing Netdata&lt;/strong> on your systems.This is key design direction for Netdata, since we want Netdata to be able to help even if you install it mid-crisis, while you have an incident at hand.&lt;/p></description></item><item><title>systemd journal logs</title><link>https://www.netdata.cloud/blog/systemd-journal-logs/</link><pubDate>Mon, 09 Oct 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/systemd-journal-logs/</guid><description>&lt;p>&lt;em>“Why bother with it? I let it run in the background and focus on more important DevOps work.”&lt;/em>
— a random DevOps Engineer at Reddit r/devops&lt;/p>
&lt;p>In an era where technology is evolving at breakneck speeds, it&amp;rsquo;s easy to overlook the tools that are right under our noses. One such underutilized powerhouse is the &lt;strong>&lt;code>systemd&lt;/code> journal&lt;/strong>. For many, it&amp;rsquo;s a mere tool to check the status of systemd service units or to tail the most recent events (journalctl -f). Others who do mainly container work, ignore even its existence.&lt;/p></description></item><item><title>Netdata Cloud On Prem</title><link>https://www.netdata.cloud/blog/netdata-cloud-on-prem/</link><pubDate>Tue, 26 Sep 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-cloud-on-prem/</guid><description>&lt;p>We at &lt;strong>Netdata&lt;/strong> understand that &lt;a href="https://blog.netdata.cloud/future-of-infrastructure-monitoring/">infrastructure monitoring&lt;/a> can be a complex maze—high costs, specialized skill sets, scalability, data silos, and more. That&amp;rsquo;s why we have always aimed to streamline and modernize this critical operation. Today, we&amp;rsquo;re thrilled to announce the launch of &lt;a href="https://www.netdata.cloud/contact-us/?subject=on-prem">Netdata Cloud On Prem&lt;/a>, a ground-breaking solution designed for robust &lt;strong>on-prem infrastructure monitoring&lt;/strong> - it comes with all the &lt;strong>Netdata Cloud&lt;/strong> features you love but fully on prem.&lt;/p>
&lt;h2 id="netdata-cloud-on-prem">&lt;strong>Netdata Cloud On-Prem&lt;/strong>&lt;/h2>
&lt;p>While &lt;a href="https://www.netdata.cloud/">Netdata Cloud&lt;/a> never stores any of your metric data on the cloud and just streams it ephemerally while you view a chart, the demand for on premise infrastructure monitoring has never been more pressing. Many large enterprises, governmental organizations, research institutes and critical infrastructures require a level of data privacy, security, and customization that only an on-prem solution can offer.&lt;/p></description></item><item><title>Netdata QoS Classes monitoring</title><link>https://www.netdata.cloud/blog/netdata-qos-monitoring/</link><pubDate>Tue, 26 Sep 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-qos-monitoring/</guid><description>&lt;p>Netdata monitors &lt;code>tc&lt;/code> QoS classes for all interfaces.&lt;/p>
&lt;p>If you also use &lt;a href="http://firehol.org/tutorial/fireqos-new-user/">FireQOS&lt;/a> it will collect interface and class names.&lt;/p>
&lt;p>There is a &lt;a href="https://raw.githubusercontent.com/netdata/netdata/master/collectors/tc.plugin/tc-qos-helper.sh.in">shell helper&lt;/a> for this (all parsing is done by the plugin in &lt;code>C&lt;/code> code - this shell script is just a configuration for the command to run to get &lt;code>tc&lt;/code> output).&lt;/p>
&lt;p>The source of the tc plugin is &lt;a href="https://raw.githubusercontent.com/netdata/netdata/master/collectors/tc.plugin/plugin_tc.c">here&lt;/a>. It is somewhat complex, because a state machine was needed to keep track of all the &lt;code>tc&lt;/code> classes, including the pseudo classes tc dynamically creates.&lt;/p></description></item><item><title>Netdata, Prometheus, Grafana Stack</title><link>https://www.netdata.cloud/blog/netdata-prometheus-grafana-stack/</link><pubDate>Tue, 26 Sep 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-prometheus-grafana-stack/</guid><description>&lt;p>In this blog, we will walk you through the basics of getting Netdata, Prometheus and Grafana all working together and
&lt;a href="https://www.netdata.cloud/blog/web-servers-and-their-performance/">monitoring your application servers&lt;/a>. This article will be using docker on your local workstation. We will be working
with docker in an ad-hoc way, launching containers that run &lt;code>/bin/bash&lt;/code> and attaching a TTY to them. We use docker here
in a purely academic fashion and do not condone running Netdata in a container. We pick this method so individuals
without &lt;a href="https://www.netdata.cloud/academy/what-is-cloud-management-how-to-maximize-efficiency/">cloud accounts&lt;/a> or access to VMs can try this out and for it&amp;rsquo;s speed of deployment.&lt;/p></description></item><item><title>Process Monitoring vs Console Tools: A Comparison</title><link>https://www.netdata.cloud/blog/netdata-processes-monitoring-comparison-with-console-tools/</link><pubDate>Tue, 26 Sep 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-processes-monitoring-comparison-with-console-tools/</guid><description>&lt;p>Netdata reads &lt;code>/proc/&amp;lt;pid&amp;gt;/stat&lt;/code> for all processes, once per second and extracts &lt;code>utime&lt;/code> and
&lt;code>stime&lt;/code> (user and system cpu utilization), much like all the console tools do.&lt;/p>
&lt;p>But it also extracts &lt;code>cutime&lt;/code> and &lt;code>cstime&lt;/code> that account the user and system time of the exit children of each process.
By keeping a map in memory of the whole process tree, it is capable of assigning the right time to every process, taking
into account all its exited children.&lt;/p></description></item><item><title>Our first ML based anomaly alert</title><link>https://www.netdata.cloud/blog/our-first-ml-based-anomaly-alert/</link><pubDate>Wed, 13 Sep 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/our-first-ml-based-anomaly-alert/</guid><description>&lt;p>Over the last few years we have slowly and methodically been building out the &lt;a href="https://learn.netdata.cloud/docs/ml-and-troubleshooting/">ML based capabilities&lt;/a> of the Netdata agent, dogfooding and iterating as we go. To date, these features have mostly been somewhat reactive and tools to aid once you are already troubleshooting.&lt;/p>
&lt;p>Now we feel we are ready to take a first gentle step into some more proactive use cases, starting with a &lt;a href="https://github.com/netdata/netdata/pull/14687">simple node level anomaly rate alert&lt;/a>.&lt;/p></description></item><item><title>Anomaly Rate By Type</title><link>https://www.netdata.cloud/blog/anomaly-rate-by-type/</link><pubDate>Wed, 30 Aug 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/anomaly-rate-by-type/</guid><description>&lt;p>We have &lt;a href="https://github.com/netdata/netdata/pull/15856">recently added&lt;/a> a more detailed anomaly rate chart to Netdata that breaks out the overall &lt;a href="https://learn.netdata.cloud/docs/ml-and-troubleshooting/machine-learning-ml-powered-anomaly-detection#node-anomaly-rate">node anomaly rate&lt;/a> by type, this lets you more easily see what parts of your infrastructure might be experiencing an uptick in anomalies when you see the overall node anomaly rate increase.&lt;/p>
&lt;h2 id="what-is-type">What is &lt;code>type&lt;/code>?&lt;/h2>
&lt;p>&lt;code>type&lt;/code> is generally the prefix of the chart id in Netdata and controls where charts live within the menu on the overview page, for example the &lt;code>mem.available&lt;/code> chart has a type of &lt;code>mem&lt;/code> which in part controls why it lives under the &amp;ldquo;Memory&amp;rdquo; section of the menu.&lt;/p></description></item><item><title>Release 1.41: Brand-New UI For Agents &amp; Parents</title><link>https://www.netdata.cloud/blog/netdata-version-1.41/</link><pubDate>Thu, 20 Jul 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-version-1.41/</guid><description>&lt;p>Netdata Agents and Parents now have a new UI!&lt;/p>
&lt;p>Checkout the release meetup video or read on to learn more about the new UI and other features in this release.&lt;/p>

 &lt;div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
 &lt;iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/WCUn4-LneCw?autoplay=0&amp;amp;controls=1&amp;amp;end=0&amp;amp;loop=0&amp;amp;mute=0&amp;amp;start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video">&lt;/iframe>
 &lt;/div>

&lt;ul>
&lt;li>&lt;a href="#v1410-netdata-open-source-growth">Netdata Growth &lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1410-release-highlights">Release Highlights&lt;/a>
&lt;ul>
&lt;li>&lt;strong>&lt;a href="#v1410-one-dashboard">New Agent Dashboard!&lt;/a>&lt;/strong>&lt;/li>
&lt;li>&lt;a href="#v1410-netdata-assistant">Netdata Assistant&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1410-netdata-freeipmi">New FreeIPMI collector for monitoring enterprise hardware&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1410-netdata-apps">Netdata Detects FDs Leaking&lt;/a>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;a href="#v1410-acknowledgements">Acknowledgements &lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1410-contributions">Contributions&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1410-contributions-collectors">Collectors&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1410-contributions-documentation">Documentation &lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1410-contributions-packaging">Packaging/Installation&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1410-contributions-health">Health&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1410-contributions-exporting">Exporting&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1410-contributions-other">Other Notable Changes&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1410-deprecation-notice">Deprecation notice&lt;/a>
&lt;ul>
&lt;li>&lt;a href="#v1410-deprected-in-this-release">Deprecated in this release&lt;/a>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;a href="#v1410-netdata-release-meetup">Netdata Release Meetup&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1410-support-options">Support options&lt;/a>&lt;/li>
&lt;/ul>
&lt;p>Steady to our schedule, this is another great Netdata release!&lt;/p></description></item><item><title>Netdata Assistant: Your AI-Powered Troubleshooting Sidekick</title><link>https://www.netdata.cloud/blog/netdata-assistant/</link><pubDate>Fri, 14 Jul 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-assistant/</guid><description>&lt;p>Hey there! We&amp;rsquo;re excited to share a new troubleshooting feature we have added to Netdata, the Netdata Assistant. We&amp;rsquo;ve built this tool to help you troubleshoot more effectively and with less stress. Let&amp;rsquo;s dive in.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="whats-the-netdata-assistant">What&amp;rsquo;s the Netdata Assistant?&lt;/h2>
&lt;p>The Netdata Assistant is an AI tool that uses large language models and our community&amp;rsquo;s knowledge to guide you during troubleshooting.&lt;/p>
&lt;p>Here&amp;rsquo;s a scenario. It&amp;rsquo;s 3 am and you get an alert. Instead of scrambling to Google what&amp;rsquo;s going on, you can just click on the assistant button. The Netdata Assistant will give you the lowdown on the alert, why it&amp;rsquo;s happening, and why you should care. It&amp;rsquo;ll also guide you on how to troubleshoot it and even offer some handy web links for more info, if you&amp;rsquo;re interested.&lt;/p></description></item><item><title>Hidden Costs Of Monitoring: Uncovering Expenses &amp; Solutions</title><link>https://www.netdata.cloud/blog/hidden-costs-of-monitoring/</link><pubDate>Fri, 07 Jul 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/hidden-costs-of-monitoring/</guid><description>&lt;p>When it comes to monitoring IT infrastructure, the &lt;a href="https://www.netdata.cloud/pricing/">costs you see on the price tag&lt;/a> of the tool are often just the tip of the iceberg. Below the waterline, a mass of hidden costs can lurk, which can significantly affect the total cost of ownership.&lt;/p>
&lt;!--truncate-->
&lt;p>In this blogpost we will cover the analysis of two traditional monitoring domains, &lt;a href="https://www.netdata.cloud/open-source/">Open Source observability&lt;/a> and Commercial Centralized observability solutions, focusing the direct and indirect impacts when implementing these solution. In summary:&lt;/p></description></item><item><title>Netdata &amp; Ansible example: ML demo room</title><link>https://www.netdata.cloud/blog/ml-demo-ansible-configuration-management/</link><pubDate>Fri, 07 Jul 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/ml-demo-ansible-configuration-management/</guid><description>&lt;p>We are always trying to lower the barrier to entry when it comes to monitoring and observability and one place we have consistently witnessed some pain from users is around adopting and approaching &lt;a href="https://www.atlassian.com/microservices/microservices-architecture/configuration-management">configuration management&lt;/a> tools and practices as your infrastructure grows and becomes more complex.&lt;/p>
&lt;p>To that end, we have begun recently publishing our own &lt;a href="https://github.com/netdata/community/tree/main/configuration-management/ansible-ml-demo">little example ansible project&lt;/a> used to maintain and manage the servers used in our public &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/rooms/machine-learning/overview">Machine Learning Demo room&lt;/a>.&lt;/p></description></item><item><title>Netdata Parents (Streaming and Replication)</title><link>https://www.netdata.cloud/blog/netdata-parents-streaming-replication/</link><pubDate>Fri, 30 Jun 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-parents-streaming-replication/</guid><description>&lt;h2 id="what-are-they-and-why-do-we-need-them">What are they and why do we need them?&lt;/h2>
&lt;p>A “Parent” is a Netdata Agent, like the ones we install on all our systems, but is configured as a central node that receives, stores and processes metrics data from other Netdata “Child” nodes in our infrastructure.&lt;/p>
&lt;p>Netdata Parents are flexible. You can have one big active-active cluster of Netdata Parents, or you can spread a lot of independent Parents across the infrastructure.&lt;/p></description></item><item><title>Release 1.40: Summary Tiles, Silencing &amp; ML Tweaks</title><link>https://www.netdata.cloud/blog/netdata-version-1.40/</link><pubDate>Wed, 14 Jun 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-version-1.40/</guid><description>&lt;p>Another release of the Netdata Monitoring solution is here!&lt;/p>

 &lt;div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
 &lt;iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/2VkWIZB8S30?autoplay=0&amp;amp;controls=1&amp;amp;end=0&amp;amp;loop=0&amp;amp;mute=0&amp;amp;start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video">&lt;/iframe>
 &lt;/div>

&lt;!--truncate-->
&lt;ul>
&lt;li>&lt;a href="#v1400-netdata-open-source-growth">Netdata Growth&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1400-release-highlights">Release Highlights&lt;/a>
&lt;ul>
&lt;li>&lt;strong>&lt;a href="#v1400-visualization-summary-dashboards">Dashboard Sections&amp;rsquo; Summary Tiles&lt;/a>&lt;/strong>&lt;br/>
Added summary tiles to most sections of the fully-automated dashboards, to provide an instant view of the most important metrics for each section.&lt;/li>
&lt;li>&lt;strong>&lt;a href="#v1400-alert-notification-silencing">Silencing of Cloud Alert Notifications&lt;/a>&lt;/strong>&lt;br/>
Maintenance window coming up? Active issue being checked? Use the Alert notification silencing engine to mute your notifications.&lt;/li>
&lt;li>&lt;strong>&lt;a href="#v1400-ml-extended-training">Machine Learning - Extended Training to 24 Hours&lt;/a>&lt;/strong>&lt;br/>
Netdata now trains multiple models per metric, to learn the behavior of each metric for the last 24 hours. Trained models are persisted on disk and are loaded back on Netdata restart.&lt;/li>
&lt;li>&lt;strong>&lt;a href="#v1400-streaming">Rewritten SSL Support for the Agent&lt;/a>&lt;/strong>&lt;br/>
Netdata Agent now features a new SSL layer that allows it to reliably use SSL on all its features, including the API and Streaming.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;a href="#v1400-alerts">Alerts and Notifications&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1400-visualizations">Visualizations / Charts and Dashboards&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1400-packaging-split">Preliminary steps to split native packages&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1400-acknowledgements">Acknowledgements&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1400-contributions">Contributions&lt;/a>
&lt;ul>
&lt;li>&lt;a href="#v1400-contributions-collectors">Collectors&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1400-contributions-documentation">Documentation&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1400-contributions-packaging">Packaging / Installation&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1400-contributions-streaming">Streaming&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1400-contributions-health">Health&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1400-contributions-exporting">Exporting&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1400-contributions-ml">ML&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1400-contributions-other">Other notable changes&lt;/a>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;a href="#v1400-deprecation-notice">Deprecation notice&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1400-cloud-recommended-version">Cloud recommended version&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1400-release-meetup">Release meetup&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1400-support-options">Support options&lt;/a>&lt;/li>
&lt;li>&lt;a href="#v1400-running-survey">Running survey&lt;/a>&lt;/li>
&lt;/ul>
&lt;h2 id="netdata-growth-a-idv1400-netdata-open-source-growtha">Netdata Growth &lt;a id="v1400-netdata-open-source-growth">&lt;/a>&lt;/h2>
&lt;p>🚀 Our community growth is increasing steadily. ❤️ Thank you! Your love and acceptance give us the energy and passion to work harder to simplify and make monitoring easier, more effective and more fun to use.&lt;/p></description></item><item><title>How Netdata's ML-based Anomaly Detection Works</title><link>https://www.netdata.cloud/blog/how-netdatas-ml-based-anomaly-detection-works/</link><pubDate>Tue, 23 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/how-netdatas-ml-based-anomaly-detection-works/</guid><description>&lt;p>&lt;img src="../2023-05-23-how-netdatas-ml-based-anomaly-detection-works/img/img.png" alt="title image">&lt;/p>
&lt;p>How does Netdata&amp;rsquo;s &lt;a href="https://learn.netdata.cloud/docs/troubleshooting-and-machine-learning/machine-learning-ml-powered-anomaly-detection">machine learning (ML) based anomaly detection&lt;/a> actually work? Read on to find out!&lt;/p>
&lt;!--truncate-->
&lt;h2 id="design-considerations">Design considerations&lt;/h2>
&lt;p>Lets first start with some of the key design considerations and principles of Netdata&amp;rsquo;s anomaly detection (&lt;em>and some comments in parenthesis along the way&lt;/em>):&lt;/p>
&lt;ol>
&lt;li>We don&amp;rsquo;t have any labels or examples of previous anomalies. This means we are in an &lt;a href="https://en.wikipedia.org/wiki/Unsupervised_learning">unsupervised setting&lt;/a> (&lt;em>best we can try to do is learn what &amp;ldquo;normal&amp;rdquo; data looks like assuming the collected data is &amp;ldquo;mostly&amp;rdquo; normal&lt;/em>).&lt;/li>
&lt;li>Needs to be lightweight and run on the agent (&lt;em>or a parent&lt;/em>).
&lt;ul>
&lt;li>Need to be very careful of impact on CPU overhead when training and scoring (&lt;em>lots of cheap models are better than a few expensive and heavy ones&lt;/em>).&lt;/li>
&lt;li>Models themselves need to be small so as to not drastically increase the agents memory footprint (&lt;em>model objects need to be small for storage&lt;/em>).&lt;/li>
&lt;li>This has implications for the ML formulation (&lt;em>sorry - no deep learning models yet) and its implementation (we need to be surgical and optimized&lt;/em>).&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Needs to scale for thousands of metrics and score in realtime every second as metrics are collected.
&lt;ul>
&lt;li>Typical Netdata nodes have thousands of metrics and we want to be able to score every metric every second with minimal latency overhead (&lt;em>we need to use sensible approaches to training like spreading the training cost over a wide training window&lt;/em>).&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Needs to be able to handle a wide variety of metrics.
&lt;ul>
&lt;li>There is no single perfect model or approach for all types of metrics so we need a good all rounder that can work well enough across any and all different types of time seres metrics (&lt;em>for any given metric of course you could handcraft a better model but thats not feasible here, we need something like a &amp;ldquo;weak learners&amp;rdquo; approach of lots of generally useful models adding up to &amp;ldquo;more than the sum of their parts&amp;rdquo;&lt;/em>).&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>Needs to be written in C or C++ as that is the language of the Netdata agent (&lt;em>We are using &lt;a href="https://github.com/davisking/dlib">dlib&lt;/a> for the current implementation&lt;/em>).&lt;/li>
&lt;li>We Need to be very careful about taking big or complex dependencies if using third party libraries.
&lt;ul>
&lt;li>We want to be able to easily build and deploy Netdata on any Linux system without having to worry about installing or managing complex dependencies (&lt;em>we need to be careful of more complex algorithms that would have larger dependencies and potentially limit where Netdata can run&lt;/em>).&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ol>
&lt;p>The above considerations are important and useful to keep in mind as we explore the system in more detail.&lt;/p></description></item><item><title>Revolutionizing Ops Centers With Real-Time Monitoring</title><link>https://www.netdata.cloud/blog/revolutionizing-operations-centers-real-time-monitoring-solution/</link><pubDate>Fri, 19 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/revolutionizing-operations-centers-real-time-monitoring-solution/</guid><description>&lt;p>&lt;img src="../2023-05-19-revolutionizing-operations-centers-real-time-monitoring-solution/img/stacked-netdata.png" alt="stacked-netdata">&lt;/p>
&lt;p>In today&amp;rsquo;s fast-paced digital landscape, 24-hour operations centers play a crucial role in managing and monitoring large-scale infrastructures. These centers must be equipped with an effective monitoring solution that addresses their unique needs, enabling them to respond quickly to incidents and maintain optimal system performance. Netdata, a comprehensive monitoring solution, has been designed to meet these critical requirements with its advanced capabilities and recent enhancements.&lt;/p>
&lt;p>In this article, we will explore how Netdata&amp;rsquo;s powerful features can transform the way 24-hour operations centers monitor and manage their complex environments, leading to improved incident detection, faster troubleshooting, and better overall system performance.&lt;/p></description></item><item><title>The Future Of Infrastructure Monitoring: Scalability &amp; AI</title><link>https://www.netdata.cloud/blog/future-of-infrastructure-monitoring/</link><pubDate>Fri, 19 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/future-of-infrastructure-monitoring/</guid><description>&lt;p>In this blog post, we will explore the importance of scalability, automation, and AI in the evolving landscape of &lt;a href="https://www.netdata.cloud/blog/what-is-infrastructure-monitoring/">infrastructure monitoring&lt;/a>. We will examine how Netdata&amp;rsquo;s innovative solution aligns with these emerging trends, and how it can empower organizations to effectively manage their modern IT infrastructure.&lt;/p>
&lt;!--truncate-->
&lt;p>In today&amp;rsquo;s increasingly complex IT landscape, the need for efficient and reliable infrastructure monitoring has never been more critical. With the &lt;a href="https://medium.com/capital-one-tech/the-microservices-paradox-e55d5af2fda5">proliferation of microservices&lt;/a>, distributed systems, and cloud-native applications, managing and monitoring the performance of these rapidly evolving environments has become a significant challenge. As a result, infrastructure monitoring solutions must adapt to keep pace with these changes and deliver the insights necessary to maintain optimal performance.&lt;/p></description></item><item><title>Monitoring Multi-Cloud &amp; Hybrid-Cloud Infrastructures</title><link>https://www.netdata.cloud/blog/monitoring-multi-cloud-hybrid-cloud/</link><pubDate>Tue, 16 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/monitoring-multi-cloud-hybrid-cloud/</guid><description>&lt;p>The advent of multi-cloud and hybrid-cloud architectures has created new opportunities for organizations to leverage best-in-class features from various cloud service providers. However, these complex environments present their own unique challenges, especially when it comes to monitoring and managing performance.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="visibility">Visibility&lt;/h2>
&lt;p>The visibility challenge in multi-cloud and hybrid-cloud environments often stems from the use of disparate monitoring tools that are native to each cloud provider. While these native tools (like Amazon CloudWatch, Google Cloud Monitoring, and Azure Monitor) are excellent within their respective ecosystems, they don&amp;rsquo;t necessarily play well together when it comes to consolidating data and providing a comprehensive, unified view of your entire infrastructure.&lt;/p></description></item><item><title>Cloud Optimization: Cost, Performance &amp; Resource Strategies</title><link>https://www.netdata.cloud/blog/mastering-cloud-optimization/</link><pubDate>Sun, 14 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/mastering-cloud-optimization/</guid><description>&lt;p>Cloud optimization is the ongoing process of analyzing, configuring, and refining cloud environments to improve performance, reduce costs, and align resource usage with business needs. As cloud adoption grows, organizations must move beyond cost-cutting alone and treat optimization as a strategic practice.&lt;/p>
&lt;h2 id="cloud-optimization-strategies-to-achieve-business-goals">Cloud Optimization Strategies To Achieve Business Goals&lt;/h2>
&lt;p>Cloud optimization strategies generally focus on cost control, performance enhancement, and efficient resource utilization. These strategies range from selecting the right cloud service model (IaaS, PaaS, or SaaS), right-sizing your resources, adopting a &lt;a href="https://www.netdata.cloud/academy/what-is-cloud-management-how-to-maximize-efficiency/">multi-cloud approach&lt;/a>, automating processes, and investing in robust monitoring tools that can reliably reveal resources utilization and help you ensure that services are tailored to meet business objectives.&lt;/p></description></item><item><title>Migrating To Cloud: Key Challenges &amp; Best Practices</title><link>https://www.netdata.cloud/blog/migrating-to-cloud-key-challenges-best-practices/</link><pubDate>Sun, 14 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/migrating-to-cloud-key-challenges-best-practices/</guid><description>&lt;p>Embarking on a cloud migration journey? Grasp the obstacles and arm yourself with best practices for a smooth transition. Success lies in understanding, planning, and adapting.&lt;/p>
&lt;!--truncate-->
&lt;p>As we continue to advance further into the 21st century, businesses of all sizes are finding themselves in the midst of a digital revolution. At the heart of this transformation lies cloud migration, a process that has become a critical strategic decision for organizations aiming to remain competitive, innovative, and responsive to fluctuating market dynamics.&lt;/p></description></item><item><title>Transform Monitoring With A Machine Learning Approach</title><link>https://www.netdata.cloud/blog/transform-monitoring-ml-first-approach/</link><pubDate>Thu, 11 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/transform-monitoring-ml-first-approach/</guid><description>&lt;p>Unlocking the full potential of monitoring through ML integration, anomaly detection, and innovative scoring engines.&lt;/p>
&lt;!--truncate-->
&lt;p>Machine Learning has been making waves in various industries, but its adoption in the monitoring and observability space has been slower than expected. Many “ML” features remain gimmicky and do not provide actual real world value to users that encourages their further use.&lt;/p>
&lt;p>At Netdata, we firmly believe that ML is crucial for monitoring, and we&amp;rsquo;ve taken an ML-first approach to provide users with powerful tools and insights. In this blog post, we&amp;rsquo;ll discuss the reasons behind our belief in ML, how we&amp;rsquo;ve integrated ML into our charts and visualizations, our query engines, the scoring engine we&amp;rsquo;ve built, and how these innovations enable metrics correlations and anomaly advisor.&lt;/p></description></item><item><title>The Future of Monitoring is Automated and Opinionated</title><link>https://www.netdata.cloud/blog/the-future-of-monitoring-is-automated-and-opinionated/</link><pubDate>Tue, 09 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/the-future-of-monitoring-is-automated-and-opinionated/</guid><description>&lt;p>So, you think you monitor your infra?&lt;/p>
&lt;!-- truncate -->
&lt;p>As humanity increasingly relies on technology, &lt;a href="https://www.netdata.cloud/blog/future-of-infrastructure-monitoring/">the need for reliable and efficient infrastructure monitoring solutions has never been greater&lt;/a>.&lt;/p>
&lt;p>However, most businesses don&amp;rsquo;t take this seriously. They make poor choices that soon trap their best talent, the people who should be propelling them ahead of their competition.&lt;/p>
&lt;p>Consider this: most of the world believes that each company needs to dedicate time, talent, and money to configure and set up the monitoring of their web servers and database servers from scratch!&lt;/p></description></item><item><title>Release 1.39.0: A new era for monitoring charts.</title><link>https://www.netdata.cloud/blog/netdata-version-1.39/</link><pubDate>Mon, 08 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-version-1.39/</guid><description>&lt;p>Another release of the Netdata Monitoring solution is here!&lt;/p>

 &lt;div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
 &lt;iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/dU4GJjpeb3I?autoplay=0&amp;amp;controls=1&amp;amp;end=0&amp;amp;loop=0&amp;amp;mute=0&amp;amp;start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video">&lt;/iframe>
 &lt;/div>

&lt;ul>
&lt;li>&lt;a href="#v1390-netdata-open-source-growth">Netdata open-source growth&lt;/a>&lt;/li>
&lt;li>&lt;a href="#release-highlights">Release highlights&lt;/a>
&lt;ul>
&lt;li>&lt;strong>&lt;a href="#netdata-charts-v30">Netdata Charts v3.0&lt;/a>&lt;/strong>
A new era for monitoring charts. Powerful, fast, easy to use. Instantly understand the dataset behind any chart. Slice, dice, filter and pivot the data in any way possible!&lt;/li>
&lt;li>&lt;strong>&lt;a href="#windows-support">Windows support&lt;/a>&lt;/strong>
Windows hosts are now first-class citizens. You can now enjoy out-of-the-box monitoring of over 200 metrics from your Windows systems and the services that run on them.&lt;/li>
&lt;li>&lt;a href="#virtual-nodes-and-custom-labels">Virtual nodes and custom labels&lt;/a>
You now have access to more monitoring superpowers for managing medium to large infrastructures. With custom labels and virtual hosts, you can easily organize your infrastructure and ensure that troubleshooting is more efficient.&lt;/li>
&lt;li>&lt;a href="#major-upcoming-changes">Major upcoming changes&lt;/a>
Separate packages for data collection plugins, mandatory &lt;code>zlib&lt;/code>, no upgrades of existing installs from versions prior to v1.11.&lt;/li>
&lt;li>&lt;a href="#bar-charts-for-functions">Bar charts for functions&lt;/a>&lt;/li>
&lt;li>&lt;a href="#opsgenie-notifications-for-business-plan-users">Opsgenie notifications for Business Plan users&lt;/a>
Business plan users can now seamlessly integrate Netdata with their Atlassian Opsgenie alerting and on call management system.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;a href="#data-collection">Data Collection&lt;/a>
&lt;ul>
&lt;li>&lt;a href="#containers-and-vms-cgroups">Containers and VMs CGROUPS&lt;/a>&lt;/li>
&lt;li>&lt;a href="#docker">Docker&lt;/a>&lt;/li>
&lt;li>&lt;a href="#kubernetes">Kubernetes&lt;/a>&lt;/li>
&lt;li>&lt;a href="#kernel-tracesmetrics-ebpf">Kernel traces/metrics eBPF&lt;/a>&lt;/li>
&lt;li>&lt;a href="#disk-space-monitoring">Disk Space Monitoring&lt;/a>&lt;/li>
&lt;li>&lt;a href="#os-provided-metrics-procplugin">OS Provided Metrics proc.plugin&lt;/a>&lt;/li>
&lt;li>&lt;a href="#postgresql">PostgreSQL&lt;/a>&lt;/li>
&lt;li>&lt;a href="#dns-query">DNS Query&lt;/a>&lt;/li>
&lt;li>&lt;a href="#http-endpoint-check">HTTP endpoint check&lt;/a>&lt;/li>
&lt;li>&lt;a href="#elasticsearch-and-opensearch">Elasticsearch and OpenSearch&lt;/a>&lt;/li>
&lt;li>&lt;a href="#dnsmasq-dns-forwarder">Dnsmasq DNS Forwarder&lt;/a>&lt;/li>
&lt;li>&lt;a href="#envoy">Envoy&lt;/a>&lt;/li>
&lt;li>&lt;a href="#files-and-directories">Files and directories&lt;/a>&lt;/li>
&lt;li>&lt;a href="#rabbitmq">RabbitMQ&lt;/a>&lt;/li>
&lt;li>&lt;a href="#chartsdplugin">charts.d.plugin&lt;/a>&lt;/li>
&lt;li>&lt;a href="#anomalies">Anomalies&lt;/a>&lt;/li>
&lt;li>&lt;a href="#generic-structured-data-pandas">Generic structured data with Pandas&lt;/a>&lt;/li>
&lt;li>&lt;a href="#generic-prometheus-collector">Generic Prometheus collector&lt;/a>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;a href="#alerts-and-notifications">Alerts and Notifications&lt;/a>
&lt;ul>
&lt;li>&lt;a href="#notifications">Notifications&lt;/a>
&lt;ul>
&lt;li>&lt;a href="#improved-email-alert-notifications">Improved email alert notifications&lt;/a>&lt;/li>
&lt;li>&lt;a href="#receive-only-notifications-for-unreachable-nodes">Receive only notifications for unreachable nodes&lt;/a>&lt;/li>
&lt;li>&lt;a href="#ntfy-agent-alert-notifications">ntfy agent alert notifications&lt;/a>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;a href="#enhanced-real-time-alert-synchronization-on-netdata-cloud">Enhanced Real-Time Alert Synchronization on Netdata Cloud&lt;/a>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;a href="#visualizations--charts-and-dashboards">Visualizations / Charts and Dashboards&lt;/a>
&lt;ul>
&lt;li>&lt;a href="#events-feed">Events Feed&lt;/a>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;a href="#machine-learning">Machine Learning&lt;/a>&lt;/li>
&lt;li>&lt;a href="#installation-and-packaging">Installation and Packaging&lt;/a>
&lt;ul>
&lt;li>&lt;a href="#improved-linux-compatibility">Improved Linux compatibility&lt;/a>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;a href="#administration">Administration&lt;/a>
&lt;ul>
&lt;li>&lt;a href="#new-way-to-retrieve-netdataconf">New way to retrieve netdata.conf&lt;/a>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;a href="#documentation-and-demos">Documentation and Demos&lt;/a>&lt;/li>
&lt;li>&lt;a href="#deprecation-notice">Deprecation notice&lt;/a>
&lt;ul>
&lt;li>&lt;a href="#deprecated-in-this-release">Deprecated in this release&lt;/a>&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;a href="#netdata-agent-release-meetup">Netdata Agent Release Meetup&lt;/a>&lt;/li>
&lt;li>&lt;a href="#support-options">Support options&lt;/a>&lt;/li>
&lt;li>&lt;a href="#running-survey">Running survey&lt;/a>&lt;/li>
&lt;li>&lt;a href="#acknowledgements">Acknowledgements&lt;/a>&lt;/li>
&lt;/ul>
&lt;h2 id="netdata-open-source-growth">Netdata open-source growth&lt;/h2>
&lt;!-- Retrieve most of these stats from netdata/netdata/README.md badges -->
&lt;ul>
&lt;li>Over 62,000 GitHub Stars&lt;/li>
&lt;li>Over 1.5 million online nodes&lt;/li>
&lt;li>Almost 92 million sessions served&lt;/li>
&lt;li>Over 600 thousand total nodes in Netdata Cloud&lt;/li>
&lt;/ul>
&lt;h2 id="release-highlights">Release highlights&lt;/h2>
&lt;h3 id="netdata-charts-v30">Netdata Charts v3.0&lt;/h3>
&lt;p>We are excited to announce Netdata Charts v3.0 and the NIDL framework. These are currently available at Netdata Cloud. At the next Netdata release, the agent dashboard will be replaced to also use the same charts.&lt;/p></description></item><item><title>Infinite Scalability: Monitoring Without Limits</title><link>https://www.netdata.cloud/blog/netdata-inifinite-scalability/</link><pubDate>Thu, 04 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-inifinite-scalability/</guid><description>&lt;p>Scalability is crucial for monitoring systems as it ensures that they can accommodate growth, maintain performance, provide flexibility, optimize costs, enhance fault tolerance, and support informed decision-making, all of which are critical for effective infrastructure management.&lt;/p>
&lt;!--truncate-->
&lt;p>Most monitoring solutions struggle with scalability, mainly because of:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>High data volume and velocity&lt;/strong>: Monitoring systems generate vast amounts of data and as the infrastructure grows, so does the volume and velocity of these data.&lt;/li>
&lt;li>&lt;strong>Resource constraints&lt;/strong>: Scalability requires efficient resource utilization, leading to bottlenecks and performance issues as the monitored environment grows.&lt;/li>
&lt;li>&lt;strong>Architectural limitations&lt;/strong>: Monitoring systems are usually designed with certain architectural constraints that limit their scalability. Most open source solutions rely on monolithic or centralized architectures that can become overwhelmed at scale.&lt;/li>
&lt;/ol>
&lt;p>For open source solutions scalability has always been a challenge, increasing their complexity significantly (check for example the scalability issues of Prometheus), while for commercial solutions it usually results in increased data collection to visualization latency and cost.&lt;/p></description></item><item><title>Monitoring Disks: Workload, Latency &amp; Saturation</title><link>https://www.netdata.cloud/blog/monitoring-disks-understanding-workload-performance-utilisation-saturation-latency/</link><pubDate>Thu, 04 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/monitoring-disks-understanding-workload-performance-utilisation-saturation-latency/</guid><description>&lt;p>&lt;img src="../2023-05-04-monitoring-disks-understanding-workload-performance-utilisation-saturation-latency/img/stacked-netdata.png" alt="stacked-netdata">&lt;/p>
&lt;p>Netdata provides a comprehensive set of charts that can help you understand the workload, performance, utilization, saturation, latency, responsiveness, and maintenance activities of your disks.
In this blog we will focus on monitoring disks as block devices, not as filesystems or mount points.&lt;/p>
&lt;!-- truncate -->
&lt;p>The &lt;code>Disks&lt;/code> section in the &lt;code>Overview&lt;/code> tab contains all the charts that are mentioned in this blog post.
&lt;img src="../2023-05-04-monitoring-disks-understanding-workload-performance-utilisation-saturation-latency/img/disks-overview.png" alt="Disks-Overview">&lt;/p>
&lt;h2 id="disk-workload-and-performance">Disk Workload and Performance&lt;/h2>
&lt;p>Netdata charts for monitoring the workload and the throughput of your disks:&lt;/p></description></item><item><title>Understanding Huge Pages</title><link>https://www.netdata.cloud/blog/understanding-huge-pages/</link><pubDate>Thu, 04 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/understanding-huge-pages/</guid><description>&lt;p>Memory-intensive applications can benefit from &lt;a href="https://www.netdata.cloud/academy/what-is-application-performance-monitoring-apm/">improved performance&lt;/a> by using huge pages, as they can reduce TLB pressure and memory fragmentation, and lower the memory management overhead overall. Developers should consider using HugeTLBfs in their mmap() and shmget() calls to take advantage of huge pages.&lt;/p>
&lt;p>Transparent Huge Pages (THP) is a Linux kernel feature that provides some of the benefits of huge pages without requiring any development effort. However, THP can cause latency in many applications. Although kernel developers are actively working to address these issues, many system administrators prefer to disable THP altogether.&lt;/p></description></item><item><title>Unlock the Secrets of Kernel Memory Usage</title><link>https://www.netdata.cloud/blog/unlock-the-secrets-of-kernel-memory-usage/</link><pubDate>Thu, 04 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/unlock-the-secrets-of-kernel-memory-usage/</guid><description>&lt;p>&lt;img src="../2023-05-04-unlock-the-secrets-of-kernel-memory-usage/img/stacked-netdata.png" alt="stacked-netdata">&lt;/p>
&lt;p>The &lt;code>mem.kernel&lt;/code> chart in Netdata provides insight into the memory usage of &lt;a href="https://www.netdata.cloud/academy/what-are-the-differences-between-bpf-and-ebpf-an-overview/">various kernel subsystems&lt;/a> and mechanisms. By understanding these dimensions and their technical details, you can monitor your system&amp;rsquo;s kernel memory usage and identify potential issues or inefficiencies. Monitoring these dimensions can help you ensure that your system is running efficiently and provide valuable insights into the performance of your kernel and memory subsystem.&lt;/p>
&lt;p>&lt;img src="../2023-05-04-unlock-the-secrets-of-kernel-memory-usage/img/mem-kernel.png" alt="mem-kernel">&lt;/p>
&lt;!-- truncate -->
&lt;h2 id="slab">Slab&lt;/h2>
&lt;p>The &lt;a href="https://en.wikipedia.org/wiki/Slab_allocation">slab allocator&lt;/a> is a memory management mechanism introduced by Jeff Bonwick in 1994 to manage &lt;a href="https://www.netdata.cloud/vsphere-monitoring/">memory allocation&lt;/a> for kernel objects. The main purpose of the slab allocator is to reduce memory fragmentation and improve the speed of memory allocation/deallocation. The slab allocator groups objects of the same size into &amp;ldquo;slabs&amp;rdquo; and caches the objects to speed up future allocations.&lt;/p></description></item><item><title>Entropy In Cryptography: Key To Security &amp; Randomness</title><link>https://www.netdata.cloud/blog/understanding-entropy-the-key-to-secure-cryptography-and-randomness/</link><pubDate>Wed, 03 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/understanding-entropy-the-key-to-secure-cryptography-and-randomness/</guid><description>&lt;p>&lt;img src="../2023-05-03-understanding-entropy-the-key-to-secure-cryptography-and-randomness/img/stacked-netdata.png" alt="stacked-netdata">&lt;/p>
&lt;p>&lt;a href="https://en.wikipedia.org/wiki/Entropy_(computing)">Entropy&lt;/a> is a measure of the randomness or unpredictability of data. In the context of cryptography, entropy is used to generate random numbers or keys that are essential for secure communication and encryption. Without a good source of entropy, cryptographic protocols can become vulnerable to attacks that exploit the predictability of the generated keys.&lt;/p>
&lt;!-- truncate -->
&lt;h2 id="what-is-entropy">What is Entropy?&lt;/h2>
&lt;p>In most operating systems, entropy is generated by collecting random events from various sources, such as hardware interrupts, mouse movements, keyboard presses, and disk activity. These events are fed into a pool of entropy, which is then used to generate random numbers when needed.&lt;/p></description></item><item><title>Context Switching &amp; Its Impact On System Performance</title><link>https://www.netdata.cloud/blog/understanding-context-switching-and-its-impact-on-system-performance/</link><pubDate>Tue, 02 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/understanding-context-switching-and-its-impact-on-system-performance/</guid><description>&lt;p>&lt;img src="../2023-05-02-understanding-context-switching-and-its-impact-on-system-performance/img/stacked-netdata.png" alt="stacked-netdata">&lt;/p>
&lt;p>Context switching is the process of switching the CPU from one process, task or thread to another. In a multitasking operating system, such as Linux, the CPU has to switch between multiple processes or threads in order to keep the system running smoothly. This is necessary because each CPU core without hyperthreading can only execute one process or thread at a time. If there are many processes or threads running simultaneously, and very few CPU cores available to handle them, the system is forced to make more context switches to balance the CPU resources among them.&lt;/p></description></item><item><title>Linux CPU Consumption, Load &amp; Pressure Explained</title><link>https://www.netdata.cloud/blog/understanding-linux-cpu-consumption-load-and-pressure-for-performance-optimisation/</link><pubDate>Tue, 02 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/understanding-linux-cpu-consumption-load-and-pressure-for-performance-optimisation/</guid><description>&lt;p>&lt;img src="../2023-05-02-understanding-linux-cpu-consumption-load-and-pressure-for-performance-optimisation/img/stacked-netdata.png" alt="stacked-netdata">&lt;/p>
&lt;p>As a system administrator, understanding how your Linux system&amp;rsquo;s CPU is being utilized is crucial for identifying bottlenecks and &lt;a href="https://www.netdata.cloud/academy/what-is-cardinality-in-databases-a-comprehensive-guide/">optimizing performance&lt;/a>. In this blog post, we&amp;rsquo;ll dive deep into the world of Linux CPU consumption, load, and pressure, and discuss how to use these metrics effectively to identify issues and improve your system&amp;rsquo;s performance.&lt;/p>
&lt;!-- truncate -->
&lt;h2 id="cpu-consumption-and-utilization">CPU Consumption and Utilization&lt;/h2>
&lt;p>CPU consumption refers to the amount of processing power being used by applications running on your system. The &lt;code>system.cpu&lt;/code> chart in Netdata represents the Total CPU utilization of your Linux system, broken down into different dimensions. Each dimension provides insight into how the CPU is being used by various tasks and processes. Here&amp;rsquo;s a brief explanation of each dimension:&lt;/p></description></item><item><title>Server Uptime Monitoring: Core Benefits For High Performance</title><link>https://www.netdata.cloud/blog/server-uptime-monitoring-why-do-we-need-it/</link><pubDate>Tue, 02 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/server-uptime-monitoring-why-do-we-need-it/</guid><description>&lt;p>&lt;img src="../2023-05-02-server-uptime-monitoring-why-do-we-need-it/img/stacked-netdata.png" alt="Server Uptime Monitoring: Core Benefits For High Performance">&lt;/p>
&lt;p>Server uptime monitoring tracks the availability and reliability of servers within your infrastructure.&lt;/p>
&lt;!-- truncate -->
&lt;h2 id="what-is-server-uptime-monitoring">What Is Server Uptime Monitoring?&lt;/h2>
&lt;p>Server uptime monitoring is the process of continuously tracking the operational status of your servers to ensure optimal performance and availability for users.&lt;/p>
&lt;p>With Netdata, you gain access to real-time, high-resolution monitoring that goes beyond basic checks, providing a detailed overview of your entire infrastructure.&lt;/p></description></item><item><title>Swap Memory: When &amp; How To Use It On Production VMs</title><link>https://www.netdata.cloud/blog/swap-memory-when-and-how-to-use-it-on-your-production-systems-or-cloud-provided-vms/</link><pubDate>Tue, 02 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/swap-memory-when-and-how-to-use-it-on-your-production-systems-or-cloud-provided-vms/</guid><description>&lt;p>&lt;img src="../2023-05-02-swap-memory-when-to-use-in-production-systems/img/stacked-netdata.png" alt="Swap Memory: Its Use On Production Systems &amp;amp; Cloud-Provided VMs">&lt;/p>
&lt;p>Swap memory, also known as virtual memory, is a space on a hard disk that is used to supplement the physical memory (RAM) of a computer. The swap space is used when the system runs out of physical memory, and it moves less frequently accessed data from RAM to the hard disk, freeing up space in RAM for more frequently accessed data. But should swap memory be enabled on production systems and &lt;a href="https://www.netdata.cloud/vsphere-monitoring/">cloud-provided virtual machines&lt;/a> (VMs)? Let&amp;rsquo;s explore the pros and cons.&lt;/p></description></item><item><title>Understanding Interrupts, Softirqs, and Softnet in Linux</title><link>https://www.netdata.cloud/blog/understanding-interrupts-softirqs-and-softnet-in-linux/</link><pubDate>Tue, 02 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/understanding-interrupts-softirqs-and-softnet-in-linux/</guid><description>&lt;p>&lt;img src="../2023-05-02-understanding-interrupts-softirqs-and-softnet-in-linux/img/stacked-netdata.png" alt="stacked-netdata">&lt;/p>
&lt;p>Interrupts, softirqs, and softnet are all critical parts of the Linux kernel that can impact system performance. In this blog post, we&amp;rsquo;ll explore their usefulness, and discuss how to monitor them using Netdata for both bare-metal servers and VMs.&lt;/p>
&lt;!-- truncate -->
&lt;h2 id="what-are-interrupts">What are Interrupts?&lt;/h2>
&lt;p>Interrupts are signals generated by hardware devices to indicate that they require attention from the CPU. Hardware devices can generate interrupts for a variety of reasons, including data transmission or reception, input/output operations, and other activities. When an interrupt is generated, the CPU stops what it is doing and handles the interrupt. Interrupts can have a significant impact on system performance, especially if there are a high number of interrupts occurring.&lt;/p></description></item><item><title>Understanding System Processes States</title><link>https://www.netdata.cloud/blog/understanding-system-processes-states/</link><pubDate>Tue, 02 May 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/understanding-system-processes-states/</guid><description>&lt;p>&lt;img src="../2023-05-02-understanding-system-processes-states/img/stacked-netdata.png" alt="stacked-netdata">&lt;/p>
&lt;p>The different states of system processes are essential to understanding how a computer system works. Each state represents a specific point in a process&amp;rsquo;s life cycle and can impact system performance and stability.&lt;/p>
&lt;!-- truncate -->
&lt;h2 id="process-states">Process States&lt;/h2>
&lt;p>Netdata&amp;rsquo;s &lt;code>system.processes_state&lt;/code> chart provides a view of these states, allowing users to monitor system performance in real-time:&lt;/p>
&lt;p>&lt;img src="../2023-05-02-understanding-system-processes-states/img/system-processes.png" alt="system-processes">&lt;/p>
&lt;ol>
&lt;li>
&lt;p>&lt;strong>Running:&lt;/strong> A process is in the Running state when it is actively using the CPU and executing instructions. This state is resource-intensive and can lead to performance issues if there are too many Running processes, causing CPU contention and system slowdowns. Processes in the Running state are prioritized using scheduling algorithms to improve system performance.&lt;/p></description></item><item><title>Why Scalable Monitoring Matters For Modern Systems</title><link>https://www.netdata.cloud/blog/why-scalable-monitoring-is-essential-for-modern-distributed-systems/</link><pubDate>Wed, 26 Apr 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/why-scalable-monitoring-is-essential-for-modern-distributed-systems/</guid><description>&lt;p>&lt;img src="../2023-04-26-why-scalable-monitoring-is-essential/img/stacked-netdata.png" alt="stacked-netdata">&lt;/p>
&lt;p>It&amp;rsquo;s becoming increasingly common to discuss the importance of scalability in monitoring solutions and how it can impact the performance and reliability of distributed systems.&lt;/p>
&lt;!-- truncate -->
&lt;p>In today&amp;rsquo;s rapidly evolving technological landscape, organizations are increasingly relying on distributed systems to power their operations. These systems consist of multiple interconnected components that work together to deliver a cohesive experience. They can span across different geographic locations, and often involve a combination of &lt;a href="https://www.netdata.cloud/product/cloud-on-premises/">on-premises, cloud&lt;/a>, and &lt;a href="https://www.netdata.cloud/solutions/technologies/docker-monitoring/">container-based environments&lt;/a>. As such, effectively managing these complex systems is critical to ensuring optimal performance, reliability, and security.&lt;/p></description></item><item><title>Netdata's AI Insights &amp; Rapid Diagnostics</title><link>https://www.netdata.cloud/blog/netdata-ai-insights-rapid-diagnostics/</link><pubDate>Wed, 19 Apr 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-ai-insights-rapid-diagnostics/</guid><description>&lt;p>Introduction to Netdata&amp;rsquo;s new visualisation providing AI Insights, supporting Rapid Diagnostics.
&lt;img src="https://user-images.githubusercontent.com/96257330/233125254-f93c9520-0a3f-4844-8d43-1f3202a5e411.png" alt="logo">&lt;/p>
&lt;!--truncate-->
&lt;h2 id="a-new-era-in-monitoring-systems-dashboards">A New Era in Monitoring Systems Dashboards&lt;/h2>
&lt;p>We&amp;rsquo;re thrilled to share an important upgrade to Netdata: &lt;strong>AI Insights &amp;amp; Rapid Diagnostics&lt;/strong>, a technology aiming to redefine what we expect from a monitoring system.&lt;/p>
&lt;h2 id="challenges-with-traditional-monitoring-dashboards">Challenges with Traditional Monitoring Dashboards&lt;/h2>
&lt;p>Traditional monitoring systems rely on a query language to help engineers create dashboards and alerts. While these languages offer power and flexibility, they come with several challenges that make monitoring and troubleshooting more complex and time-consuming:&lt;/p></description></item><item><title>Remote UNIX System Monitoring Using Net-SNMP</title><link>https://www.netdata.cloud/blog/remote-unix-monitoring-with-net-snmp/</link><pubDate>Wed, 12 Apr 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/remote-unix-monitoring-with-net-snmp/</guid><description>&lt;p>&lt;img src="../2023-04-12-remote-unix-monitoring-with-net-snmp/img/img.jpg" alt="img">&lt;/p>
&lt;p>Need to monitor a UNIX-like system, but can’t install Netdata on it? With our SNMP collector and Net-SNMP,
you can get basic system information with just a bit of relatively quick and easy configuration.&lt;/p>
&lt;!-- truncate -->
&lt;h2 id="what-is-snmp">What is SNMP?&lt;/h2>
&lt;p>The Simple Network Management Protocol, commonly known as SNMP, is a relatively lightweight protocol designed for
monitoring and configuration management for network appliances like switches, routers or gateways. However, it can also
be used for those purposes on almost any UNIX-like system thanks to the &lt;a href="http://www.net-snmp.org/">Net-SNMP project&lt;/a>.&lt;/p></description></item><item><title>Anomaly Rates in the Menu!</title><link>https://www.netdata.cloud/blog/anomaly-rates-in-the-menu/</link><pubDate>Wed, 29 Mar 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/anomaly-rates-in-the-menu/</guid><description>&lt;p>The menu (on the &lt;a href="https://learn.netdata.cloud/docs/getting-started/monitor-your-infrastructure/home-overview-and-single-node-view#overview-and-single-node-view">overview or single node tab&lt;/a>) now has an &lt;a href="https://learn.netdata.cloud/docs/troubleshooting-and-machine-learning/machine-learning-ml-powered-anomaly-detection#anomaly-rate">anomaly rate&lt;/a> button built into it that, for the entire visible window or a highlighted time range, shows the maximum chart anomaly rate within each section.&lt;/p>
&lt;p>Read on to learn more about this new feature!&lt;/p>
&lt;iframe width="560" height="315" src="https://www.youtube.com/embed/PgVh_MFHMb0?si=F2Mq6wIxHJWaHykZ" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen>&lt;/iframe>
&lt;h2 id="wait-what-is-an-anomaly-rate">Wait, what is an anomaly rate?&lt;/h2>
&lt;p>Netdata is the only monitoring agent that natively (for every metric, with zero config and sane defaults) produces anomaly rates in addition to just collecting raw metrics.&lt;/p></description></item><item><title>Introducing the Netdata demo space</title><link>https://www.netdata.cloud/blog/netdata-demo/</link><pubDate>Fri, 24 Mar 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-demo/</guid><description>&lt;p>&lt;img src="https://user-images.githubusercontent.com/24860547/201481889-0cf8e192-683f-4a80-9b96-4f69dd85490f.png" alt="image">&lt;/p>
&lt;p>Introducing Netdata&amp;rsquo;s Demo Space, a quick and easy way to experience monitoring environments before you set them up yourself.&lt;/p>
&lt;!--truncate-->
&lt;p>At Netdata, we are always striving to provide the best monitoring experience for our users. We understand that adopting a new monitoring solution can sometimes be challenging, especially when you&amp;rsquo;re unsure of how it will fit your specific environment. That&amp;rsquo;s why we&amp;rsquo;re excited to announce the Netdata Demo Space!&lt;/p></description></item><item><title>Upcoming Changes to Plugins in Native Packages</title><link>https://www.netdata.cloud/blog/split-plugin-packages/</link><pubDate>Wed, 15 Mar 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/split-plugin-packages/</guid><description>&lt;p>At Netdata, we’re committed to trying to make Netdata work as well as possible for our users. Sometimes though,
that means changing things in ways that aren’t exactly seamless. Such a change is coming soon for users of our
native DEB and RPM packages, and this blog post will explain what’s happening, why we’re doing it, and what
it means for our users.&lt;/p>
&lt;!-- truncate -->
&lt;h2 id="whats-changing">What’s changing?&lt;/h2>
&lt;p>Starting shortly after the v1.39.0 release of the Netdata Agent, we will be splitting most of our external
data collection plugins out to their own individual packages instead of bundling them all in the main &lt;code>netdata&lt;/code>
package. We already have this type of split for our CUPS and FreeIPMI plugins, and this new change will extend
that to also provide separate packages for the following plugins:&lt;/p></description></item><item><title>Windows Server Monitoring Improvements</title><link>https://www.netdata.cloud/blog/windows-monitoring-improvements/</link><pubDate>Mon, 13 Mar 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/windows-monitoring-improvements/</guid><description>&lt;p>&lt;a href="https://www.netdata.cloud/blog/web-servers-and-their-performance/">Monitor your Windows server and applications&lt;/a> running on it with Netdata - simple, powerful and free.&lt;/p>
&lt;!--truncate-->
&lt;p>Hey Netdata community,&lt;/p>
&lt;p>We have some exciting news for you: we’re launching our new and updated &lt;a href="https://learn.netdata.cloud/docs/data-collection/monitor-anything/System%20Metrics/Windows-machines">Windows collectors&lt;/a> with the goal of making the &lt;a href="https://www.netdata.cloud/windows-monitoring/">Windows monitoring experience&lt;/a> as seamless as possible 🎉&lt;/p>
&lt;p>We know that Windows monitoring has been a long time ask from many of you, and we’ve been working hard to make it easier than ever to monitor your Windows metrics with Netdata.&lt;/p></description></item><item><title>Anomaly detection on Prometheus metrics</title><link>https://www.netdata.cloud/blog/anomaly-detection-on-prometheus-metrics/</link><pubDate>Wed, 01 Mar 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/anomaly-detection-on-prometheus-metrics/</guid><description>&lt;p>&lt;img src="../2023-03-01-anomaly-detection-on-prometheus-metrics/img/img.png" alt="img">&lt;/p>
&lt;p>We have recently extended the native machine learning (ML) based anomaly detection &lt;a href="https://learn.netdata.cloud/guides/monitor/anomaly-detection">capabilities&lt;/a> of Netdata to &lt;a href="https://github.com/netdata/netdata/issues/14218">support all metrics&lt;/a>, regardless on their collection frequency (&lt;code>update every&lt;/code>).&lt;/p>
&lt;p>Previously only metrics collected every second were supported, but now Netdata can run anomaly detection out of the box with zero config on metrics with any collection frequency.&lt;/p>
&lt;p>This post will illustrate an example of what this means using &lt;a href="https://prometheus.io/">Prometheus&lt;/a> metrics (via the &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/prometheus#gsc.tab=0">Netdata Prometheus collector&lt;/a>) since they typically have a default collection frequency of 10 seconds.&lt;/p></description></item><item><title>Monitor any SQL metrics with Netdata (and Pandas ❤️)</title><link>https://www.netdata.cloud/blog/monitor-any-sql-metrics-with-netdata/</link><pubDate>Wed, 22 Feb 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/monitor-any-sql-metrics-with-netdata/</guid><description>&lt;p>&lt;img src="../2023-02-22-monitor-any-sql-metrics-with-netdata/img/img.png" alt="img">&lt;/p>
&lt;p>We recently got this great feedback from a dear user in our &lt;a href="https://discord.com/channels/847502280503590932/1075370683393118278/1075723915265069106">Discord&lt;/a>:&lt;/p>
&lt;blockquote>
&lt;p>I would really like to use Netdata to monitor custom internal metrics that come from SQL, not a fan of having 10 diff systems doing essentially the same thing as is, Netdata is pretty much all there in that regard, just needs a few extra features.&lt;/p>
&lt;/blockquote>
&lt;p>This is great and exactly what we want, a clear problem or improvement we could make to help make that users monitoring life a little easier.&lt;/p></description></item><item><title>Introducing Netdata Functions (↑ Top)</title><link>https://www.netdata.cloud/blog/netdata-functions/</link><pubDate>Wed, 15 Feb 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-functions/</guid><description>&lt;p>Netdata is committed to making it simpler and easier for everyone to monitor and troubleshoot their infrastructure. With that goal in mind, we&amp;rsquo;re excited to announce the launch of our new &amp;ldquo;Functions&amp;rdquo; feature (↑ Top), which allows Netdata Agent collectors to expose &amp;ldquo;functions&amp;rdquo; that can be executed in run-time and on-demand.&lt;/p>
&lt;!--truncate-->
&lt;h3 id="what-are-netdata-functions">What are Netdata functions?&lt;/h3>
&lt;p>Netdata has always been synonymous with real time monitoring and automated dashboards, with the recent introduction of &amp;ldquo;functions&amp;rdquo;, there&amp;rsquo;s now a new way for users to troubleshoot their infrastructure. &lt;strong>A function, in the context of Netdata, is a routine or script that can be invoked to run on a node and retrieve useful information, which is then displayed in the Netdata cloud dashboard.&lt;/strong>&lt;/p></description></item><item><title>Introducing Netdata Paid Subscriptions</title><link>https://www.netdata.cloud/blog/introducing-netdata-paid-subscriptions/</link><pubDate>Fri, 10 Feb 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/introducing-netdata-paid-subscriptions/</guid><description>&lt;p>All Netdata functionality is and will be available for free forever in the Community Plan. Paid tiers include features targeted for businesses and users who would need to customise their monitoring solution with different levels of user access, extra notification mechanisms, customer support and more.&lt;/p>
&lt;!--truncate-->
&lt;p>&lt;strong>Hello Netdata community&lt;/strong>,&lt;/p>
&lt;p>We are excited to announce that we are introducing new &lt;strong>Paid Subscriptions&lt;/strong> to Netdata as of Wednesday, 22nd of February 2023.&lt;/p></description></item><item><title>Release 1.38: Dramatic Performance &amp; Stability Gains</title><link>https://www.netdata.cloud/blog/netdata-version-1.38/</link><pubDate>Mon, 06 Feb 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-version-1.38/</guid><description>&lt;p>Another release of the Netdata Monitoring solution is here!&lt;/p>

 &lt;div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
 &lt;iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/2EjKicsRYxw?autoplay=0&amp;amp;controls=1&amp;amp;end=0&amp;amp;loop=0&amp;amp;mute=0&amp;amp;start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video">&lt;/iframe>
 &lt;/div>

&lt;ul>
&lt;li>&lt;a href="#v1380-release-highlights">Release Highlights&lt;/a>
&lt;ul>
&lt;li>
&lt;p>&lt;strong>&lt;a href="#v1380-dbenginev2">DBENGINE v2&lt;/a>&lt;/strong>
The new open-source database engine for Netdata Agents, offering huge performance, scalability and stability improvements, with a fraction of memory footprint!&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>&lt;a href="#v1380-functions">FUNCTION: Processes&lt;/a>&lt;/strong>
Netdata beyond metrics! We added the ability for &lt;strong>runtime functions&lt;/strong>, that can be implemented by any data collection plugin, to offer unlimited visibility to anything, even not-metrics, that can be valuable while troubleshooting.&lt;/p></description></item><item><title>Extending Netdata's anomaly detection training window</title><link>https://www.netdata.cloud/blog/extending-anomaly-detection-training-window/</link><pubDate>Thu, 02 Feb 2023 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/extending-anomaly-detection-training-window/</guid><description>&lt;p>We have been busy at work under the hood of the Netdata agent to introduce new capabilities that let you extend the &amp;ldquo;training window&amp;rdquo; used by Netdata&amp;rsquo;s &lt;a href="https://learn.netdata.cloud/docs/nightly/setup/configure-machine-learning-ml-powered-anomaly-detection">native anomaly detection capabilities&lt;/a>.&lt;/p>
&lt;p>This blog post will discuss one of these improvements to help you reduce &amp;ldquo;&lt;a href="https://en.wikipedia.org/wiki/False_positives_and_false_negatives#False_positive_error">false positives&lt;/a>&amp;rdquo; by essentially extending the training window by using the new (beautifully named) &lt;code>number of models per dimension&lt;/code> configuration parameter.&lt;/p>
&lt;h2 id="background">Background&lt;/h2>
&lt;p>One of the most important considerations of our native anomaly detection capabilities is the overhead of running the training and scoring computations required to train thousands of models (one per metric) and produce &lt;a href="https://learn.netdata.cloud/docs/nightly/setup/configure-machine-learning-ml-powered-anomaly-detection#anomaly-bit">anomaly bits&lt;/a> every second based on those trained models.&lt;/p></description></item><item><title>Release 1.37.1: Patch Release For Security Issues</title><link>https://www.netdata.cloud/blog/netdata-version-1.37.1/</link><pubDate>Mon, 05 Dec 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-version-1.37.1/</guid><description>&lt;p>Netdata v1.37.1 is a patch release to address issues discovered since v1.37.0. Refer to the &lt;a href="https://github.com/netdata/netdata/releases/tag/v1.37.0">v.1.37.0 release notes&lt;/a> for the full scope of that release.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="release-v1371">Release v1.37.1&lt;/h2>
&lt;p>Netdata v1.37.1 is a patch release to address issues discovered since v1.37.0. Refer to the &lt;a href="https://github.com/netdata/netdata/releases/tag/v1.37.0">v.1.37.0 release notes&lt;/a> for the full scope of that release.&lt;/p>
&lt;p>The v1.37.1 patch release fixes the following issues:&lt;/p>
&lt;ul>
&lt;li>Parent agent crash when many children instances (re)connect at the same time, causing simultaneous SSL re-initialization (&lt;a href="https://github.com/netdata/netdata/pull/14076">PR #14076&lt;/a>).&lt;/li>
&lt;li>Agent crash during dbengine database file rotation while a page is being read while being deleted (&lt;a href="https://github.com/netdata/netdata/pull/14081">PR #14081&lt;/a>).&lt;/li>
&lt;li>Agent crash on metrics page alignment when metrics were stopped being collected for a long time and then started again (&lt;a href="https://github.com/netdata/netdata/pull/14086">PR #14086&lt;/a>).&lt;/li>
&lt;li>Broken Fedora native packages (&lt;a href="https://github.com/netdata/netdata/pull/14082">PR #14082&lt;/a>).&lt;/li>
&lt;li>Fix dbengine backfilling statistics (&lt;a href="https://github.com/netdata/netdata/pull/14074">PR #14074&lt;/a>).&lt;/li>
&lt;/ul>
&lt;p>In addition, the release contains the following optimizations and improvements:&lt;/p></description></item><item><title>Release 1.37: Infinite Scalability &amp; Database Tiering</title><link>https://www.netdata.cloud/blog/netdata-version-1.37/</link><pubDate>Wed, 30 Nov 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-version-1.37/</guid><description>&lt;p>Another release of the Netdata Monitoring solution is here!&lt;/p>
&lt;p>We focused on these key areas:&lt;/p>
&lt;p>Infinite scalability of the Netdata Ecosystem&lt;/p>
&lt;p>Default Database Tiering, offering months of data retention for typical Netdata Agent installations with default settings and years of data retention for dedicated Netdata Parents.&lt;/p>
&lt;p>Overview Dashboards at Netdata Cloud got a ton of improvements to allow slicing and dicing of data directly on the UI and overcome the limitations of the web technology when thousands of charts are presented on one page.&lt;/p></description></item><item><title>Monitor &amp; Troubleshoot ISP Performance With Netdata</title><link>https://www.netdata.cloud/blog/speedtest-monitoring/</link><pubDate>Mon, 28 Nov 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/speedtest-monitoring/</guid><description>&lt;p>Find out how to monitor your Internet speed and quality and how well your ISP is performing.&lt;/p>
&lt;p>&lt;img src="https://user-images.githubusercontent.com/24860547/204470316-4682e442-6df1-4c77-b1e4-fdd96dd404f0.jpg" alt="logo">&lt;/p>
&lt;!--truncate-->
&lt;h2 id="what-factors-affect-my-internet-speed">What Factors Affect My Internet Speed?&lt;/h2>
&lt;p>Several factors can influence your internet speed, ranging from your ISP&amp;rsquo;s infrastructure to your home setup. Here&amp;rsquo;s a breakdown:&lt;/p>
&lt;h3 id="1-isp-plan--bandwidth">1. ISP Plan &amp;amp; Bandwidth&lt;/h3>
&lt;p>The speed you experience depends on the plan you choose from your ISP. Higher-tier plans offer more bandwidth, which means faster speeds for downloading, streaming, and gaming.&lt;/p></description></item><item><title>How to monitor node reboots?</title><link>https://www.netdata.cloud/blog/monitoring-node-reboots/</link><pubDate>Thu, 17 Nov 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/monitoring-node-reboots/</guid><description>&lt;p>Monitoring the health and status of nodes and servers is a critical part of effective infrastructure monitoring.&lt;/p>
&lt;p>&lt;img src="https://user-images.githubusercontent.com/96257330/202475049-22838a0b-73b1-485b-8416-5fd49d6ccb53.png" alt="logo">&lt;/p>
&lt;!--truncate-->
&lt;h2 id="how-to-monitor-node-reboots">How to monitor node reboots?&lt;/h2>
&lt;p>One of the most critical tasks of monitoring an infrastructure is to check the health of its servers/nodes. In most cases, this results in setting up a &amp;ldquo;Hardware manager&amp;rdquo; from the hardware vendor delivering these servers or setting up an SNMP (or similar) agent to continuously monitor the availability of the server and report when there is a reboot / failure.&lt;/p></description></item><item><title>How To Mute Alerts During Maintenance Windows</title><link>https://www.netdata.cloud/blog/mute-alerts/</link><pubDate>Thu, 03 Nov 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/mute-alerts/</guid><description>&lt;p>The health management APIs in Netdata allows teams to eliminate unnecessary alerting during scheduled maintenance, testing, auto scaling events, and instance reboots.&lt;/p>
&lt;!--truncate-->
&lt;p>For all SREs, it is absolutely crucial to filter out expected events during maintenance windows and quickly pinpoint critical issues in your infrastructure. Every minute is crucial while dealing with troubleshooting issues and any distractions that may hijack the troubleshooting process should be subdued.
The health &lt;a href="https://www.netdata.cloud/blog/iot-monitoring-challenges/">management APIs&lt;/a> in Netdata allows teams to eliminate unnecessary alerting during scheduled maintenance, testing, &lt;a href="https://www.netdata.cloud/blog/server-uptime-monitoring-why-do-we-need-it/">auto scaling events&lt;/a>, and instance reboots.&lt;/p></description></item><item><title>Monitor indoor air quality with Airthings and Netdata</title><link>https://www.netdata.cloud/blog/airquality/</link><pubDate>Wed, 02 Nov 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/airquality/</guid><description>&lt;p>Monitoring indoor air quality with &lt;a href="https://www.airthings.com/">Airthings&lt;/a> and Netdata. Understanding and measuring common contaminants and pollutants reduces your risk of air quality health concerns.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="indoor-air-quality-and-what-to-monitor">Indoor air quality and what to monitor&lt;/h2>
&lt;p>Indoor air quality can be a crucial influence on your health, wellbeing and productivity.&lt;/p>
&lt;p>Understanding and measuring common contaminants and pollutants is the first step towards reducing your risk of air quality health concerns.&lt;/p>
&lt;p>Airthings is a company that makes great air quality sensors that measure a wide variety of different variables including:&lt;/p></description></item><item><title>Monitor KSM performance with Netdata</title><link>https://www.netdata.cloud/blog/ksm/</link><pubDate>Tue, 01 Nov 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/ksm/</guid><description>&lt;p>Monitoring KSM (Kernel Same-page Merging) performance at deduping memory shared across VMs.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="kernel-same-page-merging-ksm">Kernel Same-page Merging (KSM)&lt;/h2>
&lt;p>Linux kernels store memory in &lt;strong>pages&lt;/strong> which are moved in and out of memory as a single block. On most Linux architectures pages are 4096 bytes. &lt;strong>KSM&lt;/strong> (Kernel Same-page Merging) is a kernel feature that scans memory looking for pages with identical content, and then de-duplicates them. The most common use-case where such duplicate pages occur is on hosts running multiple virtual machines (VMs).&lt;/p></description></item><item><title>Monitoring &amp; troubleshooting Cassandra with Netdata</title><link>https://www.netdata.cloud/blog/cassandra-monitoring-part2/</link><pubDate>Sat, 29 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/cassandra-monitoring-part2/</guid><description>&lt;p>How to monitor and troubleshoot Cassandra with Netdata.&lt;/p>
&lt;p>&lt;img src="https://user-images.githubusercontent.com/24860547/198524087-37dda416-a9a9-4c55-b379-0f46e990f83f.png" alt="logo">&lt;/p>
&lt;!--truncate-->
&lt;p>&lt;em>&lt;strong>Note&lt;/strong>: This post is the second part of a Cassandra monitoring series. Be sure to read our first entry &lt;a href="https://blog.netdata.cloud/cassandra-monitoring-part1">here&lt;/a>.&lt;/em>&lt;/p>
&lt;h2 id="monitoring-cassandra-with-netdata">Monitoring Cassandra with Netdata&lt;/h2>
&lt;p>Netdata’s &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/cassandra">Cassandra collector documentation&lt;/a> explains how to set it up to collect metrics automatically.&lt;/p>
&lt;p>Once you have followed the instructions in the docs and have installed and configured Netdata on the Cassandra cluster you are ready to start monitoring and troubleshooting. Check out the &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/rooms/cassandra/overview">Cassandra demo room&lt;/a> to interact with the charts, metrics and other functionality described here.&lt;/p></description></item><item><title>How to monitor and fix Database bloats in PostgreSQL?</title><link>https://www.netdata.cloud/blog/postgresql-database-bloat/</link><pubDate>Fri, 28 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/postgresql-database-bloat/</guid><description>&lt;p>Database bloat is disk space that was used by a table or index and is available for reuse by the database but has not been reclaimed. Bloat is created when deleting or updating tables and indexes. Here&amp;rsquo;s how to deal with it!&lt;/p>
&lt;!--truncate-->
&lt;h2 id="what-is-database-bloat">What is Database bloat?&lt;/h2>
&lt;p>Database bloat is disk space that was used by a table or index and is available for reuse by the database but has not been reclaimed. Bloat is created when deleting or updating tables and indexes.&lt;/p></description></item><item><title>Cassandra Monitoring: Key Metrics &amp; Best Practices</title><link>https://www.netdata.cloud/blog/cassandra-monitoring-part1/</link><pubDate>Thu, 27 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/cassandra-monitoring-part1/</guid><description>&lt;p>What are the important Cassandra metrics to monitor and how to monitor them.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="what-is-cassandra--why-use-it">What Is Cassandra &amp;amp; Why Use It&lt;/h2>
&lt;p>Cassandra is an open-source, distributed, wide-column NoSQL database management system written in Java. Cassandra was originally developed by &lt;a href="https://twitter.com/hedvigeng">Avinash Lakshmanan&lt;/a> and &lt;a href="https://twitter.com/pmalik">Prashant Malik&lt;/a> at Facebook and then released as open source, eventually becoming part of the &lt;a href="https://www.netdata.cloud/apache-monitoring/">Apache&lt;/a> project.&lt;/p>
&lt;p>&lt;a href="https://www.netdata.cloud/integrations/data-collection/databases/cassandra/">Cassandra is a NoSQL database&lt;/a> - NoSQL (also known as &amp;ldquo;not only SQL&amp;rdquo;) databases do not require data to be stored in tabular format. They provide flexible schemas and scale easily with large amounts of data and high user loads.&lt;/p></description></item><item><title>How to find out which application is causing server load</title><link>https://www.netdata.cloud/blog/server-load/</link><pubDate>Wed, 26 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/server-load/</guid><description>&lt;p>We often hear the term load used to describe the state of a server or a device, but we&amp;rsquo;re here to tell you what it means, precisely, and how to monitor it.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="what-is-server-load">What is server load?&lt;/h2>
&lt;p>We often hear the term &amp;ldquo;load&amp;rdquo; used to describe the state of a server or a &lt;a href="https://www.netdata.cloud/blog/iot-monitoring-challenges/">device&lt;/a>. But what does it really mean?&lt;/p>
&lt;p>System load is a measure of the amount of computational work that a system performs. An overloaded system, by definition, isn&amp;rsquo;t able to complete all its
tasks per schedule - this affects the performance and productivity of the system. And while &amp;ldquo;load&amp;rdquo; often gets conflated with CPU usage there&amp;rsquo;s a lot more to it.&lt;/p></description></item><item><title>How to monitor the disk usage on your infrastructure</title><link>https://www.netdata.cloud/blog/disk-usage/</link><pubDate>Tue, 25 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/disk-usage/</guid><description>&lt;p>The most important part of disk usage monitoring is to check the utilization of each filesystem and each mount point which can reveal existing or impending issues with the storage space on your infrastructure.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="what-does-disk-usage-du-mean">What Does Disk Usage (DU) Mean?&lt;/h2>
&lt;p>Disk usage (DU) refers to the portion or percentage of computer storage that is currently in use. It contrasts with disk space or &lt;a href="https://www.netdata.cloud/blog/server-uptime-monitoring-why-do-we-need-it/">capacity&lt;/a>,
which is the total amount of space that a given disk is capable of storing. Disk usage is a crucial metric to any computing system,
as it gives the user the information needed not only for storage, but also software requirements and overall operation. Although it usually
refers to a computer’s hard disk, it may also refer to external storage, such as a USB drive or compact disc (CD).&lt;/p></description></item><item><title>7 types of Redis latency and how to fix it</title><link>https://www.netdata.cloud/blog/7-types-of-redis-latency/</link><pubDate>Mon, 24 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/7-types-of-redis-latency/</guid><description>&lt;p>Redis is designed to be fast. In most cases, it is. However, there are times when Redis may be slow, due to network issues, disk latency, or other factors. When this happens, it is important to be able to &lt;a href="https://www.netdata.cloud/blog/server-uptime-monitoring-why-do-we-need-it/">detect the slow down&lt;/a> and investigate the cause of Redis latency.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="redis--latency">Redis &amp;amp; latency&lt;/h2>
&lt;p>Latency is the maximum delay between the time a client issues a command and the time the reply to the command is received by the client. Redis has strict requirements on average and worst case latency.&lt;/p></description></item><item><title>How to monitor systemd service liveness</title><link>https://www.netdata.cloud/blog/systemd-service-liveness/</link><pubDate>Fri, 21 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/systemd-service-liveness/</guid><description>&lt;p>The life of a sysadmin or SRE is often difficult, but occasionally very simple things can make a huge difference. Basic monitoring of your systemd services is one of those simple things, which we sometimes overlook. The simplest question one would want to know is if the thing that’s supposed to be running is actually running at all. If you use systemd services, you can guarantee an answer to that question within minutes using Netdata.&lt;/p></description></item><item><title>Web Servers Monitoring: Key Metrics &amp; Strategies</title><link>https://www.netdata.cloud/blog/web-servers-and-their-performance/</link><pubDate>Thu, 20 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/web-servers-and-their-performance/</guid><description>&lt;!--truncate-->
&lt;h2 id="the-importance-of-monitoring-web-servers">The Importance Of Monitoring Web Servers&lt;/h2>
&lt;p>Web servers are among the most important components in modern &lt;a href="https://www.netdata.cloud/blog/future-of-infrastructure-monitoring/">IT infrastructures&lt;/a>. They host the websites, web services, and &lt;a href="https://www.netdata.cloud/dotnet-monitoring/">web applications&lt;/a> that we use on a daily basis. Social networking, media streaming, software as a service (SaaS), and other activities wouldn’t be possible without the use of web servers. And with the advent of cloud computing and the movement of more services online, &lt;a href="https://www.netdata.cloud/blog/server-uptime-monitoring-why-do-we-need-it/">web servers and their monitoring are only becoming more important&lt;/a>. Given the extensive usage of Web servers, Sysadmins and SREs should monitor web servers as a key aspect for performance. &lt;/p></description></item><item><title>Using Pandas In Python: Data Analysis &amp; Performance Insights</title><link>https://www.netdata.cloud/blog/pandas-python/</link><pubDate>Wed, 19 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/pandas-python/</guid><description>&lt;p>Netdata just got a &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/python.d.plugin/pandas" target="_blank" rel="noopener">Pandas collector&lt;/a>.&lt;/p>
&lt;!--truncate-->
&lt;p>Pandas is a de-facto standard in reading and processing most types of structured data in Python so if you have some csv/json/xml data, either locally or via some HTTP endpoint, containing metrics you&amp;rsquo;d like to monitor, chances are you can now easily do this by leveraging the Pandas collector without having to develop your own custom collector as you might have in the past.&lt;/p></description></item><item><title>How to monitor HTTP endpoints</title><link>https://www.netdata.cloud/blog/http-endpoints/</link><pubDate>Mon, 17 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/http-endpoints/</guid><description>&lt;p>The &lt;a href="https://en.wikipedia.org/wiki/Hypertext_Transfer_Protocol">HTTP protocol&lt;/a> has become the de facto standard application layer protocol of the internet. From publicly available web sites and APIs to “inter-process” communications in REST based microservice architectures or large &lt;a href="https://en.wikipedia.org/wiki/Service-oriented_architecture">Service Oriented Architectures&lt;/a> based on &lt;a href="https://en.wikipedia.org/wiki/SOAP">SOAP&lt;/a>, you find HTTP being used again and again, due to its simplicity and our familiarity with it. How many protocols can you name that have &lt;a href="https://imgur.com/gallery/4KqWq">memes&lt;/a> for their status codes? Of course, such a popular protocol has endless pages written about how to properly monitor the services that rely on it, with many options specific to every use case.&lt;!--truncate--> What you will learn here is how to get your basics done in monitoring HTTP endpoints, so you can be up and running in a few minutes, monitoring all HTTP services in your &lt;a href="https://www.netdata.cloud/blog/future-of-infrastructure-monitoring/">entire infrastructure&lt;/a>. &lt;/p></description></item><item><title>How to monitor DNS query response time</title><link>https://www.netdata.cloud/blog/dns-query-response-time/</link><pubDate>Wed, 12 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/dns-query-response-time/</guid><description>&lt;p>DNS (Domain Name System) servers translate standard language web addresses to their actual IP addresses for network access.&lt;/p>
&lt;p>&lt;a href="https://xiaolishen.medium.com/the-dns-lookup-journey-240e9a5d345c">&lt;strong>DNS Lookup Journey&lt;/strong>&lt;/a>
&lt;img src="../wp-archive/uploads/2022/10/DNS-1.png" alt="">&lt;/p>
&lt;!--truncate-->
&lt;p>DNS response time is the time it takes a Domain Name Server to receive the request for a domain name’s IP address, process it, and return the IP address to the browser or application requesting it. When it comes to DNS response times, the lower the better, and generally values less than 100ms are considered to be in the acceptable range (depending on the application).&lt;/p></description></item><item><title>Why is data replication important?</title><link>https://www.netdata.cloud/blog/why-is-data-replication-important/</link><pubDate>Wed, 12 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/why-is-data-replication-important/</guid><description>&lt;p>High availability. This is what every monitoring tool needs to ensure that you never compromise on IT infrastructure visibility.&lt;!--truncate--> On top of high availability, do you really want to enable all available features on your production system? It is important for the monitoring tool to have a low footprint on your CPU consumption and memory usage. Let’s dive deeper into the recommended way of configuring Netdata to ensure high availability and a low resource footprint through data replication.&lt;/p></description></item><item><title>How to monitor host reachability</title><link>https://www.netdata.cloud/blog/how-to-monitor-host-reachability/</link><pubDate>Mon, 10 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/how-to-monitor-host-reachability/</guid><description>&lt;p>Most sysadmins and developers have at some point used a few of the popular &lt;a href="https://www.tecmint.com/linux-networking-commands" target="_blank" rel="noopener">Linux networking commands&lt;/a> or their Windows equivalents to answer the common questions of &lt;a href="https://www.netdata.cloud/blog/web-servers-and-their-performance/">host reachability&lt;/a> - that is, whether a host or service is reachable and how fast it responds.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="common-approaches-to-reachability">Common approaches to reachability&lt;/h2>
&lt;p>One of the simplest, common checks, is to simply &lt;code>ping&lt;/code> a host to verify that it’s reachable from where you issue the command, and to see the total time it takes for the host to receive your request.&lt;/p></description></item><item><title>Introducing the Netdata Source Plugin for Grafana</title><link>https://www.netdata.cloud/blog/introducing-netdata-source-plugin-for-grafana/</link><pubDate>Fri, 07 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/introducing-netdata-source-plugin-for-grafana/</guid><description>&lt;p>&lt;img src="../wp-archive/uploads/2022/09/postgresql_dash-600x354.png" alt="sample-dashboard">&lt;/p>
&lt;p>The open-source community is about to benefit greatly from Netdata&amp;rsquo;s new Grafana data source plugin, which makes use of a powerful data collection engine.&lt;/p>
&lt;!--truncate-->
&lt;p>This new plugin maximizes the troubleshooting capabilities of Netdata in Grafana, making them more widely available. Some of the key capabilities provided to you with this plugin include the following:&lt;/p>
&lt;ul>
 	&lt;li>Real-time monitoring with single-second granularity.&lt;/li>
 	&lt;li>Installation and out-of-the-box integrations available in seconds from one line of code.&lt;/li>
 	&lt;li>2,000+ metrics from across your whole Infrastructure, with insightful metadata associated with them.&lt;/li>
 	&lt;li>Access to our fresh ML metrics (anomaly rates) - exposing our ML capabilities at the edge!&lt;/li>
&lt;/ul>
&lt;h2 id="why-did-we-decide-to-do-it">Why did we decide to do it?&lt;/h2>
&lt;p>We are huge fans of Open-Source culture. Open-source is deeply rooted in Netdata&amp;rsquo;s DNA. Because of this, at Netdata, we don’t really buy into the “single pane of glass” or “observability platform” buzzwords. The reality is that things are just more complicated than that in real life.&lt;/p></description></item><item><title>How to filter metrics by label?</title><link>https://www.netdata.cloud/blog/how-to-filter-metrics-by-label/</link><pubDate>Thu, 06 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/how-to-filter-metrics-by-label/</guid><description>&lt;p>It is sometimes easy to get lost in the mountain of metrics and infinite number of dimensions when working with an infrastructure monitoring tool. Being able to filter metrics by label and visualize only what is relevant to the current scope of monitoring &amp;amp;troubleshooting, becomes absolutely crucial to the success of SREs, Sysadmins and DevOps professionals.&lt;/p>
&lt;!--truncate-->
&lt;p>The Netdata &lt;a href="https://staging1--netdata-docusaurus.netlify.app/docs/getting-started/netdata-in-a-pane">chart label filtering feature&lt;/a> supports grouping by and filtering each chart based on labels (key/value pairs) applicable to the context and provides fine-grain capability on slicing the data / metrics.&lt;/p></description></item><item><title>Missing indexes in PostgreSQL? How to quickly identify it</title><link>https://www.netdata.cloud/blog/missing-indexes-in-postgresql/</link><pubDate>Wed, 05 Oct 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/missing-indexes-in-postgresql/</guid><description>&lt;p>While working on improving the &lt;a href="https://netdata.cloud/postgresql-monitoring/">Netdata PostgreSQL collector&lt;/a>, we were monitoring our production PostgreSQL instance and something caught our attention immediately. The rows fetched ratio seemed really, really low for one particular database&amp;hellip; there were missing indexes in PostgreSQL!&lt;/p>
&lt;!--truncate-->
&lt;p>&lt;b>Rows fetched ratio&lt;/b> is the percentage of rows that contain data needed to execute the query (rows fetched), out of the total number of rows scanned (rows returned). A low value indicates that the database is performing extra work by scanning a large number of rows that aren’t required to process the query.&lt;/p></description></item><item><title>Data Collection Strategies For Infrastructure</title><link>https://www.netdata.cloud/blog/data-collection-strategies/</link><pubDate>Tue, 06 Sep 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/data-collection-strategies/</guid><description>&lt;!--truncate-->
&lt;p>Monitoring and troubleshooting; unfortunately, these terms are still used interchangeably, which can lead to misunderstandings about data collection strategies.&lt;/p>
&lt;p>In this article we aim to clarify some important definitions, processes, and common data collection strategies for monitoring solutions. We will specify the limitations of the described strategies, as well as key benefits which can potentially be also used for troubleshooting needs.&lt;/p>
&lt;p>&lt;strong>IT infrastructure monitoring&lt;/strong> is a business process of collecting and analyzing data over a period of time to improve business results.&lt;/p></description></item><item><title>How Netdata’s Machine Learning works</title><link>https://www.netdata.cloud/blog/how-netdatas-machine-learning-works/</link><pubDate>Thu, 01 Sep 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/how-netdatas-machine-learning-works/</guid><description>&lt;p>Following on from the &lt;a href="https://www.netdata.cloud/blog/introducing-anomaly-advisor-unsupervised-anomaly-detection-in-netdata" target="_blank" rel="noopener">recent launch&lt;/a> of our &lt;a href="https://learn.netdata.cloud/docs/cloud/insights/anomaly-advisor" target="_blank" rel="noopener">Anomaly Advisor&lt;/a> feature, and in keeping with &lt;a href="https://www.netdata.cloud/blog/our-approach-to-machine-learning/" target="_blank" rel="noopener">our approach to machine learning&lt;/a>, &lt;a href="https://github.com/netdata/netdata/blob/master/ml/notebooks/netdata_anomaly_detection_deepdive.ipynb" target="_blank" rel="noopener">here&lt;/a> is a detailed Python notebook outlining exactly how the machine learning powering the Anomaly Advisor actually works under the hood.&lt;/p>
&lt;!--truncate-->
&lt;p>Or if you&amp;rsquo;d rather watch a video walkthrough of the notebook then check out below.&lt;/p>
&lt;iframe width="560" height="315" src="https://www.youtube.com/embed/L1xleckyuDQ?si=rptYzWE-eLlhSL9x" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen>&lt;/iframe>
&lt;p>Try it for yourself, &lt;a href="https://learn.netdata.cloud/docs/cloud/get-started" target="_blank" rel="noopener">get started&lt;/a> by &lt;a href="https://app.netdata.cloud/?utm_source=blog&amp;amp;utm_content=how_netdata_ml_works" target="_blank" rel="noopener">signing in to Netdata&lt;/a> and connecting a node. Once initial models have been trained (usually after the agent has about one hour of data, zero configuration needed), you&amp;rsquo;ll be able to start exploring in the &lt;a href="https://learn.netdata.cloud/docs/cloud/insights/anomaly-advisor" target="_blank" rel="noopener">Anomaly Advisor&lt;/a> tab of Netdata.&lt;/p></description></item><item><title>Anomaly rate in every chart</title><link>https://www.netdata.cloud/blog/anomaly-rate-in-every-chart/</link><pubDate>Thu, 23 Jun 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/anomaly-rate-in-every-chart/</guid><description>&lt;p>A month ago, we introduced unsupervised ML &amp;amp; Anomaly Detection in Netdata, the &lt;a href="https://www.netdata.cloud/blog/introducing-anomaly-advisor-unsupervised-anomaly-detection-in-netdata/">Anomaly Advisor&lt;/a>. Today, we’re happy to announce that we’re bringing anomaly rates to every chart in Netdata Cloud. Anomaly information is no longer limited to the Anomalies tab and will be accessible to you from the Overview and Single Node View tabs as well. This will make your troubleshooting journey easier, as you will have the anomaly rates for any metric available with a single click. Whichever metric or chart you&amp;rsquo;re exploring will be instant.&lt;/p></description></item><item><title>Metric Correlations on the Agent</title><link>https://www.netdata.cloud/blog/metric-correlations-on-the-agent/</link><pubDate>Wed, 15 Jun 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/metric-correlations-on-the-agent/</guid><description>&lt;p>As of &lt;a href="https://github.com/netdata/netdata/releases/tag/v1.35.0" target="_blank" rel="noopener">&lt;code>v1.35.0&lt;/code>&lt;/a> the Netdata Agent can now run &lt;a href="https://learn.netdata.cloud/docs/cloud/insights/metric-correlations" target="_blank" rel="noopener">Metric Correlations&lt;/a> (MC) itself. This means that, for nodes with MC enabled, the Metric Correlations feature just got a whole lot faster!&lt;/p>
&lt;!--truncate-->
&lt;p>The Netdata Metric Correlations feature uses a &lt;a href="https://en.wikipedia.org/wiki/Kolmogorov%E2%80%93Smirnov_test#Two-sample_Kolmogorov%E2%80%93Smirnov_test" target="_blank" rel="noopener">Two Sample Kolmogorov-Smirnov test&lt;/a> to look for which metrics have a significant distributional change around a highlighted window of interest. This can be useful when you are interested in short term &amp;ldquo;&lt;a href="https://en.wikipedia.org/wiki/Change_detection" target="_blank" rel="noopener">change detection&lt;/a>&amp;rdquo; and want to try answer the question &amp;ldquo;what else changed around this time?&amp;rdquo;.&lt;/p></description></item><item><title>Anomaly Advisor: Unsupervised Anomaly Detection</title><link>https://www.netdata.cloud/blog/introducing-anomaly-advisor-unsupervised-anomaly-detection-in-netdata/</link><pubDate>Thu, 26 May 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/introducing-anomaly-advisor-unsupervised-anomaly-detection-in-netdata/</guid><description>&lt;p>Today we are excited to launch one of our flagship ML assisted troubleshooting features in Netdata – the Anomaly Advisor.&lt;/p>
&lt;p>The Anomaly Advisor builds on earlier work to introduce unsupervised &lt;a href="https://github.com/netdata/netdata/blob/master/ml/README.md">anomaly detection&lt;/a> capabilities into the &lt;a href="https://www.netdata.cloud/agent/">Netdata Agent&lt;/a> from &lt;a href="https://github.com/netdata/netdata/releases/tag/v1.32.0">v1.32.0&lt;/a> onwards.&lt;/p>
&lt;h2 id="getting-started">Getting Started&lt;/h2>
&lt;p>Once you &lt;a href="https://learn.netdata.cloud/docs/configure/machine-learning#configuration">enable ML&lt;/a> on your nodes, each node will begin producing an &amp;ldquo;&lt;a href="https://learn.netdata.cloud/docs/configure/machine-learning#anomaly-bit---100--anomalous-0--normal">Anomaly Bit&lt;/a>&amp;rdquo; every second in addition to raw metric values. This anomaly bit will be 1 when the trained ML models consider recent raw data for a metric to look anomalous or 0 when things look &amp;rsquo;normal&amp;rsquo;. The Anomaly Advisor leverages this information to enable seamless space or room level anomaly detection out of the box with minimal configuration.&lt;/p></description></item><item><title>Monitoring without Cooperation: Kubernetes</title><link>https://www.netdata.cloud/blog/monitoring-without-cooperation-kubernetes/</link><pubDate>Fri, 20 May 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/monitoring-without-cooperation-kubernetes/</guid><description>&lt;!--truncate-->
&lt;iframe width="560" height="315" src="https://www.youtube.com/embed/J2kdSTRJzV4?si=Rqv4gpqz_wZKusaZ" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen>&lt;/iframe>
&lt;p>Imagine this: You are an engineer at a startup. You are responsible for keeping all the applications running smoothly and safely in production. At first, you have things under control, but soon enough things start getting more complex. The system has grown by hundreds of new server nodes, new services and containers are showing up seemingly every day, and you hear there’s a project to transition everything to Kubernetes (you’ve heard the term “cloud native” so often that the phrase shows up in your dreams).&lt;/p></description></item><item><title>Kubernetes Throttling Doesn’t Have To Suck. Let Us Help!</title><link>https://www.netdata.cloud/blog/kubernetes-throttling-doesnt-have-to-suck-let-us-help/</link><pubDate>Tue, 03 May 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/kubernetes-throttling-doesnt-have-to-suck-let-us-help/</guid><description>&lt;p>CPU limits are probably the most misunderstood concept in Kubernetes CPU resources allocation and management.&lt;/p>
&lt;!--truncate-->
&lt;p>A lot of engineers advise the use of CPU limits on every container as Kubernetes best practice. Unfortunately, as we will prove below, they are wrong: CPU limits should rarely be used, if used at all!&lt;/p>
&lt;p>But why? What are the reasons that even senior DevOps engineers with vast experience in the field advise the use of CPU limits?&lt;/p></description></item><item><title>Troubleshooting Alerts the Right Way: As a Team</title><link>https://www.netdata.cloud/blog/troubleshooting-alerts/</link><pubDate>Thu, 28 Apr 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/troubleshooting-alerts/</guid><description>&lt;!--truncate-->
&lt;p>At Netdata, we love two things more than anything else: &lt;/p>
&lt;ol>
 	&lt;li> Troubleshooting and,&lt;/li>
 	&lt;li> Making things easier for you! &lt;/li>
&lt;/ol>
Our goal is to make troubleshooting and monitoring as seamless as possible with the open-source Agent. This includes giving you pre-configured alerts so that you get notified immediately when a disruption occurs.
&lt;p>The Netdata Agent comes with over 250 pre-configured and optimized alerts. But we want you to be the master of your infrastructure monitoring by:&lt;/p></description></item><item><title>CNCF Live: Machine Learning Anomaly Detection</title><link>https://www.netdata.cloud/blog/cncf-live-power-up-your-machine-learning-automated-anomaly-detection/</link><pubDate>Wed, 27 Apr 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/cncf-live-power-up-your-machine-learning-automated-anomaly-detection/</guid><description>&lt;h2 id="join-us-live-to-talk-ml">Join Us Live to Talk ML&lt;/h2>
&lt;p>Join ML Lead Andrew Maguire and Product Manager Shyam Sreevalsan on the 23rd of June at 5pm UTC for the Netdata Machine Learning Meetup, which will be livestreamed on Youtube. In this session, we will demo and preview new and future features, along with a Q&amp;amp;A, and discuss the following topics:&lt;/p>
&lt;ul>
 	&lt;li>The role of Machine Learning in DevOps Infrastructure Monitoring &amp;amp; Troubleshooting&lt;/li>
 	&lt;li>The Netdata way of approaching Machine Learning&lt;/li>
 	&lt;li>Challenges of building Macvhine learning solutions that are useful and user friendly.&lt;/li>
&lt;/ul>
Feel free to ask questions or share ideas on &lt;a href="https://discord.gg/ZyeDHQTdaW">our Community Discord&lt;/a>, or &lt;a href="https://(https://www.meetup.com/netdata-infrastructure-monitoring-meetup-group/events/286243158/">RSVP to the event&lt;/a>. We look forward to seeing you then!
&lt;h2 id="cncf-live-power-up-your-machine-learning---automated-anomaly-detection">CNCF Live: Power up your machine learning - Automated anomaly detection&lt;/h2>
&lt;p>Our Analytics &amp;amp; ML lead Andrew Maguire recently had a chance to share our new &lt;a href="https://community.netdata.cloud/t/anomaly-advisor-beta-launch/2717">Anomaly Advisor&lt;/a> feature with the wider CNCF community. In his demonstration he did some light chaos engineering (using &lt;a href="https://www.gremlin.com/">Gremlin&lt;/a> and &lt;a href="https://wiki.ubuntu.com/Kernel/Reference/stress-ng">stress-ng&lt;/a>) to generate some real anomalies on his infrastructure and watch how it all played out in the Anomaly Advisor in Netdata Cloud.&lt;/p></description></item><item><title>The Netdata Way of Troubleshooting</title><link>https://www.netdata.cloud/blog/the-netdata-way-of-troubleshooting/</link><pubDate>Mon, 04 Apr 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/the-netdata-way-of-troubleshooting/</guid><description>&lt;p>Together with you, our fabulous community, Netdata is changing the way the world thinks of high fidelity monitoring - and we are gaining momentum.&lt;/p>
&lt;p>Our chief troublemaker and CEO, Costa Tsaousis,  is the pioneer and architect of this revolution that’s brewing in the monitoring and troubleshooting space.&lt;/p>
&lt;p>Watch him explain the &lt;strong>Netdata way of troubleshooting&lt;/strong>:&lt;/p>
&lt;iframe width="560" height="315" src="https://www.youtube.com/embed/ExjwwrgXvPg?si=IrsEp9PqbFFWiegM" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen>&lt;/iframe>
&lt;h2 id="why-did-the-world-need-a-new-infrastructure-monitoring-tool-0018">Why did the world need a new infrastructure monitoring tool? (00:18)&lt;/h2>
&lt;p>Every great hero needs an origin story. In this section, Costa explains the frustrating conditions in the infrastructure and monitoring space that lead to the conception, development, and subsequent success of the Netdata open-source Agent and Netdata Cloud.&lt;/p></description></item><item><title>Our Approach to Machine Learning</title><link>https://www.netdata.cloud/blog/our-approach-to-machine-learning/</link><pubDate>Fri, 25 Mar 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/our-approach-to-machine-learning/</guid><description>&lt;p>There is a lot of buzz in the world of machine learning (ML) and as a layperson it can be hard to keep up with it all. Therefore, we decided to write down some of our thoughts and musings on how &lt;b>we&lt;/b> are approaching ML at Netdata.&lt;/p>
&lt;!--truncate-->
&lt;h2 id="our-approach-to-machine-learning-ml">Our Approach to Machine Learning (ML)&lt;/h2>
&lt;p>We’ll touch on the current state of applied ML in industry in general, and zoom in on ML in the monitoring industry. We’ll discuss how we can leverage “good honest ML” to punch above our weight and add some useful and novel features for our users over the next few years.&lt;/p></description></item><item><title>Engineering Team Best Practices For Cloud Monitoring</title><link>https://www.netdata.cloud/blog/netdata-troubleshooting-show-engineering-team-best-practices-in-working-with-cloud/</link><pubDate>Wed, 02 Mar 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-troubleshooting-show-engineering-team-best-practices-in-working-with-cloud/</guid><description>&lt;!--truncate-->
&lt;figure class="wp-block-image size-large">&lt;img src="../wp-archive/uploads/2022/03/group_promo-1200x675.png" alt="" class="wp-image-16147"/>&lt;/figure>
&lt;h2 id="panel-engineering-team-best-practices-in-working-with-cloud">Panel: Engineering Team Best Practices in Working with Cloud&lt;/h2>
&lt;p>&lt;a href="https://youtube.com/playlist?list=PL-P-gAHfL2KMN-LHSmQFjpEwl31rQYvam" target="_blank">&lt;strong>The Netdata Troubleshooting Show&lt;/strong>&lt;/a>| Season 1: Episode 2 | Thursday, March 3rd at 9 am PST (UTC/GMT -8)&lt;/p>
&lt;p>&lt;strong>[&lt;a href="https://youtu.be/zY3DiRJ_DYc" target="_blank">Join Live&lt;/a>]&lt;/strong>For engineering teams, working in the cloud has never been easier, but also more complex. Enter a new world of remote working with distributed global teams, new tech challenges, new business realities, security, monitoring, and more.&lt;/p>
&lt;p>Join our dynamic panel of experts live&lt;strong>(&lt;/strong>&lt;a href="https://youtu.be/zY3DiRJ_DYc" target="_blank">&lt;strong>watch here&lt;/strong>&lt;/a>&lt;strong>)&lt;/strong>as they talk about engineering team best practices for those working in the cloud in 2022.&lt;/p></description></item><item><title>Netdata Meetup | Install &amp; Monitor From Scratch | Guide</title><link>https://www.netdata.cloud/blog/netdata-meetup-real-world-scenario-on-how-to-install-and-monitor-from-scratch/</link><pubDate>Wed, 02 Mar 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-meetup-real-world-scenario-on-how-to-install-and-monitor-from-scratch/</guid><description>&lt;!--truncate-->
&lt;p>&lt;img src="../wp-archive/uploads/2022/03/netdata-meetup-1-1200x676.png" alt="">&lt;/p>
&lt;p>Announcing Netdata Meetups: Virtual How-to Live Show – Join us Friday, March 4th&lt;/p>
&lt;p>&lt;a href="https://youtu.be/lBd0-TFJGAY">Join us live this Friday!&lt;/a> We are launching the brand new Netdata Meetups! Join us to learn more about Netdata. This how-to virtual series will help you go deeper with Netdata and as a bonus we will add in tips and tricks for monitoring and troubleshooting. We can’t wait to see you on Friday, March 4th at 9 am PST (GMT -8).&lt;/p></description></item><item><title>Netdata Cloud’s New Architecture</title><link>https://www.netdata.cloud/blog/netdata-clouds-new-architecture/</link><pubDate>Thu, 10 Feb 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-clouds-new-architecture/</guid><description>&lt;!--truncate-->
&lt;figure class="wp-block-image size-full">&lt;img class="wp-image-16154" src="../wp-archive/uploads/2022/03/New_arch_notification.jpeg" alt="" />&lt;/figure>
&lt;p>In version v1.32 of Netdata, we announced a remarkable new update that we are extremely proud of; Netdata Cloud now runs on the most reliable and stable backend that we’ve ever built. &lt;/p>
&lt;h2 id="migration-to-the-new-netdata-cloud-architecture">Migration to the new Netdata Cloud architecture&lt;/h2>
&lt;p>To give you the best experience of Netdata Cloud, we started migrating nodes running on the old architecture to the new one. Most users don’t have to take any action on their part. If you need to take action, you will see the pop-up window above in Netdata Cloud.  &lt;/p></description></item><item><title>Meet The Netdata Community: eBPF Hero Thiago</title><link>https://www.netdata.cloud/blog/meet-the-netdata-community-every-company-needs-an-ebpf-superhero-like-thiago/</link><pubDate>Fri, 04 Feb 2022 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/meet-the-netdata-community-every-company-needs-an-ebpf-superhero-like-thiago/</guid><description>&lt;!--truncate-->
&lt;p>&lt;img src="../wp-archive/uploads/2022/03/thiago-2.png" alt="">&lt;/p>
&lt;p>Our ongoing Meet the community series focuses on global Netdata community members. In this installment, learn about Netdata staff member and Software Engineer Thiago Marques, who is hard at work building an eBPF.plugin.&lt;/p>
&lt;h3 id="introduce-yourself-what-you-do-and-your-current-job-role-at-netdata">Introduce yourself, what you do, and your current job role at Netdata.&lt;/h3>
&lt;p>I am Thiago Marques, a C developer who works with the data collector team. My primary responsibility in the company is to develop eBPF.plugin.&lt;/p></description></item><item><title>Meet The Netdata Community: Rupok Chowdhury Protik</title><link>https://www.netdata.cloud/blog/meet-the-netdata-community-rupok-chowdhury-protik-takes-software-engineering-photography-to-new-heights/</link><pubDate>Tue, 07 Dec 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/meet-the-netdata-community-rupok-chowdhury-protik-takes-software-engineering-photography-to-new-heights/</guid><description>&lt;!--truncate-->
&lt;figure class="wp-block-image size-full">&lt;img src="../wp-archive/uploads/2022/03/Meet-the-community.png" alt="" class="wp-image-16164"/>&lt;/figure>
&lt;p>In our ongoing Meet the community series, we focus on global Netdata community members. In this installment, learn about Software Engineer and photographer Rupok Chowdhury Protik, including how they use Netdata.&lt;/p>
&lt;p>&lt;strong>Introduce yourself, what you do, and your current job role.&lt;/strong>&lt;/p>
&lt;p>Hi, I’m Rupok Chowdhury Protik. I’m a PHP Developer with a Master’s degree in Software Engineering. My stack is mainly LAMP/LEMP. I used to work as the Senior Web Application Developer in a multinational company and then later joined Incsub LLC, a WordPress-focused company with people from more than 50 countries. I am currently working in Incsub LLC as a DevOps Support Engineer.&lt;/p></description></item><item><title>Netdata Cloud Show: The Observability Panel</title><link>https://www.netdata.cloud/blog/netdata-cloud-show-the-observability-panel/</link><pubDate>Tue, 07 Dec 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-cloud-show-the-observability-panel/</guid><description>&lt;!--truncate-->
&lt;figure class="wp-block-image size-large">&lt;img class="wp-image-16169" src="../wp-archive/uploads/2022/03/Netdata-Cloud-Panel-1200x674.png" alt="" />&lt;/figure>
&lt;h2 id="in-2022-60-of-security-incidents-will-involve-third-partiesstrong-forresterstrong">“In 2022, 60% of security incidents will involve third parties.”&lt;strong>– Forrester&lt;/strong>&lt;/h2>
&lt;figure class="wp-block-embed is-type-rich is-provider-embed-handler wp-block-embed-embed-handler wp-embed-aspect-16-9 wp-has-aspect-ratio">
&lt;div class="wp-block-embed__wrapper">https://www.youtube.com/embed/kILpPCRlVD0&lt;/div>
&lt;/figure>
&lt;p>&lt;strong>This Thursday, December 9th at 1 PM EST (UTC-5)&lt;/strong>Join us live this Thursday for the first-ever Netdata Cloud Show, where our all-star panel will be discussing the latest trends around observability. &lt;/p>
&lt;p>This will be broadcast live on Netdata &lt;a href="https://youtu.be/kILpPCRlVD0">YouTube&lt;/a>, &lt;a href="https://twitter.com/linuxnetdata">Twitter&lt;/a>, &lt;a href="https://www.facebook.com/linuxnetdata/">Facebook&lt;/a>, and &lt;a href="https://discord.gg/kUk3nCmbtx">Discord&lt;/a>.&lt;/p>
&lt;p>&lt;strong>Observability Panel:&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Boris Zaikin&lt;/strong>, Senior Software and Cloud Architect at Nordcloud. &lt;a href="https://twitter.com/boriszzn">@boriszzn&lt;/a>&lt;/li>
&lt;li>&lt;strong>Samir Behara&lt;/strong>, Platform Architect at EBSCO Industries, Inc. &lt;a href="https://twitter.com/samirbehara">@samirbehara&lt;/a>&lt;/li>
&lt;li>&lt;strong>Ralph Meijer&lt;/strong>, VP of Technology at Netdata. &lt;a href="https://twitter.com/ralphm">@ralphm&lt;/a>&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Topics Covered Include:&lt;/strong>&lt;/p></description></item><item><title>All-new Netdata Cloud Charts 2.0</title><link>https://www.netdata.cloud/blog/all-new-netdata-cloud-charts-2-0/</link><pubDate>Tue, 30 Nov 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/all-new-netdata-cloud-charts-2-0/</guid><description>&lt;!--truncate-->
&lt;figure class="wp-block-image size-large">&lt;img src="../wp-archive/uploads/2022/03/Netdata-Charts-2.0-1200x704.png" alt="" class="wp-image-16212"/>&lt;/figure>
&lt;p>Netdata excels in collecting, storing, and organizing metrics in out-of-the-box dashboards for powerful troubleshooting. We are now doubling down on this by transforming data into even more effective visualizations, helping you make the most sense out of all your metrics for increased observability.&lt;/p>
&lt;p>The new Netdata Charts provide a ton of useful information and we invite you to further explore our new charts from a design and development perspective. As always, it’s our goal to be as open and transparent as possible with our users on all things Netdata, including the ins and outs of how Netdata is built, why we make certain design decisions (driven by you of course!), and where we are heading as we continue to grow.&lt;/p></description></item><item><title>Meet The Netdata Community | Bastien, SysAdmin &amp; Dog Lover</title><link>https://www.netdata.cloud/blog/meet-the-netdata-community-learn-about-bastien-a-system-and-network-admin-who-loves-dogs/</link><pubDate>Tue, 30 Nov 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/meet-the-netdata-community-learn-about-bastien-a-system-and-network-admin-who-loves-dogs/</guid><description>&lt;!--truncate-->
&lt;figure class="wp-block-image size-full">&lt;img src="../wp-archive/uploads/2022/03/Screen-Shot-2021-11-30-at-3.10.06-PM.png" alt="" class="wp-image-16223"/>&lt;/figure>
&lt;p>&lt;strong>Tell people about yourself and what you do.&lt;/strong>&lt;/p>
&lt;p>My name is Bastien. I’m 28, a System and Network Admin, currently SRE in a team of two for my company. We’re working with AWS and Kubernetes to host our web apps.&lt;/p>
&lt;p>&lt;strong>How are you using Netdata and what do you like so far?&lt;/strong>&lt;/p>
&lt;p>I’m currently in the process of moving to Netdata to monitor our nine Kubernetes clusters and some standalone VMs. What I particularly love about Netdata is the simplicity of installation and configuration and the number of perfect default metrics and graphs.&lt;/p></description></item><item><title>How to extend the Geth collector</title><link>https://www.netdata.cloud/blog/how-to-extend-the-geth-collector/</link><pubDate>Mon, 13 Sep 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/how-to-extend-the-geth-collector/</guid><description>&lt;!--truncate-->
&lt;figure class="wp-block-image size-large">&lt;img src="../wp-archive/uploads/2022/03/Geth-collector-diagram-1-1200x796.png" alt="" class="wp-image-16282"/>&lt;/figure>
&lt;p>This is the the last of a 2-part blog post series regarding Netdata and Geth. If you missed the first, be sure to check it out &lt;a href="https://hackmd.io/J1x1WA-bR0a8gQeAmVdLFw" target="_blank" rel="noreferrer noopener">here&lt;/a>.&lt;/p>
&lt;p>Geth is short for Go-Ethereum and is the official implementation of the Ethereum Client in Go. Currently it’s one of the most widely used implementations and a core piece of infrastructure for the Ethereum ecosystem.&lt;/p>
&lt;p>With this proof of concept I wanted to showcase how easy it really is to gather data from any Prometheus endpoint and visualize them in Netdata. This has the added benefit of leveraging all the other features of Netdata, namely it’s per-second data collection, automatic deployment and configuration and superb system monitoring.&lt;/p></description></item><item><title>How to monitor the Geth node in under 5 minutes</title><link>https://www.netdata.cloud/blog/how-to-monitor-the-geth-node-in-under-5-minutes/</link><pubDate>Mon, 13 Sep 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/how-to-monitor-the-geth-node-in-under-5-minutes/</guid><description>&lt;!--truncate-->
&lt;figure class="wp-block-image size-large">&lt;img src="../wp-archive/uploads/2022/03/How-Geth-affect-the-CPU-charts-1200x629.png" alt="" class="wp-image-16229"/>&lt;/figure>
&lt;p>This piece is a blog post version of a workshop I gave at &lt;a href="https://ethcc.io/" target="_blank" rel="noreferrer noopener">EthCC&lt;/a> about monitoring an Ethereum Node using Netdata.&lt;/p>
&lt;p>&lt;strong>Disclaimer:&lt;/strong> Although we use Netdata, this guide is generic. We talk about metrics that can be surfaced by many other tools, such as Prometheus/Grafana or Datadog.&lt;/p>
&lt;p>The contents are as follows:&lt;/p>
&lt;ul>&lt;li class="">Introduction to Ethereum Nodes&lt;/li>&lt;li class="">What is Netdata&lt;/li>&lt;li class="">How to monitor a system that runs go-ethereum (Geth)&lt;/li>&lt;li class="">How to monitor go-ethereum (Geth)&lt;/li>&lt;/ul>
&lt;h2 id="ethereum-nodes">Ethereum Nodes&lt;/h2>
&lt;p>Running a node is no small feat, as it requires increasingly more and more resources to store the state of the blockchain and quickly process new transactions.&lt;/p></description></item><item><title>Root cause analysis using Metric Correlations</title><link>https://www.netdata.cloud/blog/root-cause-analysis-using-metric-correlations/</link><pubDate>Fri, 03 Sep 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/root-cause-analysis-using-metric-correlations/</guid><description>&lt;!--truncate-->
&lt;figure class="wp-block-image size-large">&lt;img src="../wp-archive/uploads/2022/03/Screen-Shot-2021-09-03-at-1.43.32-PM-1-1200x608.png" alt="" class="wp-image-16297"/>&lt;/figure>
&lt;p>As complexity of systems and applications continue to evolve and change, the number of metrics that need to be monitored grows in parallel. Whether you’re on a DevOps team, an SRE, or a developer building the code yourself, many of these components may be fragmented across your infrastructure, making it increasingly difficult to identify the root cause when experiencing downtime or abnormal behavior. To help solve this challenge, we built the &lt;a href="https://learn.netdata.cloud/docs/cloud/insights/metric-correlations">Metric Correlations&lt;/a> feature – an automated analysis tool that evaluates all your metrics to identify which have changed the most within a given period of interest.&lt;/p></description></item><item><title>How To Monitor Disks &amp; Filesystems With eBPF</title><link>https://www.netdata.cloud/blog/how-to-monitor-your-disks-and-filesystems-now-also-with-ebpf/</link><pubDate>Mon, 16 Aug 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/how-to-monitor-your-disks-and-filesystems-now-also-with-ebpf/</guid><description>&lt;!--truncate-->
&lt;div class="et_pb_module et_pb_text et_pb_text_0 et_pb_text_align_left et_pb_bg_layout_light">
&lt;div class="et_pb_text_inner">
&lt;img class="alignnone size-medium wp-image-16371" src="../wp-archive/uploads/2021/08/eBPF-monitoring-1536x932-1-600x364.png" alt="" width="600" height="364" />
&lt;h2 id="introduction-to-ebpf">Introduction to eBPF&lt;/h2>
&lt;/div>
&lt;/div>
&lt;div class="et_pb_module et_pb_text et_pb_text_1 et_pb_text_align_left et_pb_bg_layout_light">
&lt;div class="et_pb_text_inner">
&lt;p>Current IT monitoring software lacks the necessary metrics for minimizing downtime for systems and applications. Most provide system and application metrics but there is much more than this required for properly monitoring your infrastructure. With &lt;a title="eBPF" href="https://ebpf.io/" target="_blank" rel="noopener">eBPF&lt;/a> there is a technological advancement that allows monitoring software to provide rich information from the Linux kernel and present it. eBPF monitoring, specifically, provides a better understanding of what exactly is occurring on internal systems, which helps to identify where performance improvements can be made.&lt;/p></description></item><item><title>Netdata Is Launching Its Discord Server</title><link>https://www.netdata.cloud/blog/netdata-is-launching-its-discord-server/</link><pubDate>Tue, 22 Jun 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-is-launching-its-discord-server/</guid><description>&lt;!--truncate-->
&lt;p>It’s been a long time since our last community update, rest assured that we have been hard at work here at Netdata.&lt;/p>
&lt;p>Community building is hard, especially when you have such a venerable community like the one here at Netdata, where hundreds of contributors have contributed to creating one of the best monitoring solutions that exist.&lt;/p>
&lt;p>Last year we started to concentrate working on consolidating the community by integrating the various platforms where people come together to talk about Netdata. In that effort, we restarted our community forums in November, reworking the categories and creating a support channel, where community members help each other.&lt;/p></description></item><item><title>Netdata v1.31.0</title><link>https://www.netdata.cloud/blog/netdata-v1-31/</link><pubDate>Wed, 19 May 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-v1-31/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-medium wp-image-16402" src="../wp-archive/uploads/2021/05/v1.31.01-1-600x375.png" alt="" width="600" height="375" />
&lt;p>Give a warm welcome to Netdata v1.31.0, which features:&lt;/p>
&lt;ul>
 	&lt;li aria-level="1">&lt;strong>Re-packaged and redesigned dashboard&lt;/strong>: A more informational and feature-rich “frame” for your monitoring and troubleshooting sessions.&lt;/li>
 	&lt;li aria-level="1">&lt;strong>eBPF expands into the directory cache&lt;/strong>: Monitor whether your services or applications are properly using Linux’s memory management for the best performance and minimal disk I/O.&lt;/li>
 	&lt;li aria-level="1">&lt;strong>Machine learning-powered collectors&lt;/strong>: Detect anomalies using only your own data and minimal resource utilization on your monitored nodes.&lt;/li>
 	&lt;li aria-level="1">&lt;strong>An improved Netdata learning experience&lt;/strong>: A timeline of new content, refreshed visuals, and a newly-open sourced repository.&lt;/li>
&lt;/ul>
&lt;h2 id="h_1876248811621353710666">Re-packaged and redesigned dashboard&lt;/h2>
We re-packaged and redesigned portions of the dashboard to improve the overall experience. Part of this effort is better handling of dashboard code during installation—anyone using third-party packages (such as the Netdata Homebrew formula) will start seeing new features and the new designs starting today.
&lt;p>For those who aren’t using third-party packages (thank you!), your installation process will still get a little bit faster.&lt;/p></description></item><item><title>Kubernetes monitoring and troubleshooting made simple</title><link>https://www.netdata.cloud/blog/kubernetes-monitoring-troubleshooting/</link><pubDate>Wed, 05 May 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/kubernetes-monitoring-troubleshooting/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone wp-image-16419 size-large" src="../wp-archive/uploads/2022/03/kubernetes-monitoring-troubleshooting-1200x828.png" alt="" width="1200" height="828" />
&lt;p>Infrastructure monitoring was difficult enough when entire businesses ran off a few &lt;a href="https://www.netdata.cloud/academy/bare-metal-server/">bare metal servers&lt;/a> in a dusty, forgotten closet. Other IT infrastructure monitoring tools fell short, unable to provide complete and granular-enough metrics in real time, even when we were only dealing with a handful of systems responsible for running every part of the application stack. They were hard to configure, especially for the non-gurus out there, and didn’t provide the high-resolution metrics the gurus needed to make data-driven troubleshooting decisions.&lt;/p></description></item><item><title>Netdata 1.30.0</title><link>https://www.netdata.cloud/blog/netdata-1-30-0/</link><pubDate>Thu, 01 Apr 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-1-30-0/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-medium wp-image-16427" src="../wp-archive/uploads/2022/03/v1.30.0-600x338.png" alt="" width="600" height="338" />
&lt;p>We’re excited to introduce Netdata 1.30.0, which features:&lt;/p>
&lt;ul>
 	&lt;li>&lt;a href="https://staging-www.netdata.cloud/blog/release-1-30-0/#h_13754275791617303933429" target="_blank" rel="noopener">&lt;strong>ACLK-NG&lt;/strong>&lt;/a>: A new custom library for streaming metrics data on demand, written entirely in-house, that’s 4x faster than libmosquitto/libwebsockets.&lt;/li>
 	&lt;li>&lt;a href="https://staging-www.netdata.cloud/blog/release-1-30-0/#h_51169176821617303905806">&lt;strong>Opt-in product telemetry with PostHog&lt;/strong>&lt;/a>&lt;strong>:&lt;/strong> Goodbye, Google Analytics. Hello, self-hosted instance of PostHog.&lt;/li>
 	&lt;li>&lt;strong>&lt;a href="https://staging-www.netdata.cloud/blog/release-1-30-0/#h_612104678161617303939018">Deeper Linux kernel monitoring with eBPF&lt;/a>&lt;/strong>: Expanding our reach into the Linux kernel with page cache and synchronization syscall monitoring.&lt;/li>
 	&lt;li>&lt;strong>&lt;a href="https://staging-www.netdata.cloud/blog/release-1-30-0/#h_327502801221617303950492">Smarter preconfigured alarms&lt;/a>&lt;/strong>: Better (and less noisy) defaults, better information.&lt;/li>
 	&lt;li>&lt;strong>&lt;a href="https://staging-www.netdata.cloud/blog/release-1-30-0/#h_282794194271617303959644">Developer environment&lt;/a>&lt;/strong>: Contribute to Netdata via a Docker image and VSCode integration.&lt;/li>
 	&lt;li>&lt;strong>&lt;a href="https://staging-www.netdata.cloud/blog/release-1-30-0/#h_740997254311617303965795">Documentation improvements &amp;amp; tutorials&lt;/a>&lt;/strong>: Better standards for editing files and restarting Netdata, plus brand-new tutorials.&lt;/li>
&lt;/ul>
&lt;h2 id="h_13754275791617303933429">ACLK-NG&lt;/h2>
The ACLK-NG is a new, faster method of securely connecting a node running Netdata to Netdata Cloud. In our internal testing, it’s 4x faster than our previous implementation, which uses &lt;a href="https://github.com/netdata/mosquitto" target="_blank" rel="noopener">libmosquitto&lt;/a> and &lt;a href="https://github.com/warmcat/libwebsockets" target="_blank" rel="noopener">libwebsockets&lt;/a>.
&lt;table id="tablepress-10" class="tablepress tablepress-id-10 tablepress-responsive" >
&lt;thead>
&lt;tr class="row-1 odd">
&lt;th class="column-2" colspan="2">ACLK-NG&lt;/th>
&lt;th class="column-4" colspan="2">ACLK Legacy&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody class="row-hover">
&lt;tr class="row-2 even">
&lt;td class="column-1">&lt;/td>
&lt;td class="column-2">time (s)&lt;/td>
&lt;td class="column-3">MB/s&lt;/td>
&lt;td class="column-4">time (s)&lt;/td>
&lt;td class="column-5">MB/s&lt;/td>
&lt;/tr>
&lt;tr class="row-3 odd">
&lt;td class="column-1">Run 1&lt;/td>
&lt;td class="column-2">1.30&lt;/td>
&lt;td class="column-3">76.92&lt;/td>
&lt;td class="column-4">5.30&lt;/td>
&lt;td class="column-5">18.87&lt;/td>
&lt;/tr>
&lt;tr class="row-4 even">
&lt;td class="column-1">Run 2&lt;/td>
&lt;td class="column-2">1.31&lt;/td>
&lt;td class="column-3">76.34&lt;/td>
&lt;td class="column-4">5.29&lt;/td>
&lt;td class="column-5">18.90&lt;/td>
&lt;/tr>
&lt;tr class="row-5 odd">
&lt;td class="column-1">Run 3&lt;/td>
&lt;td class="column-2">1.38&lt;/td>
&lt;td class="column-3">72.46&lt;/td>
&lt;td class="column-4">5.27&lt;/td>
&lt;td class="column-5">18.98&lt;/td>
&lt;/tr>
&lt;tr class="row-6 even">
&lt;td class="column-1">Run 4&lt;/td>
&lt;td class="column-2">1.27&lt;/td>
&lt;td class="column-3">78.74&lt;/td>
&lt;td class="column-4">5.40&lt;/td>
&lt;td class="column-5">18.52&lt;/td>
&lt;/tr>
&lt;tr class="row-7 odd">
&lt;td class="column-1">Run 5&lt;/td>
&lt;td class="column-2">1.24&lt;/td>
&lt;td class="column-3">80.65&lt;/td>
&lt;td class="column-4">5.46&lt;/td>
&lt;td class="column-5">18.32&lt;/td>
&lt;/tr>
&lt;tr class="row-8 even">
&lt;td class="column-1">&lt;/td>
&lt;td class="column-2">&lt;strong>1.30&lt;/strong>&lt;/td>
&lt;td class="column-3">&lt;strong>77.02&lt;/strong>&lt;/td>
&lt;td class="column-4">&lt;strong>5.34&lt;/strong>&lt;/td>
&lt;td class="column-5">&lt;strong>18.72&lt;/strong>&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
With ACLK-NG enabled, you’ll get a snappier experience in Netdata Cloud, as there will be far less latency between requests for metrics and the subsequent response from individual nodes.
&lt;p>To enable ACLK-NG right now, update your nodes with the &lt;code>&amp;ndash;aclk-ng&lt;/code> option:&lt;/p></description></item><item><title>Container deployment showdown: Docker or Kubernetes?</title><link>https://www.netdata.cloud/blog/container-deployment-showdown-docker-or-kubernetes/</link><pubDate>Wed, 24 Mar 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/container-deployment-showdown-docker-or-kubernetes/</guid><description>&lt;!--truncate-->
&lt;div class="et_pb_module et_pb_text et_pb_text_0 et_pb_text_align_left et_pb_bg_layout_light">
&lt;div class="et_pb_text_inner">
&lt;img class="alignnone wp-image-16433 size-large" src="../wp-archive/uploads/2022/03/Kubernetes_vs_Docker-1200x828.png" alt="" width="1200" height="828" />
&lt;p>Monitoring the current state and performance of applications is critical for IT Ops and DevOps teams alike. Understanding the health of an application is one of the most effective ways of anticipating potential bottlenecks or slowdowns, yet it’s one of the largest challenges faced by many organizations that build and deploy software. This is largely due to applications’ distributed and diversified nature. A single outage has the potential to interrupt entire processes that, at times, can interfere with business as a whole and result in a negative effect on the bottom line.&lt;/p></description></item><item><title>5 DevOps best practices to reinforce with monitoring tools</title><link>https://www.netdata.cloud/blog/devops-best-practices-monitoring-tools/</link><pubDate>Thu, 18 Feb 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/devops-best-practices-monitoring-tools/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-medium wp-image-16441" src="../wp-archive/uploads/2022/03/devops-best-practices-monitoring-tools-v2-600x414.png" alt="" width="600" height="414" />
&lt;p>As part of a modern software development team, you’re asked to do a lot. You’re supposed to build faster, release more frequently, crush bugs, and integrate testing suites along the way. You’re supposed to implement and practice a strong DevOps culture, &lt;a href="https://github.com/upgundecha/howtheysre" target="_blank" rel="noopener noreferrer">read entire novels&lt;/a> about SRE best practices, go &lt;a href="https://en.wikipedia.org/wiki/Agile_software_development" target="_blank" rel="noopener noreferrer">agile&lt;/a>, or add a bunch of Scrum ceremonies to everyone’s calendar. Every week, the industry recommends that you “&lt;a href="https://devops.com/devops-shift-left-avoid-failure/" target="_blank" rel="noopener noreferrer">shift-left&lt;/a>” another part of the &lt;a href="https://staging-www.netdata.cloud/blog/agile-static-analysis/" target="_blank" rel="noopener noreferrer">DevOps pipeline&lt;/a>, to the point where you’re supposed to handle everything from unit testing to production deployment optimization from day one.&lt;/p></description></item><item><title>Introduction to StatsD</title><link>https://www.netdata.cloud/blog/introduction-to-statsd/</link><pubDate>Wed, 03 Feb 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/introduction-to-statsd/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-medium wp-image-16449" src="../wp-archive/uploads/2022/03/StatsD-600x414.png" alt="" width="600" height="414" />
&lt;p>StatsD is an industry-standard technology stack for monitoring applications and instrumenting any piece of software to deliver custom metrics. The StatsD architecture is based on delivering the metrics via UDP packets from any application to a central statsD server. Although the original StatsD server was written in Node.js, there are many implementations today, with Netdata being one of them.&lt;/p>
&lt;p>StatsD makes it easier for you to instrument your applications, delivering value around three main pillars: open-source, control, and modularity. That’s a real windfall for full-stack developers who need to code quickly, troubleshoot application issues on the fly, and often don’t have the necessary background knowledge to use complex monitoring platforms.&lt;/p></description></item><item><title>Actionable Alerts With Fewer False Positives</title><link>https://www.netdata.cloud/blog/actionable-intelligent-alerts/</link><pubDate>Thu, 21 Jan 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/actionable-intelligent-alerts/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-medium wp-image-16460" src="../wp-archive/uploads/2022/03/intelligent-alarms-600x413.png" alt="" width="600" height="413" />
&lt;p>Think about any sport or competitive activity, whether that’s football or a spelling bee. They always feature at least one person who acts as a moderator, referee, or judge. With their domain expertise, this person watches everyone’s behavior and constantly compares that against a set of rules. If someone crosses that threshold, they blow a whistle or throw up a flag. They are, in effect, saying that things have gone from &lt;strong>OK&lt;/strong> to &lt;strong>not OK&lt;/strong>.&lt;/p></description></item><item><title>Four key metrics for responding to IT incidents and failures</title><link>https://www.netdata.cloud/blog/four-key-metrics-for-responding-to-it-incidents-and-failures/</link><pubDate>Thu, 07 Jan 2021 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/four-key-metrics-for-responding-to-it-incidents-and-failures/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone wp-image-16474 size-large" src="../wp-archive/uploads/2021/01/DevOps-Metrics-1200x862.png" alt="" width="1200" height="862" />
&lt;div class="et_pb_module et_pb_text et_pb_text_0 et_pb_text_align_left et_pb_bg_layout_light">
&lt;div class="et_pb_text_inner">
&lt;p>If you’re a veteran in this space, you probably understand the many incident response metrics and concepts, along with the many (at times exasperating) acronyms. For those new to the space, or even those with years of experience, the &lt;a title="terminology" href="https://en.wikipedia.org/wiki/List_of_computing_and_IT_abbreviations" target="_blank" rel="noopener noreferrer">terminology&lt;/a> is often overwhelming.&lt;/p>
&lt;p>If you’re one of those people who’s struggling to navigate through the world of DevOps metrics, we’ve created this article for you. In this post, we’ll cover four main incident response metrics: MTTA, MTTR, MTBF, and MTTF. Learn what these acronyms mean, how you can use them, and how you can tie these KPIs into your Netdata experience.&lt;/p></description></item><item><title>Netdata Year in Review 2020</title><link>https://www.netdata.cloud/blog/netdata-year-in-review-2020/</link><pubDate>Fri, 18 Dec 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-year-in-review-2020/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-medium wp-image-16482" src="../wp-archive/uploads/2022/03/Roadmap-Header-600x322.png" alt="" width="600" height="322" />
&lt;p>Looking back at the unprecedented challenges we faced together in 2020, we’d like to extend our thanks to the community of people who have continued to work towards Netdata’s mission of simplifying monitoring and troubleshooting for everyone. Let’s review some of this year’s highlights.&lt;/p>
&lt;table>
&lt;tbody>
&lt;tr>
&lt;td width="50%">
&lt;table>
&lt;tbody>
&lt;tr>
&lt;td>
&lt;h3>Community&lt;/h3>
&lt;/td>
&lt;td>More than &lt;a title="https://github.com/netdata/netdata/graphs/contributors" href="https://github.com/netdata/netdata/graphs/contributors" target="_blank" rel="noopener noreferrer">400 contributors&lt;/a> have helped us grow.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Nearly 50,000 GitHub &lt;a title="stargazers" href="https://github.com/netdata/netdata/stargazers" target="_blank" rel="noopener noreferrer">stargazers&lt;/a> follow our progress.&lt;/td>
&lt;td>Our community on &lt;a title="GitHub" href="https://github.com/netdata/netdata/" target="_blank" rel="noopener noreferrer">GitHub&lt;/a> and our &lt;a title="forums" href="https://community.netdata.cloud/" target="_blank" rel="noopener noreferrer">forums&lt;/a> has grown to 4,000 strong.&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;/td>
&lt;td width="50%">&lt;img class="wp-image-16480 aligncenter" src="../wp-archive/uploads/2022/03/community.png" alt="" width="227" height="225" />&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&amp;nbsp;
&lt;table>
&lt;tbody>
&lt;tr>
&lt;td width="50%">&lt;img class="size-full wp-image-16492 aligncenter" src="../wp-archive/uploads/2020/12/stargazers.png" alt="" width="225" height="226" />&lt;/td>
&lt;td width="50%">
&lt;table>
&lt;tbody>
&lt;tr>
&lt;td>
&lt;h3>Netdata Cloud&lt;/h3>
&lt;/td>
&lt;td>Launched in May, now with more than 25,000 users &lt;a title="registered" href="https://app.netdata.cloud/" target="_blank" rel="noopener noreferrer">registered&lt;/a>.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Continuous delivery of new features included custom dashboards, Metric Correlations, overview page, and centralized alarm notifications.&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;table>
&lt;tbody>
&lt;tr>
&lt;td width="50%">
&lt;table>
&lt;tbody>
&lt;tr>
&lt;td>
&lt;h3>Netdata Agent&lt;/h3>
&lt;/td>
&lt;td>9 dot &lt;a title="releases" href="https://github.com/netdata/netdata/releases" target="_blank" rel="noopener noreferrer">releases&lt;/a>, with support for eBPF, Prometheus metrics, &amp;amp; k8s service discovery.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>More than 300 bug fixes, nearly 1,500 closed &lt;a title="issues" href="https://github.com/netdata/netdata/issues" target="_blank" rel="noopener noreferrer">issues&lt;/a>, and almost 3,200 commits.&lt;/td>
&lt;td>Nearly 400 new features or &lt;a title="improvements" href="https://github.com/netdata/netdata/" target="_blank" rel="noopener noreferrer">improvements&lt;/a>.&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;/td>
&lt;td width="50%">&lt;img class="wp-image-16499 size-full aligncenter" src="../wp-archive/uploads/2020/12/Github-1.png" alt="" width="256" height="256" />&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&amp;nbsp;
&lt;table>
&lt;tbody>
&lt;tr>
&lt;td width="50%">&lt;img class="wp-image-16501 size-full aligncenter" src="../wp-archive/uploads/2020/12/funding-2.png" alt="" width="225" height="225" />&lt;/td>
&lt;td width="50%">
&lt;table>
&lt;tbody>
&lt;tr>
&lt;td>
&lt;h3>Company&lt;/h3>
&lt;/td>
&lt;td>&lt;a title="Raised" href="https://staging-www.netdata.cloud/news/netdata-extends-series-a-funding/">Raised&lt;/a> $14.2M, extending the total raised to $31M.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Recognized in &lt;a title="Forbes Cloud 100 Rising Stars 2020" href="https://staging-www.netdata.cloud/blog/forbes-cloud-100-rising-stars-2020/">Forbes Cloud 100 Rising Stars 2020&lt;/a>, &lt;a title="2020 Stratus Awards" href="https://staging-www.netdata.cloud/news/">2020 Stratus Awards&lt;/a> as Cloud Disruptor.&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&amp;nbsp;
&lt;p> &lt;/p></description></item><item><title>Centralize Infrastructure With Alarm Notifications</title><link>https://www.netdata.cloud/blog/cloud-alarm-notifications/</link><pubDate>Thu, 17 Dec 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/cloud-alarm-notifications/</guid><description>&lt;!--truncate-->
&lt;div class="et_pb_module et_pb_text et_pb_text_0 et_pb_text_align_left et_pb_bg_layout_light">
&lt;div class="et_pb_text_inner">
&lt;img class="alignnone wp-image-16509 size-large" src="../wp-archive/uploads/2020/12/Central-Alarm-Notifications-1200x828.png" alt="" width="1200" height="828" />
&lt;p>Netdata is architected on every level, across both the open-source Netdata Agent and Netdata Cloud, to help you own every layer of your monitoring experience. With this design, all metrics data collected by the Netdata Agent stays distributed on your node, but you also leverage Netdata Cloud’s dashboards and multi-node visualizations to view the health and performance of an entire infrastructure from a single application.&lt;/p></description></item><item><title>Community Update: Discourse, Community Efforts</title><link>https://www.netdata.cloud/blog/community-update-discourse/</link><pubDate>Wed, 02 Dec 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/community-update-discourse/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone wp-image-16522 size-full" src="../wp-archive/uploads/2022/03/Community-update_-Discourse-community-efforts.png" alt="" width="681" height="470" />
&lt;p>Open source and community have always been in the DNA of Netdata, with the Agent starting as a very popular open-source project. Since then, a lot has changed, with Netdata maturing into a company, and the Netdata Agent finding its place as an open-source project in a wider offering that redesigns the monitoring experience from the ground up.&lt;/p>
&lt;p>While we had a very active &lt;a href="https://github.com/netdata/netdata/" target="_blank" rel="noopener noreferrer">GitHub repository&lt;/a>, with the majority of the Netdata Agent’s original team actively moderating the discussions and talking with users, we started a more concentrated initiative to manage our community in spring 2020. We launched our first forum using &lt;a href="https://nodebb.org/" target="_blank" rel="noopener noreferrer">NodeBB&lt;/a>, a great open-source project, and the community grew substantially, outgrowing the forum software. At the same time, we were able to identify areas of friction in the community journey. Removing that friction became a centerpiece of our strategy during Q3 2020.&lt;/p></description></item><item><title>StackPulse: Automated Incident Remediation Workflow</title><link>https://www.netdata.cloud/blog/netdata-stackpulse-remediation/</link><pubDate>Wed, 02 Dec 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-stackpulse-remediation/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-large wp-image-16535" src="../wp-archive/uploads/2022/03/netdata-stackpulse-1200x826.png" alt="" width="1200" height="826" />
&lt;p>Teams of all types use Netdata to monitor the health of their nodes with preconfigured alarms and real-time interactive visualizations, and when incidents happen, they troubleshoot issues with thousands of per-second metrics on &lt;a href="https://staging-www.netdata.cloud/cloud/" target="_blank" rel="noopener noreferrer">Netdata Cloud&lt;/a>. But based on the complexity of the team and the infrastructure they monitor, some parts of their &lt;a href="https://staging-www.netdata.cloud/incident-management/" target="_blank" rel="noopener noreferrer">incident management&lt;/a>, such as pre-planned communication and escalation processes, or even automated remediation, need to happen outside of the Netdata ecosystem.&lt;/p></description></item><item><title>What is DevOps?</title><link>https://www.netdata.cloud/blog/what-is-devops/</link><pubDate>Thu, 19 Nov 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/what-is-devops/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-full wp-image-16550" src="../wp-archive/uploads/2022/03/what-is-devops.png" alt="" width="1024" height="600" />
&lt;p>In software development, it’s important to have a team dedicated to ensuring all systems and applications maintain maximum performance and uptime. Establishing processes that limit system and application slowdowns and outages while expediting the product release process is often done through a developer operations team, also known as &lt;em>dev ops&lt;/em> or &lt;em>DevOps&lt;/em>.&lt;/p>
&lt;p>DevOps teams are responsible for improving communication and collaboration across engineering teams to increase an organization’s ability to efficiently deliver products and services that serve the end-user. In this post, we’ll describe the different elements of DevOps, including what DevOps is, how it works, and how Netdata helps DevOps teams succeed.&lt;/p></description></item><item><title>Community Repository: Consul, Ansible &amp; ML Recipes</title><link>https://www.netdata.cloud/blog/welcome-to-netdatas-community-repository-consul-ansible-ml/</link><pubDate>Wed, 11 Nov 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/welcome-to-netdatas-community-repository-consul-ansible-ml/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-large wp-image-16555" src="../wp-archive/uploads/2022/03/netdata-community-repository-1200x825.png" alt="" width="1200" height="825" />
&lt;p>On our journey to democratize monitoring, we are proud to have open source at the core of both our products and our company values. What started as a project out of frustration for lack of existing alternatives (see &lt;a href="https://www.rexfeng.com/blog/2016/01/anger-driven-development/" target="_blank" rel="noopener noreferrer">anger-driven development&lt;/a>), quickly became one of the most starred open-source projects on all of GitHub.&lt;/p>
&lt;p>Fast-forward a couple of years later, and the Netdata Agent, our open-source monitoring agent, is maturing as the best single-node monitoring experience, offering unparalleled efficiency and thousands of metrics, per-second. At the same time, we have gathered a considerable community on our GitHub repository and new forums.&lt;/p></description></item><item><title>How Netdata gets you from 0 to monitoring in minutes</title><link>https://www.netdata.cloud/blog/how-netdata-gets-you-from-0-to-monitoring-in-minutes/</link><pubDate>Wed, 11 Nov 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/how-netdata-gets-you-from-0-to-monitoring-in-minutes/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-large wp-image-16562" src="../wp-archive/uploads/2022/03/0conf-1200x826.png" alt="" width="1200" height="826" />
&lt;p>Netdata is zero-configuration monitoring. It’s a principle that we’ve stood behind since the project’s beginning, when it was only our CEO Costa trying to solve a &lt;a href="https://staging-www.netdata.cloud/blog/why-netdata-is-free/">“painful, real-world problem,”&lt;/a> and it’s one we stand by today. Our insistence on zero-configuration guides every product decision we make, every grooming process, and every React component our frontend teams design.&lt;/p>
&lt;p>In fact, zero-configuration is the exact reason why Netdata’s dashboard is &lt;a href="https://staging-www.netdata.cloud/blog/netdata-agent-dashboard/">open and accessible by default&lt;/a>. It’s how Netdata gets you from 0 monitoring to thousands of metrics, collected every second and visualized in real time, in a matter of minutes.&lt;/p></description></item><item><title>Netdata’s dashboard: open by default and secure by design</title><link>https://www.netdata.cloud/blog/netdata-agent-dashboard/</link><pubDate>Wed, 28 Oct 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-agent-dashboard/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-large wp-image-16569" src="../wp-archive/uploads/2022/03/netdata-dashboard-open-secure-min-1200x826.png" alt="" width="1200" height="826" />
&lt;p>Let’s talk through a scenario: You have a Linux-based VM running on DigitalOcean (aka a Droplet), and you install Netdata on it using our &lt;a href="https://learn.netdata.cloud/docs/get#install-the-netdata-agent" target="_blank" rel="noopener noreferrer">recommended kickstart script&lt;/a>. As the installation process winds down, the Droplet starts up the Netdata Agent’s web server and serves the local Agent web dashboard on port 19999. You navigate to the dashboard using your browser of choice, check out per-second metrics updating in real time in a few of the hundreds of preconfigured visualizations, then realize…&lt;/p></description></item><item><title>Real-Time Infrastructure Monitoring Now In Netdata Cloud</title><link>https://www.netdata.cloud/blog/bringing-rich-and-real-time-infrastructure-monitoring-to-netdata-cloud/</link><pubDate>Thu, 22 Oct 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/bringing-rich-and-real-time-infrastructure-monitoring-to-netdata-cloud/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-large wp-image-16578" src="../wp-archive/uploads/2022/03/Correlation_charts-1200x830.png" alt="" width="1200" height="830" />
&lt;p>The Netdata Agent is well-equipped to solve monitoring and troubleshooting challenges for single nodes. We love that the Agent is so valuable to our users, but Netdata Cloud is designed for infrastructure monitoring. That’s why we’re working so hard to offer even more capabilities and help users monitor and troubleshoot infrastructures of all sizes, entirely for free!&lt;/p>
&lt;p>With the new Cloud Overview, you get every real-time chart and metric you need to understand the status of your infrastructure, explore, and troubleshoot, in a single view. We designed the Overview on one existing and beloved feature and another entirely new one that we’re very excited to launch for the first time.&lt;/p></description></item><item><title>The reality of Netdata’s long-term metrics storage database</title><link>https://www.netdata.cloud/blog/the-reality-of-netdatas-long-term-metrics-storage-database/</link><pubDate>Mon, 12 Oct 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/the-reality-of-netdatas-long-term-metrics-storage-database/</guid><description>&lt;!--truncate-->
&lt;p>The perception that Netdata is only capable of short-term metrics storage is a myth. It’s a pervasive myth we still see in blog posts and through community engagement, despite it being false for more than a year.&lt;/p>
&lt;p> &lt;/p>
&lt;p>However, like all myths, this one on metrics storage began with a kernel of truth. When Netdata first flourished as an &lt;a title="open-source project" href="https://github.com/netdata/netdata" target="_blank" rel="noopener noreferrer">open-source project&lt;/a> in 2017 and 2018, the default metrics database was RAM-only. You could configure this database’s size, but for many users, that size was limited by the amount of RAM they were willing to allocate for metrics storage. We also kept the default value low to ensure Netdata worked efficiently on all hardware and a variety of operating systems.&lt;/p></description></item><item><title>Software Extensibility Is Key To Adoption</title><link>https://www.netdata.cloud/blog/software-extensibility-is-key-to-adoption/</link><pubDate>Fri, 25 Sep 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/software-extensibility-is-key-to-adoption/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-full wp-image-16613" src="../wp-archive/uploads/2022/03/Software-Extensibility-Blog-Post.png" alt="" width="969" height="638" />
&lt;p>As with most commercial products, software today is mass produced for reasons of simple economics. But making the same software work for as many people as possible while also meeting the unique requirements different people and organizations have is a challenging task.&lt;/p>
&lt;p>The strategies used today to provide extensibility at cost and at the level required by various enterprises are very similar to the ones used in the 80s and 90s, when PCs first took off and captured an enormous amount of the home computing market, and when open source software started gaining popularity.&lt;/p></description></item><item><title>Investing in Netdata: a growth story</title><link>https://www.netdata.cloud/blog/investing-in-netdata-a-growth-story/</link><pubDate>Tue, 22 Sep 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/investing-in-netdata-a-growth-story/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-large wp-image-16618" src="../wp-archive/uploads/2022/03/Series-A-funding-extension-1200x571.png" alt="" width="1200" height="571" />
&lt;p>I’m excited to announce an &lt;a title="extension to Netdata’s series A funding" href="https://staging-www.netdata.cloud/news/netdata-extends-series-a-funding/" target="_blank" rel="noopener noreferrer">extension to Netdata’s series A funding &lt;/a>in the amount of $14.2M, bringing the total amount of funding to $31M. We’re thrilled to share the news; the additional funding will help us continue building the future of health monitoring and performance troubleshooting. In case you missed it, our mission is to &lt;a title="redefine infrastructure monitoring" href="https://staging-www.netdata.cloud/blog/redefining-monitoring-netdata/" target="_blank" rel="noopener noreferrer">redefine infrastructure monitoring&lt;/a>. Our unique approach to building the right solution with and for the community is no easy task.&lt;/p></description></item><item><title>Metric Correlations: Detect Patterns &amp; Anomalies</title><link>https://www.netdata.cloud/blog/netdata-cloud-metric-correlations/</link><pubDate>Wed, 16 Sep 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-cloud-metric-correlations/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-large wp-image-16623" src="../wp-archive/uploads/2022/03/Cloud-Correlations@2x-1200x826.png" alt="" width="1200" height="826" />
&lt;p>Today, we are excited to launch our first Netdata Cloud Insights feature, Metric Correlations, developed for discovering underlying issues more quickly and identifying the root cause more efficiently. Read on to learn more about our approach to developing this new feature, how it works, and the many benefits you’ll find incorporating this into your team’s troubleshooting workflow.&lt;/p>
&lt;h2>Some background&lt;/h2>
Let’s start with a bit of a disclaimer. It seems machine learning (ML) (or “Artificial Intelligence,” if you are looking for more LinkedIn likes) has gone mainstream in the last few years, and we are probably by now somewhere near the “Peak of Inflated Expectations” on the &lt;a title="hype cycle" href="https://en.wikipedia.org/wiki/Hype_cycle" target="_blank" rel="noopener noreferrer">hype cycle&lt;/a>. It is in this context that we want to be clear about what our goals are in this space and our approach to releasing data-driven features that draw on techniques from statistics and ML. In short, we want to be clear, open, realistic, and avoid buzzwords at all costs!
&lt;p>Over the next 12 months, we are hoping to begin building a layer of intelligence&lt;sup>&lt;a href="https://staging-www.netdata.cloud/blog/netdata-cloud-metric-correlations/#1">1&lt;/a>&lt;/sup> throughout Netdata (both Cloud and Agent) to assist with “&lt;a title="human in the loop" href="https://hai.stanford.edu/blog/humans-loop-design-interactive-ai-systems" target="_blank" rel="noopener noreferrer">human in the loop&lt;/a>” troubleshooting, mainly to help users more easily surface slowdowns, anomalies, or other issues and lower your &lt;a title="cognitive load" href="https://en.wikipedia.org/wiki/Cognitive_load" target="_blank" rel="noopener noreferrer">cognitive load&lt;/a>&lt;sup>&lt;a href="https://staging-www.netdata.cloud/blog/netdata-cloud-metric-correlations/#2">2&lt;/a>&lt;/sup> as you troubleshoot using Netdata. Simply put, we’re working to streamline your mean time to resolution (MTTR).&lt;/p></description></item><item><title>Netdata Agent v1.25 &amp; Cloud Enhancements</title><link>https://www.netdata.cloud/blog/release-1-25/</link><pubDate>Wed, 16 Sep 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/release-1-25/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-large wp-image-16635" src="../wp-archive/uploads/2022/03/Agent-Release-v1.25@2x-1200x600.png" alt="" width="1200" height="600" />
&lt;p>The &lt;a title="v1.25.0" href="https://github.com/netdata/netdata/releases" target="_blank" rel="noopener noreferrer">v1.25.0&lt;/a> release of the Netdata Agent delivers on our commitment to make our metrics collection, visualization, and troubleshooting platform more stable and usable. We enhanced our recently-added &lt;a title="Prometheus collector" href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/prometheus" target="_blank" rel="noopener noreferrer">Prometheus collector&lt;/a> with user-configurable filtering and grouping, made dramatic improvements to the reliability of the Agent-Cloud link that streams metrics on-demand to your browser when you use &lt;a title="Netdata Cloud" href="https://app.netdata.cloud/" target="_blank" rel="noopener noreferrer">Netdata Cloud&lt;/a>, and more.&lt;/p>
&lt;p>Let’s jump in and look at each improvement.&lt;/p></description></item><item><title>Netdata named to the Forbes Cloud 100 Rising Stars</title><link>https://www.netdata.cloud/blog/forbes-cloud-100-rising-stars-2020/</link><pubDate>Wed, 16 Sep 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/forbes-cloud-100-rising-stars-2020/</guid><description>&lt;!--truncate-->
&lt;img class=" wp-image-16648 alignleft" src="../wp-archive/uploads/2022/03/Cloud1002020-RisingStars-SMALL.png" alt="" width="390" height="471" />
We’re excited to announce that we’ve been named to the&lt;strong> &lt;a title="Forbes 2020 Cloud 100 Rising Stars" href="https://www.forbes.com/sites/kenrickcai/2020/09/16/cloud-100-rising-stars-2020/" target="_blank" rel="noopener noreferrer">Forbes 2020 Cloud 100 Rising Stars&lt;/a>.&lt;/strong> This is a list of the top 100 private cloud companies in the world, published by Forbes in partnership with &lt;a title="Bessemer Venture Partners" href="https://www.bvp.com/" target="_blank" rel="noopener noreferrer">Bessemer Venture Partners&lt;/a> and &lt;a title="Salesforce Ventures" href="https://www.salesforce.com/company/ventures/" target="_blank" rel="noopener noreferrer">Salesforce Ventures&lt;/a>. The 20 Rising Stars represent young, high-growth and category-leading cloud companies who are poised to join the &lt;a title="Cloud 100" href="https://www.forbes.com/cloud100/" target="_blank" rel="noopener noreferrer">Cloud 100&lt;/a> ranks.
We are extremely honored to be recognized amongst our most-promising peers. This is a testament to our momentum we’ve built in partnership with our passionate community of users and contributors worldwide, as well as a validation of our community-first, open-source approach to democratizing infrastructure monitoring and troubleshooting.
The major milestones we’ve hit along the way this year include more than 5,000 new users added each day, more than 3 million users total worldwide, and nearly 50,000 GitHub stars, making Netdata the fourth most starred project in the Cloud Native Computing Foundation landscape. We’ve also launched our monitoring service, Netdata Cloud, which has seen very strong growth since its introduction in May, with more than 15,000 users registered so far. We couldn’t have done it without you!
The Forbes 2020 &lt;a title="Cloud 100" href="https://www.forbes.com/cloud100" target="_blank" rel="noopener noreferrer">Cloud 100&lt;/a> and 20 Rising Stars lists will also appear in the September 2020 issue of &lt;em>Forbes&lt;/em> magazine.</description></item><item><title>The Netdata Community Powered by NodeBB</title><link>https://www.netdata.cloud/blog/the-netdata-community-powered-by-nodebb/</link><pubDate>Thu, 13 Aug 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/the-netdata-community-powered-by-nodebb/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-large wp-image-16658" src="../wp-archive/uploads/2022/03/NodeBB-1-1200x877.png" alt="" width="1200" height="877" />
&lt;p>We recently adopted &lt;a title="NodeBB" href="https://nodebb.org/" target="_blank" rel="noopener noreferrer">NodeBB&lt;/a> as our software of choice for building &lt;a title="the Netdata Community" href="https://community.netdata.cloud/" target="_blank" rel="noopener noreferrer">the Netdata Community&lt;/a>. We have &lt;a title="many good reasons" href="https://staging-www.netdata.cloud/blog/the-netdata-community/" target="_blank" rel="noopener noreferrer">many good reasons&lt;/a> for why we wanted to provide our community with a proper home online, but I wanted to cover some of the technical reasons for choosing NodeBB for our platform, and the many parallels between the NodeBB and Netdata projects, which was certainly a driving force behind this decision.&lt;/p></description></item><item><title>Release 1.24: Prometheus Collector &amp; Multi-Host DB</title><link>https://www.netdata.cloud/blog/release-1-24/</link><pubDate>Mon, 10 Aug 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/release-1-24/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-full wp-image-16663" src="../wp-archive/uploads/2022/03/1.24-release-2.png" alt="" width="683" height="469" />
&lt;p>The v1.24.0 release of the Netdata Agent brings enhancements to the breadth of metrics we collect with a new Prometheus/OpenMetrics collector and enhanced storage and querying with a new multi-host database mode. Let’s take a look at each of these enhancements.&lt;/p>
&lt;h2>Instantly access thousands of metrics with the new Prometheus/OpenMetrics collector&lt;/h2>
This release broadens our commitment to open standards, interoperability, and extensibility with a new generic Prometheus collector that works seamlessly with any application that makes its metrics available in the &lt;a href="https://prometheus.io/docs/instrumenting/exposition_formats/">Prometheus&lt;/a>/&lt;a href="https://github.com/OpenObservability/OpenMetrics">OpenMetrics&lt;/a> exposition format, including support for Windows 10 via &lt;a href="https://github.com/prometheus-community/windows_exporter">windows_exporter&lt;/a>. Netdata will autodetect &lt;a href="https://github.com/netdata/go.d.plugin/blob/master/config/go.d/prometheus.conf">over 600 Prometheus endpoints&lt;/a> and instantly generate charts with all the exposed metrics, meaningfully visualized.
&lt;p>You can also quickly and easily configure the collector with the names and URLs of additional Prometheus endpoints to instantly view automatically generated charts with all the exposed metrics, meaningfully visualized within Netdata at the same high-granularity, per-second frequency you expect, all in real time. To learn more about how to configure, check out our &lt;a title="documentation" href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/prometheus" target="_blank" rel="noopener noreferrer">documentation&lt;/a>.&lt;/p></description></item><item><title>Introducing the all-new Netdata Cloud</title><link>https://www.netdata.cloud/blog/introducing-the-all-new-netdata-cloud/</link><pubDate>Wed, 29 Jul 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/introducing-the-all-new-netdata-cloud/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-large wp-image-16672" src="../wp-archive/uploads/2022/03/All-New-Cloud-1200x712.png" alt="" width="1200" height="712" />
&lt;p>In case you missed it, we released an all-new version of Netdata Cloud in May. &lt;a title="Netdata Cloud" href="https://staging-www.netdata.cloud/cloud/">Netdata Cloud&lt;/a> is a free service that can be accessed from any browser and provides you with a consolidated view of your entire infrastructure.&lt;/p>
&lt;p>Netdata Cloud works differently from other monitoring solutions. Most solutions limit the number and frequency of metrics because they rely on architectures that aggregate data. Netdata Cloud, however, streams limited metadata from each node running the Netdata Agent, keeping you in control of the data on your systems. The advantage of this architecture is that there is &lt;strong>no limit&lt;/strong> to the number or frequency of metrics, regardless of the scale or complexity of your IT infrastructure. You can truly monitor every metric, from every system and application, across your entire infrastructure, in real time. For free.&lt;/p></description></item><item><title>Sysadmin Day 2020: IT Heroes and Homelabs</title><link>https://www.netdata.cloud/blog/sysadmin-day-2020-it-heroes-and-homelabs/</link><pubDate>Mon, 27 Jul 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/sysadmin-day-2020-it-heroes-and-homelabs/</guid><description>&lt;!--truncate-->
&lt;p dir="ltr" lang="en">&lt;img class="alignnone size-large wp-image-16685" src="../wp-archive/uploads/2022/03/Group-1-2-1200x825.png" alt="" width="1200" height="825" />&lt;/p>
&lt;p dir="ltr" lang="en">Sysadmin Day 2020 is right around the corner and we’d like to show our appreciation for all the sysadmins out there who keep IT humming along and come to the rescue to resolve critical issues day in and day out. This year, we’re celebrating all week long by hosting an IT Heroes and Homelabs contest. &lt;strong>Join the celebration by retweeting our post with the hashtag #SysadminDay #NetdataWin, and we’ll enter you in a drawing to win some Netdata swag!&lt;/strong>&lt;/p></description></item><item><title>The Netdata Community</title><link>https://www.netdata.cloud/blog/the-netdata-community/</link><pubDate>Mon, 27 Jul 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/the-netdata-community/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-large wp-image-16700" src="../wp-archive/uploads/2022/03/Athens-company-meetup-1_2019-scaled-e1588700274782-1200x747.jpeg" alt="" width="1200" height="747" />
&lt;p>Netdata users and contributors comprise a large, global, but somewhat fragmented community – or a set of communities. You can find us on IRC (#netdata on freenode), on &lt;a href="https://www.reddit.com/r/netdata/">Reddit&lt;/a>, on &lt;a href="https://twitter.com/linuxnetdata">social media&lt;/a>, and, of course, on &lt;a href="https://github.com/netdata/netdata">GitHub&lt;/a>, where the main open-source Netdata project repo lives. And yes, you can find us on other platforms as well.&lt;/p>
&lt;p> &lt;/p>
&lt;p>GitHub is a great way to get in touch with the Netdata team and project contributors to tell them about bugs or to discuss new features. But we realized that bug reports are not a conversation that works for everyone. A more informal and easier-to-access communication channel would provide a better way for the community to congregate and engage, and would also provide an easier way for everybody to talk to the team. We wanted to provide the community a home.&lt;/p></description></item><item><title>Why Netdata picked VerneMQ</title><link>https://www.netdata.cloud/blog/why-netdata-picked-vernemq/</link><pubDate>Tue, 14 Jul 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/why-netdata-picked-vernemq/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-full wp-image-16705" src="../wp-archive/uploads/2022/03/Blog-Why-Netdata-Picked-VerneMQ.jpeg" alt="" width="683" height="470" />
&lt;p>In 2019, the Netdata team already knew that a Netdata Cloud solution in the form of an online platform would greatly complement Netdata’s distributed monitoring by making it much easier to organize large infrastructures and by enabling new ways for teams to collaborate. The old node registry available at the time wasn’t enough for Netdata’s users.&lt;/p>
&lt;p>Building an online platform, even one that does not directly process users’ metrics, is challenging. But less challenging than it was even a few years ago, since the technology stack has improved greatly over the years.&lt;/p></description></item><item><title>What is Infrastructure Monitoring?</title><link>https://www.netdata.cloud/blog/what-is-infrastructure-monitoring/</link><pubDate>Tue, 30 Jun 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/what-is-infrastructure-monitoring/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-medium wp-image-16343" src="../wp-archive/uploads/2022/03/Blog-What_is_Infrastructure_Monitoring_Header-600x450.png" alt="" width="600" height="450" />
&lt;p>IT is advancing blazingly fast. To keep up with architectural changes and hybrid environments, it’s more important than ever to maintain efficient infrastructure monitoring and troubleshooting. Adding to the complexity is the increase of distributed systems, comprised of many components and services. For IT teams to effectively manage monitoring modern infrastructure, it’s necessary to have the right practices and tools in place that enable teams to do their jobs as quickly as possible with fewer resources.&lt;/p></description></item><item><title>Release 1.23: Kubernetes &amp; eBPF Observability</title><link>https://www.netdata.cloud/blog/release-1-23/</link><pubDate>Thu, 25 Jun 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/release-1-23/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-large wp-image-16712" src="../wp-archive/uploads/2022/03/Agent-1.23-release-1200x900.png" alt="" width="1200" height="900" />
&lt;p>As adoption of container infrastructure grows in popularity, we’re continuing to focus on the most effective ways users can quickly and easily deploy container monitoring to instantly get access to deep, real-time insights. Agent release 1.23 introduces service discovery for Kubernetes clusters, monitoring for individual nodes, and eBPF monitoring per application on an event frequency for quickly identifying the root cause. Quickly shed light on your infrastructure performance with these new features!&lt;/p></description></item><item><title>The role of shift-left testing in an agile environment</title><link>https://www.netdata.cloud/blog/agile-static-analysis/</link><pubDate>Tue, 14 Apr 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/agile-static-analysis/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-large wp-image-16723" src="../wp-archive/uploads/2022/03/Netdata-Security-use-with-Static-Analysis-1200x899.png" alt="" width="1200" height="899" />
&lt;p>With the rapid growth of security threats to infrastructure, it’s more important than ever to proactively address vulnerabilities. As an open-source project, built on the trust of users and contributors, Netdata has security concerns at its core.&lt;/p>
&lt;p>Because we’re committed to code security and quality, we apply &lt;a href="https://agilemanifesto.org/">Agile principles&lt;/a> throughout the software development process. A component of this includes regular static analysis. Through continuous, automated testing, we’re able to move quickly to keep up with end-user requests without compromising our source code.&lt;/p></description></item><item><title>Release 1.21: New Collectors &amp; Faster Exporters</title><link>https://www.netdata.cloud/blog/release-1-21/</link><pubDate>Mon, 06 Apr 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/release-1-21/</guid><description>&lt;!--truncate-->
&lt;div class="et_pb_module et_pb_text et_pb_text_0 et_pb_text_align_left et_pb_bg_layout_light">
&lt;div class="et_pb_text_inner">
&lt;img class="alignnone size-full wp-image-16737" src="../wp-archive/uploads/2022/03/release-1.21.0.png" alt="" width="1200" height="600" />
&lt;p>We’re in the middle of a scary, uncertain time, and we hope those of you reading are staying safe and healthy.&lt;/p>
&lt;p>Despite the current challenges, the 40+ members of the &lt;a title="Netdata Remote Working" href="https://staging-www.netdata.cloud/blog/culture/netdata-remote-working/">remote-first Netdata&lt;/a> team have been hard at work on the next version of the Netdata Agent: v1.21.0.&lt;/p>
&lt;p>This release is foundational: While we do have fantastic new collectors and three new ways to export your metrics for long-term storage, many of the most significant changes aren’t even those you’ll notice. While they may be beneath the hood, they’re going to power some amazing new features, UX improvements, and design overhauls.&lt;/p></description></item><item><title>Creating A Thriving, Agile, Remote Team</title><link>https://www.netdata.cloud/blog/netdata-remote-working/</link><pubDate>Wed, 25 Mar 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-remote-working/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-large wp-image-16750" src="../wp-archive/uploads/2022/03/netdata-remote-working_01-1200x899.png" alt="" width="1200" height="899" />
&lt;p>The coronavirus (COVID-19) pandemic has forced many organizations to take unprecedented steps towards remote working. As a fully distributed team, we’ve faced the common challenges of remote work. Based on our experience from our very beginning in 2018, all but a few of these organizations new to remote working will face hurdles to overcome and may try to revert to colocation as soon as possible. Remote working is hard, even when it’s carefully planned and executed. When the transition is rushed and seen as a necessary, temporary inconvenience, challenges are all but inevitable.&lt;/p></description></item><item><title>The Netdata Culture and People</title><link>https://www.netdata.cloud/blog/netdata-culture-people/</link><pubDate>Mon, 23 Mar 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/netdata-culture-people/</guid><description>&lt;!--truncate-->
&lt;p>&lt;img src="../wp-archive/uploads/2022/03/people-culture_01.png" alt="">&lt;/p>
&lt;p>There are many things I absolutely love about Netdata, but I’m most proud of our people and culture. Some words about this unique experience are long overdue.&lt;/p>
&lt;p>In a career that spans over two decades and six other companies of various sizes, nothing compares to the satisfaction of working in a company like ours. My answer to the canned interview question, “Where do you see yourself in 5 years”, was always the same: I don’t care; I just want to be solving problems and working with good people, real professionals, who I can trust and respect. In retrospect, I was missing another huge part of the equation, which is to mention the kind of company I wanted to work for. “Culture eats strategy for breakfast” is a cliche. More importantly, bad culture devours people’s souls; it sucks out any creative energy one may have, reducing engagement and, therefore, productivity. Short-term wins at the expense of company culture guarantee huge losses in the long run.&lt;/p></description></item><item><title>Contribute to Netdata’s machine learning efforts!</title><link>https://www.netdata.cloud/blog/contribute-machine-learning/</link><pubDate>Mon, 16 Mar 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/contribute-machine-learning/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-full wp-image-16783" src="../wp-archive/uploads/2022/03/contribute-machine-learning.png" alt="" width="991" height="1072" />
&lt;p>Netdata contributors have greatly influenced the growth of our company and are essential to our success. The time and expertise that contributors volunteer are fundamental to our goal of helping you build extraordinary infrastructures. We highly value end-user feedback during product development, which is why we’re looking to involve you in progressing our machine learning (ML) efforts! &lt;span id="more-2975">&lt;/span>As we are continually looking for ways to improve and enhance Netdata, we are starting to explore how we can leverage machine learning to introduce new product features. Our main focus at the moment is around automated &lt;a href="https://en.wikipedia.org/wiki/Anomaly_detection">anomaly detection&lt;/a>. This is a really interesting and challenging problem (high volume, high dimensional data, lack of ground truth labels, and so on), but we should be able to use some of the metrics monitored by Netdata to deliver new, awesome product features and user experiences (AI is the &lt;a href="https://www.gsb.stanford.edu/insights/andrew-ng-why-ai-new-electricity">new electricity&lt;/a>, after all 😃). However, developing ML-driven product features is quite different than traditional software development (see steps 1 to 7 in the picture above). Mainly, this is because you never really know what specific data transformations, problem formulation, and sets of algorithms will work best in advance. (&lt;a href="https://www.kdnuggets.com/2019/09/no-free-lunch-data-science.html">Here&lt;/a> is a good article explaining things, and if you really want to go down a rabbit hole, check out this &lt;a href="https://ai.stackexchange.com/questions/15650/what-are-the-implications-of-the-no-free-lunch-theorem-for-machine-learning">Stack Overflow question&lt;/a> and this &lt;a href="https://www.quora.com/What-does-the-No-Free-Lunch-theorem-mean-for-machine-learning-In-what-ways-do-popular-ML-algorithms-overcome-the-limitations-set-by-this-theorem">Quora thread&lt;/a>). Ideally, you first need to prototype your solution “in the lab” on some data you have already collected and do a few iterations of data → problem formulation → prototype. This process gives you a level of confidence in what you are doing (and some data to back it up) to move on to the even-more-complicated step of going from prototype to production. At Netdata, we are currently trying to get to step 4, where we can first prototype some solutions on real-world data and come up with ways to measure progress.&lt;/p></description></item><item><title>Linux eBPF monitoring with Netdata</title><link>https://www.netdata.cloud/blog/linux-ebpf-monitoring-with-netdata/</link><pubDate>Fri, 21 Feb 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/linux-ebpf-monitoring-with-netdata/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-full wp-image-16799" src="../wp-archive/uploads/2022/03/linux-ebpf-monitoring-netdata.png" alt="" width="1200" height="600" />
&lt;p>Your application isn’t finished when you’ve closed the last &lt;code>if&lt;/code> block and you lined up all the brackets. There’s a whole other world of testing, debugging, and optimization that you haven’t even touched yet.&lt;/p>
&lt;p>To help you more safely step into that complex phase of making your application &lt;em>even better&lt;/em>, we’ve just released a brand-new eBPF collector in &lt;a href="https://staging-www.netdata.cloud/blog/product/release-1.20/">v1.20 of Netdata&lt;/a>. With this collector enabled, you can monitor real-time metrics of Linux kernel functions and actions from the very same monitoring and troubleshooting dashboard you use for watching entire systems, or even entire infrastructures.&lt;/p></description></item><item><title>Release 1.20: Kernel Monitoring &amp; Infra-Wide Labels</title><link>https://www.netdata.cloud/blog/release-1-20/</link><pubDate>Fri, 21 Feb 2020 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/release-1-20/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-full wp-image-16790" src="../wp-archive/uploads/2022/03/release-1.20.0.png" alt="" width="1200" height="600" />
&lt;p>In Netdata’s first major release of 2020, we’re introducing two new features on the opposite ends of the monitoring spectrum.&lt;/p>
&lt;p>On one hand, we’re releasing an eBPF collector, which lets you collect, monitor, and visualize incredibly precise metrics straight from the Linux kernel. On the other, we added the ability to label agents to help you organize entire infrastructures and see &lt;em>every&lt;/em> important piece of information about streaming nodes in one place.&lt;/p></description></item><item><title>Redefining monitoring with Netdata (and how it came to be)</title><link>https://www.netdata.cloud/blog/redefining-monitoring-with-netdata/</link><pubDate>Thu, 19 Dec 2019 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/redefining-monitoring-with-netdata/</guid><description>&lt;!--truncate-->
&lt;p>&lt;img src="../wp-archive/uploads/2019/12/redefining-monitoring-netdata_01.png" alt="">&lt;/p>
&lt;h2 id="how-netdata-was-born">How Netdata was born&lt;/h2>
&lt;p>In 2013, I worked for a company that relied on financial transactions. We had a very simple SLA: complete all financial transactions within 3 seconds.&lt;/p>
&lt;p>We were migrating the infrastructure from colocated (physical servers) to the cloud (VMs). The transition was not smooth. We had a lot of issues on the cloud side, which we couldn’t even detect. Business metrics were randomly reporting significant loss of volume and a very bad SLA, but at the operational level we saw no issues—everything seemed to be working perfectly. Traces were showing a large delay in several transactions, but there were no failures.&lt;/p></description></item><item><title>Release 1.19: Web Log Parsing &amp; Unit Testing</title><link>https://www.netdata.cloud/blog/release-1-19/</link><pubDate>Wed, 27 Nov 2019 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/release-1-19/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-full wp-image-16837" src="../wp-archive/uploads/2022/03/release-1.19.0.png" alt="" width="1200" height="600" />
&lt;p>Network monitoring is complex, which is why we’re developing a monitoring tool that will drastically increase DevOps productivity. This release is all about improving Netdata’s day-in, day-out performance. We’re working hard to make deploy enhancements that help engineers make faster, smarter decisions about their systems.&lt;/p>
&lt;p> &lt;/p>
&lt;p>v1.19 of Netdata delivers a vastly improved way to collect, parse, and understand the health and performance of any service or application that runs through an Apache or Nginx web server.&lt;/p></description></item><item><title>Agile Team Safety Harness With cmocka &amp; FOSS</title><link>https://www.netdata.cloud/blog/agile-team-cmocka-foss/</link><pubDate>Tue, 26 Nov 2019 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/agile-team-cmocka-foss/</guid><description>&lt;!--truncate-->
&lt;p>Netdata is made up from agile teams who are deeply committed to improving the usability of our product. We want to respond to our users and introduce in-demand features. Working directly with our community is the best way to make Netdata better.&lt;/p>
&lt;p> &lt;/p>
&lt;p>But we face the same the dilemma as all agile teams: &lt;strong>How do we do this safely?&lt;/strong>&lt;/p>
&lt;p>Safety means that we can move quickly without compromising the quality of our code. Because we want to move quickly, engage with our users’ desires, and keep quality high, we’re becoming very serious about adopting unit testing in our work.&lt;/p></description></item><item><title>Release 1.18: What’s new with the database engine?</title><link>https://www.netdata.cloud/blog/release-1-18/</link><pubDate>Sat, 19 Oct 2019 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/release-1-18/</guid><description>&lt;!--truncate-->
&lt;img class="alignnone size-full wp-image-16859" src="../wp-archive/uploads/2022/03/release-1.18.0.png" alt="" width="1200" height="600" />
&lt;p>As your infrastructure grows more complex, storing long-term metrics becomes difficult and costly to retain. Your team stars to limit the amount of historical data they archive, causing gaps in coverage. Anomalies start to slip through the cracks.&lt;/p>
&lt;p>Version 1.18 of Netdata aims to solve the monitoring metrics storage problem once and for all.&lt;/p>
&lt;p>Aside from 5 new collectors, 16 bug fixes, 27 improvements, and 20 documentation updates, here’s what you need to know.&lt;/p></description></item><item><title>Release 1.17: Collection frequency gets flexible</title><link>https://www.netdata.cloud/blog/release-1-17/</link><pubDate>Mon, 09 Sep 2019 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/release-1-17/</guid><description>&lt;!--truncate-->
&lt;p>The next version of Netdata has arrived! Aside from dozens of quality-of-life and papercut fixes, we’ve launched some new features we know you’ll be excited to use straight away.&lt;/p>
&lt;p>Let’s dive in.&lt;/p>
&lt;h2>What’s new?&lt;/h2>
Release v1.17.0 contains 38 bug fixes, 33 improvements, and 20 documentation updates.
&lt;p>You can, of course, view the full list at the &lt;a href="https://github.com/netdata/netdata/releases/tag/v1.17.0">v1.17.0 release notes&lt;/a> on GitHub. But, let’s talk details on a few of the improvements and changes most requested by the Netdata community.&lt;/p></description></item><item><title>How and why we’re bringing long-term storage to Netdata</title><link>https://www.netdata.cloud/blog/db-engine/</link><pubDate>Wed, 07 Aug 2019 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/db-engine/</guid><description>&lt;!--truncate-->
&lt;p>We’ve built a lot of amazing things into the open-source &lt;a href="https://github.com/netdata/netdata">Netdata&lt;/a> monitoring system. But, no matter how far we’ve come, we’ll always be proud of how little RAM it uses.&lt;/p>
&lt;p>Right now, Netdata stores metrics in your system’s RAM using a ridiculously efficient database. It only saves or loads historical metrics from disk when you restart it. With this system, Netdata can be both low-resource and exhaustive in its collection of real-time metrics.&lt;/p></description></item><item><title>Release 1.16.0: Smarter binaries and built-in TLS</title><link>https://www.netdata.cloud/blog/release-1-16/</link><pubDate>Fri, 19 Jul 2019 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/release-1-16/</guid><description>&lt;!--truncate-->
&lt;p>We’re excited to launch release v1.16.0 of the open-source &lt;a href="https://github.com/netdata/netdata/">Netdata monitoring agent&lt;/a>, which delivers real-time health monitoring and performance troubleshooting to nearly any system or application.&lt;/p>
&lt;p>This release also contains 40 bug fixes, 31 improvements, and 20 documentation updates—if you’d like to see the full list, check out the &lt;a href="https://github.com/netdata/netdata/releases/tag/v1.16.0">full release notes&lt;/a>.&lt;/p>
&lt;p>Details aside, I know people are going to be most curious about the big changes we’ve just delivered to Netdata—let’s dive in.&lt;/p></description></item><item><title>Open Source Contributions: Supporting The Community</title><link>https://www.netdata.cloud/blog/open-source-contributions/</link><pubDate>Tue, 02 Jul 2019 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/blog/open-source-contributions/</guid><description>&lt;!--truncate-->
&lt;p>Netdata &lt;em>must&lt;/em> be doing something right when it comes to inspiring contributions. Our &lt;a href="https://github.com/netdata/netdata">open-source, distributed monitoring agent&lt;/a> has &lt;img src="https://img.shields.io/github/stars/netdata/netdata.svg" alt="GitHub stars" /> on GitHub and has seen contributions from hundreds of people: &lt;img src="https://img.shields.io/github/contributors/netdata/netdata.svg" alt="GitHub contributors" />. We’ve even hired a handful of our contributors to work full-time on making the Netdata ecosystem even more powerful.&lt;/p>
&lt;p> &lt;/p>
&lt;p>The community is passionate about what we’re building, and they’re actively interested in making it work better for their particular needs.&lt;/p></description></item><item><title/><link>https://www.netdata.cloud/devopsdaysgeneva/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/devopsdaysgeneva/</guid><description/></item><item><title/><link>https://www.netdata.cloud/value/calculator/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/value/calculator/</guid><description>&lt;p>// ============================================
// NETDATA ROI CALCULATOR - LOGIC &amp;amp; VARIABLES
// ============================================&lt;/p>
&lt;p>// INPUT VARIABLES
// ============================================
hosts = integer (10 to 10,000+) // Number of monitored hosts/nodes
mttr = integer (5 to 240 minutes) // Current average MTTR for critical incidents
incidents = integer (1 to 50+) // Number of critical incidents per month
currentCosts = integer ($0 to $1M+) // Annual monitoring tool costs
teamSize = integer (5 to 200+) // Size of engineering/DevOps team
engineerCost = integer ($50 to $200+) // Average hourly cost per engineer
impactGroup = enum [&amp;ldquo;engineering&amp;rdquo;, &amp;ldquo;customers&amp;rdquo;, &amp;ldquo;both&amp;rdquo;] // Who is impacted by downtime
revenueImpact = integer ($0 to $1M+) // Revenue loss per hour of downtime&lt;/p></description></item><item><title>.NET Framework Monitoring</title><link>https://www.netdata.cloud/monitoring-101/dotnet-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/dotnet-monitoring/</guid><description>&lt;h2 id="net-framework">.NET Framework&lt;/h2>
&lt;p>.NET Framework is a software development platform that runs on Windows. &lt;a href="https://dotnet.microsoft.com/en-us/download/dotnet-framework">It provides a common set of libraries and tools that developers can use to build different types of applications, such as web apps, desktop apps, mobile apps, games, and more&lt;/a>.&lt;/p>
&lt;h2 id="net-framework-monitoring">.NET Framework Monitoring&lt;/h2>
&lt;p>One of the challenges of developing and maintaining .NET applications is to ensure their performance, reliability, and security. To do that, developers need to monitor various aspects of their applications, such as code execution, memory usage, exceptions, requests, errors, and more. Monitoring .NET applications can help developers identify and troubleshoot issues before they affect the end users.&lt;/p></description></item><item><title>1-Wire Sensors</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/1-wire-sensors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/1-wire-sensors/</guid><description/></item><item><title>1-Wire Sensors Monitoring</title><link>https://www.netdata.cloud/monitoring-101/1wiresensors-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/1wiresensors-monitoring/</guid><description>&lt;h2 id="what-are-a-1-wire-sensors">What are a 1-Wire Sensors?&lt;/h2>
&lt;p>These sensors monitor temperature. On Linux these are supported by the wire, w1_gpio, and w1_therm modules. Currently temperature sensors are supported and automatically detected.&lt;/p>
&lt;h2 id="monitoring-1-wire-sensors-with-netdata">Monitoring 1-Wire Sensors with Netdata&lt;/h2>
&lt;p>Netdata auto discovers hundreds of services, and for those it doesn&amp;rsquo;t turning on manual discovery is a one line configuration. For more information on configuring Netdata for 1-Wiresensors monitoring please read the collector &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/python.d.plugin/w1sensor/">documentation&lt;/a>.&lt;/p>
&lt;p>Netdata has a public &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/">demo space&lt;/a> (no login required) where you can explore different monitoring use-cases and get a feel for Netdata.&lt;/p></description></item><item><title>1-Wire Sensors Monitoring</title><link>https://www.netdata.cloud/monitoring-101/w1sensor-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/w1sensor-monitoring/</guid><description>&lt;h2 id="1-wire-sensors-monitoring">1-Wire Sensors Monitoring&lt;/h2>
&lt;h3 id="what-is-1-wire-sensors">What Is 1-Wire Sensors?&lt;/h3>
&lt;p>1-Wire Sensors are a type of sensor technology used primarily for temperature monitoring through a simple communication protocol. These sensors are commonly employed in environments where real-time temperature tracking is crucial, such as server rooms, laboratories, or manufacturing lines.&lt;/p>
&lt;h3 id="monitoring-1-wire-sensors-with-netdata">Monitoring 1-Wire Sensors With Netdata&lt;/h3>
&lt;p>Netdata offers an advanced and intuitive platform for monitoring 1-Wire Sensors. By leveraging the Netdata &lt;strong>&lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/w1sensor/">1-Wire Sensors monitoring tool&lt;/a>&lt;/strong>, you can achieve real-time visibility into your environmental conditions, enabling you to respond swiftly to changes that could affect your infrastructure.&lt;/p></description></item><item><title>2023 Use Case Survey and Giveaway</title><link>https://www.netdata.cloud/2023-use-case-survey-giveaway-terms-and-conditions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/2023-use-case-survey-giveaway-terms-and-conditions/</guid><description/></item><item><title>2Wcom GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/2wcom-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/2wcom-gmbh-snmp-traps/</guid><description/></item><item><title>2Wire Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/2wire-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/2wire-inc-snmp-traps/</guid><description/></item><item><title>3Com</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/3com/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/3com/</guid><description/></item><item><title>3Com Huawei</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/3com-huawei/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/3com-huawei/</guid><description/></item><item><title>3Com SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/3com-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/3com-snmp-traps/</guid><description/></item><item><title>3Par Data SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/3par-data-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/3par-data-snmp-traps/</guid><description/></item><item><title>4D Server</title><link>https://www.netdata.cloud/integrations/data-collection/databases/4d-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/4d-server/</guid><description/></item><item><title>4D Server Monitoring</title><link>https://www.netdata.cloud/monitoring-101/4d_server-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/4d_server-monitoring/</guid><description>&lt;h2 id="4d-server-monitoring">4D Server Monitoring&lt;/h2>
&lt;h3 id="what-is-4d-server">What Is 4D Server?&lt;/h3>
&lt;p>4D Server is a database server that seamlessly combines a powerful SQL engine with innovative application development capabilities. It is renowned for its flexibility and efficiency in managing complex applications and databases. Whether you are building applications, handling large databases, or performing detailed data analysis, 4D Server provides the necessary infrastructure to manage your needs effectively.&lt;/p>
&lt;h3 id="monitoring-4d-server-with-netdata">Monitoring 4D Server With Netdata&lt;/h3>
&lt;p>Monitoring 4D Server is essential for maintaining its health and performance. With Netdata, you can easily monitor 4D Server using the integrated capabilities of the go.d.plugin. This monitoring tool leverages an openmetrics (Prometheus) exporter, allowing you to gather essential metrics seamlessly. The Netdata platform can ingest data from any Prometheus exporter, providing users with automated dashboards, real-time alerts, and comprehensive insights—all without the need for a standalone Prometheus server or Grafana setup. Explore how robust 4D Server monitoring can enhance your system management.&lt;/p></description></item><item><title>4Rf Communications Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/4rf-communications-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/4rf-communications-ltd-snmp-traps/</guid><description/></item><item><title>8430FT modem</title><link>https://www.netdata.cloud/integrations/data-collection/networking/8430ft-modem/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/8430ft-modem/</guid><description/></item><item><title>8430FT Modem Monitoring</title><link>https://www.netdata.cloud/monitoring-101/8430ft-modem-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/8430ft-modem-monitoring/</guid><description>&lt;h2 id="8430ft-modem-monitoring">8430FT Modem Monitoring&lt;/h2>
&lt;h3 id="what-is-8430ft-modem">What Is 8430FT Modem?&lt;/h3>
&lt;p>The 8430FT modem is a crucial component within network stacks, providing connectivity and data transfer capabilities in various technical environments. Understanding how this modem operates allows IT professionals, network engineers, and developers to ensure optimal performance and reliability.&lt;/p>
&lt;h3 id="monitoring-8430ft-modem-with-netdata">Monitoring 8430FT Modem With Netdata&lt;/h3>
&lt;p>For monitoring the 8430FT modem, Netdata utilizes an advanced openmetrics (Prometheus) exporter. By leveraging the &lt;a href="https://github.com/dernasherbrezon/8430ft_exporter">8430FT Exporter&lt;/a>, Netdata can efficiently collect and visualize essential metrics from the modem, aiding in network performance diagnostics. Unlike traditional setups, Netdata can ingest data from any Prometheus exporter, which means that users do not need an entire Prometheus server or Grafana to obtain actionable insights. With Netdata, these metrics are instantly transformed into automated dashboards and alerts, expediting the troubleshooting process.&lt;/p></description></item><item><title>A10</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/a10/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/a10/</guid><description/></item><item><title>A10 Networks Previously Raksha Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/a10-networks-previously-raksha-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/a10-networks-previously-raksha-networks-inc-snmp-traps/</guid><description/></item><item><title>A10 Thunder</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/a10-thunder/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/a10-thunder/</guid><description/></item><item><title>Abb Power Protection S.A. SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/abb-power-protection-s.a.-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/abb-power-protection-s.a.-snmp-traps/</guid><description/></item><item><title>Ablerex Electronic Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ablerex-electronic-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ablerex-electronic-co-ltd-snmp-traps/</guid><description/></item><item><title>Acc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/acc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/acc-snmp-traps/</guid><description/></item><item><title>Accedian Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/accedian-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/accedian-inc-snmp-traps/</guid><description/></item><item><title>Accelerated Concepts Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/accelerated-concepts-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/accelerated-concepts-inc-snmp-traps/</guid><description/></item><item><title>Access Points</title><link>https://www.netdata.cloud/integrations/data-collection/networking/access-points/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/access-points/</guid><description/></item><item><title>Access Points Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ap-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ap-monitoring/</guid><description>&lt;h2 id="access-points-monitoring">Access Points Monitoring&lt;/h2>
&lt;h3 id="what-is-access-points">What Is Access Points?&lt;/h3>
&lt;p>Access Points, commonly referred to as APs, are pivotal devices in wireless networks, serving as gateways for clients to connect to larger networks. They facilitate seamless communication by acting as a bridge between wireless clients and the wired network infrastructure.&lt;/p>
&lt;h3 id="monitoring-access-points-with-netdata">Monitoring Access Points With Netdata&lt;/h3>
&lt;p>Monitoring Access Points effectively is crucial for ensuring optimal network performance and reliability. With &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/ap/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata&amp;rsquo;s Access Points monitoring tool&lt;/a>, you can gain deep insights into the operations of your wireless network infrastructure. Netdata&amp;rsquo;s platform allows for real-time monitoring and troubleshooting, offering a comprehensive view of various performance metrics. You can &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">explore our live demo&lt;/a> to see how it works in real time.&lt;/p></description></item><item><title>Accton Technology SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/accton-technology-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/accton-technology-snmp-traps/</guid><description/></item><item><title>Accuenergy Canada Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/accuenergy-canada-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/accuenergy-canada-inc-snmp-traps/</guid><description/></item><item><title>Acer SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/acer-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/acer-snmp-traps/</guid><description/></item><item><title>Acksys SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/acksys-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/acksys-snmp-traps/</guid><description/></item><item><title>Acme Packet SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/acme-packet-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/acme-packet-snmp-traps/</guid><description/></item><item><title>Actidata Company SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/actidata-company-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/actidata-company-snmp-traps/</guid><description/></item><item><title>Active Directory</title><link>https://www.netdata.cloud/integrations/data-collection/applications/active-directory/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/active-directory/</guid><description/></item><item><title>Active Directory Certificate Service</title><link>https://www.netdata.cloud/integrations/data-collection/applications/active-directory-certificate-service/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/active-directory-certificate-service/</guid><description/></item><item><title>Active Directory Federation Service</title><link>https://www.netdata.cloud/integrations/data-collection/applications/active-directory-federation-service/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/active-directory-federation-service/</guid><description/></item><item><title>Active Directory Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ad-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ad-monitoring/</guid><description>&lt;h2 id="active-directory">Active Directory&lt;/h2>
&lt;p>Active Directory (AD) is a tool from Microsoft that helps you manage Windows networks. It comes with most Windows Server systems and does many things. At first, Active Directory only helped you control domains. But now, it also helps you with many other things related to identity.&lt;/p>
&lt;p>&lt;a href="https://learn.microsoft.com/en-us/windows-server/identity/ad-fs/design/ad-fs-requirements">It stores information about objects on the network, such as users, computers, groups, printers, and policies&lt;/a>. &lt;a href="https://learn.microsoft.com/en-us/previous-versions/windows/it-pro/windows-server-2012-R2-and-2012/dn781428(v=ws.11)">It also provides methods for authenticating and authorizing users and devices&lt;/a>. Active Directory can be deployed on-premises or in the cloud with Azure Active Directory (Azure AD).&lt;/p></description></item><item><title>ActiveMQ</title><link>https://www.netdata.cloud/integrations/data-collection/databases/activemq/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/activemq/</guid><description/></item><item><title>ActiveMQ advisory topics: hidden destinations consuming resources</title><link>https://www.netdata.cloud/guides/activemq/activemq-advisory-topics-overhead/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-advisory-topics-overhead/</guid><description>&lt;h1 id="activemq-advisory-topics-hidden-destinations-consuming-resources">ActiveMQ advisory topics: hidden destinations consuming resources&lt;/h1>
&lt;p>List every topic on a busy ActiveMQ Classic broker and count how many start with &lt;code>ActiveMQ.Advisory.&lt;/code>. On a broker with a thousand queues, you can easily find a thousand or more advisory topics, each a real destination with its own JMX MBeans, each receiving non-persistent messages for broker events.&lt;/p>
&lt;p>This is not a bug. Advisory topics are how the broker publishes internal events: connections opening and closing, consumers and producers starting and stopping, destinations hitting their memory limit, messages expiring. Tools and clients can subscribe to observe the broker. The problem is that the default behavior scales with destination count, not with how much of this information anyone consumes, and the cost is paid in MBean count, heap, GC pressure, and inflated throughput counters.&lt;/p></description></item><item><title>ActiveMQ authentication failures: credential rotation fallout and brute force</title><link>https://www.netdata.cloud/guides/activemq/activemq-authentication-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-authentication-failures/</guid><description>&lt;h1 id="activemq-authentication-failures-credential-rotation-fallout-and-brute-force">ActiveMQ authentication failures: credential rotation fallout and brute force&lt;/h1>
&lt;p>Your broker log starts filling with &lt;code>SecurityException: User name [X] or password is invalid&lt;/code> or &lt;code>Authentication failed for user&lt;/code>. Sometimes it is a slow drip from a single client. Sometimes it is a flood from dozens of source IPs. Both look similar in the log, but they are very different incidents: one is usually a forgotten service after a credential rotation, the other is either a production outage (real consumers and producers dropped) or an active attack.&lt;/p></description></item><item><title>ActiveMQ authorization denied: authenticated users hitting Not authorized</title><link>https://www.netdata.cloud/guides/activemq/activemq-authorization-denied/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-authorization-denied/</guid><description>&lt;h1 id="activemq-authorization-denied-authenticated-users-hitting-not-authorized">ActiveMQ authorization denied: authenticated users hitting Not authorized&lt;/h1>
&lt;p>Your ActiveMQ broker log is filling with lines like &lt;code>User orders-svc is not authorized to write to: queue://PAYMENTS.INBOUND&lt;/code> or &lt;code>Not authorized to create: topic://ActiveMQ.Advisory.Connection&lt;/code>. The client connected fine. Authentication passed. But every send, consume, or destination creation attempt is rejected.&lt;/p>
&lt;p>This is an authorization failure, not an authentication failure. The user proved who they are; the broker is saying they cannot do the specific thing they attempted. That distinction drives the whole diagnosis: you are not chasing bad passwords, you are chasing a mismatch between the user the client authenticated as, the operation it attempted, and the authorization entries the broker loaded.&lt;/p></description></item><item><title>ActiveMQ broker down: telling a crashed broker from a hung one</title><link>https://www.netdata.cloud/guides/activemq/activemq-broker-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-broker-down/</guid><description>&lt;h1 id="activemq-broker-down-telling-a-crashed-broker-from-a-hung-one">ActiveMQ broker down: telling a crashed broker from a hung one&lt;/h1>
&lt;p>Your pager says the ActiveMQ broker is down. The first question is not &amp;ldquo;how do I restart it&amp;rdquo; but &amp;ldquo;what kind of down is this.&amp;rdquo; A crashed broker (process gone, port closed) and a hung broker (process alive, port possibly open, nothing moving) have different causes, different evidence, and different safe responses. Restarting a hung broker without capturing diagnostics destroys the evidence. Assuming a crashed broker will come straight back up ignores the KahaDB recovery window, where the process runs and the port may even bind, but no client gets served for minutes to hours.&lt;/p></description></item><item><title>ActiveMQ broker won't start: port conflicts, store recovery, and lock contention</title><link>https://www.netdata.cloud/guides/activemq/activemq-broker-wont-start/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-broker-wont-start/</guid><description>&lt;h1 id="activemq-broker-wont-start-port-conflicts-store-recovery-and-lock-contention">ActiveMQ broker won&amp;rsquo;t start: port conflicts, store recovery, and lock contention&lt;/h1>
&lt;p>The broker process is running, or maybe it is not, and clients cannot connect. You restart the service, and it still does not come up. Or worse: the process starts, the log goes quiet, and ten minutes later you are still staring at a broker that has not opened its transport connectors.&lt;/p>
&lt;p>&amp;ldquo;Won&amp;rsquo;t start&amp;rdquo; in ActiveMQ Classic covers four distinct failure modes that look identical from the outside: something else is holding port 61616 or 8161, KahaDB is replaying a large journal and simply is not done yet, the store is corrupted after an unclean shutdown and the broker is refusing to start, or a store lock in a shared-storage HA pair is held by the wrong broker and this one is waiting forever. Each has a different fix, and applying the wrong one (especially deleting store files or force-clearing locks) can turn a slow start into permanent message loss.&lt;/p></description></item><item><title>ActiveMQ connection and session leak: clients that never close</title><link>https://www.netdata.cloud/guides/activemq/activemq-connection-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-connection-leak/</guid><description>&lt;h1 id="activemq-connection-and-session-leak-clients-that-never-close">ActiveMQ connection and session leak: clients that never close&lt;/h1>
&lt;p>The broker&amp;rsquo;s connection count used to sit around 400. Last month it was 600. This week it is 1,100, and nobody deployed anything new. The JVM thread count and open file descriptor count are climbing in lockstep. Eventually, usually at the worst time, the broker hits its FD limit and starts rejecting every new connection, including the healthy clients.&lt;/p>
&lt;p>This is the classic ActiveMQ connection and session leak: clients that open JMS Connections (or Sessions) and never close them. Each leaked connection holds a socket, a file descriptor, one or more transport threads on the default TCP transport, and broker-side state. One JMS Connection can hold many Sessions, and each Session can hold many consumers and producers, so the leak compounds at each layer.&lt;/p></description></item><item><title>ActiveMQ consumers connected but not acknowledging: the zombie consumer</title><link>https://www.netdata.cloud/guides/activemq/activemq-consumer-not-acking/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-consumer-not-acking/</guid><description>&lt;h1 id="activemq-consumers-connected-but-not-acknowledging-the-zombie-consumer">ActiveMQ consumers connected but not acknowledging: the zombie consumer&lt;/h1>
&lt;p>The queue has consumers. &lt;code>ConsumerCount&lt;/code> is exactly where it should be. But dequeue rate has collapsed to near zero, &lt;code>InFlightCount&lt;/code> is pinned at the prefetch limit, and the backlog keeps growing. From the broker&amp;rsquo;s perspective everything is subscribed; from the business&amp;rsquo;s perspective nothing is being processed.&lt;/p>
&lt;p>This is the zombie consumer: a consumer that holds a live connection and a full prefetch buffer but never acknowledges. ActiveMQ dispatched the prefetch window of messages to it, and the broker will not send that consumer more until acks come back. If it is the only consumer, the queue stalls while looking fully staffed.&lt;/p></description></item><item><title>ActiveMQ CVE-2023-46604: the OpenWire deserialization RCE and how to detect exposure</title><link>https://www.netdata.cloud/guides/activemq/activemq-cve-2023-46604-openwire-rce/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-cve-2023-46604-openwire-rce/</guid><description>&lt;h1 id="activemq-cve-2023-46604-the-openwire-deserialization-rce-and-how-to-detect-exposure">ActiveMQ CVE-2023-46604: the OpenWire deserialization RCE and how to detect exposure&lt;/h1>
&lt;p>CVE-2023-46604 is a CVSS 10.0 unauthenticated remote code execution vulnerability in the ActiveMQ Classic OpenWire protocol marshaller. If an attacker can open a TCP connection to your OpenWire transport connector (default port 61616), they can send a crafted packet that causes the broker to instantiate an arbitrary class on the classpath. No credentials are required. It was exploited in the wild before and immediately after disclosure in October 2023, and unpatched, internet-reachable brokers are still being found.&lt;/p></description></item><item><title>ActiveMQ default credentials and exposed web console: admin/admin on port 8161</title><link>https://www.netdata.cloud/guides/activemq/activemq-default-credentials-exposure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-default-credentials-exposure/</guid><description>&lt;h1 id="activemq-default-credentials-and-exposed-web-console-adminadmin-on-port-8161">ActiveMQ default credentials and exposed web console: admin/admin on port 8161&lt;/h1>
&lt;p>ActiveMQ Classic ships with a web console on port 8161 and well-known default credentials: admin/admin. Combined with a management interface reachable from the network, that equals full broker control for anyone who finds the port: browse and purge queues, move messages, read message bodies, create destinations, and in some versions reach the Jolokia JMX REST API for much worse.&lt;/p></description></item><item><title>ActiveMQ deserialization errors: ClassNotFoundException and the serializable packages whitelist</title><link>https://www.netdata.cloud/guides/activemq/activemq-deserialization-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-deserialization-errors/</guid><description>&lt;h1 id="activemq-deserialization-errors-classnotfoundexception-and-the-serializable-packages-whitelist">ActiveMQ deserialization errors: ClassNotFoundException and the serializable packages whitelist&lt;/h1>
&lt;p>&lt;code>ClassNotFoundException&lt;/code>, &lt;code>InvalidClassException&lt;/code>, or &lt;code>StreamCorruptedException&lt;/code> in the ActiveMQ broker log means messages are failing to process. Sometimes this is mundane: a producer and consumer compiled against different versions of a class, or an ObjectMessage payload class that is not on the broker&amp;rsquo;s classpath. Sometimes it is the signature of a Java deserialization exploit attempt, and the correct response is a security incident, not a config tweak.&lt;/p></description></item><item><title>ActiveMQ destination explosion: dynamic destinations, MBean bloat, and GC pressure</title><link>https://www.netdata.cloud/guides/activemq/activemq-destination-explosion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-destination-explosion/</guid><description>&lt;h1 id="activemq-destination-explosion-dynamic-destinations-mbean-bloat-and-gc-pressure">ActiveMQ destination explosion: dynamic destinations, MBean bloat, and GC pressure&lt;/h1>
&lt;p>An ActiveMQ Classic broker that has been running for months shows a specific signature: heap usage climbs steadily even when message volume is flat, GC pauses get longer and more frequent, the web console and JMX queries feel sluggish, and eventually the broker OOMs or slides into a GC death spiral. Queue depths look normal. Enqueue and dequeue rates look normal. Nothing about the messaging workload explains it.&lt;/p></description></item><item><title>ActiveMQ disk full on the KahaDB partition: write failures and store corruption risk</title><link>https://www.netdata.cloud/guides/activemq/activemq-disk-full-kahadb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-disk-full-kahadb/</guid><description>&lt;h1 id="activemq-disk-full-on-the-kahadb-partition-write-failures-and-store-corruption-risk">ActiveMQ disk full on the KahaDB partition: write failures and store corruption risk&lt;/h1>
&lt;p>The KahaDB partition is at 100%. The broker log shows journal write failures, persistent producers have stopped, and you are in the worst ActiveMQ failure mode: the one where freeing space may not be enough, because an in-flight write at the moment the disk filled can leave the journal or index corrupt.&lt;/p>
&lt;p>This is not the same incident as &lt;code>StorePercentUsage&lt;/code> hitting 100%. &lt;code>StorePercentUsage&lt;/code> measures KahaDB against the configured &lt;code>storeUsage&lt;/code> limit in &lt;code>activemq.xml&lt;/code>. Disk full measures the partition against physical capacity with &lt;code>df&lt;/code>. They are independent limits, and either can fire first. If &lt;code>storeUsage&lt;/code> is larger than the partition, the OS runs out of space while the broker still thinks it has headroom, and the failure arrives with no warning from any ActiveMQ metric. This guide covers the OS-level case; for the configured-limit case, see &lt;a href="https://www.netdata.cloud/guides/activemq/activemq-store-is-full/">ActiveMQ store is full&lt;/a>.&lt;/p></description></item><item><title>ActiveMQ DLQ never expires: setting TTL so the dead-letter queue stops leaking storage</title><link>https://www.netdata.cloud/guides/activemq/activemq-dlq-no-expiration/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-dlq-no-expiration/</guid><description>&lt;h1 id="activemq-dlq-never-expires-setting-ttl-so-the-dead-letter-queue-stops-leaking-storage">ActiveMQ DLQ never expires: setting TTL so the dead-letter queue stops leaking storage&lt;/h1>
&lt;p>Your &lt;code>ActiveMQ.DLQ&lt;/code> has been sitting at some non-zero depth for months. Nobody looks at it because &amp;ldquo;it&amp;rsquo;s just the DLQ.&amp;rdquo; Meanwhile &lt;code>StorePercentUsage&lt;/code> creeps up a fraction of a percent a day, the &lt;code>db-*.log&lt;/code> journal files keep accumulating, and one morning the broker hits 100% store usage and blocks every persistent producer on the bus.&lt;/p>
&lt;p>This is the default behavior of ActiveMQ Classic, not a bug. Messages sent to the dead-letter queue have no TTL. They accumulate forever, and every one of them is a live reference that pins KahaDB journal files against garbage collection. The fix has two parts: configure expiration on the DLQ so the leak stops, and drain-and-investigate the existing backlog instead of just deleting it. Every DLQ message is a failed business transaction; purging blind throws away the evidence.&lt;/p></description></item><item><title>ActiveMQ enqueue outpacing dequeue: reading the rate imbalance before the backlog</title><link>https://www.netdata.cloud/guides/activemq/activemq-enqueue-dequeue-imbalance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-enqueue-dequeue-imbalance/</guid><description>&lt;h1 id="activemq-enqueue-outpacing-dequeue-reading-the-rate-imbalance-before-the-backlog">ActiveMQ enqueue outpacing dequeue: reading the rate imbalance before the backlog&lt;/h1>
&lt;p>On a healthy ActiveMQ Classic broker, enqueue rate and dequeue rate track each other within normal burst variability. When enqueue outpaces dequeue for more than a few minutes, the broker is accumulating a backlog whether or not QueueSize has moved enough to alarm yet. The rate delta is the leading indicator; QueueSize, MemoryPercentUsage, and StorePercentUsage are the lagging confirmations. If you wait for the backlog gauges to fire, you have already lost your runway.&lt;/p></description></item><item><title>ActiveMQ expired message count climbing: TTL expiry and silent correctness loss</title><link>https://www.netdata.cloud/guides/activemq/activemq-expired-messages/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-expired-messages/</guid><description>&lt;h1 id="activemq-expired-message-count-climbing-ttl-expiry-and-silent-correctness-loss">ActiveMQ expired message count climbing: TTL expiry and silent correctness loss&lt;/h1>
&lt;p>A queue looks fine. QueueSize is stable, dequeue rate is non-zero, no alerts on memory or store. Then someone asks why a customer order never processed, and you find &lt;code>ExpiredCount&lt;/code> on the destination has been climbing for days. Every tick of that counter is a message the broker threw away or shuffled into the DLQ because its TTL ran out before a consumer got to it. Nothing crashed. Nothing paged. The data is just gone.&lt;/p></description></item><item><title>ActiveMQ GC pause death spiral: long pauses, heartbeat timeouts, and reconnect storms</title><link>https://www.netdata.cloud/guides/activemq/activemq-gc-pause-death-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-gc-pause-death-spiral/</guid><description>&lt;h1 id="activemq-gc-pause-death-spiral-long-pauses-heartbeat-timeouts-and-reconnect-storms">ActiveMQ GC pause death spiral: long pauses, heartbeat timeouts, and reconnect storms&lt;/h1>
&lt;p>The broker is up. The port is listening. But clients keep dropping, reconnecting, dropping again, and every cycle makes things worse. Connection count charts show a sawtooth: a sudden drop, a spike, another drop. In the JVM, heap usage sits above 90% after every major GC and full GC pauses are climbing into the multi-second range.&lt;/p>
&lt;p>This is the GC pause death spiral, a positive-feedback loop and one of the characteristic failure archetypes of ActiveMQ Classic. Heap pressure produces long GC pauses. A long pause freezes every broker thread, including the ones answering client keepalives. Clients exceed &lt;code>wireFormat.maxInactivityDuration&lt;/code> (default 30000 ms) and disconnect. After the pause ends, every client reconnects at once, and each reconnect allocates new connection, session, consumer, and subscription objects. That allocation spike raises heap pressure further, the next GC pause is longer, and more clients time out. The broker oscillates between being frozen in GC and being hammered by reconnect storms until it is effectively down.&lt;/p></description></item><item><title>ActiveMQ HA failover stuck: standby cannot acquire the store lock</title><link>https://www.netdata.cloud/guides/activemq/activemq-ha-failover-stuck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-ha-failover-stuck/</guid><description>&lt;h1 id="activemq-ha-failover-stuck-standby-cannot-acquire-the-store-lock">ActiveMQ HA failover stuck: standby cannot acquire the store lock&lt;/h1>
&lt;p>The active broker died. The standby broker&amp;rsquo;s process is running, the JVM looks healthy, and yet no client can connect anywhere. &lt;code>ss&lt;/code> shows nothing listening on 61616 on either node. The failover you designed for is not happening.&lt;/p>
&lt;p>This is the shared-storage master/slave stall: the standby is blocked waiting to acquire the lock file in the KahaDB directory, and until it gets that lock it will not start its transport connectors. That behavior is by design. What is not by design is the lock never becoming available. The usual suspects are a stale lock file left behind by a crashed active, an unhealthy NFS lock daemon, or a SAN mount that is unavailable on the standby.&lt;/p></description></item><item><title>ActiveMQ InactivityIOException: Channel was inactive for too long</title><link>https://www.netdata.cloud/guides/activemq/activemq-channel-inactive-too-long/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-channel-inactive-too-long/</guid><description>&lt;h1 id="activemq-inactivityioexception-channel-was-inactive-for-too-long">ActiveMQ InactivityIOException: Channel was inactive for too long&lt;/h1>
&lt;p>Your producers or consumers are dropping JMS connections with this in the stack trace:&lt;/p>
&lt;pre tabindex="0">&lt;code>org.apache.activemq.transport.InactivityIOException: Channel was inactive for too long
&lt;/code>&lt;/pre>&lt;p>Sometimes it is one client. Sometimes every client on the broker disconnects within the same second and reconnects in a burst. The exception points at the connection, but the connection is almost never the root cause. Something on one side of the wire stopped producing traffic for longer than the OpenWire inactivity timeout, and the other side declared it dead.&lt;/p></description></item><item><title>ActiveMQ InFlightCount high: prefetch full, acks stalled, and zombie consumers</title><link>https://www.netdata.cloud/guides/activemq/activemq-inflight-count-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-inflight-count-high/</guid><description>&lt;h1 id="activemq-inflightcount-high-prefetch-full-acks-stalled-and-zombie-consumers">ActiveMQ InFlightCount high: prefetch full, acks stalled, and zombie consumers&lt;/h1>
&lt;p>A queue shows QueueSize near zero, consumers are connected, and nothing is being processed. Dequeue rate is flat, message age is rising upstream, and the only number that looks wrong is InFlightCount, sitting at a suspiciously round value. This is the zombie consumer pattern: the broker dispatched messages into consumer prefetch buffers, and those messages are never coming back as acknowledgments.&lt;/p></description></item><item><title>ActiveMQ JVM heap exhaustion: OutOfMemoryError and the OOM kill</title><link>https://www.netdata.cloud/guides/activemq/activemq-jvm-heap-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-jvm-heap-exhaustion/</guid><description>&lt;h1 id="activemq-jvm-heap-exhaustion-outofmemoryerror-and-the-oom-kill">ActiveMQ JVM heap exhaustion: OutOfMemoryError and the OOM kill&lt;/h1>
&lt;p>The broker was fine an hour ago. Now the process is gone, or it is alive but frozen: clients disconnected, the web console hangs, and the log shows &lt;code>java.lang.OutOfMemoryError&lt;/code> or, if it runs in a container, nothing at all because the kernel killed the JVM with SIGKILL (exit code 137, i.e. 128+9).&lt;/p>
&lt;p>This is ActiveMQ JVM heap exhaustion. It is distinct from the broker&amp;rsquo;s internal &lt;code>memoryUsage&lt;/code> hitting 100%. That limit triggers producer flow control and blocks &lt;code>send()&lt;/code> calls; this one kills or freezes the JVM itself. The confusing part is that the two are independent counters. You can have &lt;code>MemoryPercentUsage&lt;/code> at 50% with plenty of headroom while JVM heap is at 97% and the next full GC never finishes. Non-message allocations (destination metadata, MBeans, connection and session state, cursor overhead, transport buffers) live on the heap too, and ActiveMQ&amp;rsquo;s flow control accounting does not protect them.&lt;/p></description></item><item><title>ActiveMQ JVM thread count climbing: transport threads, NIO, and thread leaks</title><link>https://www.netdata.cloud/guides/activemq/activemq-thread-count-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-thread-count-growing/</guid><description>&lt;h1 id="activemq-jvm-thread-count-climbing-transport-threads-nio-and-thread-leaks">ActiveMQ JVM thread count climbing: transport threads, NIO, and thread leaks&lt;/h1>
&lt;p>The broker&amp;rsquo;s JVM thread count has been climbing for hours or days. Nothing has broken yet, but the trend only goes one direction, and it ends with context-switch overhead, native memory pressure from thread stacks, and eventually a broker that cannot accept new connections or gets OOM-killed despite a healthy heap.&lt;/p>
&lt;p>A rising thread count on ActiveMQ Classic is not one problem. With the default blocking TCP transport, thread count is a proxy for connection count, so growth usually means connections are growing. When thread count climbs without connection growth, you have an actual thread leak, and the diagnostic path is completely different. Telling the two apart takes about two minutes with JMX.&lt;/p></description></item><item><title>ActiveMQ KahaDB corruption: the broker won't start after an unclean shutdown</title><link>https://www.netdata.cloud/guides/activemq/activemq-kahadb-corruption/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-kahadb-corruption/</guid><description>&lt;h1 id="activemq-kahadb-corruption-the-broker-wont-start-after-an-unclean-shutdown">ActiveMQ KahaDB corruption: the broker won&amp;rsquo;t start after an unclean shutdown&lt;/h1>
&lt;p>The broker was killed (kill -9, power loss, OOM killer, container limit hit) and now it will not come back. Either the process exits during startup, or it sits there with the log stuck partway through KahaDB recovery and port 61616 never accepts a client. The log mentions &lt;code>db.data&lt;/code>, a &lt;code>db-*.log&lt;/code> journal file, or an &lt;code>IOException&lt;/code> about a missing data file.&lt;/p></description></item><item><title>ActiveMQ KahaDB db.data index bloat: slow lookups and slow startup recovery</title><link>https://www.netdata.cloud/guides/activemq/activemq-kahadb-index-large/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-kahadb-index-large/</guid><description>&lt;h1 id="activemq-kahadb-dbdata-index-bloat-slow-lookups-and-slow-startup-recovery">ActiveMQ KahaDB db.data index bloat: slow lookups and slow startup recovery&lt;/h1>
&lt;p>You restart the broker after an unclean shutdown and it sits there for 20, 30, 40 minutes before it accepts a single client connection. Or the broker is running, but persistent message dispatch feels sluggish and disk reads on the KahaDB partition are elevated for no obvious reason. You look at the KahaDB directory and the journal files are not the problem. The problem is &lt;code>db.data&lt;/code>, the B-tree index, sitting at multiple gigabytes.&lt;/p></description></item><item><title>ActiveMQ KahaDB journal files not deleted: one unacked message pinning a 32MB log</title><link>https://www.netdata.cloud/guides/activemq/activemq-kahadb-journal-files-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-kahadb-journal-files-growing/</guid><description>&lt;h1 id="activemq-kahadb-journal-files-not-deleted-one-unacked-message-pinning-a-32mb-log">ActiveMQ KahaDB journal files not deleted: one unacked message pinning a 32MB log&lt;/h1>
&lt;p>Your application queues are empty or nearly empty. Consumers are connected and processing. Yet the KahaDB directory keeps growing, &lt;code>db-*.log&lt;/code> files pile up, &lt;code>StorePercentUsage&lt;/code> creeps upward, and disk free space trends toward zero. This is the classic KahaDB journal pinning problem, and it confuses operators precisely because the visible queue state looks healthy.&lt;/p>
&lt;p>The mechanism is simple and unforgiving: a KahaDB journal file (default 32MB) is only reclaimable when &lt;strong>every&lt;/strong> message stored in that file has been acknowledged. One unacknowledged message anywhere in the file pins the entire file on disk. If that one message is sitting in the DLQ, held by an offline durable subscriber, or stuck in a dead consumer&amp;rsquo;s prefetch buffer, the file stays, and new files keep being written behind it.&lt;/p></description></item><item><title>ActiveMQ KahaDB journal write latency: fsync as the persistent-throughput ceiling</title><link>https://www.netdata.cloud/guides/activemq/activemq-store-write-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-store-write-latency/</guid><description>&lt;h1 id="activemq-kahadb-journal-write-latency-fsync-as-the-persistent-throughput-ceiling">ActiveMQ KahaDB journal write latency: fsync as the persistent-throughput ceiling&lt;/h1>
&lt;p>A producer is sending persistent messages to ActiveMQ Classic and throughput is far below what the network, CPU, and broker configuration suggest should be possible. Sends are not failing. There are no exceptions. Each &lt;code>send()&lt;/code> just takes a few milliseconds, or tens of milliseconds, and the aggregate rate caps out no matter how many producers you add.&lt;/p>
&lt;p>This is almost always the KahaDB journal fsync. For persistent messages, ActiveMQ Classic writes the message to the KahaDB write-ahead journal and acknowledges the producer only after the fsync completes. That fsync is the critical path for every persistent send, and its latency sets a hard ceiling on persistent throughput. When the underlying storage gets slow, persistent messaging gets slower in direct proportion to the fsync latency.&lt;/p></description></item><item><title>ActiveMQ memory limit reached: MemoryPercentUsage at 100% and the flow-control cliff</title><link>https://www.netdata.cloud/guides/activemq/activemq-memory-limit-reached/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-memory-limit-reached/</guid><description>&lt;h1 id="activemq-memory-limit-reached-memorypercentusage-at-100-and-the-flow-control-cliff">ActiveMQ memory limit reached: MemoryPercentUsage at 100% and the flow-control cliff&lt;/h1>
&lt;p>Your producers have stopped sending. No exceptions, no timeouts, no errors on the producer side. The &lt;code>send()&lt;/code> calls just hang, and upstream services pile up threads waiting on the broker. In the broker log you find the line operators search for at 3 a.m.:&lt;/p>
&lt;pre tabindex="0">&lt;code>Usage Manager Memory Usage ... reached memory limit
&lt;/code>&lt;/pre>&lt;p>&lt;code>MemoryPercentUsage&lt;/code> on the Broker MBean is at 100%. This is ActiveMQ Classic&amp;rsquo;s own memory accounting, not JVM heap, and 100% is not a slowdown. It is a cliff edge. At 99% everything works. At 100% every producer is flow-controlled: the broker stops reading from producer sockets, TCP backpressure builds, and sends block silently until memory frees up.&lt;/p></description></item><item><title>ActiveMQ MemoryPercentUsage climbing: reading the flow-control leading indicator</title><link>https://www.netdata.cloud/guides/activemq/activemq-memory-percent-usage-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-memory-percent-usage-high/</guid><description>&lt;h1 id="activemq-memorypercentusage-climbing-reading-the-flow-control-leading-indicator">ActiveMQ MemoryPercentUsage climbing: reading the flow-control leading indicator&lt;/h1>
&lt;p>Your ActiveMQ broker&amp;rsquo;s &lt;code>MemoryPercentUsage&lt;/code> gauge is at 65% and climbing. Nothing is broken yet. Producers are sending, consumers are consuming, queue depths look tolerable. But the trend line only goes one direction, and you know what happens at 100%: producer flow control activates, every producer&amp;rsquo;s &lt;code>send()&lt;/code> call blocks silently, and upstream services start hanging with no error message anywhere.&lt;/p>
&lt;p>This is the right moment to act, and the gauge gives you more information than most operators extract from it. The absolute value tells you where you are. The rate of climb tells you how much time you have. The surrounding signals tell you why. This article is about reading all three before the cliff edge.&lt;/p></description></item><item><title>ActiveMQ memoryUsage vs JVM heap: the two memory budgets teams confuse</title><link>https://www.netdata.cloud/guides/activemq/activemq-systemusage-memory-vs-heap/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-systemusage-memory-vs-heap/</guid><description>&lt;h1 id="activemq-memoryusage-vs-jvm-heap-the-two-memory-budgets-teams-confuse">ActiveMQ memoryUsage vs JVM heap: the two memory budgets teams confuse&lt;/h1>
&lt;p>An ActiveMQ Classic broker can hit 100% &lt;code>MemoryPercentUsage&lt;/code> and silently block every producer while the JVM heap sits at 40%. The same broker, on a different day, can die from &lt;code>OutOfMemoryError&lt;/code> while &lt;code>MemoryPercentUsage&lt;/code> shows 50% and every queue-depth dashboard looks calm. Both incidents look contradictory until you understand that ActiveMQ tracks two separate memory budgets, and only one of them is the JVM&amp;rsquo;s.&lt;/p></description></item><item><title>ActiveMQ Monitoring</title><link>https://www.netdata.cloud/monitoring-101/activemq-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/activemq-monitoring/</guid><description>&lt;h2 id="activemq-monitoring">ActiveMQ Monitoring&lt;/h2>
&lt;h3 id="what-is-activemq">What Is ActiveMQ?&lt;/h3>
&lt;p>ActiveMQ is a popular open-source message broker designed for message-oriented middleware. It is known for its versatility in handling various messaging protocols, supporting the Java Messaging Service (JMS) API, and its ability to run as a standalone broker or in a fully distributed enterprise setup. For technical users like DevOps, SREs, and developers, ActiveMQ allows seamless integration of numerous applications, facilitating reliable and asynchronous message exchanges. &lt;a href="https://activemq.apache.org/">Learn more about ActiveMQ&lt;/a>.&lt;/p></description></item><item><title>ActiveMQ monitoring checklist: the signals every production broker needs</title><link>https://www.netdata.cloud/guides/activemq/activemq-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-monitoring-checklist/</guid><description>&lt;h1 id="activemq-monitoring-checklist-the-signals-every-production-broker-needs">ActiveMQ monitoring checklist: the signals every production broker needs&lt;/h1>
&lt;p>Most ActiveMQ monitoring setups fail in one of two ways. Either they only watch process liveness and queue depth, and discover producer flow control from angry users. Or they export every JMX attribute into a dashboard nobody reads, and the one signal that mattered is buried on page four.&lt;/p>
&lt;p>This checklist organizes the signals that matter into four maturity levels, from &amp;ldquo;is the broker alive&amp;rdquo; to &amp;ldquo;can you prove end-to-end message flow works.&amp;rdquo; Each level lists the signal, where it comes from, and the threshold that makes it actionable. Scope is ActiveMQ Classic (5.x and 6.x Classic stream). Artemis has a different store, memory model, and MBean tree; do not apply these thresholds to it.&lt;/p></description></item><item><title>ActiveMQ monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/activemq/activemq-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-monitoring-maturity-model/</guid><description>&lt;h1 id="activemq-monitoring-maturity-model-from-survival-to-expert">ActiveMQ monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most ActiveMQ outages are not caused by exotic failure modes. They are caused by signals nobody was watching: a DLQ that grew for three weeks, a memory limit that hit 100% and silently blocked every producer, an offline durable subscription that pinned journal files until the disk filled. The signals were available the whole time. The team had not instrumented them yet.&lt;/p></description></item><item><title>ActiveMQ network bridge connected but not forwarding: broken demand forwarding</title><link>https://www.netdata.cloud/guides/activemq/activemq-network-bridge-not-forwarding/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-network-bridge-not-forwarding/</guid><description>&lt;h1 id="activemq-network-bridge-connected-but-not-forwarding-broken-demand-forwarding">ActiveMQ network bridge connected but not forwarding: broken demand forwarding&lt;/h1>
&lt;p>The bridge is up. JMX shows the network bridge MBeans, the broker log shows a successful connection to the remote broker, and yet messages produced on broker A sit in the queue while consumers on broker B wait idle. No errors on either side. Each broker&amp;rsquo;s local dashboard looks healthy.&lt;/p>
&lt;p>This is broken demand forwarding in an ActiveMQ Classic Network of Brokers. A network bridge is not a pipe. It forwards messages only when it knows a consumer exists on the remote side, and it learns that through advisory messages. When the bridge is connected but the demand signal is not flowing, you get a silent stall: backlog on one broker, starvation on the other, nothing alarming on a per-broker view.&lt;/p></description></item><item><title>ActiveMQ network bridge down: a Network of Brokers partition and store-and-forward backlog</title><link>https://www.netdata.cloud/guides/activemq/activemq-network-bridge-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-network-bridge-down/</guid><description>&lt;h1 id="activemq-network-bridge-down-a-network-of-brokers-partition-and-store-and-forward-backlog">ActiveMQ network bridge down: a Network of Brokers partition and store-and-forward backlog&lt;/h1>
&lt;p>In a Network of Brokers, messages cross between ActiveMQ Classic brokers over network bridges. When a bridge drops, the failure is asymmetric: the origin broker keeps accepting messages for destinations whose consumers live on the remote broker, and those messages pile up locally. The remote broker keeps running with consumers connected and nothing to consume. Both brokers look healthy in isolation. Only cross-broker correlation reveals the partition.&lt;/p></description></item><item><title>ActiveMQ Network of Brokers replay storm: burst forwarding after a bridge reconnect</title><link>https://www.netdata.cloud/guides/activemq/activemq-nob-replay-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-nob-replay-storm/</guid><description>&lt;h1 id="activemq-network-of-brokers-replay-storm-burst-forwarding-after-a-bridge-reconnect">ActiveMQ Network of Brokers replay storm: burst forwarding after a bridge reconnect&lt;/h1>
&lt;p>A network bridge between two ActiveMQ brokers drops. Producers keep sending to the origin broker, and because ActiveMQ networks do reliable store-and-forward, the messages pile up locally. When the bridge reconnects, the accumulated backlog is replayed across the bridge in a burst. The receiving broker, idle a second ago, is now ingesting the backlog at wire speed. Its memory usage spikes, and if the spike reaches 100%, producer flow control activates on the receiver and producers there block silently.&lt;/p></description></item><item><title>ActiveMQ offline durable subscriber pending messages: the silent storage leak</title><link>https://www.netdata.cloud/guides/activemq/activemq-durable-subscriber-pending/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-durable-subscriber-pending/</guid><description>&lt;h1 id="activemq-offline-durable-subscriber-pending-messages-the-silent-storage-leak">ActiveMQ offline durable subscriber pending messages: the silent storage leak&lt;/h1>
&lt;p>StorePercentUsage has been climbing for three weeks. Your queues are draining, consumer counts look right, dequeue rates track enqueue rates, and yet the KahaDB directory keeps growing and journal files keep stacking up. Nothing in the queue-depth metrics explains it.&lt;/p>
&lt;p>The usual suspect is not a queue at all. It is a durable topic subscription whose subscriber went away and never came back. Every persistent message published to that topic is still being written to the store on behalf of that subscription, and it will keep happening forever: ActiveMQ Classic has no automatic expiration for offline durable subscribers by default.&lt;/p></description></item><item><title>ActiveMQ oldest message age: the queue latency depth alone cannot show</title><link>https://www.netdata.cloud/guides/activemq/activemq-message-age-oldest-pending/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-message-age-oldest-pending/</guid><description>&lt;h1 id="activemq-oldest-message-age-the-queue-latency-depth-alone-cannot-show">ActiveMQ oldest message age: the queue latency depth alone cannot show&lt;/h1>
&lt;p>Every ActiveMQ dashboard has queue depth on it. Almost none show the age of the oldest pending message, which is the number that actually maps to your SLA. A queue holding 50 messages whose head has been waiting 45 minutes is a much bigger problem than a queue holding 50,000 messages whose head is 2 seconds old and draining fast. Depth alone cannot tell you which of those two situations you are looking at.&lt;/p></description></item><item><title>ActiveMQ per-destination DLQ: IndividualDeadLetterStrategy vs the shared ActiveMQ.DLQ</title><link>https://www.netdata.cloud/guides/activemq/activemq-per-destination-dlq/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-per-destination-dlq/</guid><description>&lt;h1 id="activemq-per-destination-dlq-individualdeadletterstrategy-vs-the-shared-activemqdlq">ActiveMQ per-destination DLQ: IndividualDeadLetterStrategy vs the shared ActiveMQ.DLQ&lt;/h1>
&lt;p>Out of the box, ActiveMQ Classic routes every undeliverable message from every queue and topic into a single queue: &lt;code>ActiveMQ.DLQ&lt;/code>. One noisy destination with a poison-message problem mixes its failures with everyone else&amp;rsquo;s, and the only way to attribute a message back to its source is to browse the DLQ and inspect the &lt;code>JMSDestination&lt;/code> property on each entry. On a busy broker that is slow, expensive, and usually done too late.&lt;/p></description></item><item><title>ActiveMQ per-destination memory usage: one noisy queue blocking every producer</title><link>https://www.netdata.cloud/guides/activemq/activemq-per-destination-memory-limit/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-per-destination-memory-limit/</guid><description>&lt;h1 id="activemq-per-destination-memory-usage-one-noisy-queue-blocking-every-producer">ActiveMQ per-destination memory usage: one noisy queue blocking every producer&lt;/h1>
&lt;p>Every producer on the broker is stuck in &lt;code>send()&lt;/code>. Broker-level &lt;code>MemoryPercentUsage&lt;/code> reads 100, the broker log shows the usage manager hitting its memory limit, and upstream services are timing out. But when you list queue depths, almost every queue is empty. One queue holds nearly all the pending messages.&lt;/p>
&lt;p>That one queue has consumed the broker&amp;rsquo;s shared memory pool. Every pending message is charged against its destination&amp;rsquo;s memory accounting and against the broker-wide system memory limit. If either limit is reached, flow control activates. With no per-destination memory limit set, ActiveMQ Classic does not throttle just that queue&amp;rsquo;s producers: it throttles everyone&amp;rsquo;s. This is the default behavior, not a bug.&lt;/p></description></item><item><title>ActiveMQ poison message loop: redelivery storms and DLQ growth</title><link>https://www.netdata.cloud/guides/activemq/activemq-poison-message-redelivery/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-poison-message-redelivery/</guid><description>&lt;h1 id="activemq-poison-message-loop-redelivery-storms-and-dlq-growth">ActiveMQ poison message loop: redelivery storms and DLQ growth&lt;/h1>
&lt;p>A queue looks busy. Consumers are connected, dequeue counters are moving, and yet the business workflow has stalled: orders are not completing, messages are getting older, and somewhere in the broker a queue called &lt;code>ActiveMQ.DLQ&lt;/code> is quietly filling up. This is the poison message loop, and it is one of the most misread failure modes in ActiveMQ Classic because the headline metrics look healthy while useful work has stopped.&lt;/p></description></item><item><title>ActiveMQ prefetch limit: unacked hoarding versus idle consumers</title><link>https://www.netdata.cloud/guides/activemq/activemq-prefetch-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-prefetch-tuning/</guid><description>&lt;h1 id="activemq-prefetch-limit-unacked-hoarding-versus-idle-consumers">ActiveMQ prefetch limit: unacked hoarding versus idle consumers&lt;/h1>
&lt;p>You have four consumers on a queue, producers are sending steadily, and yet throughput is a quarter of what you expect. &lt;code>ConsumerCount&lt;/code> says 4. &lt;code>QueueSize&lt;/code> looks small. &lt;code>DequeueCount&lt;/code> is crawling. The broker looks healthy and the consumers look connected, but three of them are doing nothing.&lt;/p>
&lt;p>This is the classic prefetch mismatch. The ActiveMQ prefetch limit controls how many messages the broker pushes to a consumer before it requires acknowledgments back. It is a client-side buffer the broker fills eagerly. With the default of 1000 for queues, the first consumer to connect can absorb the entire visible backlog into its prefetch buffer, where those messages are invisible to load balancing, no longer read as pending depth, and pinned against broker memory until they are acked.&lt;/p></description></item><item><title>ActiveMQ producer flow control: why send() hangs and producers block silently</title><link>https://www.netdata.cloud/guides/activemq/activemq-producer-flow-control/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-producer-flow-control/</guid><description>&lt;p>Your application is healthy. No exceptions, no error logs, no timeouts. But requests are piling up and every thread that touches the message producer is stuck inside &lt;code>send()&lt;/code>. Thread dumps show all your producer threads parked in the JMS client, waiting on a socket write that never completes. Nothing is wrong on the client. Nothing looks wrong on the broker either, until you check one number: &lt;code>MemoryPercentUsage&lt;/code> is at 100.&lt;/p></description></item><item><title>ActiveMQ queue with zero consumers: ConsumerCount at zero and a growing backlog</title><link>https://www.netdata.cloud/guides/activemq/activemq-queue-no-consumers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-queue-no-consumers/</guid><description>&lt;h1 id="activemq-queue-with-zero-consumers-consumercount-at-zero-and-a-growing-backlog">ActiveMQ queue with zero consumers: ConsumerCount at zero and a growing backlog&lt;/h1>
&lt;p>A production queue shows &lt;code>ConsumerCount=0&lt;/code>, &lt;code>QueueSize&lt;/code> is climbing, and &lt;code>EnqueueCount&lt;/code> keeps ticking up while &lt;code>DequeueCount&lt;/code> is flat. Nobody is draining the queue. Every message that arrives stays in the broker, charged against destination memory and, for persistent messages, written into the KahaDB journal. Left alone, this ends in one of two places: broker memory hits 100% and producer flow control silently blocks every producer, or the store fills and persistent messaging halts entirely.&lt;/p></description></item><item><title>ActiveMQ QueueSize growing: reading the broker's backlog gauge correctly</title><link>https://www.netdata.cloud/guides/activemq/activemq-queue-size-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-queue-size-growing/</guid><description>&lt;h1 id="activemq-queuesize-growing-reading-the-brokers-backlog-gauge-correctly">ActiveMQ QueueSize growing: reading the broker&amp;rsquo;s backlog gauge correctly&lt;/h1>
&lt;p>QueueSize is the first number every operator looks at, and the one most often misread. The common mental model is &amp;ldquo;messages waiting in the queue,&amp;rdquo; but that is not what the broker reports. QueueSize is a broker-maintained gauge, not enqueued-minus-dequeued, and it includes messages already dispatched to consumers but not yet acknowledged. Two queues showing QueueSize=5000 can be in completely different states, and a queue showing QueueSize=0 can still be in trouble.&lt;/p></description></item><item><title>ActiveMQ reconnection storm: transport accept spikes and CPU on handshakes</title><link>https://www.netdata.cloud/guides/activemq/activemq-reconnection-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-reconnection-storm/</guid><description>&lt;h1 id="activemq-reconnection-storm-transport-accept-spikes-and-cpu-on-handshakes">ActiveMQ reconnection storm: transport accept spikes and CPU on handshakes&lt;/h1>
&lt;p>The broker is up. Queues are draining. But the accept rate on the OpenWire connector is running at ten times baseline, JVM thread count and open file descriptors are climbing in a sawtooth, and CPU is pinned even though message throughput is flat or down. Clients are connecting, disconnecting, and reconnecting in a tight loop, and every one of those connections costs the broker real work: a TCP accept, an optional TLS handshake, wire-format negotiation, and a new transport thread.&lt;/p></description></item><item><title>ActiveMQ redelivery rate climbing: rollbacks, nacks, and retry storms</title><link>https://www.netdata.cloud/guides/activemq/activemq-redelivery-rate-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-redelivery-rate-high/</guid><description>&lt;h1 id="activemq-redelivery-rate-climbing-rollbacks-nacks-and-retry-storms">ActiveMQ redelivery rate climbing: rollbacks, nacks, and retry storms&lt;/h1>
&lt;p>Your dequeue rate looks busy, but useful work is not getting done. Messages are dispatched, rolled back or nacked, and dispatched again. The broker spends its cycles retrying instead of making progress, and the redelivery counters climb. This is the &lt;a href="https://www.netdata.cloud/guides/activemq/activemq-how-it-works-in-production/">poison message replay storm pattern&lt;/a> in its early stage, and it leads two things you do not want: dead letter queue growth and silent correctness loss.&lt;/p></description></item><item><title>ActiveMQ shared-storage HA split-brain: two brokers holding the store lock</title><link>https://www.netdata.cloud/guides/activemq/activemq-ha-split-brain/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-ha-split-brain/</guid><description>&lt;h1 id="activemq-shared-storage-ha-split-brain-two-brokers-holding-the-store-lock">ActiveMQ shared-storage HA split-brain: two brokers holding the store lock&lt;/h1>
&lt;p>In a shared-storage HA pair, exactly one ActiveMQ Classic broker is supposed to hold the KahaDB store lock and run active. The second broker sits in a polling loop, waiting for the lock, with its transport connectors down. Split-brain is the state where that invariant breaks: both brokers believe they are active, both accept client connections, and both write to the same KahaDB store on shared storage.&lt;/p></description></item><item><title>ActiveMQ store is full: StorePercentUsage at 100% and persistent messaging halted</title><link>https://www.netdata.cloud/guides/activemq/activemq-store-is-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-store-is-full/</guid><description>&lt;h1 id="activemq-store-is-full-storepercentusage-at-100-and-persistent-messaging-halted">ActiveMQ store is full: StorePercentUsage at 100% and persistent messaging halted&lt;/h1>
&lt;p>StorePercentUsage has hit 100% on the broker MBean and persistent producers have stopped. Their &lt;code>send()&lt;/code> calls are either blocked in flow control or being rejected, depending on your destination policy. Non-persistent traffic may still flow, which makes the outage look partial from the outside.&lt;/p>
&lt;p>This is the disk-side equivalent of the memory wall. Where MemoryPercentUsage at 100% blocks producers against the configured memory limit, StorePercentUsage at 100% blocks persistent producers against the configured &lt;code>storeUsage&lt;/code> limit in &lt;code>activemq.xml&lt;/code>. The broker log shows a &amp;ldquo;Persistent store is Full&amp;rdquo; message, and producers stall silently unless you configured &lt;code>sendFailIfNoSpaceAfterTimeout&lt;/code>.&lt;/p></description></item><item><title>ActiveMQ store usage climbing: the store exhaustion spiral</title><link>https://www.netdata.cloud/guides/activemq/activemq-store-usage-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-store-usage-growing/</guid><description>&lt;h1 id="activemq-store-usage-climbing-the-store-exhaustion-spiral">ActiveMQ store usage climbing: the store exhaustion spiral&lt;/h1>
&lt;p>StorePercentUsage was 62% last month. Last week it was 74%. This morning it crossed 80% and your alert fired. Queue depths look normal, consumers are connected, dequeue rates look healthy. Nothing is obviously on fire, but the persistent store keeps growing and nobody knows why.&lt;/p>
&lt;p>This is the store exhaustion spiral, one of the most common slow-burn failure patterns in ActiveMQ Classic. Unlike the memory-pressure cascade, which develops in minutes, the store spiral develops over days or weeks. Messages accumulate in the KahaDB journal faster than they are acknowledged, journal files pile up on disk, and the store limit creeps closer. When StorePercentUsage reaches 100%, the broker stops accepting persistent messages entirely and every persistent producer blocks. The warning signs were visible for weeks.&lt;/p></description></item><item><title>ActiveMQ StorePercentUsage vs actual disk space: the limit that fires after the disk is already full</title><link>https://www.netdata.cloud/guides/activemq/activemq-store-usage-vs-disk-space/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-store-usage-vs-disk-space/</guid><description>&lt;h1 id="activemq-storepercentusage-vs-actual-disk-space-the-limit-that-fires-after-the-disk-is-already-full">ActiveMQ StorePercentUsage vs actual disk space: the limit that fires after the disk is already full&lt;/h1>
&lt;p>Your broker just stopped accepting persistent messages. Producers are blocked or erroring, the KahaDB partition is at 100% disk usage, and yet your dashboard shows &lt;code>StorePercentUsage&lt;/code> at 62%. The alert you set on store usage never fired. This is one of the most common ActiveMQ monitoring traps: &lt;code>StorePercentUsage&lt;/code> and actual disk usage are two different counters measured against two different denominators, and if you have not reconciled them, one of them is lying to you.&lt;/p></description></item><item><title>ActiveMQ temp store full: non-persistent overflow and silent message drops</title><link>https://www.netdata.cloud/guides/activemq/activemq-temp-store-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-temp-store-full/</guid><description>&lt;h1 id="activemq-temp-store-full-non-persistent-overflow-and-silent-message-drops">ActiveMQ temp store full: non-persistent overflow and silent message drops&lt;/h1>
&lt;p>Non-persistent messages in ActiveMQ Classic live in broker memory. When memory fills, the broker spills them to an on-disk overflow area called the temp store (default &lt;code>data/tmp_storage&lt;/code>). When the temp store fills, non-persistent messaging breaks. Depending on your destination policies, it can break in the worst possible way: messages are silently discarded while producers see no error and consumers just see less traffic.&lt;/p></description></item><item><title>ActiveMQ temporary destination leak: request-reply queues that never close</title><link>https://www.netdata.cloud/guides/activemq/activemq-temporary-destination-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-temporary-destination-leak/</guid><description>&lt;h1 id="activemq-temporary-destination-leak-request-reply-queues-that-never-close">ActiveMQ temporary destination leak: request-reply queues that never close&lt;/h1>
&lt;p>You open the broker&amp;rsquo;s JMX view and the TemporaryQueues attribute lists hundreds or thousands of entries. The count only goes up. Nothing in the application logs looks wrong, request-reply calls mostly work, but the destination count climbs with traffic and never comes back down. That is a temporary destination leak.&lt;/p>
&lt;p>Temporary destinations back the classic JMS request-reply pattern: a client creates a TemporaryQueue, sends a request with the temp queue as the reply-to, and waits for the response. In a healthy system these destinations cycle constantly: created, used, deleted. When the count grows monotonically with request volume, something in that lifecycle is broken, and the broker is accumulating MBeans, metadata, and advisory traffic for destinations that will never be used again.&lt;/p></description></item><item><title>ActiveMQ Too many open files: file descriptor exhaustion and refused connections</title><link>https://www.netdata.cloud/guides/activemq/activemq-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-too-many-open-files/</guid><description>&lt;h1 id="activemq-too-many-open-files-file-descriptor-exhaustion-and-refused-connections">ActiveMQ Too many open files: file descriptor exhaustion and refused connections&lt;/h1>
&lt;p>The broker log shows &lt;code>Too many open files&lt;/code> on the accept path, new clients cannot connect, and existing clients may start timing out. In the worst case the broker also fails to open a new KahaDB journal file during rotation, and now you have a store integrity problem that looks, at first glance, like a disk failure. It is not. The broker has run out of file descriptors.&lt;/p></description></item><item><title>ActiveMQ transport connector not accepting: a live broker that refuses one protocol</title><link>https://www.netdata.cloud/guides/activemq/activemq-transport-connector-not-accepting/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-transport-connector-not-accepting/</guid><description>&lt;h1 id="activemq-transport-connector-not-accepting-a-live-broker-that-refuses-one-protocol">ActiveMQ transport connector not accepting: a live broker that refuses one protocol&lt;/h1>
&lt;p>The broker process is running. The JVM is up, OpenWire clients on 61616 are producing and consuming, and your process-level health check is green. But the AMQP clients on 5672 cannot connect, and their reconnect loops are filling your application logs.&lt;/p>
&lt;p>This partial failure is worse than a full outage in one specific way: most monitoring does not see it. A TCP check against one port, or a process-alive check, tells you nothing about the other four connectors. ActiveMQ Classic runs each transport connector (OpenWire 61616, AMQP 5672, STOMP 61613, MQTT 1883, WebSocket 61614) as its own accept path, and each one can fail independently while the rest of the broker looks healthy.&lt;/p></description></item><item><title>ActiveMQ.DLQ growing: dead letter queue accumulation and poison messages</title><link>https://www.netdata.cloud/guides/activemq/activemq-dlq-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-dlq-growing/</guid><description>&lt;h1 id="activemqdlq-growing-dead-letter-queue-accumulation-and-poison-messages">ActiveMQ.DLQ growing: dead letter queue accumulation and poison messages&lt;/h1>
&lt;p>The queue named &lt;code>ActiveMQ.DLQ&lt;/code> is growing and nobody is consuming from it. That is the default dead letter queue: it has no consumers, no TTL, and no automatic cleanup, so every message that lands there stays there forever.&lt;/p>
&lt;p>Every message in the DLQ is two problems at once. It is a failed business transaction that some consumer gave up on after exhausting redeliveries. And it is a storage leak: DLQ messages are never acknowledged, and in KahaDB a single unacknowledged message pins its entire journal file. A DLQ that grows quietly for weeks is a common root cause behind store exhaustion, which eventually halts persistent messaging on the whole broker.&lt;/p></description></item><item><title>Actona Technologies Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/actona-technologies-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/actona-technologies-inc-snmp-traps/</guid><description/></item><item><title>Adaptec Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/adaptec-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/adaptec-inc-snmp-traps/</guid><description/></item><item><title>Adaptec RAID</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/adaptec-raid/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/adaptec-raid/</guid><description/></item><item><title>Adaptec RAID Monitoring</title><link>https://www.netdata.cloud/monitoring-101/adaptecraid-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/adaptecraid-monitoring/</guid><description>&lt;h2 id="adaptec-raid-monitoring">Adaptec RAID Monitoring&lt;/h2>
&lt;h3 id="what-is-adaptec-raid">What Is Adaptec RAID?&lt;/h3>
&lt;p>Adaptec RAID is a trusted storage solution designed to manage and protect your data using Redundant Array of Independent Disks (RAID) technology. These controllers are widely used for enhancing storage reliability and performance by spreading data across multiple disks.&lt;/p>
&lt;h3 id="monitoring-adaptec-raid-with-netdata">Monitoring Adaptec RAID With Netdata&lt;/h3>
&lt;p>To monitor Adaptec RAID efficiently, Netdata offers a comprehensive monitoring tool that integrates seamlessly with your system. The &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/adaptec_raid/">$name monitoring tool&lt;/a> helps you gain insights into the status and health of your RAID arrays in real time, enabling proactive maintenance and swift troubleshooting.&lt;/p></description></item><item><title>Adtran SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/adtran-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/adtran-snmp-traps/</guid><description/></item><item><title>Adva AG Optical Networking SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/adva-ag-optical-networking-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/adva-ag-optical-networking-snmp-traps/</guid><description/></item><item><title>Advanced Fibre Communications Afc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/advanced-fibre-communications-afc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/advanced-fibre-communications-afc-snmp-traps/</guid><description/></item><item><title>Advantech Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/advantech-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/advantech-co-ltd-snmp-traps/</guid><description/></item><item><title>Aerohive Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aerohive-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aerohive-networks-inc-snmp-traps/</guid><description/></item><item><title>Aethra SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aethra-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aethra-snmp-traps/</guid><description/></item><item><title>Affirmed Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/affirmed-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/affirmed-networks-inc-snmp-traps/</guid><description/></item><item><title>Agent SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/agent-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/agent-snmp-traps/</guid><description/></item><item><title>Agere Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/agere-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/agere-systems-inc-snmp-traps/</guid><description/></item><item><title>Agfeo GmbH Co KG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/agfeo-gmbh-co-kg-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/agfeo-gmbh-co-kg-snmp-traps/</guid><description/></item><item><title>Aginode Germany GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aginode-germany-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aginode-germany-gmbh-snmp-traps/</guid><description/></item><item><title>Airespace Inc Formerly Black Storm Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/airespace-inc-formerly-black-storm-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/airespace-inc-formerly-black-storm-networks-snmp-traps/</guid><description/></item><item><title>Alamos FE2 server</title><link>https://www.netdata.cloud/integrations/data-collection/applications/alamos-fe2-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/alamos-fe2-server/</guid><description/></item><item><title>Alamos FE2 server Monitoring</title><link>https://www.netdata.cloud/monitoring-101/alamos_fe2-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/alamos_fe2-monitoring/</guid><description>&lt;h2 id="alamos-fe2-server-monitoring">Alamos FE2 server Monitoring&lt;/h2>
&lt;h3 id="what-is-alamos-fe2-server">What Is Alamos FE2 server?&lt;/h3>
&lt;p>The Alamos FE2 server is a comprehensive platform that demands diligent monitoring to ensure optimal performance and health. This server acts as a vital component in various infrastructures requiring consistent oversight and data collection for efficient operations.&lt;/p>
&lt;h3 id="monitoring-alamos-fe2-server-with-netdata">Monitoring Alamos FE2 server With Netdata&lt;/h3>
&lt;p>Monitoring the Alamos FE2 server is streamlined with Netdata, which employs an openmetrics (Prometheus) exporter. Netdata is capable of ingesting data from any Prometheus exporter, a significant advantage over traditional setups requiring a dedicated Prometheus server or Grafana for visualization. By leveraging Netdata, users obtain automated dashboards, intelligent alerts, and real-time insights designed to keep the Alamos FE2 server running smoothly.&lt;/p></description></item><item><title>Alaxala Networks Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/alaxala-networks-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/alaxala-networks-corporation-snmp-traps/</guid><description/></item><item><title>Albal Ingenieros S A SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/albal-ingenieros-s-a-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/albal-ingenieros-s-a-snmp-traps/</guid><description/></item><item><title>Albentia Systems S A SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/albentia-systems-s-a-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/albentia-systems-s-a-snmp-traps/</guid><description/></item><item><title>Albis Technologies Ltd Formerly Siemens Switzerland Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/albis-technologies-ltd-formerly-siemens-switzerland-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/albis-technologies-ltd-formerly-siemens-switzerland-ltd-snmp-traps/</guid><description/></item><item><title>Alcatel Lucent</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/alcatel-lucent/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/alcatel-lucent/</guid><description/></item><item><title>Alcatel Lucent ENT</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/alcatel-lucent-ent/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/alcatel-lucent-ent/</guid><description/></item><item><title>Alcatel Lucent Enterprise SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/alcatel-lucent-enterprise-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/alcatel-lucent-enterprise-snmp-traps/</guid><description/></item><item><title>Alcatel Lucent IND</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/alcatel-lucent-ind/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/alcatel-lucent-ind/</guid><description/></item><item><title>Alcatel Lucent Omni Access WLC</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/alcatel-lucent-omni-access-wlc/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/alcatel-lucent-omni-access-wlc/</guid><description/></item><item><title>Alcatel-Lucent BGP</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/alcatel-lucent-bgp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/alcatel-lucent-bgp/</guid><description/></item><item><title>Ale USA Inc Omniswitch SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ale-usa-inc-omniswitch-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ale-usa-inc-omniswitch-snmp-traps/</guid><description/></item><item><title>Alebra Technologies Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/alebra-technologies-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/alebra-technologies-inc-snmp-traps/</guid><description/></item><item><title>Alerta</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/alerta/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/alerta/</guid><description/></item><item><title>Allaire Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/allaire-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/allaire-corporation-snmp-traps/</guid><description/></item><item><title>Allied Telesis Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/allied-telesis-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/allied-telesis-inc-snmp-traps/</guid><description/></item><item><title>Allot Communications SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/allot-communications-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/allot-communications-snmp-traps/</guid><description/></item><item><title>Alpine Linux</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/alpine-linux/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/alpine-linux/</guid><description/></item><item><title>Alpine Optoelectronics Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/alpine-optoelectronics-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/alpine-optoelectronics-inc-snmp-traps/</guid><description/></item><item><title>Alteon Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/alteon-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/alteon-networks-inc-snmp-traps/</guid><description/></item><item><title>Altergy Systems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/altergy-systems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/altergy-systems-snmp-traps/</guid><description/></item><item><title>Alvarion Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/alvarion-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/alvarion-ltd-snmp-traps/</guid><description/></item><item><title>AM2320</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/am2320/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/am2320/</guid><description/></item><item><title>AM2320 Monitoring</title><link>https://www.netdata.cloud/monitoring-101/am2320-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/am2320-monitoring/</guid><description>&lt;h2 id="what-is-am2320">What is AM2320?&lt;/h2>
&lt;p>AM2320 is a temperature and humidity sensor that can be used to measure environment conditions. It features a range of features including a high accuracy of ± 0.5°C and low power consumption. It is suitable for a variety of applications, making it a versatile choice for your project.&lt;/p>
&lt;h2 id="monitoring-am2320-with-netdata">Monitoring AM2320 with Netdata&lt;/h2>
&lt;p>The prerequisites for monitoring AM2320 with Netdata are to have AM2320 and &lt;a href="https://learn.netdata.cloud/docs/cloud/get-started/">Netdata installed&lt;/a> on your system.&lt;/p></description></item><item><title>Amazon CloudWatch</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/amazon-cloudwatch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/amazon-cloudwatch/</guid><description/></item><item><title>Amazon Linux</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/amazon-linux/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/amazon-linux/</guid><description/></item><item><title>Amazon SNS</title><link>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/amazon-sns/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/amazon-sns/</guid><description/></item><item><title>AMD CPU &amp; GPU</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/amd-cpu--gpu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/amd-cpu--gpu/</guid><description/></item><item><title>AMD CPU &amp; GPU Monitoring</title><link>https://www.netdata.cloud/monitoring-101/amd_smi-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/amd_smi-monitoring/</guid><description>&lt;h2 id="amd-cpu--gpu-monitoring">AMD CPU &amp;amp; GPU Monitoring&lt;/h2>
&lt;h3 id="what-is-amd-cpu--gpu">What Is AMD CPU &amp;amp; GPU?&lt;/h3>
&lt;p>AMD CPUs and GPUs form the core of many computational systems, driving performance in everything from gaming computers to data centers that power cloud applications. The AMD System Management Interface (SMI) is crucial for monitoring the performance and health of these processors, helping users optimize hardware performance through precise metrics.&lt;/p>
&lt;h3 id="monitoring-amd-cpu--gpu-with-netdata">Monitoring AMD CPU &amp;amp; GPU With Netdata&lt;/h3>
&lt;p>Monitoring AMD CPU &amp;amp; GPU is essential for ensuring that your systems operate efficiently and without interruption. Netdata leverages the &lt;a href="https://github.com/amd/amd_smi_exporter">AMD SMI Exporter&lt;/a> to monitor these devices. By using an openmetrics (Prometheus) exporter, Netdata can seamlessly ingest metrics without the need for a Prometheus server or Grafana. Once connected, users can enjoy automated dashboards and alerts, making it a comprehensive AMD CPU &amp;amp; GPU monitoring tool.&lt;/p></description></item><item><title>AMD GPU</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/amd-gpu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/amd-gpu/</guid><description/></item><item><title>American Power Conversion Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/american-power-conversion-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/american-power-conversion-corp-snmp-traps/</guid><description/></item><item><title>Amperion Incorporated SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/amperion-incorporated-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/amperion-incorporated-snmp-traps/</guid><description/></item><item><title>An D Cz SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/an-d-cz-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/an-d-cz-snmp-traps/</guid><description/></item><item><title>Ancor Communications SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ancor-communications-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ancor-communications-snmp-traps/</guid><description/></item><item><title>Andover Controls Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/andover-controls-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/andover-controls-corporation-snmp-traps/</guid><description/></item><item><title>Anue</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/anue/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/anue/</guid><description/></item><item><title>Anuesystems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/anuesystems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/anuesystems-snmp-traps/</guid><description/></item><item><title>Aol Netscape Communications Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aol-netscape-communications-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aol-netscape-communications-corp-snmp-traps/</guid><description/></item><item><title>Ap Nederland B.V. SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ap-nederland-b.v.-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ap-nederland-b.v.-snmp-traps/</guid><description/></item><item><title>Apache</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/apache/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/apache/</guid><description/></item><item><title>Apache /server-status exposed: the reconnaissance leak hiding in mod_status</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-server-status-exposed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-server-status-exposed/</guid><description>&lt;h1 id="apache-server-status-exposed-the-reconnaissance-leak-hiding-in-mod_status">Apache /server-status exposed: the reconnaissance leak hiding in mod_status&lt;/h1>
&lt;p>A scanner, pentest, or audit just flagged your Apache server: &lt;code>/server-status&lt;/code> is reachable from the internet. mod_status renders a live view of your server&amp;rsquo;s internals, and anyone who can load it can watch your traffic in near real time.&lt;/p>
&lt;p>The operational problem is that mod_status is also the best monitoring source Apache has. BusyWorkers, the scoreboard, requests per second: all of it comes from &lt;code>/server-status?auto&lt;/code>. So the fix is not &amp;ldquo;disable mod_status.&amp;rdquo; The fix is to restrict who can reach it, verify the restriction holds on every vhost, and keep scraping it locally.&lt;/p></description></item><item><title>Apache %D includes client transfer time: the latency alert that cries wolf</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-percent-d-client-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-percent-d-client-latency/</guid><description>&lt;h1 id="apache-d-includes-client-transfer-time-the-latency-alert-that-cries-wolf">Apache %D includes client transfer time: the latency alert that cries wolf&lt;/h1>
&lt;p>Your p99 latency alert fired at 03:00. You pull the access log, sort by &lt;code>%D&lt;/code>, and find requests with durations of 400, 600, even 800 seconds. The server looks fine: CPU normal, workers mostly idle, backend healthy. The &amp;ldquo;slow&amp;rdquo; requests are large file downloads to clients on slow connections. Apache served the response instantly; the client took thirteen minutes to read it.&lt;/p></description></item><item><title>Apache %T vs %D vs %{ms}T: stop rounding latency to whole seconds</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-percent-t-vs-d/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-percent-t-vs-d/</guid><description>&lt;h1 id="apache-t-vs-d-vs-mst-stop-rounding-latency-to-whole-seconds">Apache %T vs %D vs %{ms}T: stop rounding latency to whole seconds&lt;/h1>
&lt;p>Open your Apache access log and look at the request duration field. If every line shows 0, your monitoring pipeline is probably fine. Your LogFormat is the problem. The &lt;code>%T&lt;/code> directive logs time to serve the request in whole seconds, so any request that completes in under one second logs as 0. On a fleet where most requests finish in well under 100ms, &lt;code>%T&lt;/code> collapses the entire latency distribution into a single useless value.&lt;/p></description></item><item><title>Apache 401/403 floods: credential stuffing, scanner probes, and access denials</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-brute-force-401-403/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-brute-force-401-403/</guid><description>&lt;h1 id="apache-401403-floods-credential-stuffing-scanner-probes-and-access-denials">Apache 401/403 floods: credential stuffing, scanner probes, and access denials&lt;/h1>
&lt;p>Your access log is filling with 401s against a login endpoint, or 403s across paths like &lt;code>/.env&lt;/code> and &lt;code>/wp-admin&lt;/code>. Some of this is the normal background radiation of the public internet. Some of it is an active credential stuffing run against your users. The status code alone does not tell you which; the shape of the traffic does.&lt;/p>
&lt;p>The three patterns look different once you group the log lines:&lt;/p></description></item><item><title>Apache 500 Internal Server Error: modules, handlers, and misconfiguration</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-500-internal-server-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-500-internal-server-error/</guid><description>&lt;h1 id="apache-500-internal-server-error-modules-handlers-and-misconfiguration">Apache 500 Internal Server Error: modules, handlers, and misconfiguration&lt;/h1>
&lt;p>Unlike the 502/503/504 family, which Apache generates while proxying to a broken or exhausted backend, a 500 is generated inside Apache itself or by the handler Apache invoked: a module crashed, a CGI script failed, a rewrite rule looped, an &lt;code>.htaccess&lt;/code> directive is invalid, or the server hit a permission problem reaching the content.&lt;/p>
&lt;p>That distinction is the diagnostic strategy. For a 502 you look at the backend. For a 500 you look at Apache&amp;rsquo;s own error log, because every 500 Apache emits has a corresponding error-log line with the real cause. The access log tells you &lt;em>that&lt;/em> it happened and &lt;em>which URL&lt;/em> triggered it. The error log tells you &lt;em>why&lt;/em>.&lt;/p></description></item><item><title>Apache 502 Bad Gateway: a backend that returned an invalid response</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-502-bad-gateway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-502-bad-gateway/</guid><description>&lt;h1 id="apache-502-bad-gateway-a-backend-that-returned-an-invalid-response">Apache 502 Bad Gateway: a backend that returned an invalid response&lt;/h1>
&lt;p>You are seeing 502 Bad Gateway responses from Apache, usually in bursts, often on some requests and not others. Users report intermittent failures. Apache is running, the port is open, static content may even work. The proxied application path is the thing failing.&lt;/p>
&lt;p>A 502 from Apache is almost never an Apache bug. In a mod_proxy deployment, 502 means Apache acted as a reverse proxy, forwarded the request to a backend, and got back something it could not use: a garbled response, an empty response, a connection closed mid-reply, or a connection that refused to carry the request at all. Apache is surfacing a backend signal.&lt;/p></description></item><item><title>Apache 503 Service Unavailable: worker exhaustion versus proxy pool exhaustion</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-503-service-unavailable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-503-service-unavailable/</guid><description>&lt;h1 id="apache-503-service-unavailable-worker-exhaustion-versus-proxy-pool-exhaustion">Apache 503 Service Unavailable: worker exhaustion versus proxy pool exhaustion&lt;/h1>
&lt;p>Your access log is filling with 503s and users are reporting the site is down. The status code tells you almost nothing: Apache returns 503 for two root causes that look identical in the access log but need completely different fixes.&lt;/p>
&lt;p>The first is frontend worker exhaustion. Every worker slot in the scoreboard is busy, Apache has hit MaxRequestWorkers, and new connections queue in the kernel backlog until they time out. The second is a mod_proxy failure: a backend is down, marked errored, or the per-child proxy connection pool is too small, so Apache refuses to forward requests even though its own workers are mostly idle.&lt;/p></description></item><item><title>Apache 504 Gateway Timeout: slow backends, ProxyTimeout, and worker pile-up</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-504-gateway-timeout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-504-gateway-timeout/</guid><description>&lt;h1 id="apache-504-gateway-timeout-slow-backends-proxytimeout-and-worker-pile-up">Apache 504 Gateway Timeout: slow backends, ProxyTimeout, and worker pile-up&lt;/h1>
&lt;p>A 504 on an Apache-proxied path means a gateway in the request chain timed out. There are two cases:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Apache generated the 504&lt;/strong>: mod_proxy waited longer than the effective proxy timeout for the backend.&lt;/li>
&lt;li>&lt;strong>The backend generated the 504&lt;/strong>: Apache received that status and passed it through, often because the backend was proxying to a dead dependency of its own.&lt;/li>
&lt;/ol>
&lt;p>The access log alone does not distinguish those cases. Correlate the 504 with Apache&amp;rsquo;s error log.&lt;/p></description></item><item><title>Apache 5xx error rate: 500 vs 502 vs 503 vs 504 and what each one means</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-5xx-error-rate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-5xx-error-rate/</guid><description>&lt;p>Your Apache 5xx error rate just spiked. The status code is the first fork in the diagnostic path, and it is a sharp one: a 500 means the failure happened inside Apache or a module, while 502, 503, and 504 are mod_proxy telling you something about the relationship between Apache and a backend. Treating them as one bucket of &amp;ldquo;server errors&amp;rdquo; wastes the best lead you have.&lt;/p>
&lt;p>There is a second trap. A sustained 5xx rate almost always means something real is broken, because 5xx is server-side failure by definition. But the reverse is not true: a proxied application that returns error pages with HTTP 200 is invisible to status-based monitoring. Low 5xx is necessary, not sufficient, evidence of health.&lt;/p></description></item><item><title>Apache accepting deprecated TLS: TLS 1.0/1.1 and weak-cipher exposure</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-tls-protocol-anomalies/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-tls-protocol-anomalies/</guid><description>&lt;h1 id="apache-accepting-deprecated-tls-tls-1011-and-weak-cipher-exposure">Apache accepting deprecated TLS: TLS 1.0/1.1 and weak-cipher exposure&lt;/h1>
&lt;p>If a compliance scan flags &amp;ldquo;TLS 1.0 enabled&amp;rdquo; or &amp;ldquo;weak cipher suites supported&amp;rdquo; on an Apache vhost, the finding is almost always real. Apache&amp;rsquo;s default &lt;code>SSLProtocol&lt;/code> in the 2.4 branch is &lt;code>all -SSLv3&lt;/code>, so TLS 1.0 and TLS 1.1 are offered unless someone explicitly removed them. A vhost that was correct three years ago may still be accepting protocol versions that PCI DSS and modern browser policy treat as deprecated.&lt;/p></description></item><item><title>Apache AH00484: server reached MaxRequestWorkers setting - worker pool exhausted</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-max-request-workers-reached/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-max-request-workers-reached/</guid><description>&lt;h1 id="apache-ah00484-server-reached-maxrequestworkers-setting---worker-pool-exhausted">Apache AH00484: server reached MaxRequestWorkers setting - worker pool exhausted&lt;/h1>
&lt;p>You found this line in the Apache error log:&lt;/p>
&lt;pre tabindex="0">&lt;code>AH00484: server reached MaxRequestWorkers setting, consider raising the MaxRequestWorkers setting
&lt;/code>&lt;/pre>&lt;p>This is Apache explicitly telling you that every worker slot in its pool was occupied and it had nowhere to put a new connection. From that moment, new connections queue in the kernel&amp;rsquo;s TCP listen backlog, and once the backlog fills, clients get connection refused. Users experience timeouts or a dead site while the Apache process looks perfectly healthy in &lt;code>ps&lt;/code>.&lt;/p></description></item><item><title>Apache AH00558: Could not reliably determine the server's fully qualified domain name</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-ah00558-server-name/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-ah00558-server-name/</guid><description>&lt;h1 id="apache-ah00558-could-not-reliably-determine-the-servers-fully-qualified-domain-name">Apache AH00558: Could not reliably determine the server&amp;rsquo;s fully qualified domain name&lt;/h1>
&lt;p>You restart Apache, run &lt;code>apachectl configtest&lt;/code>, or check the error log after a deploy, and you see it again:&lt;/p>
&lt;pre tabindex="0">&lt;code>AH00558: httpd: Could not reliably determine the server&amp;#39;s fully qualified domain name, using 127.0.1.1. Set the &amp;#39;ServerName&amp;#39; directive globally to suppress this message
&lt;/code>&lt;/pre>&lt;p>The short answer: Apache is still serving traffic. This warning does not stop startup, does not drop requests, and is not a security issue. It means no global &lt;code>ServerName&lt;/code> directive is set, so Apache guessed a name for itself and is telling you what it guessed.&lt;/p></description></item><item><title>Apache and nf_conntrack: table full and the invisible packet drops</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-nf-conntrack-table-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-nf-conntrack-table-full/</guid><description>&lt;h1 id="apache-and-nf_conntrack-table-full-and-the-invisible-packet-drops">Apache and nf_conntrack: table full and the invisible packet drops&lt;/h1>
&lt;p>Users report timeouts and intermittent connection failures, the load balancer is flapping the backend out of rotation, and yet everything on the Apache host looks fine. BusyWorkers are normal. The 5xx rate is flat. The listen backlog is empty. CPU and memory are unremarkable. Apache is not logging errors, because from Apache&amp;rsquo;s point of view, nothing is wrong.&lt;/p>
&lt;p>The failure is one layer below Apache, in the kernel&amp;rsquo;s netfilter connection tracking subsystem. When the &lt;code>nf_conntrack&lt;/code> table fills up, the kernel emits &lt;code>nf_conntrack: table full, dropping packet&lt;/code> and starts silently discarding packets. New SYN packets never reach the TCP stack, so they never reach the listen backlog, so they never reach an Apache worker. Clients see their SYNs dropped and retransmit until they give up: a timeout, not a refusal, which makes it look like a network problem.&lt;/p></description></item><item><title>Apache backend response time: telling 'Apache is slow' from 'the backend is slow'</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-backend-response-time/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-backend-response-time/</guid><description>&lt;h1 id="apache-backend-response-time-telling-apache-is-slow-from-the-backend-is-slow">Apache backend response time: telling &amp;lsquo;Apache is slow&amp;rsquo; from &amp;rsquo;the backend is slow&amp;rsquo;&lt;/h1>
&lt;p>Users report the site is slow. Apache is up, requests eventually complete. Someone says &amp;ldquo;Apache is slow,&amp;rdquo; someone else says &amp;ldquo;the app is slow,&amp;rdquo; and both are guessing, because from the outside the two are indistinguishable: the client just sees a slow response.&lt;/p>
&lt;p>When Apache proxies to a backend, the &lt;code>%D&lt;/code> value in your access log bundles three things together: Apache&amp;rsquo;s own processing, time waiting for the backend, and time transferring the response to the client. A slow backend, a saturated Apache, and a client on a bad connection all produce the same inflated &lt;code>%D&lt;/code>. Without separate instrumentation per component, you cannot attribute the latency, and teams routinely burn hours tuning Apache when the database is the problem.&lt;/p></description></item><item><title>Apache balancer member in error state: reading balancer-manager and failover</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-balancer-member-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-balancer-member-error/</guid><description>&lt;h1 id="apache-balancer-member-in-error-state-reading-balancer-manager-and-failover">Apache balancer member in error state: reading balancer-manager and failover&lt;/h1>
&lt;p>You opened balancer-manager, or someone pasted a screenshot into the incident channel, and one of your BalancerMembers shows &lt;code>Err&lt;/code> in the status column. Traffic is still flowing, but the pool is quietly running on fewer backends than you think. If enough members flip to error state, Apache stops proxying entirely and returns 503s even though httpd itself is healthy.&lt;/p>
&lt;p>This state is mod_proxy_balancer doing its job: it detected failures against a backend and pulled that member out of rotation so requests stop dying on it. The problem is that the balancer tells you almost nothing about why, and it keeps the member sidelined on its own retry schedule regardless of whether the backend has recovered.&lt;/p></description></item><item><title>Apache BusyWorkers and IdleWorkers: reading worker utilization from mod_status</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-busyworkers-idleworkers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-busyworkers-idleworkers/</guid><description>&lt;h1 id="apache-busyworkers-and-idleworkers-reading-worker-utilization-from-mod_status">Apache BusyWorkers and IdleWorkers: reading worker utilization from mod_status&lt;/h1>
&lt;p>&lt;code>BusyWorkers&lt;/code> and &lt;code>IdleWorkers&lt;/code> are the two most-quoted numbers from Apache&amp;rsquo;s mod_status output, and the two most frequently misread. Together they are the primary saturation gauge for the server: how much of the worker pool is currently occupied. Read them wrong and you either miss the onset of worker exhaustion or page someone for healthy autoscaling churn.&lt;/p>
&lt;p>This guide covers what the two counters actually count, how to turn them into a utilization ratio you can alert on, why Apache degrades at a cliff edge rather than gradually, and how to distinguish the two very different situations that both show &lt;code>IdleWorkers: 0&lt;/code>.&lt;/p></description></item><item><title>Apache capped by systemd: TasksMax, LimitNOFILE, and MemoryMax override your config</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-systemd-tasksmax/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-systemd-tasksmax/</guid><description>&lt;h1 id="apache-capped-by-systemd-tasksmax-limitnofile-and-memorymax-override-your-config">Apache capped by systemd: TasksMax, LimitNOFILE, and MemoryMax override your config&lt;/h1>
&lt;p>You raised &lt;code>MaxRequestWorkers&lt;/code>, restarted Apache, and it still refuses to spawn more workers. The error log shows no &lt;code>AH00484&lt;/code>, the host has free CPU and RAM, and connections queue and time out. Or Apache was OOM-killed even though &lt;code>free&lt;/code> showed gigabytes available. Or you hit &amp;ldquo;Too many open files&amp;rdquo; at a fraction of the limit you set in &lt;code>/etc/security/limits.conf&lt;/code>.&lt;/p></description></item><item><title>Apache child pid exit signal Segmentation fault: crashing workers</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-segmentation-fault/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-segmentation-fault/</guid><description>&lt;h1 id="apache-child-pid-exit-signal-segmentation-fault-crashing-workers">Apache child pid exit signal Segmentation fault: crashing workers&lt;/h1>
&lt;p>You found this line in the error log:&lt;/p>
&lt;pre tabindex="0">&lt;code>[core:notice] [pid 1234] AH00052: child pid 5678 exit signal Segmentation fault (11)
&lt;/code>&lt;/pre>&lt;p>A child process died on a memory access violation. The parent logged the exit status and spawned a replacement. That is Apache&amp;rsquo;s designed recovery path, and a single occurrence with no user impact is that mechanism working. But a segfault is never &amp;ldquo;normal noise&amp;rdquo;: a bug executed in Apache, a loaded module, or a shared library, and any segfault in production deserves root-cause analysis. Multiple per minute means the service is actively degrading.&lt;/p></description></item><item><title>Apache CLOSE_WAIT and TIME_WAIT: connection leaks versus normal churn</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-connection-states-close-wait/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-connection-states-close-wait/</guid><description>&lt;h1 id="apache-close_wait-and-time_wait-connection-leaks-versus-normal-churn">Apache CLOSE_WAIT and TIME_WAIT: connection leaks versus normal churn&lt;/h1>
&lt;p>You ran &lt;code>ss -tan&lt;/code> on an Apache host and saw thousands of connections that are not ESTABLISHED. Some are in TIME_WAIT, some in CLOSE_WAIT, and the numbers look alarming. The first question is not &amp;ldquo;how do I get rid of them&amp;rdquo; but &amp;ldquo;which of these is actually a problem&amp;rdquo;.&lt;/p>
&lt;p>The two states look similar in a socket listing but mean opposite things. TIME_WAIT is the kernel doing its job after a connection closes cleanly; it is normal churn on any busy web server. CLOSE_WAIT means the remote peer closed the connection and Apache never finished closing its side; a sustained, growing CLOSE_WAIT count is an application-side leak that consumes file descriptors and, eventually, workers.&lt;/p></description></item><item><title>Apache configtest and Include wildcards: catching bad config before it bites</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-config-test-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-config-test-failed/</guid><description>&lt;h1 id="apache-configtest-and-include-wildcards-catching-bad-config-before-it-bites">Apache configtest and Include wildcards: catching bad config before it bites&lt;/h1>
&lt;p>&lt;code>apachectl configtest&lt;/code> (equivalent to &lt;code>httpd -t&lt;/code>) parses your Apache configuration and reports &lt;code>Syntax OK&lt;/code> or a specific error. It is the cheapest safety check in your toolchain, and it is the main thing standing between a bad edit and the stale-config trap: the state where you believe a new configuration is live, but Apache rejected it at reload time and is still running the old one.&lt;/p></description></item><item><title>Apache CPU saturation: TLS handshakes, mod_deflate, mod_rewrite, and mod_security</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-cpu-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-cpu-saturation/</guid><description>&lt;h1 id="apache-cpu-saturation-tls-handshakes-mod_deflate-mod_rewrite-and-mod_security">Apache CPU saturation: TLS handshakes, mod_deflate, mod_rewrite, and mod_security&lt;/h1>
&lt;p>Apache serving static content is almost never CPU-bound. The request path (accept, read, sendfile, log) is cheap. When httpd processes start eating cores, the cause is nearly always something doing per-request or per-connection computation: TLS handshakes, mod_deflate compression, mod_rewrite regex evaluation, or mod_security rule processing.&lt;/p>
&lt;p>The symptom is gradual, not a cliff. Latency creeps up as CPU saturates. Unlike worker exhaustion, which fails hard at 100% utilization, CPU saturation degrades service progressively, which makes it easy to ignore until p99 latency is unacceptable.&lt;/p></description></item><item><title>Apache error log monitoring: severity levels, AH codes, and what to alert on</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-error-log-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-error-log-monitoring/</guid><description>&lt;p>Metrics tell you that something is wrong with Apache. The error log tells you what. When &lt;code>BusyWorkers&lt;/code> pegs at &lt;code>MaxRequestWorkers&lt;/code>, the scoreboard shows saturation, but the error log hands you &lt;code>AH00484&lt;/code>: &amp;ldquo;server reached MaxRequestWorkers setting.&amp;rdquo; When a child dies, the process table shows a respawn; the error log shows &amp;ldquo;Segmentation fault.&amp;rdquo;&lt;/p>
&lt;p>This guide covers how to read the log efficiently: the severity hierarchy, the AH code scheme introduced in 2.4, the patterns worth alerting on, and the configuration gotchas that silently suppress the messages you need.&lt;/p></description></item><item><title>Apache graceful reload ran the old config: the silent stale-configuration trap</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-graceful-reload-stale-config/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-graceful-reload-stale-config/</guid><description>&lt;h1 id="apache-graceful-reload-ran-the-old-config-the-silent-stale-configuration-trap">Apache graceful reload ran the old config: the silent stale-configuration trap&lt;/h1>
&lt;p>You edited a vhost, ran &lt;code>apachectl graceful&lt;/code>, saw no error, and moved on. Hours later someone notices the new TLS certificate is not being served, the new redirect is missing, or the old &lt;code>ProxyPass&lt;/code> target is still receiving traffic. Apache never went down and no alert fired. The reload failed and the old configuration kept running.&lt;/p>
&lt;p>This is one of Apache&amp;rsquo;s worst failure modes because nothing looks broken. The server keeps serving and request rates stay normal. The only thing wrong is that the running configuration is not the configuration on disk, which quietly invalidates every assumption you make while debugging the next incident.&lt;/p></description></item><item><title>Apache graceful restart pile-up: G states, stacked generations, and doubled memory</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-graceful-restart-pileup/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-graceful-restart-pileup/</guid><description>&lt;h1 id="apache-graceful-restart-pile-up-g-states-stacked-generations-and-doubled-memory">Apache graceful restart pile-up: G states, stacked generations, and doubled memory&lt;/h1>
&lt;p>Memory on an Apache host climbs in steps, each step landing a few minutes after a deploy, a config push, or a log rotation. The scoreboard shows an unusual number of &lt;code>G&lt;/code> (gracefully finishing) workers. &lt;code>ps&lt;/code> shows far more httpd children than &lt;code>MaxRequestWorkers&lt;/code> should allow. Nothing is erroring yet, but the host is drifting toward swap, and the next restart will make it worse.&lt;/p></description></item><item><title>Apache HTTPD monitoring checklist: the signals every production web server needs</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-monitoring-checklist/</guid><description>&lt;h1 id="apache-httpd-monitoring-checklist-the-signals-every-production-web-server-needs">Apache HTTPD monitoring checklist: the signals every production web server needs&lt;/h1>
&lt;p>This is a working checklist for engineers running Apache HTTPD in production. It organizes the signals worth collecting into four maturity levels, from &amp;ldquo;is the process alive&amp;rdquo; through &amp;ldquo;which worker is stuck and why.&amp;rdquo; Use it to audit an existing setup for gaps or to build one without over-instrumenting on day one.&lt;/p>
&lt;p>Two things before the list. First, every saturation signal in Apache is interpreted through the active Multi-Processing Module (MPM), so the checklist starts there. Second, the levels are cumulative: Level 2 assumes Level 1 is in place. Do not skip ahead. The scoreboard state distribution is useless if nobody is watching the error log for &lt;code>AH00484&lt;/code>.&lt;/p></description></item><item><title>Apache HTTPD monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-monitoring-maturity-model/</guid><description>&lt;h1 id="apache-httpd-monitoring-maturity-model-from-survival-to-expert">Apache HTTPD monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most teams monitor Apache at whatever level their last outage forced them to reach. A disk fills up, the site goes down, and &amp;ldquo;check log disk space&amp;rdquo; appears in the runbook. A backend hangs, workers exhaust, and &amp;ldquo;watch BusyWorkers&amp;rdquo; gets added. This model organizes that organic growth into four deliberate levels so you can see what you have, what you are missing, and which blind spot will produce your next incident.&lt;/p></description></item><item><title>Apache keepalive consuming workers: KeepAliveTimeout, the K state, and MPM choice</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-keepalive-consuming-workers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-keepalive-consuming-workers/</guid><description>&lt;h1 id="apache-keepalive-consuming-workers-keepalivetimeout-the-k-state-and-mpm-choice">Apache keepalive consuming workers: KeepAliveTimeout, the K state, and MPM choice&lt;/h1>
&lt;p>The symptom looks like a capacity problem: Apache logs &lt;code>AH00484: server reached MaxRequestWorkers setting&lt;/code>, new connections start queuing, and users see slow responses or 503s. But the request rate is low, and nothing in CPU, memory, or bandwidth explains it. Then you open the scoreboard and see it: a wall of &lt;code>K&lt;/code> states. Most of your workers are not serving requests. They are parked on idle keepalive connections, waiting for a next request that may never come.&lt;/p></description></item><item><title>Apache killed by the OOM killer: the memory exhaustion cascade</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-oom-killer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-oom-killer/</guid><description>&lt;h1 id="apache-killed-by-the-oom-killer-the-memory-exhaustion-cascade">Apache killed by the OOM killer: the memory exhaustion cascade&lt;/h1>
&lt;p>Your monitoring says Apache is up. The parent process exists, the port is listening, but the site is down or intermittently dead. When you finally check the kernel log, there it is: &lt;code>Out of memory: Killed process ... (httpd)&lt;/code>. Not once. Dozens of times, at regular intervals, going back hours.&lt;/p>
&lt;p>This is the Apache OOM death spiral. The kernel OOM killer terminates httpd child processes because total memory demand exceeded RAM. The Apache parent survives, respawns the children, the new children immediately start serving queued requests and allocating memory, and the OOM killer kills them again. The service looks &amp;ldquo;running&amp;rdquo; while serving little or nothing.&lt;/p></description></item><item><title>Apache listen queue overflow: Recv-Q growth, ListenBacklog, and refused connections</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-listen-queue-overflow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-listen-queue-overflow/</guid><description>&lt;h1 id="apache-listen-queue-overflow-recv-q-growth-listenbacklog-and-refused-connections">Apache listen queue overflow: Recv-Q growth, ListenBacklog, and refused connections&lt;/h1>
&lt;p>Users report &amp;ldquo;connection refused&amp;rdquo; or intermittent timeouts, but Apache is running, the port is open, and a TCP connect from localhost sometimes works. The load balancer flaps the backend in and out of rotation. Nothing in the access log explains it, because the failing connections never got far enough to be logged.&lt;/p>
&lt;p>This is the signature of listen queue overflow. Connections complete the TCP handshake in the kernel and sit in the accept queue waiting for an Apache worker to call &lt;code>accept()&lt;/code>. When workers cannot keep up, the queue fills, and the kernel starts ignoring or resetting new connection attempts. The server looks up. It is effectively down for a slice of arriving connections.&lt;/p></description></item><item><title>Apache ListenBacklog vs net.core.somaxconn: the silently truncated accept queue</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-somaxconn-listenbacklog/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-somaxconn-listenbacklog/</guid><description>&lt;h1 id="apache-listenbacklog-vs-netcoresomaxconn-the-silently-truncated-accept-queue">Apache ListenBacklog vs net.core.somaxconn: the silently truncated accept queue&lt;/h1>
&lt;p>You tuned Apache for burst absorption. You set &lt;code>ListenBacklog 2048&lt;/code> (or left the default 511, reasoning it was generous). Then a traffic spike arrived, workers saturated, and connections were refused far earlier than your capacity model predicted. The scoreboard and &lt;code>MaxRequestWorkers&lt;/code> get the blame, but the real culprit is often one layer down: the kernel quietly rewrote your backlog at &lt;code>listen()&lt;/code> time and never told anyone.&lt;/p></description></item><item><title>Apache log rotation losing lines: logrotate copytruncate vs graceful reopen</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-log-rotation-copytruncate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-log-rotation-copytruncate/</guid><description>&lt;h1 id="apache-log-rotation-losing-lines-logrotate-copytruncate-vs-graceful-reopen">Apache log rotation losing lines: logrotate copytruncate vs graceful reopen&lt;/h1>
&lt;p>You notice it after the fact: the access log has a gap. Requests you know happened, because the application processed them and the client got a response, never appear in any log file. Not in the current log, not in the rotated one. The gap lines up with the moment logrotate ran.&lt;/p>
&lt;p>Or the opposite symptom: after rotation, Apache keeps writing to the old file. &lt;code>access.log.1&lt;/code> grows for days while &lt;code>access.log&lt;/code> stays empty. Or every night at rotation time you see a blip of dropped connections and a spike of 499s or client resets, because the postrotate script does a hard restart instead of a graceful one.&lt;/p></description></item><item><title>Apache log stall deadlock: workers stuck in the L state, throughput at zero</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-log-stall-deadlock/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-log-stall-deadlock/</guid><description>&lt;h1 id="apache-log-stall-deadlock-workers-stuck-in-the-l-state-throughput-at-zero">Apache log stall deadlock: workers stuck in the L state, throughput at zero&lt;/h1>
&lt;p>The Apache parent process is running. The port is open. A TCP connect succeeds. But nothing is served: requests hang, throughput is at or near zero, and the scoreboard is a wall of &lt;code>L&lt;/code> characters. Every worker has finished its request and is now blocked writing the log line for it.&lt;/p>
&lt;p>Workers write access log entries synchronously at the end of the request cycle, before returning to the idle pool. If that write blocks, because the log filesystem is full or a piped logging program has stopped reading, the worker never becomes available again. New workers accept new connections, finish those requests, and block on the same write. Within minutes the entire worker pool is frozen in the Logging state.&lt;/p></description></item><item><title>Apache MaxConnectionsPerChild: bounding leaky modules by recycling children</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-maxconnectionsperchild/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-maxconnectionsperchild/</guid><description>&lt;h1 id="apache-maxconnectionsperchild-bounding-leaky-modules-by-recycling-children">Apache MaxConnectionsPerChild: bounding leaky modules by recycling children&lt;/h1>
&lt;p>The symptom usually arrives as a slow memory problem, not a clean Apache failure. Per-child RSS climbs for hours or days, total Apache memory grows linearly, swap starts to move, and then the kernel OOM killer begins shooting httpd children. The parent respawns them, the new children leak too, and the host enters a respawn and OOM loop.&lt;/p>
&lt;p>When that pattern is present, &lt;code>MaxConnectionsPerChild 0&lt;/code> is often part of the story. The default value of 0 means children never recycle. A leaky module, or plain APR pool fragmentation in a long-lived prefork child, gets unlimited time to grow. Setting a finite value, commonly 5000 to 10000, forces each child to exit after handling that many connections so the parent can replace it with a fresh, smaller process.&lt;/p></description></item><item><title>Apache MaxRequestWorkers tuning: sizing the worker pool against memory</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-maxrequestworkers-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-maxrequestworkers-tuning/</guid><description>&lt;h1 id="apache-maxrequestworkers-tuning-sizing-the-worker-pool-against-memory">Apache MaxRequestWorkers tuning: sizing the worker pool against memory&lt;/h1>
&lt;p>MaxRequestWorkers is the single most misconfigured directive in Apache HTTPD. The failure pattern is always the same: someone picks a round number like 1000, deploys it on a 4GB server running mod_php children at 50MB RSS each, and the arithmetic (1000 x 50MB = 50GB of potential demand on 4GB of RAM) guarantees an OOM cascade the first time traffic reaches the limit. The setting feels like a performance knob. It is a memory budget.&lt;/p></description></item><item><title>Apache Monitoring</title><link>https://www.netdata.cloud/monitoring-101/apache-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/apache-monitoring/</guid><description>&lt;h2 id="apache-monitoring">Apache Monitoring&lt;/h2>
&lt;h3 id="what-is-apache">What Is Apache?&lt;/h3>
&lt;p>Apache is one of the most widely used web servers in the world. Officially known as the Apache HTTP Server, it was developed and is maintained by an open community under the auspices of the Apache Software Foundation. Since its inception, Apache has grown into an essential component of the web infrastructure. For more details, visit the &lt;a href="https://httpd.apache.org/">official Apache website&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-apache-with-netdata">Monitoring Apache With Netdata&lt;/h3>
&lt;p>Monitoring Apache is crucial to ensuring optimal web server performance, resource utilization, and improving user experiences. Netdata&amp;rsquo;s real-time monitoring solution provides instant insights into the performance and status of your Apache server. The Netdata Agent, when deployed, offers a wide array of metrics and visualizations to help you keep track of connections, request rates, bandwidth usage, and more. For a full demonstration, check out our &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">Live Demo&lt;/a>.&lt;/p></description></item><item><title>Apache No space left on device: a full log disk that stops the server serving</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-no-space-left-on-device/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-no-space-left-on-device/</guid><description>&lt;h1 id="apache-no-space-left-on-device-a-full-log-disk-that-stops-the-server-serving">Apache No space left on device: a full log disk that stops the server serving&lt;/h1>
&lt;p>The error string operators usually search for is &lt;code>(28)No space left on device&lt;/code> in the Apache error log. Sometimes it appears at startup and Apache refuses to run. The nastier version appears at runtime: the site goes dark, but the process is alive, the port is open, and TCP connections are accepted. Health checks that only test the socket say everything is fine.&lt;/p></description></item><item><title>Apache OCSP stapling failure: the silent handshake that adds client latency</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-ocsp-stapling-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-ocsp-stapling-failure/</guid><description>&lt;h1 id="apache-ocsp-stapling-failure-the-silent-handshake-that-adds-client-latency">Apache OCSP stapling failure: the silent handshake that adds client latency&lt;/h1>
&lt;p>OCSP stapling failure is a textbook &amp;ldquo;silently catastrophic&amp;rdquo; signal: nothing in the error log at default levels, no 5xx, no worker pile-up, and yet every new TLS handshake is slower than it should be. When Apache cannot fetch or attach a stapled OCSP response, each client that wants revocation information makes its own OCSP request to the CA&amp;rsquo;s responder, adding hundreds of milliseconds before the handshake completes. The server looks healthy from the inside. The slowness only exists on the client side.&lt;/p></description></item><item><title>Apache open-file limits: raising ulimit and systemd LimitNOFILE correctly</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-ulimit-nofile/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-ulimit-nofile/</guid><description>&lt;h1 id="apache-open-file-limits-raising-ulimit-and-systemd-limitnofile-correctly">Apache open-file limits: raising ulimit and systemd LimitNOFILE correctly&lt;/h1>
&lt;p>The default per-process open-file limit on most Linux distributions is 1024, far too low for production Apache. Every client connection, every backend proxy connection, every log file, and every pipe consumes a file descriptor in every child process, and 1024 runs out long before &lt;code>MaxRequestWorkers&lt;/code> does. When a child hits the limit, the failure is a cliff edge: &amp;ldquo;Too many open files&amp;rdquo; in the error log, failed accepts, failed backend connections, and intermittent 5xx responses, often in only some children at first, which makes the symptom look random.&lt;/p></description></item><item><title>Apache per-child RSS climbing: the memory-leak slow death</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-memory-leak-rss-growth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-memory-leak-rss-growth/</guid><description>&lt;h1 id="apache-per-child-rss-climbing-the-memory-leak-slow-death">Apache per-child RSS climbing: the memory-leak slow death&lt;/h1>
&lt;p>Your Apache server has been up for three weeks. Nothing is erroring. Request rate is normal, latency is fine, the scoreboard looks healthy. But free memory keeps shrinking, swap usage is creeping up, and last night the OOM killer took out two httpd children. The parent respawned them, and the cycle started over.&lt;/p>
&lt;p>This is the memory-leak slow death: a module loaded into Apache (most commonly mod_php, mod_perl, or a custom module) leaks a small amount of memory per request. Each child&amp;rsquo;s RSS grows monotonically over hours or days. Nothing looks broken until the system runs out of memory, the OOM killer starts shooting children, and the parent respawns fresh ones that also leak. You are now in a kill-and-respawn spiral.&lt;/p></description></item><item><title>Apache piped logging failures: when rotatelogs dies and workers SIGPIPE</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-piped-logging-rotatelogs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-piped-logging-rotatelogs/</guid><description>&lt;h1 id="apache-piped-logging-failures-when-rotatelogs-dies-and-workers-sigpipe">Apache piped logging failures: when rotatelogs dies and workers SIGPIPE&lt;/h1>
&lt;p>Piped logging, where &lt;code>ErrorLog&lt;/code> or &lt;code>CustomLog&lt;/code> points at a program instead of a file (&lt;code>CustomLog &amp;quot;| /usr/bin/rotatelogs /var/log/httpd/access_log.%Y-%m-%d 86400&amp;quot; combined&lt;/code>), adds a new failure mode to Apache: the log program itself. When rotatelogs (or whatever sits at the end of the pipe) dies, wedges, or cannot write because its target disk is full, the pipe breaks. Log writes then fail silently, or worse, the Apache children writing to the broken pipe receive SIGPIPE and can crash.&lt;/p></description></item><item><title>Apache proxy connection refused: AH01114 / (111) Connection refused to backend</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-proxy-connection-refused/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-proxy-connection-refused/</guid><description>&lt;h1 id="apache-proxy-connection-refused-ah01114--111-connection-refused-to-backend">Apache proxy connection refused: AH01114 / (111) Connection refused to backend&lt;/h1>
&lt;p>Your error log is filling with lines like these:&lt;/p>
&lt;pre tabindex="0">&lt;code>AH00957: HTTP: attempt to connect to 127.0.0.1:8080 (localhost) failed
AH01114: HTTP: failed to make connection to backend: localhost
&lt;/code>&lt;/pre>&lt;p>Clients see 502 or 503 responses. Apache itself is running fine. The failure is on the outbound leg: &lt;code>mod_proxy&lt;/code> tried to open a TCP connection to the backend and the kernel told it no.&lt;/p></description></item><item><title>Apache Pulsar</title><link>https://www.netdata.cloud/integrations/data-collection/databases/apache-pulsar/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/apache-pulsar/</guid><description/></item><item><title>Apache Pulsar abandoned subscriptions: cursor leaks that pin storage forever</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-subscription-cursor-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-subscription-cursor-leak/</guid><description>&lt;h1 id="apache-pulsar-abandoned-subscriptions-cursor-leaks-that-pin-storage-forever">Apache Pulsar abandoned subscriptions: cursor leaks that pin storage forever&lt;/h1>
&lt;p>Bookie disk usage climbs steadily across the cluster, but publish rates are flat and there is no traffic spike. Per-subscription backlogs look fine for the subscriptions you know about. When you try to delete an old topic to reclaim space, the operation fails with a message about active subscriptions. The topic has not had a consumer in weeks.&lt;/p>
&lt;p>This is the signature of a silent subscription cursor leak. Applications create dynamically-named durable subscriptions and then disconnect without unsubscribing. Each abandoned subscription leaves behind a cursor pinned at its last acknowledged position. Pulsar cannot delete any message after that cursor&amp;rsquo;s position until the subscription acknowledges past it, which never happens because no consumer is connected. Storage grows monotonically, and the growth is invisible in per-subscription backlog dashboards because those dashboards only track subscriptions the team knows about.&lt;/p></description></item><item><title>Apache Pulsar active connections climbing: connection leaks and file descriptor exhaustion</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-active-connections-climbing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-active-connections-climbing/</guid><description>&lt;h1 id="apache-pulsar-active-connections-climbing-connection-leaks-and-file-descriptor-exhaustion">Apache Pulsar active connections climbing: connection leaks and file descriptor exhaustion&lt;/h1>
&lt;p>&lt;code>pulsar_active_connections&lt;/code> has been drifting upward for weeks. Not spiking, not crashing, just rising a few connections a day. Then one morning new producers start failing to connect, or the broker dies with &lt;code>OutOfDirectMemoryError&lt;/code>, and the postmortem shows the leak was visible the whole time.&lt;/p>
&lt;p>Every TCP connection costs the broker one file descriptor and a slice of Netty direct memory. Connections that are established but never closed accumulate silently until one of those two resources hits its ceiling, at which point the failure is abrupt: refused connections or a crash, with no graceful degradation in between.&lt;/p></description></item><item><title>Apache Pulsar authentication failures: expired tokens, wrong credentials, and brute force</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-authentication-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-authentication-failures/</guid><description>&lt;h1 id="apache-pulsar-authentication-failures-expired-tokens-wrong-credentials-and-brute-force">Apache Pulsar authentication failures: expired tokens, wrong credentials, and brute force&lt;/h1>
&lt;p>Authentication failures in Apache Pulsar surface through &lt;code>pulsar_authentication_failures_total&lt;/code> and &lt;code>AuthenticationException&lt;/code> entries in broker logs. In a stable environment, this counter sits near zero. When it spikes, the temporal pattern matters more than the absolute volume: a single burst from a known client after a credential rotation is normal; sporadic low-rate failures from production IPs point to misconfiguration; a sustained flood from unknown sources signals brute force or credential compromise.&lt;/p></description></item><item><title>Apache Pulsar authorization failures: the AuthorizationException only the logs will show</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-authorization-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-authorization-failures/</guid><description>&lt;h1 id="apache-pulsar-authorization-failures-the-authorizationexception-only-the-logs-will-show">Apache Pulsar authorization failures: the AuthorizationException only the logs will show&lt;/h1>
&lt;p>A producer or consumer authenticates successfully but cannot produce, consume, or manage a topic. The client receives an &lt;code>AuthorizationException&lt;/code> or a generic permission error. Your dashboard shows nothing unusual because Pulsar exposes no authorization failure counter in its Prometheus metrics. The only evidence is in broker logs.&lt;/p>
&lt;p>This is a common blind spot. Authentication failures have a dedicated metric (&lt;code>pulsar_authentication_failures_total&lt;/code>), but authorization failures do not. A client with valid credentials but missing role permissions fails silently from a metrics perspective. You find out when an application team reports errors or when someone greps the broker log.&lt;/p></description></item><item><title>Apache Pulsar AutoRecovery stalled: under-replicated ledgers that never heal</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-autorecovery-stalled/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-autorecovery-stalled/</guid><description>&lt;h1 id="apache-pulsar-autorecovery-stalled-under-replicated-ledgers-that-never-heal">Apache Pulsar AutoRecovery stalled: under-replicated ledgers that never heal&lt;/h1>
&lt;p>Under-replicated ledgers are the durability canary for your Pulsar cluster. When a bookie fails, ledger fragments on it lose redundancy and the under-replicated count spikes. AutoRecovery is supposed to drive that count back to zero. When the count stays flat or grows instead of trending down, recovery has stalled.&lt;/p>
&lt;p>A stall is not slow recovery. Slow recovery means the count is decreasing, just not fast enough. A stall means it is frozen or climbing. The diagnostic paths and fixes are completely different.&lt;/p></description></item><item><title>Apache Pulsar backlog age vs size: the latency depth alone cannot show</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-backlog-age-vs-size/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-backlog-age-vs-size/</guid><description>&lt;h1 id="apache-pulsar-backlog-age-vs-size-the-latency-depth-alone-cannot-show">Apache Pulsar backlog age vs size: the latency depth alone cannot show&lt;/h1>
&lt;p>Backlog size tells you how much data is waiting. Backlog age tells you how long it has been waiting. Monitoring only one leaves you blind to entire classes of consumer health failures.&lt;/p>
&lt;p>A stable backlog of 1,000 entries is unremarkable if those entries are 5 seconds old. The same 1,000 entries from 5 hours ago means a cursor is stuck, a consumer is down, or an application has silently stopped acknowledging. Size alone cannot distinguish these scenarios. Age can.&lt;/p></description></item><item><title>Apache Pulsar backlog quota exceeded: producers held or rejected when consumers stall</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-backlog-quota-exceeded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-backlog-quota-exceeded/</guid><description>&lt;h1 id="apache-pulsar-backlog-quota-exceeded-producers-held-or-rejected-when-consumers-stall">Apache Pulsar backlog quota exceeded: producers held or rejected when consumers stall&lt;/h1>
&lt;p>Producers are timing out or receiving errors, but the root cause is not on the producer side. A subscription backlog has crossed a configured quota, and the backlog policy has kicked in. Depending on the policy, producers are now blocked, rejected, or the oldest messages are being silently deleted. The symptom is a producer outage or data loss, but the root cause is a consumer that stopped keeping up.&lt;/p></description></item><item><title>Apache Pulsar bookie add-entry queue not draining: writes arriving faster than the disk can commit</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-add-entry-in-progress-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-add-entry-in-progress-growing/</guid><description>&lt;h1 id="apache-pulsar-bookie-add-entry-queue-not-draining-writes-arriving-faster-than-the-disk-can-commit">Apache Pulsar bookie add-entry queue not draining: writes arriving faster than the disk can commit&lt;/h1>
&lt;p>When &lt;code>bookkeeper_server_ADD_ENTRY_IN_PROGRESS&lt;/code> grows and does not drain to near-zero within seconds of a traffic burst, the bookie write path is saturated. Writes are arriving faster than the journal disk can fsync them, or the write thread is blocked by GC. Left unaddressed, this causes broker-side timeouts, producer failures, and throughput collapse.&lt;/p>
&lt;p>This metric is a gauge, not a counter. Absolute value matters less than trend. Spikes during traffic bursts are normal; sustained positive growth is not. The queue fills and stays filled because the bookie cannot commit entries as fast as they arrive. By the time broker publish latency spikes, the bookie has already been saturated for seconds or minutes.&lt;/p></description></item><item><title>Apache Pulsar bookie disk filling: runway to read-only and how to reclaim space</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-bookie-disk-filling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-bookie-disk-filling/</guid><description>&lt;h1 id="apache-pulsar-bookie-disk-filling-runway-to-read-only-and-how-to-reclaim-space">Apache Pulsar bookie disk filling: runway to read-only and how to reclaim space&lt;/h1>
&lt;p>Bookie disk usage decides whether your cluster can accept writes. When a bookie&amp;rsquo;s ledger directories fill past the configured threshold (default 95%), the bookie transitions to read-only mode and stops accepting new entries. If enough bookies go read-only, the write quorum for affected ledgers cannot be satisfied, and producers start seeing errors.&lt;/p>
&lt;p>BookKeeper needs disk headroom to compact entry logs and reclaim space from deleted ledgers. &lt;!-- TODO: verify whether major compaction is actually suspended at diskUsageWarnThreshold (default 0.90) by default, or whether suspension only occurs at diskUsageThreshold (default 0.95) when isForceGCAllowWhenNoSpace=false --> At 95% (&lt;code>diskUsageThreshold&lt;/code>), the bookie goes read-only and suspends GC entirely. At that point, shortening retention or deleting topics will not reclaim space because the compaction that rewrites entry logs is itself suspended.&lt;/p></description></item><item><title>Apache Pulsar bookie failure cascade: recovery I/O that topples surviving bookies</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-bookie-failure-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-bookie-failure-cascade/</guid><description>&lt;h1 id="apache-pulsar-bookie-failure-cascade-recovery-io-that-topples-surviving-bookies">Apache Pulsar bookie failure cascade: recovery I/O that topples surviving bookies&lt;/h1>
&lt;p>Bookies are failing one at a time. Each time one drops, AutoRecovery starts replicating its ledgers to the survivors. The recovery reads and writes add I/O load to bookies already handling foreground traffic. Journal sync latency climbs, publish latency follows, then the next bookie starts timing out. It fails, and the cycle accelerates.&lt;/p>
&lt;p>This is the bookie failure cascade. The root cause is not the initial bookie failure; it is the recovery mechanism competing with production traffic on bookies near their I/O ceiling. With default settings, AutoRecovery triggers immediately (&lt;code>lostBookieRecoveryDelay=0&lt;/code>) and reads entries in batches of 100 (&lt;code>rereplicationEntryBatchSize=100&lt;/code>) from surviving bookies, then writes them locally.&lt;/p></description></item><item><title>Apache Pulsar bookie journal and ledger storage on one disk: the #1 architecture mistake</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-journal-storage-shared-disk/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-journal-storage-shared-disk/</guid><description>&lt;h1 id="apache-pulsar-bookie-journal-and-ledger-storage-on-one-disk-the-1-architecture-mistake">Apache Pulsar bookie journal and ledger storage on one disk: the #1 architecture mistake&lt;/h1>
&lt;p>Every BookKeeper bookie has two storage responsibilities with fundamentally different I/O profiles. The journal is a write-ahead log: sequential writes, fsync&amp;rsquo;d per entry or batch, on the critical path of every producer acknowledgment. Entry logs and ledger indexes (managed by DbLedgerStorage, the default storage backend) store message payloads and serve random reads whenever consumers catch up on historical data.&lt;/p></description></item><item><title>Apache Pulsar bookie read latency high: catch-up reads competing with the write path</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-bookie-read-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-bookie-read-latency-high/</guid><description>&lt;h1 id="apache-pulsar-bookie-read-latency-high-catch-up-reads-competing-with-the-write-path">Apache Pulsar bookie read latency high: catch-up reads competing with the write path&lt;/h1>
&lt;p>When &lt;code>bookkeeper_server_READ_ENTRY_REQUEST&lt;/code> or &lt;code>bookie_BOOKIE_READ_ENTRY&lt;/code> P99 latency rises well above baseline, the cause is usually not an isolated disk problem. It is consumers draining backlog (catch-up reads) generating a high volume of cold reads that bypass both the broker&amp;rsquo;s managed ledger cache and the bookie&amp;rsquo;s read cache, hitting ledger storage disks directly. If journal and ledger directories share a physical device, those read I/O operations compete with journal fsyncs and the write path degrades too.&lt;/p></description></item><item><title>Apache Pulsar bookie read-only: disk full and bookie_SERVER_STATUS at zero</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-bookie-read-only/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-bookie-read-only/</guid><description>&lt;h1 id="apache-pulsar-bookie-read-only-disk-full-and-bookie_server_status-at-zero">Apache Pulsar bookie read-only: disk full and bookie_SERVER_STATUS at zero&lt;/h1>
&lt;p>&lt;code>bookie_SERVER_STATUS == 0&lt;/code> means the bookie has transitioned to read-only mode and is no longer accepting writes. In most cases this is a self-protective response to ledger disk usage crossing &lt;code>diskUsageThreshold&lt;/code> (default 0.95). The bookie keeps serving reads, but any topic whose ensemble includes this bookie may fail to meet its write quorum, causing broker-side write errors and producer timeouts.&lt;/p></description></item><item><title>Apache Pulsar broker down: telling a dead broker from a fenced one</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-broker-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-broker-down/</guid><description>&lt;h1 id="apache-pulsar-broker-down-telling-a-dead-broker-from-a-fenced-one">Apache Pulsar broker down: telling a dead broker from a fenced one&lt;/h1>
&lt;p>Your alerting fired: the broker&amp;rsquo;s HTTP admin endpoint on :8080 has been unreachable for more than two minutes, and the broker was previously running, so this is not a fresh deploy or a rolling restart. Pager says &amp;ldquo;broker down.&amp;rdquo; That phrase hides two very different incidents.&lt;/p>
&lt;p>The first is a hard death: the JVM crashed, got OOM-killed, or the host failed. The process is gone and the OS will tell you so in seconds. The second is a fenced broker: the process is alive, maybe even accepting TCP connections, but it has lost its ZooKeeper session and with it ownership of every namespace bundle it served. From the client&amp;rsquo;s perspective the broker is down. From the process table&amp;rsquo;s perspective it is fine. The fix, the blast radius, and the forensics are completely different for each.&lt;/p></description></item><item><title>Apache Pulsar broker GC death spiral: heap pressure, stop-the-world pauses, and lost topic ownership</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-broker-gc-death-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-broker-gc-death-spiral/</guid><description>&lt;h1 id="apache-pulsar-broker-gc-death-spiral-heap-pressure-stop-the-world-pauses-and-lost-topic-ownership">Apache Pulsar broker GC death spiral: heap pressure, stop-the-world pauses, and lost topic ownership&lt;/h1>
&lt;p>A broker is losing topic ownership, recovering, and losing it again. Clients reconnect in bursts. Lookup failures climb. Publish latency is erratic even though traffic is flat and CPU looks busy for no obvious reason. The broker process never actually dies, which is why restarts and instance health checks keep &amp;ldquo;fixing&amp;rdquo; it for ten minutes at a time.&lt;/p></description></item><item><title>Apache Pulsar broker hotspot: one broker owning far more topics than the rest</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-broker-hotspot-topic-count/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-broker-hotspot-topic-count/</guid><description>&lt;h1 id="apache-pulsar-broker-hotspot-one-broker-owning-far-more-topics-than-the-rest">Apache Pulsar broker hotspot: one broker owning far more topics than the rest&lt;/h1>
&lt;p>You have a multi-broker Pulsar cluster where one broker carries a disproportionate share of topic ownership. Cluster-wide averages look acceptable, but that single broker shows elevated GC pressure, higher publish latency, more active connections, or a larger heap footprint than its peers. The imbalance may have built for hours or days without triggering an alert because aggregate metrics hid it.&lt;/p></description></item><item><title>Apache Pulsar broker lookup failures: new clients cannot find their topic</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-broker-lookup-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-broker-lookup-failures/</guid><description>&lt;h1 id="apache-pulsar-broker-lookup-failures-new-clients-cannot-find-their-topic">Apache Pulsar broker lookup failures: new clients cannot find their topic&lt;/h1>
&lt;p>Your dashboards look fine. Messages are flowing, publish latency is normal, consumers are draining backlog. But the new service you just deployed cannot start: its producer hangs or fails trying to connect to its topic. Restarting it does not help. Existing clients are unaffected.&lt;/p>
&lt;p>This is the classic Pulsar grey failure. Topic lookup is how every producer and consumer discovers which broker owns its topic. Existing connections keep working because they already resolved ownership. New traffic fails when lookups cannot resolve. Process health checks, message rates, and bookie metrics can all look healthy until a deploy, restart, scale event, or broker bounce forces mass re-lookup.&lt;/p></description></item><item><title>Apache Pulsar bundle unload thrashing: the load balancer that never converges</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-bundle-unload-thrashing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-bundle-unload-thrashing/</guid><description>&lt;h1 id="apache-pulsar-bundle-unload-thrashing-the-load-balancer-that-never-converges">Apache Pulsar bundle unload thrashing: the load balancer that never converges&lt;/h1>
&lt;p>The primary signal is &lt;code>pulsar_lb_unload_bundle_total&lt;/code>, a counter exposed on the broker Prometheus metrics endpoint. In steady state it barely moves: bundle ownership changes are rare, ideally less than once per hour. When this counter climbs past one unload per minute and there is no rolling upgrade, no broker restart, and no planned maintenance, the load balancer is thrashing.&lt;/p>
&lt;p>Each unload transfers a namespace bundle (a hash-range slice of topics) from one broker to another. Each transfer drops the TCP connections for every producer and consumer on those topics. Clients reconnect automatically, but reconnection generates metadata operations against ZooKeeper, new ledger creation in BookKeeper, and a latency blip measured in tens of milliseconds per affected topic. When unloads happen every few seconds, those blips stack into sustained degradation: elevated publish latency, client reconnection storms, and metadata store pressure that feeds back into more instability.&lt;/p></description></item><item><title>Apache Pulsar consumers connected but not acknowledging: the silent consumer stall</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-consumer-stalled-no-acks/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-consumer-stalled-no-acks/</guid><description>&lt;h1 id="apache-pulsar-consumers-connected-but-not-acknowledging-the-silent-consumer-stall">Apache Pulsar consumers connected but not acknowledging: the silent consumer stall&lt;/h1>
&lt;p>Your Pulsar cluster looks healthy. Brokers are up, bookies are writable, producers are publishing at normal rates. But for one topic or subscription, &lt;code>pulsar_rate_out&lt;/code> has gone flat. Messages are flowing in but nothing is coming out. The backlog is climbing. Consumers show as connected, but the acknowledgment rate is zero.&lt;/p>
&lt;p>No error surfaces in broker logs. The broker serves connections and bookie writes are fast. The problem lives in the consumer application or in the interaction between consumer configuration and broker dispatch policy.&lt;/p></description></item><item><title>Apache Pulsar dead letter queue filling: maxRedeliveryCount and the DLQ topic</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-dead-letter-queue-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-dead-letter-queue-growing/</guid><description>&lt;h1 id="apache-pulsar-dead-letter-queue-filling-maxredeliverycount-and-the-dlq-topic">Apache Pulsar dead letter queue filling: maxRedeliveryCount and the DLQ topic&lt;/h1>
&lt;p>The dead letter topic is where Pulsar parks messages that exhausted their retry budget. When you set &lt;code>DeadLetterPolicy.builder().maxRedeliverCount(N).build()&lt;/code> on a Shared or Key_Shared subscription, the broker stops redelivering a message after N failed attempts, routes it to &lt;code>{topic}-{subscription}-DLQ&lt;/code>, and auto-acks it in the origin subscription so backlog clears.&lt;/p>
&lt;p>A growing DLQ arrival rate is a ledger of application processing failures. A few poison messages parked for inspection is the DLQ working as designed. A steady stream of thousands of messages per second is a systemic failure. Unmonitored, the DLQ topic itself becomes the incident: messages accumulate, storage grows, and nobody notices until disk fills or a downstream team asks why their data is missing.&lt;/p></description></item><item><title>Apache Pulsar ensemble, write quorum, and ack quorum: what E, Qw, and Qa actually guarantee</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-write-ack-quorum-explained/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-write-ack-quorum-explained/</guid><description>&lt;h1 id="apache-pulsar-ensemble-write-quorum-and-ack-quorum-what-e-qw-and-qa-actually-guarantee">Apache Pulsar ensemble, write quorum, and ack quorum: what E, Qw, and Qa actually guarantee&lt;/h1>
&lt;p>Every persistent message in Apache Pulsar passes through BookKeeper&amp;rsquo;s quorum system, governed by three parameters: ensemble size (E), write quorum (Qw), and ack quorum (Qa). Configured per namespace and inherited by topics, they determine how many bookies receive each entry, how many must confirm a durable write before the producer is acknowledged, and how the cluster behaves when bookies fail, slow down, or are taken offline.&lt;/p></description></item><item><title>Apache Pulsar entry log GC falling behind: reclaimed space that never comes back</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-entry-log-gc-lagging/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-entry-log-gc-lagging/</guid><description>&lt;h1 id="apache-pulsar-entry-log-gc-falling-behind-reclaimed-space-that-never-comes-back">Apache Pulsar entry log GC falling behind: reclaimed space that never comes back&lt;/h1>
&lt;p>A bookie&amp;rsquo;s ledger disk is filling. Retention policies have deleted ledgers, TTL has expired messages, and cursors have advanced past the data. By every logical measure, the space should be free. But disk usage keeps climbing, and the bookie is heading toward read-only.&lt;/p>
&lt;p>BookKeeper does not store each ledger in a separate file. Entry logs are large append-only files that interleave entries from many ledgers, written sequentially as they arrive. When a ledger is deleted, its entries remain physically embedded in the entry log alongside data from other ledgers that are still active. The only way to reclaim that space is compaction: the garbage collector reads the live entries from an old entry log, writes them into a new file, then deletes the old one. If compaction stalls, throttles to a crawl, or never triggers, dead data accumulates indefinitely.&lt;/p></description></item><item><title>Apache Pulsar geo-replication backlog: replication lag and your real RPO window</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-geo-replication-backlog/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-geo-replication-backlog/</guid><description>&lt;h1 id="apache-pulsar-geo-replication-backlog-replication-lag-and-your-real-rpo-window">Apache Pulsar geo-replication backlog: replication lag and your real RPO window&lt;/h1>
&lt;p>Geo-replication in Pulsar is asynchronous by design. Messages are persisted locally and acknowledged to producers before they reach the remote cluster. Every message sitting in the replication backlog is a message that would be lost if you failed over right now. The replication backlog is your recovery point objective (RPO) window, measured in real time.&lt;/p>
&lt;p>The common operational mistake is checking whether replication is &amp;ldquo;working&amp;rdquo; (connected, producing to the remote cluster) without checking how far behind it is. A replicator can be fully connected and actively shipping messages while being hours behind. The dashboard shows green. The failover plan assumes near-zero data loss. The reality is data loss measured in hours.&lt;/p></description></item><item><title>Apache Pulsar geo-replication disconnected: a remote cluster that stopped receiving</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-geo-replication-disconnected/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-geo-replication-disconnected/</guid><description>&lt;h1 id="apache-pulsar-geo-replication-disconnected-a-remote-cluster-that-stopped-receiving">Apache Pulsar geo-replication disconnected: a remote cluster that stopped receiving&lt;/h1>
&lt;p>The symptom is unmistakable: &lt;code>pulsar_replication_disconnected_count&lt;/code> is non-zero for all replicators to a specific remote cluster, &lt;code>pulsar_replication_connected_count&lt;/code> has dropped to zero, and &lt;code>pulsar_replication_backlog&lt;/code> is growing at the rate of local publish throughput. The remote cluster has stopped receiving messages. Local producers and consumers continue working normally because Pulsar geo-replication is asynchronous. Messages persist locally first, then replicate. The damage is silent and cumulative.&lt;/p></description></item><item><title>Apache Pulsar journal force write queue growing: the earliest write-saturation signal</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-journal-force-write-queue-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-journal-force-write-queue-growing/</guid><description>&lt;h1 id="apache-pulsar-journal-force-write-queue-growing-the-earliest-write-saturation-signal">Apache Pulsar journal force write queue growing: the earliest write-saturation signal&lt;/h1>
&lt;p>&lt;code>bookie_journal_JOURNAL_FORCE_WRITE_QUEUE_SIZE&lt;/code> measures the depth of pending fsync batches inside a BookKeeper bookie&amp;rsquo;s journal write path. In a healthy system this gauge sits at or near zero, draining between write groups. When it sustains a non-zero depth, the journal disk cannot commit writes durably fast enough to keep up with incoming traffic.&lt;/p>
&lt;p>This signal rises before &lt;code>bookie_journal_JOURNAL_SYNC&lt;/code> latency spikes and before &lt;code>bookkeeper_server_ADD_ENTRY_IN_PROGRESS&lt;/code> grows. If you are not watching the force write queue, your first indication of write-path saturation will be broker publish latency degradation or producer timeouts, which means you are already two or three steps into the cascade.&lt;/p></description></item><item><title>Apache Pulsar ledger rollover latency spikes: the periodic blip that is usually normal</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-ledger-rollover-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-ledger-rollover-latency/</guid><description>&lt;h1 id="apache-pulsar-ledger-rollover-latency-spikes-the-periodic-blip-that-is-usually-normal">Apache Pulsar ledger rollover latency spikes: the periodic blip that is usually normal&lt;/h1>
&lt;p>You see a periodic sub-second spike in publish latency on your Pulsar brokers. It happens at regular intervals, lasts a fraction of a second, and then latency returns to baseline. No errors, no bookie failures, no GC pauses. The spike is visible in P99 publish latency but barely touches P50. It correlates across topics on the same broker but not across the entire cluster.&lt;/p></description></item><item><title>Apache Pulsar LedgerFencedException: split ownership and fencing loops</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-ledger-fenced-exception/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-ledger-fenced-exception/</guid><description>&lt;h1 id="apache-pulsar-ledgerfencedexception-split-ownership-and-fencing-loops">Apache Pulsar LedgerFencedException: split ownership and fencing loops&lt;/h1>
&lt;p>LedgerFencedException appears when a broker tries to write to a BookKeeper ledger that another broker has already fenced. The error string is explicit: &amp;ldquo;Ledger has been fenced off. Some other client must have opened it to read.&amp;rdquo;&lt;/p>
&lt;p>In most production environments, this is expected noise during planned failover. It becomes a problem when fencing occurs without a corresponding bundle transfer, or when two brokers fence each other&amp;rsquo;s ledgers in a loop.&lt;/p></description></item><item><title>Apache Pulsar managed ledger cache miss rate high: consumer reads falling through to bookies</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-managed-ledger-cache-misses/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-managed-ledger-cache-misses/</guid><description>&lt;h1 id="apache-pulsar-managed-ledger-cache-miss-rate-high-consumer-reads-falling-through-to-bookies">Apache Pulsar managed ledger cache miss rate high: consumer reads falling through to bookies&lt;/h1>
&lt;p>The managed ledger cache is the broker&amp;rsquo;s off-heap buffer of recently written entries. Connected consumers read from it at sub-millisecond latency. On a cache miss, the broker issues a read to the BookKeeper bookie ensemble holding that ledger segment, which adds disk I/O, network bandwidth consumption between broker and bookie, and consumer read latency.&lt;/p>
&lt;p>A sustained cache miss rate above 20% after warmup means the broker is serving a meaningful fraction of consumer reads from storage instead of memory. The extra bookie read traffic loads storage disks that also serve journal writes. If bookie read latency rises enough, publish latency follows. This is the leading edge of the Backlog Cascade: consumer reads and producer writes converging on the same physical disks.&lt;/p></description></item><item><title>Apache Pulsar message redelivery storm: poison messages and consumers making no progress</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-redelivery-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-redelivery-storm/</guid><description>&lt;h1 id="apache-pulsar-message-redelivery-storm-poison-messages-and-consumers-making-no-progress">Apache Pulsar message redelivery storm: poison messages and consumers making no progress&lt;/h1>
&lt;p>Subscription redelivery rate is climbing. Consumers are connected, the dispatch rate (&lt;code>pulsar_rate_out&lt;/code>) looks healthy, but messages keep cycling back without being acknowledged. The backlog may even appear stable or zero, because &lt;code>pulsar_subscription_back_log&lt;/code> measures messages not yet dispatched, not messages dispatched but unacked. The definitive signal is &lt;code>pulsar_subscription_msg_rate_redeliver&lt;/code> trending upward relative to dispatch. When redelivery exceeds 10% of your dispatch rate, something is wrong. When it approaches 100%, you have zero forward progress: the broker burns memory and network re-sending messages that will never succeed.&lt;/p></description></item><item><title>Apache Pulsar messages expiring before consumers read them: TTL and silent loss</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-message-ttl-expiry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-message-ttl-expiry/</guid><description>&lt;h1 id="apache-pulsar-messages-expiring-before-consumers-read-them-ttl-and-silent-loss">Apache Pulsar messages expiring before consumers read them: TTL and silent loss&lt;/h1>
&lt;p>When &lt;code>pulsar_subscription_msg_rate_expired&lt;/code> is non-zero on a topic where every message matters, messages are being silently deleted before consumers can read them. No errors appear in consumer logs. No producer failures occur. The broker&amp;rsquo;s TTL mechanism acknowledges messages on behalf of the subscription without ever delivering them to the consumer application.&lt;/p>
&lt;p>The backlog can hide the problem. TTL expiry trims unacked messages from subscription cursors, so backlog can appear stable or declining while real consumer lag grows underneath. Operators monitoring backlog size alone see a healthy-looking number. The data loss is invisible unless you specifically monitor the expiry rate.&lt;/p></description></item><item><title>Apache Pulsar metadata store latency: the leading indicator before every cluster-wide failure</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-metadata-store-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-metadata-store-latency-high/</guid><description>&lt;h1 id="apache-pulsar-metadata-store-latency-the-leading-indicator-before-every-cluster-wide-failure">Apache Pulsar metadata store latency: the leading indicator before every cluster-wide failure&lt;/h1>
&lt;p>Metadata store latency is the earliest signal that a Pulsar cluster is heading toward a widespread outage. In most 3.x deployments the metadata store is ZooKeeper. &lt;!-- TODO: verify Oxia introduction version --> Pulsar 3.3.0 introduced experimental Oxia support as an eventual replacement. Every broker, bookie, and load balancer operation that touches cluster topology, bundle ownership, schema lookups, ledger metadata, or cursor persistence routes through this store. When round-trip latency for those operations rises, the symptoms appear downstream in brokers and bookies, but the root cause is upstream.&lt;/p></description></item><item><title>Apache Pulsar Monitoring</title><link>https://www.netdata.cloud/monitoring-101/pulsar-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/pulsar-monitoring/</guid><description>&lt;h2 id="apache-pulsar-monitoring">Apache Pulsar Monitoring&lt;/h2>
&lt;h3 id="what-is-apache-pulsar">What Is Apache Pulsar?&lt;/h3>
&lt;p>&lt;a href="https://pulsar.apache.org/">Apache Pulsar&lt;/a> is an open-source distributed messaging and streaming platform. It provides a unified messaging model and is designed to handle high-throughput, low-latency workloads. With its multi-tenancy, geo-replication, and seamless scalability features, Pulsar is a comprehensive solution tailored for both messaging and streaming use cases.&lt;/p>
&lt;h3 id="monitoring-apache-pulsar-with-netdata">Monitoring Apache Pulsar With Netdata&lt;/h3>
&lt;p>Monitoring Apache Pulsar is crucial for ensuring the system’s health and performance. Netdata offers a powerful and intuitive way to monitor Pulsar, leveraging its robust capabilities to deliver real-time insights. The &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/pulsar/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata Pulsar monitoring tool&lt;/a> taps into Pulsar&amp;rsquo;s &lt;a href="https://pulsar.apache.org/docs/en/deploy-monitoring/#broker-stats">Prometheus endpoint&lt;/a> to collect a wide range of metrics that help you keep your messaging system running smoothly.&lt;/p></description></item><item><title>Apache Pulsar monitoring checklist: the signals every production cluster needs</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-monitoring-checklist/</guid><description>&lt;h1 id="apache-pulsar-monitoring-checklist-the-signals-every-production-cluster-needs">Apache Pulsar monitoring checklist: the signals every production cluster needs&lt;/h1>
&lt;p>Pulsar fails differently from most systems you operate. A cluster can report every process &amp;ldquo;up&amp;rdquo; while the write path stalls on a bookie journal disk, while a broker GC-pauses its ZooKeeper session away, or while a subscription silently freezes because its unacked message count hit a limit. Monitoring that only checks &amp;ldquo;is the broker running&amp;rdquo; misses almost every real Pulsar incident.&lt;/p></description></item><item><title>Apache Pulsar monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-monitoring-maturity-model/</guid><description>&lt;h1 id="apache-pulsar-monitoring-maturity-model-from-survival-to-expert">Apache Pulsar monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most Pulsar outages do not arrive without warning. The journal force write queue grows before publish latency spikes. ZooKeeper latency drifts upward for days before brokers lose session ownership. Direct memory climbs for weeks before the broker dies with an OutOfDirectMemoryError while heap dashboards look fine. The signals were there. The team just was not collecting them yet.&lt;/p>
&lt;p>This is a four-level maturity model for Pulsar monitoring, from the bare minimum that tells you the cluster is alive to the deep signals operators add after their second or third incident. Use it to audit your current coverage and decide what to instrument next. The levels are cumulative: each one assumes everything below it is already in place.&lt;/p></description></item><item><title>Apache Pulsar NotEnoughBookiesException: new ledgers cannot be created and writes fail</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-notenoughbookies-exception/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-notenoughbookies-exception/</guid><description>&lt;h1 id="apache-pulsar-notenoughbookiesexception-new-ledgers-cannot-be-created-and-writes-fail">Apache Pulsar NotEnoughBookiesException: new ledgers cannot be created and writes fail&lt;/h1>
&lt;p>&lt;code>NotEnoughBookiesException&lt;/code> (BookKeeper error code -6, surfaced as &lt;code>BKNotEnoughBookiesException&lt;/code>) fires when the ensemble placement policy cannot satisfy the ensemble size (E) and write quorum (Qw) requirements for a new ledger. The broker requests a new ledger on size or time rollover, during topic recovery, and during compaction. If the policy cannot form a valid ensemble, ledger creation fails, the managed ledger has nowhere to write, and messages cannot be persisted durably.&lt;/p></description></item><item><title>Apache Pulsar OutOfDirectMemoryError: the off-heap crash JVM heap dashboards never show</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-outofdirectmemory-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-outofdirectmemory-error/</guid><description>&lt;h1 id="apache-pulsar-outofdirectmemoryerror-the-off-heap-crash-jvm-heap-dashboards-never-show">Apache Pulsar OutOfDirectMemoryError: the off-heap crash JVM heap dashboards never show&lt;/h1>
&lt;p>The broker is down. Your JVM dashboard shows heap at 55%, GC pauses normal, no heap OOM anywhere. Then someone opens the broker log and finds the real cause:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-text" data-lang="text">&lt;span style="display:flex;">&lt;span>io.netty.util.internal.OutOfDirectMemoryError: failed to allocate 16777216 byte(s) of direct memory (used: ..., max: ...)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>or the plainer JDK variant, &lt;code>java.lang.OutOfMemoryError: Direct buffer memory&lt;/code>. In recent Pulsar versions the process may have exited deliberately: PulsarByteBufAllocator catches the allocation failure and, with the default &lt;code>-Dpulsar.allocator.exit_on_oom=true&lt;/code>, kills the JVM rather than limp along half-broken. Either way, the broker is dead in the part your heap monitoring was not watching.&lt;/p></description></item><item><title>Apache Pulsar publish latency high: reading pulsar_broker_publish_latency P99</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-publish-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-publish-latency-high/</guid><description>&lt;h1 id="apache-pulsar-publish-latency-high-reading-pulsar_broker_publish_latency-p99">Apache Pulsar publish latency high: reading pulsar_broker_publish_latency P99&lt;/h1>
&lt;p>Your alert fired on sustained P99 elevation of &lt;code>pulsar_broker_publish_latency&lt;/code> above 2x the rolling baseline. Producers are seeing slow acknowledgements.&lt;/p>
&lt;p>The metric &lt;code>pulsar_broker_publish_latency&lt;/code> is a Summary metric exposed on the broker Prometheus endpoint. It measures the time from when the broker receives a message from a producer through the BookKeeper write path (write quorum Qw, ack quorum Qa) and back to the client callback. It is broker-side only: it excludes client-to-broker network time and producer-side batching delay. The Summary type provides quantiles at 0.5, 0.95, 0.99, 0.999, 0.9999, and 1.0. Alert on P99. P50 can look healthy while P99 is spiking, and it is the tail that triggers producer timeouts.&lt;/p></description></item><item><title>Apache Pulsar subscription backlog growing: consumers falling behind producers</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-subscription-backlog-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-subscription-backlog-growing/</guid><description>&lt;h1 id="apache-pulsar-subscription-backlog-growing-consumers-falling-behind-producers">Apache Pulsar subscription backlog growing: consumers falling behind producers&lt;/h1>
&lt;p>A growing subscription backlog means producers are publishing faster than consumers can acknowledge. The absolute backlog size matters less than its trajectory. A high but stable backlog is normal for lagged or replay consumers. A monotonically increasing backlog is a problem regardless of size: it consumes bookie disk until the bookie goes read-only, the backlog quota trips and throttles producers, or retention and TTL silently delete messages the consumer never saw.&lt;/p></description></item><item><title>Apache Pulsar throttled connections: the broker shedding load under pressure</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-throttled-connections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-throttled-connections/</guid><description>&lt;h1 id="apache-pulsar-throttled-connections-the-broker-shedding-load-under-pressure">Apache Pulsar throttled connections: the broker shedding load under pressure&lt;/h1>
&lt;p>Your dashboard shows &lt;code>pulsar_broker_throttled_connections&lt;/code> climbing from zero, producers are complaining about elevated send latency, and some clients are timing out on new operations. The broker process is up, the health endpoint returns 200, and heap looks fine. The broker is not broken. It is defending itself.&lt;/p>
&lt;p>Connection throttling is a protective mechanism. When the broker&amp;rsquo;s internal send queues on a connection back up beyond a configured ceiling, it stops reading new requests from that TCP connection (it disables auto-read on the Netty channel) until the backlog drains. The throttled connection count tells you how many connections are currently in that state. In a healthy cluster this gauge is zero. Any sustained non-zero value means the broker is at or near its capacity for the work being pushed through it.&lt;/p></description></item><item><title>Apache Pulsar TLS certificate expiry: the silent, total outage no metric warns you about</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-tls-certificate-expiry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-tls-certificate-expiry/</guid><description>&lt;h1 id="apache-pulsar-tls-certificate-expiry-the-silent-total-outage-no-metric-warns-you-about">Apache Pulsar TLS certificate expiry: the silent, total outage no metric warns you about&lt;/h1>
&lt;p>When every component in a Pulsar cluster uses TLS (brokers, bookies, ZooKeeper, proxies, clients), a single expired certificate causes immediate handshake failures. Existing connections may persist briefly, but every new connection attempt fails the TLS handshake. Producers cannot publish. Consumers cannot subscribe. Brokers cannot reach bookies. Brokers cannot reach ZooKeeper.&lt;/p>
&lt;p>Pulsar does not expose a metric for certificate expiry. There is no &lt;code>pulsar_cert_days_until_expiry&lt;/code> gauge. The first signals you see are indirect: authentication failures spike, connections drop, and SSL handshake exceptions fill the logs. By that point, the outage is already happening. The only reliable defense is external certificate expiry monitoring that alerts you days or weeks before the cert becomes invalid.&lt;/p></description></item><item><title>Apache Pulsar topic ownership oscillation: brokers fighting over the same bundle</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-topic-ownership-oscillation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-topic-ownership-oscillation/</guid><description>&lt;h1 id="apache-pulsar-topic-ownership-oscillation-brokers-fighting-over-the-same-bundle">Apache Pulsar topic ownership oscillation: brokers fighting over the same bundle&lt;/h1>
&lt;p>Broker A takes ownership of a namespace bundle, opens managed ledgers, and begins serving clients. The load balancer sees A as overloaded and unloads the bundle. Broker B picks it up, fences the ledgers, creates new ones, and starts serving. The balancer now sees B as overloaded and moves it back. The cycle repeats indefinitely.&lt;/p>
&lt;p>Each iteration forces ledger fencing, new ledger creation, ZK metadata writes, and client disconnection and reconnection storms. Producers see intermittent publish latency spikes. New clients fail lookups during the brief ownership gap between fencing and re-acquisition. The cluster looks healthy in aggregate because no single broker is down, but the affected topics are in a constant state of transition.&lt;/p></description></item><item><title>Apache Pulsar unacked messages at the limit: the silent dispatch freeze</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-unacked-messages-dispatch-freeze/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-unacked-messages-dispatch-freeze/</guid><description>&lt;h1 id="apache-pulsar-unacked-messages-at-the-limit-the-silent-dispatch-freeze">Apache Pulsar unacked messages at the limit: the silent dispatch freeze&lt;/h1>
&lt;p>Consumers are connected. Producers are publishing. Backlog looks flat or zero. But no messages are being processed.&lt;/p>
&lt;p>Pulsar brokers enforce a per-subscription limit on unacknowledged messages (&lt;code>maxUnackedMessagesPerSubscription&lt;/code>, default 200,000) and a per-consumer limit (&lt;code>maxUnackedMessagesPerConsumer&lt;/code>, default 50,000). When unacked messages hit either ceiling, the broker stops dispatching new messages to that subscription or consumer. No error is returned to the client. No exception is thrown. Dispatch pauses silently.&lt;/p></description></item><item><title>Apache Pulsar under-replicated ledgers: data at risk after a bookie failure</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-under-replicated-ledgers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-under-replicated-ledgers/</guid><description>&lt;h1 id="apache-pulsar-under-replicated-ledgers-data-at-risk-after-a-bookie-failure">Apache Pulsar under-replicated ledgers: data at risk after a bookie failure&lt;/h1>
&lt;p>You lost a bookie. The broker layer kept serving traffic because your write quorum absorbed the failure. But now &lt;code>auditor_NUM_UNDER_REPLICATED_LEDGERS&lt;/code> is climbing, and it is not coming back down. Every minute it stays elevated is a minute your data exists on fewer copies than your replication factor demands. If another bookie fails before AutoRecovery finishes rereplicating, those ledgers are gone permanently.&lt;/p></description></item><item><title>Apache Pulsar write stall: bookie journal fsync latency and the blocked write path</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-journal-write-stall/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-journal-write-stall/</guid><description>&lt;h1 id="apache-pulsar-write-stall-bookie-journal-fsync-latency-and-the-blocked-write-path">Apache Pulsar write stall: bookie journal fsync latency and the blocked write path&lt;/h1>
&lt;p>The write stall is the most common performance failure in Pulsar. It starts at the bookie journal disk and cascades upward through the write path until producers are blocked or timing out.&lt;/p>
&lt;p>Every persistent message write in Pulsar must survive a journal fsync before the bookie acknowledges it. The broker waits for ack quorum (Qa) acknowledgments before acknowledging the producer. When the journal disk cannot sync fast enough, pending fsync operations accumulate, add-entry operations queue up, brokers hold connections open waiting for acks, client buffers fill, and throughput collapses.&lt;/p></description></item><item><title>Apache Pulsar ZooKeeper session cascade: reconnect storms and the thundering herd</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-zookeeper-session-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-zookeeper-session-cascade/</guid><description>&lt;h1 id="apache-pulsar-zookeeper-session-cascade-reconnect-storms-and-the-thundering-herd">Apache Pulsar ZooKeeper session cascade: reconnect storms and the thundering herd&lt;/h1>
&lt;p>Multiple brokers lose their ZooKeeper sessions within seconds of each other. Bundle ownership churns across the cluster as surviving brokers acquire orphaned bundles. Every connected producer and consumer simultaneously discovers its topic has moved or its connection has dropped, triggering a mass reconnection event. Each reconnection generates topic lookups, ephemeral node registrations, and watch re-establishments that hit the already overloaded ZK ensemble. Latency rises further, more sessions expire, and the cascade tightens.&lt;/p></description></item><item><title>Apache Pulsar ZooKeeper session expired: brokers fenced and topics reassigned</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-zookeeper-session-expired/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-zookeeper-session-expired/</guid><description>&lt;h1 id="apache-pulsar-zookeeper-session-expired-brokers-fenced-and-topics-reassigned">Apache Pulsar ZooKeeper session expired: brokers fenced and topics reassigned&lt;/h1>
&lt;p>A broker&amp;rsquo;s ZooKeeper session expires when it fails to send heartbeats within the session timeout window. The expiry triggers immediate fencing of the broker, loss of all namespace bundle ownership, and reassignment of those bundles to surviving brokers. Clients connected to the fenced broker see errors like &amp;ldquo;Topic is temporarily unavailable&amp;rdquo; or &amp;ldquo;Attempting to add producer to a fenced topic&amp;rdquo; until the topics are picked up elsewhere.&lt;/p></description></item><item><title>Apache Pulsar ZooKeeper watch explosion: thousands of watches killing metadata latency</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-zookeeper-watch-explosion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-zookeeper-watch-explosion/</guid><description>&lt;h1 id="apache-pulsar-zookeeper-watch-explosion-thousands-of-watches-killing-metadata-latency">Apache Pulsar ZooKeeper watch explosion: thousands of watches killing metadata latency&lt;/h1>
&lt;p>Topic lookups slow down. Bundle ownership transfers stall. Broker sessions flicker between connected and disconnected. ZooKeeper reports thousands, sometimes tens of thousands, of registered watches. This is a watch explosion, and the root cause is rarely in ZooKeeper itself. It is in client connection churn.&lt;/p>
&lt;p>Watches accumulate when consumers disconnect and reconnect in waves. Each reconnection registers metadata watches for topic ownership, subscription state, and policy changes. During a mass reconnection event (broker restart, network blip, load balancer cycle), thousands of ephemeral znodes are created and deleted in rapid succession. Each creation and deletion triggers watch registration and notification. ZK processes requests sequentially, so as notification load grows, request latency climbs. Once latency exceeds session timeout, brokers lose sessions, triggering bundle unloads and another round of client reconnections. The feedback loop tightens until the cluster thrashes.&lt;/p></description></item><item><title>Apache request latency: p50/p95/p99 from the access log %D field</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-request-latency-percentiles/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-request-latency-percentiles/</guid><description>&lt;h1 id="apache-request-latency-p50p95p99-from-the-access-log-d-field">Apache request latency: p50/p95/p99 from the access log %D field&lt;/h1>
&lt;p>&amp;ldquo;Apache is slow&amp;rdquo; is one of the least actionable statements in operations. The follow-up questions are always: slow for whom, on which URLs, and how slow at the tail? Averages cannot answer any of those. A server where every request takes 200ms and a server where half the requests take 5ms and half take 400ms have the same mean latency and completely different operational problems.&lt;/p></description></item><item><title>Apache request rate dropped to near zero: silent load shedding and its causes</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-request-rate-drop/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-request-rate-drop/</guid><description>&lt;h1 id="apache-request-rate-dropped-to-near-zero-silent-load-shedding-and-its-causes">Apache request rate dropped to near zero: silent load shedding and its causes&lt;/h1>
&lt;p>Your dashboard shows requests per second falling off a cliff, but nobody changed anything, the load balancer still reports healthy demand, and there is no error spike to point at. The most natural reading, &amp;ldquo;traffic went away,&amp;rdquo; is usually wrong.&lt;/p>
&lt;p>The trap is in how the number is produced. The request rate most operators watch comes from mod_status &lt;code>Total Accesses&lt;/code>, which counts completed requests. It is not a measure of arriving demand. When Apache cannot complete requests, because every worker is stuck or a backend has stopped answering, the completed-request rate collapses even while clients are still hammering the front door. The server has not lost its traffic. It is silently shedding load, and the metric you trust is the last place that shows up.&lt;/p></description></item><item><title>Apache reverse proxy worker busy: proxy connection pool exhaustion (AH01136)</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-proxy-pool-exhausted/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-proxy-pool-exhausted/</guid><description>&lt;h1 id="apache-reverse-proxy-worker-busy-proxy-connection-pool-exhaustion-ah01136">Apache reverse proxy worker busy: proxy connection pool exhaustion (AH01136)&lt;/h1>
&lt;p>You are seeing intermittent 503s on proxied endpoints under what looks like moderate load. The error log shows &lt;code>AH01136: Reverse proxy worker busy&lt;/code>. BusyWorkers is nowhere near MaxRequestWorkers, CPU and memory are fine, and requests keep failing. Raising MaxRequestWorkers changed nothing.&lt;/p>
&lt;p>That is the signature of proxy connection pool exhaustion. The bottleneck is not Apache&amp;rsquo;s worker pool. It is the smaller, less visible pool of backend connections that mod_proxy maintains per child process, and its default size is too small for production load.&lt;/p></description></item><item><title>Apache scoreboard states explained: what _ S R W K D C L G tell you</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-scoreboard-states/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-scoreboard-states/</guid><description>&lt;h1 id="apache-scoreboard-states-explained-what-_-s-r-w-k-d-c-l-g-tell-you">Apache scoreboard states explained: what _ S R W K D C L G tell you&lt;/h1>
&lt;p>When someone asks &amp;ldquo;why is Apache slow?&amp;rdquo;, the first place to look is not CPU, not memory, not the access log. It is the scoreboard. The scoreboard is a shared-memory segment where every worker slot records what it is doing right now, one character per slot. It is what &lt;code>mod_status&lt;/code> reads, and it is the most diagnostic structure Apache exposes.&lt;/p></description></item><item><title>Apache slow backend cascade: how one slow upstream starves the whole worker pool</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-slow-backend-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-slow-backend-cascade/</guid><description>&lt;h1 id="apache-slow-backend-cascade-how-one-slow-upstream-starves-the-whole-worker-pool">Apache slow backend cascade: how one slow upstream starves the whole worker pool&lt;/h1>
&lt;p>Users report the site down. The load balancer has pulled the node from rotation. You SSH in expecting to find Apache melting, and instead find a perfectly calm process: normal CPU, normal memory, no crash dumps in the error log. The port is listening, but new connections hang or get refused.&lt;/p>
&lt;p>This is the slow backend cascade: the classic &amp;ldquo;Apache outage&amp;rdquo; in reverse-proxy deployments. Apache is not broken. Every worker is blocked waiting on a backend that has gone slow, and the frontend has run out of execution slots as a consequence.&lt;/p></description></item><item><title>Apache Slowloris: R-state workers, slow-read attacks, and mod_reqtimeout</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-slowloris-attack/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-slowloris-attack/</guid><description>&lt;h1 id="apache-slowloris-r-state-workers-slow-read-attacks-and-mod_reqtimeout">Apache Slowloris: R-state workers, slow-read attacks, and mod_reqtimeout&lt;/h1>
&lt;p>Your Apache server stopped serving requests, but nothing looks broken. CPU is normal. Memory is normal. The backend is healthy when you test it directly. Request rate is flat or dropping. Yet clients time out, and the load balancer is pulling the node from rotation.&lt;/p>
&lt;p>This is the Slowloris signature: a slow-read denial of service where an attacker opens many connections and drips request data byte by byte, holding each worker in the Reading (R) state indefinitely. With enough slow connections, the entire worker pool is consumed reading requests that never complete. The server is not overloaded. It is held hostage.&lt;/p></description></item><item><title>Apache SSL certificate expired: the total, preventable HTTPS outage</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-certificate-expired/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-certificate-expired/</guid><description>&lt;h1 id="apache-ssl-certificate-expired-the-total-preventable-https-outage">Apache SSL certificate expired: the total, preventable HTTPS outage&lt;/h1>
&lt;p>Every HTTPS client is failing at once. Browsers show NET::ERR_CERT_DATE_INVALID, API clients throw certificate validation errors, monitoring probes time out, and your access logs have gone quiet because no TLS handshake ever completes. Apache itself is running fine: the process is up, workers are idle, and port 443 accepts TCP connections. The failure is at the TLS layer, before any HTTP request exists.&lt;/p></description></item><item><title>Apache SSL session cache: shmcb sizing, sharing, and silent eviction</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-ssl-session-cache/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-ssl-session-cache/</guid><description>&lt;h1 id="apache-ssl-session-cache-shmcb-sizing-sharing-and-silent-eviction">Apache SSL session cache: shmcb sizing, sharing, and silent eviction&lt;/h1>
&lt;p>A full TLS handshake is the most CPU-intensive thing Apache does per connection. Session resumption exists to avoid paying that cost on every reconnect: a client that recently completed a handshake can resume instead of redoing the asymmetric cryptography. When resumption silently stops working, every connection pays full price, and the only symptom is rising CPU with no error in the log.&lt;/p></description></item><item><title>Apache swap thrashing: the latency cliff before the OOM kill</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-swap-thrashing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-swap-thrashing/</guid><description>&lt;h1 id="apache-swap-thrashing-the-latency-cliff-before-the-oom-kill">Apache swap thrashing: the latency cliff before the OOM kill&lt;/h1>
&lt;p>The symptom arrives before the failure. P95 latency doubles, then triples, across every endpoint at once, including static files that should be served from page cache. No 5xx spike yet. No OOM kills in &lt;code>dmesg&lt;/code> yet. The Apache parent process is fine, the scoreboard shows workers in normal states, and every request is suddenly slow. This is the swap window: the period between &amp;ldquo;memory is tight&amp;rdquo; and &amp;ldquo;the kernel starts killing children,&amp;rdquo; where Apache is technically up but effectively degraded.&lt;/p></description></item><item><title>Apache TLS handshake CPU: broken session resumption and the HTTPS CPU wall</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-tls-handshake-cpu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-tls-handshake-cpu/</guid><description>&lt;h1 id="apache-tls-handshake-cpu-broken-session-resumption-and-the-https-cpu-wall">Apache TLS handshake CPU: broken session resumption and the HTTPS CPU wall&lt;/h1>
&lt;p>The symptom looks like a capacity problem: Apache processes pinned near 100% CPU, request latency climbing, and no obvious cause in the access log. Traffic is up, but not absurdly. The workers are not stuck on a backend. The scoreboard is busy but not exhausted. And yet the box is melting.&lt;/p>
&lt;p>On HTTPS-heavy servers, the usual suspect is the TLS handshake. A full TLS handshake with an RSA-2048 certificate costs on the order of 15 ms of CPU time per connection. Session resumption cuts that cost by roughly 10x. When resumption works, repeat clients skip the expensive public-key operation entirely. When resumption is broken, every connection pays the full price, and at a few hundred new connections per second the math stops working.&lt;/p></description></item><item><title>Apache Too many open files: file descriptor exhaustion (EMFILE)</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-too-many-open-files/</guid><description>&lt;h1 id="apache-too-many-open-files-file-descriptor-exhaustion-emfile">Apache Too many open files: file descriptor exhaustion (EMFILE)&lt;/h1>
&lt;p>Your error log starts showing lines like:&lt;/p>
&lt;pre tabindex="0">&lt;code>(24)Too many open files: AH00035: access to /some/path failed
&lt;/code>&lt;/pre>&lt;p>or proxy requests begin failing with &lt;code>AH01114&lt;/code> connection errors, and users see intermittent 5xx responses. Sometimes only some requests fail. Sometimes the site looks fine for hours and then breaks at peak. The process is running, the port is open, and nothing in CPU or memory looks wrong.&lt;/p></description></item><item><title>Apache unexpected restarts: SIGTERM vs SIGUSR1, crashes, and OOM respawns</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-unexpected-restarts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-unexpected-restarts/</guid><description>&lt;h1 id="apache-unexpected-restarts-sigterm-vs-sigusr1-crashes-and-oom-respawns">Apache unexpected restarts: SIGTERM vs SIGUSR1, crashes, and OOM respawns&lt;/h1>
&lt;p>&lt;code>mod_status&lt;/code> shows &lt;code>ServerUptimeSeconds&lt;/code> at 340 again, and it was 340 an hour ago too. Something is restarting Apache, and unless you deployed at that exact moment, it is not you. The error log says &lt;code>resuming normal operations&lt;/code> every few minutes, or &lt;code>caught SIGTERM&lt;/code>, or nothing at all between restarts. Each signature points at a different failure mode.&lt;/p>
&lt;p>Apache restarts fall into three buckets: intentional graceful reloads (SIGUSR1), hard restarts (SIGTERM or SIGHUP, often from logrotate or config management), and crash respawns (child segfaults, OOM kills, or the parent dying and being restarted by systemd or a wrapper script). The first is usually harmless. The second drops every in-flight connection. The third means something is actively wrong and will keep happening until you find it.&lt;/p></description></item><item><title>Apache with mod_php: why per-child memory explodes and when to move to PHP-FPM</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-mod-php-memory/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-mod-php-memory/</guid><description>&lt;h1 id="apache-with-mod_php-why-per-child-memory-explodes-and-when-to-move-to-php-fpm">Apache with mod_php: why per-child memory explodes and when to move to PHP-FPM&lt;/h1>
&lt;p>Each &lt;code>apache2&lt;/code> or &lt;code>httpd&lt;/code> child sits at 80, 120, sometimes 200MB of RSS, the sum of all children creeps toward total RAM, and the OOM killer starts picking off workers at peak traffic. &lt;code>MaxRequestWorkers&lt;/code> is already set conservatively, but it does not matter: per-child memory is so large that any worker count high enough to serve your traffic exceeds what the machine can hold.&lt;/p></description></item><item><title>APC</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/apc/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/apc/</guid><description/></item><item><title>APC Netbotz</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/apc-netbotz/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/apc-netbotz/</guid><description/></item><item><title>APC PDU</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/apc-pdu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/apc-pdu/</guid><description/></item><item><title>APC UPS</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/apc-ups/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/apc-ups/</guid><description/></item><item><title>APC UPS</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/apc-ups/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/apc-ups/</guid><description/></item><item><title>APC UPS Monitoring</title><link>https://www.netdata.cloud/monitoring-101/apcupsd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/apcupsd-monitoring/</guid><description>&lt;h2 id="apc-ups-monitoring">APC UPS Monitoring&lt;/h2>
&lt;h3 id="what-is-apc-ups">What Is APC UPS?&lt;/h3>
&lt;p>APC UPS, or Uninterruptible Power Supply systems from &lt;a href="https://www.apc.com">APC&lt;/a>, are crucial components in ensuring the continuity and protection of power to critical infrastructure. These devices ensure that your servers, data centers, and critical systems stay powered even during a power outage, safeguarding your operations against data loss and downtime.&lt;/p>
&lt;h3 id="monitoring-apc-ups-with-netdata">Monitoring APC UPS With Netdata&lt;/h3>
&lt;p>Monitoring platforms like Netdata allow you to keep an eye on the health, performance, and capacity of your APC UPS units in real-time. Utilizing the &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/apcupsd/?utm_source=website&amp;amp;utm_content=monitoring101">Apcupsd&lt;/a> daemon, Netdata provides comprehensive visibility into UPS performance and alerts you to any issues so they can be addressed promptly.&lt;/p></description></item><item><title>Aperto Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aperto-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aperto-networks-snmp-traps/</guid><description/></item><item><title>APIcast</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/apicast/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/apicast/</guid><description/></item><item><title>APIcast Monitoring</title><link>https://www.netdata.cloud/monitoring-101/apicast-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/apicast-monitoring/</guid><description>&lt;h2 id="apicast-monitoring">APIcast Monitoring&lt;/h2>
&lt;h3 id="what-is-apicast">What Is APIcast?&lt;/h3>
&lt;p>APIcast is a powerful API management solution designed for API traffic across various web and cloud applications. It acts as a gateway, controlling how APIs are exposed, secured, and consumed. Built by 3scale, APIcast allows organizations to manage and monetize their APIs effectively, ensuring robust security and operational efficiency. Discover more about &lt;a href="https://github.com/3scale/apicast">APIcast here&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-apicast-with-netdata">Monitoring APIcast With Netdata&lt;/h3>
&lt;p>To effectively monitor APIcast, Netdata leverages an openmetrics (Prometheus) exporter. This allows Netdata to ingest data from any Prometheus exporter and provide automated dashboards and alerts without necessitating a Prometheus server or Grafana setup. This seamless integration helps in gaining real-time insights into APIcast&amp;rsquo;s performance, ensuring optimal operation. By using the Netdata infrastructure monitoring tool, APIcast users can experience detailed, user-friendly dashboards and benefit from real-time alerts, all designed to keep your API services performing at their peak.&lt;/p></description></item><item><title>Apple Computer Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/apple-computer-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/apple-computer-inc-snmp-traps/</guid><description/></item><item><title>Applications</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/applications/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/applications/</guid><description/></item><item><title>Applied Innovation Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/applied-innovation-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/applied-innovation-inc-snmp-traps/</guid><description/></item><item><title>Apply to Become a Netdata Partner</title><link>https://www.netdata.cloud/partnership-contact/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/partnership-contact/</guid><description/></item><item><title>AppOptics</title><link>https://www.netdata.cloud/integrations/exporters/appoptics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/appoptics/</guid><description/></item><item><title>Aptis Communications Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aptis-communications-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aptis-communications-inc-snmp-traps/</guid><description/></item><item><title>Arbor Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/arbor-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/arbor-networks-snmp-traps/</guid><description/></item><item><title>Arch Linux</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/arch-linux/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/arch-linux/</guid><description/></item><item><title>Areca Technology Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/areca-technology-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/areca-technology-corporation-snmp-traps/</guid><description/></item><item><title>Argus Technologies SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/argus-technologies-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/argus-technologies-snmp-traps/</guid><description/></item><item><title>Aricent Communication Holdings Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aricent-communication-holdings-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aricent-communication-holdings-ltd-snmp-traps/</guid><description/></item><item><title>Arista</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/arista/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/arista/</guid><description/></item><item><title>Arista BGP</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/arista-bgp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/arista-bgp/</guid><description/></item><item><title>Arista Networks Inc Formerly Arastra Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/arista-networks-inc-formerly-arastra-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/arista-networks-inc-formerly-arastra-inc-snmp-traps/</guid><description/></item><item><title>Arista Switch</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/arista-switch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/arista-switch/</guid><description/></item><item><title>Armillaire Technologies SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/armillaire-technologies-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/armillaire-technologies-snmp-traps/</guid><description/></item><item><title>ARP / IP Neighbor Topology</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/arp---ip-neighbor-topology/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/arp---ip-neighbor-topology/</guid><description/></item><item><title>ARP cache staleness: when IP-to-MAC mapping goes bad</title><link>https://www.netdata.cloud/guides/network/network-arp-cache-staleness/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-arp-cache-staleness/</guid><description>&lt;h1 id="arp-cache-staleness-when-ip-to-mac-mapping-goes-bad">ARP cache staleness: when IP-to-MAC mapping goes bad&lt;/h1>
&lt;p>Hosts on the same subnet stop reaching each other after a VM live-migrates, a container gets rescheduled, or a firewall fails over. ICMP works from some hosts but not others. TCP sessions hang or reset. The data plane is healthy, but the ARP cache on one or more hosts holds a stale IP-to-MAC mapping.&lt;/p>
&lt;p>ARP cache staleness is the gap between when a MAC address changes and when every interested host learns about the change. On Linux, this gap is governed by the neighbor (NUD) state machine and its timing parameters. On Windows Vista and later, the neighbor cache follows the same RFC 4861 model. Both platforms default to roughly the same reachable time window: about 15 to 45 seconds before an entry transitions to a stale state, followed by a probe sequence that adds several more seconds before resolution or eviction.&lt;/p></description></item><item><title>Arris Interactive LLC SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/arris-interactive-llc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/arris-interactive-llc-snmp-traps/</guid><description/></item><item><title>Arrowpoint Communications Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/arrowpoint-communications-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/arrowpoint-communications-inc-snmp-traps/</guid><description/></item><item><title>Artel Video Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/artel-video-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/artel-video-systems-inc-snmp-traps/</guid><description/></item><item><title>Artem Gmbhmichael Marsanu Catrinel Catrinescu SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/artem-gmbhmichael-marsanu-catrinel-catrinescu-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/artem-gmbhmichael-marsanu-catrinel-catrinescu-snmp-traps/</guid><description/></item><item><title>Aruba</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/aruba/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/aruba/</guid><description/></item><item><title>Aruba A Hewlett Packard Enterprise Company SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aruba-a-hewlett-packard-enterprise-company-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aruba-a-hewlett-packard-enterprise-company-snmp-traps/</guid><description/></item><item><title>Aruba Access Point</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/aruba-access-point/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/aruba-access-point/</guid><description/></item><item><title>Aruba Clearpass</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/aruba-clearpass/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/aruba-clearpass/</guid><description/></item><item><title>Aruba CX Switch</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/aruba-cx-switch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/aruba-cx-switch/</guid><description/></item><item><title>Aruba Mobility Controller</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/aruba-mobility-controller/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/aruba-mobility-controller/</guid><description/></item><item><title>Aruba Switch</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/aruba-switch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/aruba-switch/</guid><description/></item><item><title>Aruba Wireless Controller</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/aruba-wireless-controller/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/aruba-wireless-controller/</guid><description/></item><item><title>Asante Technology SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/asante-technology-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/asante-technology-snmp-traps/</guid><description/></item><item><title>Ascend Communications Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ascend-communications-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ascend-communications-inc-snmp-traps/</guid><description/></item><item><title>Ascom Sweden AB SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ascom-sweden-ab-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ascom-sweden-ab-snmp-traps/</guid><description/></item><item><title>Asentria Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/asentria-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/asentria-corporation-snmp-traps/</guid><description/></item><item><title>Asetek SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/asetek-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/asetek-snmp-traps/</guid><description/></item><item><title>Askey Computer Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/askey-computer-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/askey-computer-corp-snmp-traps/</guid><description/></item><item><title>ASP.NET</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/asp.net/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/asp.net/</guid><description/></item><item><title>Astaro AG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/astaro-ag-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/astaro-ag-snmp-traps/</guid><description/></item><item><title>Asymmetric routing: why your path and latency measurements lie</title><link>https://www.netdata.cloud/guides/network/network-asymmetric-routing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-asymmetric-routing/</guid><description>&lt;h1 id="asymmetric-routing-why-your-path-and-latency-measurements-lie">Asymmetric routing: why your path and latency measurements lie&lt;/h1>
&lt;p>Your monitoring says the path is fine. Ping latency is normal, traceroute shows a clean route, and interface counters look healthy. But applications are slow, TCP sessions stall or reset, and users are complaining. Your tools are measuring only half the path.&lt;/p>
&lt;p>In asymmetric routing, traffic from host A to host B takes one path (P1) while return traffic from B to A takes a different path (P2). When P2 is degraded, congested, or broken, your measurements average the healthy forward path with the impaired return path. Every acknowledgment and response is fighting through a bad route while the aggregate looks acceptable.&lt;/p></description></item><item><title>At T SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/at-t-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/at-t-snmp-traps/</guid><description/></item><item><title>Ateme SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ateme-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ateme-snmp-traps/</guid><description/></item><item><title>Aten International Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aten-international-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aten-international-co-ltd-snmp-traps/</guid><description/></item><item><title>Atlas Computer Equipment Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/atlas-computer-equipment-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/atlas-computer-equipment-inc-snmp-traps/</guid><description/></item><item><title>Atm Forum SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/atm-forum-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/atm-forum-snmp-traps/</guid><description/></item><item><title>Atmel Hellas S A N SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/atmel-hellas-s-a-n-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/atmel-hellas-s-a-n-snmp-traps/</guid><description/></item><item><title>Atto Technology Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/atto-technology-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/atto-technology-inc-snmp-traps/</guid><description/></item><item><title>Audiocodes Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/audiocodes-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/audiocodes-ltd-snmp-traps/</guid><description/></item><item><title>Audiocodes Mediant SBC</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/audiocodes-mediant-sbc/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/audiocodes-mediant-sbc/</guid><description/></item><item><title>Audit log gaps: detecting syslog/trap tampering or loss</title><link>https://www.netdata.cloud/guides/network/network-audit-log-gap/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-audit-log-gap/</guid><description>&lt;h1 id="audit-log-gaps-detecting-syslogtrap-tampering-or-loss">Audit log gaps: detecting syslog/trap tampering or loss&lt;/h1>
&lt;p>An audit log gap is any period where expected syslog messages or SNMP traps from a network device fail to arrive at the collector. UDP syslog on port 514 and SNMP traps on port 162 are fire-and-forget transports with no delivery guarantee. The kernel silently drops datagrams when socket buffers fill, and the application layer never sees the loss. TCP syslog can stall under collector backpressure. Most gaps are operational: network loss, device-side buffer overflow, or logging subsystem failure. The difficulty is distinguishing those from deliberate log suppression after compromise.&lt;/p></description></item><item><title>Auditec S.A. SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/auditec-s.a.-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/auditec-s.a.-snmp-traps/</guid><description/></item><item><title>Auspex Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/auspex-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/auspex-systems-inc-snmp-traps/</guid><description/></item><item><title>Austin Hughes Electronics Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/austin-hughes-electronics-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/austin-hughes-electronics-ltd-snmp-traps/</guid><description/></item><item><title>AuthLog</title><link>https://www.netdata.cloud/integrations/data-collection/applications/authlog/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/authlog/</guid><description/></item><item><title>AuthLog Monitoring</title><link>https://www.netdata.cloud/monitoring-101/authlog-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/authlog-monitoring/</guid><description>&lt;h2 id="authlog-monitoring">AuthLog Monitoring&lt;/h2>
&lt;h3 id="what-is-authlog">What Is AuthLog?&lt;/h3>
&lt;p>AuthLog is a tool designed for monitoring authentication logs to gain security insights and facilitate efficient access management. Its primary function is to track and analyze authentication attempts, providing crucial data for assessing security posture and identifying unauthorized access attempts.&lt;/p>
&lt;h3 id="monitoring-authlog-with-netdata">Monitoring AuthLog With Netdata&lt;/h3>
&lt;p>To monitor AuthLog effectively, Netdata uses an openmetrics (Prometheus) exporter, the &lt;a href="https://github.com/woblerr/authlog_exporter">AuthLog Exporter&lt;/a>. Netdata&amp;rsquo;s robust monitoring capabilities allow it to ingest data from any Prometheus exporter. This means you can set up automated dashboards, alerts, and more without the need for a Prometheus server or Grafana. By utilizing Netdata, users can leverage real-time monitoring of their AuthLog metrics to ensure comprehensive security coverage and quick response to potential threats.&lt;/p></description></item><item><title>Availant SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/availant-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/availant-snmp-traps/</guid><description/></item><item><title>Avamar SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/avamar-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/avamar-snmp-traps/</guid><description/></item><item><title>Avaya</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avaya/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avaya/</guid><description/></item><item><title>Avaya Aura Media Server</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avaya-aura-media-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avaya-aura-media-server/</guid><description/></item><item><title>Avaya Cajun Switch</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avaya-cajun-switch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avaya-cajun-switch/</guid><description/></item><item><title>Avaya Communication SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/avaya-communication-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/avaya-communication-snmp-traps/</guid><description/></item><item><title>Avaya Media Gateway</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avaya-media-gateway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avaya-media-gateway/</guid><description/></item><item><title>Avaya Nortel Ethernet Routing Switch</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avaya-nortel-ethernet-routing-switch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avaya-nortel-ethernet-routing-switch/</guid><description/></item><item><title>Aventail Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aventail-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aventail-corporation-snmp-traps/</guid><description/></item><item><title>Aviat Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aviat-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aviat-networks-snmp-traps/</guid><description/></item><item><title>Avici Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/avici-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/avici-systems-inc-snmp-traps/</guid><description/></item><item><title>Avista Labs Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/avista-labs-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/avista-labs-inc-snmp-traps/</guid><description/></item><item><title>Avocent ACS</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avocent-acs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avocent-acs/</guid><description/></item><item><title>Avocent Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/avocent-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/avocent-corporation-snmp-traps/</guid><description/></item><item><title>Avtech</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avtech/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avtech/</guid><description/></item><item><title>Avtech Roomalert 32S</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avtech-roomalert-32s/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avtech-roomalert-32s/</guid><description/></item><item><title>Avtech Roomalert 3E</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avtech-roomalert-3e/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avtech-roomalert-3e/</guid><description/></item><item><title>Avtech Roomalert 3S</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avtech-roomalert-3s/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/avtech-roomalert-3s/</guid><description/></item><item><title>Avtech Software Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/avtech-software-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/avtech-software-inc-snmp-traps/</guid><description/></item><item><title>Aware Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aware-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/aware-inc-snmp-traps/</guid><description/></item><item><title>AWS EC2 Compute instances</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/aws-ec2-compute-instances/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/aws-ec2-compute-instances/</guid><description/></item><item><title>AWS EC2 Compute Instances Monitoring</title><link>https://www.netdata.cloud/monitoring-101/aws_ec2-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/aws_ec2-monitoring/</guid><description>&lt;h2 id="aws-ec2-compute-instances-monitoring">AWS EC2 Compute Instances Monitoring&lt;/h2>
&lt;h3 id="what-is-aws-ec2-compute-instances">What Is AWS EC2 Compute Instances?&lt;/h3>
&lt;p>AWS EC2 (Elastic Compute Cloud) instances are a key element of Amazon Web Services, providing scalable computing capacity in the cloud. They eliminate the need to invest in hardware, allowing companies to focus on their applications while benefiting from the flexibility of cloud infrastructure.&lt;/p>
&lt;h3 id="monitoring-aws-ec2-compute-instances-with-netdata">Monitoring AWS EC2 Compute Instances With Netdata&lt;/h3>
&lt;p>Netdata offers a comprehensive platform for monitoring AWS EC2 Compute Instances, highlighting the critical metrics that can significantly impact performance and cost management. Utilizing an openmetrics (Prometheus) exporter, Netdata can seamlessly ingest data from any Prometheus exporter, delivering automated dashboards, alerts, and more, all without the need for a Prometheus server or Grafana. This powerful capability allows technical users to monitor AWS EC2 with precision and minimum friction.&lt;/p></description></item><item><title>AWS ECS Containers</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/aws-ecs-containers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/aws-ecs-containers/</guid><description/></item><item><title>AWS IP Ranges</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/aws-ip-ranges/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/aws-ip-ranges/</guid><description/></item><item><title>AWS Kinesis</title><link>https://www.netdata.cloud/integrations/exporters/aws-kinesis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/aws-kinesis/</guid><description/></item><item><title>AWS Quota</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/aws-quota/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/aws-quota/</guid><description/></item><item><title>AWS Quota Monitoring</title><link>https://www.netdata.cloud/monitoring-101/aws_quota-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/aws_quota-monitoring/</guid><description>&lt;h2 id="aws-quota-monitoring">AWS Quota Monitoring&lt;/h2>
&lt;h3 id="what-is-aws-quota">What Is AWS Quota?&lt;/h3>
&lt;p>AWS Quota refers to the limits set on the number of resources and API requests you can use in Amazon Web Services. Effective management of AWS Quotas ensures optimal resource usage, cost management, and prevents disruption in services due to resource limitations.&lt;/p>
&lt;h3 id="monitoring-aws-quota-with-netdata">Monitoring AWS Quota With Netdata&lt;/h3>
&lt;p>To monitor AWS Quota with Netdata, you can leverage the &lt;a href="https://github.com/emylincon/aws_quota_exporter">aws_quota_exporter&lt;/a>, an openmetrics (prometheus) exporter. Netdata is designed to seamlessly ingest data from any Prometheus exporter, providing you with powerful, automated dashboards and alerts without the need for setting up a Prometheus server or Grafana. Netdata’s out-of-the-box integration simplifies the complexity associated with setting up AWS Quota monitoring, enabling you to focus on resource management and optimization.&lt;/p></description></item><item><title>AWS RDS</title><link>https://www.netdata.cloud/integrations/data-collection/databases/aws-rds/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/aws-rds/</guid><description/></item><item><title>AWS RDS Monitoring</title><link>https://www.netdata.cloud/monitoring-101/aws_rds-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/aws_rds-monitoring/</guid><description>&lt;h2 id="aws-rds-monitoring">AWS RDS Monitoring&lt;/h2>
&lt;h3 id="what-is-aws-rds">What Is AWS RDS?&lt;/h3>
&lt;p>Amazon RDS (Relational Database Service) is a managed cloud database service that simplifies database administration tasks such as hardware provisioning, database setup, patching, and backups. It enables organizations to efficiently scale databases while managing performance, availability, and security. AWS RDS supports various database engines, including MySQL, PostgreSQL, MariaDB, Oracle, and SQL Server, making it a versatile choice for cloud-based database solutions.&lt;/p>
&lt;h3 id="monitoring-aws-rds-with-netdata">Monitoring AWS RDS With Netdata&lt;/h3>
&lt;p>Monitoring AWS RDS is critical for ensuring the performance, availability, and reliability of your cloud databases. Netdata provides a powerful tool for monitoring AWS RDS using an openmetrics (Prometheus) exporter. Netdata can ingest real-time metrics from any Prometheus exporter, including the &lt;a href="https://github.com/percona/rds_exporter">rds_exporter&lt;/a>, without needing a separate Prometheus server or Grafana setup.&lt;/p></description></item><item><title>AWS Secrets Manager</title><link>https://www.netdata.cloud/integrations/all/aws-secrets-manager/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/all/aws-secrets-manager/</guid><description/></item><item><title>AWS SNS</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/aws-sns/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/aws-sns/</guid><description/></item><item><title>Axis Communications AB SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/axis-communications-ab-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/axis-communications-ab-snmp-traps/</guid><description/></item><item><title>Azure API Management</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-api-management/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-api-management/</guid><description/></item><item><title>Azure App Service</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-app-service/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-app-service/</guid><description/></item><item><title>Azure Application Gateway</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-application-gateway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-application-gateway/</guid><description/></item><item><title>Azure Application Insights</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-application-insights/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-application-insights/</guid><description/></item><item><title>Azure Cache for Redis</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-cache-for-redis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-cache-for-redis/</guid><description/></item><item><title>Azure Cognitive Services</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-cognitive-services/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-cognitive-services/</guid><description/></item><item><title>Azure Container Apps</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-container-apps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-container-apps/</guid><description/></item><item><title>Azure Container Instances</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-container-instances/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-container-instances/</guid><description/></item><item><title>Azure Container Registry</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-container-registry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-container-registry/</guid><description/></item><item><title>Azure Cosmos DB Account</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-cosmos-db-account/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-cosmos-db-account/</guid><description/></item><item><title>Azure Data Explorer</title><link>https://www.netdata.cloud/integrations/exporters/azure-data-explorer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/azure-data-explorer/</guid><description/></item><item><title>Azure Data Explorer Cluster</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-data-explorer-cluster/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-data-explorer-cluster/</guid><description/></item><item><title>Azure Data Factory</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-data-factory/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-data-factory/</guid><description/></item><item><title>Azure Event Grid Topic</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-event-grid-topic/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-event-grid-topic/</guid><description/></item><item><title>Azure Event Hub</title><link>https://www.netdata.cloud/integrations/exporters/azure-event-hub/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/azure-event-hub/</guid><description/></item><item><title>Azure Event Hubs Namespace</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-event-hubs-namespace/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-event-hubs-namespace/</guid><description/></item><item><title>Azure ExpressRoute Circuit</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-expressroute-circuit/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-expressroute-circuit/</guid><description/></item><item><title>Azure ExpressRoute Gateway</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-expressroute-gateway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-expressroute-gateway/</guid><description/></item><item><title>Azure Firewall</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-firewall/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-firewall/</guid><description/></item><item><title>Azure Front Door</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-front-door/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-front-door/</guid><description/></item><item><title>Azure Functions</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-functions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-functions/</guid><description/></item><item><title>Azure IoT Hub</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-iot-hub/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-iot-hub/</guid><description/></item><item><title>Azure IP Ranges</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/azure-ip-ranges/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/azure-ip-ranges/</guid><description/></item><item><title>Azure Key Vault</title><link>https://www.netdata.cloud/integrations/all/azure-key-vault/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/all/azure-key-vault/</guid><description/></item><item><title>Azure Key Vault</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-key-vault/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-key-vault/</guid><description/></item><item><title>Azure Kubernetes Service Cluster</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-kubernetes-service-cluster/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-kubernetes-service-cluster/</guid><description/></item><item><title>Azure Load Balancer</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-load-balancer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-load-balancer/</guid><description/></item><item><title>Azure Log Analytics Workspace</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-log-analytics-workspace/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-log-analytics-workspace/</guid><description/></item><item><title>Azure Logic Apps Workflow</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-logic-apps-workflow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-logic-apps-workflow/</guid><description/></item><item><title>Azure Machine Learning Workspace</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-machine-learning-workspace/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-machine-learning-workspace/</guid><description/></item><item><title>Azure Monitor</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-monitor/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-monitor/</guid><description/></item><item><title>Azure MySQL Flexible Server</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-mysql-flexible-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-mysql-flexible-server/</guid><description/></item><item><title>Azure NAT Gateway</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-nat-gateway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-nat-gateway/</guid><description/></item><item><title>Azure PostgreSQL Flexible Server</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-postgresql-flexible-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-postgresql-flexible-server/</guid><description/></item><item><title>Azure Service Bus Namespace</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-service-bus-namespace/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-service-bus-namespace/</guid><description/></item><item><title>Azure SQL Database</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-sql-database/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-sql-database/</guid><description/></item><item><title>Azure SQL Elastic Pool</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-sql-elastic-pool/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-sql-elastic-pool/</guid><description/></item><item><title>Azure SQL Managed Instance</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-sql-managed-instance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-sql-managed-instance/</guid><description/></item><item><title>Azure Storage Account</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-storage-account/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-storage-account/</guid><description/></item><item><title>Azure Stream Analytics Job</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-stream-analytics-job/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-stream-analytics-job/</guid><description/></item><item><title>Azure Synapse Analytics Workspace</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-synapse-analytics-workspace/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-synapse-analytics-workspace/</guid><description/></item><item><title>Azure Virtual Machine</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-virtual-machine/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-virtual-machine/</guid><description/></item><item><title>Azure Virtual Machine Scale Set</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-virtual-machine-scale-set/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-virtual-machine-scale-set/</guid><description/></item><item><title>Azure VPN Gateway</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-vpn-gateway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/azure-vpn-gateway/</guid><description/></item><item><title>B A T M Advance Technologies SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/b-a-t-m-advance-technologies-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/b-a-t-m-advance-technologies-snmp-traps/</guid><description/></item><item><title>Bachmann GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bachmann-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bachmann-gmbh-snmp-traps/</guid><description/></item><item><title>Bancomm SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bancomm-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bancomm-snmp-traps/</guid><description/></item><item><title>Barco Control Rooms SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/barco-control-rooms-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/barco-control-rooms-snmp-traps/</guid><description/></item><item><title>Barix AG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/barix-ag-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/barix-ag-snmp-traps/</guid><description/></item><item><title>Barracuda</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/barracuda/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/barracuda/</guid><description/></item><item><title>Barracuda Cloudgen</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/barracuda-cloudgen/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/barracuda-cloudgen/</guid><description/></item><item><title>Barracuda Networks AG Previous Was Phion Information Technologies SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/barracuda-networks-ag-previous-was-phion-information-technologies-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/barracuda-networks-ag-previous-was-phion-information-technologies-snmp-traps/</guid><description/></item><item><title>Barracuda Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/barracuda-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/barracuda-networks-inc-snmp-traps/</guid><description/></item><item><title>Bay Technical Associates SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bay-technical-associates-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bay-technical-associates-snmp-traps/</guid><description/></item><item><title>BCache</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/bcache/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/bcache/</guid><description/></item><item><title>Bdt GmbH Co KG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bdt-gmbh-co-kg-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bdt-gmbh-co-kg-snmp-traps/</guid><description/></item><item><title>Beanstalk</title><link>https://www.netdata.cloud/integrations/data-collection/databases/beanstalk/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/beanstalk/</guid><description/></item><item><title>Beanstalk Monitoring</title><link>https://www.netdata.cloud/monitoring-101/beanstalk-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/beanstalk-monitoring/</guid><description>&lt;h2 id="beanstalk-monitoring">Beanstalk Monitoring&lt;/h2>
&lt;h3 id="what-is-beanstalk">What Is Beanstalk?&lt;/h3>
&lt;p>Beanstalk is a simple and fast work queue. It is designed to reduce the complexity of creating distributed applications and focuses on speed and reliability. Beanstalk helps developers manage background jobs easily, allowing processes to be executed asynchronously.&lt;/p>
&lt;h3 id="monitoring-beanstalk-with-netdata">Monitoring Beanstalk With Netdata&lt;/h3>
&lt;p>Monitor Beanstalk effortlessly using Netdata&amp;rsquo;s &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/beanstalk/?utm_source=website&amp;amp;utm_content=monitoring101">beanstalk monitoring tool&lt;/a>. Netdata provides real-time performance metrics and analytics, enabling you to keep a close eye on your Beanstalk servers and ensure they are running smoothly. Whether you are dealing with data collection, message brokers, or distributed systems, optimizing performance and preventing resource bottlenecks has never been easier with Netdata.&lt;/p></description></item><item><title>Beijing Raisecom Scientific Technology Development Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/beijing-raisecom-scientific-technology-development-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/beijing-raisecom-scientific-technology-development-co-ltd-snmp-traps/</guid><description/></item><item><title>Bekarts International SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bekarts-international-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bekarts-international-snmp-traps/</guid><description/></item><item><title>Bellcore SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bellcore-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bellcore-snmp-traps/</guid><description/></item><item><title>Benu Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/benu-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/benu-networks-inc-snmp-traps/</guid><description/></item><item><title>Beronet GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/beronet-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/beronet-gmbh-snmp-traps/</guid><description/></item><item><title>Best Power A Division Of General Signal Power Systems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/best-power-a-division-of-general-signal-power-systems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/best-power-a-division-of-general-signal-power-systems-snmp-traps/</guid><description/></item><item><title>Better Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/better-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/better-networks-snmp-traps/</guid><description/></item><item><title>BGP flapping: why a peer keeps resetting and how to find the cause</title><link>https://www.netdata.cloud/guides/network/network-bgp-flapping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-bgp-flapping/</guid><description>&lt;h1 id="bgp-flapping-why-a-peer-keeps-resetting-and-how-to-find-the-cause">BGP flapping: why a peer keeps resetting and how to find the cause&lt;/h1>
&lt;p>A BGP peer cycling between Established and Idle is sending a specific signal. The session tears down because one side sent a NOTIFICATION message, and that message carries an error code and subcode that pinpoints the cause. Most monitoring watches only the FSM state (up or down) and ignores the NOTIFICATION payload, so the operator sees flapping without knowing why.&lt;/p></description></item><item><title>BGP NOTIFICATION and Cease messages: what each subcode is telling you</title><link>https://www.netdata.cloud/guides/network/network-bgp-notification-cease/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-bgp-notification-cease/</guid><description>&lt;h1 id="bgp-notification-and-cease-messages-what-each-subcode-is-telling-you">BGP NOTIFICATION and Cease messages: what each subcode is telling you&lt;/h1>
&lt;p>A BGP NOTIFICATION in your router log is a peer telling you why it tore down the session. The message carries an error code and an error subcode. Those two numbers tell you whether you are looking at a maintenance window, a route leak, a prefix-limit hit, a CPU-starved control plane, or a BFD-triggered teardown.&lt;/p>
&lt;p>Cease (code 6) is the most common NOTIFICATION. Its subcodes, defined in RFC 4486 and extended by RFC 8538 and RFC 9384, hold most of the diagnostic value. Codes 2 through 5 appear less often but point to distinct failure classes: parameter mismatch, malformed updates, hold-timer expiry, and FSM errors.&lt;/p></description></item><item><title>BGP Peering Topology</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/bgp-peering-topology/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/bgp-peering-topology/</guid><description/></item><item><title>BGP RIB and FIB growth: monitoring route-table size before it bites</title><link>https://www.netdata.cloud/guides/network/network-bgp-rib-fib-growth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-bgp-rib-fib-growth/</guid><description>&lt;h1 id="bgp-rib-and-fib-growth-monitoring-route-table-size-before-it-bites">BGP RIB and FIB growth: monitoring route-table size before it bites&lt;/h1>
&lt;p>The global BGP routing table grows every year. In 2026, the IPv4 default-free zone sits at approximately 940,000 prefixes, with IPv6 adding roughly 190,000 more. &lt;!-- TODO: verify 2026 DFZ sizes against current Potaroo/APNIC data; these may be conservative given the 2024 IPv4 baseline was already around 950k --> These numbers increase steadily, and the hardware that programs forwarding decisions from them has finite capacity. When that capacity runs out, new routes do not get installed in the forwarding plane. Traffic to affected destinations blackholes. The BGP session stays Established the entire time.&lt;/p></description></item><item><title>BGP route leak and hijack: the detection signals and alerts that matter</title><link>https://www.netdata.cloud/guides/network/network-bgp-route-leak-hijack/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-bgp-route-leak-hijack/</guid><description>&lt;h1 id="bgp-route-leak-and-hijack-the-detection-signals-and-alerts-that-matter">BGP route leak and hijack: the detection signals and alerts that matter&lt;/h1>
&lt;p>A BGP route leak or hijack does not tear down your session. The peering stays Established, keepalives flow, and the session &amp;ldquo;up&amp;rdquo; indicator stays green. What changes is which prefixes your network believes are reachable, through which origin AS, and via what path. Traffic is silently misrouted or blackholed while the session looks healthy.&lt;/p>
&lt;p>BGP has no built-in authentication of route ownership. Any AS can announce any prefix. Whether other networks accept the announcement depends on their filtering, and filtering is inconsistently deployed. Approximately 50% of routable IP prefixes carry a Route Origin Authorization (ROA), and only about 6.5% of Internet users sit behind networks that actively reject RPKI-invalid routes. &lt;!-- TODO: verify exact adoption percentages as of current date --> That gap is where leaks and hijacks propagate globally before anyone notices.&lt;/p></description></item><item><title>BGP session Established but stale: detecting silent route loss</title><link>https://www.netdata.cloud/guides/network/network-bgp-session-stale/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-bgp-session-stale/</guid><description>&lt;h1 id="bgp-session-established-but-stale-detecting-silent-route-loss">BGP session Established but stale: detecting silent route loss&lt;/h1>
&lt;p>Your BGP session to a transit provider or iBGP peer is Established, but destinations are unreachable. The RIB is missing prefixes from that peer, or the routes it has are stale. No NOTIFICATION was sent, no session flap occurred, and your monitoring trusts the FSM state.&lt;/p>
&lt;p>This is the &amp;ldquo;Established but stale&amp;rdquo; pattern. KEEPALIVEs are still exchanged at the TCP level, but the UPDATE exchange has stopped. The peer stopped sending routes, a middlebox is silently dropping UPDATE packets, or Graceful Restart is holding the session open after the remote side went down.&lt;/p></description></item><item><title>Bharti Telesoft International Pvt Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bharti-telesoft-international-pvt-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bharti-telesoft-international-pvt-ltd-snmp-traps/</guid><description/></item><item><title>BIND 'no more recursive clients: quota reached': the recursive-clients circuit breaker</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-no-more-recursive-clients/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-no-more-recursive-clients/</guid><description>&lt;h1 id="bind-no-more-recursive-clients-quota-reached-the-recursive-clients-circuit-breaker">BIND &amp;ldquo;no more recursive clients: quota reached&amp;rdquo;: the recursive-clients circuit breaker&lt;/h1>
&lt;p>The log message &lt;code>no more recursive clients (N/M): quota reached&lt;/code> means BIND has run out of slots for in-flight recursive queries, and new queries are failing. Resolution for your clients is already degraded or broken.&lt;/p>
&lt;p>The &lt;code>recursive-clients&lt;/code> option (default 1000) caps the number of concurrent upstream fetches the resolver can have outstanding. A soft quota at 90 percent (default 900) starts shedding load before the hard limit. Once the hard limit is reached, every new recursive query returns SERVFAIL, including queries for domains whose upstream nameservers are healthy.&lt;/p></description></item><item><title>BIND 'too many open files': file descriptor exhaustion and silently dropped queries</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-too-many-open-files/</guid><description>&lt;h1 id="bind-too-many-open-files-file-descriptor-exhaustion-and-silently-dropped-queries">BIND &amp;rsquo;too many open files&amp;rsquo;: file descriptor exhaustion and silently dropped queries&lt;/h1>
&lt;p>BIND logs &amp;ldquo;too many open files&amp;rdquo; or &amp;ldquo;socket: file descriptor exceeds limit&amp;rdquo; and starts dropping queries. UDP health checks may still pass. Zone transfers fail intermittently. The &lt;code>rndc&lt;/code> control channel becomes sluggish or unresponsive. Clients see random timeouts that look like upstream nameserver problems. This is file descriptor exhaustion, and the symptoms masquerade as network or disk issues.&lt;/p></description></item><item><title>BIND 9 Monitoring</title><link>https://www.netdata.cloud/monitoring-101/bind9-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/bind9-monitoring/</guid><description>&lt;h2 id="what-is-bind-9">What is BIND 9?&lt;/h2>
&lt;p>BIND 9 is a flexible, full-featured open source DNS system.&lt;/p>
&lt;h2 id="monitoring-bind-9-with-netdata">Monitoring BIND 9 with Netdata&lt;/h2>
&lt;p>The prerequisites for monitoring &lt;a href="https://www.isc.org/bind/">BIND 9&lt;/a> with Netdata are to have BIND and &lt;a href="https://learn.netdata.cloud/docs/cloud/get-started/">Netdata installed&lt;/a> on your system.&lt;/p>
&lt;p>Netdata auto discovers hundreds of services, and for those it doesn&amp;rsquo;t turning on manual discovery is a one line configuration. For more information on configuring Netdata for BIND 9 monitoring please read the collector &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/bind/">documentation&lt;/a>.&lt;/p>
&lt;p>You should now see the &lt;code>bind&lt;/code> section on the Overview tab in Netdata Cloud already populated with charts about all the metrics you care about.&lt;/p></description></item><item><title>BIND cache eviction storms: DeleteLRU, an undersized max-cache-size, and the pressure spiral</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-cache-eviction-deletelru/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-cache-eviction-deletelru/</guid><description>&lt;h1 id="bind-cache-eviction-storms-deletelru-an-undersized-max-cache-size-and-the-pressure-spiral">BIND cache eviction storms: DeleteLRU, an undersized max-cache-size, and the pressure spiral&lt;/h1>
&lt;p>DeleteLRU is rising fast while DeleteTTL barely moves. Cache hit ratio is dropping. Outbound recursive queries are climbing. RecursClients is trending toward its limit. This is a cache eviction storm: a self-reinforcing loop that can end in SERVFAIL for any query requiring recursion.&lt;/p>
&lt;p>The cache is too small for the working set. BIND evicts entries by LRU before their TTLs expire. Each eviction forces a cache miss on the next query for that name, triggering an outbound recursive fetch. More fetches mean more concurrent recursive clients, higher latency per resolution, and less cache room as new entries from upstream compete with entries still under eviction pressure.&lt;/p></description></item><item><title>BIND cache hit ratio dropping: the leading edge of recursive pain</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-cache-hit-ratio-dropping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-cache-hit-ratio-dropping/</guid><description>&lt;h1 id="bind-cache-hit-ratio-dropping-the-leading-edge-of-recursive-pain">BIND cache hit ratio dropping: the leading edge of recursive pain&lt;/h1>
&lt;p>A dropping cache hit ratio is rarely the problem itself. It is the leading indicator that something is about to get worse. When fewer queries are answered from cache, each miss consumes a recursive-client slot, adds latency, and increases exposure to upstream slowness. On a busy resolver, a sustained hit-ratio decline from 95% to 80% can roughly triple outbound query volume and push recursive-client utilization into the danger zone.&lt;/p></description></item><item><title>BIND clients-per-query and max-clients-per-query: duplicate recursion for popular names</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-clients-per-query/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-clients-per-query/</guid><description>&lt;h1 id="bind-clients-per-query-and-max-clients-per-query-duplicate-recursion-for-popular-names">BIND clients-per-query and max-clients-per-query: duplicate recursion for popular names&lt;/h1>
&lt;p>When hundreds of clients query the same domain at the same instant, BIND does not send hundreds of identical recursive queries upstream. It sends one fetch and attaches the waiting clients to it. Two configuration knobs control how many waiters can attach before BIND starts dropping the overflow: &lt;code>clients-per-query&lt;/code> (the soft limit, default 10) and &lt;code>max-clients-per-query&lt;/code> (the hard ceiling, default 100).&lt;/p></description></item><item><title>BIND cold cache after restart: the warming storm and elevated upstream load</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-cold-cache-warming/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-cold-cache-warming/</guid><description>&lt;p>After a &lt;code>named&lt;/code> restart, the recursive cache is empty. Every incoming query is a cache miss. On a resolver doing 50,000 qps, that means at least 50,000 outbound recursive queries per second to upstream authoritative servers &amp;ndash; a 10x increase over steady-state outbound volume. This is the cache-warming storm, and it lasts 30 to 60 minutes.&lt;/p>
&lt;p>Three things happen simultaneously: upstream load spikes, the &lt;code>recursive-clients&lt;/code> table fills toward its limit, and cache hit ratio starts at zero and climbs. For the first 30 to 60 seconds, some queries may return SERVFAIL while root priming completes and authoritative zones finish loading. All of this is expected behavior.&lt;/p></description></item><item><title>BIND CPU saturation: single-core bottlenecks, DNSSEC crypto, and per-thread contention</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-cpu-single-core-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-cpu-single-core-saturation/</guid><description>&lt;h1 id="bind-cpu-saturation-single-core-bottlenecks-dnssec-crypto-and-per-thread-contention">BIND CPU saturation: single-core bottlenecks, DNSSEC crypto, and per-thread contention&lt;/h1>
&lt;p>BIND CPU saturation often hides from aggregate monitoring. Because &lt;code>named&lt;/code> runs one worker thread per CPU core through netmgr, aggregate process CPU can read 25% on a 4-core host while a single core is pinned at 100%. Standard monitoring that checks total CPU utilization sees a healthy daemon. Clients see latency spikes and intermittent timeouts.&lt;/p>
&lt;p>The core diagnostic challenge: BIND&amp;rsquo;s statistics channel reports no per-thread CPU data. You must measure at the OS level with &lt;code>mpstat&lt;/code>, &lt;code>pidstat&lt;/code>, or direct &lt;code>/proc&lt;/code> inspection. Without per-core visibility, the symptom presents as unexplained latency with no obvious cause.&lt;/p></description></item><item><title>BIND DNSSEC failing from clock drift: NTP, RRSIG inception/expiry windows, and SERVFAIL</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-dnssec-validation-clock-drift/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-dnssec-validation-clock-drift/</guid><description>&lt;h1 id="bind-dnssec-failing-from-clock-drift-ntp-rrsig-inceptionexpiry-windows-and-servfail">BIND DNSSEC failing from clock drift: NTP, RRSIG inception/expiry windows, and SERVFAIL&lt;/h1>
&lt;p>Signed domains start returning SERVFAIL while unsigned domains resolve normally. The &lt;code>named&lt;/code> process is running, CPU is moderate, and the cache hit ratio has not collapsed. There was no recent &lt;code>rndc reload&lt;/code> or configuration change. The failure appeared gradually over minutes or hours, not all at once.&lt;/p>
&lt;p>This is the fingerprint of a DNSSEC validation failure caused by clock drift. Every RRSIG record carries an inception timestamp (when the signature becomes valid) and an expiry timestamp (when it stops being valid). BIND compares these against the local system clock. If the clock drifts far enough outside the RRSIG validity window, valid signatures appear expired or not-yet-valid, and BIND returns SERVFAIL for the entire domain.&lt;/p></description></item><item><title>BIND DNSSEC validation failing: 'broken trust chain', ValFail, and SERVFAIL for signed domains</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-broken-trust-chain/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-broken-trust-chain/</guid><description>&lt;h1 id="bind-dnssec-validation-failing-broken-trust-chain-valfail-and-servfail-for-signed-domains">BIND DNSSEC validation failing: &amp;lsquo;broken trust chain&amp;rsquo;, ValFail, and SERVFAIL for signed domains&lt;/h1>
&lt;p>Your resolver is returning SERVFAIL for major signed domains. Logs show &lt;code>broken trust chain resolving 'example.com'&lt;/code> across dozens of unrelated domains. The per-view &lt;code>ValFail&lt;/code> counter is climbing. Unsigned domains resolve normally, and authoritative zones you serve locally still answer.&lt;/p>
&lt;p>This is a DNSSEC validation failure on the resolver side. BIND is rejecting signed answers because something in the chain of trust is broken locally. When the breakage is local (clock drift, stale trust anchors, corrupted managed-keys), every signed domain fails simultaneously. When it is upstream (an individual zone with expired signatures), only that zone is affected.&lt;/p></description></item><item><title>BIND dnssec-validation disabled: the security regression that 'fixes' SERVFAIL</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-dnssec-validation-disabled/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-dnssec-validation-disabled/</guid><description>&lt;h1 id="bind-dnssec-validation-disabled-the-security-regression-that-fixes-servfail">BIND dnssec-validation disabled: the security regression that &amp;lsquo;fixes&amp;rsquo; SERVFAIL&lt;/h1>
&lt;p>SERVFAIL on signed domains is a real production problem. Users cannot reach major services because .com, .org, google.com, and countless other zones are DNSSEC-signed. When validation fails, those names stop resolving. The pressure to restore service is immediate.&lt;/p>
&lt;p>The fastest way to make the SERVFAIL disappear is one line in &lt;code>named.conf&lt;/code>:&lt;/p>
&lt;pre tabindex="0">&lt;code>dnssec-validation no;
&lt;/code>&lt;/pre>&lt;p>After a reload, every signed domain resolves again. The monitoring dashboard goes green. The ticket closes. Everything works, including cache-poisoning attacks. The resolver now accepts forged responses for every query it processes.&lt;/p></description></item><item><title>BIND dynamic update failures: UpdateFail, denied updates, and TSIG drift</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-dynamic-update-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-dynamic-update-failures/</guid><description>&lt;h1 id="bind-dynamic-update-failures-updatefail-denied-updates-and-tsig-drift">BIND dynamic update failures: UpdateFail, denied updates, and TSIG drift&lt;/h1>
&lt;p>A rising UpdateFail counter on a BIND authoritative server means the dynamic update pipeline is broken. The counter does not tell you why. You need to correlate it with the security log, the zone serial, and the journal file state to narrow the cause.&lt;/p>
&lt;p>Failures cluster into three families:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Authentication failures&lt;/strong>: TSIG key drift, wrong key names, clock skew past the 300-second fudge window.&lt;/li>
&lt;li>&lt;strong>Authorization failures&lt;/strong>: the zone&amp;rsquo;s &lt;code>allow-update&lt;/code> or &lt;code>update-policy&lt;/code> does not grant the presented key permission to modify the zone.&lt;/li>
&lt;li>&lt;strong>Journal and filesystem failures&lt;/strong>: disk exhaustion, permission errors, &lt;code>.jnl&lt;/code> corruption from manual zone edits.&lt;/li>
&lt;/ul>
&lt;p>Each produces a distinct signal pattern in the logs and counters. Dynamic updates are denied by default. A zone must explicitly include &lt;code>allow-update&lt;/code> or &lt;code>update-policy&lt;/code> to accept updates. These two directives are mutually exclusive; configuring both is a configuration error and named will refuse to load the zone.&lt;/p></description></item><item><title>BIND forwarding loops: recursion that never terminates and burns recursive slots</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-forwarding-loop/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-forwarding-loop/</guid><description>&lt;h1 id="bind-forwarding-loops-recursion-that-never-terminates-and-burns-recursive-slots">BIND forwarding loops: recursion that never terminates and burns recursive slots&lt;/h1>
&lt;p>Recursive-clients is climbing toward its limit. SERVFAIL responses are rising across unrelated domains. QueryTimeout counters are ticking up in the per-view resolver stats. It looks like a textbook recursive resolution cascade: an upstream nameserver is slow or unreachable, in-flight queries are piling up, and BIND&amp;rsquo;s circuit breaker is about to trip.&lt;/p>
&lt;p>But when you run &lt;code>rndc recursing&lt;/code> to identify the culprit, the upstream IP addresses are not external nameservers. They are your own infrastructure. Another BIND resolver you control, or this very server.&lt;/p></description></item><item><title>BIND inline signing silently failed: missing keys and a zone served unsigned</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-inline-signing-silent-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-inline-signing-silent-failure/</guid><description>&lt;h1 id="bind-inline-signing-silently-failed-missing-keys-and-a-zone-served-unsigned">BIND inline signing silently failed: missing keys and a zone served unsigned&lt;/h1>
&lt;p>Your BIND authoritative server is running. Queries return answers. The zone loads without error. But validating resolvers worldwide are returning SERVFAIL for your domain, and there is no error line in your logs. The zone is configured for inline signing, but it is being served unsigned.&lt;/p>
&lt;p>When a zone has &lt;code>inline-signing yes;&lt;/code> with &lt;code>auto-dnssec maintain;&lt;/code> or &lt;code>dnssec-policy default;&lt;/code>, BIND should sign the zone transparently and maintain valid RRSIG records. If the signing key files are missing from the key-directory, BIND loads the unsigned zone and serves it as-is. No startup error. No log message. No statistics counter.&lt;/p></description></item><item><title>BIND journal (.jnl) corruption: dynamic-update and IXFR failures that block zone load</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-journal-corruption/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-journal-corruption/</guid><description>&lt;h1 id="bind-journal-jnl-corruption-dynamic-update-and-ixfr-failures-that-block-zone-load">BIND journal (.jnl) corruption: dynamic-update and IXFR failures that block zone load&lt;/h1>
&lt;p>A zone fails to load after a restart or &lt;code>rndc reload&lt;/code>. The &lt;code>named&lt;/code> process is running, other zones respond normally, and &lt;code>named-checkzone&lt;/code> reports the zone file is syntactically valid. The problem is the journal file (&lt;code>.jnl&lt;/code>) alongside the zone file, which records pending dynamic updates and drives IXFR between primary and secondary servers.&lt;/p>
&lt;p>Journal corruption has two operational faces. The acute case: &lt;code>named&lt;/code> starts but refuses to serve the affected zone, logging &amp;ldquo;journal rollforward failed: journal out of sync with zone.&amp;rdquo; The subtle case: IXFR transfers keep failing and falling back to full AXFR. The zone stays current, but each refresh pulls the entire zone over TCP instead of just the diff, wasting bandwidth and CPU.&lt;/p></description></item><item><title>BIND lame delegations: 'lame server resolving' and nameservers that are not authoritative</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-lame-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-lame-server/</guid><description>&lt;h1 id="bind-lame-delegations-lame-server-resolving-and-nameservers-that-are-not-authoritative">BIND lame delegations: &amp;rsquo;lame server resolving&amp;rsquo; and nameservers that are not authoritative&lt;/h1>
&lt;p>The &lt;code>lame server resolving&lt;/code> message in BIND&amp;rsquo;s &lt;code>lame-servers&lt;/code> log category means the resolver contacted a delegated nameserver that answered but was not authoritative for the zone it was supposed to serve. Each lame encounter wastes a fetch cycle and adds latency. On busy resolvers with high query volume for affected zones, the cost compounds. The per-view &lt;code>Lame&lt;/code> counter in resolver statistics tracks how often this happens.&lt;/p></description></item><item><title>BIND managed-keys and trust anchors: KSK rollover, RFC 5011, and a stale root key</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-managed-keys-trust-anchor/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-managed-keys-trust-anchor/</guid><description>&lt;h1 id="bind-managed-keys-and-trust-anchors-ksk-rollover-rfc-5011-and-a-stale-root-key">BIND managed-keys and trust anchors: KSK rollover, RFC 5011, and a stale root key&lt;/h1>
&lt;p>DNSSEC validation depends on a chain of trust anchored at the root zone&amp;rsquo;s Key Signing Key (KSK). BIND maintains that anchor automatically using RFC 5011 trust anchor management, stored in a managed-keys database. When it breaks, every signed domain on the internet fails validation simultaneously, and the resolver returns SERVFAIL for the majority of real-world queries.&lt;/p></description></item><item><title>BIND max-cache-size: sizing the resolver cache without triggering the OOM killer</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-max-cache-size-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-max-cache-size-tuning/</guid><description>&lt;h1 id="bind-max-cache-size-sizing-the-resolver-cache-without-triggering-the-oom-killer">BIND max-cache-size: sizing the resolver cache without triggering the OOM killer&lt;/h1>
&lt;p>&lt;code>named&lt;/code> disappears from the process table. &lt;code>dmesg&lt;/code> shows the OOM killer selected it. systemd restarts it, the cache is cold, every query triggers recursion, upstream load spikes, and RSS climbs again. Within hours the cycle repeats. The root cause is often not a memory leak. It is the default &lt;code>max-cache-size&lt;/code>.&lt;/p>
&lt;p>Since BIND 9.11, the default &lt;code>max-cache-size&lt;/code> for views with &lt;code>recursion yes&lt;/code> is 90% of physical memory. On a dedicated resolver with 16 GB of RAM, that is 14.4 GB for the cache alone. On a shared box or a mixed-role server that also serves authoritative zones, that default guarantees the OOM killer will eventually visit.&lt;/p></description></item><item><title>BIND monitoring checklist: the signals every production resolver and authoritative server needs</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-monitoring-checklist/</guid><description>&lt;h1 id="bind-monitoring-checklist-the-signals-every-production-resolver-and-authoritative-server-needs">BIND monitoring checklist: the signals every production resolver and authoritative server needs&lt;/h1>
&lt;p>BIND (&lt;code>named&lt;/code>) processes queries through a pipeline: receive packet, parse, ACL/view check, cache lookup (if recursive), zone lookup (if authoritative), recursive fetch on cache miss, apply RPZ/DNSSEC, serialize response, send. Every signal below maps to a stage in that pipeline or a resource it competes for: CPU, memory, file descriptors, network buffers, source ports.&lt;/p>
&lt;p>The levels are cumulative. Most signals come from the statistics channel (JSON at &lt;code>/json/v1/server&lt;/code>, explicitly configured in &lt;code>named.conf&lt;/code>). The &lt;code>rndc stats&lt;/code> file is an alternative but appends indefinitely and can fill disk if not rotated. Commands below assume a single &lt;code>named&lt;/code> process (standard deployment; BIND is multithreaded, not multiprocess).&lt;/p></description></item><item><title>BIND monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-monitoring-maturity-model/</guid><description>&lt;h1 id="bind-monitoring-maturity-model-from-survival-to-expert">BIND monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most BIND deployments catch complete outages but miss the slow-burn failures that cause real incidents: cache pressure spirals, recursive client exhaustion, DNSSEC signature expiry, and kernel-level UDP drops that BIND itself never sees.&lt;/p>
&lt;p>This article maps four monitoring maturity levels, from Survival to Expert. Each level adds signals that catch failure patterns invisible to the previous one. Use this as an inventory checklist: identify your current level, then decide which signals to add next based on whether you run a recursive resolver, an authoritative-only server, or a mixed-role deployment.&lt;/p></description></item><item><title>BIND named killed by the OOM killer: memory exhaustion and the cold-restart storm</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-named-oom-killed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-named-oom-killed/</guid><description>&lt;h1 id="bind-named-killed-by-the-oom-killer-memory-exhaustion-and-the-cold-restart-storm">BIND named killed by the OOM killer: memory exhaustion and the cold-restart storm&lt;/h1>
&lt;p>&lt;code>named&lt;/code> vanished from the process table. &lt;code>systemctl status named&lt;/code> shows &lt;code>inactive (dead)&lt;/code> or a restart timestamp from minutes ago. The likely cause: the Linux OOM killer terminated &lt;code>named&lt;/code> when system memory was exhausted.&lt;/p>
&lt;p>The cycle is self-reinforcing. BIND&amp;rsquo;s RSS grows steadily (excessive cache allocation, oversized RPZ, or a version-specific leak) until the kernel OOM killer selects &lt;code>named&lt;/code> as the victim. All DNS resolution fails instantly. If systemd restarts the service, the cold cache forces every query through recursion, creating a warming storm that spikes upstream load. If the memory condition persists, the cycle repeats: start, grow, OOM, kill, restart.&lt;/p></description></item><item><title>BIND named RSS climbing: cache growth, allocator fragmentation, and real leaks</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-memory-growth-rss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-memory-growth-rss/</guid><description>&lt;h1 id="bind-named-rss-climbing-cache-growth-allocator-fragmentation-and-real-leaks">BIND named RSS climbing: cache growth, allocator fragmentation, and real leaks&lt;/h1>
&lt;p>&lt;code>named&lt;/code> RSS trending upward after warmup is the leading indicator before an OOM kill. The hard part is distinguishing three conditions that look identical on an RSS chart: cache-driven growth that will plateau under &lt;code>max-cache-size&lt;/code>, allocator fragmentation that inflates RSS without a real leak, and genuine unbounded growth from a bug or misconfiguration.&lt;/p>
&lt;p>RSS that climbs during a traffic burst or cache warmup and stays high is normal allocator behavior. BIND&amp;rsquo;s allocator (jemalloc on most builds) holds freed blocks for reuse rather than returning them to the OS via &lt;code>munmap&lt;/code>. The operational question is not &amp;ldquo;is RSS high?&amp;rdquo; but &amp;ldquo;is RSS still growing, and how much runway remains?&amp;rdquo;&lt;/p></description></item><item><title>BIND NOTIFY not propagating: secondaries not refreshing when the primary changes</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-notify-not-received/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-notify-not-received/</guid><description>&lt;h1 id="bind-notify-not-propagating-secondaries-not-refreshing-when-the-primary-changes">BIND NOTIFY not propagating: secondaries not refreshing when the primary changes&lt;/h1>
&lt;p>You updated a zone on the primary, reloaded with &lt;code>rndc reload&lt;/code>, and confirmed the primary is serving the new serial. The secondaries still show the old one. Queries against them return stale data for minutes or hours. The zone is not broken, just delayed.&lt;/p>
&lt;p>This is NOTIFY not reaching the secondaries. NOTIFY (RFC 1996) is a UDP push from the primary telling secondaries to check for a new SOA serial immediately, rather than waiting for the refresh timer. When NOTIFY is lost, blocked, or misconfigured, the secondary learns about changes only when its SOA refresh timer fires (commonly 1 hour for typical SOA values). The data looks delayed, not broken. The refresh timer eventually catches up, which is why this rarely pages, but it can mask more serious transfer problems if the refresh timer is the only thing keeping secondaries current.&lt;/p></description></item><item><title>BIND NXDOMAIN spike: DGA malware, water torture, and Windows suffix search lists</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-nxdomain-spike/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-nxdomain-spike/</guid><description>&lt;h1 id="bind-nxdomain-spike-dga-malware-water-torture-and-windows-suffix-search-lists">BIND NXDOMAIN spike: DGA malware, water torture, and Windows suffix search lists&lt;/h1>
&lt;p>The &lt;code>QryNXDOMAIN&lt;/code> counter jumping above 3x baseline is a common alert trigger, but the raw rate tells you almost nothing. Two resolvers with identical NXDOMAIN rates can be in completely different states: one healthy, one under active attack.&lt;/p>
&lt;p>The signal that matters is query-name cardinality and entropy. If the same names repeat, the spike is benign (Windows suffix search lists, new client rollouts). If nearly every query name is unique and random, you are looking at a water torture attack or DGA malware beaconing. This distinction determines whether you page someone at 3 a.m. or close the alert.&lt;/p></description></item><item><title>BIND open recursive resolver: DNS amplification abuse and allow-recursion posture</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-open-resolver-amplification/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-open-resolver-amplification/</guid><description>&lt;h1 id="bind-open-recursive-resolver-dns-amplification-abuse-and-allow-recursion-posture">BIND open recursive resolver: DNS amplification abuse and allow-recursion posture&lt;/h1>
&lt;p>An open recursive resolver answers recursive DNS queries from any source address. BIND configured with &lt;code>allow-recursion { any; }&lt;/code> serves legitimate clients, but it also serves attackers: spoofed source IPs turn your resolver into a DDoS amplification relay, where small queries produce large responses directed at victims you have never heard of.&lt;/p>
&lt;p>The abuse can be invisible. If the attack volume is small relative to your legitimate traffic, query rates look normal, cache hit ratios look healthy, and SERVFAIL rates stay flat. The only evidence is in the configuration itself and in subtle traffic pattern shifts: elevated ANY or TXT query shares, high source-IP diversity, or responses going to networks that have no business querying your resolver.&lt;/p></description></item><item><title>BIND query logging in production: the gradual performance bottleneck teams forget to turn off</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-query-logging-performance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-query-logging-performance/</guid><description>&lt;h1 id="bind-query-logging-in-production-the-gradual-performance-bottleneck-teams-forget-to-turn-off">BIND query logging in production: the gradual performance bottleneck teams forget to turn off&lt;/h1>
&lt;p>BIND query logging is a debugging tool, not a monitoring strategy. Teams enable it during an incident or security investigation, then forget to disable it. The result is a performance degradation so gradual that it gets attributed to traffic growth, hardware aging, or &amp;ldquo;BIND being slow.&amp;rdquo; By the time someone connects the dots, the resolver has been losing throughput for weeks.&lt;/p></description></item><item><title>BIND random subdomain (water torture) attack: an NXDOMAIN flood that bypasses the cache</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-random-subdomain-attack/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-random-subdomain-attack/</guid><description>&lt;h1 id="bind-random-subdomain-water-torture-attack-an-nxdomain-flood-that-bypasses-the-cache">BIND random subdomain (water torture) attack: an NXDOMAIN flood that bypasses the cache&lt;/h1>
&lt;p>Your BIND resolver is up, the process is running, UDP and TCP both answer on port 53. But clients across the network are experiencing slow DNS or outright SERVFAIL. NXDOMAIN rate has spiked to several times baseline. Cache hit ratio is collapsing. Recursive clients are climbing toward the hard limit. Upstream query rate has ballooned to approach or exceed the inbound rate.&lt;/p></description></item><item><title>BIND RecursClients climbing toward the limit: reading the recursive saturation gauge</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-recursive-clients-climbing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-recursive-clients-climbing/</guid><description>&lt;h1 id="bind-recursclients-climbing-toward-the-limit-reading-the-recursive-saturation-gauge">BIND RecursClients climbing toward the limit: reading the recursive saturation gauge&lt;/h1>
&lt;p>RecursClients is the number of queries currently in flight, each waiting for an upstream nameserver to respond. When this gauge climbs toward the configured &lt;code>recursive-clients&lt;/code> limit, BIND is running out of slots to start new recursive lookups. Past the hard limit, every new recursive query receives SERVFAIL.&lt;/p>
&lt;p>The metric is a gauge, not a rate. The absolute number means little without the configured ceiling next to it: 300 is comfortable against a limit of 1000 and dangerous against a limit of 350. Track RecursClients as a percentage of &lt;code>recursive-clients&lt;/code>, not as a raw count, and watch the daily peak for runway estimation.&lt;/p></description></item><item><title>BIND recursive resolution cascade: one slow upstream taking down all resolution</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-recursion-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-recursion-cascade/</guid><description>&lt;h1 id="bind-recursive-resolution-cascade-one-slow-upstream-taking-down-all-resolution">BIND recursive resolution cascade: one slow upstream taking down all resolution&lt;/h1>
&lt;p>SERVFAIL responses are flooding your recursive resolver. Users cannot resolve dozens of unrelated domains. But &lt;code>named&lt;/code> is running, CPU looks moderate, and the authoritative zones on the same instance are still answering fine. This is the recursive resolution cascade: a single slow or unreachable upstream authoritative server consumes all available &lt;code>recursive-clients&lt;/code> slots, and BIND returns SERVFAIL for queries that have nothing to do with the failing upstream.&lt;/p></description></item><item><title>BIND REFUSED responses: ACL denials, recursion policy, and clients that get locked out</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-refused/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-refused/</guid><description>&lt;h1 id="bind-refused-responses-acl-denials-recursion-policy-and-clients-that-get-locked-out">BIND REFUSED responses: ACL denials, recursion policy, and clients that get locked out&lt;/h1>
&lt;p>REFUSED (DNS rcode 5) is a deliberate policy decision: BIND received the query, parsed it, and chose not to answer. The server is alive, listening, and processing queries. It is configured to reject this particular query from this particular source.&lt;/p>
&lt;p>The actionable scenario: legitimate client subnets that previously resolved names or queried your zones suddenly start receiving REFUSED, typically after a configuration change. The fix is almost always an ACL or view mismatch, not a restart.&lt;/p></description></item><item><title>BIND resolver NumFetch per view: per-view recursive pressure in split-horizon setups</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-numfetch-per-view/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-numfetch-per-view/</guid><description>&lt;h1 id="bind-resolver-numfetch-per-view-per-view-recursive-pressure-in-split-horizon-setups">BIND resolver NumFetch per view: per-view recursive pressure in split-horizon setups&lt;/h1>
&lt;p>RecursClients is a global NSStats counter that reports the total number of recursive clients awaiting resolution across all views. It does not break down which view is consuming the pool. When RecursClients climbs to 847 out of 1000, you know the pool is stressed but not whether the pressure is evenly distributed or whether one view is responsible for most of it.&lt;/p></description></item><item><title>BIND resolver RTT distribution shifting high: upstream nameserver degradation</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-upstream-rtt-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-upstream-rtt-high/</guid><description>&lt;h1 id="bind-resolver-rtt-distribution-shifting-high-upstream-nameserver-degradation">BIND resolver RTT distribution shifting high: upstream nameserver degradation&lt;/h1>
&lt;p>The QryRTT histogram in BIND&amp;rsquo;s per-view resolver statistics is shifting toward higher millisecond buckets. What was once dominated by RTT10 and RTT100 (sub-100ms responses from upstream authoritative servers) is now accumulating in RTT500, RTT800, and RTT1600. In severe cases, most outbound recursive queries land in the overflow RTT1600+ bucket.&lt;/p>
&lt;p>This metric is not client-perceived latency. The QryRTT counters measure the round-trip time of BIND&amp;rsquo;s outbound recursive queries to upstream authoritative nameservers. A rightward shift means the servers BIND depends on for cache misses are taking longer to respond. BIND has no native inbound client latency histogram; if you need end-to-end latency from the client perspective, use dnstap or external measurement.&lt;/p></description></item><item><title>BIND Response Rate Limiting (RRL): RateDropped, RateSlipped, and throttled legitimate clients</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-rrl-rate-limiting/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-rrl-rate-limiting/</guid><description>&lt;h1 id="bind-response-rate-limiting-rrl-ratedropped-rateslipped-and-throttled-legitimate-clients">BIND Response Rate Limiting (RRL): RateDropped, RateSlipped, and throttled legitimate clients&lt;/h1>
&lt;p>You see sustained non-zero &lt;code>RateDropped&lt;/code> or &lt;code>RateSlipped&lt;/code> counters in BIND&amp;rsquo;s statistics channel. Either RRL is absorbing a real DNS amplification or flood attack, or the configuration is too aggressive and silently dropping or truncating responses to legitimate clients. BIND&amp;rsquo;s RRL counters do not distinguish attacker from legitimate client. A dropped response is a dropped response, whether the source was a spoofed botnet node or a real resolver. Telling the difference requires correlating the counters with traffic patterns, source IP distribution, and TCP reachability.&lt;/p></description></item><item><title>BIND RPZ rewrites: threat-interception counters and reading a malware outbreak</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-rpz-rewrites/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-rpz-rewrites/</guid><description>&lt;h1 id="bind-rpz-rewrites-threat-interception-counters-and-reading-a-malware-outbreak">BIND RPZ rewrites: threat-interception counters and reading a malware outbreak&lt;/h1>
&lt;p>Response Policy Zones (RPZ) let BIND rewrite DNS responses that match policy rules: blocked, dropped, redirected, or passed through with an exception marker. The &lt;code>RPZRewrites&lt;/code> counter in BIND&amp;rsquo;s Name Server Statistics tracks how often this happens. When that counter spikes, it is usually the first telemetry signal that something inside your network is reaching known-bad infrastructure.&lt;/p>
&lt;p>This article covers what &lt;code>RPZRewrites&lt;/code> actually counts, what it silently excludes, how to distinguish a real malware outbreak from background noise, and what RPZ costs in query performance and memory. It assumes RPZ is already configured. For the broader BIND monitoring framework, see &lt;a href="https://www.netdata.cloud/guides/bind-dns/bind-dns-how-it-works-in-production/">How BIND actually works in production: a mental model for operators&lt;/a>.&lt;/p></description></item><item><title>BIND RRSIG expiry on authoritative zones: the silent signing time bomb</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-rrsig-expiry-authoritative/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-rrsig-expiry-authoritative/</guid><description>&lt;h1 id="bind-rrsig-expiry-on-authoritative-zones-the-silent-signing-time-bomb">BIND RRSIG expiry on authoritative zones: the silent signing time bomb&lt;/h1>
&lt;p>Your zone is down. Validating resolvers worldwide return SERVFAIL for every query to your domain. But the authoritative server looks fine: &lt;code>named&lt;/code> is running, port 53 is listening, the statistics channel shows normal traffic. No errors in the logs. No alerts fired.&lt;/p>
&lt;p>The problem is expired RRSIG signatures. BIND&amp;rsquo;s authoritative server does not validate its own signatures at serve time. It will serve expired RRSIGs indefinitely, and every validating resolver that queries your zone will reject the response. The outage is invisible from the server itself.&lt;/p></description></item><item><title>BIND secondary zone expired: the SOA expire timer runs out and the zone returns SERVFAIL</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-secondary-zone-expired/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-secondary-zone-expired/</guid><description>&lt;h1 id="bind-secondary-zone-expired-the-soa-expire-timer-runs-out-and-the-zone-returns-servfail">BIND secondary zone expired: the SOA expire timer runs out and the zone returns SERVFAIL&lt;/h1>
&lt;p>A single zone on your BIND secondary starts returning SERVFAIL. Other zones on the same server answer normally. The &lt;code>named&lt;/code> process is up, CPU and memory look fine, and port 53 responds to health checks. Your monitoring shows green because it checks process liveness or queries a different zone. Only clients asking for that one zone are failing.&lt;/p></description></item><item><title>BIND SERVFAIL responses: what a DNS SERVFAIL actually means and how to trace the cause</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-servfail/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-servfail/</guid><description>&lt;h1 id="bind-servfail-responses-what-a-dns-servfail-actually-means-and-how-to-trace-the-cause">BIND SERVFAIL responses: what a DNS SERVFAIL actually means and how to trace the cause&lt;/h1>
&lt;p>SERVFAIL (RCODE 2) means BIND attempted resolution and could not return a usable answer. It is not NXDOMAIN (name does not exist), REFUSED (policy denial), or FORMERR (malformed query). The process is alive, port 53 is open, the query was received, and the answer is failure.&lt;/p>
&lt;p>This makes SERVFAIL invisible to binary health checks. A BIND server returning 100% SERVFAIL to every client still passes process-liveness and port-check probes. SERVFAIL is a symptom with many possible causes: upstream nameserver timeouts, DNSSEC validation failures, recursive-clients exhaustion, broken delegation, a zone that failed to load, or a configuration error. Tracing it requires correlating multiple BIND signals.&lt;/p></description></item><item><title>BIND SOA expire runway: the countdown that predicts a silent secondary-zone outage</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-soa-expire-runway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-soa-expire-runway/</guid><description>&lt;h1 id="bind-soa-expire-runway-the-countdown-that-predicts-a-silent-secondary-zone-outage">BIND SOA expire runway: the countdown that predicts a silent secondary-zone outage&lt;/h1>
&lt;p>A BIND secondary serving stale data is not broken. It answers queries with old but functional records, and from the outside it looks healthy. The danger is the countdown underneath: the SOA expire timer, ticking from the last successful zone transfer. When it reaches zero, the secondary stops serving the zone and returns SERVFAIL or REFUSED for every query against it.&lt;/p></description></item><item><title>BIND SOA serial mismatch: a secondary serving stale data behind the primary</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-soa-serial-mismatch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-soa-serial-mismatch/</guid><description>&lt;h1 id="bind-soa-serial-mismatch-a-secondary-serving-stale-data-behind-the-primary">BIND SOA serial mismatch: a secondary serving stale data behind the primary&lt;/h1>
&lt;p>The secondary&amp;rsquo;s SOA serial is behind the primary. Clients querying the secondary get stale A values, missing new entries, deleted hosts still resolving. The server responds NOERROR because its zone data is internally consistent. It is not current.&lt;/p>
&lt;p>This is the precursor to zone expiry, and it is silent. The secondary serves stale but functional answers for the entire SOA expire period, typically 1 to 4 weeks. Monitoring that checks &amp;ldquo;can I resolve this zone?&amp;rdquo; passes. Health checks on port 53 pass. Everything looks fine except the data is wrong. The only signal is the serial number gap between primary and secondary, and BIND&amp;rsquo;s statistics channel does not expose it. You must probe externally.&lt;/p></description></item><item><title>BIND tcp-clients exhaustion: the TCP connection limit, transfers, and truncation fallback</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-tcp-clients-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-tcp-clients-exhaustion/</guid><description>&lt;h1 id="bind-tcp-clients-exhaustion-the-tcp-connection-limit-transfers-and-truncation-fallback">BIND tcp-clients exhaustion: the TCP connection limit, transfers, and truncation fallback&lt;/h1>
&lt;p>BIND is up. &lt;code>rndc status&lt;/code> shows &amp;ldquo;running.&amp;rdquo; Your UDP health check returns NOERROR in under 2ms. But zone transfers are failing, some DNSSEC-validated queries time out, and clients receiving large responses report intermittent connection refused on TCP/53. If your monitoring only probes UDP, you will not know anything is wrong until a secondary&amp;rsquo;s zone expires or a downstream resolver escalates a ticket.&lt;/p></description></item><item><title>BIND TSIG failure on zone transfer: BADKEY, BADTIME, and refused transfers</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-tsig-badkey-transfer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-tsig-badkey-transfer/</guid><description>&lt;h1 id="bind-tsig-failure-on-zone-transfer-badkey-badtime-and-refused-transfers">BIND TSIG failure on zone transfer: BADKEY, BADTIME, and refused transfers&lt;/h1>
&lt;p>Zone transfers between your BIND primary and secondary have stopped. The secondary is falling behind on SOA serial, and the logs show TSIG verification failures: BADKEY, BADTIME, or BADSIG. The transfer is refused, and the secondary silently drifts toward zone expiry.&lt;/p>
&lt;p>The failure hides in the &lt;code>security&lt;/code> and &lt;code>xfer-in&lt;/code> logging categories, not in the main query path. A secondary can serve stale data for days or weeks until the SOA expire timer runs out, at which point it stops serving the zone entirely and returns SERVFAIL. Nothing alerts until the zone disappears.&lt;/p></description></item><item><title>BIND UDP packet-rate saturation: softirq, single-core bottlenecks, and pps limits</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-packet-rate-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-packet-rate-saturation/</guid><description>&lt;p>DNS queries are timing out for some clients. BIND statistics look clean: query rate is steady, SERVFAIL is low, cache hit ratio is normal. The link is not saturated. BIND logs show nothing unusual. But clients keep reporting intermittent failures with no apparent correlation to domain, client subnet, or time of day.&lt;/p>
&lt;p>If you have ruled out upstream issues, cache pressure, and DNSSEC failures, the problem may be in the kernel. When the host cannot process UDP packets fast enough, the kernel drops them before BIND reads them from the socket. BIND has no visibility into these drops: no log entry, no statistics counter, no error. The only evidence lives in kernel-level UDP counters that most DNS monitoring setups never collect.&lt;/p></description></item><item><title>BIND UDP receive buffer errors: the invisible query loss named never logs</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-udp-receive-buffer-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-udp-receive-buffer-errors/</guid><description>&lt;h1 id="bind-udp-receive-buffer-errors-the-invisible-query-loss-named-never-logs">BIND UDP receive buffer errors: the invisible query loss named never logs&lt;/h1>
&lt;p>Clients report intermittent DNS timeouts. Retries are up. Your monitoring says &lt;code>named&lt;/code> is healthy: the process is running, port 53 is open, functional queries succeed, and BIND&amp;rsquo;s statistics channel shows reasonable response rates with no elevated SERVFAIL. You suspect upstream nameserver problems, network path issues, or client-side misconfiguration. None of those investigations turn up anything.&lt;/p>
&lt;p>The problem may be happening between the kernel and BIND, in a layer where &lt;code>named&lt;/code> has zero visibility. When the kernel&amp;rsquo;s UDP receive buffer overflows, packets are silently dropped before BIND ever reads them from the socket. There is no log entry, no statistics counter increment, no error of any kind inside &lt;code>named&lt;/code>. The only evidence lives in kernel-level counters that most BIND monitoring setups never collect.&lt;/p></description></item><item><title>BIND UDP vs TCP query ratio: truncation, EDNS negotiation, and TCP fallback spikes</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-udp-tcp-ratio/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-udp-tcp-ratio/</guid><description>&lt;h1 id="bind-udp-vs-tcp-query-ratio-truncation-edns-negotiation-and-tcp-fallback-spikes">BIND UDP vs TCP query ratio: truncation, EDNS negotiation, and TCP fallback spikes&lt;/h1>
&lt;p>A rising TCP share of DNS queries is one of the most informative early warning signals in BIND. TCP normally carries under 5% of query volume on a recursive resolver. When that ratio shifts upward, something has changed in the resolution path: responses are being truncated, EDNS negotiation is failing, zone transfers are spiking, or a firewall is interfering with DNS traffic.&lt;/p></description></item><item><title>BIND unauthorized zone transfer attempts: AXFR/IXFR from sources not on allow-transfer</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-unauthorized-zone-transfer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-unauthorized-zone-transfer/</guid><description>&lt;h1 id="bind-unauthorized-zone-transfer-attempts-axfrixfr-from-sources-not-on-allow-transfer">BIND unauthorized zone transfer attempts: AXFR/IXFR from sources not on allow-transfer&lt;/h1>
&lt;p>Your BIND logs show denied AXFR or IXFR requests from unknown source IPs. The entries appear under the &lt;code>security&lt;/code> category at &lt;code>error&lt;/code> severity:&lt;/p>
&lt;pre tabindex="0">&lt;code>client 198.51.100.42#53124: zone transfer &amp;#39;example.com/AXFR/IN&amp;#39; denied
&lt;/code>&lt;/pre>&lt;p>Denied attempts mean your ACL is working. The critical question is whether any unauthorized transfer succeeded, because a successful AXFR exfiltrates the entire zone: every record, internal hostnames, SRV targets, TXT metadata, and the full infrastructure topology.&lt;/p></description></item><item><title>BIND upstream query timeouts: QueryTimeout, retries, and the 30-second resolution stall</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-upstream-timeouts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-upstream-timeouts/</guid><description>&lt;h1 id="bind-upstream-query-timeouts-querytimeout-retries-and-the-30-second-resolution-stall">BIND upstream query timeouts: QueryTimeout, retries, and the 30-second resolution stall&lt;/h1>
&lt;p>When a BIND resolver sends a recursive query upstream and gets no response, it waits. The default wait is 10 seconds, and BIND may retry the query up to 3 times before giving up. A single failed resolution can occupy a recursive-client slot for 30 seconds or more. If the upstream failure is broad enough, those slots fill, the recursive-clients limit is reached, and every new recursive query starts returning SERVFAIL.&lt;/p></description></item><item><title>BIND zone not loaded after reload: 'loading from master file failed' and the silent zone outage</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-zone-not-loaded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-zone-not-loaded/</guid><description>&lt;h1 id="bind-zone-not-loaded-after-reload-loading-from-master-file-failed-and-the-silent-zone-outage">BIND zone not loaded after reload: &amp;rsquo;loading from master file failed&amp;rsquo; and the silent zone outage&lt;/h1>
&lt;p>You ran &lt;code>rndc reload&lt;/code> after editing a zone file. &lt;code>rndc status&lt;/code> reports the server is running. systemd says the service is active. Your monitoring confirms the process is alive and port 53 is open. But something is wrong with one specific zone.&lt;/p>
&lt;p>&lt;code>named&lt;/code> does not crash when a single zone fails to load. It logs the error and continues serving the zones that loaded successfully. Two outcomes are possible for the affected zone:&lt;/p></description></item><item><title>BIND zone transfer failed (AXFR/IXFR): reading xfer-in failures before the zone goes stale</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-zone-transfer-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-zone-transfer-failed/</guid><description>&lt;h1 id="bind-zone-transfer-failed-axfrixfr-reading-xfer-in-failures-before-the-zone-goes-stale">BIND zone transfer failed (AXFR/IXFR): reading xfer-in failures before the zone goes stale&lt;/h1>
&lt;p>A secondary BIND server&amp;rsquo;s zone transfers are failing. The zone is still being served, clients are still getting answers, and your health checks are still green. The secondary is serving stale data from the last successful transfer, and the SOA expire timer is counting down. When it reaches zero, the secondary stops serving the zone and returns SERVFAIL or REFUSED for every query.&lt;/p></description></item><item><title>Bintec Communications GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bintec-communications-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bintec-communications-gmbh-snmp-traps/</guid><description/></item><item><title>bio-rd / RIPE RIS</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/bio-rd---ripe-ris/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/bio-rd---ripe-ris/</guid><description/></item><item><title>Bird Routing Daemon</title><link>https://www.netdata.cloud/integrations/data-collection/networking/bird-routing-daemon/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/bird-routing-daemon/</guid><description/></item><item><title>Bird Routing Daemon Monitoring</title><link>https://www.netdata.cloud/monitoring-101/bird-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/bird-monitoring/</guid><description>&lt;h2 id="bird-routing-daemon-monitoring">Bird Routing Daemon Monitoring&lt;/h2>
&lt;h3 id="what-is-bird-routing-daemon">What Is Bird Routing Daemon?&lt;/h3>
&lt;p>The Bird Routing Daemon is a comprehensive software for managing BGP and other routing protocols. It&amp;rsquo;s favored in the networking community for its scalability and flexibility. With Bird, administrators can maintain robust and efficient network routing essential for data flow across various network architectures.&lt;/p>
&lt;h3 id="monitoring-bird-routing-daemon-with-netdata">Monitoring Bird Routing Daemon With Netdata&lt;/h3>
&lt;p>Netdata provides an efficient Bird Routing Daemon monitoring tool that leverages an openmetrics (Prometheus) exporter for data collection. This integration allows seamless ingestion of Bird Routing Daemon metrics without needing an additional Prometheus server or Grafana setup. Netdata offers automated dashboards and alerts out-of-the-box, ensuring you have real-time insight into your network&amp;rsquo;s performance and reliability.&lt;/p></description></item><item><title>Bird Technologies SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bird-technologies-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bird-technologies-snmp-traps/</guid><description/></item><item><title>Bke A S SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bke-a-s-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bke-a-s-snmp-traps/</guid><description/></item><item><title>Blackbox</title><link>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/blackbox/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/blackbox/</guid><description/></item><item><title>Blackbox Monitoring</title><link>https://www.netdata.cloud/monitoring-101/blackbox-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/blackbox-monitoring/</guid><description>&lt;h2 id="blackbox-monitoring">Blackbox Monitoring&lt;/h2>
&lt;h3 id="what-is-blackbox">What Is Blackbox?&lt;/h3>
&lt;p>Blackbox is a versatile monitoring tool designed to test endpoints via HTTP, DNS, TCP and ICMP protocols. As a member of the Prometheus ecosystem, Blackbox enables the continuous evaluation of uptime and response times, providing critical insights into external service availability. By simulating requests to your key infrastructure components, Blackbox ensures you are the first to know when issues occur.&lt;/p>
&lt;h3 id="monitoring-blackbox-with-netdata">Monitoring Blackbox With Netdata&lt;/h3>
&lt;p>To effectively monitor Blackbox, Netdata integrates with the Prometheus ecosystem via an openmetrics (Prometheus) exporter. This means you can leverage the &lt;a href="https://github.com/prometheus/blackbox_exporter">Blackbox exporter&lt;/a> to gather metrics seamlessly. With Netdata, you can ingest data from any Prometheus exporter, streamlining your monitoring setup without needing a full Prometheus server or Grafana dashboard stack. Netdata automatically provides dashboards and alerts, ensuring you have real-time insights into your monitored services.&lt;/p></description></item><item><title>Blade Network Technologies Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/blade-network-technologies-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/blade-network-technologies-inc-snmp-traps/</guid><description/></item><item><title>blk_update_request: I/O error, dev nvme0n1: reading NVMe I/O errors in the kernel log</title><link>https://www.netdata.cloud/guides/nvme/nvme-blk-update-request-io-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-blk-update-request-io-error/</guid><description>&lt;h1 id="blk_update_request-io-error-dev-nvme0n1-reading-nvme-io-errors-in-the-kernel-log">blk_update_request: I/O error, dev nvme0n1: reading NVMe I/O errors in the kernel log&lt;/h1>
&lt;p>&lt;code>dmesg&lt;/code> shows lines like this:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-text" data-lang="text">&lt;span style="display:flex;">&lt;span>blk_update_request: I/O error, dev nvme0n1, sector 12345678
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Maybe one line. Maybe thousands. Maybe interleaved with &lt;code>nvme nvme0: Resetting controller&lt;/code> or &lt;code>nvme nvme0: Removing&lt;/code>. The operational question is always the same: is this a single bad block the drive surfaced on a read, or is the device, the controller, or the PCIe link underneath it failing?&lt;/p></description></item><item><title>Blue Coat Licensing</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/licensing-monitoring/blue-coat-licensing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/licensing-monitoring/blue-coat-licensing/</guid><description/></item><item><title>Blue Coat Systems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/blue-coat-systems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/blue-coat-systems-snmp-traps/</guid><description/></item><item><title>Bluecat Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bluecat-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bluecat-networks-snmp-traps/</guid><description/></item><item><title>Bluecat Server</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/bluecat-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/bluecat-server/</guid><description/></item><item><title>Bluecoat Proxysg</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/bluecoat-proxysg/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/bluecoat-proxysg/</guid><description/></item><item><title>Blueflood</title><link>https://www.netdata.cloud/integrations/exporters/blueflood/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/blueflood/</guid><description/></item><item><title>Bluesocket Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bluesocket-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bluesocket-inc-snmp-traps/</guid><description/></item><item><title>Bmc Software SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bmc-software-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bmc-software-snmp-traps/</guid><description/></item><item><title>BMP (BGP Monitoring Protocol)</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/bmp-bgp-monitoring-protocol/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/bmp-bgp-monitoring-protocol/</guid><description/></item><item><title>BOINC</title><link>https://www.netdata.cloud/integrations/data-collection/applications/boinc/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/boinc/</guid><description/></item><item><title>BOINC Monitoring</title><link>https://www.netdata.cloud/monitoring-101/boinc-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/boinc-monitoring/</guid><description>&lt;h2 id="boinc-monitoring">BOINC Monitoring&lt;/h2>
&lt;h3 id="what-is-boinc">What Is BOINC?&lt;/h3>
&lt;p>The Berkeley Open Infrastructure for Network Computing (BOINC) is a platform for volunteer computing and is used by various scientific projects. It allows users to donate their computing resources to assist in complex calculations and data analysis. Understanding the underlying metrics of BOINC is crucial for maintaining optimal performance and ensuring that resources are effectively utilized.&lt;/p>
&lt;h3 id="monitoring-boinc-with-netdata">Monitoring BOINC With Netdata&lt;/h3>
&lt;p>Netdata provides an intuitive and powerful way to monitor your BOINC instances. Our BOINC monitoring tool collects real-time data and allows you to visualize important metrics to identify trends and potential issues. With Netdata, you can gain insights into task counts and manage distributed computing efforts more effectively.&lt;/p></description></item><item><title>Borderware Technologies Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/borderware-technologies-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/borderware-technologies-inc-snmp-traps/</guid><description/></item><item><title>BOSH</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/bosh/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/bosh/</guid><description/></item><item><title>BOSH Monitoring</title><link>https://www.netdata.cloud/monitoring-101/bosh-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/bosh-monitoring/</guid><description>&lt;h2 id="bosh-monitoring">BOSH Monitoring&lt;/h2>
&lt;h3 id="what-is-bosh">What Is BOSH?&lt;/h3>
&lt;p>BOSH is an open-source tool used for release engineering, deployment, lifecycle management, and monitoring of distributed systems. As a pivotal component in cloud orchestration, BOSH simplifies the deployment of cloud infrastructure, ensuring scalability and resilience. It&amp;rsquo;s an essential tool for DevOps, SRE, developers, IT admins, and IT engineers who manage complex application environments.&lt;/p>
&lt;h3 id="monitoring-bosh-with-netdata">Monitoring BOSH With Netdata&lt;/h3>
&lt;p>To monitor BOSH effectively, Netdata employs a powerful openmetrics (Prometheus) exporter called the &lt;a href="https://github.com/bosh-prometheus/bosh_exporter">BOSH exporter&lt;/a>. This integration allows Netdata to ingest data directly from any Prometheus exporter. With Netdata, you get automated dashboards and real-time alerts, all without needing a separate Prometheus server or Grafana setup. This makes Netdata a highly efficient BOSH monitoring tool, allowing for a seamless and comprehensive monitoring experience.&lt;/p></description></item><item><title>Brand Communications Limited SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/brand-communications-limited-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/brand-communications-limited-snmp-traps/</guid><description/></item><item><title>Bridgewave Communications SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bridgewave-communications-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bridgewave-communications-snmp-traps/</guid><description/></item><item><title>Broadband Access Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/broadband-access-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/broadband-access-systems-inc-snmp-traps/</guid><description/></item><item><title>Broadcom Limited SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/broadcom-limited-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/broadcom-limited-snmp-traps/</guid><description/></item><item><title>Broadsoft Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/broadsoft-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/broadsoft-inc-snmp-traps/</guid><description/></item><item><title>Brocade</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/brocade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/brocade/</guid><description/></item><item><title>Brocade Communication Systems Inc Formerly Foundry Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/brocade-communication-systems-inc-formerly-foundry-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/brocade-communication-systems-inc-formerly-foundry-networks-inc-snmp-traps/</guid><description/></item><item><title>Brocade Communications Systems Inc Formerly Mcdata Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/brocade-communications-systems-inc-formerly-mcdata-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/brocade-communications-systems-inc-formerly-mcdata-corporation-snmp-traps/</guid><description/></item><item><title>Brocade Communications Systems Inc Formerly Nuview Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/brocade-communications-systems-inc-formerly-nuview-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/brocade-communications-systems-inc-formerly-nuview-inc-snmp-traps/</guid><description/></item><item><title>Brocade Communications Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/brocade-communications-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/brocade-communications-systems-inc-snmp-traps/</guid><description/></item><item><title>Brocade FC Switch</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/brocade-fc-switch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/brocade-fc-switch/</guid><description/></item><item><title>Brother</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/brother/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/brother/</guid><description/></item><item><title>Brother NET Printer</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/brother-net-printer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/brother-net-printer/</guid><description/></item><item><title>Bti Photonic Systems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bti-photonic-systems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bti-photonic-systems-snmp-traps/</guid><description/></item><item><title>BTRFS</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/btrfs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/btrfs/</guid><description/></item><item><title>BungeeCord</title><link>https://www.netdata.cloud/integrations/data-collection/applications/bungeecord/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/bungeecord/</guid><description/></item><item><title>BungeeCord Monitoring</title><link>https://www.netdata.cloud/monitoring-101/bungeecord-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/bungeecord-monitoring/</guid><description>&lt;h2 id="bungeecord-monitoring">BungeeCord Monitoring&lt;/h2>
&lt;h3 id="what-is-bungeecord">What Is BungeeCord?&lt;/h3>
&lt;p>BungeeCord is a popular proxy server for the Minecraft game, used to connect multiple servers together for seamless gameplay. It allows players to switch between different Minecraft servers without needing to relog, making it an essential tool for managing complex Minecraft networks.&lt;/p>
&lt;h3 id="monitoring-bungeecord-with-netdata">Monitoring BungeeCord With Netdata&lt;/h3>
&lt;p>To monitor BungeeCord, Netdata employs an openmetrics (Prometheus) exporter. This approach leverages the &lt;a href="https://github.com/weihao/bungeecord-prometheus-exporter">BungeeCord Prometheus Exporter&lt;/a>, allowing Netdata to track important server metrics seamlessly. With Netdata, you can ingest data from any Prometheus exporter, offering automated dashboards, alerting, and more. Remarkably, this doesn&amp;rsquo;t require a Prometheus server or Grafana, making it simple and efficient.&lt;/p></description></item><item><title>Bytesphere LLC SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bytesphere-llc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/bytesphere-llc-snmp-traps/</guid><description/></item><item><title>C C Power Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/c-c-power-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/c-c-power-inc-snmp-traps/</guid><description/></item><item><title>Cable Television Laboratories Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cable-television-laboratories-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cable-television-laboratories-inc-snmp-traps/</guid><description/></item><item><title>Cacheflow Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cacheflow-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cacheflow-inc-snmp-traps/</guid><description/></item><item><title>Cacti SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cacti-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cacti-snmp-traps/</guid><description/></item><item><title>Cadant Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cadant-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cadant-inc-snmp-traps/</guid><description/></item><item><title>CAIDA Routeviews Prefix-to-AS</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/caida-routeviews-prefix-to-as/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/caida-routeviews-prefix-to-as/</guid><description/></item><item><title>Calix Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/calix-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/calix-networks-snmp-traps/</guid><description/></item><item><title>Cambium Networks Limited Formerly Pipinghot Networks Limited SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cambium-networks-limited-formerly-pipinghot-networks-limited-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cambium-networks-limited-formerly-pipinghot-networks-limited-snmp-traps/</guid><description/></item><item><title>Carel SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/carel-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/carel-snmp-traps/</guid><description/></item><item><title>Cascade Communications Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cascade-communications-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cascade-communications-corp-snmp-traps/</guid><description/></item><item><title>Cassandra</title><link>https://www.netdata.cloud/integrations/data-collection/databases/cassandra/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/cassandra/</guid><description/></item><item><title>Cassandra adding and removing nodes safely: vnodes, tokens, and cleanup</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-adding-removing-nodes-safely/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-adding-removing-nodes-safely/</guid><description>&lt;h1 id="cassandra-adding-and-removing-nodes-safely-vnodes-tokens-and-cleanup">Cassandra adding and removing nodes safely: vnodes, tokens, and cleanup&lt;/h1>
&lt;p>Expanding or contracting a Cassandra cluster redistributes token ownership, triggers bulk streaming, and leaves stale data that cleanup must reclaim. Skip cleanup after a bootstrap, run it while a node is joining, or misconfigure token allocation, and you waste disk space, create hot spots, and risk silent inconsistency.&lt;/p>
&lt;p>This guide covers adding and removing nodes with virtual nodes (vnodes) enabled, the default for Cassandra 3.x through 5.x. Commands and paths assume standard packaged Cassandra on Linux. Perform topology changes one node at a time.&lt;/p></description></item><item><title>Cassandra Batch too large warning: oversized batches and coordinator OOM</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-batch-too-large-warning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-batch-too-large-warning/</guid><description>&lt;h1 id="cassandra-batch-too-large-warning-oversized-batches-and-coordinator-oom">Cassandra Batch too large warning: oversized batches and coordinator OOM&lt;/h1>
&lt;p>Oversized &lt;code>BEGIN BATCH&lt;/code> statements cause &lt;code>Batch for [ks.table] is of size N, exceeding specified threshold of M by ...&lt;/code> warnings and coordinator OutOfMemoryError. Unlike single-partition batches, which provide atomicity within one partition, multi-partition batches force the coordinator to hold mutation buffers for every affected partition until all replicas acknowledge. When the buffer grows large enough, it triggers heap pressure, long GC pauses, and eventual OOM.&lt;/p></description></item><item><title>Cassandra clock skew: how NTP drift silently corrupts data</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-clock-skew-data-corruption/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-clock-skew-data-corruption/</guid><description>&lt;h1 id="cassandra-clock-skew-how-ntp-drift-silently-corrupts-data">Cassandra clock skew: how NTP drift silently corrupts data&lt;/h1>
&lt;p>Cassandra&amp;rsquo;s last-write-wins conflict resolution is simple, fast, and unforgiving. Every write carries a timestamp; when replicas disagree, the highest timestamp wins. The database assumes larger timestamps correspond to later wall-clock events. When node clocks drift, that assumption collapses. A write that happened first can carry a later timestamp, or a later write an earlier one. The result is not a timeout, an &lt;code>UnavailableException&lt;/code>, or an &lt;code>ERROR&lt;/code> log entry. It is silent data loss, permanent shadowing of valid writes, or the sudden return of deleted rows.&lt;/p></description></item><item><title>Cassandra commitlog disk full: segment exhaustion and forced flushes</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-commitlog-disk-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-commitlog-disk-full/</guid><description>&lt;h1 id="cassandra-commitlog-disk-full-segment-exhaustion-and-forced-flushes">Cassandra commitlog disk full: segment exhaustion and forced flushes&lt;/h1>
&lt;p>WriteTimeoutException from client drivers, a node still UP in gossip, and a commitlog volume nearing 100% with climbing pending tasks in &lt;code>nodetool info&lt;/code> mean commitlog segment exhaustion. Unlike data disk exhaustion, which slows compaction, commitlog pressure blocks the write path directly: every mutation must be durably appended to the WAL before acknowledgment.&lt;/p>
&lt;p>Cassandra recycles commitlog segments only after all memtables they reference are flushed to SSTables. When the flush pipeline cannot keep pace, segments accumulate until the total size exceeds &lt;code>commitlog_total_space&lt;/code> (or &lt;code>commitlog_total_space_in_mb&lt;/code> on older versions) or the filesystem fills. Cassandra then forces flushes of every dirty column family referenced in the oldest segment to free space. If flushes are already backed up, this cascades into blocked segment allocation, dropped mutations, and depending on &lt;code>commit_failure_policy&lt;/code>, a node that stops accepting writes entirely.&lt;/p></description></item><item><title>Cassandra commitlog pending tasks: write-path I/O pressure</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-commitlog-pending-tasks/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-commitlog-pending-tasks/</guid><description>&lt;h1 id="cassandra-commitlog-pending-tasks-write-path-io-pressure">Cassandra commitlog pending tasks: write-path I/O pressure&lt;/h1>
&lt;p>Sustained non-zero CommitLog PendingTasks means a Cassandra node&amp;rsquo;s write path is backing up. Every write must be appended to the commitlog and synced to disk before the coordinator acknowledges it. When the fsync thread cannot keep up, mutations queue. This starts as elevated write latency; if the queue persists, it forces emergency memtable flushes, overwhelms the flush and compaction pipeline, and produces dropped mutations.&lt;/p></description></item><item><title>Cassandra compaction death spiral: when writes outrun compaction throughput</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-compaction-death-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-compaction-death-spiral/</guid><description>&lt;h1 id="cassandra-compaction-death-spiral-when-writes-outrun-compaction-throughput">Cassandra compaction death spiral: when writes outrun compaction throughput&lt;/h1>
&lt;p>P99 read latency climbs while disk utilisation on the data volume pins near 100%. &lt;code>nodetool compactionstats&lt;/code> shows pending tasks rising hour over hour, and &lt;code>nodetool tablestats&lt;/code> reports a growing SSTable count. Writes stay fast; reads slow down. This is the compaction death spiral: writes exceed compaction throughput, SSTables accumulate, and read amplification rises.&lt;/p>
&lt;p>Unlike a sudden node crash, this failure is gradual. A background queue grows a little each day. Once disk I/O saturates, the cycle self-reinforces: compaction falls further behind, reads consult more files, latency spikes, and the backlog deepens. By the time client SLAs breach, recovery can take hours.&lt;/p></description></item><item><title>Cassandra compaction strategies: STCS vs LCS vs TWCS vs UCS</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-choosing-compaction-strategy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-choosing-compaction-strategy/</guid><description>&lt;h1 id="cassandra-compaction-strategies-stcs-vs-lcs-vs-twcs-vs-ucs">Cassandra compaction strategies: STCS vs LCS vs TWCS vs UCS&lt;/h1>
&lt;p>Compaction merges immutable SSTables, discards tombstones, and reclaims disk space. The strategy assigned to a table controls the tradeoff between write amplification and read amplification, and it determines how much temporary disk headroom you must preserve. A fit strategy keeps SSTable counts low and latency predictable. A mismatch creates compaction debt: creeping P99 read latency first, then disk space exhaustion, and finally write rejections when compaction cannot reclaim space fast enough to keep up with flushes.&lt;/p></description></item><item><title>Cassandra compaction stuck: large partitions blocking a compaction thread</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-stuck-compaction/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-stuck-compaction/</guid><description>&lt;h1 id="cassandra-compaction-stuck-large-partitions-blocking-a-compaction-thread">Cassandra compaction stuck: large partitions blocking a compaction thread&lt;/h1>
&lt;p>&lt;code>nodetool compactionstats&lt;/code> shows a compaction on one table that has not moved past the same byte offset for hours. The progress percentage is frozen, the pending queue behind it is growing, and read latency on that table is creeping up. This is not a slow disk. A single large partition has monopolized a compaction thread, turning background maintenance into a bottleneck that threatens node stability.&lt;/p></description></item><item><title>Cassandra consistency levels explained: QUORUM, ONE, LOCAL_QUORUM, and EACH_QUORUM</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-consistency-levels-explained/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-consistency-levels-explained/</guid><description>&lt;h1 id="cassandra-consistency-levels-explained-quorum-one-local_quorum-and-each_quorum">Cassandra consistency levels explained: QUORUM, ONE, LOCAL_QUORUM, and EACH_QUORUM&lt;/h1>
&lt;p>Consistency level (CL) balances latency, availability, and correctness. It does not control how many replicas store data; replication factor (RF) does. CL controls how many replicas must acknowledge a read or write before the coordinator returns to the client. The wrong choice produces UnavailableException during rolling restarts that should be safe, or leaves replicas inconsistent for hours after a write is acknowledged at CL=ONE. The four CLs that define most production topologies are ONE, QUORUM, LOCAL_QUORUM, and EACH_QUORUM.&lt;/p></description></item><item><title>Cassandra CorruptSSTableException and FSError: disk failure and recovery</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-storage-exceptions-corrupt-sstable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-storage-exceptions-corrupt-sstable/</guid><description>&lt;h1 id="cassandra-corruptsstableexception-and-fserror-disk-failure-and-recovery">Cassandra CorruptSSTableException and FSError: disk failure and recovery&lt;/h1>
&lt;p>A Cassandra node stops serving traffic or refuses to start. In &lt;code>system.log&lt;/code> you see &lt;code>org.apache.cassandra.io.sstable.CorruptSSTableException&lt;/code>, &lt;code>FSError&lt;/code>, or a JVM shutdown triggered by a filesystem exception. These indicate disk failure, filesystem corruption, or irreversible SSTable damage, not retryable application bugs.&lt;/p>
&lt;p>Because SSTables are immutable, a corrupt file cannot be patched. The node either stops serving the data or shuts down, depending on &lt;code>disk_failure_policy&lt;/code>. Recovery requires at least one healthy replica. Without that, corruption is data loss.&lt;/p></description></item><item><title>Cassandra disk space exhaustion: emergency recovery when the data volume fills</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-disk-space-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-disk-space-exhaustion/</guid><description>&lt;h1 id="cassandra-disk-space-exhaustion-emergency-recovery-when-the-data-volume-fills">Cassandra disk space exhaustion: emergency recovery when the data volume fills&lt;/h1>
&lt;p>A Cassandra node that runs out of disk space does not degrade gracefully. Compaction halts because it cannot allocate temporary space to merge SSTables. Old SSTables are never deleted. Writes append to the commitlog until segment allocation blocks. At that point the node rejects mutations and &lt;code>CommitLog.WaitingOnSegmentAllocation&lt;/code> climbs. You may see &lt;code>No space left on device&lt;/code> errors while the data volume still reports a few percent free, because Cassandra&amp;rsquo;s internal headroom requirements are stricter than the filesystem.&lt;/p></description></item><item><title>Cassandra dropped mutations: silent write loss and load shedding</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-dropped-mutations/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-dropped-mutations/</guid><description>&lt;h1 id="cassandra-dropped-mutations-silent-write-loss-and-load-shedding">Cassandra dropped mutations: silent write loss and load shedding&lt;/h1>
&lt;p>Your application logs successful writes, but reads return stale or missing data. An alert fires on &lt;code>DroppedMessage&lt;/code> rate for &lt;code>MUTATION&lt;/code> scope. The client never received an error, yet a replica discarded the write after it sat in the &lt;code>MutationStage&lt;/code> queue past timeout. This is Cassandra load shedding. Silent write loss occurs whenever not enough other replicas succeed to meet the consistency level.&lt;/p></description></item><item><title>Cassandra dropped reads and other messages: reading nodetool tpstats Dropped</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-dropped-reads-and-messages/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-dropped-reads-and-messages/</guid><description>&lt;h1 id="cassandra-dropped-reads-and-other-messages-reading-nodetool-tpstats-dropped">Cassandra dropped reads and other messages: reading nodetool tpstats Dropped&lt;/h1>
&lt;p>When &lt;code>nodetool tpstats&lt;/code> reports non-zero values in the Dropped section, the node discarded internal messages that exceeded their stage timeout. These counters are cumulative since JVM startup, not rates. A non-zero value warrants investigation: the timeout is defined in &lt;code>cassandra.yaml&lt;/code> by settings such as &lt;code>read_request_timeout_in_ms&lt;/code> and &lt;code>write_request_timeout_in_ms&lt;/code>, so a drop means the message sat in the queue for seconds.&lt;/p>
&lt;p>Dropped messages are a lagging indicator. Correlate the drop type with the matching thread pool pending count, disk I/O latency, and GC pause duration to find the root cause.&lt;/p></description></item><item><title>Cassandra frequent memtable flushes: small SSTables and compaction burden</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-frequent-memtable-flushes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-frequent-memtable-flushes/</guid><description>&lt;h1 id="cassandra-frequent-memtable-flushes-small-sstables-and-compaction-burden">Cassandra frequent memtable flushes: small SSTables and compaction burden&lt;/h1>
&lt;p>When &lt;code>MemtableFlushWriter&lt;/code> pending tasks climb on Cassandra nodes, SSTables often land on disk at only tens of megabytes and multiply quickly. Within hours, &lt;code>PendingCompactions&lt;/code> rises, read latency drifts, and disk I/O stays pinned despite flat write throughput. Frequent memtable flushing under memory pressure creates compaction debt and read amplification. Once started, the cluster enters a feedback loop that is hard to unwind without targeted tuning.&lt;/p></description></item><item><title>Cassandra GC death spiral: long pauses, gossip flapping, and recovery</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-gc-death-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-gc-death-spiral/</guid><description>&lt;h1 id="cassandra-gc-death-spiral-long-pauses-gossip-flapping-and-recovery">Cassandra GC death spiral: long pauses, gossip flapping, and recovery&lt;/h1>
&lt;p>You are paged because a Cassandra node is flapping between UP and DOWN in &lt;code>nodetool status&lt;/code>, client timeouts are rising, and system logs show &lt;code>GCInspector&lt;/code> warnings. The node has not crashed. It is stuck in a GC death spiral: heap pressure produces long pauses, gossip marks the node DOWN, and the resulting retry and hint traffic creates even more heap pressure when the node recovers. It can start with a single large partition read, a misconfigured cache, or an oversized batch statement, and escalates until the node is effectively useless. Catch it early by watching the GC floor and gossip stability together, not just process uptime.&lt;/p></description></item><item><title>Cassandra GC pauses too long: diagnosing G1 stop-the-world pauses</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-gc-pauses-too-long/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-gc-pauses-too-long/</guid><description>&lt;h1 id="cassandra-gc-pauses-too-long-diagnosing-g1-stop-the-world-pauses">Cassandra GC pauses too long: diagnosing G1 stop-the-world pauses&lt;/h1>
&lt;p>&lt;code>ReadTimeoutException&lt;/code> and &lt;code>WriteTimeoutException&lt;/code> from clients, &lt;code>GCInspector&lt;/code> warnings in &lt;code>system.log&lt;/code>, and nodes flapping between &lt;code>UP&lt;/code> and &lt;code>DOWN&lt;/code> in &lt;code>nodetool status&lt;/code> without a JVM restart mean G1 is producing long stop-the-world pauses. Root causes include promotion pressure, humongous objects, or allocation bursts. Left unchecked, one node&amp;rsquo;s pauses trigger gossip failures, retries, and hint replay that drive cluster-wide degradation.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>G1GC is the default collector for Cassandra 4.x on JDK 11+. During a stop-the-world pause, every thread freezes, including gossip, native transport, and compaction. Cassandra logs &lt;code>GCInspector&lt;/code> warnings when a pause exceeds the configured threshold, commonly 500 ms &lt;!-- TODO: verify default gc_warn_threshold_in_ms in target versions -->. Pauses longer than ~2 seconds cause gossip rounds to be missed; under the default phi accrual failure detector threshold of 8, sustained pauses of tens of seconds result in the node being marked DOWN &lt;!-- TODO: verify exact pause duration for conviction at default phi threshold -->. While the JVM is paused, mutations queue, reads stall, hints accumulate on peers, and clients retry. On recovery, hint replay and retry bursts raise allocation pressure, creating a self-reinforcing spiral.&lt;/p></description></item><item><title>Cassandra gossip flapping: nodes bouncing UP and DOWN</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-gossip-flapping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-gossip-flapping/</guid><description>&lt;h1 id="cassandra-gossip-flapping-nodes-bouncing-up-and-down">Cassandra gossip flapping: nodes bouncing UP and DOWN&lt;/h1>
&lt;p>&lt;code>nodetool status&lt;/code> shows a node flipping between &lt;code>UN&lt;/code> and &lt;code>DN&lt;/code>, or multiple nodes doing it in sequence. Each transition forces the cluster to replay hints, recalculate read repair, and propagate gossip state. When a node transitions more than three times in thirty minutes, you are dealing with gossip flapping. It is almost always JVM heap pressure or GC pauses misdiagnosed as a network problem.&lt;/p></description></item><item><title>Cassandra heap pressure: sizing the JVM heap and tuning G1GC</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-heap-pressure-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-heap-pressure-tuning/</guid><description>&lt;h1 id="cassandra-heap-pressure-sizing-the-jvm-heap-and-tuning-g1gc">Cassandra heap pressure: sizing the JVM heap and tuning G1GC&lt;/h1>
&lt;p>Cassandra runs as a single JVM process per node. Every write path allocation, memtable mutation, read merge buffer, and cache entry lives on the heap. When the heap is undersized or GC is left at JVM defaults, stop-the-world pauses freeze gossip, client requests, and compaction. A pause longer than roughly 18 seconds (the default phi accrual threshold is 8) causes peers to mark the node DOWN, which triggers hinted handoff, replay storms, and client retries that worsen memory pressure: the GC death spiral.&lt;/p></description></item><item><title>Cassandra hint overflow: max_hint_window expiry and silent data divergence</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-hint-overflow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-hint-overflow/</guid><description>&lt;h1 id="cassandra-hint-overflow-max_hint_window-expiry-and-silent-data-divergence">Cassandra hint overflow: max_hint_window expiry and silent data divergence&lt;/h1>
&lt;p>You restart a node after a four-hour outage. Gossip converges, &lt;code>nodetool status&lt;/code> shows &lt;code>UN&lt;/code>, and clients reconnect. Reads at consistency level &lt;code>ONE&lt;/code> return stale data. The node has never been repaired.&lt;/p>
&lt;p>The problem is hint overflow: the outage lasted longer than &lt;code>max_hint_window_in_ms&lt;/code> (default three hours), so coordinators stopped saving hints after the window expired. Writes accepted during the final hour of the outage are missing from that replica. Coordinator logs show no errors; write acknowledgments succeeded because other replicas responded. Only anti-entropy repair closes the gap. Without it, the missing data sits on that replica indefinitely, surfacing as inconsistent reads or resurrected deletes.&lt;/p></description></item><item><title>Cassandra hints accumulating: hinted handoff backlog and replay storms</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-hints-accumulating/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-hints-accumulating/</guid><description>&lt;h1 id="cassandra-hints-accumulating-hinted-handoff-backlog-and-replay-storms">Cassandra hints accumulating: hinted handoff backlog and replay storms&lt;/h1>
&lt;p>A coordinator node triggers a disk alert for &lt;code>/var/lib/cassandra/hints&lt;/code>, or a replica returning from maintenance flaps UP/DOWN under a write burst it never requested. Hints let writes succeed when a replica is temporarily unreachable, but a large backlog turns that safety net into a secondary failure. Hints consume coordinator disk, expire after &lt;code>max_hint_window_in_ms&lt;/code>, and can synchronize into a replay storm that overwhelms a recovering node.&lt;/p></description></item><item><title>Cassandra hot partition: when one key saturates a replica set</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-hot-partition/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-hot-partition/</guid><description>&lt;h1 id="cassandra-hot-partition-when-one-key-saturates-a-replica-set">Cassandra hot partition: when one key saturates a replica set&lt;/h1>
&lt;p>One or two nodes run hot while the rest of the cluster idles. Client P99 latency doubles or triples, but the average looks fine. Timeouts cluster on a subset of hosts, and &lt;code>nodetool status&lt;/code> shows uneven load that does not match token ring expectations. When you trace requests, a single partition key consumes a disproportionate share of reads or writes. This is a hot partition: the partitioner mapped one key to a narrow token range, and the replicas owning that range are saturated.&lt;/p></description></item><item><title>Cassandra java.lang.OutOfMemoryError: Java heap space - causes and recovery</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-out-of-memory-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-out-of-memory-error/</guid><description>&lt;h1 id="cassandra-javalangoutofmemoryerror-java-heap-space---causes-and-recovery">Cassandra java.lang.OutOfMemoryError: Java heap space - causes and recovery&lt;/h1>
&lt;p>Your Cassandra log shows &lt;code>java.lang.OutOfMemoryError: Java heap space&lt;/code>. The node stops responding to client requests, gossip marks it DOWN, and the JVM may crash. Because Cassandra runs as a single JVM process, heap exhaustion freezes every in-memory subsystem: memtables, prepared statement caches, bloom filter summaries, and in-flight request buffers.&lt;/p>
&lt;p>The heap may climb toward its limit for hours, or a single massive allocation from a large partition read or oversized batch can push it over immediately. Recovery depends on whether the root cause is capacity, data modeling, or a traffic flood.&lt;/p></description></item><item><title>Cassandra killed by the Linux OOM killer: off-heap memory and RSS</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-off-heap-oom-kill/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-off-heap-oom-kill/</guid><description>&lt;h1 id="cassandra-killed-by-the-linux-oom-killer-off-heap-memory-and-rss">Cassandra killed by the Linux OOM killer: off-heap memory and RSS&lt;/h1>
&lt;p>The JVM heap chart shows 50% utilization and a flat line. There is no &lt;code>OutOfMemoryError&lt;/code>. Then the Cassandra process vanishes. &lt;code>dmesg&lt;/code> shows the OOM killer terminated the JVM: &lt;code>Killed process 12345 (java)&lt;/code>. The JVM heap metric does not include native allocations: bloom filters, compression metadata, index summaries, direct buffers, and chunk cache. When heap plus off-heap RSS exceeds available RAM, the kernel kills the process. This guide covers how to confirm that pattern, reduce off-heap footprint, and prevent recurrence.&lt;/p></description></item><item><title>Cassandra large partition pathology: Compacting large partition warnings and reads</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-large-partition/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-large-partition/</guid><description>&lt;h1 id="cassandra-large-partition-pathology-compacting-large-partition-warnings-and-reads">Cassandra large partition pathology: Compacting large partition warnings and reads&lt;/h1>
&lt;p>If &lt;code>system.log&lt;/code> prints &lt;code>Compacting large partition ks/table:key (N bytes)&lt;/code> and reads to the affected table are spiking, a single partition has crossed &lt;code>compaction_large_partition_warning_threshold_mb&lt;/code> (default 100 MB). Oversized partitions force the node to deserialize a large in-memory index on reads, rewrite the entire partition during compaction, and move it atomically during streaming or repair. Left alone, the partition grows until it triggers GC pauses, gossip flapping, and cascading retries. Use this guide to confirm the offending key, measure the blast radius, and choose between immediate relief and a data-model fix.&lt;/p></description></item><item><title>Cassandra lightweight transaction contention: Paxos round-trips and CAS latency</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-lwt-contention/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-lwt-contention/</guid><description>&lt;h1 id="cassandra-lightweight-transaction-contention-paxos-round-trips-and-cas-latency">Cassandra lightweight transaction contention: Paxos round-trips and CAS latency&lt;/h1>
&lt;p>A normal Cassandra write reaches a coordinator, gets appended to the commitlog and memtable on the replicas, and returns. One network round-trip. A lightweight transaction (LWT) using &lt;code>IF NOT EXISTS&lt;/code>, &lt;code>IF EXISTS&lt;/code>, or &lt;code>IF &amp;lt;condition&amp;gt;&lt;/code> triggers Paxos consensus instead: four round-trips of coordination overhead, plus serialization of contending requests against the same partition. If your monitoring only tracks aggregate read and write latency, LWT tail latency is invisible until clients time out.&lt;/p></description></item><item><title>Cassandra Monitoring</title><link>https://www.netdata.cloud/monitoring-101/cassandra-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/cassandra-monitoring/</guid><description>&lt;h2 id="cassandra-monitoring">Cassandra Monitoring&lt;/h2>
&lt;p>Apache Cassandra, an open-source NoSQL database management system renowned for its scalability and high availability, is pivotal to many organizations running large-scale data infrastructures. Efficiently monitor Cassandra to ensure its optimal performance and reliability.&lt;/p>
&lt;h3 id="what-is-cassandra">What Is Cassandra?&lt;/h3>
&lt;p>Cassandra is a highly scalable, distributed NoSQL database system designed to manage large amounts of structured data across many commodity servers, providing high availability with no single point of failure. It&amp;rsquo;s specifically used for critical applications needing reliability at large scale. &lt;a href="https://cassandra.apache.org/_/index.html">Learn more about Cassandra&lt;/a>.&lt;/p></description></item><item><title>Cassandra monitoring checklist: the signals every production cluster needs</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-monitoring-checklist/</guid><description>&lt;h1 id="cassandra-monitoring-checklist-the-signals-every-production-cluster-needs">Cassandra monitoring checklist: the signals every production cluster needs&lt;/h1>
&lt;p>The four maturity levels below are cumulative. Do not instrument level 2 until level 1 is visible and alerted.&lt;/p>
&lt;p>Cassandra&amp;rsquo;s peer-to-peer architecture and LSM storage engine create failure modes generic infrastructure monitoring misses. A node can be UP in gossip and accepting CQL connections while dropping mutations or accumulating compaction debt that only surfaces hours later. These signals catch liveness, performance, saturation, and consistency failures: GC death spirals, compaction avalanches, tombstone storms, and silent data divergence.&lt;/p></description></item><item><title>Cassandra monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-monitoring-maturity-model/</guid><description>&lt;h1 id="cassandra-monitoring-maturity-model-from-survival-to-expert">Cassandra monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Cassandra exposes JMX MBeans, virtual tables, and log signals. Without priority, teams miss compaction debt or drown in noise. This model structures production monitoring into four cumulative levels. Each level adds signals that reduce mean time to detection and catch the failures that dominate Cassandra incidents: data resurrection from missed repair, compaction death spirals, and GC-induced gossip flapping.&lt;/p>
&lt;p>Audit your current instrumentation against these levels. See the &lt;a href="https://www.netdata.cloud/guides/cassandra/cassandra-monitoring-checklist/">Cassandra monitoring checklist&lt;/a> for a condensed signal inventory.&lt;/p></description></item><item><title>Cassandra native transport not running: node UP in gossip but refusing CQL clients</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-native-transport-not-running/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-native-transport-not-running/</guid><description>&lt;h1 id="cassandra-native-transport-not-running-node-up-in-gossip-but-refusing-cql-clients">Cassandra native transport not running: node UP in gossip but refusing CQL clients&lt;/h1>
&lt;p>A node reports &lt;code>UN&lt;/code> in &lt;code>nodetool status&lt;/code> but rejects CQL connections on port 9042. Gossip and replication are healthy; the failure is isolated to the native transport layer.&lt;/p>
&lt;p>Because the node remains in the token ring, it continues to handle internode replication, gossip, and streaming. Applications see it as down; the cluster sees it as up. The JMX attribute &lt;code>NativeTransportRunning&lt;/code> on &lt;code>org.apache.cassandra.db:type=StorageService&lt;/code> is &lt;code>false&lt;/code> while gossip heartbeats continue. The usual triggers are &lt;code>nodetool disablebinary&lt;/code> left active after maintenance, or a firewall blocking TCP 9042.&lt;/p></description></item><item><title>Cassandra network partition and split brain: detection and reconciliation</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-network-partition-split-brain/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-network-partition-split-brain/</guid><description>&lt;h1 id="cassandra-network-partition-and-split-brain-detection-and-reconciliation">Cassandra network partition and split brain: detection and reconciliation&lt;/h1>
&lt;p>A partial network partition does not always stop the cluster. Each isolated subset often continues to serve traffic, accept writes, and report itself healthy while marking the other side as DOWN. By the time you notice contradictory gossip views or resurrected data, the two sides have diverged for hours. A healed partition is not self-resolving. If both sides accepted writes, last-write-wins semantics combined with even modest clock skew can silently overwrite valid data. Cross-DC deployments are most vulnerable because WAN latency already strains the phi accrual failure detector and inter-DC gossip paths have more single points of failure.&lt;/p></description></item><item><title>Cassandra node showing DN in nodetool status: gossip, phi, and recovery</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-node-down-nodetool-status/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-node-down-nodetool-status/</guid><description>&lt;h1 id="cassandra-node-showing-dn-in-nodetool-status-gossip-phi-and-recovery">Cassandra node showing DN in nodetool status: gossip, phi, and recovery&lt;/h1>
&lt;p>You run &lt;code>nodetool status&lt;/code> and one of your nodes shows &lt;code>DN&lt;/code> (Down/Normal). Before you restart anything, understand that this output is the local node&amp;rsquo;s opinion, not global truth. In Cassandra&amp;rsquo;s peer-to-peer architecture, every node runs its own phi accrual failure detector over gossip heartbeats. A &lt;code>DN&lt;/code> mark means this specific observer has not heard from the target for roughly 18 seconds at default settings. Another node in the same cluster may still show the same target as &lt;code>UN&lt;/code>.&lt;/p></description></item><item><title>Cassandra node stuck in joining (UJ): bootstrap diagnosis</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-bootstrap-stuck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-bootstrap-stuck/</guid><description>&lt;h1 id="cassandra-node-stuck-in-joining-uj-bootstrap-diagnosis">Cassandra node stuck in joining (UJ): bootstrap diagnosis&lt;/h1>
&lt;p>You add a node to the ring, run &lt;code>nodetool status&lt;/code>, and see it stuck in &lt;code>UJ&lt;/code> (Up/Joining) for hours. The cluster sees it in gossip, but it never transitions to &lt;code>UN&lt;/code> (Up/Normal). Client drivers do not route traffic to it, so the expansion has not added usable capacity. Until the state changes, the node is a ghost member: visible to the ring but unable to serve reads or writes for its assigned token ranges.&lt;/p></description></item><item><title>Cassandra Not enough space for compaction: STCS space amplification and recovery</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-not-enough-space-for-compaction/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-not-enough-space-for-compaction/</guid><description>&lt;h1 id="cassandra-not-enough-space-for-compaction-stcs-space-amplification-and-recovery">Cassandra Not enough space for compaction: STCS space amplification and recovery&lt;/h1>
&lt;p>&lt;code>Not enough space for compaction&lt;/code> in &lt;code>system.log&lt;/code> means STCS has hit a structural space-amplification limit. Disk usage may already be above 50%. Cassandra aborts the compaction, skips the tier, and leaves tombstones and old versions unmerged. SSTable count rises, read amplification increases, and free space stops being reclaimed. Left unchecked, this enters a compaction death spiral that ends in write rejection or disk exhaustion. Recovery requires immediate free space, targeted cleanup, and a headroom plan that accounts for transient STCS amplification.&lt;/p></description></item><item><title>Cassandra OperationTimedOutException: client-side timeouts vs server timeouts</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-operation-timed-out-exception/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-operation-timed-out-exception/</guid><description>&lt;h1 id="cassandra-operationtimedoutexception-client-side-timeouts-vs-server-timeouts">Cassandra OperationTimedOutException: client-side timeouts vs server timeouts&lt;/h1>
&lt;p>&lt;code>OperationTimedOutException&lt;/code> (driver 3.x) or &lt;code>DriverTimeoutException&lt;/code> (driver 4.x) means the driver gave up before the coordinator responded. This is distinct from a server-side &lt;code>ReadTimeout&lt;/code> or &lt;code>WriteTimeout&lt;/code>, where the coordinator explicitly errors because replicas were too slow. Confusing the two leads to tuning the wrong timeout, masking a server capacity problem or adding unnecessary client latency. Use server-side signals and driver behavior to tell them apart.&lt;/p></description></item><item><title>Cassandra pending compactions growing: the compaction backlog runbook</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-pending-compactions-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-pending-compactions-growing/</guid><description>&lt;h1 id="cassandra-pending-compactions-growing-the-compaction-backlog-runbook">Cassandra pending compactions growing: the compaction backlog runbook&lt;/h1>
&lt;p>Pending tasks climbing in &lt;code>nodetool compactionstats&lt;/code> is normal after a bulk load under STCS, but when the number trends upward for hours it signals that your node is producing SSTables faster than compaction can merge them.&lt;/p>
&lt;p>This is the leading indicator of the compaction death spiral. Left unchecked, the backlog drives read amplification up, saturates disk I/O, and eventually exhausts disk space as temporary compaction files accumulate. Writes often stay fast while reads degrade, masking the problem.&lt;/p></description></item><item><title>Cassandra quorum loss: when too many replicas are down to satisfy the CL</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-quorum-loss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-quorum-loss/</guid><description>&lt;h1 id="cassandra-quorum-loss-when-too-many-replicas-are-down-to-satisfy-the-cl">Cassandra quorum loss: when too many replicas are down to satisfy the CL&lt;/h1>
&lt;p>Your application is throwing &lt;code>UnavailableException&lt;/code>. Writes and reads fail immediately, not timing out. The coordinator checked the token ring, counted live replicas, and rejected the request because it cannot satisfy the consistency level. No client retry will fix this server-side. The cluster is in quorum loss and will not recover until enough replicas are restored.&lt;/p>
&lt;p>When the count of down nodes for a token range exceeds &lt;code>floor(RF/2)&lt;/code>, every request at &lt;code>QUORUM&lt;/code> or stronger fails instantly. Surviving nodes may still serve &lt;code>CL=ONE&lt;/code> requests for ranges that have a live replica, but any operation requiring a majority is an outage for those partitions. This is a PAGE-level composite failure pattern that demands immediate node recovery followed by repair.&lt;/p></description></item><item><title>Cassandra read latency spikes: P99 vs P50 and proxyhistograms</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-read-latency-spikes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-read-latency-spikes/</guid><description>&lt;h1 id="cassandra-read-latency-spikes-p99-vs-p50-and-proxyhistograms">Cassandra read latency spikes: P99 vs P50 and proxyhistograms&lt;/h1>
&lt;p>Your application is timing out on Cassandra reads, but a quick glance at average latency looks acceptable. This is the percentile trap. In Cassandra, read latency follows a long-tail distribution: most requests are fast, but a fraction hit slow replicas, large partitions, or GC pauses and become orders of magnitude slower. The critical first step is knowing whether the spike lives at the coordinator or the replica, and whether it is systemic or isolated to the tail. &lt;code>nodetool proxyhistograms&lt;/code> and &lt;code>nodetool tablehistograms&lt;/code> answer exactly that.&lt;/p></description></item><item><title>Cassandra ReadTimeoutException: diagnosing coordinator read timeouts</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-read-timeout-exception/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-read-timeout-exception/</guid><description>&lt;h1 id="cassandra-readtimeoutexception-diagnosing-coordinator-read-timeouts">Cassandra ReadTimeoutException: diagnosing coordinator read timeouts&lt;/h1>
&lt;p>The driver throws &lt;code>ReadTimeoutException&lt;/code> when the coordinator fails to gather enough replica responses within &lt;code>read_request_timeout_in_ms&lt;/code> (default 5000 ms). This is a server-side timeout. It maps directly to the JMX metric &lt;code>ClientRequest,scope=Read,name=Timeouts&lt;/code>.&lt;/p>
&lt;p>This is not &lt;code>OperationTimedOutException&lt;/code>, which fires on the driver&amp;rsquo;s socket timeout. It is also not &lt;code>UnavailableException&lt;/code>, which means not enough replicas were alive to attempt the read. Here, replicas are alive but too slow. The read may have executed partially on some replicas, yet the coordinator could not assemble a response that met the consistency level within the window.&lt;/p></description></item><item><title>Cassandra repair failing or stuck: partial repairs and how to verify completion</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-repair-failing-or-stuck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-repair-failing-or-stuck/</guid><description>&lt;h1 id="cassandra-repair-failing-or-stuck-partial-repairs-and-how-to-verify-completion">Cassandra repair failing or stuck: partial repairs and how to verify completion&lt;/h1>
&lt;p>&lt;code>nodetool repair&lt;/code> returning to the prompt, or hanging at 99 percent, are both dangerous because the worst outcome is silent: a partial repair. A partial repair anti-compacts some token ranges and skips others, leaving inconsistency while looking complete. Because incremental repair marks SSTables as repaired during anti-compaction, a session that fails mid-range leaves later ranges unrepaired while earlier ones are already marked done. No built-in alert fires when only forty percent of ranges were covered. If unrepaired ranges contain tombstones, deleted data resurrects once &lt;code>gc_grace_seconds&lt;/code> passes. This guide shows how to verify completion, diagnose stuck sessions, and recover without causing a cascading I/O incident.&lt;/p></description></item><item><title>Cassandra repair not running: the silent gap that resurrects deleted data</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-repair-not-running/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-repair-not-running/</guid><description>&lt;h1 id="cassandra-repair-not-running-the-silent-gap-that-resurrects-deleted-data">Cassandra repair not running: the silent gap that resurrects deleted data&lt;/h1>
&lt;p>Deleted rows reappear in application queries, or &lt;code>system_distributed.repair_history&lt;/code> shows a last success older than your &lt;code>gc_grace_seconds&lt;/code> window. In Cassandra, a repair that is not running, not completing, or not scheduled is an active data integrity risk.&lt;/p>
&lt;p>Cassandra uses tombstones to track deletions and TTL expirations. Those tombstones must survive on every replica until anti-entropy repair propagates the delete. Once &lt;code>gc_grace_seconds&lt;/code> passes, compaction drops tombstones. If repair has not finished for a table within that window, nodes that missed the original delete retain live data while the rest have discarded the tombstone. The result is zombie data resurrection.&lt;/p></description></item><item><title>Cassandra repair overload: when anti-entropy repair causes the outage it prevents</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-repair-overload/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-repair-overload/</guid><description>&lt;h1 id="cassandra-repair-overload-when-anti-entropy-repair-causes-the-outage-it-prevents">Cassandra repair overload: when anti-entropy repair causes the outage it prevents&lt;/h1>
&lt;p>A full &lt;code>nodetool repair&lt;/code> started during peak traffic can spike P99 read latency from milliseconds to hundreds of milliseconds, trigger write timeouts, and cause nodes to flap between UP and DOWN in gossip. The repair job meant to prevent inconsistency becomes the cause of the outage.&lt;/p>
&lt;p>Anti-entropy repair is a heavy distributed scan, not a background task. It reads all local data to build Merkle trees, exchanges hashes with replicas, and streams differing ranges. On a multi-terabyte node this means terabytes of sequential disk reads, heavy CPU hashing, and gigabits of network traffic. Without dedicated headroom, repair competes with the commitlog, memtable flushes, compaction, and client requests for disk bandwidth, CPU, and network. The result is thread pool backpressure, dropped messages, GC pressure from Merkle tree construction, and eventually gossip failure as the node becomes unresponsive.&lt;/p></description></item><item><title>Cassandra Scanned over N tombstones warning: finding the offending query</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-scanned-over-tombstones-warning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-scanned-over-tombstones-warning/</guid><description>&lt;h1 id="cassandra-scanned-over-n-tombstones-warning-finding-the-offending-query">Cassandra Scanned over N tombstones warning: finding the offending query&lt;/h1>
&lt;p>Your application latency may still look acceptable at the median, but &lt;code>system.log&lt;/code> is filling with warnings like &lt;code>Read 500 live rows and 12000 tombstone cells for query SELECT ...&lt;/code> or &lt;code>Scanned over 5000 tombstones&lt;/code>. These messages mean a single read is sifting through thousands of delete markers to return a small result set. Each tombstone consumes CPU, disk I/O, and heap memory during the merge phase. Left alone, the same query will eventually cross &lt;code>tombstone_failure_threshold&lt;/code> (default 100000) and be aborted by the coordinator. The log line gives you the query text, but the real operational work is deciding whether the root cause is a missing repair, a compaction backlog, or a data model mismatch, and then confirming which partition is the actual source of the dead data.&lt;/p></description></item><item><title>Cassandra schema disagreement: nodetool describecluster shows multiple versions</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-schema-disagreement/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-schema-disagreement/</guid><description>&lt;h1 id="cassandra-schema-disagreement-nodetool-describecluster-shows-multiple-versions">Cassandra schema disagreement: nodetool describecluster shows multiple versions&lt;/h1>
&lt;p>&lt;code>nodetool describecluster&lt;/code> should report exactly one schema version UUID cluster-wide. When the &lt;code>Schema versions&lt;/code> section lists more than one UUID, the cluster is in schema disagreement. DDL (&lt;code>CREATE&lt;/code>, &lt;code>ALTER&lt;/code>, &lt;code>DROP&lt;/code>) fails or hangs until every node converges. DML and reads against existing tables are unaffected.&lt;/p>
&lt;p>Transient disagreement lasting less than five minutes during or immediately after a DDL change is normal. Cassandra propagates schema mutations asynchronously via gossip. If multiple versions persist beyond five minutes, you have a stuck node, a partitioned peer, or a migration stage backlog.&lt;/p></description></item><item><title>Cassandra secondary index pitfalls: when 2i scatters across the ring</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-secondary-index-pitfalls/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-secondary-index-pitfalls/</guid><description>&lt;h1 id="cassandra-secondary-index-pitfalls-when-2i-scatters-across-the-ring">Cassandra secondary index pitfalls: when 2i scatters across the ring&lt;/h1>
&lt;p>A query on an indexed column behaves differently in Cassandra than in a relational database. Without a partition key to anchor the request, the coordinator cannot hash the value to a token range and choose the right replicas. It fans out to every node in the cluster, waits for each local index scan to complete, and collates the results. Adding nodes can make indexed queries slower because scatter is linear with cluster size.&lt;/p></description></item><item><title>Cassandra snapshots silently consuming disk: hard links and clearsnapshot</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-snapshots-consuming-disk/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-snapshots-consuming-disk/</guid><description>&lt;h1 id="cassandra-snapshots-silently-consuming-disk-hard-links-and-clearsnapshot">Cassandra snapshots silently consuming disk: hard links and clearsnapshot&lt;/h1>
&lt;p>&lt;code>df&lt;/code> shows usage climbing toward 90%, but &lt;code>nodetool info&lt;/code> Load does not explain the gap. If commitlog and hints are normal and compaction backlog is small, the missing space is likely a snapshot taken days or weeks ago. Cassandra snapshots use hard links, so they allocate no additional blocks at creation. After compaction deletes live SSTables, the snapshot links become the sole owners of old blocks and silently consume gigabytes.&lt;/p></description></item><item><title>Cassandra streaming failures: stalled bootstrap, decommission, and rebuild</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-streaming-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-streaming-failures/</guid><description>&lt;h1 id="cassandra-streaming-failures-stalled-bootstrap-decommission-and-rebuild">Cassandra streaming failures: stalled bootstrap, decommission, and rebuild&lt;/h1>
&lt;p>Bootstrapping stuck in JOINING for six hours. Decommission streaming with zero byte progress after eight hours. Rebuild failed, leaving incomplete token ranges. These are streaming failures.&lt;/p>
&lt;p>Streaming moves SSTables between nodes during topology changes and repair. When a session fails, the topology is left incomplete. When it stalls with no progress for more than thirty minutes, the node stays in a transitional state and may not handle traffic correctly. Root causes are usually network timeouts, source node bottlenecks, or configuration mismatches. Unlike client request timeouts, streaming failures do not always surface as explicit errors; a session can hang silently while the control channel stays open.&lt;/p></description></item><item><title>Cassandra thread pool pending and blocked tasks: SEDA backpressure</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-thread-pool-pending-blocked/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-thread-pool-pending-blocked/</guid><description>&lt;h1 id="cassandra-thread-pool-pending-and-blocked-tasks-seda-backpressure">Cassandra thread pool pending and blocked tasks: SEDA backpressure&lt;/h1>
&lt;p>You run &lt;code>nodetool tpstats&lt;/code> and see non-zero values in the Pending or Blocked columns. On a healthy Cassandra node, request-stage pools like MUTATION and READ should show zero pending tasks in steady state. When pending climbs and stays above zero, the SEDA pipeline is backing up. If Blocked also rises, the node has moved from queuing to rejecting work.&lt;/p>
&lt;p>This is not a transient spike you can ignore. Sustained pending tasks on the MUTATION or READ stages add latency to every client request. Blocked tasks mean the queue is full and the node is actively shedding load. In the GOSSIP stage, even a small pending backlog is an emergency: it means gossip is falling behind, which leads to false DOWN marking across the cluster.&lt;/p></description></item><item><title>Cassandra tombstone storm: delete-heavy tables and read latency collapse</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-tombstone-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-tombstone-storm/</guid><description>&lt;h1 id="cassandra-tombstone-storm-delete-heavy-tables-and-read-latency-collapse">Cassandra tombstone storm: delete-heavy tables and read latency collapse&lt;/h1>
&lt;p>DELETEs and TTL expirations write tombstones. Read latency bifurcates: P50 stays flat while P99 spikes. Logs show queries scanning thousands of tombstones; eventually clients see queries abort after crossing &lt;code>tombstone_failure_threshold&lt;/code>. Disk space does not shrink despite deletions. This is a tombstone storm.&lt;/p>
&lt;p>Cassandra does not remove deleted data immediately. A DELETE inserts a tombstone marker that persists until compaction purges it after &lt;code>gc_grace_seconds&lt;/code> elapses and all replicas have been repaired. When tombstones scatter across many SSTables, every read must scan and merge them, generating temporary heap objects and escalating GC pressure. Only a subset of partitions may be affected, which is why P50 stays flat while tail latency explodes.&lt;/p></description></item><item><title>Cassandra TombstoneOverwhelmingException: reads aborted by tombstone_failure_threshold</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-tombstone-overwhelming-exception/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-tombstone-overwhelming-exception/</guid><description>&lt;h1 id="cassandra-tombstoneoverwhelmingexception-reads-aborted-by-tombstone_failure_threshold">Cassandra TombstoneOverwhelmingException: reads aborted by tombstone_failure_threshold&lt;/h1>
&lt;p>TombstoneOverwhelmingException is a hard stop. Cassandra aborts the read after scanning more tombstones than &lt;code>tombstone_failure_threshold&lt;/code> allows. The default is 100000 tombstones per query. When this exception hits a production table, client reads fail outright.&lt;/p>
&lt;p>Without the threshold, a tombstone-heavy read pins cores, saturates disk I/O, and triggers GC pauses that cascade into gossip flapping or OOM. The exception is a circuit breaker. It is also a signal that tombstones are accumulating faster than compaction can purge them, or that the data model is generating them faster than the storage engine can remove them.&lt;/p></description></item><item><title>Cassandra Too many open files: file descriptor exhaustion</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-too-many-open-files/</guid><description>&lt;h1 id="cassandra-too-many-open-files-file-descriptor-exhaustion">Cassandra Too many open files: file descriptor exhaustion&lt;/h1>
&lt;p>Your application starts logging connection timeouts or &lt;code>Unable to connect&lt;/code> errors, yet &lt;code>nodetool status&lt;/code> still marks the node as UP. Inside the Cassandra system log, you see &lt;code>java.io.IOException: Too many open files&lt;/code> or compaction tasks failing with &lt;code>Cannot open enough files&lt;/code>. This is file descriptor exhaustion. The node operates normally until it hits the process ulimit, then it cannot open new SSTables, accept client sockets, or maintain internode connections. The result is a node that looks alive in gossip but is functionally unable to serve reads, writes, or background compaction.&lt;/p></description></item><item><title>Cassandra too many SSTables per table: read amplification and how to fix it</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-too-many-sstables/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-too-many-sstables/</guid><description>&lt;h1 id="cassandra-too-many-sstables-per-table-read-amplification-and-how-to-fix-it">Cassandra too many SSTables per table: read amplification and how to fix it&lt;/h1>
&lt;p>Your Cassandra read P99 latency is climbing. Queries that used to take milliseconds now time out. The application reports intermittent &lt;code>ReadTimeoutException&lt;/code>. You check &lt;code>nodetool tablestats&lt;/code> and see one table sitting at 200 SSTables. For LCS, that is a catastrophe. For STCS, anything sustained above 50 means compaction has fallen behind.&lt;/p>
&lt;p>Each extra SSTable adds a bloom filter check and a potential disk seek to every read. The read path must merge fragments from memtables plus every SSTable that might contain the partition. When SSTable counts balloon, the node spends more time checking filters and seeking than returning data. This is read amplification.&lt;/p></description></item><item><title>Cassandra TTL tombstone accumulation: time-series tables and TWCS</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-ttl-tombstone-accumulation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-ttl-tombstone-accumulation/</guid><description>&lt;h1 id="cassandra-ttl-tombstone-accumulation-time-series-tables-and-twcs">Cassandra TTL tombstone accumulation: time-series tables and TWCS&lt;/h1>
&lt;p>You set a TTL on every row expecting old data to vanish. Instead, disk usage climbs, read latency spikes, and your logs fill with tombstone warnings.&lt;/p>
&lt;p>An expired TTL does not delete data immediately. Cassandra writes a tombstone that survives until compaction runs and the tombstone exceeds gc_grace_seconds (default 864000 seconds, or 10 days). In time-series workloads ingesting millions of TTL&amp;rsquo;d rows per hour, the wrong compaction strategy turns this bookkeeping into a production incident.&lt;/p></description></item><item><title>Cassandra UnavailableException: not enough replicas for the consistency level</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-unavailable-exception/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-unavailable-exception/</guid><description>&lt;h1 id="cassandra-unavailableexception-not-enough-replicas-for-the-consistency-level">Cassandra UnavailableException: not enough replicas for the consistency level&lt;/h1>
&lt;p>Cassandra threw &lt;code>UnavailableException&lt;/code>. The coordinator rejected the request immediately without contacting replicas. No mutation occurred, and retrying at the same consistency level against the same coordinator fails until enough replicas are &lt;code>UN&lt;/code> in the coordinator&amp;rsquo;s gossip view. This is a topology problem, not a performance problem.&lt;/p>
&lt;p>This exception is qualitatively different from &lt;code>TimeoutException&lt;/code>. A timeout means enough replicas were alive but responded too slowly. An unavailable means the coordinator never sent the request because the topology could not satisfy the consistency level.&lt;/p></description></item><item><title>Cassandra WriteTimeoutException: coordinator write timeouts and writeType</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-write-timeout-exception/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-write-timeout-exception/</guid><description>&lt;h1 id="cassandra-writetimeoutexception-coordinator-write-timeouts-and-writetype">Cassandra WriteTimeoutException: coordinator write timeouts and writeType&lt;/h1>
&lt;p>A &lt;code>WriteTimeoutException&lt;/code> means the coordinator did not receive enough replica acknowledgments before &lt;code>write_request_timeout_in_ms&lt;/code> expired. It does not mean the write failed; one or more replicas may have already persisted the mutation, so the outcome is ambiguous. The &lt;code>writeType&lt;/code> field determines whether a client-side retry is safe.&lt;/p>
&lt;p>Distinguish between a slow replica and an unavailable one. Understand the idempotency contract for the write type, and correlate coordinator timeouts with replica-side saturation. The default &lt;code>write_request_timeout_in_ms&lt;/code> is 2000 ms. If replicas cannot append to the commitlog and update the memtable within that window, the coordinator throws this exception.&lt;/p></description></item><item><title>Cassandra zombie data resurrection: gc_grace_seconds and unrepaired tombstones</title><link>https://www.netdata.cloud/guides/cassandra/cassandra-data-resurrection-gc-grace/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/cassandra-data-resurrection-gc-grace/</guid><description>&lt;h1 id="cassandra-zombie-data-resurrection-gc_grace_seconds-and-unrepaired-tombstones">Cassandra zombie data resurrection: gc_grace_seconds and unrepaired tombstones&lt;/h1>
&lt;p>A query returns rows that were deleted weeks ago. Application logs show no errors. All nodes report UP in &lt;code>nodetool status&lt;/code>. Compaction is running and disk usage looks normal. The data is back because a tombstone was compacted away before anti-entropy repair verified that every replica saw the delete. Once the tombstone is gone, live data on unrepaired replicas is treated as authoritative. The next repair or read repair streams it back to the tombstone-less nodes.&lt;/p></description></item><item><title>Castle Rock Computing SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/castle-rock-computing-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/castle-rock-computing-snmp-traps/</guid><description/></item><item><title>Cato Networks</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cato-networks/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cato-networks/</guid><description/></item><item><title>Cato Networks Topology</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/cato-networks-topology/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/cato-networks-topology/</guid><description/></item><item><title>CDP Topology</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/cdp-topology/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/cdp-topology/</guid><description/></item><item><title>Ce T SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ce-t-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ce-t-snmp-traps/</guid><description/></item><item><title>Celery</title><link>https://www.netdata.cloud/integrations/data-collection/applications/celery/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/celery/</guid><description/></item><item><title>Celery Monitoring</title><link>https://www.netdata.cloud/monitoring-101/celery-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/celery-monitoring/</guid><description>&lt;h2 id="celery-monitoring">Celery Monitoring&lt;/h2>
&lt;h3 id="what-is-celery">What Is Celery?&lt;/h3>
&lt;p>Celery is an open-source distributed task queue system designed to handle asynchronous tasks effectively. It’s built on a robust, flexible architecture and is commonly used in web development to manage scheduled tasks and ease resource consumption. It functions using a producer/consumer pattern, with workers executing jobs and a message broker dispatching them. Celery is praised for its simplicity, scalability, and language-agnostic nature, making it a favorite among developers seeking reliable task management and distributed computing frameworks.&lt;/p></description></item><item><title>Centec Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/centec-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/centec-networks-inc-snmp-traps/</guid><description/></item><item><title>Centillion Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/centillion-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/centillion-networks-inc-snmp-traps/</guid><description/></item><item><title>CentOS</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/centos/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/centos/</guid><description/></item><item><title>CentOS Stream</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/centos-stream/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/centos-stream/</guid><description/></item><item><title>Centrum Communications Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/centrum-communications-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/centrum-communications-inc-snmp-traps/</guid><description/></item><item><title>Ceph</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ceph/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ceph/</guid><description/></item><item><title>Ceph backfill_toofull: recovery blocked because target OSDs are full</title><link>https://www.netdata.cloud/guides/ceph/ceph-backfill-toofull/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-backfill-toofull/</guid><description>&lt;h1 id="ceph-backfill_toofull-recovery-blocked-because-target-osds-are-full">Ceph backfill_toofull: recovery blocked because target OSDs are full&lt;/h1>
&lt;p>A PG in &lt;code>backfill_toofull&lt;/code> is not making progress toward &lt;code>active+clean&lt;/code>. The source OSD has the data, CRUSH has chosen the target, but the target refuses the reservation because it is above &lt;code>backfillfull_ratio&lt;/code>. Client reads and writes still succeed; the redundancy gap is not closing.&lt;/p>
&lt;p>The cluster is usually not at the hard &lt;code>full&lt;/code> ratio (0.95 default). It is at &lt;code>backfillfull&lt;/code> (0.90 default) on one or more individual OSDs, and that is enough to halt recovery for every PG whose up set includes them. Symptoms: HEALTH_WARN (or HEALTH_ERR on older releases), a flat degraded-object count, and recovery bytes/sec near zero.&lt;/p></description></item><item><title>Ceph blocked ops: client I/O stuck behind a single slow OSD</title><link>https://www.netdata.cloud/guides/ceph/ceph-blocked-ops/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-blocked-ops/</guid><description>&lt;h1 id="ceph-blocked-ops-client-io-stuck-behind-a-single-slow-osd">Ceph blocked ops: client I/O stuck behind a single slow OSD&lt;/h1>
&lt;p>A client writes an object. The primary OSD forwards sub-ops to its replicas, then waits for every replica to ack before it acks the client. If any one OSD in the acting set stalls, the whole op stalls. After &lt;code>osd_op_complaint_time&lt;/code> (default 30 seconds) the OSD logs a slow op and the monitor raises &lt;code>SLOW_OPS&lt;/code>. Past that boundary the op is effectively blocked, not merely slow. Clients time out and retry, which puts more ops in flight against the same stuck OSD.&lt;/p></description></item><item><title>Ceph BLUEFS_SPILLOVER: RocksDB metadata spilling onto the slow device</title><link>https://www.netdata.cloud/guides/ceph/ceph-bluestore-db-spillover/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-bluestore-db-spillover/</guid><description>&lt;h1 id="ceph-bluefs_spillover-rocksdb-metadata-spilling-onto-the-slow-device">Ceph BLUEFS_SPILLOVER: RocksDB metadata spilling onto the slow device&lt;/h1>
&lt;p>A small subset of OSDs shows periodic commit-latency spikes and slow ops while the rest of the cluster looks healthy. Capacity metrics are normal. SMART is clean. &lt;code>ceph -s&lt;/code> reports &lt;code>HEALTH_WARN&lt;/code>, and &lt;code>ceph health detail&lt;/code> returns something like:&lt;/p>
&lt;pre tabindex="0">&lt;code>BLUEFS_SPILLOVER
 3 OSDs spilled over ~18 GiB metadata from &amp;#39;db&amp;#39; device
 (e.g. osd.12 spilled over 6.1 GiB metadata from &amp;#39;db&amp;#39; device)
&lt;/code>&lt;/pre>&lt;p>BlueStore&amp;rsquo;s RocksDB metadata has outgrown its dedicated fast DB partition (SSD/NVMe) and is spilling onto the slow HDD data partition. Compaction that took milliseconds on flash now takes seconds on spinning disk. Between compaction cycles the OSD looks fine; during compaction it stalls. This is a cliff edge, not gradual degradation: the moment &lt;code>slow_used_bytes&lt;/code> goes nonzero, latency steps up by one to two orders of magnitude on the affected OSDs.&lt;/p></description></item><item><title>Ceph BlueStore allocator fragmentation: rising latency at moderate fullness</title><link>https://www.netdata.cloud/guides/ceph/ceph-bluestore-fragmentation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-bluestore-fragmentation/</guid><description>&lt;h1 id="ceph-bluestore-allocator-fragmentation-rising-latency-at-moderate-fullness">Ceph BlueStore allocator fragmentation: rising latency at moderate fullness&lt;/h1>
&lt;p>Write latency climbs on OSDs that are only moderately full. Capacity dashboards look healthy, the cluster sits well below &lt;code>nearfull&lt;/code>, and SMART is clean. But a subset of OSDs shows rising commit and apply latency, slow ops begin accumulating, and the affected OSDs are the long-lived ones carrying heavy overwrite workloads. This is BlueStore allocator fragmentation: the free-space structure BlueStore walks on every allocation has become shredded into small, non-contiguous free extents, so each new write needs a longer search before a suitable run of blocks is found.&lt;/p></description></item><item><title>Ceph BlueStore RocksDB compaction stalls: periodic latency spikes</title><link>https://www.netdata.cloud/guides/ceph/ceph-bluestore-compaction-stall/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-bluestore-compaction-stall/</guid><description>&lt;h1 id="ceph-bluestore-rocksdb-compaction-stalls-periodic-latency-spikes">Ceph BlueStore RocksDB compaction stalls: periodic latency spikes&lt;/h1>
&lt;p>Periodic latency spikes on Ceph OSDs that appear and clear on their own schedule are a signature of BlueStore RocksDB compaction. The cluster reports brief bursts of slow ops on a small set of OSDs, commit latency climbs for seconds to minutes, then returns to baseline without intervention. Health checks may briefly show &lt;code>SLOW_OPS&lt;/code>&lt;!-- TODO: verify whether a distinct BLUESTORE_SLOW_OP_ALERT health check exists on Reef and later, or whether operators only see SLOW_OPS --> before clearing.&lt;/p></description></item><item><title>Ceph capacity death spiral: an OSD fails and recovery has nowhere to go</title><link>https://www.netdata.cloud/guides/ceph/ceph-capacity-death-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-capacity-death-spiral/</guid><description>&lt;h1 id="ceph-capacity-death-spiral-an-osd-fails-and-recovery-has-nowhere-to-go">Ceph capacity death spiral: an OSD fails and recovery has nowhere to go&lt;/h1>
&lt;p>An OSD fails on a cluster you have been running at 80% or higher. &lt;code>ceph -s&lt;/code> shows recovery starting, then &lt;code>ceph health detail&lt;/code> lists &lt;code>OSD_NEARFULL&lt;/code> and &lt;code>BACKFILL_TOOFULL&lt;/code>. The degraded PG count climbs, then flatlines. Recovery bytes per second sits near zero. Nothing is healing, and you are one more failure away from data loss.&lt;/p>
&lt;p>This is the Ceph capacity death spiral. It is not a single fault: it is the intersection of tight capacity, CRUSH imbalance, and the hard thresholds Ceph uses to protect itself. Recovery needs spare space on the surviving OSDs. When those OSDs are already past the backfillfull ratio (default 0.90), Ceph refuses to push more data onto them, and backfill stalls for every PG that would target them.&lt;/p></description></item><item><title>Ceph CephFS client eviction (EBLACKLISTED): blocklisted clients and failed I/O</title><link>https://www.netdata.cloud/guides/ceph/ceph-mds-client-eviction/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-mds-client-eviction/</guid><description>&lt;h1 id="ceph-cephfs-client-eviction-eblacklisted-blocklisted-clients-and-failed-io">Ceph CephFS client eviction (EBLACKLISTED): blocklisted clients and failed I/O&lt;/h1>
&lt;p>CephFS clients suddenly see &lt;code>EBLACKLISTED&lt;/code> errors on open file handles. Reads and writes that worked seconds ago fail; new operations against the same mount return the same error code. The MDS has forcibly evicted the client and added its address to the cluster-wide OSD blocklist. Every OSD now refuses traffic from that client address, and the mount is effectively dead until the blocklist is cleared or expires.&lt;/p></description></item><item><title>Ceph client latency vs OSD latency: fast disks, slow clients</title><link>https://www.netdata.cloud/guides/ceph/ceph-client-vs-osd-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-client-vs-osd-latency/</guid><description>&lt;h1 id="ceph-client-latency-vs-osd-latency-fast-disks-slow-clients">Ceph client latency vs OSD latency: fast disks, slow clients&lt;/h1>
&lt;p>You ran &lt;code>ceph osd perf&lt;/code>, the &lt;code>commit_latency&lt;/code> and &lt;code>apply_latency&lt;/code> columns all look healthy, and you declared the cluster fast. Then a client team opens a ticket: writes are taking 50ms, 100ms, sometimes timing out. You re-check &lt;code>ceph osd perf&lt;/code>. Still fast. The OSDs are not the bottleneck, but the clients are still slow.&lt;/p>
&lt;p>&lt;code>ceph osd perf&lt;/code>, the &lt;code>ceph_osd_*_latency_ms&lt;/code> Prometheus metrics, and &lt;code>rbd perf image iostat&lt;/code> all measure time inside the OSD data path. They do not measure the round trip the client experiences. Client-visible latency is the sum of OSD processing, public network RTT, replication round-trips to peer OSDs, CRUSH map recalculation during flapping, and queuing delay on the OSD&amp;rsquo;s public-facing messenger. A cluster can show 2ms commit latency on every OSD while clients see 50ms response times because the public network is saturated, or because public and cluster traffic share one NIC.&lt;/p></description></item><item><title>Ceph deep scrub performance impact: I/O saturation that mimics an incident</title><link>https://www.netdata.cloud/guides/ceph/ceph-deep-scrub-impact/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-deep-scrub-impact/</guid><description>&lt;h1 id="ceph-deep-scrub-performance-impact-io-saturation-that-mimics-an-incident">Ceph deep scrub performance impact: I/O saturation that mimics an incident&lt;/h1>
&lt;p>A Ceph cluster suddenly shows elevated apply latency across many OSDs. Slow ops tick up. Client write latency degrades. The dashboard looks like the start of a real incident. But cluster health is HEALTH_OK and all PGs are active+clean. Before you start chasing a failing disk or a network partition, check whether a deep scrub is running.&lt;/p>
&lt;p>Deep scrub reads and checksums every byte of every object in a PG. On HDD-backed OSDs, this is a sequential read pass across the entire device. With default scheduling, it happens to every PG roughly once a week. If the schedule concentrates scrubs onto the same OSDs at the same time, or if your scrub window is narrow, deep scrub will saturate disk I/O and produce symptoms that look identical to a real performance incident.&lt;/p></description></item><item><title>Ceph degraded objects: reduced redundancy and the race against a second failure</title><link>https://www.netdata.cloud/guides/ceph/ceph-degraded-objects/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-degraded-objects/</guid><description>&lt;h1 id="ceph-degraded-objects-reduced-redundancy-and-the-race-against-a-second-failure">Ceph degraded objects: reduced redundancy and the race against a second failure&lt;/h1>
&lt;p>Degraded objects in Ceph are the most direct measure of reduced redundancy: somewhere in the cluster, at least one object has fewer live replicas (or fewer erasure-coded fragments) than its pool requires. The cluster is still serving I/O, but the protection against a follow-on failure has thinned for those objects.&lt;/p>
&lt;p>&lt;code>ceph_num_objects_degraded&lt;/code> is not alarming on its own. Every OSD failure, restart, and CRUSH reweight produces a transient spike that recovery is designed to heal. The real signal is the slope: is the count falling back toward zero, or has it plateaued while the cluster is still degraded?&lt;/p></description></item><item><title>Ceph failed_repair: when ceph pg repair cannot fix the inconsistency</title><link>https://www.netdata.cloud/guides/ceph/ceph-pg-failed-repair/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-pg-failed-repair/</guid><description>&lt;h1 id="ceph-failed_repair-when-ceph-pg-repair-cannot-fix-the-inconsistency">Ceph failed_repair: when ceph pg repair cannot fix the inconsistency&lt;/h1>
&lt;p>A &lt;code>failed_repair&lt;/code> PG state means Ceph attempted to repair a data inconsistency that scrub or deep-scrub detected, and could not. The PG is now &lt;code>active+clean+inconsistent+failed_repair&lt;/code>, the cluster is at HEALTH_WARN (typically via the &lt;code>OSD_SCRUB_ERRORS&lt;/code> health check), and the inconsistency will not heal on its own. Reads still succeed from consistent replicas, but the cluster has stopped trying to fix the divergence until you intervene.&lt;/p></description></item><item><title>Ceph FS_DEGRADED: standby MDS failed to take over a rank</title><link>https://www.netdata.cloud/guides/ceph/ceph-fs-degraded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-fs-degraded/</guid><description>&lt;h1 id="ceph-fs_degraded-standby-mds-failed-to-take-over-a-rank">Ceph FS_DEGRADED: standby MDS failed to take over a rank&lt;/h1>
&lt;p>&lt;code>FS_DEGRADED&lt;/code> fires when at least one CephFS rank is &lt;code>failed&lt;/code> or &lt;code>damaged&lt;/code> and a standby did not promote. Clients can usually still reach the filesystem through surviving ranks, but you are running without the failover reserve the MDS cluster was sized to provide.&lt;/p>
&lt;p>&lt;code>FS_DEGRADED&lt;/code> is the precursor to &lt;code>MDS_ALL_DOWN&lt;/code>. If the last active rank fails before you restore a healthy standby, CephFS becomes fully unavailable. The playbook classifies &lt;code>FS_DEGRADED&lt;/code> active for more than 120 seconds as a TICKET. If CephFS is a primary storage interface, treat 120 seconds as the upper bound on response time, not a soft target.&lt;/p></description></item><item><title>Ceph health detail: mapping ceph_health_detail checks to a cause</title><link>https://www.netdata.cloud/guides/ceph/ceph-health-detail/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-health-detail/</guid><description>&lt;h1 id="ceph-health-detail-mapping-ceph_health_detail-checks-to-a-cause">Ceph health detail: mapping ceph_health_detail checks to a cause&lt;/h1>
&lt;p>&lt;code>ceph_health_detail&lt;/code> is the bridge between the umbrella status an alert fires on (&lt;code>ceph_health_status&lt;/code>: 0=HEALTH_OK, 1=HEALTH_WARN, 2=HEALTH_ERR) and the specific fault you have to fix. Each health check is a separate gauge with &lt;code>name&lt;/code>, &lt;code>severity&lt;/code>, and &lt;code>message&lt;/code> labels; value &lt;code>1&lt;/code> means active, &lt;code>0&lt;/code> means inactive. The &lt;code>name&lt;/code> label is the check code (&lt;code>OSD_FULL&lt;/code>, &lt;code>PG_AVAILABILITY&lt;/code>, &lt;code>MON_CLOCK_SKEW&lt;/code>, etc.) and is what you build alert routing, dashboards, and post-incident timelines around.&lt;/p></description></item><item><title>Ceph HEALTH_ERR: reading the umbrella status and finding the real fault</title><link>https://www.netdata.cloud/guides/ceph/ceph-health-err/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-health-err/</guid><description>&lt;h1 id="ceph-health_err-reading-the-umbrella-status-and-finding-the-real-fault">Ceph HEALTH_ERR: reading the umbrella status and finding the real fault&lt;/h1>
&lt;p>&lt;code>ceph_health_status == 2&lt;/code> is the most severe top-level status Ceph reports, and one of the most over-paged signals in production. The reason is structural: HEALTH_ERR is an umbrella aggregation. It tells you that one or more child health checks crossed into ERR severity. It does not tell you which subsystem failed, whether I/O is actually blocked, or whether the condition will self-resolve when PGs finish peering.&lt;/p></description></item><item><title>Ceph HEALTH_WARN: which warnings are noise and which are structural</title><link>https://www.netdata.cloud/guides/ceph/ceph-health-warn/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-health-warn/</guid><description>&lt;h1 id="ceph-health_warn-which-warnings-are-noise-and-which-are-structural">Ceph HEALTH_WARN: which warnings are noise and which are structural&lt;/h1>
&lt;p>HEALTH_WARN is an umbrella status, not a single condition. It fires during normal recovery after an OSD restart, and it also fires when the cluster is one step from data loss. Both look identical if you only watch the top-level status. The common mistake is blanket-silencing WARN because it fires too often during recovery, then missing the structural warnings that signal real danger.&lt;/p></description></item><item><title>Ceph LARGE_OMAP_OBJECTS: the RGW bucket-index OMAP storm and resharding</title><link>https://www.netdata.cloud/guides/ceph/ceph-large-omap-objects/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-large-omap-objects/</guid><description>&lt;h1 id="ceph-large_omap_objects-the-rgw-bucket-index-omap-storm-and-resharding">Ceph LARGE_OMAP_OBJECTS: the RGW bucket-index OMAP storm and resharding&lt;/h1>
&lt;p>The &lt;code>LARGE_OMAP_OBJECTS&lt;/code> health warning fires when a deep scrub finds a RADOS object carrying more OMAP keys than the configured threshold. On clusters running the RADOS Gateway (RGW), the most common trigger is the bucket-index shard: a single shard that has accumulated too many entries because a bucket holds millions of objects with too few index shards.&lt;/p>
&lt;p>Once a shard crosses the threshold, the symptoms cluster around two areas. RGW clients see slow LIST responses and slow multipart coordination on the affected bucket. The OSD hosting the shard shows elevated commit latency, slow ops, and in the worst case BlueStore RocksDB pressure that resembles a compaction stall. The OSD usually remains &lt;code>up+in&lt;/code>, so cluster-level health stays at WARN rather than ERR, which makes the problem easy to overlook until a user complains about a stuck &lt;code>ListObjects&lt;/code> call.&lt;/p></description></item><item><title>Ceph MDS_ALL_DOWN: CephFS is completely unavailable</title><link>https://www.netdata.cloud/guides/ceph/ceph-mds-all-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-mds-all-down/</guid><description>&lt;h1 id="ceph-mds_all_down-cephfs-is-completely-unavailable">Ceph MDS_ALL_DOWN: CephFS is completely unavailable&lt;/h1>
&lt;p>MDS_ALL_DOWN means no active Metadata Server (MDS) rank exists for a CephFS filesystem. Every CephFS client depending on that filesystem blocks on metadata operations. This is a hard outage for CephFS, not a degradation. The health check &lt;code>ceph_health_detail{name=&amp;quot;MDS_ALL_DOWN&amp;quot;}&lt;/code> goes active the moment the Monitor has no active rank to assign for that filesystem.&lt;/p>
&lt;p>The metadata itself is journaled in a RADOS pool. Unless MDS_DAMAGED is also firing, the metadata is intact. The outage is availability, not loss. Recovery means restoring at least one active MDS, after which clients reconnect and resume.&lt;/p></description></item><item><title>Ceph MDS_CACHE_OVERSIZED: cache above limit and the cap recall that follows</title><link>https://www.netdata.cloud/guides/ceph/ceph-mds-cache-oversized/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-mds-cache-oversized/</guid><description>&lt;h1 id="ceph-mds_cache_oversized-cache-above-limit-and-the-cap-recall-that-follows">Ceph MDS_CACHE_OVERSIZED: cache above limit and the cap recall that follows&lt;/h1>
&lt;!-- TODO: verify the canonical health-check name. Ceph sources use MDS_HEALTH_CACHE_OVERSIZED; the shorter MDS_CACHE_OVERSIZED form is used throughout this article as operator shorthand. The MDS_CLIENT_RECALL check name should also be verified against the running version. -->
&lt;p>&lt;code>MDS_CACHE_OVERSIZED&lt;/code> fires when an active Metadata Server daemon&amp;rsquo;s in-memory inode and capability cache crosses &lt;code>mds_cache_memory_limit&lt;/code> * &lt;code>mds_health_cache_threshold&lt;/code>. It is almost always workload-driven: a client has touched enough files in a short enough window that the MDS is caching more metadata than its budget allows.&lt;/p></description></item><item><title>Ceph MDS_DAMAGED: metadata journal or cache corruption</title><link>https://www.netdata.cloud/guides/ceph/ceph-mds-damaged/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-mds-damaged/</guid><description>&lt;h1 id="ceph-mds_damaged-metadata-journal-or-cache-corruption">Ceph MDS_DAMAGED: metadata journal or cache corruption&lt;/h1>
&lt;p>MDS_DAMAGED means the CephFS Metadata Server has found damaged metadata, either in its journal or in on-disk structures read from the metadata pool, and has deliberately refused to continue serving the rank. CephFS may be partially or fully unavailable, and recovery is not the usual automatic failover path. A standby that takes over would replay the same suspect journal, so the cluster parks the rank until a human intervenes.&lt;/p></description></item><item><title>Ceph MON_CLOCK_SKEW: clock drift between monitors and election churn</title><link>https://www.netdata.cloud/guides/ceph/ceph-mon-clock-skew/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-mon-clock-skew/</guid><description>&lt;h1 id="ceph-mon_clock_skew-clock-drift-between-monitors-and-election-churn">Ceph MON_CLOCK_SKEW: clock drift between monitors and election churn&lt;/h1>
&lt;p>The &lt;code>MON_CLOCK_SKEW&lt;/code> health check fires when the leader monitor detects clock drift beyond &lt;code>mon_clock_drift_allowed&lt;/code> (default 0.05 seconds, 50 milliseconds) on any monitor in the quorum. It raises &lt;code>HEALTH_WARN&lt;/code>, not &lt;code>HEALTH_ERR&lt;/code>, and on its own it does not stop client I/O. But it is the precursor to a class of failure that does: repeated Paxos elections, slow map distribution, and, if the skew grows, monitor quorum loss.&lt;/p></description></item><item><title>Ceph MON_DOWN: a monitor out of quorum and reduced redundancy</title><link>https://www.netdata.cloud/guides/ceph/ceph-mon-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-mon-down/</guid><description>&lt;h1 id="ceph-mon_down-a-monitor-out-of-quorum-and-reduced-redundancy">Ceph MON_DOWN: a monitor out of quorum and reduced redundancy&lt;/h1>
&lt;p>&lt;code>MON_DOWN&lt;/code> fires when one or more monitor daemons are not part of the active quorum. The common case is a 3-monitor cluster with one monitor gone: quorum still holds with 2 of 3, but the cluster has lost redundancy and tolerates zero further monitor loss before consensus collapses. The Prometheus Ceph mixin surfaces this as &lt;code>CephMonDown&lt;/code> (warning) and escalates to &lt;code>CephMonDownQuorumAtRisk&lt;/code> (critical) when the number of down monitors equals the minimum quorum count.&lt;/p></description></item><item><title>Ceph monitor election storm: monitors that cannot hold a stable quorum</title><link>https://www.netdata.cloud/guides/ceph/ceph-mon-election-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-mon-election-storm/</guid><description>&lt;h1 id="ceph-monitor-election-storm-monitors-that-cannot-hold-a-stable-quorum">Ceph monitor election storm: monitors that cannot hold a stable quorum&lt;/h1>
&lt;p>A Ceph monitor election storm is what happens when the MON cluster cannot complete and hold an election. Each round of Paxos leader election starts, partially completes, then restarts before the new leader can commit any map updates. The election epoch counter climbs, the leader name changes moment to moment, and &lt;code>ceph -s&lt;/code> itself starts taking several seconds to return because even reading the current map requires a responsive leader.&lt;/p></description></item><item><title>Ceph monitor quorum lost: the cluster can no longer update its maps</title><link>https://www.netdata.cloud/guides/ceph/ceph-mon-quorum-lost/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-mon-quorum-lost/</guid><description>&lt;h1 id="ceph-monitor-quorum-lost-the-cluster-can-no-longer-update-its-maps">Ceph monitor quorum lost: the cluster can no longer update its maps&lt;/h1>
&lt;p>Ceph monitor quorum loss is a PAGE condition. Without a majority of MONs agreeing through Paxos, the cluster cannot commit any map update: no OSD up/down transitions, no PG state changes, no pool edits, no CRUSH adjustments. Existing clients keep running on cached maps for a while, but new client connections fail immediately, and as soon as those cached maps go stale, in-flight I/O stalls.&lt;/p></description></item><item><title>Ceph Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ceph-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ceph-monitoring/</guid><description>&lt;h2 id="ceph-monitoring">Ceph Monitoring&lt;/h2>
&lt;h3 id="what-is-ceph">What Is Ceph?&lt;/h3>
&lt;p>Ceph is a distributed storage system designed to provide excellent performance, reliability, and scalability. It achieves this by distributing data across multiple storage devices and ensuring data redundancy. As an open-source project, Ceph has become a popular choice for managing petabytes of data because of its self-healing and self-managing capabilities.&lt;/p>
&lt;h3 id="monitoring-ceph-with-netdata">Monitoring Ceph With Netdata&lt;/h3>
&lt;p>Monitoring Ceph clusters effectively is crucial to maintaining their optimal performance and ensuring data availability. Netdata offers a comprehensive Ceph monitoring tool that provides real-time insights into your Ceph infrastructure. By using &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/ceph/">Netdata&amp;rsquo;s collector for Ceph&lt;/a>, you can access granular metrics for your entire Ceph cluster, including individual Pools and OSDs.&lt;/p></description></item><item><title>Ceph monitoring checklist: the signals every production cluster needs</title><link>https://www.netdata.cloud/guides/ceph/ceph-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-monitoring-checklist/</guid><description>&lt;h1 id="ceph-monitoring-checklist-the-signals-every-production-cluster-needs">Ceph monitoring checklist: the signals every production cluster needs&lt;/h1>
&lt;p>This checklist defines the signals a production Ceph cluster needs, organized by monitoring maturity. Each level is a superset of the previous one; higher-level signals rely on lower-level context to distinguish real failures from normal churn. The four levels are survival, operational, mature, and expert. You cannot skip levels in practice.&lt;/p>
&lt;p>Two caveats apply across all levels. First, &lt;code>ceph_health_status&lt;/code> is an umbrella, not a complete picture: HEALTH_OK does not guarantee performance, and HEALTH_WARN covers both expected churn (active recovery) and structural problems (nearfull, noout trap, scrub inconsistency). Drill into &lt;code>ceph_health_detail&lt;/code> before treating WARN as noise. Second, per-OSD and per-pool metrics matter more than cluster averages. A cluster at 60% average capacity with one OSD at 85% is closer to trouble than the average suggests.&lt;/p></description></item><item><title>Ceph monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/ceph/ceph-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-monitoring-maturity-model/</guid><description>&lt;h1 id="ceph-monitoring-maturity-model-from-survival-to-expert">Ceph monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Ceph&amp;rsquo;s health surface is unusually broad for a storage system. A cluster can report HEALTH_OK while a single OSD runs at 20x the latency of its peers, while the BlueStore DB partition spills to slow media, or while the PG autoscaler splits placement groups under live client load. The gap between &amp;ldquo;is the cluster up&amp;rdquo; and &amp;ldquo;is the cluster observable enough to operate&amp;rdquo; is where most preventable outages live.&lt;/p></description></item><item><title>Ceph nearfull: the 85% warning that decides whether the cluster can heal</title><link>https://www.netdata.cloud/guides/ceph/ceph-osd-nearfull/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-osd-nearfull/</guid><description>&lt;h1 id="ceph-nearfull-the-85-warning-that-decides-whether-the-cluster-can-heal">Ceph nearfull: the 85% warning that decides whether the cluster can heal&lt;/h1>
&lt;p>&lt;code>OSD_NEARFULL&lt;/code> or &lt;code>POOL_NEAR_FULL&lt;/code> in &lt;code>ceph health detail&lt;/code> means &lt;code>HEALTH_WARN&lt;/code>. Clients are still reading and writing, and nothing looks broken yet. The nearfull ratio (default 0.85) is not a polite reminder to plan storage. It is the point where Ceph warns it is running out of the spare space it needs to heal itself.&lt;/p>
&lt;p>One OSD failure at 85% can push surviving OSDs past backfillfull (0.90), at which point backfills refuse to start and recovery stalls. If another OSD fails in that window, you are left with degraded PGs that cannot be recovered, and the next stop is &lt;code>OSD_FULL&lt;/code> at 95% where all client writes return ENOSPC.&lt;/p></description></item><item><title>Ceph noout flag left set: the most common preventable outage</title><link>https://www.netdata.cloud/guides/ceph/ceph-noout-flag-set/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-noout-flag-set/</guid><description>&lt;h1 id="ceph-noout-flag-left-set-the-most-common-preventable-outage">Ceph noout flag left set: the most common preventable outage&lt;/h1>
&lt;p>The noout flag is the single most common source of preventable Ceph outages. An operator sets it before maintenance to stop Ceph from marking down OSDs as out, then forgets to unset it. The cluster keeps serving I/O, health stays at HEALTH_WARN (not ERR), and nothing visible breaks. But recovery is now disabled by policy. The next OSD failure stacks on top of the first, and degraded PGs that should have healed days ago are still degraded.&lt;/p></description></item><item><title>Ceph norecover and nobackfill: recovery intentionally, or accidentally, stopped</title><link>https://www.netdata.cloud/guides/ceph/ceph-norecover-nobackfill/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-norecover-nobackfill/</guid><description>&lt;h1 id="ceph-norecover-and-nobackfill-recovery-intentionally-or-accidentally-stopped">Ceph norecover and nobackfill: recovery intentionally, or accidentally, stopped&lt;/h1>
&lt;p>The symptom is familiar: PGs are stuck in &lt;code>degraded&lt;/code> or &lt;code>undersized&lt;/code> states, the cluster is &lt;code>HEALTH_WARN&lt;/code>, but recovery throughput is zero. OSD load, network, and capacity all look normal. The cause is often two cluster-wide flags sitting in the OSD map: &lt;code>norecover&lt;/code> and &lt;code>nobackfill&lt;/code>. These flags are legitimate tools for protecting client I/O during recovery storms or maintenance, and a common source of &amp;ldquo;forgotten flag&amp;rdquo; incidents alongside &lt;code>noout&lt;/code>.&lt;/p></description></item><item><title>Ceph OSD commit and apply latency: reading per-OSD latency outliers</title><link>https://www.netdata.cloud/guides/ceph/ceph-osd-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-osd-latency-high/</guid><description>&lt;h1 id="ceph-osd-commit-and-apply-latency-reading-per-osd-latency-outliers">Ceph OSD commit and apply latency: reading per-OSD latency outliers&lt;/h1>
&lt;p>Two metrics tell you more about per-OSD health than almost anything else in Ceph: &lt;code>ceph_osd_commit_latency_ms&lt;/code> and &lt;code>ceph_osd_apply_latency_ms&lt;/code>. They are per-OSD gauges exported by the manager&amp;rsquo;s Prometheus module, labeled by &lt;code>ceph_daemon&lt;/code>, and they measure the OSD&amp;rsquo;s internal I/O time, not the latency your clients experience end-to-end. A cluster can show healthy aggregate latency while one OSD quietly runs 5x slower than its peers of the same device class. The clients whose objects land on that OSD see tail latency spikes; everyone else sees normal performance. Cluster-wide averages hide this. Per-OSD comparison surfaces it.&lt;/p></description></item><item><title>Ceph OSD down: telling a dead disk apart from a network blip</title><link>https://www.netdata.cloud/guides/ceph/ceph-osd-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-osd-down/</guid><description>&lt;h1 id="ceph-osd-down-telling-a-dead-disk-apart-from-a-network-blip">Ceph OSD down: telling a dead disk apart from a network blip&lt;/h1>
&lt;p>&lt;code>OSD_DOWN&lt;/code> fires. &lt;code>ceph_osd_up&lt;/code> flipped to 0. The 600-second countdown to OUT and recovery has begun.&lt;/p>
&lt;p>Your first job: figure out whether this is a dead disk (act now, plan a replacement) or a network blip (wait, verify, let the OSD come back). Acting on the wrong diagnosis wastes disk and network I/O on a needless backfill, or worse, leaves a failing disk in service past the point where SMART was already warning you.&lt;/p></description></item><item><title>Ceph OSD flapping: OSDs cycling up and down and the peering storm that follows</title><link>https://www.netdata.cloud/guides/ceph/ceph-osd-flapping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-osd-flapping/</guid><description>&lt;h1 id="ceph-osd-flapping-osds-cycling-up-and-down-and-the-peering-storm-that-follows">Ceph OSD flapping: OSDs cycling up and down and the peering storm that follows&lt;/h1>
&lt;p>OSD flapping is a failure cascade. One OSD misses heartbeats, peers mark it down, its PGs start peering and recovering elsewhere, the OSD comes back, peering reverses, and the cycle repeats. Each flap mints a new OSD map epoch that every OSD in the cluster must process. The peering overhead from a single flapping OSD can slow dozens of healthy OSDs enough that they also miss heartbeats, and the cascade spreads.&lt;/p></description></item><item><title>Ceph OSD fullness imbalance: one OSD full while the cluster average looks fine</title><link>https://www.netdata.cloud/guides/ceph/ceph-osd-fullness-imbalance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-osd-fullness-imbalance/</guid><description>&lt;h1 id="ceph-osd-fullness-imbalance-one-osd-full-while-the-cluster-average-looks-fine">Ceph OSD fullness imbalance: one OSD full while the cluster average looks fine&lt;/h1>
&lt;p>Cluster writes are blocked. &lt;code>ceph status&lt;/code> shows &lt;code>HEALTH_ERR&lt;/code> with &lt;code>OSD_FULL&lt;/code> active. But &lt;code>ceph df&lt;/code> reports the cluster at 65% utilized. Both are correct.&lt;/p>
&lt;p>The Ceph full ratio (default 0.95) is enforced per-OSD, not cluster-wide. When any single OSD crosses that threshold, Ceph refuses writes for every PG that OSD serves. Because CRUSH distributes PGs across OSDs, one full OSD can block writes to a large fraction of PGs even when ninety-nine other OSDs have ample free space.&lt;/p></description></item><item><title>Ceph OSD heartbeat timeouts: the osd_heartbeat_grace precursor to flapping</title><link>https://www.netdata.cloud/guides/ceph/ceph-osd-heartbeat-timeout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-osd-heartbeat-timeout/</guid><description>&lt;h1 id="ceph-osd-heartbeat-timeouts-the-osd_heartbeat_grace-precursor-to-flapping">Ceph OSD heartbeat timeouts: the osd_heartbeat_grace precursor to flapping&lt;/h1>
&lt;p>Ceph OSDs declare each other down based on a simple rule: if a peer does not respond to heartbeat pings within &lt;code>osd_heartbeat_grace&lt;/code> (default 20 seconds), it is reported to the monitors as unresponsive. Two peers from different failure domains must agree before the monitor marks the OSD down. This is the mechanism that turns a transient stall into an OSD state change, and repeated near-misses against this 20-second window are the most reliable leading indicator of OSD flapping.&lt;/p></description></item><item><title>Ceph OSD up/down vs in/out: the four states and mon_osd_down_out_interval</title><link>https://www.netdata.cloud/guides/ceph/ceph-osd-up-down-in-out/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-osd-up-down-in-out/</guid><description>&lt;h1 id="ceph-osd-updown-vs-inout-the-four-states-and-mon_osd_down_out_interval">Ceph OSD up/down vs in/out: the four states and mon_osd_down_out_interval&lt;/h1>
&lt;p>Most Ceph operations references describe an OSD as &amp;ldquo;down&amp;rdquo; as if it were a single state. It is not. An OSD carries two independent flags: &lt;code>up&lt;/code>/&lt;code>down&lt;/code> (is the daemon alive?) and &lt;code>in&lt;/code>/&lt;code>out&lt;/code> (does CRUSH place data on it?). Combined, that produces four states, and the most dangerous one, &lt;code>down+in&lt;/code>, is invisible if you only look at one flag.&lt;/p>
&lt;p>The two flags are linked by a clock. After an OSD goes &lt;code>down&lt;/code>, the monitors wait &lt;code>mon_osd_down_out_interval&lt;/code> (default 600 seconds) before automatically flipping it to &lt;code>out&lt;/code>. That 10-minute window is the gap between &amp;ldquo;OSD stopped&amp;rdquo; and &amp;ldquo;recovery starts.&amp;rdquo; That timer, and the flags on either side of it, are the vocabulary every recovery, flapping, and noout runbook assumes.&lt;/p></description></item><item><title>Ceph OSD_FULL: all writes stopped at the 95% full ratio</title><link>https://www.netdata.cloud/guides/ceph/ceph-osd-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-osd-full/</guid><description>&lt;h1 id="ceph-osd_full-all-writes-stopped-at-the-95-full-ratio">Ceph OSD_FULL: all writes stopped at the 95% full ratio&lt;/h1>
&lt;p>The cluster suddenly stops accepting writes. Client applications report ENOSPC errors. &lt;code>ceph status&lt;/code> returns &lt;code>HEALTH_ERR&lt;/code>. &lt;code>ceph health detail&lt;/code> shows &lt;code>OSD_FULL&lt;/code> active. Reads still succeed, but every write, update, and delete fails cluster-wide.&lt;/p>
&lt;p>This is a hard stop, not a throttle. Ceph refuses all write operations once any OSD crosses the configured &lt;code>full_ratio&lt;/code> (default 0.95). CRUSH spreads every PG across multiple OSDs, so a single full OSD can block writes to hundreds of PGs even when cluster-average utilization looks moderate.&lt;/p></description></item><item><title>Ceph osd_memory_target: cache eviction, RSS growth, and OSD OOM kills</title><link>https://www.netdata.cloud/guides/ceph/ceph-osd-memory-target/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-osd-memory-target/</guid><description>&lt;h1 id="ceph-osd_memory_target-cache-eviction-rss-growth-and-osd-oom-kills">Ceph osd_memory_target: cache eviction, RSS growth, and OSD OOM kills&lt;/h1>
&lt;p>An OSD repeatedly killed by the kernel OOM killer, or a host where OSDs flap up and down shortly after systemd restarts them, very often traces back to one knob: &lt;code>osd_memory_target&lt;/code>. The same knob set too low for the working set produces a quieter failure: read latency climbs on HDD-backed OSDs as BlueStore evicts onodes and RocksDB block-cache pages the workload actually needs.&lt;/p></description></item><item><title>Ceph PG degraded: fewer replicas than the pool size, and when it matters</title><link>https://www.netdata.cloud/guides/ceph/ceph-pg-degraded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-pg-degraded/</guid><description>&lt;h1 id="ceph-pg-degraded-fewer-replicas-than-the-pool-size-and-when-it-matters">Ceph PG degraded: fewer replicas than the pool size, and when it matters&lt;/h1>
&lt;p>A degraded placement group in Ceph has fewer copies of some objects than the pool&amp;rsquo;s configured &lt;code>size&lt;/code>. When an OSD goes down or is removed, every PG that had a replica on that OSD drops below target and enters the &lt;code>degraded&lt;/code> state. The cluster can still serve I/O as long as the surviving replica count is at or above &lt;code>min_size&lt;/code>, but redundancy is reduced until recovery rebuilds the missing copies.&lt;/p></description></item><item><title>Ceph PG down: no surviving replica for reads or writes</title><link>https://www.netdata.cloud/guides/ceph/ceph-pg-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-pg-down/</guid><description>&lt;h1 id="ceph-pg-down-no-surviving-replica-for-reads-or-writes">Ceph PG down: no surviving replica for reads or writes&lt;/h1>
&lt;p>A placement group in the &lt;code>down&lt;/code> state has no replica that can serve I/O. Reads and writes to objects in that PG fail. Clients see EIO or ENODEV, and &lt;code>ceph health detail&lt;/code> reports &lt;code>PG_AVAILABILITY&lt;/code> with one or more PGs flagged &lt;code>down&lt;/code>. Data is unavailable, not just degraded.&lt;/p>
&lt;p>A PG can briefly pass through &lt;code>down&lt;/code> during peering after an OSD failure, before activating on a surviving replica. The 300-second sustain on &lt;code>sum(ceph_pg_down) &amp;gt; 0&lt;/code> filters that transient window. Once a PG has been &lt;code>down&lt;/code> past the sustain, no OSD in the acting set can serve it, and the cluster will not heal it without operator action.&lt;/p></description></item><item><title>Ceph PG incomplete: placement groups that cannot serve I/O</title><link>https://www.netdata.cloud/guides/ceph/ceph-pg-incomplete/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-pg-incomplete/</guid><description>&lt;h1 id="ceph-pg-incomplete-placement-groups-that-cannot-serve-io">Ceph PG incomplete: placement groups that cannot serve I/O&lt;/h1>
&lt;p>An &lt;code>incomplete&lt;/code> placement group means the PG cannot find enough authoritative data to serve reads or writes. This is a data-availability failure, and it does not self-resolve the way a transient &lt;code>peering&lt;/code> or &lt;code>recovering&lt;/code> state does.&lt;/p>
&lt;p>Alert on &lt;code>sum(ceph_pg_incomplete) &amp;gt; 0&lt;/code> sustained for more than 300 seconds, summed across all pools. The 300 second sustain filters out cold-start peering after a cluster-wide restart, which typically completes within 60-120 seconds. Anything still &lt;code>incomplete&lt;/code> after five minutes is genuinely stuck.&lt;/p></description></item><item><title>Ceph PG inconsistent (OSD_SCRUB_ERRORS): scrub found replica divergence</title><link>https://www.netdata.cloud/guides/ceph/ceph-pg-inconsistent/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-pg-inconsistent/</guid><description>&lt;h1 id="ceph-pg-inconsistent-osd_scrub_errors-scrub-found-replica-divergence">Ceph PG inconsistent (OSD_SCRUB_ERRORS): scrub found replica divergence&lt;/h1>
&lt;p>A scrub or deep-scrub finished comparing replicas for a placement group and found they disagree. The cluster surfaces this as &lt;code>OSD_SCRUB_ERRORS&lt;/code> (often paired with &lt;code>PG_DAMAGED&lt;/code>) and the affected PG sits in &lt;code>active+clean+inconsistent&lt;/code>. Client reads still succeed because Ceph serves them from a consistent replica, but at least one copy in the acting set is corrupt.&lt;/p>
&lt;p>The danger is not the symptom. Reads work, and Ceph did what it was designed to do: detect silent divergence. The danger is that the corruption was found, not fixed, and the window during which an uncorrupted replica survives is your margin of safety. If the OSD holding the good copy fails before you repair, the object becomes unreadable or unwritable.&lt;/p></description></item><item><title>Ceph PG stale: the monitor has not heard from the PG's OSDs</title><link>https://www.netdata.cloud/guides/ceph/ceph-pg-stale/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-pg-stale/</guid><description>&lt;h1 id="ceph-pg-stale-the-monitor-has-not-heard-from-the-pgs-osds">Ceph PG stale: the monitor has not heard from the PG&amp;rsquo;s OSDs&lt;/h1>
&lt;p>A &lt;code>stale&lt;/code> placement group means the monitor cluster has stopped trusting the last status it received for that PG. The acting primary OSD that should be reporting state has gone silent, so the MON cannot confirm whether the PG is active, recovering, or unavailable. The recorded state is preserved but flagged as untrusted.&lt;/p>
&lt;p>This differs from &lt;code>down&lt;/code>. A &lt;code>down&lt;/code> PG means the cluster has positively confirmed that no replica can serve I/O. A &lt;code>stale&lt;/code> PG means the cluster does not know the current state because the reporting chain broke. Both can co-occur, but the response differs: &lt;code>down&lt;/code> is confirmed replica loss; &lt;code>stale&lt;/code> is a communication or reporting failure.&lt;/p></description></item><item><title>Ceph PG undersized: fewer copies than the pool wants to place</title><link>https://www.netdata.cloud/guides/ceph/ceph-pg-undersized/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-pg-undersized/</guid><description>&lt;h1 id="ceph-pg-undersized-fewer-copies-than-the-pool-wants-to-place">Ceph PG undersized: fewer copies than the pool wants to place&lt;/h1>
&lt;p>A Ceph cluster reports &lt;code>ceph health detail&lt;/code> listing PGs in an &lt;code>undersized&lt;/code> state. The cluster serves I/O at reduced redundancy. The warning does not resolve on its own and recovery stalls. A structural block prevents CRUSH from placing the configured number of replicas.&lt;/p>
&lt;p>Unlike &lt;code>peering&lt;/code> or &lt;code>recovering&lt;/code>, &lt;code>undersized&lt;/code> is not transient. The acting set has fewer OSDs than the pool&amp;rsquo;s configured &lt;code>size&lt;/code> (the replication factor for replicated pools, or the k+m sum for erasure-coded pools). Ceph runs the PG at reduced redundancy because CRUSH cannot find enough distinct failure domains to satisfy the rule. The fix is rarely to wait. It requires changing capacity, topology, or the CRUSH rule.&lt;/p></description></item><item><title>Ceph PG_NOT_DEEP_SCRUBBED: scrub verification debt and undetected bit rot</title><link>https://www.netdata.cloud/guides/ceph/ceph-not-scrubbed-in-time/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-not-scrubbed-in-time/</guid><description>&lt;h1 id="ceph-pg_not_deep_scrubbed-scrub-verification-debt-and-undetected-bit-rot">Ceph PG_NOT_DEEP_SCRUBBED: scrub verification debt and undetected bit rot&lt;/h1>
&lt;p>&lt;code>PG_NOT_DEEP_SCRUBBED&lt;/code> means Ceph&amp;rsquo;s proactive integrity check is falling behind. Deep scrub is the only mechanism that reads every byte of every object on every replica and verifies byte-for-byte consistency. When PGs miss their deep-scrub window repeatedly, silent corruption from bit rot, DRAM errors, firmware bugs, and incomplete writes after power loss accumulates without any signal.&lt;/p>
&lt;p>The companion check &lt;code>PG_NOT_SCRUBBED&lt;/code> covers the lighter daily scrub, which compares object metadata across replicas. Both checks surface through &lt;code>ceph health detail&lt;/code> and the &lt;code>ceph_health_detail&lt;/code> Prometheus metric exposed by the MGR module. There is no per-PG &amp;ldquo;overdue&amp;rdquo; gauge in the standard metrics pipeline, which is why scrub debt often goes unmonitored until someone runs &lt;code>ceph -s&lt;/code> and sees a &lt;code>HEALTH_WARN&lt;/code> they do not recognize.&lt;/p></description></item><item><title>Ceph PGs stuck peering: the slow peering loop after a mass restart</title><link>https://www.netdata.cloud/guides/ceph/ceph-pg-peering-stuck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-pg-peering-stuck/</guid><description>&lt;h1 id="ceph-pgs-stuck-peering-the-slow-peering-loop-after-a-mass-restart">Ceph PGs stuck peering: the slow peering loop after a mass restart&lt;/h1>
&lt;p>After a full cluster power cycle, mass OSD restart, or wide maintenance event, peering should finish within a minute or two. When it does not, &lt;code>ceph -s&lt;/code> shows a long list of PGs in &lt;code>peering&lt;/code> or &lt;code>peering+activating&lt;/code>, the count is not decreasing, and client I/O for those PGs is hung. HEALTH_WARN or HEALTH_ERR can sit there for 20, 30, or 60 minutes with no visible progress.&lt;/p></description></item><item><title>Ceph recovery stalled: degraded PGs that are not healing</title><link>https://www.netdata.cloud/guides/ceph/ceph-recovery-stalled/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-recovery-stalled/</guid><description>&lt;h1 id="ceph-recovery-stalled-degraded-pgs-that-are-not-healing">Ceph recovery stalled: degraded PGs that are not healing&lt;/h1>
&lt;p>&lt;code>ceph -s&lt;/code> shows &lt;code>HEALTH_WARN&lt;/code> or &lt;code>HEALTH_ERR&lt;/code>, the cluster reports non-zero &lt;code>degraded&lt;/code> or &lt;code>undersized&lt;/code> placement groups, and clients still appear served. The degraded PG count is not climbing, which feels like progress. It is not. A flat degraded count with zero recovery rate is one of the most dangerous states a Ceph cluster can sit in: nothing is healing, and the operator assumes the system is working through it.&lt;/p></description></item><item><title>Ceph recovery storm: rebuild traffic starving client I/O</title><link>https://www.netdata.cloud/guides/ceph/ceph-recovery-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-recovery-storm/</guid><description>&lt;h1 id="ceph-recovery-storm-rebuild-traffic-starving-client-io">Ceph recovery storm: rebuild traffic starving client I/O&lt;/h1>
&lt;p>A Ceph recovery storm occurs when the cluster&amp;rsquo;s self-healing machinery starves the workloads it is supposed to serve. After an OSD failure, host loss, or bulk OSD addition, recovery and backfill traffic floods the same disks, network links, and OSD CPU that client I/O depends on. Client latency climbs 10x to 100x, applications time out, and retries add more load on top of the recovery stream.&lt;/p></description></item><item><title>Ceph RGW failed requests: aborted request rate and what it actually counts</title><link>https://www.netdata.cloud/guides/ceph/ceph-rgw-failed-requests/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-rgw-failed-requests/</guid><description>&lt;h1 id="ceph-rgw-failed-requests-aborted-request-rate-and-what-it-actually-counts">Ceph RGW failed requests: aborted request rate and what it actually counts&lt;/h1>
&lt;p>&lt;code>ceph_rgw_failed_req&lt;/code> is one of the most commonly misread RADOS Gateway metrics. Operators see &amp;ldquo;failed&amp;rdquo; in the name and assume it counts HTTP 4xx and 5xx responses. It does not. It counts requests where the client connection was aborted before the response completed. That distinction changes who you page, where you look, and which signals you correlate with.&lt;/p>
&lt;h2 id="what-it-is-and-why-it-matters">What it is and why it matters&lt;/h2>
&lt;p>&lt;code>ceph_rgw_failed_req&lt;/code> is a per-instance counter exposed by each RGW daemon, scraped from the admin socket and labeled with &lt;code>instance_id&lt;/code> (alongside &lt;code>ceph_daemon&lt;/code>, &lt;code>instance&lt;/code>, &lt;code>job&lt;/code>). Its sibling counter is &lt;code>ceph_rgw_req&lt;/code>, the total request count. The operational signal is the ratio between the two: a sustained failed-request rate above 5% of total traffic for more than 5 minutes is the playbook&amp;rsquo;s TICKET condition.&lt;/p></description></item><item><title>Ceph RGW garbage-collection backlog: deleted data still consuming space</title><link>https://www.netdata.cloud/guides/ceph/ceph-rgw-gc-backlog/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-rgw-gc-backlog/</guid><description>&lt;h1 id="ceph-rgw-garbage-collection-backlog-deleted-data-still-consuming-space">Ceph RGW garbage-collection backlog: deleted data still consuming space&lt;/h1>
&lt;p>You deleted several terabytes of S3 objects, but &lt;code>ceph df&lt;/code> shows raw usage barely moving. The cluster is approaching nearfull and the write freeze is coming. The deletes returned 204 to clients, so they succeeded from the S3 layer&amp;rsquo;s perspective, but the underlying RADOS objects are still on disk, queued behind the RADOS Gateway garbage collector.&lt;/p>
&lt;p>This is one of the quietest contributors to &amp;ldquo;the cluster is full but we deleted everything&amp;rdquo;. RGW does not free object data inline on delete. It marks the head object as deleted, enqueues the data objects (tail segments, multipart parts) for asynchronous garbage collection, and relies on a background GC thread on each gateway to drain the queue. When that thread stops making progress, deleted data keeps consuming capacity indefinitely.&lt;/p></description></item><item><title>Ceph RGW GET/PUT latency: S3 request latency and queue length</title><link>https://www.netdata.cloud/guides/ceph/ceph-rgw-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-rgw-latency/</guid><description>&lt;h1 id="ceph-rgw-getput-latency-s3-request-latency-and-queue-length">Ceph RGW GET/PUT latency: S3 request latency and queue length&lt;/h1>
&lt;p>When users report slow S3 GET or PUT responses, the RGW daemon is rarely the root cause. RADOS Gateway is a stateless HTTP frontend that translates REST calls into RADOS object operations. Its observed latency is dominated by the time those underlying operations take, plus whatever queuing happens when the gateway has more in-flight work than it can drain.&lt;/p>
&lt;p>The RGW perf counters expose two distinct kinds of signal: per-operation latency accumulators (&lt;code>ceph_rgw_op_get_obj_lat_sum/_count&lt;/code> and &lt;code>ceph_rgw_op_put_obj_lat_sum/_count&lt;/code>) for the S3 operations themselves, plus queue gauges (&lt;code>ceph_rgw_qlen&lt;/code> and &lt;code>ceph_rgw_qactive&lt;/code>) that show whether requests are piling up inside the daemon. Treating those signals together is the difference between &amp;ldquo;S3 is slow&amp;rdquo; and &amp;ldquo;S3 is slow because one OSD hosting a bucket index shard is in OMAP collapse.&amp;rdquo;&lt;/p></description></item><item><title>Ceph RGW orphaned multipart uploads: space that disappears from view</title><link>https://www.netdata.cloud/guides/ceph/ceph-rgw-multipart-orphans/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-rgw-multipart-orphans/</guid><description>&lt;h1 id="ceph-rgw-orphaned-multipart-uploads-space-that-disappears-from-view">Ceph RGW orphaned multipart uploads: space that disappears from view&lt;/h1>
&lt;p>Pool capacity climbs toward nearfull but bucket listings and &lt;code>radosgw-admin bucket stats&lt;/code> cannot account for it. The obvious explanations (client growth, snapshots, OMAP bloat) do not fit. This pattern is frequently caused by orphaned multipart upload parts in the RADOS Gateway (RGW) data pool.&lt;/p>
&lt;p>S3 multipart uploads create a manifest, upload parts into the bucket&amp;rsquo;s data pool, then call &lt;code>CompleteMultipartUpload&lt;/code> to assemble the final object. When that final step never happens, when the client retries uploads of the same key, or when RGW aborts an upload incompletely, the individual parts remain in RADOS. They occupy raw space but no live bucket index entry references them.&lt;/p></description></item><item><title>Ceph slow requests (SLOW_OPS): operations blocked past osd_op_complaint_time</title><link>https://www.netdata.cloud/guides/ceph/ceph-slow-requests/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-slow-requests/</guid><description>&lt;h1 id="ceph-slow-requests-slow_ops-operations-blocked-past-osd_op_complaint_time">Ceph slow requests (SLOW_OPS): operations blocked past osd_op_complaint_time&lt;/h1>
&lt;p>&lt;code>ceph_healthcheck_slow_ops&lt;/code> greater than zero means operations on the cluster have crossed &lt;code>osd_op_complaint_time&lt;/code> (default 30s) and are stuck, not merely slow. The &lt;code>SLOW_OPS&lt;/code> health check surfaces them as a warning, and the metric itself is a live gauge pulled from &lt;code>ceph health detail&lt;/code>.&lt;/p>
&lt;p>Slow ops are a symptom of something downstream blocking I/O: a failing disk, a saturated BlueStore RocksDB DB, network timeouts between OSDs, or heavy deep-scrub on HDD during a maintenance window. The right first move is to read where in the pipeline each slow op is stuck before changing any config.&lt;/p></description></item><item><title>Ceph too many PGs per OSD: mon_max_pg_per_osd, peering cost, and sizing</title><link>https://www.netdata.cloud/guides/ceph/ceph-too-many-pgs-per-osd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-too-many-pgs-per-osd/</guid><description>&lt;h1 id="ceph-too-many-pgs-per-osd-mon_max_pg_per_osd-peering-cost-and-sizing">Ceph too many PGs per OSD: mon_max_pg_per_osd, peering cost, and sizing&lt;/h1>
&lt;p>PG count per OSD is one of the few Ceph sizing decisions with direct cost in both directions. Too many PGs per OSD means more memory and CPU spent on peering, recovery scans, and PG log maintenance across more logical units. Too few means CRUSH cannot spread data evenly, producing hotspots on specific OSDs while others sit idle.&lt;/p>
&lt;p>&lt;code>mon_max_pg_per_osd&lt;/code> (default 250&lt;!-- TODO: verify whether default was 200 before Luminous 12.2.10 -->) is the cluster&amp;rsquo;s failsafe against runaway PG counts. When an OSD approaches this number, Ceph raises &lt;code>TOO_MANY_PGS&lt;/code> and blocks new pool creation, &lt;code>pg_num&lt;/code> increases, and replication factor changes. This is a guard rail, not a performance target. The existing cluster keeps serving I/O; only topology changes that would add PGs are blocked.&lt;/p></description></item><item><title>Ceph unfound objects: the cluster cannot locate a surviving copy</title><link>https://www.netdata.cloud/guides/ceph/ceph-unfound-objects/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-unfound-objects/</guid><description>&lt;h1 id="ceph-unfound-objects-the-cluster-cannot-locate-a-surviving-copy">Ceph unfound objects: the cluster cannot locate a surviving copy&lt;/h1>
&lt;p>When a placement group enters &lt;code>recovery_unfound&lt;/code> or &lt;code>backfill_unfound&lt;/code>, the cluster has objects it knows should exist but cannot locate on any OSD that is currently up and in. &lt;code>ceph health detail&lt;/code> surfaces this as the &lt;code>OBJECT_UNFOUND&lt;/code> check, and the cluster-wide gauge &lt;code>ceph_num_objects_unfound&lt;/code> rises above zero. Every known replica or erasure-coded chunk is on an OSD that is down, destroyed, or has not yet been probed.&lt;/p></description></item><item><title>Cerent Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cerent-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cerent-corporation-snmp-traps/</guid><description/></item><item><title>Chateau Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/chateau-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/chateau-systems-inc-snmp-traps/</guid><description/></item><item><title>Chatsworth PDU</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/chatsworth-pdu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/chatsworth-pdu/</guid><description/></item><item><title>Check Point Software Technologies Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/check-point-software-technologies-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/check-point-software-technologies-ltd-snmp-traps/</guid><description/></item><item><title>Checkpoint</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/checkpoint/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/checkpoint/</guid><description/></item><item><title>Checkpoint Licensing</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/licensing-monitoring/checkpoint-licensing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/licensing-monitoring/checkpoint-licensing/</guid><description/></item><item><title>Cherokee International Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cherokee-international-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cherokee-international-corporation-snmp-traps/</guid><description/></item><item><title>Cheyenne Software SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cheyenne-software-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cheyenne-software-snmp-traps/</guid><description/></item><item><title>Chia</title><link>https://www.netdata.cloud/integrations/data-collection/applications/chia/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/chia/</guid><description/></item><item><title>Chia Monitoring</title><link>https://www.netdata.cloud/monitoring-101/chia-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/chia-monitoring/</guid><description>&lt;h2 id="chia-monitoring">Chia Monitoring&lt;/h2>
&lt;h3 id="what-is-chia">What Is Chia?&lt;/h3>
&lt;p>Chia is a blockchain and smart transaction platform designed to be eco-friendly through its innovative use of proof-of-space-and-time. It enables users to participate in the blockchain by utilizing free disk space instead of relying on energy-intensive proof-of-work.&lt;/p>
&lt;h3 id="monitoring-chia-with-netdata">Monitoring Chia With Netdata&lt;/h3>
&lt;p>Monitoring your Chia setup is essential for ensuring optimal performance and troubleshooting issues before they escalate. Netdata offers a Chia monitoring tool through a Prometheus exporter &lt;a href="https://github.com/chia-network/chia-exporter">available here&lt;/a>. Netdata stands out because it can ingest data from any &lt;a href="https://learn.netdata.cloud/docs/collecting-metrics/generic-collecting-metrics/prometheus-endpoint/?utm_source=website&amp;amp;utm_content=monitoring101">Prometheus exporter&lt;/a>, turning this data into automated dashboards and alerts without requiring additional software such as Prometheus servers or Grafana. By using Netdata, you gain powerful insights into your Chia operations with minimal setup.&lt;/p></description></item><item><title>Chippcom SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/chippcom-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/chippcom-snmp-traps/</guid><description/></item><item><title>Chloride SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/chloride-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/chloride-snmp-traps/</guid><description/></item><item><title>Christ Elektronik CLM5IP Monitoring</title><link>https://www.netdata.cloud/monitoring-101/clm5ip-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/clm5ip-monitoring/</guid><description>&lt;h2 id="christ-elektronik-clm5ip-monitoring">Christ Elektronik CLM5IP Monitoring&lt;/h2>
&lt;h3 id="what-is-christ-elektronik-clm5ip">What Is Christ Elektronik CLM5IP?&lt;/h3>
&lt;p>The Christ Elektronik CLM5IP power panel is an advanced instrument designed for monitoring and controlling power systems. It is widely used in industrial and IoT environments for its ability to collect extensive power usage data, ensure efficient energy management, and optimize performance.&lt;/p>
&lt;h3 id="monitoring-christ-elektronik-clm5ip-with-netdata">Monitoring Christ Elektronik CLM5IP With Netdata&lt;/h3>
&lt;p>Monitoring the Christ Elektronik CLM5IP power panel can provide invaluable insights into the energy usage and health of your systems. By using Netdata as a monitoring tool, you can leverage OpenMetrics through a Prometheus exporter, such as the &lt;a href="https://github.com/christmann/clm5ip_exporter/">Christmann CLM5IP Exporter&lt;/a>, to collect metrics effortlessly. Netdata allows you to ingest data from any Prometheus exporter, providing automated dashboards, real-time alerts, and comprehensive analytics, all without the need for a Prometheus server or Grafana. This makes Netdata a versatile and powerful tool for monitoring Christ Elektronik CLM5IP.&lt;/p></description></item><item><title>Christ Elektronik CLM5IP power panel</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/christ-elektronik-clm5ip-power-panel/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/christ-elektronik-clm5ip-power-panel/</guid><description/></item><item><title>Chromatis Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/chromatis-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/chromatis-networks-inc-snmp-traps/</guid><description/></item><item><title>Chronix</title><link>https://www.netdata.cloud/integrations/exporters/chronix/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/chronix/</guid><description/></item><item><title>Chrony</title><link>https://www.netdata.cloud/integrations/data-collection/networking/chrony/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/chrony/</guid><description/></item><item><title>Chrony Monitoring</title><link>https://www.netdata.cloud/monitoring-101/chrony-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/chrony-monitoring/</guid><description>&lt;h2 id="chrony-monitoring">Chrony Monitoring&lt;/h2>
&lt;h3 id="what-is-chrony">What Is Chrony?&lt;/h3>
&lt;p>Chrony is a versatile software suite for maintaining the system clock accuracy on your Linux distributions. Acting as an implementation of the Network Time Protocol (NTP), Chrony ensures your systems are synchronized to accurate time sources, which is crucial for various server operations. It is highly adaptive, providing swift synchronization capabilities, making it particularly beneficial for systems with sporadic network connectivity or those that endure extreme delays.&lt;/p></description></item><item><title>Chrysalis</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/chrysalis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/chrysalis/</guid><description/></item><item><title>Chrysalis Luna HSM</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/chrysalis-luna-hsm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/chrysalis-luna-hsm/</guid><description/></item><item><title>Ciena Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ciena-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ciena-corporation-snmp-traps/</guid><description/></item><item><title>Cilium Agent</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/cilium-agent/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/cilium-agent/</guid><description/></item><item><title>Cilium Agent Monitoring</title><link>https://www.netdata.cloud/monitoring-101/cilium_agent-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/cilium_agent-monitoring/</guid><description>&lt;h2 id="cilium-agent-monitoring">Cilium Agent Monitoring&lt;/h2>
&lt;h3 id="what-is-cilium-agent">What Is Cilium Agent?&lt;/h3>
&lt;p>Cilium is an open-source software package that provides networking, security, and load balancing for cloud-native environments like Kubernetes. The Cilium Agent is a key component within this framework, responsible for enforcing policies and facilitating complex networking requirements with simplicity and transparency.&lt;/p>
&lt;h3 id="monitoring-cilium-agent-with-netdata">Monitoring Cilium Agent With Netdata&lt;/h3>
&lt;p>To monitor Cilium Agent efficiently, Netdata employs an OpenMetrics (Prometheus) exporter. This integration allows Netdata to ingest valuable monitoring data directly from any Prometheus exporter. With Netdata, you don’t need a separate Prometheus server or Grafana setup. The tool provides automated dashboards and real-time alerts, streamlining the process of monitoring Cilium Agent metrics like network security and connectivity.&lt;/p></description></item><item><title>Cilium Operator</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/cilium-operator/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/cilium-operator/</guid><description/></item><item><title>Cilium Operator Monitoring</title><link>https://www.netdata.cloud/monitoring-101/cilium_operator-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/cilium_operator-monitoring/</guid><description>&lt;h2 id="cilium-operator-monitoring">Cilium Operator Monitoring&lt;/h2>
&lt;h3 id="what-is-cilium-operator">What Is Cilium Operator?&lt;/h3>
&lt;p>Cilium Operator is a crucial component for Kubernetes network security management. It enhances Kubernetes networking by providing secure and transparent connections within your cluster. Whether you are a DevOps engineer or a Site Reliability Engineer (SRE), understanding the intricacies of the Cilium Operator and monitoring its performance can significantly enhance the security and efficiency of your deployments.&lt;/p>
&lt;h3 id="monitoring-cilium-operator-with-netdata">Monitoring Cilium Operator With Netdata&lt;/h3>
&lt;p>To effectively monitor Cilium Operator, Netdata uses an openmetrics (Prometheus) exporter. This seamless integration allows Netdata to ingest data from any Prometheus exporter. By doing so, you can access automated dashboards, alerts, and more—all without needing a separate Prometheus server or Grafana. This makes Netdata an ideal Cilium Operator monitoring tool, combining ease of use with powerful capabilities.&lt;/p></description></item><item><title>Cilium Proxy</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/cilium-proxy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/cilium-proxy/</guid><description/></item><item><title>Cilium Proxy Monitoring</title><link>https://www.netdata.cloud/monitoring-101/cilium_proxy-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/cilium_proxy-monitoring/</guid><description>&lt;h2 id="cilium-proxy-monitoring">Cilium Proxy Monitoring&lt;/h2>
&lt;h3 id="what-is-cilium-proxy">What Is Cilium Proxy?&lt;/h3>
&lt;p>Cilium Proxy is a key component in modern network security and performance for Kubernetes environments. It acts as an intermediary that controls and processes requests and responses between different microservices, ensuring secure and efficient communication.&lt;/p>
&lt;h3 id="monitoring-cilium-proxy-with-netdata">Monitoring Cilium Proxy With Netdata&lt;/h3>
&lt;p>Monitoring Cilium Proxy is crucial for maintaining optimal network performance and security. Netdata offers an intuitive solution to monitor Cilium Proxy using an openmetrics (Prometheus) exporter. Netdata can seamlessly ingest data from any Prometheus exporter, providing automated dashboards, alerts, and insightful metrics without the need for a Prometheus server or Grafana. This allows DevOps teams to quickly troubleshoot issues and maintain seamless network operations using Netdata&amp;rsquo;s unparalleled real-time monitoring capabilities.&lt;/p></description></item><item><title>Cirpack SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cirpack-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cirpack-snmp-traps/</guid><description/></item><item><title>Cisco</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco/</guid><description/></item><item><title>Cisco 3850</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-3850/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-3850/</guid><description/></item><item><title>Cisco Access Point</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-access-point/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-access-point/</guid><description/></item><item><title>Cisco ASA</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-asa/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-asa/</guid><description/></item><item><title>Cisco ASR</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-asr/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-asr/</guid><description/></item><item><title>Cisco BGP</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/cisco-bgp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/cisco-bgp/</guid><description/></item><item><title>Cisco Catalyst</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-catalyst/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-catalyst/</guid><description/></item><item><title>Cisco Catalyst WLC</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-catalyst-wlc/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-catalyst-wlc/</guid><description/></item><item><title>Cisco Csr1000V</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-csr1000v/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-csr1000v/</guid><description/></item><item><title>Cisco Firepower</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-firepower/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-firepower/</guid><description/></item><item><title>Cisco Firepower ASA</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-firepower-asa/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-firepower-asa/</guid><description/></item><item><title>Cisco ICM</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-icm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-icm/</guid><description/></item><item><title>Cisco Ironport Email</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-ironport-email/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-ironport-email/</guid><description/></item><item><title>Cisco ISE</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-ise/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-ise/</guid><description/></item><item><title>Cisco ISR</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-isr/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-isr/</guid><description/></item><item><title>Cisco ISR 4431</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-isr-4431/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-isr-4431/</guid><description/></item><item><title>Cisco Legacy WLC</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-legacy-wlc/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-legacy-wlc/</guid><description/></item><item><title>Cisco Licensing</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/licensing-monitoring/cisco-licensing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/licensing-monitoring/cisco-licensing/</guid><description/></item><item><title>Cisco Load Balancer</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-load-balancer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-load-balancer/</guid><description/></item><item><title>Cisco Meraki (Cloud Controller)</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-meraki-cloud-controller/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-meraki-cloud-controller/</guid><description/></item><item><title>Cisco NCS</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-ncs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-ncs/</guid><description/></item><item><title>Cisco Nexus</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-nexus/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-nexus/</guid><description/></item><item><title>Cisco SB</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-sb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-sb/</guid><description/></item><item><title>Cisco UC Virtual Machine</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-uc-virtual-machine/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-uc-virtual-machine/</guid><description/></item><item><title>Cisco UCS</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-ucs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-ucs/</guid><description/></item><item><title>Cisco WAN Optimizer</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-wan-optimizer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cisco-wan-optimizer/</guid><description/></item><item><title>Ciscosystems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ciscosystems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ciscosystems-snmp-traps/</guid><description/></item><item><title>Citrix</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/citrix/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/citrix/</guid><description/></item><item><title>Citrix Netscaler</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/citrix-netscaler/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/citrix-netscaler/</guid><description/></item><item><title>Citrix Netscaler SDX</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/citrix-netscaler-sdx/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/citrix-netscaler-sdx/</guid><description/></item><item><title>Citrix Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/citrix-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/citrix-systems-inc-snmp-traps/</guid><description/></item><item><title>City Com B.V. SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/city-com-b.v.-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/city-com-b.v.-snmp-traps/</guid><description/></item><item><title>ClamAV daemon</title><link>https://www.netdata.cloud/integrations/data-collection/applications/clamav-daemon/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/clamav-daemon/</guid><description/></item><item><title>ClamAV Daemon Monitoring</title><link>https://www.netdata.cloud/monitoring-101/clamd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/clamd-monitoring/</guid><description>&lt;h2 id="clamav-daemon-monitoring">ClamAV Daemon Monitoring&lt;/h2>
&lt;h3 id="what-is-clamav-daemon">What Is ClamAV Daemon?&lt;/h3>
&lt;p>ClamAV Daemon is part of the ClamAV antivirus software toolkit, developed for detecting threats and viruses on various platforms. It plays a crucial role in ensuring the security of IT infrastructures by scanning files and attachments in real-time to prevent malware and other malicious threats from affecting your systems.&lt;/p>
&lt;h3 id="monitoring-clamav-daemon-with-netdata">Monitoring ClamAV Daemon With Netdata&lt;/h3>
&lt;p>Monitoring the ClamAV Daemon is seamless with &lt;a href="https://www.netdata.cloud">Netdata&lt;/a>, as it utilizes an openmetrics (&lt;em>Prometheus&lt;/em>) exporter to fetch real-time metrics. With Netdata, you can ingest data from any Prometheus exporter providing an automated setup with dashboards, alerts, and in-depth analytics, all without the need for setting up a Prometheus server or Grafana.&lt;/p></description></item><item><title>Clamscan results</title><link>https://www.netdata.cloud/integrations/data-collection/applications/clamscan-results/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/clamscan-results/</guid><description/></item><item><title>Clamscan Results Monitoring</title><link>https://www.netdata.cloud/monitoring-101/clamscan-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/clamscan-monitoring/</guid><description>&lt;h2 id="clamscan-results-monitoring">Clamscan Results Monitoring&lt;/h2>
&lt;h3 id="what-is-clamscan-results-monitoring">What Is Clamscan Results Monitoring?&lt;/h3>
&lt;p>Clamscan is a command-line utility used in conjunction with ClamAV for detecting malware on a system. Monitoring Clamscan results is crucial to ensure the effectiveness of your anti-malware efforts, track performance metrics, and maintain the security of your network. By keeping an eye on the data provided by Clamscan, IT admins, DevOps, and security personnel can act promptly in response to vulnerabilities or infections.&lt;/p></description></item><item><title>Clarent Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/clarent-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/clarent-corporation-snmp-traps/</guid><description/></item><item><title>Clash</title><link>https://www.netdata.cloud/integrations/data-collection/networking/clash/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/clash/</guid><description/></item><item><title>Clash Monitoring</title><link>https://www.netdata.cloud/monitoring-101/clash-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/clash-monitoring/</guid><description>&lt;h2 id="clash-monitoring">Clash Monitoring&lt;/h2>
&lt;h3 id="what-is-clash">What Is Clash?&lt;/h3>
&lt;p>Clash is a versatile and powerful proxy server designed for efficient network traffic management. It serves as an essential component in managing and troubleshooting network performance in today&amp;rsquo;s complex IT environments. With various features designed to seamlessly handle traffic routing, Clash ensures that data flows are optimized, which is vital for maintaining network health and performance.&lt;/p>
&lt;h3 id="monitoring-clash-with-netdata">Monitoring Clash With Netdata&lt;/h3>
&lt;p>To effectively monitor Clash, Netdata offers a seamless integration using an openmetrics (prometheus) exporter. This integration allows technical teams—be it DevOps, SREs, Developers, IT admins, or engineers—to gain comprehensive insights into the performance of their Clash instances. Netdata can ingest data from any Prometheus exporter, providing users with automated dashboards, real-time alerts, and more, without the necessity of having a dedicated Prometheus server or Grafana setup. This makes it easier for users looking for a $name monitoring tool to implement and manage network performance monitoring efficiently.&lt;/p></description></item><item><title>Classifiers</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/classifiers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/classifiers/</guid><description/></item><item><title>Clavister AB SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/clavister-ab-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/clavister-ab-snmp-traps/</guid><description/></item><item><title>Clickarrray Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/clickarrray-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/clickarrray-networks-inc-snmp-traps/</guid><description/></item><item><title>ClickHouse</title><link>https://www.netdata.cloud/integrations/data-collection/databases/clickhouse/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/clickhouse/</guid><description/></item><item><title>ClickHouse active part count growing: reading MaxPartCountForPartition before it pages</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-active-part-count-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-active-part-count-growing/</guid><description>&lt;h1 id="clickhouse-active-part-count-growing-reading-maxpartcountforpartition-before-it-pages">ClickHouse active part count growing: reading MaxPartCountForPartition before it pages&lt;/h1>
&lt;p>Rising &lt;code>MaxPartCountForPartition&lt;/code> is the leading indicator for the most common ClickHouse production failure: parts accumulating faster than background merges can consolidate them. A single partition crossing 500 active parts means you have hours, not days, before inserts delay and eventually fail with &lt;code>TOO_MANY_PARTS&lt;/code>.&lt;/p>
&lt;p>The thresholds are per-partition. A table with ten partitions at fifty parts each is healthy; one partition at 950 parts is approaching throttling. Projections create hidden parts inside the same table that count toward the same limits. Materialized views route inserts to separate target tables that can hit their own limits independently.&lt;/p></description></item><item><title>ClickHouse ALTER UPDATE/DELETE overuse: why mutations are not row updates</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-alter-update-delete-overuse/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-alter-update-delete-overuse/</guid><description>&lt;h1 id="clickhouse-alter-updatedelete-overuse-why-mutations-are-not-row-updates">ClickHouse ALTER UPDATE/DELETE overuse: why mutations are not row updates&lt;/h1>
&lt;p>Your inserts are slowing down. &lt;code>system.merges&lt;/code> shows long-running background tasks with &lt;code>is_mutation = 1&lt;/code>, while the active part count climbs toward the &lt;code>parts_to_delay_insert&lt;/code> threshold. Write latency rises and &lt;code>DelayedInserts&lt;/code> increases even though memory, disk, and CPU are not exhausted. The culprit is usually an application treating ClickHouse as an OLTP store, issuing &lt;code>ALTER TABLE UPDATE&lt;/code> or &lt;code>ALTER TABLE DELETE&lt;/code> as routine operations.&lt;/p></description></item><item><title>ClickHouse async inserts: when async_insert fixes too-many-parts and when it hides it</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-async-inserts-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-async-inserts-tuning/</guid><description>&lt;h1 id="clickhouse-async-inserts-when-async_insert-fixes-too-many-parts-and-when-it-hides-it">ClickHouse async inserts: when async_insert fixes too-many-parts and when it hides it&lt;/h1>
&lt;p>Async inserts move batching responsibility from the client to the ClickHouse server. The server buffers rows and flushes them as larger blocks instead of persisting every INSERT as a separate part. When the root cause is an unbatchable client, this stops the small-inserts anti-pattern from flooding the merge pool. When the root cause is high-cardinality partitioning or sustained over-ingestion, async inserts relocate the crisis into an in-memory buffer that loses data on crash and hides backpressure from the application.&lt;/p></description></item><item><title>ClickHouse authentication failures: system.session_log, brute force, and credential drift</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-authentication-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-authentication-failures/</guid><description>&lt;h1 id="clickhouse-authentication-failures-systemsession_log-brute-force-and-credential-drift">ClickHouse authentication failures: system.session_log, brute force, and credential drift&lt;/h1>
&lt;p>You notice a spike in failed connection attempts to ClickHouse: a security scanner flags repeated TCP 9000 probes, or an application logs connection timeouts after a secrets rotation. ClickHouse exposes authentication events through &lt;code>system.session_log&lt;/code>, but only if the feature is enabled. Without it, fallback to server error logs and &lt;code>system.query_log&lt;/code> exceptions.&lt;/p>
&lt;p>The failures split into two patterns. Malicious: brute-force or credential-scanning campaigns against exposed TCP 9000 or HTTP 8123. Operational drift: a rotated password not updated in a client config, or a deployment shipping an old connection string. Distinguish them fast. Block an external attacker at the network layer; fix credential drift on the client.&lt;/p></description></item><item><title>ClickHouse background pool saturation: when merges and mutations starve</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-background-pool-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-background-pool-saturation/</guid><description>&lt;h1 id="clickhouse-background-pool-saturation-when-merges-and-mutations-starve">ClickHouse background pool saturation: when merges and mutations starve&lt;/h1>
&lt;p>Insert latency is climbing and &lt;code>DelayedInserts&lt;/code> is ticking up. &lt;code>system.merges&lt;/code> shows every slot occupied, yet &lt;code>system.parts&lt;/code> keeps growing. When the background merge and mutation pool saturates, new merges queue instead of starting, parts accumulate, and the distance to insert rejections shrinks fast. Distinguish true thread starvation from I/O-bound stalls, identify when mutations are the culprit, and relieve pressure before inserts fail.&lt;/p></description></item><item><title>ClickHouse cannot connect to ZooKeeper/Keeper: diagnosing the coordination layer</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-keeper-connection-lost/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-keeper-connection-lost/</guid><description>&lt;h1 id="clickhouse-cannot-connect-to-zookeeperkeeper-diagnosing-the-coordination-layer">ClickHouse cannot connect to ZooKeeper/Keeper: diagnosing the coordination layer&lt;/h1>
&lt;p>If &lt;code>SELECT * FROM system.zookeeper WHERE path = '/'&lt;/code> fails, replicated tables flip to readonly, or &lt;code>ON CLUSTER&lt;/code> DDL hangs, the coordination layer is broken. The server stays up and &lt;code>GET /ping&lt;/code> returns &lt;code>Ok.&lt;/code>, so liveness checks miss the problem. Partial degradation can escalate to a write outage if replicated tables cannot re-establish sessions.&lt;/p>
&lt;p>ClickHouse uses the coordination service for ReplicatedMergeTree leader election, replication log queues, insert deduplication, and distributed DDL. ClickHouse Keeper typically listens on port 9181; external ZooKeeper ensembles listen on 2181. The diagnostic path differs slightly, but the symptom is the same: ClickHouse cannot reliably complete coordination operations.&lt;/p></description></item><item><title>ClickHouse checksum mismatch and broken parts: detecting data corruption</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-data-corruption-checksum-mismatch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-data-corruption-checksum-mismatch/</guid><description>&lt;h1 id="clickhouse-checksum-mismatch-and-broken-parts-detecting-data-corruption">ClickHouse checksum mismatch and broken parts: detecting data corruption&lt;/h1>
&lt;p>ClickHouse logs showing &lt;code>Checksum doesn't match&lt;/code>, &lt;code>Broken part&lt;/code>, or similar errors indicate data corruption. Affected parts move to &lt;code>system.detached_parts&lt;/code>. Queries may throw exceptions or return partial results. On replicated clusters, a replica with corrupt parts may lag because it cannot validate fetched parts. Corruption does not self-resolve. You must quarantine the bad part, identify the root cause, and rebuild the data from a healthy source or backup.&lt;/p></description></item><item><title>ClickHouse client connections climbing: TCP 9000, HTTP 8123, and connection leaks</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-connection-count-climbing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-connection-count-climbing/</guid><description>&lt;h1 id="clickhouse-client-connections-climbing-tcp-9000-http-8123-and-connection-leaks">ClickHouse client connections climbing: TCP 9000, HTTP 8123, and connection leaks&lt;/h1>
&lt;p>&lt;code>TCPConnection&lt;/code> or &lt;code>HTTPConnection&lt;/code> climbing on a ClickHouse node means each active socket consumes a file descriptor. When growth is uncorrelated with query throughput, it is usually a connection leak or pool misconfiguration rather than healthy concurrency. If the count approaches &lt;code>max_connections&lt;/code>, the server rejects new client connections. If the Linux &lt;code>nofile&lt;/code> limit is reached first, queries and merges fail with &amp;ldquo;too many open files.&amp;rdquo;&lt;/p></description></item><item><title>ClickHouse DB::Exception: Too many parts - causes and fixes</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-too-many-parts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-too-many-parts/</guid><description>&lt;h1 id="clickhouse-dbexception-too-many-parts---causes-and-fixes">ClickHouse DB::Exception: Too many parts - causes and fixes&lt;/h1>
&lt;p>When logs show &lt;code>DB::Exception: Too many parts (N). Merges are processing significantly slower than inserts&lt;/code> (exception code 252), INSERTs that succeeded yesterday are now failing. Data backs up in ingestors, message queues, or client buffers, and every retry adds load to a system already drowning in small files.&lt;/p>
&lt;p>At least one partition in a MergeTree table has exceeded the hard limit for active data parts, and background merges cannot consolidate them fast enough. Do not restart the server or raise the limit indefinitely. Reduce the part creation rate, remove blockers from the merge pipeline, and give background threads runway to catch up.&lt;/p></description></item><item><title>ClickHouse DelayedInserts climbing: the warning before too-many-parts</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-delayed-inserts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-delayed-inserts/</guid><description>&lt;h1 id="clickhouse-delayedinserts-climbing-the-warning-before-too-many-parts">ClickHouse DelayedInserts climbing: the warning before too-many-parts&lt;/h1>
&lt;p>Insert latency is climbing and &lt;code>system.events.DelayedInserts&lt;/code> is no longer flat. ClickHouse is sleeping during INSERT because at least one partition has crossed &lt;code>parts_to_delay_insert&lt;/code>. The database still accepts writes, but injects a sleep before each insert commits. This is the warning window before hard failure. If the merge backlog is not resolved, &lt;code>DelayedInserts&lt;/code> climbs until &lt;code>RejectedInserts&lt;/code> starts ticking and clients receive &lt;code>DB::Exception: Too many parts&lt;/code>.&lt;/p></description></item><item><title>ClickHouse detached parts piling up: reading system.detached_parts and reclaiming space</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-detached-parts-piling-up/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-detached-parts-piling-up/</guid><description>&lt;h1 id="clickhouse-detached-parts-piling-up-reading-systemdetached_parts-and-reclaiming-space">ClickHouse detached parts piling up: reading system.detached_parts and reclaiming space&lt;/h1>
&lt;p>Disk usage is climbing and &lt;code>system.disks&lt;/code> shows &lt;code>unreserved_space&lt;/code> shrinking, yet active part counts in &lt;code>system.parts&lt;/code> look normal. Queries are not failing, merges appear healthy, and there is no insert rejection storm. The hidden consumer is often a growing pile of &lt;strong>detached parts&lt;/strong>: directories that ClickHouse removed from the active dataset but left on the filesystem. Unlike active parts, detached parts do not participate in queries or merges, and ClickHouse never deletes them automatically. Standard monitoring that only watches &lt;code>system.parts&lt;/code> with &lt;code>active = 1&lt;/code> misses them entirely. This article explains how to read &lt;code>system.detached_parts&lt;/code>, interpret the &lt;code>reason&lt;/code> column, decide whether to reattach or discard the data, and prevent silent recurrence.&lt;/p></description></item><item><title>ClickHouse disk space collapse: why merges need free space and how the spiral starts</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-disk-space-collapse/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-disk-space-collapse/</guid><description>&lt;h1 id="clickhouse-disk-space-collapse-why-merges-need-free-space-and-how-the-spiral-starts">ClickHouse disk space collapse: why merges need free space and how the spiral starts&lt;/h1>
&lt;p>Disk usage climbs from 75% to 85% over a week, then hits 90%. A few hours later ClickHouse rejects inserts, part counts spike, and the volume is at 98%. The system did not simply run out of space. It entered a self-reinforcing spiral where background merges, the mechanism that reclaims space, stalled because they needed free space to work.&lt;/p></description></item><item><title>ClickHouse disk space monitoring: free_space, unreserved_space, and the 80% target</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-disk-space-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-disk-space-monitoring/</guid><description>&lt;h1 id="clickhouse-disk-space-monitoring-free_space-unreserved_space-and-the-80-target">ClickHouse disk space monitoring: free_space, unreserved_space, and the 80% target&lt;/h1>
&lt;p>ClickHouse disk space is not a simple capacity gauge. Merges write new parts before deleting old ones, so free space is an operational dependency of the storage engine. Insufficient disk does not just mean &amp;ldquo;running low&amp;rdquo;: background merges stall, parts accumulate, and the system enters a self-reinforcing death spiral.&lt;/p>
&lt;p>This guide covers the &lt;code>system.disks&lt;/code> metrics, why &lt;code>unreserved_space&lt;/code> matters more than &lt;code>free_space&lt;/code>, and operational targets that keep merges alive.&lt;/p></description></item><item><title>ClickHouse distributed DDL stuck: ON CLUSTER queries that never finish</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-distributed-ddl-stuck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-distributed-ddl-stuck/</guid><description>&lt;h1 id="clickhouse-distributed-ddl-stuck-on-cluster-queries-that-never-finish">ClickHouse distributed DDL stuck: ON CLUSTER queries that never finish&lt;/h1>
&lt;p>&lt;code>ON CLUSTER&lt;/code> DDL is not atomic. The initiator writes a task into a shared queue in ZooKeeper or ClickHouse Keeper under &lt;code>/clickhouse/task_queue/ddl/&lt;/code>. Each node pulls entries in order, executes the DDL locally, and reports status back. One slow, down, or misconfigured node can leave the entry unfinished forever, blocking all later DDL on that node and creating schema drift that is invisible until a replica rejects an insert or a subsequent ALTER fails.&lt;/p></description></item><item><title>ClickHouse distributed query amplification: one coordinator, many shard subqueries</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-distributed-query-amplification/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-distributed-query-amplification/</guid><description>&lt;h1 id="clickhouse-distributed-query-amplification-one-coordinator-many-shard-subqueries">ClickHouse distributed query amplification: one coordinator, many shard subqueries&lt;/h1>
&lt;p>A SELECT against a distributed table registers one entry in &lt;code>system.processes&lt;/code> on the coordinator. That query fans out subqueries to every addressed shard, waits for intermediate results, then merges them locally. True CPU, memory, and network load is multiplied by the shard count, and the slowest participating shard dictates latency.&lt;/p>
&lt;p>This amplification is invisible to cluster-average metrics. A dashboard averaging CPU or query latency across all nodes can look healthy while one shard is saturated, because the coordinator and healthy shards mask the straggler. Operators often misdiagnose the slowdown as a generalized cluster problem or a coordinator bottleneck, when the root cause is usually a single overloaded shard, a missing sharding key filter, or an expensive GLOBAL operator that broadcasts data across the network.&lt;/p></description></item><item><title>ClickHouse full table scan: partition pruning failures and the primary key</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-full-table-scan-no-partition-pruning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-full-table-scan-no-partition-pruning/</guid><description>&lt;h1 id="clickhouse-full-table-scan-partition-pruning-failures-and-the-primary-key">ClickHouse full table scan: partition pruning failures and the primary key&lt;/h1>
&lt;p>A query that returned in milliseconds yesterday now takes minutes and saturates CPU. Check whether it is reading every row before you add vCPUs or shards. A missing filter on the partition key or a mismatch with the primary key prefix forces ClickHouse to open every part and scan every granule. Symptoms are runaway &lt;code>read_rows&lt;/code> counts, tail latency spikes, and sustained CPU decompression that hardware scaling cannot fix.&lt;/p></description></item><item><title>ClickHouse insert latency rising: the leading indicator of write-pipeline trouble</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-insert-latency-rising/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-insert-latency-rising/</guid><description>&lt;h1 id="clickhouse-insert-latency-rising-the-leading-indicator-of-write-pipeline-trouble">ClickHouse insert latency rising: the leading indicator of write-pipeline trouble&lt;/h1>
&lt;p>Your ClickHouse inserts are taking longer. A query that committed in 200 ms last week is now taking 5 seconds, then 15, then 30. In most databases this signals slow disks or lock contention. In ClickHouse, sustained insert latency is the earliest operational signal that the write pipeline is congesting. It precedes &lt;code>DelayedInserts&lt;/code>, part-count alerts, and the hard stop of &lt;code>RejectedInserts&lt;/code> by minutes to hours. Wait for the error and the merge debt is already severe.&lt;/p></description></item><item><title>ClickHouse Keeper latency high: the early warning before sessions expire</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-keeper-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-keeper-latency-high/</guid><description>&lt;h1 id="clickhouse-keeper-latency-high-the-early-warning-before-sessions-expire">ClickHouse Keeper latency high: the early warning before sessions expire&lt;/h1>
&lt;p>INSERTs to replicated tables slow down, &lt;code>ON CLUSTER&lt;/code> DDL hangs, and the replication queue grows on followers. &lt;code>SELECT 1&lt;/code> and HTTP &lt;code>/ping&lt;/code> stay healthy, and non-replicated tables are fine. The culprit is usually the coordination service, not ClickHouse itself.&lt;/p>
&lt;p>Rising ZooKeeper or ClickHouse Keeper operation latency is a leading indicator. Replicated inserts, replication log updates, and distributed DDL all round-trip through Keeper. Because Keeper writes its transaction log synchronously, disk I/O on the Keeper node is the most common bottleneck. A degraded-but-connected coordination service is worse than a hard partition: it silently slows every replicated operation until sessions start expiring and replicas flip to readonly. This article explains how to read the early signals, isolate the cause, and fix it before sessions expire.&lt;/p></description></item><item><title>ClickHouse Keeper saturation spiral: too many tables, DDL storms, and cluster freeze</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-keeper-saturation-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-keeper-saturation-spiral/</guid><description>&lt;h1 id="clickhouse-keeper-saturation-spiral-too-many-tables-ddl-storms-and-cluster-freeze">ClickHouse Keeper saturation spiral: too many tables, DDL storms, and cluster freeze&lt;/h1>
&lt;p>INSERTs fail. Replicated tables flip read-only. &lt;code>ON CLUSTER&lt;/code> DDL hangs. &lt;code>curl http://localhost:8123/ping&lt;/code> still returns &lt;code>Ok.&lt;/code> and &lt;code>SELECT 1&lt;/code> still works. This is Keeper saturation: the coordination layer is choking while liveness probes give false confidence.&lt;/p>
&lt;p>The spiral starts when ZooKeeper or ClickHouse Keeper cannot keep up with metadata load. Every replicated table registers znodes and watches. Every DDL operation adds more. Coordination latency climbs until heartbeat traffic cannot complete within the negotiated session timeout; sessions expire, replicas become read-only, and writes fail. Reconnecting nodes then trigger a thundering herd.&lt;/p></description></item><item><title>ClickHouse killed by the OOM killer: RSS, max_server_memory_usage, and cgroup limits</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-oom-killed-by-kernel/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-oom-killed-by-kernel/</guid><description>&lt;h1 id="clickhouse-killed-by-the-oom-killer-rss-max_server_memory_usage-and-cgroup-limits">ClickHouse killed by the OOM killer: RSS, max_server_memory_usage, and cgroup limits&lt;/h1>
&lt;p>You restart a pod and &lt;code>kubectl describe pod&lt;/code> shows &lt;code>Reason: OOMKilled&lt;/code> with exit code 137. Inside ClickHouse, &lt;code>MemoryTracking&lt;/code> sits well below &lt;code>max_server_memory_usage&lt;/code>, and &lt;code>system.text_log&lt;/code> shows no warning. The process is gone, merges are dead, and replication queues are backing up.&lt;/p>
&lt;p>The Linux OOM killer targets RSS, not ClickHouse&amp;rsquo;s internal &lt;code>MemoryTracking&lt;/code>. Untracked allocations, jemalloc arena fragmentation, and cgroup accounting quirks create a persistent gap between what ClickHouse thinks it is using and what the kernel sees. In containerized environments, the cgroup OOM killer can evict the pod before ClickHouse ever triggers its own server-wide limit.&lt;/p></description></item><item><title>ClickHouse long-running queries: finding and killing the resource hog</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-long-running-queries/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-long-running-queries/</guid><description>&lt;h1 id="clickhouse-long-running-queries-finding-and-killing-the-resource-hog">ClickHouse long-running queries: finding and killing the resource hog&lt;/h1>
&lt;p>A query that should finish in seconds is still running after twenty minutes. Memory on the ClickHouse node is climbing, query latency has doubled, and you suspect a single query is holding resources it will never release. In ClickHouse, a long-running query can be a legitimate analytical job crunching terabytes, a Cartesian JOIN exploding in memory, or a GROUP BY that has spilled to disk and slowed to a crawl. Telling the difference determines whether you kill it or let it finish.&lt;/p></description></item><item><title>ClickHouse mark cache and uncompressed cache: reading low hit rates</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-cache-hit-rate-low/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-cache-hit-rate-low/</guid><description>&lt;h1 id="clickhouse-mark-cache-and-uncompressed-cache-reading-low-hit-rates">ClickHouse mark cache and uncompressed cache: reading low hit rates&lt;/h1>
&lt;p>ClickHouse query latency spikes after a restart, or monitoring shows a sustained drop in mark cache hit rate. Before increasing &lt;code>mark_cache_size&lt;/code>, determine whether you are seeing normal warmup or a cache that is too small for the working set.&lt;/p>
&lt;p>The mark cache stores primary key index granule positions. The uncompressed cache stores decompressed column blocks. Low hit rates in each produce different symptoms. This guide covers how to read the metrics, distinguish warmup from real problems, and decide when tuning is warranted.&lt;/p></description></item><item><title>ClickHouse Memory limit (for query) exceeded: per-query limits and GROUP BY/JOIN blowups</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-memory-limit-for-query-exceeded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-memory-limit-for-query-exceeded/</guid><description>&lt;h1 id="clickhouse-memory-limit-for-query-exceeded-per-query-limits-and-group-byjoin-blowups">ClickHouse Memory limit (for query) exceeded: per-query limits and GROUP BY/JOIN blowups&lt;/h1>
&lt;p>&lt;code>Code: 241. DB::Exception: Memory limit (for query) exceeded&lt;/code> means a single query&amp;rsquo;s allocations breached the &lt;code>max_memory_usage&lt;/code> ceiling. ClickHouse tracks memory in a hierarchy: server-wide, per-user, and per-query. The server kills the query to protect the rest of the workload.&lt;/p>
&lt;p>This differs from &lt;code>Memory limit (total) exceeded&lt;/code> (server-level pressure) and &lt;code>Memory limit (for user) exceeded&lt;/code> (profile-level pressure). A query-level breach usually stems from one of three patterns: a high-cardinality &lt;code>GROUP BY&lt;/code>, an unbounded &lt;code>DISTINCT&lt;/code>, or a &lt;code>JOIN&lt;/code> that materializes more rows than expected.&lt;/p></description></item><item><title>ClickHouse Memory limit (total) exceeded - server-wide memory pressure and fixes</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-memory-limit-total-exceeded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-memory-limit-total-exceeded/</guid><description>&lt;h1 id="clickhouse-memory-limit-total-exceeded---server-wide-memory-pressure-and-fixes">ClickHouse Memory limit (total) exceeded - server-wide memory pressure and fixes&lt;/h1>
&lt;p>&lt;code>Code: 241. DB::Exception: Memory limit (total) exceeded: would use X bytes, current RSS Y, maximum Z.&lt;/code> is the server-level cap, not a per-query limit. When ClickHouse&amp;rsquo;s &lt;code>MemoryTracking&lt;/code> hits &lt;code>max_server_memory_usage&lt;/code> (default 90% of physical RAM), the server kills the heaviest running queries to protect the process. New and existing queries fail until memory drops. Find the largest memory consumer and stop it before the OOM killer does.&lt;/p></description></item><item><title>ClickHouse memory pressure death spiral: runaway queries, retries, and OOM</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-memory-pressure-death-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-memory-pressure-death-spiral/</guid><description>&lt;h1 id="clickhouse-memory-pressure-death-spiral-runaway-queries-retries-and-oom">ClickHouse memory pressure death spiral: runaway queries, retries, and OOM&lt;/h1>
&lt;p>&lt;code>MEMORY_LIMIT_EXCEEDED&lt;/code> errors climb in the query log. Queries that normally finish in seconds now take minutes or are killed outright. The ClickHouse process is near its memory limit, but killing the heaviest query only frees capacity for a moment before another query is killed. If the application retries immediately, pressure never drops. With spill-to-disk enabled, the bottleneck shifts to disk I/O, starving background merges and slowing the whole system.&lt;/p></description></item><item><title>ClickHouse MemoryTracking vs MemoryResident: reading the memory gap correctly</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-memory-tracking-vs-rss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-memory-tracking-vs-rss/</guid><description>&lt;h1 id="clickhouse-memorytracking-vs-memoryresident-reading-the-memory-gap-correctly">ClickHouse MemoryTracking vs MemoryResident: reading the memory gap correctly&lt;/h1>
&lt;p>You finish a large batch query, open monitoring, and see ClickHouse MemoryTracking drop by 40 GB while MemoryResident barely moves. Or you watch RSS climb for hours after a restart while MemoryTracking tracks the rise steadily. Neither pattern indicates a leak. The gap between ClickHouse&amp;rsquo;s internal ledger and OS resident set size is normal: the server accounts for memory synchronously while jemalloc retains pages for reuse.&lt;/p></description></item><item><title>ClickHouse merge death spiral: when parts accumulate faster than merges consolidate</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-merge-death-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-merge-death-spiral/</guid><description>&lt;h1 id="clickhouse-merge-death-spiral-when-parts-accumulate-faster-than-merges-consolidate">ClickHouse merge death spiral: when parts accumulate faster than merges consolidate&lt;/h1>
&lt;p>Insert latency climbs. Application logs show ClickHouse throttling writes. Eventually inserts fail with &lt;code>Too many parts&lt;/code>. Disk usage rises even though ingestion volume is flat. The cluster is up but refusing writes.&lt;/p>
&lt;p>This is the merge death spiral: a self-reinforcing loop where parts accumulate faster than background merges consolidate them. Every INSERT creates immutable on-disk parts. Background merge threads combine smaller parts into larger ones to keep query performance healthy and part counts low. When insert pressure exceeds merge throughput, the backlog grows, merge overhead increases, and the system chokes on its own structure.&lt;/p></description></item><item><title>ClickHouse merge duration climbing: the leading indicator of part explosion</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-merge-duration-climbing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-merge-duration-climbing/</guid><description>&lt;h1 id="clickhouse-merge-duration-climbing-the-leading-indicator-of-part-explosion">ClickHouse merge duration climbing: the leading indicator of part explosion&lt;/h1>
&lt;p>&lt;code>system.merges&lt;/code> shows elapsed times in hours. Your dashboards show P99 merge duration climbing over the past 48 hours. Rising merge duration is the earliest signal that your cluster is heading toward a part-count crisis, typically 1 to 3 days before inserts throttle or fail entirely. &lt;!-- TODO: verify 1-3 day lead time generalizes across ingest patterns and cluster sizes -->&lt;/p></description></item><item><title>ClickHouse merges not keeping up: diagnosing a stalled or starved merge pool</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-merges-not-keeping-up/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-merges-not-keeping-up/</guid><description>&lt;h1 id="clickhouse-merges-not-keeping-up-diagnosing-a-stalled-or-starved-merge-pool">ClickHouse merges not keeping up: diagnosing a stalled or starved merge pool&lt;/h1>
&lt;p>When insert latency climbs and &lt;code>system.merges&lt;/code> is empty while active parts grow, the background merge pool is likely stalled or starved. ClickHouse relies on background merges to consolidate immutable parts after each INSERT. Without merges, parts accumulate: query scans open more files, memory pressure shifts to the mark cache and file descriptor tables, and the system approaches the &lt;code>TOO_MANY_PARTS&lt;/code> threshold.&lt;/p></description></item><item><title>ClickHouse Monitoring</title><link>https://www.netdata.cloud/monitoring-101/clickhouse-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/clickhouse-monitoring/</guid><description>&lt;h2 id="clickhouse-monitoring">ClickHouse Monitoring&lt;/h2>
&lt;h3 id="what-is-clickhouse">What Is ClickHouse?&lt;/h3>
&lt;p>&lt;a href="https://clickhouse.com/">ClickHouse&lt;/a> is a columnar database management system (DBMS) known for its high performance in managing online analytical processing (OLAP) queries. It excels in processing large volumes of data, making it a popular choice for data analytics tasks. Understanding and monitoring the performance of ClickHouse can ensure that your systems are running efficiently and reliably.&lt;/p>
&lt;h3 id="monitoring-clickhouse-with-netdata">Monitoring ClickHouse With Netdata&lt;/h3>
&lt;p>Netdata provides a comprehensive &amp;ldquo;ClickHouse monitoring tool&amp;rdquo; that allows you to monitor ClickHouse instances in real-time, with insightful visuals and metabolism into the system&amp;rsquo;s performances. You can easily integrate Netdata with ClickHouse by using the &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/clickhouse/">go.d.plugin&lt;/a>, enabling automatic discovery and effortless setup.&lt;/p></description></item><item><title>ClickHouse monitoring checklist: the signals every production cluster needs</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-monitoring-checklist/</guid><description>&lt;h1 id="clickhouse-monitoring-checklist-the-signals-every-production-cluster-needs">ClickHouse monitoring checklist: the signals every production cluster needs&lt;/h1>
&lt;p>ClickHouse failures usually begin as storage-structure debt: immutable parts accumulate faster than merges consolidate them, coordination sessions expire, or disk space drops below the threshold merges need to complete. Query latency degrades only after the crisis is hours old.&lt;/p>
&lt;p>This checklist groups monitoring signals into four maturity levels. Level 1 is the minimum viable instrumentation to avoid data loss and unavailability. Each subsequent level adds leading indicators that catch part accumulation, replication divergence, and memory pressure while they are still reversible. The signals are drawn from ClickHouse system tables and OS-level metrics. They apply to single-node, sharded, and replicated setups; replicated tables add ZooKeeper/Keeper signals that belong in Level 2.&lt;/p></description></item><item><title>ClickHouse monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-monitoring-maturity-model/</guid><description>&lt;h1 id="clickhouse-monitoring-maturity-model-from-survival-to-expert">ClickHouse monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most production ClickHouse incidents are not mysterious. They are predictable storage-structure or coordination failures that better monitoring would have surfaced hours earlier. If you are running ClickHouse at scale, you need to know whether your observability is actually catching the failure modes that matter, or just proving that the process is running.&lt;/p>
&lt;p>This maturity model is a diagnostic mirror, not a trophy case. Use it to audit your dashboards, tune alert noise, and decide what to instrument next. Each level builds on the last. Skipping levels leaves predictable gaps: teams with beautiful query-latency dashboards still get surprised by merge death spirals because they never instrumented part counts per partition.&lt;/p></description></item><item><title>ClickHouse mutation stuck: parts_to_do not decreasing and how to recover</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-mutation-stuck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-mutation-stuck/</guid><description>&lt;h1 id="clickhouse-mutation-stuck-parts_to_do-not-decreasing-and-how-to-recover">ClickHouse mutation stuck: parts_to_do not decreasing and how to recover&lt;/h1>
&lt;p>Check &lt;code>system.mutations&lt;/code> and find a mutation active for hours. &lt;code>parts_to_do&lt;/code> has not moved in thirty minutes and parts are accumulating. Queries return, but insert latency climbs. On replicated tables, each replica processes mutations independently, so a stall on one node creates silent divergence while the rest of the cluster appears healthy.&lt;/p>
&lt;p>Mutations rewrite data parts to apply &lt;code>ALTER UPDATE&lt;/code>, &lt;code>ALTER DELETE&lt;/code>, or projection changes. They run sequentially per table and share the background merge and mutation pool with regular merges. A stalled mutation blocks subsequent mutations for that table and consumes threads needed for merges. The result is merge starvation, insert delays, and eventually &lt;code>Too many parts&lt;/code> rejections.&lt;/p></description></item><item><title>ClickHouse mutations silently blocking merges: the hidden cause of part growth</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-mutations-blocking-merges/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-mutations-blocking-merges/</guid><description>&lt;h1 id="clickhouse-mutations-silently-blocking-merges-the-hidden-cause-of-part-growth">ClickHouse mutations silently blocking merges: the hidden cause of part growth&lt;/h1>
&lt;p>You are watching part counts climb on a ClickHouse node. &lt;code>system.merges&lt;/code> shows active background tasks, health checks return Ok, and the log shows no mutation errors. Yet inserts are slowing and &lt;code>MaxPartCountForPartition&lt;/code> is trending toward the delay threshold. The pool looks busy, so merges should be keeping up. They are not: some of those busy slots are mutations, and mutations starve merges silently.&lt;/p></description></item><item><title>ClickHouse No space left on device: emergency recovery when the data disk fills</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-no-space-left-on-device/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-no-space-left-on-device/</guid><description>&lt;h1 id="clickhouse-no-space-left-on-device-emergency-recovery-when-the-data-disk-fills">ClickHouse No space left on device: emergency recovery when the data disk fills&lt;/h1>
&lt;p>When ClickHouse returns &lt;code>No space left on device&lt;/code> to client inserts or the server log fills with write errors, the situation is past a simple capacity alert. ClickHouse does not degrade gradually on a full disk. Background merges halt immediately because they require temporary free space to write combined parts before removing source files. Once merges stop, small insert parts accumulate, metadata overhead grows, and disk usage accelerates. TTL-based expiration also stops because TTL cleanup is executed by merges. ClickHouse system tables such as &lt;code>system.query_log&lt;/code> and &lt;code>system.part_log&lt;/code> are MergeTree tables that can grow unbounded and consume the remaining space.&lt;/p></description></item><item><title>ClickHouse projections and hidden parts: the part count you can't see</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-projections-hidden-parts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-projections-hidden-parts/</guid><description>&lt;h1 id="clickhouse-projections-and-hidden-parts-the-part-count-you-cant-see">ClickHouse projections and hidden parts: the part count you can&amp;rsquo;t see&lt;/h1>
&lt;p>Inserts fail with &lt;code>TOO_MANY_PARTS&lt;/code> while &lt;code>system.parts&lt;/code> on the base table looks comfortable. If you monitor only the source table, you are missing projection sub-parts inside base directories and independent parts in materialized view target tables. A single insert can spawn parts across projection subdirectories and multiple downstream tables. The number in &lt;code>system.parts&lt;/code> is a floor, not a ceiling. The gap between visible and real part count is where merge crises begin.&lt;/p></description></item><item><title>ClickHouse query error rate high: reading exception codes 241, 252, and 999</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-query-error-rate-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-query-error-rate-high/</guid><description>&lt;h1 id="clickhouse-query-error-rate-high-reading-exception-codes-241-252-and-999">ClickHouse query error rate high: reading exception codes 241, 252, and 999&lt;/h1>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>ClickHouse logs failed queries to system.query_log with type = &amp;lsquo;ExceptionWhileProcessing&amp;rsquo; and a numeric exception_code. The FailedQuery counter in system.events increments for each failure. Keep the FailedQuery / Query ratio below 1%. A sustained climb above that threshold signals a systemic issue, not a few bad queries.&lt;/p>
&lt;p>The codes that dominate production incidents fall into five classes. &lt;!-- TODO: verify exception code numbers and descriptions against your ClickHouse version -->&lt;/p></description></item><item><title>ClickHouse query latency P99 spikes: tail latency, hot shards, and cold data</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-query-latency-p99-spikes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-query-latency-p99-spikes/</guid><description>&lt;h1 id="clickhouse-query-latency-p99-spikes-tail-latency-hot-shards-and-cold-data">ClickHouse query latency P99 spikes: tail latency, hot shards, and cold data&lt;/h1>
&lt;p>Your ClickHouse cluster shows healthy average query latency, but users report intermittent timeouts. P99 spikes while P50 stays flat. This divergence is tail latency: a subset of queries is dramatically slower. In distributed clusters, the culprit is often a single hot shard or a replica serving stale data. In tiered storage, a query touching cold S3 data can be one to two orders of magnitude slower.&lt;/p></description></item><item><title>ClickHouse Replica is lost: SYSTEM RESTORE REPLICA and recovering a diverged replica</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-replica-is-lost/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-replica-is-lost/</guid><description>&lt;h1 id="clickhouse-replica-is-lost-system-restore-replica-and-recovering-a-diverged-replica">ClickHouse Replica is lost: SYSTEM RESTORE REPLICA and recovering a diverged replica&lt;/h1>
&lt;p>You run &lt;code>SELECT count()&lt;/code> on two replicas and get different results. Your application returns inconsistent aggregations depending on which node answers. The replication dashboard looks green: ZooKeeper sessions are active, &lt;code>queue_size&lt;/code> is zero, and no replica is &lt;code>readonly&lt;/code>. The replica has permanently lost parts, but ZooKeeper does not know because the loss happened outside the replication log. Silent divergence hides behind healthy metrics until queries start returning wrong results.&lt;/p></description></item><item><title>ClickHouse replicas out of sync: when SELECT count() differs across replicas</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-replicas-out-of-sync-counts-differ/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-replicas-out-of-sync-counts-differ/</guid><description>&lt;h1 id="clickhouse-replicas-out-of-sync-when-select-count-differs-across-replicas">ClickHouse replicas out of sync: when SELECT count() differs across replicas&lt;/h1>
&lt;p>You run the same &lt;code>SELECT count()&lt;/code> against two replicas and get different numbers. The replica with the lower count still responds instantly, its HTTP ping returns &lt;code>Ok.&lt;/code>, its replication queue may even be empty, and its CPU and memory look fine. Standard liveness checks give false confidence while the replica serves stale or incomplete data.&lt;/p>
&lt;p>ClickHouse replication is asynchronous and coordinated through ZooKeeper or ClickHouse Keeper. A replica can lose its session, miss fetches, or silently drop data, yet continue serving reads from local disk. Distributed queries hide the divergence unless you compare row counts or partition metadata across nodes.&lt;/p></description></item><item><title>ClickHouse ReplicatedDataLoss > 0: detecting and responding to lost parts</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-replicated-data-loss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-replicated-data-loss/</guid><description>&lt;h1 id="clickhouse-replicateddataloss--0-detecting-and-responding-to-lost-parts">ClickHouse ReplicatedDataLoss &amp;gt; 0: detecting and responding to lost parts&lt;/h1>
&lt;p>&lt;code>ReplicatedDataLoss &amp;gt; 0&lt;/code> is a hard signal in ClickHouse. A nonzero value in &lt;code>system.events&lt;/code> means the server has determined that a data part is missing and cannot be retrieved from any available replica. This is not replication lag, a transient fetch failure, or the normal &lt;code>ReplicatedPartFetchesOfMerged&lt;/code> optimization.&lt;/p>
&lt;p>Queries that touch the affected part can return incomplete results or errors. The immediate risk is silent divergence between replicas, where one replica serves stale or incomplete results without failing the query. Confirm the event, identify the scope, and determine whether a healthy peer still has the part.&lt;/p></description></item><item><title>ClickHouse replication lag: absolute_delay, queue_size, and catch-up diagnosis</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-replication-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-replication-lag/</guid><description>&lt;h1 id="clickhouse-replication-lag-absolute_delay-queue_size-and-catch-up-diagnosis">ClickHouse replication lag: absolute_delay, queue_size, and catch-up diagnosis&lt;/h1>
&lt;p>You notice &lt;code>absolute_delay&lt;/code> climbing on one replica. SELECTs there return older rows than on peers, and failover to the lagging node risks losing recently inserted data. In ClickHouse, replication lag is not a single failure; it is a symptom with several distinct causes. A replica can fall behind because it cannot pull entries from the Keeper log, because it cannot fetch parts fast enough, because a mutation is blocking the queue, or because the source replica itself is too slow to serve data.&lt;/p></description></item><item><title>ClickHouse replication queue stuck: num_tries, last_exception, and dead entries</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-replication-queue-stuck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-replication-queue-stuck/</guid><description>&lt;h1 id="clickhouse-replication-queue-stuck-num_tries-last_exception-and-dead-entries">ClickHouse replication queue stuck: num_tries, last_exception, and dead entries&lt;/h1>
&lt;p>A replica can show low &lt;code>absolute_delay&lt;/code> in &lt;code>system.replicas&lt;/code> while &lt;code>system.replication_queue&lt;/code> contains entries with &lt;code>num_tries&lt;/code> in the hundreds and the same &lt;code>last_exception&lt;/code> repeating for hours. These dead entries do not self-resolve. They block merges, fetches, or mutations, causing silent divergence and stale reads. This guide shows how to read the queue correctly, identify the failure mode from the entry type, and clear the blockage without making the replica diverge further.&lt;/p></description></item><item><title>ClickHouse slow queries: diagnosis from query_log to plan to fix</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-slow-queries-diagnosis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-slow-queries-diagnosis/</guid><description>&lt;h1 id="clickhouse-slow-queries-diagnosis-from-query_log-to-plan-to-fix">ClickHouse slow queries: diagnosis from query_log to plan to fix&lt;/h1>
&lt;p>P99 latency climbs before averages move. In ClickHouse, tail latency spikes usually mean a subset of queries is hitting cold data, scanning too many parts, or spilling to disk. The goal is to separate plan problems from resource problems fast, then confirm the fix before the next batch of queries arrives.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>Slow queries show up as rising &lt;code>query_duration_ms&lt;/code> in &lt;code>system.query_log&lt;/code> for &lt;code>QueryFinish&lt;/code> events, or rising &lt;code>elapsed&lt;/code> in &lt;code>system.processes&lt;/code>. ClickHouse latency is sensitive to how many parts a query opens, whether the mark cache is warm, and whether the pipeline stays in memory.&lt;/p></description></item><item><title>ClickHouse small inserts anti-pattern: why single-row inserts melt the merge pool</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-small-inserts-anti-pattern/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-small-inserts-anti-pattern/</guid><description>&lt;h1 id="clickhouse-small-inserts-anti-pattern-why-single-row-inserts-melt-the-merge-pool">ClickHouse small inserts anti-pattern: why single-row inserts melt the merge pool&lt;/h1>
&lt;p>&lt;code>DB::Exception: Too many parts&lt;/code> spikes, query latency climbs, and the background merge pool runs flat out while making no progress. Disk, CPU, and memory look healthy. The cause is usually single-row INSERTs from the application.&lt;/p>
&lt;p>Every INSERT into a MergeTree table creates at least one immutable data part on disk. ClickHouse is optimized for batch inserts. When clients send single-row or micro-batch inserts, parts are created faster than background merges can consolidate them. This is the number one driver of part accumulation in production clusters.&lt;/p></description></item><item><title>ClickHouse system log tables eating disk: query_log, part_log, and TTL</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-system-tables-unbounded-growth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-system-tables-unbounded-growth/</guid><description>&lt;h1 id="clickhouse-system-log-tables-eating-disk-query_log-part_log-and-ttl">ClickHouse system log tables eating disk: query_log, part_log, and TTL&lt;/h1>
&lt;p>You are investigating disk growth on a ClickHouse node. User tables look reasonable, but &lt;code>system.query_log&lt;/code> or &lt;code>system.part_log&lt;/code> are consuming tens or hundreds of gigabytes. These tables use MergeTree-family engines and are subject to the same part-count and merge pressure limits as production tables. Without a TTL rule, data accumulates forever. During incidents with high query error rates or retry storms, &lt;code>query_log&lt;/code> can expand fast enough to threaten disk capacity and stall logging.&lt;/p></description></item><item><title>ClickHouse Table is in readonly mode: is_readonly on replicated tables and how to fix it</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-table-is-in-readonly-mode/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-table-is-in-readonly-mode/</guid><description>&lt;h1 id="clickhouse-table-is-in-readonly-mode-is_readonly-on-replicated-tables-and-how-to-fix-it">ClickHouse Table is in readonly mode: is_readonly on replicated tables and how to fix it&lt;/h1>
&lt;p>Inserts fail with &lt;code>TABLE_IS_READ_ONLY&lt;/code>. &lt;code>system.replicas&lt;/code> shows &lt;code>is_readonly = 1&lt;/code> for affected tables and the replica rejects writes. For replicated tables, this is a coordination failure, not a disk or memory problem: the replica has lost its session with ClickHouse Keeper or ZooKeeper, or the ensemble is unreachable.&lt;/p>
&lt;p>Reads from local parts may still succeed, which hides the failure from load balancers and monitoring probes that rely on &lt;code>SELECT 1&lt;/code> or the HTTP ping endpoint. Brief readonly states lasting seconds are normal during Keeper leader elections. Sustained readonly lasting minutes means the replica is diverging and inserts are being lost or routed elsewhere.&lt;/p></description></item><item><title>ClickHouse too many open files: file descriptors, part count, and nofile limits</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-open-file-descriptors-exhausted/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-open-file-descriptors-exhausted/</guid><description>&lt;h1 id="clickhouse-too-many-open-files-file-descriptors-part-count-and-nofile-limits">ClickHouse too many open files: file descriptors, part count, and nofile limits&lt;/h1>
&lt;p>ClickHouse aborts queries with &amp;ldquo;Too many open files&amp;rdquo; or the server process dies. Logs show errors about failing to open column files or metadata. This is not a traditional leak; it is a capacity cliff. Every active MergeTree part keeps multiple files open, background merges temporarily spike that count, and the Linux nofile limit is usually the bottleneck. The default limit of 1024 is catastrophic for production ClickHouse.&lt;/p></description></item><item><title>ClickHouse too many partitions: why over-partitioning multiplies your part count</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-too-many-partitions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-too-many-partitions/</guid><description>&lt;h1 id="clickhouse-too-many-partitions-why-over-partitioning-multiplies-your-part-count">ClickHouse too many partitions: why over-partitioning multiplies your part count&lt;/h1>
&lt;p>You are watching &lt;code>MaxPartCountForPartition&lt;/code> climb, or you have already seen &lt;code>DB::Exception: Too many parts&lt;/code>. You check the table and see a modest total of active parts. The table looks healthy, but ClickHouse rejects inserts anyway.&lt;/p>
&lt;p>The problem is not the table total. It is the partition boundary.&lt;/p>
&lt;p>ClickHouse enforces part limits per partition, not per table. A high-cardinality &lt;code>PARTITION BY&lt;/code> key like &lt;code>toYYYYMMDD(timestamp)&lt;/code> or a business identifier creates thousands of isolated part pools. Parts never merge across those pools. One hundred partitions with thirty parts each produces 3,000 parts, multiplied merge-scheduling overhead, and higher file-descriptor pressure. One partition crossing the hard limit kills inserts for that partition while the rest of the table appears idle.&lt;/p></description></item><item><title>ClickHouse Too many simultaneous queries: max_concurrent_queries and query storms</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-too-many-simultaneous-queries/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-too-many-simultaneous-queries/</guid><description>&lt;h1 id="clickhouse-too-many-simultaneous-queries-max_concurrent_queries-and-query-storms">ClickHouse Too many simultaneous queries: max_concurrent_queries and query storms&lt;/h1>
&lt;p>You run a query and ClickHouse returns &amp;ldquo;Too many simultaneous queries.&amp;rdquo; New connections either queue or fail outright. Queries that completed in seconds yesterday now time out. The server is not down, but it is not usable.&lt;/p>
&lt;p>ClickHouse is optimized for fewer, heavier analytical queries. As concurrency rises, CPU, memory, and I/O contention increase non-linearly. A small spike can become a storm because each query is greedy.&lt;/p></description></item><item><title>ClickHouse TTL not deleting data: why expired rows survive and the disk fills</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-ttl-not-deleting-data/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-ttl-not-deleting-data/</guid><description>&lt;h1 id="clickhouse-ttl-not-deleting-data-why-expired-rows-survive-and-the-disk-fills">ClickHouse TTL not deleting data: why expired rows survive and the disk fills&lt;/h1>
&lt;p>A MergeTree table with a configured TTL can still accumulate disk usage when expired rows remain visible to &lt;code>SELECT&lt;/code>. This is not a syntax or timezone error. ClickHouse enforces TTL only as a side effect of background merges. When merges stop or never start, TTL stops, and storage grows without warning.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>ClickHouse evaluates TTL expressions only during merge operations. When the merge scheduler combines parts, it checks whether rows exceeded their TTL interval. Depending on table settings, it either drops the entire part if every row expired, or rewrites the part to exclude expired rows. Both paths require a merge. If the background merge pool is saturated, disk space is too low for temporary merge output, or a partition has no new inserts to trigger merge selection, expired rows remain indefinitely. The system does not log a warning for TTL skips; the only visible symptom is growing storage.&lt;/p></description></item><item><title>ClickHouse unauthorized DROP TABLE: auditing DDL and privilege anomalies</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-unauthorized-ddl-drop-table/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-unauthorized-ddl-drop-table/</guid><description>&lt;h1 id="clickhouse-unauthorized-drop-table-auditing-ddl-and-privilege-anomalies">ClickHouse unauthorized DROP TABLE: auditing DDL and privilege anomalies&lt;/h1>
&lt;p>A production table disappears. An application returns &amp;ldquo;table does not exist.&amp;rdquo; In ClickHouse, DROP TABLE removes the table definition and MergeTree data parts from disk immediately. On replicated tables, a single &lt;code>ON CLUSTER&lt;/code> command propagates through the coordination service and can erase the table everywhere before you intervene. There is no native undo.&lt;/p>
&lt;p>This guide is for finding out what happened, determining blast radius, and closing the gaps.&lt;/p></description></item><item><title>ClickHouse ZooKeeper session has expired: causes, recovery, and tuning</title><link>https://www.netdata.cloud/guides/clickhouse/clickhouse-zookeeper-session-expired/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/clickhouse-zookeeper-session-expired/</guid><description>&lt;h1 id="clickhouse-zookeeper-session-has-expired-causes-recovery-and-tuning">ClickHouse ZooKeeper session has expired: causes, recovery, and tuning&lt;/h1>
&lt;p>When &lt;code>system.replicas&lt;/code> shows &lt;code>is_session_expired = 1&lt;/code>, the replica has lost its ZooKeeper session and stopped participating in replication. It rejects inserts, coordinated merges, and distributed DDL. Depending on quorum and load-balancer configuration, writes may shift silently to other replicas or halt for entire shards.&lt;/p>
&lt;p>&lt;code>is_session_expired&lt;/code> often appears alongside &lt;code>is_readonly = 1&lt;/code>, but the two are distinct. Session expiration means the coordination session is dead. Read-only means the replica refuses writes. Session expiration is the most severe because it breaks the replica&amp;rsquo;s contract with the cluster.&lt;/p></description></item><item><title>Cloud Foundry</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/cloud-foundry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/cloud-foundry/</guid><description/></item><item><title>Cloud Foundry Firehose</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/cloud-foundry-firehose/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/cloud-foundry-firehose/</guid><description/></item><item><title>Cloud Foundry Firehose Monitoring</title><link>https://www.netdata.cloud/monitoring-101/cloud_foundry_firebase-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/cloud_foundry_firebase-monitoring/</guid><description>&lt;h2 id="cloud-foundry-firehose-monitoring">Cloud Foundry Firehose Monitoring&lt;/h2>
&lt;h3 id="what-is-cloud-foundry-firehose">What Is Cloud Foundry Firehose?&lt;/h3>
&lt;p>Cloud Foundry Firehose is a powerful component of the Cloud Foundry platform providing a single interface to access a variety of operational data streams. It serves as the backbone for collecting metrics, logs, events, and other crucial data about the system, making it indispensable for effective cloud service management and troubleshooting.&lt;/p>
&lt;h3 id="monitoring-cloud-foundry-firehose-with-netdata">Monitoring Cloud Foundry Firehose With Netdata&lt;/h3>
&lt;p>To monitor Cloud Foundry Firehose effectively, Netdata leverages an openmetrics (Prometheus) exporter. This integration makes it possible to ingest real-time data from any Prometheus exporter, providing users with automated dashboards, alerts, and more, all without the need for a Prometheus server or Grafana. By using Netdata’s &lt;a href="https://github.com/bosh-prometheus/firehose_exporter">Cloud Foundry Firehose monitoring tool&lt;/a>, DevOps teams can visualize live metrics and troubleshoot issues much faster, enhancing the overall reliability and efficiency of the Cloud Foundry environment.&lt;/p></description></item><item><title>Cloud Foundry Monitoring</title><link>https://www.netdata.cloud/monitoring-101/cloud_foundry-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/cloud_foundry-monitoring/</guid><description>&lt;h2 id="cloud-foundry-monitoring">Cloud Foundry Monitoring&lt;/h2>
&lt;h3 id="what-is-cloud-foundry">What Is Cloud Foundry?&lt;/h3>
&lt;p>Cloud Foundry is an open-source platform as a service (PaaS) providing a choice of cloud service providers, development frameworks, and application services. With Cloud Foundry, developers can efficiently manage application deployment, lifecycle, and scaling. This platform supports a variety of languages and frameworks, including Java, Node.js, Ruby, and Python, making it versatile for many development needs.&lt;/p>
&lt;h3 id="monitoring-cloud-foundry-with-netdata">Monitoring Cloud Foundry With Netdata&lt;/h3>
&lt;p>To monitor Cloud Foundry, Netdata utilizes an openmetrics (Prometheus) exporter. This capability allows Netdata to ingest data efficiently from any Prometheus exporter, providing users with automated dashboards and alerting systems without the hassle of managing a separate Prometheus server or Grafana. This seamless integration ensures that monitoring gets technical teams actionable insights promptly, optimizing their operations and infrastructure management.&lt;/p></description></item><item><title>Cloudgenix SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cloudgenix-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cloudgenix-snmp-traps/</guid><description/></item><item><title>CloudWatch Monitoring</title><link>https://www.netdata.cloud/monitoring-101/aws_cloudwatch-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/aws_cloudwatch-monitoring/</guid><description>&lt;h2 id="cloudwatch-monitoring">CloudWatch Monitoring&lt;/h2>
&lt;h3 id="what-is-cloudwatch">What Is CloudWatch?&lt;/h3>
&lt;p>Amazon CloudWatch is a powerful monitoring and management service offered by AWS that provides data and actionable insights to monitor applications, understand and respond to system-wide performance changes, optimize resource utilization, and gain a unified view of operational health. It enables DevOps engineers, IT admins, and developers to track metrics, monitor log files, and set alarms to quickly react to potential performance issues.&lt;/p>
&lt;h3 id="monitoring-cloudwatch-with-netdata">Monitoring CloudWatch With Netdata&lt;/h3>
&lt;p>To monitor CloudWatch with ease and precision, Netdata employs the openmetrics (prometheus) exporter. The great advantage of using Netdata as a CloudWatch monitoring tool is that it can ingest data from any Prometheus exporter. This setup allows for automated dashboards and alerts, which eliminate the need for a standalone Prometheus server or Grafana. By bridging these functionalities, tools for monitoring CloudWatch become highly intuitive and efficient.&lt;/p></description></item><item><title>ClusterControl CMON</title><link>https://www.netdata.cloud/integrations/data-collection/databases/clustercontrol-cmon/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/clustercontrol-cmon/</guid><description/></item><item><title>ClusterControl CMON Monitoring</title><link>https://www.netdata.cloud/monitoring-101/cmon-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/cmon-monitoring/</guid><description>&lt;h2 id="clustercontrol-cmon-monitoring">ClusterControl CMON Monitoring&lt;/h2>
&lt;h3 id="what-is-clustercontrol-cmon">What Is ClusterControl CMON?&lt;/h3>
&lt;p>ClusterControl CMON is a robust database management platform designed to streamline the management and monitoring of database clusters. By providing critical insights into database health, performance, and security, it serves as an invaluable tool for IT administrators and developers who need to maintain high availability and performance.&lt;/p>
&lt;h3 id="monitoring-clustercontrol-cmon-with-netdata">Monitoring ClusterControl CMON With Netdata&lt;/h3>
&lt;p>Netdata offers a seamless solution to monitor ClusterControl CMON using its advanced metrics collection capabilities. To monitor ClusterControl CMON, Netdata uses an openmetrics (Prometheus) exporter, allowing you to harness the benefits of Prometheus without the complexity of setting up a Prometheus server. By ingesting data from any Prometheus exporter, Netdata provides automated dashboards, alerts, and more. This streamlined setup means you do not need additional services like Grafana, simplifying your monitoring stack while ensuring detailed insights into your ClusterControl CMON environment.&lt;/p></description></item><item><title>CockroachDB</title><link>https://www.netdata.cloud/integrations/data-collection/databases/cockroachdb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/cockroachdb/</guid><description/></item><item><title>CockroachDB /health?ready=1: load balancer checks, draining, and impaired nodes</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-health-ready-endpoint/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-health-ready-endpoint/</guid><description>&lt;h1 id="cockroachdb-healthready1-load-balancer-checks-draining-and-impaired-nodes">CockroachDB /health?ready=1: load balancer checks, draining, and impaired nodes&lt;/h1>
&lt;p>CockroachDB exposes a readiness endpoint at &lt;code>GET /health?ready=1&lt;/code> on its HTTP port (default 8080). Load balancers and Kubernetes readiness probes use it to decide whether to route SQL traffic to a node. The node returns HTTP 200 when ready and HTTP 503 when not. This signal separates &amp;ldquo;process is alive&amp;rdquo; from &amp;ldquo;this node should receive client connections.&amp;rdquo;&lt;/p>
&lt;p>The most common operator mistake is using a plain TCP check against the SQL port (26257) or the plain &lt;code>/health&lt;/code> endpoint instead of &lt;code>/health?ready=1&lt;/code>. A TCP check succeeds as long as the port is listening, telling you nothing about whether the node is draining, write-stalled, or GC-thrashing. The plain &lt;code>/health&lt;/code> endpoint returns 200 whenever the process is running, regardless of draining state. Neither is safe for routing decisions.&lt;/p></description></item><item><title>CockroachDB admission control throttling: queue depth, store-write, and capacity headroom</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-admission-control-throttling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-admission-control-throttling/</guid><description>&lt;h1 id="cockroachdb-admission-control-throttling-queue-depth-store-write-and-capacity-headroom">CockroachDB admission control throttling: queue depth, store-write, and capacity headroom&lt;/h1>
&lt;p>When admission control starts queuing requests, p99 latency climbs while throughput stays flat or degrades. The cluster hasn&amp;rsquo;t crashed, disks aren&amp;rsquo;t full, and CPU may not be saturated. An internal flow-control system decided the node or store is at capacity and started holding work back to protect itself.&lt;/p>
&lt;p>Admission control (v21.2+, enabled by default since v22.1) &lt;!-- TODO: verify store-write queue default-enablement version; KV was v22.1, store-write may have come later --> regulates work through five queues: &lt;code>kv&lt;/code>, &lt;code>sql-kv-response&lt;/code>, &lt;code>sql-sql-response&lt;/code>, &lt;code>elastic-cpu&lt;/code>, and &lt;code>store-write&lt;/code>. Each gates a different class of work with distinct triggers. Knowing which queue is deep and why is the difference between a five-minute diagnosis and a multi-hour investigation into &amp;ldquo;why is the database slow.&amp;rdquo;&lt;/p></description></item><item><title>CockroachDB backup job failures: RPO breaches, duration trends, and stuck jobs</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-backup-job-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-backup-job-failures/</guid><description>&lt;h1 id="cockroachdb-backup-job-failures-rpo-breaches-duration-trends-and-stuck-jobs">CockroachDB backup job failures: RPO breaches, duration trends, and stuck jobs&lt;/h1>
&lt;p>CockroachDB scheduled backups run through the internal jobs system, visible via &lt;code>crdb_internal.jobs&lt;/code> (backed by &lt;code>system.jobs&lt;/code>). When a scheduled backup fails silently, stalls indefinitely, or grows so slowly that it cannot complete within its interval, your recovery point objective (RPO) is at risk. The failure is insidious: backups often appear healthy until you realize the last successful completion was 36 hours ago.&lt;/p></description></item><item><title>CockroachDB certificate expired: TLS handshake failures and online rotation</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-certificate-expired/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-certificate-expired/</guid><description>&lt;h1 id="cockroachdb-certificate-expired-tls-handshake-failures-and-online-rotation">CockroachDB certificate expired: TLS handshake failures and online rotation&lt;/h1>
&lt;p>CockroachDB uses mutual TLS for every connection: node-to-node, client-to-node, and admin UI. There is no plaintext fallback and no grace period. When a certificate expires, affected connections fail immediately.&lt;/p>
&lt;p>Node certificate expiry prevents nodes from completing TLS handshakes with each other, which can look like a network partition or quorum loss. Client certificate expiry prevents applications from connecting. CA certificate expiry invalidates the entire trust chain at once, breaking every connection simultaneously.&lt;/p></description></item><item><title>CockroachDB changefeed lag: changefeed_max_behind_nanos and the GC time bomb</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-changefeed-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-changefeed-lag/</guid><description>&lt;h1 id="cockroachdb-changefeed-lag-changefeed_max_behind_nanos-and-the-gc-time-bomb">CockroachDB changefeed lag: changefeed_max_behind_nanos and the GC time bomb&lt;/h1>
&lt;p>Changefeed lag in CockroachDB starts as a consumer latency problem and ends as a disk space emergency. The &lt;code>changefeed_max_behind_nanos&lt;/code> gauge measures how far behind your CDC feeds have fallen, but the real danger is what follows: a stalled changefeed holds a protected timestamp that prevents MVCC garbage collection, and dead data accumulates silently until the disk fills. The cluster appears healthy from the SQL side. Queries are fast, nodes are up, ranges are available. But every deleted row and every overwritten value stays on disk because GC cannot reclaim the space.&lt;/p></description></item><item><title>CockroachDB clock skew cascade: how shared NTP drift causes quorum loss</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-clock-skew-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-clock-skew-cascade/</guid><description>&lt;h1 id="cockroachdb-clock-skew-cascade-how-shared-ntp-drift-causes-quorum-loss">CockroachDB clock skew cascade: how shared NTP drift causes quorum loss&lt;/h1>
&lt;p>Multiple CockroachDB nodes crashed overnight. The logs show &amp;ldquo;clock synchronization error: this node is more than 500ms away from at least half of the known nodes.&amp;rdquo; You restart them, and they crash again. Some ranges are now unavailable. The cluster is losing quorum.&lt;/p>
&lt;p>A shared NTP failure caused multiple nodes to drift past CockroachDB&amp;rsquo;s self-termination threshold in quick succession. Single-node clock skew is bad but recoverable. Multi-node skew from a shared NTP source can take down quorum faster than the cluster can heal.&lt;/p></description></item><item><title>CockroachDB clock synchronization error: this node is more than 500ms away from at least half of the known nodes</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-clock-synchronization-error-500ms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-clock-synchronization-error-500ms/</guid><description>&lt;h1 id="cockroachdb-clock-synchronization-error-this-node-is-more-than-500ms-away-from-at-least-half-of-the-known-nodes">CockroachDB clock synchronization error: this node is more than 500ms away from at least half of the known nodes&lt;/h1>
&lt;p>You see this fatal log line on a CockroachDB node:&lt;/p>
&lt;pre tabindex="0">&lt;code>clock synchronization error: this node is more than 500ms away from at least half of the known nodes
&lt;/code>&lt;/pre>&lt;p>The process exits immediately. If the node is managed by systemd, Kubernetes, or a process supervisor, it restarts and crashes again. The crash-loop continues until the clock problem is fixed.&lt;/p></description></item><item><title>CockroachDB clock_offset_meannanos high: catching clock drift before self-termination</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-clock-offset-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-clock-offset-high/</guid><description>&lt;h1 id="cockroachdb-clock_offset_meannanos-high-catching-clock-drift-before-self-termination">CockroachDB clock_offset_meannanos high: catching clock drift before self-termination&lt;/h1>
&lt;p>&lt;code>clock_offset_meannanos&lt;/code> measures the mean clock offset between a CockroachDB node and its peers. When it climbs, you are on a path that ends in silent performance degradation from widened read uncertainty windows, or a node self-terminating to preserve data consistency.&lt;/p>
&lt;p>The thresholds are unforgiving. CockroachDB uses a default &lt;code>--max-offset&lt;/code> of 500ms. A node self-terminates when its mean offset exceeds 80% of that value (400ms) relative to a majority of peers. But a constant 200ms offset, stable and below any threshold, still doubles the uncertainty interval for every read. Transactions silently restart more often, P99 read latency creeps up, and nobody suspects the clock.&lt;/p></description></item><item><title>CockroachDB compaction backlog growing: when Pebble can't keep pace with writes</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-compaction-backlog-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-compaction-backlog-growing/</guid><description>&lt;h1 id="cockroachdb-compaction-backlog-growing-when-pebble-cant-keep-pace-with-writes">CockroachDB compaction backlog growing: when Pebble can&amp;rsquo;t keep pace with writes&lt;/h1>
&lt;p>Pebble&amp;rsquo;s background compaction threads sometimes fall behind foreground writes. In CockroachDB, this does not cause immediate failure. SSTable files accumulate in Level 0 and the compaction queue, read amplification rises, and the node drifts toward write stalls. Because the database continues to serve traffic, the backlog is easy to miss until it becomes severe.&lt;/p>
&lt;p>The earliest visible sign is a gentle upward trend in Level 0 file counts or marked-for-compaction files over hours. These metrics fluctuate with workload bursts, but a sustained upward slope means the store is consuming headroom. Healthy operation requires compaction throughput at least twice the write ingestion rate. Less than that leaves no margin for bursts, MVCC garbage collection, or rebalancing.&lt;/p></description></item><item><title>CockroachDB connection storm after failover: reconnect stampedes and surviving-node overload</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-connection-storm-after-failover/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-connection-storm-after-failover/</guid><description>&lt;h1 id="cockroachdb-connection-storm-after-failover-reconnect-stampedes-and-surviving-node-overload">CockroachDB connection storm after failover: reconnect stampedes and surviving-node overload&lt;/h1>
&lt;p>When a CockroachDB node dies, every client connected to it reconnects at the same time. Without jittered backoff in client connection pools, hundreds or thousands of new connections land on the surviving nodes within seconds. The survivors are not overwhelmed by additional query load. They are overwhelmed by connection overhead: per-connection goroutines, session memory allocations, TLS handshakes, and SQL planner initialization.&lt;/p></description></item><item><title>CockroachDB consistency check failed: checksum mismatches and confirmed corruption</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-data-consistency-check-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-data-consistency-check-failure/</guid><description>&lt;h1 id="cockroachdb-consistency-check-failed-checksum-mismatches-and-confirmed-corruption">CockroachDB consistency check failed: checksum mismatches and confirmed corruption&lt;/h1>
&lt;p>When the background consistency checker detects a checksum mismatch between replicas of a range, CockroachDB logs a fatal error and the node terminates. This is not a transient condition. The database is telling you that data integrity is compromised.&lt;/p>
&lt;p>The definitive message: &amp;ldquo;consistency check failed with N inconsistent replicas&amp;rdquo; or a Pebble-level &amp;ldquo;checksum mismatch&amp;rdquo; error. Both originate from CockroachDB&amp;rsquo;s consistency checker subsystem, not from unrelated log noise. This article covers the failure mechanism, the recovery path, and the signals that help you scope the damage.&lt;/p></description></item><item><title>CockroachDB context deadline exceeded: timeouts, slow ranges, and overload</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-context-deadline-exceeded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-context-deadline-exceeded/</guid><description>&lt;h1 id="cockroachdb-context-deadline-exceeded-timeouts-slow-ranges-and-overload">CockroachDB context deadline exceeded: timeouts, slow ranges, and overload&lt;/h1>
&lt;p>&amp;ldquo;context deadline exceeded&amp;rdquo; tells you something is slow. It does not tell you what. The same error appears whether a leaseholder is overloaded, a network path is congested, L0 sublevels are climbing past write-stall territory, or the application set a 2-second statement timeout on a query that normally takes 50 milliseconds.&lt;/p>
&lt;p>Treat the error as a symptom, not a diagnosis. The diagnostic question is: which layer is slow?&lt;/p></description></item><item><title>CockroachDB CPU saturation: Raft ticking, SQL execution, and the per-node ceiling</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-cpu-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-cpu-saturation/</guid><description>&lt;h1 id="cockroachdb-cpu-saturation-raft-ticking-sql-execution-and-the-per-node-ceiling">CockroachDB CPU saturation: Raft ticking, SQL execution, and the per-node ceiling&lt;/h1>
&lt;p>CockroachDB is CPU-hungry by design. Every range runs its own Raft state machine. Every SQL statement parses, plans, and executes through Go. Compaction, encryption-at-rest, and checksumming all burn cycles. When CPU saturates, the failure mode is not a clean slowdown: Raft heartbeats get delayed, admission control starts queuing, GC pauses lengthen, and the node risks losing liveness. The cluster average can look healthy while one leaseholder melts.&lt;/p></description></item><item><title>CockroachDB detecting hot ranges: per-range QPS, CPU asymmetry, and the Hot Ranges page</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-detecting-hot-ranges/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-detecting-hot-ranges/</guid><description>&lt;h1 id="cockroachdb-detecting-hot-ranges-per-range-qps-cpu-asymmetry-and-the-hot-ranges-page">CockroachDB detecting hot ranges: per-range QPS, CPU asymmetry, and the Hot Ranges page&lt;/h1>
&lt;p>Hot ranges are the most common performance bottleneck in CockroachDB that does not surface in aggregate metrics. The leaseholder model routes all reads and writes for a range through a single node. When one range receives disproportionate traffic, that node saturates while the rest of the cluster idles. The cluster-wide CPU average looks healthy. The per-node breakdown tells a different story.&lt;/p></description></item><item><title>CockroachDB disk space running out: capacity_available trends and the 20% rule</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-disk-space-running-out/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-disk-space-running-out/</guid><description>&lt;h1 id="cockroachdb-disk-space-running-out-capacity_available-trends-and-the-20-rule">CockroachDB disk space running out: capacity_available trends and the 20% rule&lt;/h1>
&lt;p>When a CockroachDB store runs out of disk space, the node cannot accept writes, cannot compact its LSM tree (the operation that would reclaim space), and enters a downward spiral that typically requires operator intervention. The &lt;code>capacity_available&lt;/code> metric tracks remaining free space per store, but raw free space alone does not tell the full story. MVCC garbage, protected timestamps, compaction space amplification, and replica rebalancing all consume space faster than headline data growth suggests.&lt;/p></description></item><item><title>CockroachDB disk stall detected: storage_disk_stalled and node self-termination</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-disk-stall-detected/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-disk-stall-detected/</guid><description>&lt;h1 id="cockroachdb-disk-stall-detected-storage_disk_stalled-and-node-self-termination">CockroachDB disk stall detected: storage_disk_stalled and node self-termination&lt;/h1>
&lt;p>When &lt;code>storage_disk_stalled&lt;/code> goes nonzero on a CockroachDB node, the Pebble storage engine has detected that disk I/O is no longer completing within the expected time window. The node is on a countdown to self-termination. This is not a performance degradation warning. It is a safety mechanism preparing to fire.&lt;/p>
&lt;p>CockroachDB writes every committed transaction through a write-ahead log (WAL). The WAL fsync is on the critical path for every write: Raft cannot acknowledge a commit until the log entry is persisted to disk. If that fsync blocks for long enough, the node cannot process Raft heartbeats, cannot commit writes, and cannot renew its liveness record. Rather than continue operating in a state that could produce data inconsistency, CockroachDB terminates the process.&lt;/p></description></item><item><title>CockroachDB error 53200: SQL memory budget exhausted and query rejection</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-sql-memory-budget-53200/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-sql-memory-budget-53200/</guid><description>&lt;h1 id="cockroachdb-error-53200-sql-memory-budget-exhausted-and-query-rejection">CockroachDB error 53200: SQL memory budget exhausted and query rejection&lt;/h1>
&lt;p>Error 53200 is CockroachDB&amp;rsquo;s PostgreSQL-compatible signal that the per-node SQL memory budget has run out. Applications see SQLSTATE 53200 (&amp;ldquo;insufficient resources&amp;rdquo;) with text such as &amp;ldquo;memory budget exceeded&amp;rdquo; and byte counts showing what was requested, what is allocated, and what the budget allows.&lt;/p>
&lt;p>The SQL memory budget is enforced per-node, bounded by &lt;code>--max-sql-memory&lt;/code> (default 25% of system RAM). When a node&amp;rsquo;s SQL execution layer exhausts its budget, queries on that node either spill to temporary disk storage (slow) or are rejected outright with 53200. One large analytical query on a single gateway node can starve every other session connected to that node, even if the rest of the cluster has memory to spare.&lt;/p></description></item><item><title>CockroachDB file descriptor exhaustion: SSTables, connections, and ulimit</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-file-descriptor-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-file-descriptor-exhaustion/</guid><description>&lt;h1 id="cockroachdb-file-descriptor-exhaustion-sstables-connections-and-ulimit">CockroachDB file descriptor exhaustion: SSTables, connections, and ulimit&lt;/h1>
&lt;p>When CockroachDB hits its file descriptor limit, multiple subsystems fail at once. New SQL client connections are refused. SSTable file opens fail with I/O errors. Inter-node gRPC connections fail, destabilizing Raft consensus. The node may crash, stall, or refuse to start. In the logs you will see &amp;ldquo;too many open files&amp;rdquo; errors and, if the limit is below the startup minimum, the message &amp;ldquo;open file descriptor limit of X is under the minimum required Y&amp;rdquo;.&lt;/p></description></item><item><title>CockroachDB Go GC pauses high: when garbage collection threatens Raft heartbeats</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-go-gc-pauses-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-go-gc-pauses-high/</guid><description>&lt;h1 id="cockroachdb-go-gc-pauses-high-when-garbage-collection-threatens-raft-heartbeats">CockroachDB Go GC pauses high: when garbage collection threatens Raft heartbeats&lt;/h1>
&lt;p>GC pause time spiking, node liveness flapping, lease transfers churning. SQL latency oscillates between acceptable and terrible. The cluster is alive but unstable, and the oscillation pattern is the tell.&lt;/p>
&lt;p>CockroachDB runs on the Go runtime, which performs stop-the-world garbage collection pauses. When those pauses grow long enough, they block the Raft heartbeat loop. The node stops responding to heartbeats for the duration of the pause. If the pause is long enough, the cluster declares the node dead and redistributes its leases. Then the pause ends, the node recovers, leases come back, the heap grows again, and the cycle repeats.&lt;/p></description></item><item><title>CockroachDB hot range bottleneck: one leaseholder saturated while the cluster idles</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-hot-range-bottleneck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-hot-range-bottleneck/</guid><description>&lt;h1 id="cockroachdb-hot-range-bottleneck-one-leaseholder-saturated-while-the-cluster-idles">CockroachDB hot range bottleneck: one leaseholder saturated while the cluster idles&lt;/h1>
&lt;p>One node runs hot while the rest of the cluster idles. CPU on that node is 80%+, but others hover around 20-30%. SQL latency is elevated, but only for specific tables. Transaction retry rates climb with &lt;code>writetooold&lt;/code> as the dominant restart cause. No range is unavailable, no node has lost liveness, and disk I/O is within normal bounds. The cluster has aggregate capacity, but a single range is funneling all its traffic through one leaseholder.&lt;/p></description></item><item><title>CockroachDB intent accumulation cascade: abandoned transactions and intentcount growth</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-intent-accumulation-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-intent-accumulation-cascade/</guid><description>&lt;h1 id="cockroachdb-intent-accumulation-cascade-abandoned-transactions-and-intentcount-growth">CockroachDB intent accumulation cascade: abandoned transactions and intentcount growth&lt;/h1>
&lt;p>When &lt;code>intentcount&lt;/code> and &lt;code>intentbytes&lt;/code> climb and refuse to come down, your cluster is accumulating unresolved write intents from transactions that never committed or rolled back. Every subsequent transaction touching those keys must stop, resolve the intent, then proceed. At scale, the cluster spends more CPU and I/O resolving old intents than executing new work.&lt;/p>
&lt;p>This is the intent accumulation cascade. It is one of CockroachDB&amp;rsquo;s subtler failure modes because the cluster looks healthy on infrastructure metrics: nodes are live, ranges are available, disk I/O is within bounds. The damage shows up in transaction latency and throughput, not in availability signals.&lt;/p></description></item><item><title>CockroachDB LSM compaction death spiral: L0 sublevels, read amplification, and write stalls</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-lsm-compaction-death-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-lsm-compaction-death-spiral/</guid><description>&lt;h1 id="cockroachdb-lsm-compaction-death-spiral-l0-sublevels-read-amplification-and-write-stalls">CockroachDB LSM compaction death spiral: L0 sublevels, read amplification, and write stalls&lt;/h1>
&lt;p>SQL P99 latency jumps from milliseconds to seconds. KV write latency climbs. Nodes transfer leases. Logs show Pebble write stall messages. This is the LSM compaction death spiral: writes outpace the storage engine&amp;rsquo;s ability to compact data from Level 0 down the LSM tree. L0 sublevels stack up, read amplification rises, and the node eventually stalls writes to protect itself. By the time write stalls appear, the node is already at risk of losing Raft leases and appearing partially unavailable. This guide shows how to diagnose the spiral, stop it, and prevent it.&lt;/p></description></item><item><title>CockroachDB memory pressure, GC thrashing, and Raft liveness failure</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-memory-gc-liveness-thrash/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-memory-gc-liveness-thrash/</guid><description>&lt;h1 id="cockroachdb-memory-pressure-gc-thrashing-and-raft-liveness-failure">CockroachDB memory pressure, GC thrashing, and Raft liveness failure&lt;/h1>
&lt;p>A CockroachDB node loses liveness in a repeating pattern: it drops out, the cluster redistributes its leases, it recovers, reacquires leases, then drops again. Each cycle lasts seconds to minutes. Application queries see intermittent timeouts, ambiguous results, and latency spikes that correlate with the node&amp;rsquo;s oscillation. The DB Console shows the node flapping between live and not-live states.&lt;/p>
&lt;p>This is the memory pressure to GC thrashing to Raft liveness failure cascade. The Go runtime heap grows until garbage collection pauses become long enough to prevent the node from renewing its liveness heartbeat. Once the heartbeat interval lapses, the cluster marks the node dead and moves its leases. When GC completes and memory is freed, the node recovers and reacquires leases, restarting the cycle.&lt;/p></description></item><item><title>CockroachDB Monitoring</title><link>https://www.netdata.cloud/monitoring-101/cockroachdb-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/cockroachdb-monitoring/</guid><description>&lt;h2 id="cockroachdb-monitoring">CockroachDB Monitoring&lt;/h2>
&lt;p>CockroachDB is a cloud-native distributed SQL database designed to scale horizontally and survive datacenter failures. Monitoring CockroachDB is crucial for maintaining the health, performance, and reliability of your database infrastructure. Netdata provides a comprehensive CockroachDB monitoring tool that enables real-time visibility and enhanced troubleshooting capabilities to ensure optimal database operations.&lt;/p>
&lt;h3 id="what-is-cockroachdb">What Is CockroachDB?&lt;/h3>
&lt;p>CockroachDB is a distributed SQL database, similar in functionality to Google Spanner, designed to deliver strong fault tolerance, consistency, and scalability across distributed environments. Its architecture allows for seamless scaling and resilience, making it an ideal choice for modern, high-demand database operations.&lt;/p></description></item><item><title>CockroachDB monitoring checklist: the signals every production cluster needs</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-monitoring-checklist/</guid><description>&lt;h1 id="cockroachdb-monitoring-checklist-the-signals-every-production-cluster-needs">CockroachDB monitoring checklist: the signals every production cluster needs&lt;/h1>
&lt;p>CockroachDB layers SQL execution on a replicated KV store backed by Pebble LSM trees, Raft consensus, and MVCC concurrency control. Each subsystem has distinct failure modes, and interactions between them create cascades that single-signal monitoring cannot catch. A cluster can show healthy CPU, adequate disk space, and sub-millisecond SQL latency while L0 sublevels climb toward write stalls or clock offset drifts toward the self-termination threshold.&lt;/p></description></item><item><title>CockroachDB monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-monitoring-maturity-model/</guid><description>&lt;h1 id="cockroachdb-monitoring-maturity-model-from-survival-to-expert">CockroachDB monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>CockroachDB failures rarely announce themselves through a single metric. A slow disk turns into L0 compaction debt, which stalls writes, which drops Raft proposals, which makes ranges unavailable. A drifting clock raises transaction retry rates long before any node self-terminates. To run this database safely, you need layered observability that matches the system&amp;rsquo;s own layers: storage engine, Raft replication, distributed SQL, and transaction execution.&lt;/p></description></item><item><title>CockroachDB MVCC garbage growing: tombstones, GC lag, and silent disk growth</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-mvcc-garbage-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-mvcc-garbage-growing/</guid><description>&lt;h1 id="cockroachdb-mvcc-garbage-growing-tombstones-gc-lag-and-silent-disk-growth">CockroachDB MVCC garbage growing: tombstones, GC lag, and silent disk growth&lt;/h1>
&lt;p>MVCC garbage accumulation produces no errors, no latency spikes, and no user-visible symptoms until the disk fills. By then, you may be hours away from a compaction death spiral that takes the node or cluster offline.&lt;/p>
&lt;p>CockroachDB retains old MVCC versions of every key until garbage collection removes them after the GC TTL window (default 25 hours). Deletes and updates create tombstones: markers that a key was removed at a specific timestamp. These tombstones persist in the LSM tree until MVCC GC runs and Pebble compaction physically removes them downstream. If anything blocks that cleanup, dead data accumulates silently.&lt;/p></description></item><item><title>CockroachDB node liveness failure: heartbeats, lease redistribution, and flapping</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-node-liveness-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-node-liveness-failure/</guid><description>&lt;h1 id="cockroachdb-node-liveness-failure-heartbeats-lease-redistribution-and-flapping">CockroachDB node liveness failure: heartbeats, lease redistribution, and flapping&lt;/h1>
&lt;p>Lease transfer spikes, briefly unavailable ranges, and client errors such as ambiguous results or connection resets indicate node liveness failure. In the logs, nodes transition to not-live and back within seconds. When the cluster decides a node cannot renew its liveness heartbeat, it redistributes leases. If the node recovers fast enough to renew but not fast enough to stay healthy, it flaps: an oscillating state more destructive than a clean outage because it repeatedly interrupts in-flight work and prevents stable failover.&lt;/p></description></item><item><title>CockroachDB out of memory: sys_rss, --cache, --max-sql-memory, and the OOM killer</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-out-of-memory-oom/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-out-of-memory-oom/</guid><description>&lt;h1 id="cockroachdb-out-of-memory-sys_rss---cache---max-sql-memory-and-the-oom-killer">CockroachDB out of memory: sys_rss, &amp;ndash;cache, &amp;ndash;max-sql-memory, and the OOM killer&lt;/h1>
&lt;p>A CockroachDB node disappears. No graceful shutdown, no drain sequence, no error in the SQL layer. The process is gone, and &lt;code>dmesg&lt;/code> shows the kernel OOM killer selected it. Or in Kubernetes, the pod restarts with reason &lt;code>OOMKilled&lt;/code>.&lt;/p>
&lt;p>The root cause is almost always a mismatch between what CockroachDB thinks it can allocate and what the container or host actually allows. CockroachDB partitions its memory into two manually-sized pools: the Pebble block cache (&lt;code>--cache&lt;/code>) and the SQL execution budget (&lt;code>--max-sql-memory&lt;/code>). The Go garbage collector only manages the Go heap. CGo allocations, primarily the Pebble block cache and memtables, are manually managed. When the sum of these pools plus runtime overhead exceeds the container or host limit, the OOM killer intervenes.&lt;/p></description></item><item><title>CockroachDB Pebble write stalls: when the storage engine refuses writes</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-pebble-write-stalls/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-pebble-write-stalls/</guid><description>&lt;h1 id="cockroachdb-pebble-write-stalls-when-the-storage-engine-refuses-writes">CockroachDB Pebble write stalls: when the storage engine refuses writes&lt;/h1>
&lt;p>When application writes time out or return ambiguous errors, check CockroachDB logs for &lt;code>pebble: write stall&lt;/code> and watch &lt;code>storage_write_stalls&lt;/code>. Pebble pauses writes to a store when Level 0 compaction debt exceeds safe thresholds. Until compaction drains L0, that store cannot accept new writes.&lt;/p>
&lt;p>Write stalls are the most severe storage signal in CockroachDB. A brief stall during bulk loading may be harmless, but sustained stalls at one per second mean the node cannot meet its Raft obligations. The node may still serve reads, but it can lose Raft leadership because it cannot append log entries. That cascades into lease transfers and temporary range unavailability.&lt;/p></description></item><item><title>CockroachDB protected timestamp GC stall: stalled changefeeds that silently fill the disk</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-protected-timestamp-gc-stall/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-protected-timestamp-gc-stall/</guid><description>&lt;h1 id="cockroachdb-protected-timestamp-gc-stall-stalled-changefeeds-that-silently-fill-the-disk">CockroachDB protected timestamp GC stall: stalled changefeeds that silently fill the disk&lt;/h1>
&lt;p>Disk space is growing on your CockroachDB cluster and nothing explains it. &lt;code>capacity_available&lt;/code> trends downward hour after hour, but live data size has not changed. No bulk imports, no schema changes, no write spike. The SQL layer looks healthy: latency is fine, throughput is normal. But the disk is filling.&lt;/p>
&lt;p>A CDC changefeed, a backup job, or another internal operation has created a protected timestamp (PTS) record that prevents MVCC garbage collection from reclaiming old data versions. The job has stalled, paused, or entered a retry loop, but the PTS record persists. Dead MVCC data accumulates underneath it. Disk usage grows linearly until free space drops below what compaction needs, at which point L0 sublevels spike, write stalls begin, and nodes lose liveness.&lt;/p></description></item><item><title>CockroachDB Raft liveness failure cascade: slow node, lost leases, rolling unavailability</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-raft-liveness-failure-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-raft-liveness-failure-cascade/</guid><description>&lt;h1 id="cockroachdb-raft-liveness-failure-cascade-slow-node-lost-leases-rolling-unavailability">CockroachDB Raft liveness failure cascade: slow node, lost leases, rolling unavailability&lt;/h1>
&lt;p>You notice brief, repeating SQL timeouts or ambiguous result errors across your application. In the CockroachDB DB Console, one node is still marked as running, yet its range lease count is oscillating and lease transfers are spiking. This is not a clean node crash. It is a Raft liveness failure cascade: a node has become slow enough to miss Raft heartbeats, but not dead enough to fail permanently. The cluster keeps trying to move leases away and back, creating rolling windows of unavailability.&lt;/p></description></item><item><title>CockroachDB Raft snapshot storm: when followers fall too far behind</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-raft-snapshot-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-raft-snapshot-storm/</guid><description>&lt;h1 id="cockroachdb-raft-snapshot-storm-when-followers-fall-too-far-behind">CockroachDB Raft snapshot storm: when followers fall too far behind&lt;/h1>
&lt;p>Raft snapshots are CockroachDB&amp;rsquo;s fallback when incremental log replication is no longer possible. A leader sends a full copy of a range (up to 512 MiB by default) to a follower that has fallen so far behind that the log entries it needs have already been truncated. In steady state, this almost never happens. When it starts happening at scale, it creates a resource-draining cycle that can degrade an entire cluster.&lt;/p></description></item><item><title>CockroachDB range count per node too high: Raft ticking as a scaling dimension</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-range-count-too-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-range-count-too-high/</guid><description>&lt;h1 id="cockroachdb-range-count-per-node-too-high-raft-ticking-as-a-scaling-dimension">CockroachDB range count per node too high: Raft ticking as a scaling dimension&lt;/h1>
&lt;p>CPU utilization is climbing across one or more CockroachDB nodes, and the usual explanations do not fit. Query throughput is flat. Disk I/O looks healthy. Admission control is not queuing. L0 sublevels are in single digits. CPU keeps trending upward, and the only change is that the database holds more data.&lt;/p>
&lt;p>The likely cause is the scaling dimension most teams never instrument: range count per node. Every range in CockroachDB is an independent Raft consensus group. Even idle, each range&amp;rsquo;s Raft state machine ticks at a fixed interval, consuming CPU separate from query execution, compaction, or any workload-driven process. This baseline overhead grows linearly with range count per node.&lt;/p></description></item><item><title>CockroachDB range unavailable: diagnosing ranges_unavailable and recovering quorum</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-range-unavailable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-range-unavailable/</guid><description>&lt;h1 id="cockroachdb-range-unavailable-diagnosing-ranges_unavailable-and-recovering-quorum">CockroachDB range unavailable: diagnosing ranges_unavailable and recovering quorum&lt;/h1>
&lt;p>The Prometheus gauge &lt;code>ranges_unavailable&lt;/code> is the hardest of hard signals in CockroachDB. When it rises above zero and stays there, some slice of your keyspace has no leaseholder or has lost Raft quorum. Reads and writes to those ranges block or fail, and applications see errors, retries, or timeouts. If the unavailable ranges include system keyspaces, the impact is not limited to user tables. The cluster may lose the ability to heartbeat node liveness, execute schema changes, or route requests.&lt;/p></description></item><item><title>CockroachDB ReadWithinUncertaintyInterval restarts: the near-diagnostic signal of clock skew</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-readwithinuncertainty-restarts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-readwithinuncertainty-restarts/</guid><description>&lt;h1 id="cockroachdb-readwithinuncertaintyinterval-restarts-the-near-diagnostic-signal-of-clock-skew">CockroachDB ReadWithinUncertaintyInterval restarts: the near-diagnostic signal of clock skew&lt;/h1>
&lt;p>When CockroachDB reports &lt;code>readwithinuncertainty&lt;/code> as a transaction restart cause, you have a clock synchronization problem. This restart cause does not appear in meaningful quantities for any other reason. Any sustained nonzero rate, even well below the self-termination threshold, indicates that NTP or the underlying clock source is not keeping node clocks aligned.&lt;/p>
&lt;p>The error stems from CockroachDB&amp;rsquo;s Hybrid Logical Clock (HLC) design. Each transaction receives a timestamp from its gateway node&amp;rsquo;s HLC, which combines physical wall-clock time with a logical counter. When a read on one node encounters a write from a transaction that started on a different node, and the two nodes&amp;rsquo; clocks are not aligned, the database cannot determine which transaction began first. CockroachDB conservatively restarts the reading transaction at a higher timestamp. Under SERIALIZABLE isolation, this restart requires client-side retry logic and adds directly to tail latency.&lt;/p></description></item><item><title>CockroachDB replica unavailable: lost quorum and stuck Raft groups</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-replica-unavailable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-replica-unavailable/</guid><description>&lt;h1 id="cockroachdb-replica-unavailable-lost-quorum-and-stuck-raft-groups">CockroachDB replica unavailable: lost quorum and stuck Raft groups&lt;/h1>
&lt;p>&lt;code>ranges_unavailable&lt;/code> has gone nonzero. Clients are seeing &amp;ldquo;replica unavailable&amp;rdquo; errors. Some portion of your keyspace cannot be read or written. This is an active availability incident.&lt;/p>
&lt;p>In CockroachDB, every range (a ~512 MiB slice of the keyspace) is replicated across multiple nodes. A write requires quorum acknowledgment from a majority of replicas before it commits. When too few replicas are reachable, the range loses quorum and cannot serve reads or writes. The &lt;code>ranges_unavailable&lt;/code> metric tracks exactly this condition: ranges with no leaseholder or with lost Raft quorum.&lt;/p></description></item><item><title>CockroachDB result is ambiguous: ambiguous commit errors and how to handle them</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-result-is-ambiguous/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-result-is-ambiguous/</guid><description>&lt;h1 id="cockroachdb-result-is-ambiguous-ambiguous-commit-errors-and-how-to-handle-them">CockroachDB result is ambiguous: ambiguous commit errors and how to handle them&lt;/h1>
&lt;p>Your application received &lt;code>result is ambiguous&lt;/code> (PostgreSQL error code 40003, &lt;code>statement_completion_unknown&lt;/code>) from CockroachDB. The transaction may have committed. It may have been aborted. The database does not know, and neither do you.&lt;/p>
&lt;p>This is not a bug. It is a correctness property of CockroachDB&amp;rsquo;s distributed commit protocol. When the transaction coordinator loses contact with the leaseholder during the final commit step, or during the last statement of an implicit transaction, it cannot determine whether the commit succeeded. Rather than guess and risk a silent double-apply, CockroachDB surfaces the ambiguity.&lt;/p></description></item><item><title>CockroachDB RETRY_WRITE_TOO_OLD: the most common contention restart and how to kill it</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-write-too-old/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-write-too-old/</guid><description>&lt;h1 id="cockroachdb-retry_write_too_old-the-most-common-contention-restart-and-how-to-kill-it">CockroachDB RETRY_WRITE_TOO_OLD: the most common contention restart and how to kill it&lt;/h1>
&lt;p>Your cluster is healthy: nodes are live, ranges are available, disk space is fine, L0 sublevels are low. But transaction commit latency is climbing, P99 is widening, and the &lt;code>txn_restarts&lt;/code> metric shows a rising rate tagged &lt;code>writetooold&lt;/code>. Whether applications surface errors to users depends on whether their retry logic works.&lt;/p>
&lt;p>RETRY_WRITE_TOO_OLD is the most common transaction restart cause under CockroachDB&amp;rsquo;s default SERIALIZABLE isolation. It is not an infrastructure problem. It is a contention problem rooted in schema design, key patterns, and transaction scope. Adding nodes does not help because the bottleneck is logical serialization, not physical capacity. Once you identify the conflicting keys and the transaction patterns driving them, writetooold restarts are fixable.&lt;/p></description></item><item><title>CockroachDB sequential primary key hotspot: SERIAL, timestamps, and write skew</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-sequential-key-hotspot/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-sequential-key-hotspot/</guid><description>&lt;h1 id="cockroachdb-sequential-primary-key-hotspot-serial-timestamps-and-write-skew">CockroachDB sequential primary key hotspot: SERIAL, timestamps, and write skew&lt;/h1>
&lt;p>One node in your CockroachDB cluster runs hot while the others sit near idle. Write latency for specific tables climbs into hundreds of milliseconds. Transaction restarts tagged &lt;code>writetooold&lt;/code> increase steadily alongside the insert rate. The cluster has plenty of aggregate CPU, memory, and disk headroom, but the workload cannot use it.&lt;/p>
&lt;p>CockroachDB stores data in an ordered keyspace divided into ranges of approximately 512 MiB. Each range has a single leaseholder that serves all reads and coordinates all writes for keys in that range. When primary keys are sequential (SERIAL, auto-incrementing integers, timestamps), every new insert lands near the end of the keyspace, concentrated in the same range. That range&amp;rsquo;s leaseholder becomes a single-node bottleneck.&lt;/p></description></item><item><title>CockroachDB SQL error rate: XX000 internal errors and error-code triage</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-sql-error-rate-xx000/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-sql-error-rate-xx000/</guid><description>&lt;h1 id="cockroachdb-sql-error-rate-xx000-internal-errors-and-error-code-triage">CockroachDB SQL error rate: XX000 internal errors and error-code triage&lt;/h1>
&lt;p>&lt;code>sql_failure_count&lt;/code> is a single Prometheus counter with no error-code labels. A serialization conflict (40001) is expected noise in a contended workload. An internal error (XX000) is a database fault that needs immediate escalation. Treating the aggregate counter as one signal guarantees the wrong response.&lt;/p>
&lt;p>This article covers how to break that aggregate apart by error code, when XX000 internal errors warrant paging, and how to triage the other common error classes.&lt;/p></description></item><item><title>CockroachDB SQL latency high: sql_service_latency, per-fingerprint diagnosis, and the layers</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-sql-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-sql-latency-high/</guid><description>&lt;p>P99 &lt;code>sql_service_latency&lt;/code> is climbing. Applications are timing out or retrying. The DB Console overview chart shows the spike, but the aggregate histogram cannot tell you which query is slow or which subsystem is adding the latency.&lt;/p>
&lt;p>The most common diagnostic mistake is stopping at aggregate P99. A single slow query hiding among thousands of fast ones barely moves aggregate P99 but can take down an endpoint. A system-wide storage problem lifts all fingerprints together. Without per-fingerprint breakdown, you cannot distinguish &amp;ldquo;one query regressed&amp;rdquo; from &amp;ldquo;everything got slower.&amp;rdquo;&lt;/p></description></item><item><title>CockroachDB storage_l0_sublevels climbing: the earliest warning of write stalls</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-l0-sublevels-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-l0-sublevels-high/</guid><description>&lt;h1 id="cockroachdb-storage_l0_sublevels-climbing-the-earliest-warning-of-write-stalls">CockroachDB storage_l0_sublevels climbing: the earliest warning of write stalls&lt;/h1>
&lt;p>&lt;code>storage_l0_sublevels&lt;/code> climbing on a single store is the earliest actionable signal for CockroachDB storage distress. The cluster may still serve traffic with slightly elevated SQL latency and no active write stalls. L0 sublevels typically provide 10-30 minutes of warning before write stalls begin, but only when monitored per-store. &lt;!-- TODO: verify 10-30 minute lead time holds across hardware/configs -->&lt;/p>
&lt;p>As L0 sublevels rise, Pebble consults more SSTables per read, read amplification increases nonlinearly, and compaction falls behind. Admission control shapes regular traffic at 5 sublevels and elastic traffic at 1 sublevel. &lt;!-- TODO: verify exact admission control sublevel thresholds --> Past 20 sublevels, write stalls become imminent. Because this metric is per-store, a node with multiple stores can mask a hot disk behind healthy node-level aggregates.&lt;/p></description></item><item><title>CockroachDB store is full: emergency disk recovery and why deletes don't free space</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-store-is-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-store-is-full/</guid><description>&lt;h1 id="cockroachdb-store-is-full-emergency-disk-recovery-and-why-deletes-dont-free-space">CockroachDB store is full: emergency disk recovery and why deletes don&amp;rsquo;t free space&lt;/h1>
&lt;p>A CockroachDB node&amp;rsquo;s store directory is full or nearly full, and the node may have already shut itself down to protect data integrity.&lt;/p>
&lt;p>Deleting data will not help immediately. CockroachDB uses MVCC: a SQL DELETE writes tombstones at new timestamps rather than removing old data. Until those tombstones propagate through Pebble compaction, the old versions stay on disk. A large-scale delete can temporarily increase disk usage because the tombstones are written before old data is reclaimed.&lt;/p></description></item><item><title>CockroachDB too many connections: sql_conns, pooling, and goroutine pressure</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-too-many-connections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-too-many-connections/</guid><description>&lt;h1 id="cockroachdb-too-many-connections-sql_conns-pooling-and-goroutine-pressure">CockroachDB too many connections: sql_conns, pooling, and goroutine pressure&lt;/h1>
&lt;p>CockroachDB has no hard built-in connection limit comparable to PostgreSQL&amp;rsquo;s &lt;code>max_connections&lt;/code>. &lt;!-- TODO: verify whether the cluster setting sql.max_connections (or server.max_connections_per_gateway) provides an effective cap in recent versions. --> Connections accumulate until something else breaks: goroutine pressure, memory exhaustion, or file descriptor limits.&lt;/p>
&lt;p>The &lt;code>sql_conns&lt;/code> gauge tracks active pgwire (SQL) connections per node. With properly configured client-side pooling, this number stays modest. When it climbs past 1000-2000 per node, you almost certainly have a pooling problem upstream.&lt;/p></description></item><item><title>CockroachDB transaction contention storm: retries, intents, and tail latency</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-transaction-contention-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-transaction-contention-storm/</guid><description>&lt;h1 id="cockroachdb-transaction-contention-storm-retries-intents-and-tail-latency">CockroachDB transaction contention storm: retries, intents, and tail latency&lt;/h1>
&lt;p>The cluster looks healthy. Node liveness is stable, &lt;code>ranges_unavailable&lt;/code> is zero, disk I/O is within bounds, CPU is moderate across all nodes. But P99 SQL latency is climbing, the transaction abort rate is up, and applications are reporting timeout errors. The database is spending most of its energy waiting, retrying, and resolving conflicts rather than doing useful work.&lt;/p>
&lt;p>This is a transaction contention storm. The infrastructure is fine; the workload is the problem. Multiple transactions are contending for the same keys, creating serialized hot paths. Each conflict generates retries under CockroachDB&amp;rsquo;s default SERIALIZABLE isolation, which in turn leaves write intents that block subsequent transactions. Under sustained load, the retry-intent-latency cycle becomes a positive feedback loop.&lt;/p></description></item><item><title>CockroachDB transaction retry rate high: breaking down txn_restarts by cause</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-transaction-retry-rate-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-transaction-retry-rate-high/</guid><description>&lt;h1 id="cockroachdb-transaction-retry-rate-high-breaking-down-txn_restarts-by-cause">CockroachDB transaction retry rate high: breaking down txn_restarts by cause&lt;/h1>
&lt;p>When application latency creeps upward without an obvious cause, or you start seeing PostgreSQL error code 40001 (serialization failure) in application logs, the transaction retry rate is likely climbing. CockroachDB exposes &lt;code>txn_restarts&lt;/code> as a set of counters, each tagged by cause. The aggregate number tells you retries are happening. It does not tell you why.&lt;/p>
&lt;p>The five dominant causes are &lt;code>writetooold&lt;/code>, &lt;code>serializable&lt;/code>, &lt;code>readwithinuncertainty&lt;/code>, &lt;code>txnaborted&lt;/code>, and &lt;code>txnpush&lt;/code>. Three additional sub-metrics (&lt;code>asyncwritefailure&lt;/code>, &lt;code>commitdeadlineexceeded&lt;/code>, &lt;code>unknown&lt;/code>) cover edge cases. A spike dominated by &lt;code>readwithinuncertainty&lt;/code> is a clock infrastructure problem. A spike dominated by &lt;code>writetooold&lt;/code> is a schema or application contention problem. The fix for one does nothing for the other.&lt;/p></description></item><item><title>CockroachDB TransactionRetryWithProtoRefreshError: RETRY_SERIALIZABLE causes and fixes</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-transaction-retry-serializable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-transaction-retry-serializable/</guid><description>&lt;h1 id="cockroachdb-transactionretrywithprotorefresherror-retry_serializable-causes-and-fixes">CockroachDB TransactionRetryWithProtoRefreshError: RETRY_SERIALIZABLE causes and fixes&lt;/h1>
&lt;p>The application logs show &lt;code>TransactionRetryWithProtoRefreshError: TransactionRetryError: retry txn (RETRY_SERIALIZABLE - failed preemptive refresh due to a conflict: committed value on key /Table/.../0)&lt;/code>. SQLSTATE is 40001. Clients are timing out or failing outright. The cluster looks healthy: nodes are live, ranges are available, disk space is fine. The problem is in the transaction layer.&lt;/p>
&lt;p>Under SERIALIZABLE isolation (CockroachDB&amp;rsquo;s default), the database uses optimistic concurrency. It proceeds as if no conflict exists, then detects conflicts at commit or statement boundary. When a transaction&amp;rsquo;s preemptive refresh fails because another transaction already committed a write to a key the first transaction read, CockroachDB emits RETRY_SERIALIZABLE. This is expected behavior, not a bug.&lt;/p></description></item><item><title>CockroachDB under-replicated ranges: ranges_underreplicated and the healing margin</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-under-replicated-ranges/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-under-replicated-ranges/</guid><description>&lt;h1 id="cockroachdb-under-replicated-ranges-ranges_underreplicated-and-the-healing-margin">CockroachDB under-replicated ranges: ranges_underreplicated and the healing margin&lt;/h1>
&lt;p>When &lt;code>ranges_underreplicated&lt;/code> rises above zero, CockroachDB has ranges with fewer live replicas than the configured replication factor. The replicate queue is already working to heal them by transferring Raft snapshots to new target nodes, but until healing completes, those ranges have reduced or zero fault tolerance margin.&lt;/p>
&lt;p>The key operational question is not whether under-replicated ranges exist (they will, transiently, after almost any node event), but whether the cluster is healing fast enough to stay ahead of the next failure. This is the healing margin: the buffer between the current replication state and the point where a range loses quorum and becomes unavailable.&lt;/p></description></item><item><title>CockroachDB WAL fsync latency high: the most direct write-path health signal</title><link>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-wal-fsync-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/cockroachdb-wal-fsync-latency-high/</guid><description>&lt;h1 id="cockroachdb-wal-fsync-latency-high-the-most-direct-write-path-health-signal">CockroachDB WAL fsync latency high: the most direct write-path health signal&lt;/h1>
&lt;p>Elevated &lt;code>storage_wal_fsync_latency&lt;/code> means the storage device is taking too long to durably persist write-ahead log entries. Every write in CockroachDB must complete a WAL fsync before Raft can acknowledge it. Fsync latency is the gating signal for the entire write path.&lt;/p>
&lt;p>On healthy local SSDs, WAL fsync P99 should be under 10ms. Sustained values above 50ms mean the device is struggling. Above 200ms, the node is approaching disk stall territory, where CockroachDB&amp;rsquo;s built-in stall detection may terminate the process to prevent data inconsistency. Disk stalls cause self-termination before utilization metrics react: a volume reporting 60% utilization can stall on an individual fsync for seconds if the underlying hardware is failing, the I/O scheduler is misbehaving, or a cloud provider is throttling IOPS.&lt;/p></description></item><item><title>Codan Limited SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/codan-limited-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/codan-limited-snmp-traps/</guid><description/></item><item><title>Codex SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/codex-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/codex-snmp-traps/</guid><description/></item><item><title>Codima Technologies Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/codima-technologies-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/codima-technologies-ltd-snmp-traps/</guid><description/></item><item><title>Cohesity Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cohesity-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cohesity-inc-snmp-traps/</guid><description/></item><item><title>Cold-start topology: why your map is incomplete after a collector restart</title><link>https://www.netdata.cloud/guides/network/network-cold-start-topology/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-cold-start-topology/</guid><description>&lt;h1 id="cold-start-topology-why-your-map-is-incomplete-after-a-collector-restart">Cold-start topology: why your map is incomplete after a collector restart&lt;/h1>
&lt;p>You restart your flow collector or topology engine for a routine upgrade. The process comes back up cleanly. The dashboard loads. But the topology map is half-empty, endpoint positions are wrong, and within minutes someone pages you asking why a security investigation points to the wrong switch port.&lt;/p>
&lt;p>The root cause is not a bug. After a restart, the topology inference engine has no cached neighbor tables, no FDB entries, no ARP data, and potentially no flow templates. It must rebuild all of these from live polling and flow data before it can produce a reliable view. The window between restart and first complete topology ranges from a few minutes to over 30 minutes, depending on poll cadence, template refresh intervals, and which sources your topology engine fuses.&lt;/p></description></item><item><title>Collectd</title><link>https://www.netdata.cloud/integrations/data-collection/applications/collectd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/collectd/</guid><description/></item><item><title>Collectd Monitoring</title><link>https://www.netdata.cloud/monitoring-101/collectd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/collectd-monitoring/</guid><description>&lt;h2 id="collectd-monitoring">Collectd Monitoring&lt;/h2>
&lt;h3 id="what-is-collectd">What Is Collectd?&lt;/h3>
&lt;p>Collectd is a powerful tool for monitoring systems and applications. It collects performance data from your systems and networks, providing detailed insights to help troubleshoot issues and optimize performance. As a long-standing part of the open-source community, Collectd offers a comprehensive and extensible solution for those needing robust data collection capabilities.&lt;/p>
&lt;h3 id="monitoring-collectd-with-netdata">Monitoring Collectd With Netdata&lt;/h3>
&lt;p>To monitor Collectd effectively, Netdata uses an OpenMetrics (Prometheus) exporter. This integration allows Netdata to ingest data directly from any Prometheus exporter without requiring a Prometheus server or Grafana setup. With Netdata, you gain access to automated dashboards, real-time alerts, and insightful metrics visualization, making it an ideal Collectd monitoring tool. Easily connect and start monitoring your systems using &lt;a href="https://github.com/prometheus/collectd_exporter">Collectd exporter&lt;/a> combined with &lt;a href="https://app.netdata.cloud/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata&amp;rsquo;s intuitive interface&lt;/a>.&lt;/p></description></item><item><title>Collector CPU and TSDB write-queue saturation: the capacity signals</title><link>https://www.netdata.cloud/guides/network/network-collector-cpu-disk-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-collector-cpu-disk-saturation/</guid><description>&lt;h1 id="collector-cpu-and-tsdb-write-queue-saturation-the-capacity-signals">Collector CPU and TSDB write-queue saturation: the capacity signals&lt;/h1>
&lt;p>When a network monitoring collector saturates, the first visible symptom is rarely high collector CPU. It is traffic charts showing a decline during a traffic spike, an SNMP poll cycle drifting past its configured interval, or unexplained gaps in flow data. The degradation sits one to three subsystems downstream of the actual bottleneck, which is why collector-side incidents are frequently misdiagnosed.&lt;/p></description></item><item><title>Colubris Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/colubris-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/colubris-networks-inc-snmp-traps/</guid><description/></item><item><title>Com21 SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/com21-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/com21-snmp-traps/</guid><description/></item><item><title>Comap A S SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/comap-a-s-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/comap-a-s-snmp-traps/</guid><description/></item><item><title>Comet System S R O SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/comet-system-s-r-o-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/comet-system-s-r-o-snmp-traps/</guid><description/></item><item><title>Command_Timeout climbing: the drive is taking too long to respond</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-command-timeout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-command-timeout/</guid><description>&lt;h1 id="command_timeout-climbing-the-drive-is-taking-too-long-to-respond">Command_Timeout climbing: the drive is taking too long to respond&lt;/h1>
&lt;p>SMART attribute ID 188 (Command_Timeout) is climbing. Two interpretation traps cause false alarms here, and the symptom overlaps with at least four distinct failure modes ranging from benign background maintenance to imminent controller death.&lt;/p>
&lt;p>The first trap is vendor-specific raw value encoding. On Seagate drives, the raw value packs three 16-bit counters into a single 48-bit field, producing numbers in the billions that decode to single-event counts. Monitoring tools that read the raw value as one integer will page you for a single timeout during power-on sequencing.&lt;/p></description></item><item><title>Commend International GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/commend-international-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/commend-international-gmbh-snmp-traps/</guid><description/></item><item><title>Communication Company Nari Group Corporation Information Technology SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/communication-company-nari-group-corporation-information-technology-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/communication-company-nari-group-corporation-information-technology-snmp-traps/</guid><description/></item><item><title>Community</title><link>https://www.netdata.cloud/community/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/community/</guid><description/></item><item><title>Compaq SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/compaq-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/compaq-snmp-traps/</guid><description/></item><item><title>Compellent Technologies SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/compellent-technologies-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/compellent-technologies-snmp-traps/</guid><description/></item><item><title>Compex Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/compex-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/compex-inc-snmp-traps/</guid><description/></item><item><title>Computer Associates International SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/computer-associates-international-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/computer-associates-international-snmp-traps/</guid><description/></item><item><title>Computer Network Technology SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/computer-network-technology-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/computer-network-technology-snmp-traps/</guid><description/></item><item><title>Comtech Efdata Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/comtech-efdata-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/comtech-efdata-corporation-snmp-traps/</guid><description/></item><item><title>Concord Communications SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/concord-communications-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/concord-communications-snmp-traps/</guid><description/></item><item><title>Concourse</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/concourse/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/concourse/</guid><description/></item><item><title>Concourse Monitoring</title><link>https://www.netdata.cloud/monitoring-101/concourse-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/concourse-monitoring/</guid><description>&lt;h2 id="concourse-monitoring">Concourse Monitoring&lt;/h2>
&lt;h3 id="what-is-concourse">What Is Concourse?&lt;/h3>
&lt;p>Concourse is a powerful, open-source CI/CD system that automates and streamlines your application development and deployment processes. Its clean and straightforward interface, coupled with its ability to handle complex workflows, makes it a popular choice among DevOps teams looking to enhance their software release cycles.&lt;/p>
&lt;h3 id="monitoring-concourse-with-netdata">Monitoring Concourse With Netdata&lt;/h3>
&lt;p>Monitoring your Concourse environment is crucial for ensuring optimal pipeline performance and detecting issues before they affect application delivery. With Netdata, monitoring Concourse becomes seamless and efficient. Netdata leverages an openmetrics (Prometheus) exporter to collect comprehensive metrics from your Concourse instance. This means that you can ingest data directly from any Prometheus exporter, receiving automated dashboards, alerts, and more, without the need for a separate Prometheus server or Grafana setup.&lt;/p></description></item><item><title>Confmon Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/confmon-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/confmon-corp-snmp-traps/</guid><description/></item><item><title>Connection Technology Systems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/connection-technology-systems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/connection-technology-systems-snmp-traps/</guid><description/></item><item><title>Conntrack</title><link>https://www.netdata.cloud/integrations/data-collection/networking/conntrack/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/conntrack/</guid><description/></item><item><title>Consul</title><link>https://www.netdata.cloud/integrations/data-collection/applications/consul/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/consul/</guid><description/></item><item><title>Consul "ACL not found": requests rejected after a token or policy change</title><link>https://www.netdata.cloud/guides/consul/consul-acl-not-found/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-acl-not-found/</guid><description>&lt;h1 id="consul-acl-not-found-requests-rejected-after-a-token-or-policy-change">Consul &amp;ldquo;ACL not found&amp;rdquo;: requests rejected after a token or policy change&lt;/h1>
&lt;p>&amp;ldquo;ACL not found&amp;rdquo; appears as HTTP 403 responses carrying &amp;ldquo;ACL not found&amp;rdquo; or &amp;ldquo;token does not exist: ACL not found&amp;rdquo;. It surfaces in agent and server logs, sidecar injector output, and as failed service registrations, health check updates, or KV writes. The blast radius depends on which token is missing: a single application token breaks one service; a replication or agent token can break an entire datacenter&amp;rsquo;s authorization pipeline.&lt;/p></description></item><item><title>Consul "No cluster leader": every write is failing</title><link>https://www.netdata.cloud/guides/consul/consul-no-cluster-leader/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-no-cluster-leader/</guid><description>&lt;h1 id="consul-no-cluster-leader-every-write-is-failing">Consul &amp;ldquo;No cluster leader&amp;rdquo;: every write is failing&lt;/h1>
&lt;p>When Consul cannot elect a Raft leader, every write fails. The error surfaces to clients as &lt;code>rpc error making call: No cluster leader&lt;/code>, sometimes as &lt;code>no known leader&lt;/code>. Service registrations hang, KV writes return errors, sessions and ACL tokens cannot be created, and health check state stops reaching the catalog.&lt;/p>
&lt;p>Only stale reads continue to work: stale-mode API queries and DNS lookups (which default to &lt;code>allow_stale&lt;/code>) &lt;!-- TODO: verify exact Consul version where DNS allow_stale defaulted to true --> return whatever the responding server last committed. Default and consistent reads fail because they require a leader, which can surface differently to callers depending on their read mode.&lt;/p></description></item><item><title>Consul "too many open files": file descriptor exhaustion on servers</title><link>https://www.netdata.cloud/guides/consul/consul-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-too-many-open-files/</guid><description>&lt;h1 id="consul-too-many-open-files-file-descriptor-exhaustion-on-servers">Consul &amp;ldquo;too many open files&amp;rdquo;: file descriptor exhaustion on servers&lt;/h1>
&lt;p>When a Consul server hits its file descriptor ceiling, every consumer that needs a new FD fails immediately: RPC connections are refused, DNS listeners stop accepting queries, xDS streams to Envoy sidecars cannot open, and outbound health-check connections fail. Gossip probes time out and the failure detection protocol marks otherwise healthy peers as suspect or failed. A healthy-looking cluster a minute ago is now generating cascading pages.&lt;/p></description></item><item><title>Consul ACL resolution latency: token cache thrashing on every request</title><link>https://www.netdata.cloud/guides/consul/consul-acl-resolution-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-acl-resolution-latency/</guid><description>&lt;h1 id="consul-acl-resolution-latency-token-cache-thrashing-on-every-request">Consul ACL resolution latency: token cache thrashing on every request&lt;/h1>
&lt;p>Every authenticated Consul request pays an ACL resolution tax. When that tax is sub-millisecond, nobody notices. When the token cache cannot hold the working set, every request becomes a miss and that tax multiplies across DNS lookups, HTTP API calls, RPC forwarding, and Connect intention evaluation. The cluster keeps answering, but slowly, and the slowness is uniform across every authenticated path.&lt;/p></description></item><item><title>Consul anti-entropy not syncing: local agent state and the catalog drifting apart</title><link>https://www.netdata.cloud/guides/consul/consul-catalog-staleness-anti-entropy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-catalog-staleness-anti-entropy/</guid><description>&lt;h1 id="consul-anti-entropy-not-syncing-local-agent-state-and-the-catalog-drifting-apart">Consul anti-entropy not syncing: local agent state and the catalog drifting apart&lt;/h1>
&lt;p>A service is registered on the local agent. The health check is passing locally. But &lt;code>dig myservice.service.consul&lt;/code> returns nothing, the load balancer has no targets, and the catalog API shows no instances. The servers have a leader, gossip is healthy, Raft metrics look normal. The problem is between the agent and the catalog.&lt;/p>
&lt;p>Consul&amp;rsquo;s catalog is authoritative on the servers. Each agent maintains its own local state (registered services, health checks, node metadata) and periodically reconciles that state with the server catalog through a background process called anti-entropy sync. The agent treats its local view as authoritative and pushes changes to the catalog. When this sync fails, the catalog retains the last-known state it received from that agent. A service registered locally never appears cluster-wide. A service deregistered locally continues to show up in DNS and API queries.&lt;/p></description></item><item><title>Consul blocking query accumulation: leaked watches that pile up goroutines</title><link>https://www.netdata.cloud/guides/consul/consul-blocking-query-accumulation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-blocking-query-accumulation/</guid><description>&lt;h1 id="consul-blocking-query-accumulation-leaked-watches-that-pile-up-goroutines">Consul blocking query accumulation: leaked watches that pile up goroutines&lt;/h1>
&lt;p>Goroutine count on your Consul servers is creeping upward. File descriptor usage follows the same trajectory. Neither reverses during quiet periods. Eventually the server hits its FD limit or goroutine scheduling overhead degrades everything: new connections are refused, health checks stop reaching the catalog, anti-entropy sync stalls, and the cluster cascades.&lt;/p>
&lt;p>The culprit is leaked blocking queries. Consul&amp;rsquo;s watch mechanism and any client using the HTTP long-polling API (&lt;code>?index=&lt;/code> and &lt;code>?wait=&lt;/code> parameters) holds a goroutine and an FD on the server for up to the wait timeout (5 minutes default, 10 minutes maximum). A fleet of consul-template instances, application-level watchers, or service mesh control planes can open thousands of concurrent blocking queries. This is fine when queries are properly opened and closed. The leak starts when they are not: a client crashes without cleanly closing its connection, a consul-template bug prevents goroutine cleanup, a watch configuration exceeds the server&amp;rsquo;s tracking capacity, or a go-memdb WatchSet bug causes goroutine proliferation.&lt;/p></description></item><item><title>Consul catalog bloat: too many services and checks slowing everything down</title><link>https://www.netdata.cloud/guides/consul/consul-catalog-bloat/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-catalog-bloat/</guid><description>&lt;h1 id="consul-catalog-bloat-too-many-services-and-checks-slowing-everything-down">Consul catalog bloat: too many services and checks slowing everything down&lt;/h1>
&lt;p>DNS latency is creeping up. API calls to &lt;code>/v1/catalog/services&lt;/code> take longer than they used to. Server restarts that used to take 20 seconds now take minutes. The Raft snapshot on disk keeps growing, and so does server RSS. No single failure point, no stack trace, no page. Just a slow, compounding drag on everything Consul does.&lt;/p>
&lt;p>This is catalog bloat. Registered service instances and health checks grow the in-memory state store, the on-disk Raft snapshot, anti-entropy reconciliation work, and catalog-scan latency for every DNS and API query. The cost compounds week over week until a snapshot creation tips a leader election, a restart takes long enough to miss a deploy window, or DNS p99 crosses the threshold where applications time out on service discovery.&lt;/p></description></item><item><title>Consul catalog poisoning: services registered from unknown nodes</title><link>https://www.netdata.cloud/guides/consul/consul-catalog-poisoning-unknown-registration/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-catalog-poisoning-unknown-registration/</guid><description>&lt;h1 id="consul-catalog-poisoning-services-registered-from-unknown-nodes">Consul catalog poisoning: services registered from unknown nodes&lt;/h1>
&lt;p>A service instance appears in your catalog on a node you have never deployed. Its &lt;code>ServiceAddress&lt;/code> points somewhere you do not control. That is catalog poisoning: a registration that did not come from your infrastructure.&lt;/p>
&lt;p>Without strict ACLs, any process that can reach the Consul HTTP API can call the registration endpoints and add an instance for any service name. Consumers resolving discovery through Consul DNS or the catalog and health APIs will route a fraction of their traffic to that address. In a non-Connect deployment there is nothing in the data path that verifies the instance is what it claims to be.&lt;/p></description></item><item><title>Consul client rpc failed: agents alive but the catalog is going stale</title><link>https://www.netdata.cloud/guides/consul/consul-client-rpc-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-client-rpc-failed/</guid><description>&lt;h1 id="consul-client-rpc-failed-agents-alive-but-the-catalog-is-going-stale">Consul client rpc failed: agents alive but the catalog is going stale&lt;/h1>
&lt;p>A client agent&amp;rsquo;s &lt;code>consul.client.rpc.failed&lt;/code> counter ticks up. The agent process is running, gossip reports the node as &lt;code>alive&lt;/code>, local health checks execute on schedule, and &lt;code>consul members&lt;/code> lists the node as healthy. The server cluster looks fine: leadership is stable, Raft commit times are normal, and there is no election noise. Nothing on the standard dashboard is red.&lt;/p></description></item><item><title>Consul Connect "no healthy upstream": Envoy has no endpoints to route to</title><link>https://www.netdata.cloud/guides/consul/consul-envoy-no-healthy-upstream/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-envoy-no-healthy-upstream/</guid><description>&lt;h1 id="consul-connect-no-healthy-upstream-envoy-has-no-endpoints-to-route-to">Consul Connect &amp;ldquo;no healthy upstream&amp;rdquo;: Envoy has no endpoints to route to&lt;/h1>
&lt;p>When an application in a Consul Connect mesh receives HTTP 503 with &lt;code>response_flags=UH&lt;/code> in the Envoy access log, Envoy&amp;rsquo;s router rejected the request because it had zero eligible endpoints for the upstream cluster. The downstream service is fine, the network is fine, and the request never left the sidecar. Somewhere between the Envoy admin port, the local Consul agent, the Consul servers, and the upstream instances, the endpoint set has gone empty, all-unhealthy, or silently ejected.&lt;/p></description></item><item><title>Consul Connect CA rotation failure: a root roll that never finished</title><link>https://www.netdata.cloud/guides/consul/consul-connect-ca-rotation-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-connect-ca-rotation-failure/</guid><description>&lt;h1 id="consul-connect-ca-rotation-failure-a-root-roll-that-never-finished">Consul Connect CA rotation failure: a root roll that never finished&lt;/h1>
&lt;p>Consul Connect CA root rotation promotes a new root, cross-signs the new intermediate against the old root, lets both roots coexist while leaf certificates roll over, and retires the old root only after the last leaf signed by it expires. When that rollout never completes, the first thing you notice is intermittent mTLS failures between specific service pairs, hours after the rotation was triggered, with no obvious network or config change to blame.&lt;/p></description></item><item><title>Consul Connect certificate expired: mTLS handshakes failing across the mesh</title><link>https://www.netdata.cloud/guides/consul/consul-connect-certificate-expiry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-connect-certificate-expiry/</guid><description>&lt;h1 id="consul-connect-certificate-expired-mtls-handshakes-failing-across-the-mesh">Consul Connect certificate expired: mTLS handshakes failing across the mesh&lt;/h1>
&lt;p>Consul Connect relies on short-lived leaf certificates for service-to-service mTLS. When those certificates expire without rotating, the failure is cliff-edged: handshakes that worked seconds ago start failing, Envoy logs fill with TLS errors, and dependent services stop communicating. The blast radius grows over minutes to hours as each leaf certificate reaches its own expiry independently.&lt;/p>
&lt;p>The first symptom is usually intermittent: a specific upstream starts returning connection resets or 5xx errors. As more leaves expire, failures fan out across service pairs until the mesh is broadly broken. Expiry alone is a countdown with no early warning. The signal that matters earlier is renewal success: are leaves actually being re-issued, or is the rotation pipeline silently broken while the clock runs out?&lt;/p></description></item><item><title>Consul critical health checks spiking: real outage or broken checks?</title><link>https://www.netdata.cloud/guides/consul/consul-service-critical-health-checks/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-service-critical-health-checks/</guid><description>&lt;h1 id="consul-critical-health-checks-spiking-real-outage-or-broken-checks">Consul critical health checks spiking: real outage or broken checks?&lt;/h1>
&lt;p>The alert fires when &lt;code>/v1/health/state/critical&lt;/code> jumps from single digits to dozens or hundreds of checks in minutes. The page says &amp;ldquo;services are unhealthy,&amp;rdquo; but that is ambiguous: either real services are down, or the checks themselves are broken. Pick the wrong fork and you burn an hour on the wrong dependency.&lt;/p>
&lt;p>The first decision is not &amp;ldquo;what failed.&amp;rdquo; It is &amp;ldquo;is this a service-level event, a node-level event, or a monitoring-level event?&amp;rdquo; The cheapest signal is the check &lt;code>Output&lt;/code> text field, which carries the raw error from check execution. It holds strings like &lt;code>connection refused&lt;/code>, &lt;code>i/o timeout&lt;/code>, &lt;code>certificate has expired&lt;/code>. That field usually resolves the diagnosis in seconds, and most teams ignore it until their second or third incident.&lt;/p></description></item><item><title>Consul cross-datacenter query failure: prepared-query failover masking a DC outage</title><link>https://www.netdata.cloud/guides/consul/consul-cross-dc-query-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-cross-dc-query-failure/</guid><description>&lt;h1 id="consul-cross-datacenter-query-failure-prepared-query-failover-masking-a-dc-outage">Consul cross-datacenter query failure: prepared-query failover masking a DC outage&lt;/h1>
&lt;p>A Consul prepared query with cross-DC failover does exactly what it was configured to do: it silently forwards lookups to a remote DC when local results are empty. Consumers receive valid service discovery results, applications keep connecting, and no error fires. The local DC may have zero healthy instances, and nobody on the consumer side knows.&lt;/p>
&lt;p>The DNS response is identical whether nodes came from the local DC or a remote one. There is no EDNS flag, no source record, no indicator in the answer. Receiving results does not mean local health.&lt;/p></description></item><item><title>Consul DeregisterCriticalServiceAfter: instances vanishing from the catalog</title><link>https://www.netdata.cloud/guides/consul/consul-deregister-critical-service-after/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-deregister-critical-service-after/</guid><description>&lt;h1 id="consul-deregistercriticalserviceafter-instances-vanishing-from-the-catalog">Consul DeregisterCriticalServiceAfter: instances vanishing from the catalog&lt;/h1>
&lt;p>A service that was merely slow to recover has disappeared from the Consul catalog. Healthy-instance count dropped to zero. DNS returns empty results. Load balancers have no targets. Yet the service process is running, the host is up, and the Consul agent on that host is healthy and gossiping. DCSA did exactly what you configured.&lt;/p>
&lt;p>&lt;code>DeregisterCriticalServiceAfter&lt;/code> (DCSA) is an agent-side reaper that removes a service instance and its checks from the catalog once the check has been critical for a configured duration. It is deliberate garbage collection for stale entries. Leave it unset and critical checks plus their output strings accumulate unbounded in the state store, growing memory and slowing snapshots.&lt;/p></description></item><item><title>Consul DNS latency high: slow lookups stalling connections and failovers</title><link>https://www.netdata.cloud/guides/consul/consul-dns-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-dns-latency-high/</guid><description>&lt;h1 id="consul-dns-latency-high-slow-lookups-stalling-connections-and-failovers">Consul DNS latency high: slow lookups stalling connections and failovers&lt;/h1>
&lt;p>Slow Consul DNS lookups rarely look like a DNS problem at first. Applications see connection timeouts, retries, and slow startup. Load balancers see health check flapping. Service mesh sidecars see upstream resolution failures. The DNS layer is invisible to most teams until &lt;code>consul.dns.domain_query&lt;/code> crosses the page threshold.&lt;/p>
&lt;p>The first decision in a Consul DNS latency incident is structural: is the slowness in the server cluster (Raft, disk, catalog size) or in the DNS path itself (TTL, agent CPU, query complexity, downstream resolvers)? The signals are all in agent telemetry, but you have to know which ones to correlate.&lt;/p></description></item><item><title>Consul DNS SERVFAIL: service discovery is broken for your applications</title><link>https://www.netdata.cloud/guides/consul/consul-dns-servfail/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-dns-servfail/</guid><description>&lt;h1 id="consul-dns-servfail-service-discovery-is-broken-for-your-applications">Consul DNS SERVFAIL: service discovery is broken for your applications&lt;/h1>
&lt;p>SERVFAIL from Consul&amp;rsquo;s DNS interface on port 8600 means service discovery is broken for every application resolving names through it. Applications see connection refused, timeouts, and cascading failovers. The symptom is broad, but the root cause lives along one of four layers: the downstream resolver (if present), the Consul agent&amp;rsquo;s built-in DNS server, the agent-to-server RPC channel, or the server&amp;rsquo;s catalog state.&lt;/p></description></item><item><title>Consul Go GC pauses: stop-the-world stalls that disturb Raft timing</title><link>https://www.netdata.cloud/guides/consul/consul-gc-pause-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-gc-pause-high/</guid><description>&lt;h1 id="consul-go-gc-pauses-stop-the-world-stalls-that-disturb-raft-timing">Consul Go GC pauses: stop-the-world stalls that disturb Raft timing&lt;/h1>
&lt;p>Consul servers are Go binaries, and Go&amp;rsquo;s garbage collector still has stop-the-world phases despite doing most of its work concurrently. On a server with a large live heap and a high allocation rate, those pauses can stretch into hundreds of milliseconds. The Raft leader sends heartbeats from a goroutine in the same process. When a pause runs long enough, the leader stops heartbeating; followers see &lt;code>consul.raft.leader.lastContact&lt;/code> climbing toward the election timeout and start a new election. Writes block for the duration.&lt;/p></description></item><item><title>Consul goroutine count climbing: the leak behind slow resource exhaustion</title><link>https://www.netdata.cloud/guides/consul/consul-goroutine-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-goroutine-leak/</guid><description>&lt;h1 id="consul-goroutine-count-climbing-the-leak-behind-slow-resource-exhaustion">Consul goroutine count climbing: the leak behind slow resource exhaustion&lt;/h1>
&lt;p>&lt;code>consul.runtime.num_goroutines&lt;/code> is a gauge of how many concurrent operations the process is juggling: every blocking query, every gRPC or xDS stream, every health check handler runs inside a goroutine. In a healthy cluster the number is baseline-dependent but flat. When it creeps upward hour over hour, something is spawning goroutines and not cleaning them up.&lt;/p>
&lt;p>The absolute number is a distraction. A medium cluster idles between a few hundred and 20,000 goroutines; a large Connect deployment legitimately runs higher. The trend is the signal. Monotonic growth without a matching increase in services, watchers, or sidecar count is a leak, regardless of where the number sits.&lt;/p></description></item><item><title>Consul gossip encryption key mismatch: a botched keyring rotation splits the pool</title><link>https://www.netdata.cloud/guides/consul/consul-gossip-encryption-key-mismatch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-gossip-encryption-key-mismatch/</guid><description>&lt;h1 id="consul-gossip-encryption-key-mismatch-a-botched-keyring-rotation-splits-the-pool">Consul gossip encryption key mismatch: a botched keyring rotation splits the pool&lt;/h1>
&lt;p>You rotated the gossip encryption key. Nodes are dropping from the member list, showing as &amp;ldquo;failed&amp;rdquo; or &amp;ldquo;suspect&amp;rdquo; to peers, and service discovery is degrading. The network is fine. The agents are running. Raft may even be stable. But the gossip pool has split.&lt;/p>
&lt;p>The root cause is at the protocol layer. Consul&amp;rsquo;s Serf gossip protocol encrypts every message with a symmetric key. When nodes hold different keys, they cannot decrypt each other&amp;rsquo;s gossip messages. Each side sees the other as unresponsive, and failure detection kicks in. This looks exactly like a network partition, but pings between the affected hosts succeed.&lt;/p></description></item><item><title>Consul gossip flapping: nodes oscillating between alive, suspect, and failed</title><link>https://www.netdata.cloud/guides/consul/consul-gossip-flapping-suspect-nodes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-gossip-flapping-suspect-nodes/</guid><description>&lt;h1 id="consul-gossip-flapping-nodes-oscillating-between-alive-suspect-and-failed">Consul gossip flapping: nodes oscillating between alive, suspect, and failed&lt;/h1>
&lt;p>When Consul nodes rapidly flip between alive, suspect, and failed in the gossip pool, every state transition fires a membership event that downstream consumers react to. Load balancers pull endpoints in and out, consul-template reloads configurations, and Envoy sidecars receive new xDS endpoint pushes. A flapping node generates more operational noise than a cleanly failed one.&lt;/p>
&lt;p>The Serf gossip protocol converges under stable network conditions. Under packet loss, CPU starvation, or partial partitions, it oscillates instead. The symptom appears in &lt;code>consul members&lt;/code> as nodes cycling between alive, suspect, and failed; in Serf metrics as spikes in &lt;code>consul.serf.lan.member.flaps&lt;/code>; and in downstream systems as configuration reload storms that track the flap cadence.&lt;/p></description></item><item><title>Consul gossip storm after mass recovery: rejoin floods and anti-entropy spikes</title><link>https://www.netdata.cloud/guides/consul/consul-gossip-storm-mass-recovery/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-gossip-storm-mass-recovery/</guid><description>&lt;h1 id="consul-gossip-storm-after-mass-recovery-rejoin-floods-and-anti-entropy-spikes">Consul gossip storm after mass recovery: rejoin floods and anti-entropy spikes&lt;/h1>
&lt;p>After a mass recovery event (AZ healing, partition recovery, fleet-wide rolling restart), dozens or hundreds of Consul agents rejoin the LAN gossip pool in a compressed window. Each returning agent generates member-join events that every other member processes, and triggers anti-entropy sync to reconcile its local service registrations against the server catalog. The coordinated spike hits Serf message queues, catalog registration writes, and Raft commits simultaneously.&lt;/p></description></item><item><title>Consul health check flapping: the passing/critical oscillation that churns the catalog</title><link>https://www.netdata.cloud/guides/consul/consul-health-check-flapping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-health-check-flapping/</guid><description>&lt;h1 id="consul-health-check-flapping-the-passingcritical-oscillation-that-churns-the-catalog">Consul health check flapping: the passing/critical oscillation that churns the catalog&lt;/h1>
&lt;p>Every health check state transition is a catalog write. The client agent runs the check locally, detects a status change, and pushes the update to the leader. The leader commits a Raft log entry, the FSM applies it, per-service caches invalidate, blocking queries watching that service return, and every downstream consumer (load balancer, Envoy control plane, consul-template, DNS resolver) re-evaluates. One flapping check is an annoyance. A dozen flapping in unison saturates the Raft write pipeline and looks indistinguishable from a registration storm.&lt;/p></description></item><item><title>Consul intention denied: service-to-service traffic blocked by policy</title><link>https://www.netdata.cloud/guides/consul/consul-intention-denied-connection/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-intention-denied-connection/</guid><description>&lt;h1 id="consul-intention-denied-service-to-service-traffic-blocked-by-policy">Consul intention denied: service-to-service traffic blocked by policy&lt;/h1>
&lt;p>Service A can no longer reach service B. Envoy sidecars log 403 responses or connection resets. The Consul server cluster looks healthy: leader stable, Raft committing, gossip intact, certificates within their lifetime. Traffic that worked an hour ago is now blocked.&lt;/p>
&lt;p>This is almost always policy, not infrastructure. Consul Connect intentions are the mesh authorization layer, evaluated at connection establishment and enforced by Envoy RBAC filters. When an intention denies a connection, the data plane is doing exactly what it was configured to do. The diagnostic job is to determine whether the denial is correct (policy working as designed) or a misconfiguration: wrong identity, missing allow, precedence mistake, or stale cache.&lt;/p></description></item><item><title>Consul KV store saturation: using the KV store as a database</title><link>https://www.netdata.cloud/guides/consul/consul-kv-store-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-kv-store-saturation/</guid><description>&lt;h1 id="consul-kv-store-saturation-using-the-kv-store-as-a-database">Consul KV store saturation: using the KV store as a database&lt;/h1>
&lt;p>Consul KV writes are slow. Raft commit times are climbing. Service registration and health check updates lag. DNS queries for service discovery take longer than usual. This combination often points to an application treating Consul&amp;rsquo;s KV store as a general-purpose database: high-frequency writes, large values, or deep key trees.&lt;/p>
&lt;p>The problem is architectural, not a tuning issue. The KV store is part of Consul&amp;rsquo;s Raft finite state machine. Every KV write is a Raft log entry that must be committed through the leader, replicated to a quorum of servers, and applied to the in-memory state store on every server. Every KV value is included in every Raft snapshot. When an application treats the KV store as a database, the write load backs up the entire Raft pipeline, and every Consul subsystem that depends on Raft degrades with it.&lt;/p></description></item><item><title>Consul KV value too large: the 512KB limit and why approaching it hurts</title><link>https://www.netdata.cloud/guides/consul/consul-kv-value-too-large/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-kv-value-too-large/</guid><description>&lt;h1 id="consul-kv-value-too-large-the-512kb-limit-and-why-approaching-it-hurts">Consul KV value too large: the 512KB limit and why approaching it hurts&lt;/h1>
&lt;p>You pushed a value into Consul KV and got HTTP 413, or your server logs are filling with lines like &lt;code>Request body(524401 bytes) too large, max size: 524288 bytes&lt;/code>. That is the obvious failure: a single value crossed the default 512KB ceiling. The less obvious failure is that values well under the limit are still expensive, because every KV byte is replicated through Raft to every server and re-emitted in every snapshot.&lt;/p></description></item><item><title>Consul KV write latency high: every write is a Raft commit</title><link>https://www.netdata.cloud/guides/consul/consul-kv-write-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-kv-write-latency-high/</guid><description>&lt;h1 id="consul-kv-write-latency-high-every-write-is-a-raft-commit">Consul KV write latency high: every write is a Raft commit&lt;/h1>
&lt;p>A KV PUT that used to return in 20ms is now taking 500ms, 2s, or timing out entirely. Applications stall on locks, configuration writes queue, and leader election alerts may be firing. Consul&amp;rsquo;s KV is not a standalone subsystem. Every KV write is a Raft log entry that must be replicated to a quorum of servers, fsynced to disk on each, and applied to the in-memory state machine before the API call returns. KV write latency is a direct reflection of Raft commit health.&lt;/p></description></item><item><title>Consul leader election storm: repeated elections and rolling write outages</title><link>https://www.netdata.cloud/guides/consul/consul-leader-election-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-leader-election-storm/</guid><description>&lt;h1 id="consul-leader-election-storm-repeated-elections-and-rolling-write-outages">Consul leader election storm: repeated elections and rolling write outages&lt;/h1>
&lt;p>Your Consul cluster has a leader. Writes are still failing. The leader keeps changing. Applications see intermittent &amp;ldquo;no cluster leader&amp;rdquo; errors, watches reconnect, DNS returns stale results, and every few seconds a different server wins an election only to lose it again.&lt;/p>
&lt;!-- TODO: verify metric type. consul.raft.state.leader appears to be a gauge (1 when leader, 0 otherwise), not a counter. This article repeatedly describes it as a counter that "increments" or "climbs." If it is a gauge, monitoring should track transitions between 0 and 1, not counter increments. Same concern applies to consul.raft.state.candidate below. -->
&lt;p>This is a leader election storm. Unlike a clean one-time failover, the cluster never stabilizes. Each election blocks all writes for one to several seconds. When elections recur faster than the recovery window, the cluster is effectively write-unavailable while technically always having &amp;ldquo;a leader.&amp;rdquo; A naive alert on &amp;ldquo;no leader&amp;rdquo; stays silent. The real signal is recurrence: leadership transitions accumulating over time, each one a brief but real outage.&lt;/p></description></item><item><title>Consul leader stable but commits stalled: writes silently failing</title><link>https://www.netdata.cloud/guides/consul/consul-stalled-commit-index/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-stalled-commit-index/</guid><description>&lt;h1 id="consul-leader-stable-but-commits-stalled-writes-silently-failing">Consul leader stable but commits stalled: writes silently failing&lt;/h1>
&lt;p>You are paged for &amp;ldquo;Consul writes failing.&amp;rdquo; You check &lt;code>/v1/status/leader&lt;/code>: it returns an address. You check &lt;code>consul members&lt;/code>: all servers alive. Leadership has been stable for hours. Your &amp;ldquo;is there a leader?&amp;rdquo; alert is silent. Yet KV writes time out, service registrations hang, and sessions cannot be created.&lt;/p>
&lt;p>The leader is alive, holds leadership, and answers health probes. But its commit index is not advancing. Somewhere between the leader receiving a write and that write becoming visible in the FSM, the pipeline is stalled. The cluster is write-dead.&lt;/p></description></item><item><title>Consul lost quorum: Raft peers below the majority needed to elect a leader</title><link>https://www.netdata.cloud/guides/consul/consul-raft-peers-below-quorum/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-raft-peers-below-quorum/</guid><description>&lt;h1 id="consul-lost-quorum-raft-peers-below-the-majority-needed-to-elect-a-leader">Consul lost quorum: Raft peers below the majority needed to elect a leader&lt;/h1>
&lt;p>A Consul cluster has lost quorum when the number of voting Raft peers drops below the majority required to elect a leader. For a 3-server cluster that means fewer than 2 voters. For a 5-server cluster, fewer than 3. With no leader, every write fails: service registrations, health-check state updates, KV writes, session creation, ACL token creation. Reads served in stale mode still return data, but that data is frozen at the moment quorum was lost.&lt;/p></description></item><item><title>Consul Monitoring</title><link>https://www.netdata.cloud/monitoring-101/consul-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/consul-monitoring/</guid><description>&lt;h2 id="consul-monitoring">Consul Monitoring&lt;/h2>
&lt;h3 id="what-is-consul">What Is Consul?&lt;/h3>
&lt;p>Consul is a highly scalable and distributed service networking platform that enables users to manage and secure their microservices connections. Developed by HashiCorp, Consul offers key features such as service discovery, health checking, load balancing, and a secure service mesh. &lt;a href="https://www.consul.io/">Learn more about Consul&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-consul-with-netdata">Monitoring Consul With Netdata&lt;/h3>
&lt;p>Netdata provides a powerful and real-time monitoring solution for Consul, utilizing the go.d.plugin. This tool connects seamlessly to your Consul installation, enabling comprehensive monitoring by leveraging the &lt;a href="https://developer.hashicorp.com/consul/api-docs">Consul REST API&lt;/a>. By using Netdata, you gain instant access to crucial Consul metrics, helping you ensure optimal performance and quick identification of potential issues. To see Netdata in action, check our &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">Live Demo&lt;/a>.&lt;/p></description></item><item><title>Consul monitoring checklist: the signals every production cluster needs</title><link>https://www.netdata.cloud/guides/consul/consul-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-monitoring-checklist/</guid><description>&lt;h1 id="consul-monitoring-checklist-the-signals-every-production-cluster-needs">Consul monitoring checklist: the signals every production cluster needs&lt;/h1>
&lt;p>Consul runs three subsystems concurrently: Raft consensus, Serf gossip, and a catalog state machine backed by BoltDB. Each fails independently. A monitoring setup that checks only &amp;ldquo;is there a leader&amp;rdquo; and &amp;ldquo;are all members alive&amp;rdquo; will miss the most common production incidents: slow disk causing leader churn, client agents losing RPC connectivity while gossip stays green, blocking query leaks, and Connect certificate rotation failures.&lt;/p></description></item><item><title>Consul monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/consul/consul-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-monitoring-maturity-model/</guid><description>&lt;h1 id="consul-monitoring-maturity-model-from-survival-to-expert">Consul monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Consul runs three independent subsystems that must all be healthy for the cluster to function: Raft consensus, Serf gossip, and the catalog state machine with its anti-entropy sync pipeline. Most teams monitor one or two well and discover the third during an incident.&lt;/p>
&lt;p>The model is cumulative. Each level adds signals to the previous one; you cannot skip to Mature without the Operational baseline in place. A team with composite-pattern detection at Expert but no client-agent RPC monitoring at Operational is blind to the most common silent failure mode: the catalog drifting from reality while every server-side metric looks healthy.&lt;/p></description></item><item><title>Consul on EBS: burst-credit exhaustion and the sudden latency cliff</title><link>https://www.netdata.cloud/guides/consul/consul-ebs-burst-credit-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-ebs-burst-credit-exhaustion/</guid><description>&lt;h1 id="consul-on-ebs-burst-credit-exhaustion-and-the-sudden-latency-cliff">Consul on EBS: burst-credit exhaustion and the sudden latency cliff&lt;/h1>
&lt;p>Consul is stable. Raft commit times under 10ms, leadership steady, discovery in single-digit milliseconds. Then, with no deploy, no config change, and no traffic spike, writes start timing out. &lt;code>consul.raft.commitTime&lt;/code> jumps from 5ms to 500ms. Followers report &lt;code>lastContact&lt;/code> climbing toward the election timeout. Within seconds, the cluster elects a new leader, then another, then another. Every write fails with &amp;ldquo;no cluster leader.&amp;rdquo;&lt;/p></description></item><item><title>Consul Permission denied (403): authorized token, missing permission</title><link>https://www.netdata.cloud/guides/consul/consul-permission-denied-403/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-permission-denied-403/</guid><description>&lt;h1 id="consul-permission-denied-403-authorized-token-missing-permission">Consul Permission denied (403): authorized token, missing permission&lt;/h1>
&lt;p>A 403 &amp;ldquo;Permission denied&amp;rdquo; from the Consul HTTP API means the request reached the server with a token it recognizes, but policy evaluation denied the operation. The token exists, resolves, and is not expired. It lacks the rule covering the resource being touched: a service, a KV prefix, a node, an operator endpoint. This differs from &amp;ldquo;ACL not found&amp;rdquo;, where the SecretID is unknown to the server, and the two need different responses.&lt;/p></description></item><item><title>Consul raft commitTime high: the write pipeline is slowing down</title><link>https://www.netdata.cloud/guides/consul/consul-raft-commit-time-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-raft-commit-time-high/</guid><description>&lt;h1 id="consul-raft-committime-high-the-write-pipeline-is-slowing-down">Consul raft commitTime high: the write pipeline is slowing down&lt;/h1>
&lt;p>&lt;code>consul.raft.commitTime&lt;/code> is the single best indicator of Raft health on a Consul server. It is a leader-only timer that folds disk write latency, replication latency, and FSM apply time into one number. When it climbs, every Consul write slows: service registrations queue, KV writes block, health check updates lag, and consumers downstream of the catalog start timing out. The disk-bound leader cannot get log entries committed fast enough, followers lose touch, and the next step is leader elections, during which the cluster cannot commit writes at all.&lt;/p></description></item><item><title>Consul Raft data directory full: the server that can no longer write</title><link>https://www.netdata.cloud/guides/consul/consul-raft-data-dir-disk-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-raft-data-dir-disk-full/</guid><description>&lt;h1 id="consul-raft-data-directory-full-the-server-that-can-no-longer-write">Consul Raft data directory full: the server that can no longer write&lt;/h1>
&lt;p>The server&amp;rsquo;s &lt;code>data_dir&lt;/code> volume hits 100%. Raft cannot append to &lt;code>raft.db&lt;/code> and cannot fsync snapshot files. The server either refuses writes or crashes. If the affected server is the leader, the commit index stalls and followers drift toward an election.&lt;/p>
&lt;p>The most common cause is not disk hardware failure. It is a snapshot that keeps failing. Raft truncates its log only after a successful snapshot, so each failed attempt leaves &lt;code>raft.db&lt;/code> larger than it should be. Eventually the log fills the volume.&lt;/p></description></item><item><title>Consul raft lastContact rising: followers drifting toward an election</title><link>https://www.netdata.cloud/guides/consul/consul-raft-last-contact-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-raft-last-contact-high/</guid><description>&lt;h1 id="consul-raft-lastcontact-rising-followers-drifting-toward-an-election">Consul raft lastContact rising: followers drifting toward an election&lt;/h1>
&lt;p>&lt;code>consul.raft.leader.lastContact&lt;/code> is the predictive signal before a Raft election fires. It measures the elapsed time since the leader last successfully contacted each follower server. When this value rises, a follower is drifting toward the election timeout. Cross that threshold and the follower starts a new election: writes stall, clients see &amp;ldquo;no cluster leader&amp;rdquo; errors, and downstream consumers retry in a thundering herd.&lt;/p></description></item><item><title>Consul Raft log divergence: catching a corrupt follower before it wins an election</title><link>https://www.netdata.cloud/guides/consul/consul-raft-log-divergence/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-raft-log-divergence/</guid><description>&lt;h1 id="consul-raft-log-divergence-catching-a-corrupt-follower-before-it-wins-an-election">Consul Raft log divergence: catching a corrupt follower before it wins an election&lt;/h1>
&lt;p>A divergent Consul follower can gossip normally, answer Serf probes, and appear in &lt;code>consul members&lt;/code> as alive while its persisted Raft state is corrupt or inconsistent. The leader replicates to the healthy majority and clients see correct data. The divergence becomes catastrophic only when that follower wins an election and serves state the cluster never agreed on.&lt;/p></description></item><item><title>Consul registration storm: catalog churn overwhelming Raft</title><link>https://www.netdata.cloud/guides/consul/consul-catalog-churn-registration-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-catalog-churn-registration-storm/</guid><description>&lt;h1 id="consul-registration-storm-catalog-churn-overwhelming-raft">Consul registration storm: catalog churn overwhelming Raft&lt;/h1>
&lt;p>The complaint comes in as &amp;ldquo;Consul is slow.&amp;rdquo; DNS lookups lag, service discovery returns stale results, HTTP API writes take longer than usual. The cluster has a leader, gossip is healthy, no node is down. The Raft commit time metric is climbing. The root cause is often hiding in the catalog registration rate.&lt;/p>
&lt;p>Every service registration and deregistration is a Raft write. The leader appends the entry to its log, persists it (fsync), replicates to followers (who also persist), waits for quorum, and applies the entry to the in-memory state store. This is the same pipeline that handles KV writes, health-check state updates, session creation, and ACL operations. When the registration rate is high enough, it saturates the pipeline and every consumer of the consensus layer pays the cost.&lt;/p></description></item><item><title>Consul serf queue backlog: an agent falling behind on gossip</title><link>https://www.netdata.cloud/guides/consul/consul-gossip-queue-backlog/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-gossip-queue-backlog/</guid><description>&lt;h1 id="consul-serf-queue-backlog-an-agent-falling-behind-on-gossip">Consul serf queue backlog: an agent falling behind on gossip&lt;/h1>
&lt;p>A Consul agent whose &lt;code>consul.serf.queue.Event&lt;/code>, &lt;code>consul.serf.queue.Intent&lt;/code>, or &lt;code>consul.serf.queue.Query&lt;/code> metric stays above zero is no longer keeping up with gossip. In steady state all three should read zero. Transient spikes during bulk joins, leaves, and rolling restarts are expected and self-drain within a few gossip intervals. The problem begins when the spike does not drain.&lt;/p>
&lt;p>Serf&amp;rsquo;s queue depth is the observable proxy for the node&amp;rsquo;s internal health score. A sustained non-zero value means the agent is receiving gossip faster than its event loop can process. Downstream effects start subtle: failure detection latency rises, join and leave intents propagate slowly, and the node&amp;rsquo;s own probe replies arrive late at peers. Then the feedback loop engages.&lt;/p></description></item><item><title>Consul serfHealth check failing: the node-level check behind mass deregistration</title><link>https://www.netdata.cloud/guides/consul/consul-serfhealth-check-failing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-serfhealth-check-failing/</guid><description>&lt;h1 id="consul-serfhealth-check-failing-the-node-level-check-behind-mass-deregistration">Consul serfHealth check failing: the node-level check behind mass deregistration&lt;/h1>
&lt;p>Every Consul node carries an automatic health check called &lt;code>serfHealth&lt;/code>. It is not user-defined, it cannot be removed, and it reflects the node&amp;rsquo;s membership in the Serf LAN gossip pool. When this check goes critical, Consul treats the entire node and every service registered on it as unhealthy, regardless of what individual service checks report.&lt;/p>
&lt;p>This is the most commonly misdiagnosed failure in Consul. Operators see a service &amp;ldquo;go down&amp;rdquo; across all instances on a host, investigate the service checks, and find them still passing. The services are fine. The node&amp;rsquo;s gossip membership broke, and &lt;code>serfHealth&lt;/code> cascaded that failure into every service on the node.&lt;/p></description></item><item><title>Consul server in failed state: reading consul members during an incident</title><link>https://www.netdata.cloud/guides/consul/consul-server-failed-in-gossip/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-server-failed-in-gossip/</guid><description>&lt;h1 id="consul-server-in-failed-state-reading-consul-members-during-an-incident">Consul server in failed state: reading consul members during an incident&lt;/h1>
&lt;p>You ran &lt;code>consul members&lt;/code> during an incident and one of your servers shows &lt;code>failed&lt;/code>. Before fixing anything, understand what the status means. Gossip failure and Raft failure are separate subsystems. A server in &lt;code>failed&lt;/code> state is not answering gossip probes, but it may still be a voting Raft peer, or its agent process may still be running but starved of resources.&lt;/p></description></item><item><title>Consul server memory climbing toward OOM: heap, catalog, and snapshots</title><link>https://www.netdata.cloud/guides/consul/consul-server-memory-growth-oom/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-server-memory-growth-oom/</guid><description>&lt;h1 id="consul-server-memory-climbing-toward-oom-heap-catalog-and-snapshots">Consul server memory climbing toward OOM: heap, catalog, and snapshots&lt;/h1>
&lt;p>A Consul server&amp;rsquo;s RSS is climbing and OOMKill is close. The cause could be legitimate catalog growth, a goroutine leak, snapshot amplification, or a runtime pathology. Memory growth in Consul usually mixes real state growth, Go runtime behavior, and at least one subsystem leaking.&lt;/p>
&lt;p>Go&amp;rsquo;s garbage collector does not return memory to the OS immediately. RSS of roughly 2x the live heap is normal and is not a leak. What you should investigate is monotonic RSS growth that survives a GC, or live heap (&lt;code>consul.runtime.alloc_bytes&lt;/code>) that does not return to a baseline after a transient event.&lt;/p></description></item><item><title>Consul service has zero healthy instances: discovery returns nothing</title><link>https://www.netdata.cloud/guides/consul/consul-service-zero-healthy-instances/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-service-zero-healthy-instances/</guid><description>&lt;h1 id="consul-service-has-zero-healthy-instances-discovery-returns-nothing">Consul service has zero healthy instances: discovery returns nothing&lt;/h1>
&lt;p>A consumer calls &lt;code>/v1/health/service/&amp;lt;name&amp;gt;?passing=true&lt;/code> and gets &lt;code>[]&lt;/code>. DNS lookups for &lt;code>&amp;lt;name&amp;gt;.service.consul&lt;/code> return nothing. Every downstream consumer treats the service as gone. This is total unavailability for that one service, even if the rest of the Consul cluster is healthy.&lt;/p>
&lt;p>Two consumer surfaces are affected at once. API consumers see the empty array directly. DNS consumers (port 8600) get a negative response, and clients may cache it. Recovery from zero means both fixing the checks and flushing the negative caches downstream.&lt;/p></description></item><item><title>Consul session invalidation: distributed locks releasing across the cluster</title><link>https://www.netdata.cloud/guides/consul/consul-session-invalidation-lock-loss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-session-invalidation-lock-loss/</guid><description>&lt;h1 id="consul-session-invalidation-distributed-locks-releasing-across-the-cluster">Consul session invalidation: distributed locks releasing across the cluster&lt;/h1>
&lt;p>Sessions are Consul&amp;rsquo;s mechanism for binding distributed locks to node and application health. When a session is invalidated, every KV key it holds is acted on according to its &lt;code>Behavior&lt;/code>. With the default &lt;code>release&lt;/code>, the lock holder is cleared and the key&amp;rsquo;s &lt;code>ModifyIndex&lt;/code> increments. With &lt;code>delete&lt;/code>, the key is removed entirely. Every consumer watching that key reacts.&lt;/p>
&lt;p>A spike in session invalidation surfaces as a cluster-wide release of distributed locks. Leader elections fire, caches flush, configuration reloads trigger, and any logic keyed off &lt;code>ModifyIndex&lt;/code> change wakes up. The operator&amp;rsquo;s job is to identify which sessions invalidated, why, and whether the root cause is a single node, a renewal-path failure, or a gossip flap.&lt;/p></description></item><item><title>Consul slow disk causing Raft timeouts: the number-one cause of leader instability</title><link>https://www.netdata.cloud/guides/consul/consul-slow-disk-raft-timeouts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-slow-disk-raft-timeouts/</guid><description>&lt;h1 id="consul-slow-disk-causing-raft-timeouts-the-number-one-cause-of-leader-instability">Consul slow disk causing Raft timeouts: the number-one cause of leader instability&lt;/h1>
&lt;p>Your Consul cluster is cycling through leaders. Writes fail with &amp;ldquo;no cluster leader&amp;rdquo; errors. The cluster technically has a leader at any given moment, but constant re-elections make writes effectively unavailable. The logs show &lt;code>[WARN] raft: heartbeat timeout reached, starting election&lt;/code> repeating across all servers.&lt;/p>
&lt;p>The root cause is almost always the same: slow disk I/O on the volume backing the Raft data directory. Consul&amp;rsquo;s Raft log store performs an fsync on every appended entry. When the disk cannot complete those syncs quickly enough, heartbeats and commits slow down. Followers miss their heartbeat window and start new elections. The new leader lands on the same slow storage. The cycle repeats.&lt;/p></description></item><item><title>Consul snapshot size growing: state bloat, slow restores, and commit spikes</title><link>https://www.netdata.cloud/guides/consul/consul-snapshot-size-growth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-snapshot-size-growth/</guid><description>&lt;h1 id="consul-snapshot-size-growing-state-bloat-slow-restores-and-commit-spikes">Consul snapshot size growing: state bloat, slow restores, and commit spikes&lt;/h1>
&lt;p>Consul snapshots grow with the size of the FSM state: every registered service instance, every health check, every KV entry, every ACL token, every session. None of those individual writes looks expensive, so a team that registers a few new services per week will not see a spike in any single metric. What they get instead is a slow, monotonic climb in snapshot size that compounds across months until one day a snapshot save takes long enough to push Raft commit times past the heartbeat timeout, or a rejoined follower sits in snapshot restore long enough to fall out of the replication loop entirely.&lt;/p></description></item><item><title>Consul stale DNS queries: the agent is answering from cache</title><link>https://www.netdata.cloud/guides/consul/consul-dns-stale-queries/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-dns-stale-queries/</guid><description>&lt;h1 id="consul-stale-dns-queries-the-agent-is-answering-from-cache">Consul stale DNS queries: the agent is answering from cache&lt;/h1>
&lt;p>When &lt;code>consul.dns.stale_queries&lt;/code> climbs, the Consul agent is doing what it was designed to do: keep answering DNS from local state when it cannot get a fresh answer from a server. That behavior has been the DNS default since Consul 0.7. But &amp;ldquo;stale&amp;rdquo; means two very different things depending on how stale, and why.&lt;/p>
&lt;p>A stale answer two seconds old during a leader handoff is harmless availability. A stale answer five minutes old, still pointing at instances that crashed four minutes ago, is silently routing traffic to dead services. The counter alone cannot tell you which case you are in. Correlate it with agent-to-server RPC health, server health, and your consistency configuration.&lt;/p></description></item><item><title>Consul stale Raft peer: removing a failed server from the configuration</title><link>https://www.netdata.cloud/guides/consul/consul-remove-failed-raft-peer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-remove-failed-raft-peer/</guid><description>&lt;h1 id="consul-stale-raft-peer-removing-a-failed-server-from-the-configuration">Consul stale Raft peer: removing a failed server from the configuration&lt;/h1>
&lt;p>A Consul server has crashed, been decommissioned, or been replaced, but its entry still sits in the Raft configuration as a voting peer. Serf gossip already marks the node &lt;code>failed&lt;/code> or &lt;code>left&lt;/code>, yet &lt;code>GET /v1/operator/raft/configuration&lt;/code> still lists it as a voter. The dead peer keeps counting toward quorum even though it will never cast another vote.&lt;/p>
&lt;p>This is not always an emergency. A healthy cluster with autopilot enabled usually self-heals within tens of seconds. It becomes an emergency when autopilot is disabled, when the cleanup window has not elapsed, or when a second failure tips the cluster below the now-inflated quorum threshold.&lt;/p></description></item><item><title>Consul TLS not fully enforced: verify_incoming, verify_outgoing, and server hostname</title><link>https://www.netdata.cloud/guides/consul/consul-tls-verify-not-enforced/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-tls-verify-not-enforced/</guid><description>&lt;p>You have CA, certificate, and key files configured for Consul. The agent starts without errors, health checks pass, services register, DNS resolves. But the cluster may still accept plaintext or unauthenticated connections because &lt;code>verify_incoming&lt;/code>, &lt;code>verify_outgoing&lt;/code>, or &lt;code>verify_server_hostname&lt;/code> are missing, set to false, or placed in the wrong configuration stanza.&lt;/p>
&lt;p>Certificate files (&lt;code>ca_file&lt;/code>, &lt;code>cert_file&lt;/code>, &lt;code>key_file&lt;/code>) do not enforce TLS on their own. Consul requires explicit verification flags to reject plaintext connections and require client certificates. Without them, any network observer between agents and servers can read the catalog, KV values, and health check results, and any process that can reach the RPC port can impersonate a Consul agent.&lt;/p></description></item><item><title>Consul WAN federation down: a datacenter missing from the WAN pool</title><link>https://www.netdata.cloud/guides/consul/consul-wan-federation-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-wan-federation-down/</guid><description>&lt;h1 id="consul-wan-federation-down-a-datacenter-missing-from-the-wan-pool">Consul WAN federation down: a datacenter missing from the WAN pool&lt;/h1>
&lt;p>When an entire remote datacenter&amp;rsquo;s servers vanish from the WAN gossip pool, the symptom is usually downstream: cross-DC service lookups return empty results, prepared queries that failover across DCs stop working, and ACL or Connect CA replication between datacenters silently stalls. The Consul UI on the local DC may still look healthy because the local cluster is fine. The damage is in the federation layer.&lt;/p></description></item><item><title>Consul WAN link saturation: cross-DC latency triggering gossip suspicion</title><link>https://www.netdata.cloud/guides/consul/consul-wan-link-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-wan-link-saturation/</guid><description>&lt;h1 id="consul-wan-link-saturation-cross-dc-latency-triggering-gossip-suspicion">Consul WAN link saturation: cross-DC latency triggering gossip suspicion&lt;/h1>
&lt;p>Remote datacenter servers start oscillating between alive and suspect in the WAN gossip pool. Cross-DC RPCs fail with &amp;ldquo;no path to datacenter&amp;rdquo; errors. ACL replication lag in secondary DCs climbs into seconds. Cross-DC prepared queries time out. LAN Consul signals look fine, but the federation is broken.&lt;/p>
&lt;p>The usual cause is WAN link saturation. Consul&amp;rsquo;s WAN gossip pool uses a separate, more conservative set of timing defaults than LAN, but it is still vulnerable to high or spiking RTT. When the inter-DC link saturates, gossip probes time out, remote servers are marked suspect then failed, and WAN membership flaps. Each flap disrupts cross-DC RPC routing, which compounds the problem because RPC retries add traffic to the already-saturated link.&lt;/p></description></item><item><title>Consul xDS stream churn: Envoy sidecars running on stale configuration</title><link>https://www.netdata.cloud/guides/consul/consul-xds-stream-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-xds-stream-errors/</guid><description>&lt;h1 id="consul-xds-stream-churn-envoy-sidecars-running-on-stale-configuration">Consul xDS stream churn: Envoy sidecars running on stale configuration&lt;/h1>
&lt;p>Each Envoy sidecar in a Consul Connect mesh holds a long-lived gRPC xDS stream to a Consul server. That stream is the control plane: endpoint lists, intentions, and mTLS certificate rotations all flow across it. When the stream is healthy, the data plane converges within seconds of a catalog change. When it breaks or churns, the sidecar falls back to whatever Envoy cached last.&lt;/p></description></item><item><title>Contact Netdata</title><link>https://www.netdata.cloud/contact-us/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/contact-us/</guid><description/></item><item><title>containerd Containers</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/containerd-containers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/containerd-containers/</guid><description/></item><item><title>Containers</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/containers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/containers/</guid><description/></item><item><title>Conteg SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/conteg-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/conteg-snmp-traps/</guid><description/></item><item><title>Convertronic GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/convertronic-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/convertronic-gmbh-snmp-traps/</guid><description/></item><item><title>Cooler Master Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cooler-master-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cooler-master-co-ltd-snmp-traps/</guid><description/></item><item><title>Copper Mountain Communications Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/copper-mountain-communications-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/copper-mountain-communications-inc-snmp-traps/</guid><description/></item><item><title>CoreDNS</title><link>https://www.netdata.cloud/integrations/data-collection/networking/coredns/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/coredns/</guid><description/></item><item><title>CoreDNS /health vs /ready: the readiness-probe mistake that serves SERVFAIL</title><link>https://www.netdata.cloud/guides/coredns/coredns-health-vs-ready-probe/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-health-vs-ready-probe/</guid><description>&lt;h1 id="coredns-health-vs-ready-the-readiness-probe-mistake-that-serves-servfail">CoreDNS /health vs /ready: the readiness-probe mistake that serves SERVFAIL&lt;/h1>
&lt;p>Your CoreDNS pods look healthy. Liveness passes, readiness passes, the pods are Running and Ready. Yet for the first seconds or minutes after every pod start, clients get SERVFAIL for &lt;code>cluster.local&lt;/code> names. New deployments fail to resolve their dependencies, retries pile up, and then the problem vanishes on its own.&lt;/p>
&lt;p>This is the classic symptom of pointing the Kubernetes readiness probe at &lt;code>/health&lt;/code> instead of &lt;code>/ready&lt;/code>. The two endpoints answer different questions, and confusing them makes your readiness gate meaningless.&lt;/p></description></item><item><title>CoreDNS 403 from the Kubernetes API: RBAC misconfiguration breaking cluster DNS</title><link>https://www.netdata.cloud/guides/coredns/coredns-rbac-403-api-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-rbac-403-api-errors/</guid><description>&lt;h1 id="coredns-403-from-the-kubernetes-api-rbac-misconfiguration-breaking-cluster-dns">CoreDNS 403 from the Kubernetes API: RBAC misconfiguration breaking cluster DNS&lt;/h1>
&lt;p>The symptom usually arrives in one of two forms. Either CoreDNS pods never become ready after an install or upgrade, stuck at 0/1 while &lt;code>kubectl&lt;/code> shows no crash and no obvious error, or cluster DNS suddenly returns SERVFAIL for &lt;code>cluster.local&lt;/code> names while external resolution keeps working. In both cases, the smoking gun is the same: &lt;code>coredns_kubernetes_rest_client_requests_total{code=&amp;quot;403&amp;quot;}&lt;/code> is incrementing.&lt;/p>
&lt;p>A 403 from the Kubernetes API is not a transient error. It means the API server received the request, authenticated the caller, and refused it because the CoreDNS ServiceAccount lacks the RBAC permissions the kubernetes plugin needs to list and watch Services, Endpoints, and EndpointSlices. Without those watches, the plugin cannot build or maintain its in-memory record set, so it cannot answer queries for the cluster zone.&lt;/p></description></item><item><title>CoreDNS 5-second DNS timeout: the Kubernetes glibc A+AAAA conntrack race</title><link>https://www.netdata.cloud/guides/coredns/coredns-5-second-dns-timeout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-5-second-dns-timeout/</guid><description>&lt;h1 id="coredns-5-second-dns-timeout-the-kubernetes-glibc-aaaaa-conntrack-race">CoreDNS 5-second DNS timeout: the Kubernetes glibc A+AAAA conntrack race&lt;/h1>
&lt;p>Your application latency histogram has a cluster of samples at almost exactly 5 seconds. Not 4.8, not 5.3: a tight spike right at the glibc default resolver timeout. Tracing the slow calls shows they stall on name resolution. CoreDNS dashboards are green: low latency, no SERVFAIL, healthy cache hit ratio, both replicas fine.&lt;/p>
&lt;p>That combination, application-side DNS stalls at exactly the resolver timeout with clean CoreDNS metrics, is the signature of the glibc A+AAAA conntrack race. The loss is in the kernel network path between the application pod and CoreDNS, not inside CoreDNS. No amount of Corefile tuning will touch it.&lt;/p></description></item><item><title>CoreDNS all upstreams down: the forwarding black hole and healthcheck_broken</title><link>https://www.netdata.cloud/guides/coredns/coredns-all-upstreams-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-all-upstreams-down/</guid><description>&lt;h1 id="coredns-all-upstreams-down-the-forwarding-black-hole-and-healthcheck_broken">CoreDNS all upstreams down: the forwarding black hole and healthcheck_broken&lt;/h1>
&lt;p>Every forwarded query through CoreDNS is suddenly returning SERVFAIL, request latency has dropped instead of rising, and &lt;code>coredns_forward_healthcheck_broken_total&lt;/code> is climbing. This is the upstream black hole: every upstream resolver configured in the &lt;code>forward&lt;/code> plugin is failing health checks at the same time, and CoreDNS has nowhere to send external queries.&lt;/p>
&lt;p>The counter-intuitive part is the latency profile. When upstreams die hard, failures are fast. Queries do not sit in timeouts the way they do with a slow upstream. If you only alert on high latency, this failure mode sails under your dashboards until clients start reporting resolution errors.&lt;/p></description></item><item><title>CoreDNS ANY query flood: DNS amplification and reflection abuse</title><link>https://www.netdata.cloud/guides/coredns/coredns-any-query-amplification/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-any-query-amplification/</guid><description>&lt;h1 id="coredns-any-query-flood-dns-amplification-and-reflection-abuse">CoreDNS ANY query flood: DNS amplification and reflection abuse&lt;/h1>
&lt;p>You are looking at &lt;code>coredns_dns_requests_total&lt;/code> and the &lt;code>type=&amp;quot;ANY&amp;quot;&lt;/code> series has jumped from a flat zero to a sustained rate, or ANY queries have crept past a few percent of total traffic. That pattern is the classic fingerprint of DNS amplification and reflection abuse: an attacker sends small ANY queries with a spoofed source address, and your CoreDNS replies with large responses delivered to the victim&amp;rsquo;s IP. Your server is the reflector; someone else absorbs the blast.&lt;/p></description></item><item><title>CoreDNS AXFR zone transfer attempts: reconnaissance of internal service topology</title><link>https://www.netdata.cloud/guides/coredns/coredns-axfr-zone-transfer-attempts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-axfr-zone-transfer-attempts/</guid><description>&lt;h1 id="coredns-axfr-zone-transfer-attempts-reconnaissance-of-internal-service-topology">CoreDNS AXFR zone transfer attempts: reconnaissance of internal service topology&lt;/h1>
&lt;p>An AXFR query asks a DNS server for the entire contents of a zone: every record, not one. Against a Kubernetes cluster running CoreDNS, a successful zone transfer hands the requester a complete map of internal service topology: every Service name, every namespace, every headless Service endpoint, every pod hostname you publish. That is exactly the map an attacker wants before lateral movement.&lt;/p></description></item><item><title>CoreDNS cache collapse: the cold-cache thundering herd after a rollout</title><link>https://www.netdata.cloud/guides/coredns/coredns-cache-collapse-thundering-herd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-cache-collapse-thundering-herd/</guid><description>&lt;h1 id="coredns-cache-collapse-the-cold-cache-thundering-herd-after-a-rollout">CoreDNS cache collapse: the cold-cache thundering herd after a rollout&lt;/h1>
&lt;p>You just rolled out a CoreDNS config change or image bump. Within seconds, DNS latency across the cluster jumps, upstream DNS traffic spikes to several times baseline, and SERVFAILs appear in application logs. One to five minutes later, it all goes away on its own. The dashboard is green again.&lt;/p>
&lt;p>That is the cache collapse pattern: every CoreDNS pod restarted at roughly the same time, every in-memory cache emptied at once, and every client in the cluster re-queried the same names simultaneously. The flood hit your upstreams harder than they could absorb. In the worst version, upstreams return SERVFAIL, CoreDNS caches those SERVFAILs for 5 seconds each, and a brief overload amplifies into a visible outage.&lt;/p></description></item><item><title>CoreDNS cache evictions: the cache is too small for the working set</title><link>https://www.netdata.cloud/guides/coredns/coredns-cache-evictions-undersized/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-cache-evictions-undersized/</guid><description>&lt;h1 id="coredns-cache-evictions-the-cache-is-too-small-for-the-working-set">CoreDNS cache evictions: the cache is too small for the working set&lt;/h1>
&lt;p>&lt;code>coredns_cache_evictions_total&lt;/code> climbing steadily, or upstream query load and DNS latency creeping up without any change in traffic volume, both point at the same thing: the cache plugin is evicting entries before their TTL expires because it has run out of room, and queries that should have been sub-millisecond cache hits are being forwarded to upstream resolvers instead.&lt;/p>
&lt;p>This is progressive degradation, not a cliff edge. Latency rises, upstream load rises, and the hit ratio slides, slowly enough that it often goes unnoticed until someone asks why DNS got slower this quarter.&lt;/p></description></item><item><title>CoreDNS cache hit ratio dropping: latency and upstream load climbing together</title><link>https://www.netdata.cloud/guides/coredns/coredns-cache-hit-ratio-dropping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-cache-hit-ratio-dropping/</guid><description>&lt;h1 id="coredns-cache-hit-ratio-dropping-latency-and-upstream-load-climbing-together">CoreDNS cache hit ratio dropping: latency and upstream load climbing together&lt;/h1>
&lt;p>The cache hit ratio is &lt;code>coredns_cache_hits_total / coredns_cache_requests_total&lt;/code>. When it falls, every query that used to be answered from memory in under a millisecond now goes to the kubernetes plugin or out to an upstream resolver. Two things happen at once: &lt;code>coredns_dns_request_duration_seconds&lt;/code> climbs, and forwarded query volume climbs with it. That second effect is the dangerous one, because the upstream flood can degrade the upstreams themselves and turn a cache problem into a resolution outage.&lt;/p></description></item><item><title>CoreDNS conntrack table full: silent UDP packet drops with a node-wide blast radius</title><link>https://www.netdata.cloud/guides/coredns/coredns-conntrack-table-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-conntrack-table-full/</guid><description>&lt;h1 id="coredns-conntrack-table-full-silent-udp-packet-drops-with-a-node-wide-blast-radius">CoreDNS conntrack table full: silent UDP packet drops with a node-wide blast radius&lt;/h1>
&lt;p>Applications across a node start timing out on DNS lookups. Then on database connections. Then on API calls. You check CoreDNS and the dashboards are green: low latency, no SERVFAILs, healthy pods. But the timeouts keep coming, and they are not limited to DNS.&lt;/p>
&lt;p>This is conntrack table exhaustion. Every UDP DNS query that traverses Kubernetes iptables DNAT creates a connection tracking entry on the node with a default timeout of 30 seconds. At high DNS QPS, the table fills. When it does, the kernel drops new packets silently: no ICMP error, no reset, no log line the application will ever see. The kernel log says &lt;code>nf_conntrack: table full, dropping packet&lt;/code> and nothing else anywhere warns you.&lt;/p></description></item><item><title>CoreDNS Corefile parse error: startup failures and invalid plugin config</title><link>https://www.netdata.cloud/guides/coredns/coredns-corefile-parse-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-corefile-parse-error/</guid><description>&lt;h1 id="coredns-corefile-parse-error-startup-failures-and-invalid-plugin-config">CoreDNS Corefile parse error: startup failures and invalid plugin config&lt;/h1>
&lt;p>You changed the Corefile, and now one of two things has happened. Either the CoreDNS pods are in CrashLoopBackOff and cluster DNS is down, or the pods look fine but your change never took effect and the log is full of parse errors. Both are the same root event: the Corefile parser rejected your configuration. The difference is whether the parser ran at process start (fatal) or during a live reload (rejected, old config kept).&lt;/p></description></item><item><title>CoreDNS CPU throttling: CFS limits making a green dashboard lie about latency</title><link>https://www.netdata.cloud/guides/coredns/coredns-cpu-throttling-slow-responses/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-cpu-throttling-slow-responses/</guid><description>&lt;h1 id="coredns-cpu-throttling-cfs-limits-making-a-green-dashboard-lie-about-latency">CoreDNS CPU throttling: CFS limits making a green dashboard lie about latency&lt;/h1>
&lt;p>Applications are reporting intermittent DNS timeouts and multi-hundred-millisecond resolution spikes. You open the CoreDNS dashboard and everything looks fine: &lt;code>coredns_dns_request_duration_seconds&lt;/code> shows sub-millisecond cache hits, SERVFAIL is zero, QPS is normal, pods are healthy. Both things are true at the same time, and that is exactly the problem.&lt;/p>
&lt;p>When a CoreDNS pod has a tight CPU limit, the Linux CFS scheduler throttles the container: the process is runnable but the kernel refuses to schedule it until the next quota period. That stall is real latency for the client, but it never appears in any CoreDNS metric, because CoreDNS only measures time spent actively processing a query. Time spent waiting for CPU does not exist as far as &lt;code>coredns_dns_request_duration_seconds&lt;/code> is concerned.&lt;/p></description></item><item><title>CoreDNS CrashLoopBackOff: triaging loop, OOM, config, and port-bind causes</title><link>https://www.netdata.cloud/guides/coredns/coredns-crashloopbackoff-triage/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-crashloopbackoff-triage/</guid><description>&lt;h1 id="coredns-crashloopbackoff-triaging-loop-oom-config-and-port-bind-causes">CoreDNS CrashLoopBackOff: triaging loop, OOM, config, and port-bind causes&lt;/h1>
&lt;p>A CoreDNS pod in CrashLoopBackOff means the process starts, dies, and gets restarted by Kubernetes in a tightening backoff loop. While this is happening, at least one DNS replica is contributing nothing. If all replicas are crash-looping, the &lt;code>kube-dns&lt;/code> Service has no backends and cluster-wide name resolution is down: queries hang at the client until a pod recovers.&lt;/p>
&lt;p>CrashLoopBackOff is not one failure. It is a symptom with a small set of common root causes, and they need different fixes. The four you will see most often: the loop plugin detecting a forwarding loop, the OOM killer terminating the container, a Corefile that fails to parse, and a port-bind failure on startup. RBAC and Kubernetes API access failures show up in the same pod state too, usually as fatal list errors or persistent un-readiness.&lt;/p></description></item><item><title>CoreDNS DNS programming latency: how long a Service takes to become resolvable</title><link>https://www.netdata.cloud/guides/coredns/coredns-dns-programming-duration/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-dns-programming-duration/</guid><description>&lt;h1 id="coredns-dns-programming-latency-how-long-a-service-takes-to-become-resolvable">CoreDNS DNS programming latency: how long a Service takes to become resolvable&lt;/h1>
&lt;p>You create a Service, the API server accepts it, &lt;code>kubectl get svc&lt;/code> shows it, and yet for some period of time nothing in the cluster can resolve its name. That gap, from &amp;ldquo;the object exists in the API&amp;rdquo; to &amp;ldquo;CoreDNS answers for it&amp;rdquo;, is DNS programming latency, and CoreDNS exposes it as a first-class metric: &lt;code>coredns_kubernetes_dns_programming_duration_seconds&lt;/code>.&lt;/p>
&lt;p>This is the Kubernetes DNS Programming SLI. It exists because service discovery delay is a real failure mode: a rollout completes, endpoints are ready, but clients still get NXDOMAIN or stale answers for seconds or tens of seconds. From the application&amp;rsquo;s perspective the new pods are unreachable by name. From CoreDNS&amp;rsquo;s request metrics everything looks green.&lt;/p></description></item><item><title>CoreDNS forward max_concurrent rejects: the forward plugin is overwhelmed</title><link>https://www.netdata.cloud/guides/coredns/coredns-forward-max-concurrent-rejects/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-forward-max-concurrent-rejects/</guid><description>&lt;h1 id="coredns-forward-max_concurrent-rejects-the-forward-plugin-is-overwhelmed">CoreDNS forward max_concurrent rejects: the forward plugin is overwhelmed&lt;/h1>
&lt;p>&lt;code>coredns_forward_max_concurrent_rejects_total&lt;/code> is incrementing and clients are getting REFUSED responses for forwarded queries. Something is piling up between CoreDNS and your upstreams, and the rejects are the valve letting pressure escape.&lt;/p>
&lt;p>This counter only increments when the number of in-flight forwarded queries hits the &lt;code>max_concurrent&lt;/code> cap you configured. Every rejected query gets a REFUSED response, not SERVFAIL. That distinction matters for triage: REFUSED from this path is a capacity signal, not a resolution failure, and unlike NXDOMAIN it is not subject to negative caching, so it will not linger in client caches after the pressure clears.&lt;/p></description></item><item><title>CoreDNS GC pauses adding tail latency: go_gc_duration_seconds and heap pressure</title><link>https://www.netdata.cloud/guides/coredns/coredns-gc-pause-tail-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-gc-pause-tail-latency/</guid><description>&lt;h1 id="coredns-gc-pauses-adding-tail-latency-go_gc_duration_seconds-and-heap-pressure">CoreDNS GC pauses adding tail latency: go_gc_duration_seconds and heap pressure&lt;/h1>
&lt;p>Your CoreDNS dashboards look mostly fine: P50 latency is sub-millisecond, SERVFAIL is zero, cache hit ratio is healthy. But P99 is spiking to tens or hundreds of milliseconds at irregular intervals, and a small fraction of clients see DNS timeouts they cannot explain. The process never crashes. Upstream latency is clean. CPU looks acceptable on average.&lt;/p>
&lt;p>This is the classic signature of Go garbage collection stop-the-world pauses landing on DNS query latency. CoreDNS is a Go process, and every in-flight query is a goroutine. When the GC stops the world, every one of those goroutines waits. DNS is supposed to be sub-millisecond for cache hits, so a pause that would be invisible in a batch service is user-visible here. The working threshold: GC pauses over 10ms are impactful for DNS and show up as P99 spikes.&lt;/p></description></item><item><title>CoreDNS goroutine count climbing: blocked upstream calls and leaks</title><link>https://www.netdata.cloud/guides/coredns/coredns-goroutine-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-goroutine-leak/</guid><description>&lt;h1 id="coredns-goroutine-count-climbing-blocked-upstream-calls-and-leaks">CoreDNS goroutine count climbing: blocked upstream calls and leaks&lt;/h1>
&lt;p>Your &lt;code>go_goroutines&lt;/code> graph for CoreDNS is trending up. Maybe it spiked during an incident and never came back down. Maybe it has been climbing slowly for days. Either way, the question is the same: are queries piling up behind a slow upstream, or is something leaking goroutines that will never exit?&lt;/p>
&lt;p>The distinction matters because the fixes are completely different. A blocked-upstream spike resolves when the upstream recovers. A leak grows until the OOM killer ends the debate for you.&lt;/p></description></item><item><title>CoreDNS high request latency: reading P99 by zone to find the cause</title><link>https://www.netdata.cloud/guides/coredns/coredns-high-request-latency-p99/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-high-request-latency-p99/</guid><description>&lt;h1 id="coredns-high-request-latency-reading-p99-by-zone-to-find-the-cause">CoreDNS high request latency: reading P99 by zone to find the cause&lt;/h1>
&lt;p>P99 on &lt;code>coredns_dns_request_duration_seconds&lt;/code> is at 400ms when it used to sit at 15ms. Do not start with the aggregate number: aggregate latency in CoreDNS blends at least two very different workloads. Cache hits should complete in single-digit milliseconds. Forwarded and Kubernetes-backed queries depend on systems outside CoreDNS itself.&lt;/p>
&lt;p>The &lt;code>zone&lt;/code> label on the duration histogram is the fork in the road. High latency in &lt;code>cluster.local&lt;/code> points at the Kubernetes path. High latency in forwarded zones points at upstream resolvers. High latency everywhere, at every percentile, points at the CoreDNS process itself: CPU saturation or GC pressure. Until you split the histogram by zone, you are guessing.&lt;/p></description></item><item><title>CoreDNS Kubernetes API disconnect: stale records and new services going invisible</title><link>https://www.netdata.cloud/guides/coredns/coredns-kubernetes-api-disconnect/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-kubernetes-api-disconnect/</guid><description>&lt;h1 id="coredns-kubernetes-api-disconnect-stale-records-and-new-services-going-invisible">CoreDNS Kubernetes API disconnect: stale records and new services going invisible&lt;/h1>
&lt;p>The symptom that usually surfaces first is not a DNS error. It is a developer saying &amp;ldquo;I deployed the service twenty minutes ago and nothing can reach it,&amp;rdquo; while every existing workload looks fine. &lt;code>nslookup kubernetes.default&lt;/code> works. External domains resolve. CoreDNS dashboards are green. But anything created or changed recently in the cluster does not exist as far as DNS is concerned.&lt;/p></description></item><item><title>CoreDNS Loop detected and CrashLoopBackOff: the forwarding loop that kills the pod</title><link>https://www.netdata.cloud/guides/coredns/coredns-loop-detected-crashloopbackoff/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-loop-detected-crashloopbackoff/</guid><description>&lt;h1 id="coredns-loop-detected-and-crashloopbackoff-the-forwarding-loop-that-kills-the-pod">CoreDNS Loop detected and CrashLoopBackOff: the forwarding loop that kills the pod&lt;/h1>
&lt;p>Your CoreDNS pods are in CrashLoopBackOff. &lt;code>kubectl logs&lt;/code> shows a line like &lt;code>[FATAL] plugin/loop: Loop (127.0.0.1:53 -&amp;gt; :53) detected for zone &amp;quot;.&amp;quot;&lt;/code>, and the process exits a few seconds after every start. There are no CoreDNS metrics to look at, because the process dies before the metrics endpoint is ever scraped.&lt;/p>
&lt;p>The crash is intentional. The &lt;code>loop&lt;/code> plugin detected that the Corefile&amp;rsquo;s &lt;code>forward&lt;/code> target routes queries back to CoreDNS itself, and it killed the process on purpose to prevent an infinite query amplification storm. Kubernetes restarts the pod, the loop is detected again, and the cycle repeats.&lt;/p></description></item><item><title>CoreDNS memory climbing: heap growth, post-GC minima, and leak detection</title><link>https://www.netdata.cloud/guides/coredns/coredns-memory-leak-heap-growth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-memory-leak-heap-growth/</guid><description>&lt;h1 id="coredns-memory-climbing-heap-growth-post-gc-minima-and-leak-detection">CoreDNS memory climbing: heap growth, post-GC minima, and leak detection&lt;/h1>
&lt;p>CoreDNS memory has been creeping up for hours. The Grafana panel shows a sawtooth that never quite comes back down, and you need to answer one question before the pod gets OOMKilled: is this Go being Go, or is something actually leaking?&lt;/p>
&lt;p>The trap is that &lt;code>go_memstats_heap_inuse_bytes&lt;/code> is spiky by design. The heap grows, the garbage collector reclaims, the heap grows again. If you react to instantaneous peaks you will chase ghosts all night. The signal that separates normal GC behavior from a leak is the trend of the post-GC minima: the floor each sawtooth returns to. If that floor keeps rising over an hour or more and never returns to baseline, you have a genuine leak heading for the container memory limit.&lt;/p></description></item><item><title>CoreDNS memory limit and the re-list spike: sizing for the restart peak</title><link>https://www.netdata.cloud/guides/coredns/coredns-memory-limit-relist-spike/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-memory-limit-relist-spike/</guid><description>&lt;h1 id="coredns-memory-limit-and-the-re-list-spike-sizing-for-the-restart-peak">CoreDNS memory limit and the re-list spike: sizing for the restart peak&lt;/h1>
&lt;p>A CoreDNS pod that runs fine for weeks gets OOMKilled thirty seconds after a restart. Kubernetes starts it again. It re-syncs, allocates heavily, crosses the limit, and gets killed again. Restart count climbs into the hundreds, and cluster DNS capacity drops by half or more while the crash loop runs. &lt;code>kubectl describe pod&lt;/code> shows &lt;code>Last State: Terminated, Reason: OOMKilled&lt;/code>, and nothing in steady-state monitoring predicted it.&lt;/p></description></item><item><title>CoreDNS Monitoring</title><link>https://www.netdata.cloud/monitoring-101/coredns-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/coredns-monitoring/</guid><description>&lt;h2 id="coredns-monitoring">CoreDNS Monitoring&lt;/h2>
&lt;h3 id="what-is-coredns">What Is CoreDNS?&lt;/h3>
&lt;p>CoreDNS is a flexible and extensible DNS server that can serve as a cluster DNS inside Kubernetes or provide other DNS services. Built with a focus on flexibility, plugin support, and efficiency, it&amp;rsquo;s designed to facilitate the addition and combination of functionalities. You can read more about CoreDNS on their &lt;a href="https://coredns.io/">official website&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-coredns-with-netdata">Monitoring CoreDNS With Netdata&lt;/h3>
&lt;p>Netdata is an ideal tool for monitoring CoreDNS, offering real-time visibility and insightful metrics. With Netdata, you can track DNS queries, responses, and errors to ensure the robustness and performance of your DNS server. Explore our &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">Live Demo&lt;/a> to experience Netdata’s capabilities firsthand.&lt;/p></description></item><item><title>CoreDNS monitoring checklist: the signals every production resolver needs</title><link>https://www.netdata.cloud/guides/coredns/coredns-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-monitoring-checklist/</guid><description>&lt;h1 id="coredns-monitoring-checklist-the-signals-every-production-resolver-needs">CoreDNS monitoring checklist: the signals every production resolver needs&lt;/h1>
&lt;p>CoreDNS has one failure mode that defeats most monitoring setups: a pod that passes every health probe while returning SERVFAIL for every query. The &lt;code>/health&lt;/code> endpoint on port 8080 checks process liveness. It does not test DNS resolution. The &lt;code>/ready&lt;/code> endpoint on port 8181 is plugin-aware, but it does not test resolution either. Teams that alert on pod status and probe results are monitoring the wrong thing.&lt;/p></description></item><item><title>CoreDNS monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/coredns/coredns-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-monitoring-maturity-model/</guid><description>&lt;h1 id="coredns-monitoring-maturity-model-from-survival-to-expert">CoreDNS monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most CoreDNS monitoring setups are either a single &amp;ldquo;is the pod up&amp;rdquo; check or a wall of dashboards nobody reads during an incident. Neither works. The right question is not &amp;ldquo;how many metrics do we collect&amp;rdquo; but &amp;ldquo;which failure modes can we actually detect right now.&amp;rdquo; That is what a maturity model answers: it orders signals by the failures they catch, so you can see exactly which class of outage would currently reach your users before it reaches your alerts.&lt;/p></description></item><item><title>CoreDNS ndots:5 query amplification: one lookup becoming five</title><link>https://www.netdata.cloud/guides/coredns/coredns-ndots-search-domain-amplification/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-ndots-search-domain-amplification/</guid><description>&lt;h1 id="coredns-ndots5-query-amplification-one-lookup-becoming-five">CoreDNS ndots:5 query amplification: one lookup becoming five&lt;/h1>
&lt;p>CoreDNS dashboards show 50,000 QPS, the denial cache is churning, and NXDOMAIN responses are a third of all traffic, but nothing is broken. What you are looking at is probably not real demand. The default &lt;code>ndots:5&lt;/code> in every pod&amp;rsquo;s &lt;code>/etc/resolv.conf&lt;/code> expands each external lookup through the Kubernetes search domain list, so one logical lookup becomes four to six wire queries before the real name is ever tried.&lt;/p></description></item><item><title>CoreDNS NOERROR with zero answers: the resolution failure that reports success</title><link>https://www.netdata.cloud/guides/coredns/coredns-noerror-empty-answers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-noerror-empty-answers/</guid><description>&lt;h1 id="coredns-noerror-with-zero-answers-the-resolution-failure-that-reports-success">CoreDNS NOERROR with zero answers: the resolution failure that reports success&lt;/h1>
&lt;p>Your RCODE dashboards are green. NOERROR rate is healthy, SERVFAIL is near zero, health probes pass. Yet an application team reports their service cannot resolve a dependency. &lt;code>dig&lt;/code> returns instantly with &lt;code>status: NOERROR&lt;/code> and &lt;code>ANSWER SECTION: 0&lt;/code>. No address, no error, no timeout. The resolver library treats this as &amp;ldquo;the name exists but has no records&amp;rdquo; and the application fails with a connection error that looks nothing like a DNS problem.&lt;/p></description></item><item><title>CoreDNS not resolving external domains: the missing catch-all forward zone</title><link>https://www.netdata.cloud/guides/coredns/coredns-external-domains-not-resolving/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-external-domains-not-resolving/</guid><description>&lt;h1 id="coredns-not-resolving-external-domains-the-missing-catch-all-forward-zone">CoreDNS not resolving external domains: the missing catch-all forward zone&lt;/h1>
&lt;p>The symptom is specific: names inside the cluster resolve, but anything outside the cluster fails. &lt;code>kubernetes.default.svc.cluster.local&lt;/code> works. &lt;code>example.com&lt;/code> does not. Applications report DNS timeouts, &lt;code>getaddrinfo&lt;/code> failures, or intermittent external dependency errors while service discovery inside Kubernetes looks fine.&lt;/p>
&lt;p>That split is the clue. CoreDNS is not one global resolver. It matches each query to the most specific server block in the Corefile, then runs that block&amp;rsquo;s plugin chain. If &lt;code>cluster.local&lt;/code> is handled by the &lt;code>kubernetes&lt;/code> plugin but there is no catch-all block for &lt;code>.&lt;/code>, external names match nothing useful and are refused instead of forwarded.&lt;/p></description></item><item><title>CoreDNS NXDOMAIN flood from one source: DGA malware and domain enumeration</title><link>https://www.netdata.cloud/guides/coredns/coredns-nxdomain-flood-dga/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-nxdomain-flood-dga/</guid><description>&lt;h1 id="coredns-nxdomain-flood-from-one-source-dga-malware-and-domain-enumeration">CoreDNS NXDOMAIN flood from one source: DGA malware and domain enumeration&lt;/h1>
&lt;p>Your CoreDNS dashboards are green: SERVFAIL is zero, latency is fine, cache hit ratio looks normal. But the NXDOMAIN rate has jumped, and when you pull query logs, one source IP is responsible for thousands of NXDOMAIN responses per hour, querying names that do not exist.&lt;/p>
&lt;p>This is a security signal, not a CoreDNS health signal. A single source generating a sustained flood of NXDOMAIN responses is doing one of three things: cycling randomly generated domains looking for a command-and-control server (DGA malware in a compromised pod), probing which service names exist in your cluster (reconnaissance), or misbehaving in a way that merely looks like an attack (a config typo hammering a name that was never created).&lt;/p></description></item><item><title>CoreDNS NXDOMAIN vs SERVFAIL: why alerting on the wrong one buries real incidents</title><link>https://www.netdata.cloud/guides/coredns/coredns-nxdomain-vs-servfail/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-nxdomain-vs-servfail/</guid><description>&lt;h1 id="coredns-nxdomain-vs-servfail-why-alerting-on-the-wrong-one-buries-real-incidents">CoreDNS NXDOMAIN vs SERVFAIL: why alerting on the wrong one buries real incidents&lt;/h1>
&lt;p>Your CoreDNS alert fired again. The &amp;ldquo;DNS error rate&amp;rdquo; panel shows a wall of red, so you open it, see a pile of NXDOMAIN responses, and silence the alert. Meanwhile, a genuine upstream failure is producing SERVFAIL responses somewhere in that same wall of red, and nobody will notice until an application team opens a ticket.&lt;/p>
&lt;p>This is the most common CoreDNS alerting mistake: treating all non-NOERROR responses as one bucket. NXDOMAIN and SERVFAIL look similar in a naive error counter, but they mean opposite things. NXDOMAIN is the server correctly reporting that a name does not exist. SERVFAIL is the server admitting it failed to answer at all. Alerting on the sum of both guarantees that the noise (NXDOMAIN, constant and expected in Kubernetes) drowns the signal (SERVFAIL, rare and always actionable).&lt;/p></description></item><item><title>CoreDNS OOMKilled: the memory cliff and the restart crash loop</title><link>https://www.netdata.cloud/guides/coredns/coredns-oomkilled/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-oomkilled/</guid><description>&lt;h1 id="coredns-oomkilled-the-memory-cliff-and-the-restart-crash-loop">CoreDNS OOMKilled: the memory cliff and the restart crash loop&lt;/h1>
&lt;p>You found the CoreDNS pods in &lt;code>kube-system&lt;/code> with &lt;code>OOMKilled&lt;/code> as the last terminated state. Maybe one pod restarted once and recovered. Maybe both replicas are flapping and cluster DNS is degraded or down. Either way, the pod status operators paste into a search bar is the same: &lt;code>Last State: Terminated, Reason: OOMKilled&lt;/code>.&lt;/p>
&lt;p>Memory exhaustion in CoreDNS is a cliff edge, not a slowdown. There is no degradation curve where latency rises and errors climb while you watch. RSS approaches the container memory limit, the kernel OOM killer terminates the process instantly, and every client served by that pod loses DNS until a replacement is ready. The warning phase exists only in your metrics, and only if you are watching the right ones.&lt;/p></description></item><item><title>CoreDNS panics: recovered handler crashes and coredns_panics_total</title><link>https://www.netdata.cloud/guides/coredns/coredns-panics-recovered/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-panics-recovered/</guid><description>&lt;h1 id="coredns-panics-recovered-handler-crashes-and-coredns_panics_total">CoreDNS panics: recovered handler crashes and coredns_panics_total&lt;/h1>
&lt;p>You found &lt;code>coredns_panics_total&lt;/code> above zero, or you saw &lt;code>Recovered from panic&lt;/code> in the CoreDNS logs. The pods are not restarting, the health endpoint returns 200, and most DNS queries still resolve. That combination is what makes this failure mode easy to ignore and dangerous to dismiss.&lt;/p>
&lt;p>A recovered panic means a query handler crashed inside the plugin chain and CoreDNS&amp;rsquo;s recovery wrapper caught it, keeping the process alive. The server survived. The query that triggered the panic did not: the client that sent it got no answer and waited out its own timeout. Every increment of the counter is a real query that failed, plus evidence of a genuine bug in CoreDNS or one of its plugins.&lt;/p></description></item><item><title>CoreDNS per-replica divergence: why averaged metrics hide a failing pod</title><link>https://www.netdata.cloud/guides/coredns/coredns-per-replica-divergence/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-per-replica-divergence/</guid><description>&lt;h1 id="coredns-per-replica-divergence-why-averaged-metrics-hide-a-failing-pod">CoreDNS per-replica divergence: why averaged metrics hide a failing pod&lt;/h1>
&lt;p>Your CoreDNS dashboard looks slightly off. Average latency drifted up, SERVFAIL ratio is elevated but under the threshold, and nobody is being paged. Meanwhile, half the cluster&amp;rsquo;s DNS queries are slow or failing, because one of your two replicas is degraded and kube-proxy is still sending it traffic.&lt;/p>
&lt;p>This is the per-replica divergence failure mode. CoreDNS in Kubernetes is almost never a singleton: the default deployment runs two replicas behind the &lt;code>kube-dns&lt;/code> ClusterIP Service, and kube-proxy load-balances across them. Any aggregation that averages or sums across pods blends a sick replica&amp;rsquo;s numbers with healthy ones until the result looks like mild degradation instead of a partial outage.&lt;/p></description></item><item><title>CoreDNS per-upstream health check failures: degraded redundancy before total loss</title><link>https://www.netdata.cloud/guides/coredns/coredns-per-upstream-health-check-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-per-upstream-health-check-failures/</guid><description>&lt;h1 id="coredns-per-upstream-health-check-failures-degraded-redundancy-before-total-loss">CoreDNS per-upstream health check failures: degraded redundancy before total loss&lt;/h1>
&lt;p>One of your upstream DNS resolvers is failing CoreDNS&amp;rsquo;s health checks. Queries are still being answered because the remaining upstreams are carrying the forwarding load, so nothing is paging yet. This is exactly the window you want: degraded redundancy, not an outage. The next failure in this sequence is &lt;code>coredns_forward_healthcheck_broken_total&lt;/code> incrementing, at which point every upstream is marked unhealthy and you are in a real incident.&lt;/p></description></item><item><title>CoreDNS pod alive but not ready: the /ready endpoint and API sync</title><link>https://www.netdata.cloud/guides/coredns/coredns-readiness-probe-not-ready/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-readiness-probe-not-ready/</guid><description>&lt;h1 id="coredns-pod-alive-but-not-ready-the-ready-endpoint-and-api-sync">CoreDNS pod alive but not ready: the /ready endpoint and API sync&lt;/h1>
&lt;p>The pod is &lt;code>Running&lt;/code>. The process is up, the logs show CoreDNS started, maybe it is even answering some queries. But &lt;code>kubectl get pods -n kube-system&lt;/code> shows &lt;code>0/1 READY&lt;/code>, the pod is excluded from the &lt;code>kube-dns&lt;/code> Service endpoints, and your cluster is running on one fewer DNS replica than you think. If this is your only replica, or the other one is on the same node, you are one hiccup away from a cluster-wide DNS outage.&lt;/p></description></item><item><title>CoreDNS query rate dropped to zero while the process looks healthy</title><link>https://www.netdata.cloud/guides/coredns/coredns-query-rate-dropped-to-zero/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-query-rate-dropped-to-zero/</guid><description>&lt;h1 id="coredns-query-rate-dropped-to-zero-while-the-process-looks-healthy">CoreDNS query rate dropped to zero while the process looks healthy&lt;/h1>
&lt;p>The dashboard shows &lt;code>coredns_dns_requests_total&lt;/code> flatlined at zero. You check the pod: &lt;code>Running&lt;/code>, zero restarts. &lt;code>/health&lt;/code> on port 8080 returns 200. The process is fine. So either nobody in the cluster is resolving names anymore, or the traffic is dying somewhere between your clients and the CoreDNS process, and CoreDNS has no idea.&lt;/p>
&lt;p>The second possibility is the trap. CoreDNS only counts queries it actually receives. A UDP packet dropped by the kernel before it reaches the CoreDNS socket is never counted, never logged, and never reflected in any CoreDNS metric. A full conntrack table, an overflowing UDP receive buffer, a network policy isolating the pod, or a readiness failure that pulled the pod out of the Service endpoints all look identical in CoreDNS metrics: nothing.&lt;/p></description></item><item><title>CoreDNS reload failed: config drift when the new Corefile never took effect</title><link>https://www.netdata.cloud/guides/coredns/coredns-reload-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-reload-failed/</guid><description>&lt;h1 id="coredns-reload-failed-config-drift-when-the-new-corefile-never-took-effect">CoreDNS reload failed: config drift when the new Corefile never took effect&lt;/h1>
&lt;p>You pushed a Corefile change, nothing broke, and everyone moved on. Weeks later you discover the cluster is still running the old configuration. The metric that would have told you is &lt;code>coredns_reload_failed_total&lt;/code>, and it was nonzero the whole time.&lt;/p>
&lt;p>A failed CoreDNS reload is not an outage. The &lt;code>reload&lt;/code> plugin polls the Corefile every 30 seconds (with jitter) and triggers a graceful reload when the SHA512 checksum of the file changes. If the new configuration has a syntax error, an invalid plugin directive, or hits a port conflict, CoreDNS logs the error, increments &lt;code>coredns_reload_failed_total&lt;/code>, and keeps serving DNS with the old config. Queries keep flowing. The only thing that changed is that reality and your intent have quietly diverged.&lt;/p></description></item><item><title>CoreDNS response size distribution: large answers, TCP pressure, and amplification</title><link>https://www.netdata.cloud/guides/coredns/coredns-large-response-size/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-large-response-size/</guid><description>&lt;h1 id="coredns-response-size-distribution-large-answers-tcp-pressure-and-amplification">CoreDNS response size distribution: large answers, TCP pressure, and amplification&lt;/h1>
&lt;p>Most CoreDNS dashboards track QPS, latency, and SERVFAIL. Response size rarely makes the cut until something odd happens: TCP/53 traffic climbs for no obvious reason, a workload starts resolving incomplete answers, or a security review asks whether your resolvers could be used as DDoS reflectors. All three questions are answered by the same signal: the distribution of DNS response sizes CoreDNS is sending.&lt;/p></description></item><item><title>CoreDNS response truncation: the TC bit, EDNS bufsize, and TCP fallback</title><link>https://www.netdata.cloud/guides/coredns/coredns-response-truncation-tcp-fallback/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-response-truncation-tcp-fallback/</guid><description>&lt;h1 id="coredns-response-truncation-the-tc-bit-edns-bufsize-and-tcp-fallback">CoreDNS response truncation: the TC bit, EDNS bufsize, and TCP fallback&lt;/h1>
&lt;p>Most DNS traffic fits in a single UDP datagram. When it does not, the protocol has an escape hatch: the server sets the TC (truncated) bit in the response, and the client retries the same query over TCP. CoreDNS implements this correctly in the common case, but the edges are where production incidents live: firewalls that silently block TCP/53, legacy clients that do not speak EDNS0, and known CoreDNS bugs where the TC flag never reaches the client at all.&lt;/p></description></item><item><title>CoreDNS returning REFUSED: no matching zone, an ACL, or the forward concurrency limit</title><link>https://www.netdata.cloud/guides/coredns/coredns-refused-responses/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-refused-responses/</guid><description>&lt;h1 id="coredns-returning-refused-no-matching-zone-an-acl-or-the-forward-concurrency-limit">CoreDNS returning REFUSED: no matching zone, an ACL, or the forward concurrency limit&lt;/h1>
&lt;p>CoreDNS is answering queries with &lt;code>REFUSED&lt;/code> and clients cannot resolve names. Unlike SERVFAIL, which means &amp;ldquo;I tried and failed,&amp;rdquo; REFUSED means &amp;ldquo;I declined to try.&amp;rdquo; REFUSED is almost never an upstream or network problem: CoreDNS itself is deciding not to answer, and it does that for exactly three reasons.&lt;/p>
&lt;p>Two are configuration problems: no server block matches the query&amp;rsquo;s zone, or the &lt;code>acl&lt;/code> plugin is rejecting the source. One is a capacity problem: the &lt;code>forward&lt;/code> plugin&amp;rsquo;s &lt;code>max_concurrent&lt;/code> limit is shedding load. Determine which of the three you are dealing with before touching the Corefile. Editing the Corefile to fix a capacity REFUSED, or scaling replicas to fix a config REFUSED, wastes the incident.&lt;/p></description></item><item><title>CoreDNS returning SERVFAIL: the resolver is failing queries and what to check first</title><link>https://www.netdata.cloud/guides/coredns/coredns-servfail-responses/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-servfail-responses/</guid><description>&lt;h1 id="coredns-returning-servfail-the-resolver-is-failing-queries-and-what-to-check-first">CoreDNS returning SERVFAIL: the resolver is failing queries and what to check first&lt;/h1>
&lt;p>Applications are failing to resolve names and CoreDNS is answering with SERVFAIL. In Kubernetes this surfaces as connection errors everywhere at once, because nearly all in-cluster communication depends on DNS. SERVFAIL is not one failure. It is CoreDNS saying &amp;ldquo;a plugin in my chain could not answer this query,&amp;rdquo; and the cause is usually upstream DNS, the Kubernetes API, or the Corefile itself.&lt;/p></description></item><item><title>CoreDNS serve_stale: keeping resolution alive while masking upstream failure</title><link>https://www.netdata.cloud/guides/coredns/coredns-serve-stale-masking-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-serve-stale-masking-failures/</guid><description>&lt;h1 id="coredns-serve_stale-keeping-resolution-alive-while-masking-upstream-failure">CoreDNS serve_stale: keeping resolution alive while masking upstream failure&lt;/h1>
&lt;p>The &lt;code>serve_stale&lt;/code> option in the CoreDNS cache plugin is a resilience feature with a monitoring side effect that bites teams during real incidents. When it is enabled, CoreDNS answers queries from expired cache entries instead of failing them when the upstream resolver is unreachable. Clients keep resolving names through an upstream outage. Dashboards stay green. The upstream can be dead for an hour before anyone notices, because the one metric that would tell you is almost never graphed.&lt;/p></description></item><item><title>CoreDNS SERVFAIL cache amplification: a one-second blip becomes a five-second outage</title><link>https://www.netdata.cloud/guides/coredns/coredns-servfail-cache-amplification/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-servfail-cache-amplification/</guid><description>&lt;h1 id="coredns-servfail-cache-amplification-a-one-second-blip-becomes-a-five-second-outage">CoreDNS SERVFAIL cache amplification: a one-second blip becomes a five-second outage&lt;/h1>
&lt;p>Your dashboards show an upstream DNS hiccup that lasted one second. Your users report an outage that lasted five. Both are accurate, and the gap between them is the CoreDNS cache plugin doing what it was designed to do.&lt;/p>
&lt;p>CoreDNS caches SERVFAIL responses for 5 seconds by default. The cache plugin keeps separate positive (&lt;code>success&lt;/code>) and negative (&lt;code>denial&lt;/code>) caches, and SERVFAIL goes into the denial cache alongside NXDOMAIN and NODATA. The intent is sound: RFC 2308 permits caching server-failure responses to shield an already-struggling upstream from a retry storm. The side effect is that a single SERVFAIL answer for a hot record is served to every client that asks during the 5-second window, even after the upstream has fully recovered.&lt;/p></description></item><item><title>CoreDNS serving stale cluster DNS: why you need a functional freshness test</title><link>https://www.netdata.cloud/guides/coredns/coredns-stale-dns-records-freshness-test/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-stale-dns-records-freshness-test/</guid><description>&lt;h1 id="coredns-serving-stale-cluster-dns-why-you-need-a-functional-freshness-test">CoreDNS serving stale cluster DNS: why you need a functional freshness test&lt;/h1>
&lt;p>Your CoreDNS dashboards are green. Query rate normal, latency sub-millisecond, SERVFAIL near zero, cache hit ratio healthy. Meanwhile, a Service deleted an hour ago still resolves, and the new Service your team just deployed returns NXDOMAIN. Applications are failing, but nothing in your monitoring fired.&lt;/p>
&lt;p>This is the silent stale data pattern. The kubernetes plugin builds its DNS records from a watch on the API server. If that watch disconnects and cannot re-establish, CoreDNS keeps answering from its in-memory snapshot. The snapshot drifts from reality. No errors go back to clients. No performance metric moves, because the server is genuinely healthy and fast. The data is simply wrong.&lt;/p></description></item><item><title>CoreDNS slow upstream: per-upstream latency, goroutine pileup, and the to label</title><link>https://www.netdata.cloud/guides/coredns/coredns-slow-upstream-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-slow-upstream-latency/</guid><description>&lt;h1 id="coredns-slow-upstream-per-upstream-latency-goroutine-pileup-and-the-to-label">CoreDNS slow upstream: per-upstream latency, goroutine pileup, and the to label&lt;/h1>
&lt;p>Overall DNS latency is climbing. P99 is above 500ms, but P50 looks only mildly elevated. SERVFAIL is nonzero but nowhere near a full outage. The &lt;code>/health&lt;/code> endpoint returns 200, every upstream passes its health checks, and yet clients are complaining that resolution is slow. Meanwhile &lt;code>go_goroutines&lt;/code> is trending up and heap is following it.&lt;/p>
&lt;p>This is the slow upstream drag pattern: one upstream DNS server is responding slowly but still returning valid answers. Because it never actually fails, the forward plugin&amp;rsquo;s health checks never mark it down, so queries keep getting routed to it. Each of those queries holds a goroutine while it waits. Latency accumulates, goroutines accumulate, and memory follows. Left alone, this ends in either &lt;code>max_concurrent&lt;/code> rejects (if you set one) or OOM (if you did not).&lt;/p></description></item><item><title>CoreDNS too many open files: file descriptor exhaustion and refused connections</title><link>https://www.netdata.cloud/guides/coredns/coredns-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-too-many-open-files/</guid><description>&lt;h1 id="coredns-too-many-open-files-file-descriptor-exhaustion-and-refused-connections">CoreDNS too many open files: file descriptor exhaustion and refused connections&lt;/h1>
&lt;p>CoreDNS pods are logging &lt;code>too many open files&lt;/code> and DNS resolution is failing intermittently or completely. Clients see timeouts or refused connections while the process keeps running and &lt;code>/health&lt;/code> still returns 200. Restarting the pod fixes it for a while, then it comes back.&lt;/p>
&lt;p>This is file descriptor exhaustion, and it is a cliff-edge failure. The moment &lt;code>process_open_fds&lt;/code> reaches &lt;code>process_max_fds&lt;/code>, CoreDNS cannot open anything new: no upstream connections, no listening sockets, no log files, no Kubernetes API watch streams. There is no queuing and no graceful degradation. New connection attempts fail immediately and DNS breaks.&lt;/p></description></item><item><title>CoreDNS TTL=0 responses: the anti-pattern that silently bypasses the cache</title><link>https://www.netdata.cloud/guides/coredns/coredns-ttl-zero-defeats-cache/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-ttl-zero-defeats-cache/</guid><description>&lt;h1 id="coredns-ttl0-responses-the-anti-pattern-that-silently-bypasses-the-cache">CoreDNS TTL=0 responses: the anti-pattern that silently bypasses the cache&lt;/h1>
&lt;p>Your cache hit ratio is sliding. Upstream query volume is climbing. P99 latency is drifting up with it. You open the Corefile and the &lt;code>cache&lt;/code> plugin is right there, configured the way it has been for months. Nothing changed on your side, but the cache has effectively stopped working.&lt;/p>
&lt;p>A common cause is records arriving with TTL=0. When an upstream resolver, or your own zone data, answers with a zero TTL, those answers are uncacheable by contract: every client query for that name has to be forwarded again. The cache plugin is present and correct, but there is nothing for it to hold on to. It looks like a cache failure but is actually a data problem arriving through the response path.&lt;/p></description></item><item><title>CoreDNS UDP buffer errors: kernel receive-buffer overflow dropping queries</title><link>https://www.netdata.cloud/guides/coredns/coredns-udp-buffer-errors-packet-loss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-udp-buffer-errors-packet-loss/</guid><description>&lt;h1 id="coredns-udp-buffer-errors-kernel-receive-buffer-overflow-dropping-queries">CoreDNS UDP buffer errors: kernel receive-buffer overflow dropping queries&lt;/h1>
&lt;p>Applications are reporting intermittent DNS timeouts. You open the CoreDNS dashboard and everything looks fine: latency is low, SERVFAIL rate is zero, the health endpoint returns 200, and QPS is suspiciously flat. Not spiking, not zero, just lower than the client demand you know exists.&lt;/p>
&lt;p>This is the UDP buffer cliff. Incoming UDP traffic is exceeding the kernel socket receive buffer on the node, and the kernel is dropping DNS queries before CoreDNS ever reads them. Because CoreDNS only counts the packets it actually receives, every CoreDNS-level metric looks clean. The failure lives one layer below, in the kernel.&lt;/p></description></item><item><title>CoreDNS upstream connection cache misses: new connections adding latency per query</title><link>https://www.netdata.cloud/guides/coredns/coredns-upstream-connection-cache-misses/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-upstream-connection-cache-misses/</guid><description>&lt;h1 id="coredns-upstream-connection-cache-misses-new-connections-adding-latency-per-query">CoreDNS upstream connection cache misses: new connections adding latency per query&lt;/h1>
&lt;p>You are here because forwarded DNS queries got slower, or because &lt;code>coredns_proxy_conn_cache_misses_total&lt;/code> is climbing and you need to know whether it matters. It does. Every connection cache miss means CoreDNS opens a fresh connection to an upstream resolver instead of reusing one from its pool. Setup cost lands on that query, and each new connection holds a file descriptor until it is closed or reaped.&lt;/p></description></item><item><title>CoreDNS with NodeLocal DNSCache: what changes about monitoring</title><link>https://www.netdata.cloud/guides/coredns/coredns-nodelocaldns-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-nodelocaldns-monitoring/</guid><description>&lt;h1 id="coredns-with-nodelocal-dnscache-what-changes-about-monitoring">CoreDNS with NodeLocal DNSCache: what changes about monitoring&lt;/h1>
&lt;p>You deployed NodeLocal DNSCache to fix the 5-second timeout race and conntrack pressure, and it worked. DNS latency dropped, the conntrack table stopped filling, and CoreDNS QPS fell off a cliff. Then someone looked at the CoreDNS dashboard, saw near-zero traffic, and concluded CoreDNS was oversized. That conclusion is how the next incident starts.&lt;/p>
&lt;p>NodeLocal DNSCache runs a caching DNS proxy as a DaemonSet on every node. It intercepts pod DNS queries before they enter the iptables DNAT and conntrack path, so each query is answered on the local node instead of traversing NAT to a central CoreDNS pod. This eliminates the kernel race condition behind the 5-second glibc timeout and removes DNS as a conntrack consumer.&lt;/p></description></item><item><title>Coriolis Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/coriolis-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/coriolis-networks-snmp-traps/</guid><description/></item><item><title>Correlating cloud VPC flow logs with on-prem NetFlow</title><link>https://www.netdata.cloud/guides/network/network-cloud-onprem-flow-correlation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-cloud-onprem-flow-correlation/</guid><description>&lt;h1 id="correlating-cloud-vpc-flow-logs-with-on-prem-netflow">Correlating cloud VPC flow logs with on-prem NetFlow&lt;/h1>
&lt;p>Cloud flow logs and on-prem flow records share the 5-tuple concept but diverge in nearly every dimension that matters for correlation: transport, latency, sampling, timestamps, topology, and NAT visibility. Cloud providers emit VPC flow logs via push to object storage with implicit sampling and aggregation intervals measured in minutes. On-premises devices export NetFlow v5/v9, IPFIX, or sFlow over UDP with configurable sampling and near-real-time delivery.&lt;/p></description></item><item><title>Cortex</title><link>https://www.netdata.cloud/integrations/exporters/cortex/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/cortex/</guid><description/></item><item><title>Cosine Communications SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cosine-communications-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cosine-communications-snmp-traps/</guid><description/></item><item><title>Couchbase</title><link>https://www.netdata.cloud/integrations/data-collection/databases/couchbase/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/couchbase/</guid><description/></item><item><title>Couchbase Monitoring</title><link>https://www.netdata.cloud/monitoring-101/couchbase-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/couchbase-monitoring/</guid><description>&lt;h2 id="couchbase-monitoring">Couchbase Monitoring&lt;/h2>
&lt;h3 id="what-is-couchbase">What Is Couchbase?&lt;/h3>
&lt;p>Couchbase is a distributed NoSQL cloud database ideal for interactive web and mobile applications. Designed to support massive data volumes and a large number of users, Couchbase provides core database functions while enhancing performance through its advanced memory-first architecture.&lt;/p>
&lt;h3 id="monitoring-couchbase-with-netdata">Monitoring Couchbase With Netdata&lt;/h3>
&lt;p>Monitoring Couchbase’s health and performance is essential to maintaining its operations-efficiently. With Netdata, you can seamlessly integrate Couchbase monitoring and get real-time insights into vital metrics. Netdata’s intuitive dashboard provides detailed metrics, enabling you to keep a close watch on Couchbase’s performance.&lt;/p></description></item><item><title>CouchDB</title><link>https://www.netdata.cloud/integrations/data-collection/databases/couchdb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/couchdb/</guid><description/></item><item><title>CouchDB Monitoring</title><link>https://www.netdata.cloud/monitoring-101/couchdb-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/couchdb-monitoring/</guid><description>&lt;h2 id="couchdb-monitoring">CouchDB Monitoring&lt;/h2>
&lt;h3 id="what-is-couchdb">What Is CouchDB?&lt;/h3>
&lt;p>CouchDB is an open-source database developed by the Apache Software Foundation, known for its seamless storage and retrieval of JSON documents. It features a distributed architecture with easy replication, making it a popular choice for web applications and mobile apps that require offline-first functionality. Learn more on the &lt;a href="https://couchdb.apache.org/">CouchDB official site&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-couchdb-with-netdata">Monitoring CouchDB With Netdata&lt;/h3>
&lt;p>Monitoring CouchDB is crucial to ensure optimal performance and reliability of your database infrastructure. Netdata provides powerful tools for monitoring CouchDB, offering real-time insights into various metrics such as database activity, HTTP requests, active tasks, and more. Leveraging the Netdata agent, you can efficiently track these metrics and respond swiftly to potential issues.&lt;/p></description></item><item><title>CPU performance</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/cpu-performance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/cpu-performance/</guid><description/></item><item><title>Cradlepoint</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cradlepoint/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cradlepoint/</guid><description/></item><item><title>CraftBeerPi</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/craftbeerpi/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/craftbeerpi/</guid><description/></item><item><title>CraftBeerPi Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ftbeerpi-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ftbeerpi-monitoring/</guid><description>&lt;h2 id="craftbeerpi-monitoring">CraftBeerPi Monitoring&lt;/h2>
&lt;h3 id="what-is-craftbeerpi">What Is CraftBeerPi?&lt;/h3>
&lt;p>CraftBeerPi is an open-source brewing automation software designed to revolutionize the homebrewing experience. By integrating with various sensors and controllers, CraftBeerPi helps manage the entire brewing process, ensuring precision and consistency with every batch.&lt;/p>
&lt;h3 id="monitoring-craftbeerpi-with-netdata">Monitoring CraftBeerPi With Netdata&lt;/h3>
&lt;p>When it comes to ensuring the peak performance of your brewing setup, monitoring CraftBeerPi with Netdata is an optimal choice. To monitor CraftBeerPi, Netdata employs an openmetrics (Prometheus) exporter, specifically the &lt;a href="https://github.com/jo-hannes/craftbeerpi_exporter">CraftBeerPi exporter&lt;/a>. This integration allows Netdata to ingest data from any Prometheus exporter seamlessly.&lt;/p></description></item><item><title>CrateDB</title><link>https://www.netdata.cloud/integrations/exporters/cratedb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/cratedb/</guid><description/></item><item><title>Cray Communications A S SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cray-communications-a-s-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cray-communications-a-s-snmp-traps/</guid><description/></item><item><title>Cray SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cray-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cray-snmp-traps/</guid><description/></item><item><title>Croix Rouge Francaise SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/croix-rouge-francaise-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/croix-rouge-francaise-snmp-traps/</guid><description/></item><item><title>Crossbeam Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/crossbeam-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/crossbeam-systems-inc-snmp-traps/</guid><description/></item><item><title>Crowdsec</title><link>https://www.netdata.cloud/integrations/data-collection/applications/crowdsec/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/crowdsec/</guid><description/></item><item><title>Crowdsec Monitoring</title><link>https://www.netdata.cloud/monitoring-101/crowdsec-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/crowdsec-monitoring/</guid><description>&lt;h2 id="crowdsec-monitoring">Crowdsec Monitoring&lt;/h2>
&lt;h3 id="what-is-crowdsec">What Is Crowdsec?&lt;/h3>
&lt;p>Crowdsec is an open-source, collaborative security solution designed to detect and respond to threats. It emphasizes community collaboration to build a robust defense against cyber threats by leveraging shared threat intelligence.&lt;/p>
&lt;h3 id="monitoring-crowdsec-with-netdata">Monitoring Crowdsec With Netdata&lt;/h3>
&lt;p>To effectively monitor Crowdsec, Netdata utilizes an openmetrics (Prometheus) exporter. This approach allows Netdata to assimilate data from any Prometheus exporter, providing seamless integration for users without the necessity of deploying a Prometheus server or Grafana. With automatic dashboards and alerting capabilities, Netdata transforms your Crowdsec monitoring setup into a high-performance, real-time monitoring system.&lt;/p></description></item><item><title>Cryptowatch</title><link>https://www.netdata.cloud/integrations/data-collection/applications/cryptowatch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/cryptowatch/</guid><description/></item><item><title>Cryptowatch Monitoring</title><link>https://www.netdata.cloud/monitoring-101/cryptowatch-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/cryptowatch-monitoring/</guid><description>&lt;h2 id="cryptowatch-monitoring">Cryptowatch Monitoring&lt;/h2>
&lt;h3 id="what-is-cryptowatch">What Is Cryptowatch?&lt;/h3>
&lt;p>Cryptowatch is a powerful platform that provides market information for cryptocurrencies, allowing traders and analysts to track price changes, volume, and other critical metrics in real-time. It aggregates data from various exchanges to present a unified view of cryptocurrency markets, making it an indispensable tool for anyone involved in this digital asset space.&lt;/p>
&lt;h3 id="monitoring-cryptowatch-with-netdata">Monitoring Cryptowatch With Netdata&lt;/h3>
&lt;p>Netdata&amp;rsquo;s advanced monitoring capabilities make it an excellent choice for keeping tabs on Cryptowatch. By using an openmetrics (Prometheus) exporter, specifically the &lt;a href="https://github.com/nbarrientos/cryptowat_exporter">Cryptowatch Exporter&lt;/a>, Netdata can seamlessly ingest data from Cryptowatch. This process does not require a full Prometheus server or Grafana setup, providing automated dashboards and alerts out of the box. This simplicity and integration capability mark Netdata as a leading Cryptowatch monitoring tool in the market.&lt;/p></description></item><item><title>Ctc Union Technologies Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ctc-union-technologies-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ctc-union-technologies-co-ltd-snmp-traps/</guid><description/></item><item><title>Cube Optics AG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cube-optics-ag-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cube-optics-ag-snmp-traps/</guid><description/></item><item><title>CUDA out of memory with free memory available: GPU memory fragmentation</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-cuda-oom-with-free-memory/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-cuda-oom-with-free-memory/</guid><description>&lt;h1 id="cuda-out-of-memory-with-free-memory-available-gpu-memory-fragmentation">CUDA out of memory with free memory available: GPU memory fragmentation&lt;/h1>
&lt;p>Your training run or inference server has been up for hours. Then it dies with &lt;code>torch.cuda.OutOfMemoryError: CUDA out of memory&lt;/code>. You check &lt;code>nvidia-smi&lt;/code> and the GPU shows gigabytes free. The error message and the driver disagree, and the driver looks right.&lt;/p>
&lt;p>Both are right. The GPU has free memory in aggregate, but no single contiguous block large enough for the allocation that just failed. The memory is fragmented: free in total, unusable in practice.&lt;/p></description></item><item><title>CUDA out of memory: diagnosing NVIDIA GPU framebuffer exhaustion</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-cuda-out-of-memory/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-cuda-out-of-memory/</guid><description>&lt;h1 id="cuda-out-of-memory-diagnosing-nvidia-gpu-framebuffer-exhaustion">CUDA out of memory: diagnosing NVIDIA GPU framebuffer exhaustion&lt;/h1>
&lt;p>Your training job or inference server died with &lt;code>RuntimeError: CUDA out of memory. Tried to allocate X MiB&lt;/code>. Sometimes the GPU is genuinely full. More often, &lt;code>nvidia-smi&lt;/code> shows gigabytes free and the error makes no sense at first glance.&lt;/p>
&lt;p>GPU framebuffer is a cliff-edge resource. Unlike CPU memory there is no swap and no graceful degradation: when a &lt;code>cudaMalloc&lt;/code> cannot be satisfied, the allocation fails atomically and the process dies. It is also one of the most misdiagnosed GPU failures, because the number &lt;code>nvidia-smi&lt;/code> shows is not the number your framework is working with.&lt;/p></description></item><item><title>Cumulus Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cumulus-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cumulus-networks-inc-snmp-traps/</guid><description/></item><item><title>CUPS</title><link>https://www.netdata.cloud/integrations/data-collection/applications/cups/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/cups/</guid><description/></item><item><title>Current_Pending_Sector non-zero: unreadable sectors and I/O latency spikes</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-current-pending-sector/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-current-pending-sector/</guid><description>&lt;h1 id="current_pending_sector-non-zero-unreadable-sectors-and-io-latency-spikes">Current_Pending_Sector non-zero: unreadable sectors and I/O latency spikes&lt;/h1>
&lt;p>A non-zero &lt;code>Current_Pending_Sector&lt;/code> (SMART attribute ID 197) means the drive has sectors it cannot reliably read. These sectors failed a read and are waiting for either a write to trigger reallocation or an offline scan to confirm the defect. The drive has not yet remapped them to its spare pool.&lt;/p>
&lt;p>The operational impact is often misdiagnosed. Every time the filesystem touches one of these sectors, the drive firmware enters a multi-second retry loop. This shows up in &lt;code>iostat&lt;/code> as extremely high &lt;code>await&lt;/code> with low throughput and idle CPU. Teams chase this as a software freeze, a kernel bug, or a filesystem problem for hours before checking SMART data.&lt;/p></description></item><item><title>Custom</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/custom/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/custom/</guid><description/></item><item><title>Custom MMDB Database</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/custom-mmdb-database/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/custom-mmdb-database/</guid><description/></item><item><title>Cxr SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cxr-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cxr-snmp-traps/</guid><description/></item><item><title>Cyan SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cyan-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cyan-snmp-traps/</guid><description/></item><item><title>Cyber Ark SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cyber-ark-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cyber-ark-snmp-traps/</guid><description/></item><item><title>Cyber Power System Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cyber-power-system-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cyber-power-system-inc-snmp-traps/</guid><description/></item><item><title>Cyberguard Corporationdavid Rhein SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cyberguard-corporationdavid-rhein-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/cyberguard-corporationdavid-rhein-snmp-traps/</guid><description/></item><item><title>Cyberpower PDU</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cyberpower-pdu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/cyberpower-pdu/</guid><description/></item><item><title>D Link Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/d-link-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/d-link-systems-inc-snmp-traps/</guid><description/></item><item><title>Dantel Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dantel-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dantel-inc-snmp-traps/</guid><description/></item><item><title>Dantherm Cooling A S SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dantherm-cooling-a-s-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dantherm-cooling-a-s-snmp-traps/</guid><description/></item><item><title>Danware Data A S SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/danware-data-a-s-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/danware-data-a-s-snmp-traps/</guid><description/></item><item><title>Dasan Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dasan-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dasan-co-ltd-snmp-traps/</guid><description/></item><item><title>Data Domain Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/data-domain-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/data-domain-inc-snmp-traps/</guid><description/></item><item><title>Data Units Written vs rated TBW: computing SSD endurance runway</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-data-units-written-tbw/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-data-units-written-tbw/</guid><description>&lt;h1 id="data-units-written-vs-rated-tbw-computing-ssd-endurance-runway">Data Units Written vs rated TBW: computing SSD endurance runway&lt;/h1>
&lt;p>SSD endurance is finite and specified in the datasheet as a TBW (Total Bytes Written) rating. The odometer is a SMART counter, but it uses non-obvious units and counts host writes, not NAND writes. This guide covers how to read the write-volume counter from smartctl, convert it to bytes correctly, compare it against rated TBW, and project when the drive reaches its endurance limit.&lt;/p></description></item><item><title>Datadirect Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/datadirect-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/datadirect-networks-snmp-traps/</guid><description/></item><item><title>Datalogic S P A SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/datalogic-s-p-a-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/datalogic-s-p-a-snmp-traps/</guid><description/></item><item><title>Datapower Technology Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/datapower-technology-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/datapower-technology-inc-snmp-traps/</guid><description/></item><item><title>Dataprobe Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dataprobe-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dataprobe-inc-snmp-traps/</guid><description/></item><item><title>DB-IP IP Intelligence</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/db-ip-ip-intelligence/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/db-ip-ip-intelligence/</guid><description/></item><item><title>DCGM nv-hostengine down or hung: GPU telemetry goes dark or stale</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-dcgm-daemon-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-dcgm-daemon-down/</guid><description>&lt;h1 id="dcgm-nv-hostengine-down-or-hung-gpu-telemetry-goes-dark-or-stale">DCGM nv-hostengine down or hung: GPU telemetry goes dark or stale&lt;/h1>
&lt;p>Your GPU dashboards flatline, or worse, they do not. The graphs keep rendering smooth, plausible values for temperature, utilization, and power, but nothing on the node has actually changed in twenty minutes. Both symptoms point at the same component: nv-hostengine, the DCGM daemon that sits between NVML and every monitoring tool you run.&lt;/p>
&lt;p>nv-hostengine fails in two distinct ways, and only one of them is obvious. The obvious failure is a dead process: no data, gaps in every time series, alerts firing on absent metrics. The dangerous failure is an alive-but-hung daemon: the process exists, the socket accepts connections, and DCGM serves stale values from its in-memory field value cache because an underlying NVML call is blocked on the driver. Dashboards look normal. Threshold alerts stay quiet because the cached values are inside bounds. You find out when a human notices the numbers have not moved.&lt;/p></description></item><item><title>DCGM, nvidia-smi, and NVML: which NVIDIA GPU telemetry source to use</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-dcgm-vs-nvidia-smi-nvml/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-dcgm-vs-nvidia-smi-nvml/</guid><description>&lt;h1 id="dcgm-nvidia-smi-and-nvml-which-nvidia-gpu-telemetry-source-to-use">DCGM, nvidia-smi, and NVML: which NVIDIA GPU telemetry source to use&lt;/h1>
&lt;p>Every NVIDIA GPU troubleshooting session starts with a telemetry question: where do I get this number, and can I trust it? Teams routinely mix the three available sources without realizing they sit at different layers of the same stack. They grep nvidia-smi output in cron jobs, run DCGM and nvidia-smi side by side and wonder why both get slow, or alert on a DCGM field that silently returns zeros on their driver version.&lt;/p></description></item><item><title>Deal Registration</title><link>https://www.netdata.cloud/partner-deal-registrations/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/partner-deal-registrations/</guid><description/></item><item><title>Debian</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/debian/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/debian/</guid><description/></item><item><title>Debian SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/debian-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/debian-snmp-traps/</guid><description/></item><item><title>Dec SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dec-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dec-snmp-traps/</guid><description/></item><item><title>Decapsulation</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/decapsulation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/decapsulation/</guid><description/></item><item><title>Deep Sea Electronics PLC SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/deep-sea-electronics-plc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/deep-sea-electronics-plc-snmp-traps/</guid><description/></item><item><title>Del Mar Solutions Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/del-mar-solutions-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/del-mar-solutions-inc-snmp-traps/</guid><description/></item><item><title>Deliberant SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/deliberant-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/deliberant-snmp-traps/</guid><description/></item><item><title>Dell</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dell/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dell/</guid><description/></item><item><title>Dell BGP</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/dell-bgp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/dell-bgp/</guid><description/></item><item><title>Dell EMC Data Domain</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dell-emc-data-domain/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dell-emc-data-domain/</guid><description/></item><item><title>Dell EMC ScaleIO</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/dell-emc-scaleio/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/dell-emc-scaleio/</guid><description/></item><item><title>Dell EMC ScaleIO Monitoring</title><link>https://www.netdata.cloud/monitoring-101/scaleio-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/scaleio-monitoring/</guid><description>&lt;h2 id="dell-emc-scaleio-monitoring">Dell EMC ScaleIO Monitoring&lt;/h2>
&lt;h3 id="what-is-dell-emc-scaleio">What Is Dell EMC ScaleIO?&lt;/h3>
&lt;p>Dell EMC ScaleIO, also known as VxFlex OS, is a software-defined storage solution designed to turn local storage into shared block storage. ScaleIO is tailored for modern data centers, providing enhanced flexibility and efficiency in handling storage needs. You can learn more about ScaleIO &lt;a href="https://www.dell.com/en-ca/dt/storage/scaleio/scaleioreadynode.htm">here&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-scaleio-with-netdata">Monitoring ScaleIO With Netdata&lt;/h3>
&lt;p>Netdata offers a comprehensive and user-friendly platform to monitor ScaleIO instances effectively. By utilizing Netdata&amp;rsquo;s ScaleIO monitoring tools, you can collect detailed metrics from ScaleIO components via the VxFlex OS Gateway API. This enables real-time insights and actionable analytics, crucial for maintaining optimal performance.&lt;/p></description></item><item><title>Dell Force10</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dell-force10/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dell-force10/</guid><description/></item><item><title>Dell Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dell-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dell-inc-snmp-traps/</guid><description/></item><item><title>Dell OS10</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dell-os10/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dell-os10/</guid><description/></item><item><title>Dell Powerconnect</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dell-powerconnect/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dell-powerconnect/</guid><description/></item><item><title>Dell Poweredge</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dell-poweredge/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dell-poweredge/</guid><description/></item><item><title>Dell PowerStore</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/dell-powerstore/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/dell-powerstore/</guid><description/></item><item><title>Dell PowerVault ME4/ME5</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/dell-powervault-me4-me5/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/dell-powervault-me4-me5/</guid><description/></item><item><title>Dell Sonicwall</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dell-sonicwall/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dell-sonicwall/</guid><description/></item><item><title>Delta Electronics Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/delta-electronics-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/delta-electronics-inc-snmp-traps/</guid><description/></item><item><title>Delta Electronics Switzerland AG Formerly Delta Energy Systems Sweden AB SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/delta-electronics-switzerland-ag-formerly-delta-energy-systems-sweden-ab-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/delta-electronics-switzerland-ag-formerly-delta-energy-systems-sweden-ab-snmp-traps/</guid><description/></item><item><title>Deltanet AG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/deltanet-ag-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/deltanet-ag-snmp-traps/</guid><description/></item><item><title>Designer Systems Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/designer-systems-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/designer-systems-ltd-snmp-traps/</guid><description/></item><item><title>dev.cpu.0.freq</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/dev.cpu.0.freq/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/dev.cpu.0.freq/</guid><description/></item><item><title>dev.cpu.temperature</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/dev.cpu.temperature/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/dev.cpu.temperature/</guid><description/></item><item><title>Deva Broadcast Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/deva-broadcast-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/deva-broadcast-ltd-snmp-traps/</guid><description/></item><item><title>Device control-plane CPU saturation: when SNMP polling causes the spike</title><link>https://www.netdata.cloud/guides/network/network-device-control-plane-cpu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-device-control-plane-cpu/</guid><description>&lt;h1 id="device-control-plane-cpu-saturation-when-snmp-polling-causes-the-spike">Device control-plane CPU saturation: when SNMP polling causes the spike&lt;/h1>
&lt;p>The router&amp;rsquo;s control-plane CPU is pinned at 95%. SNMP polls are timing out. BGP sessions are approaching hold-time expiry. The instinct is to blame the device or suspect an attack, but the monitoring system itself is frequently the source of the load.&lt;/p>
&lt;p>The control-plane CPU handles everything that is not hardware-forwarded packet switching: the SNMP agent, BGP, OSPF, STP, the CLI, syslog, AAA, and management interfaces. When SNMP polling saturates this CPU, every control-plane function degrades at once. The symptoms look like a device problem. The cause is often on the collector side.&lt;/p></description></item><item><title>Device memory pressure: control-plane memory pools and leaks</title><link>https://www.netdata.cloud/guides/network/network-device-memory-pressure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-device-memory-pressure/</guid><description>&lt;h1 id="device-memory-pressure-control-plane-memory-pools-and-leaks">Device memory pressure: control-plane memory pools and leaks&lt;/h1>
&lt;p>A network device alerts at 85% control-plane memory utilization. Or worse: it does not alert, and you discover the problem when BGP sessions start dropping from hold-time expiry, SNMP stops responding, or the device reboots itself. Control-plane memory exhaustion degrades every process on the route processor: routing protocols, CLI, SNMP, management interfaces, logging.&lt;/p>
&lt;p>&amp;ldquo;High memory utilization&amp;rdquo; on a network device is ambiguous. Some platforms cache aggressively and report 97% used under normal operation. Some pools are expected to sit near zero free. A genuine memory leak may take weeks to manifest, making it easy to dismiss the trend until the device crashes at 3 a.m.&lt;/p></description></item><item><title>devstat</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/devstat/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/devstat/</guid><description/></item><item><title>Dialogic Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dialogic-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dialogic-corporation-snmp-traps/</guid><description/></item><item><title>Dialogic Media Gateway</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dialogic-media-gateway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dialogic-media-gateway/</guid><description/></item><item><title>Didactum Security GmbH Formerly Didactum Ltd Deutschland SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/didactum-security-gmbh-formerly-didactum-ltd-deutschland-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/didactum-security-gmbh-formerly-didactum-ltd-deutschland-snmp-traps/</guid><description/></item><item><title>Digipower Manufacturing Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/digipower-manufacturing-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/digipower-manufacturing-inc-snmp-traps/</guid><description/></item><item><title>Digital China Shanghai Networks Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/digital-china-shanghai-networks-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/digital-china-shanghai-networks-ltd-snmp-traps/</guid><description/></item><item><title>Digital Link SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/digital-link-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/digital-link-snmp-traps/</guid><description/></item><item><title>Digital Video Broadcasting Dvb SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/digital-video-broadcasting-dvb-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/digital-video-broadcasting-dvb-snmp-traps/</guid><description/></item><item><title>Discord</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/discord/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/discord/</guid><description/></item><item><title>Discord</title><link>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/discord/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/discord/</guid><description/></item><item><title>Discourse</title><link>https://www.netdata.cloud/integrations/data-collection/applications/discourse/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/discourse/</guid><description/></item><item><title>Discourse Monitoring</title><link>https://www.netdata.cloud/monitoring-101/discourse-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/discourse-monitoring/</guid><description>&lt;h2 id="discourse-monitoring">Discourse Monitoring&lt;/h2>
&lt;h3 id="what-is-discourse">What Is Discourse?&lt;/h3>
&lt;p>&lt;strong>Discourse&lt;/strong> is a modern forum software for your community. It reimagines what a discussion platform should be, offering seamless integration with existing web apps, and a mobile-optimized experience.&lt;/p>
&lt;h3 id="monitoring-discourse-with-netdata">Monitoring Discourse With Netdata&lt;/h3>
&lt;p>To monitor Discourse effectively and ensure optimal performance, you can rely on Netdata—a powerful real-time monitoring solution. Netdata leverages the Discourse Prometheus exporter to collect relevant metrics. This allows users to monitor Discourse without the need for setting up a Prometheus server or Grafana. Automated dashboards, alerts, and visual insights are readily available from Netdata&amp;rsquo;s platform.&lt;/p></description></item><item><title>Disk space</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/disk-space/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/disk-space/</guid><description/></item><item><title>Disk Statistics</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/disk-statistics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/disk-statistics/</guid><description/></item><item><title>Dismuntel S A L SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dismuntel-s-a-l-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dismuntel-s-a-l-snmp-traps/</guid><description/></item><item><title>Distributed Management Task Force Dmtf SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/distributed-management-task-force-dmtf-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/distributed-management-task-force-dmtf-snmp-traps/</guid><description/></item><item><title>Distributed Processing Technology SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/distributed-processing-technology-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/distributed-processing-technology-snmp-traps/</guid><description/></item><item><title>Dlink</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dlink/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dlink/</guid><description/></item><item><title>Dlink DGS Switch</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dlink-dgs-switch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/dlink-dgs-switch/</guid><description/></item><item><title>DMARC</title><link>https://www.netdata.cloud/integrations/data-collection/applications/dmarc/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/dmarc/</guid><description/></item><item><title>DMARC Monitoring</title><link>https://www.netdata.cloud/monitoring-101/dmarc-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/dmarc-monitoring/</guid><description>&lt;h2 id="dmarc-monitoring">DMARC Monitoring&lt;/h2>
&lt;h3 id="what-is-dmarc">What Is DMARC?&lt;/h3>
&lt;p>DMARC (Domain-based Message Authentication, Reporting &amp;amp; Conformance) is an email authentication policy and reporting protocol. It is designed to give domain owners a way to protect their domain from being used in email spoofing, phishing scams, and other cybercrimes. By implementing DMARC, organizations can significantly improve the security and credibility of their email communications.&lt;/p>
&lt;h3 id="monitoring-dmarc-with-netdata">Monitoring DMARC With Netdata&lt;/h3>
&lt;p>Netdata provides a robust DMARC monitoring tool using the openmetrics exporter from Prometheus. Specifically, it uses the &lt;a href="https://github.com/jgosmann/dmarc-metrics-exporter">dmarc-metrics-exporter&lt;/a> to collect detailed metrics that offer insights into email authentication processes. With Netdata, you can ingest data from any Prometheus exporter and instantly gain access to automated dashboards and alerts without necessitating a Prometheus server or Grafana setup. This streamlined setup helps you keep a close watch on DMARC metrics in a user-friendly manner.&lt;/p></description></item><item><title>DMCache devices</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/dmcache-devices/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/dmcache-devices/</guid><description/></item><item><title>DMCache Monitoring</title><link>https://www.netdata.cloud/monitoring-101/dmcache-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/dmcache-monitoring/</guid><description>&lt;h2 id="dmcache-monitoring">DMCache Monitoring&lt;/h2>
&lt;h3 id="what-is-dmcache">What Is DMCache?&lt;/h3>
&lt;p>DMCache, a critical component of storage systems, acts as a device mapper cache enabling systems to manage disk block caching. By leveraging DMCache, systems can significantly enhance I/O performance by storing frequently accessed disk blocks in a faster storage medium.&lt;/p>
&lt;h3 id="monitoring-dmcache-with-netdata">Monitoring DMCache With Netdata&lt;/h3>
&lt;p>Effective DMCache monitoring ensures that your storage systems are performing optimally and that resources are efficiently utilized. The &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/dmcache/">Netdata DMCache monitoring tool&lt;/a> offers real-time visibility into your DMCache devices, allowing you to track critical performance metrics and gain actionable insights through comprehensive visualizations.&lt;/p></description></item><item><title>DNS query</title><link>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/dns-query/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/dns-query/</guid><description/></item><item><title>DNS Query Monitoring</title><link>https://www.netdata.cloud/monitoring-101/dnsquery-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/dnsquery-monitoring/</guid><description>&lt;h2 id="dns-query-monitoring">DNS Query Monitoring&lt;/h2>
&lt;h3 id="what-is-dns-query">What Is DNS Query?&lt;/h3>
&lt;p>DNS Query refers to the process by which an inquiry is made to the Domain Name System to obtain information about a domain, such as the corresponding IP address. It&amp;rsquo;s a critical function that ensures users can access websites and services smoothly.&lt;/p>
&lt;h3 id="monitoring-dns-query-with-netdata">Monitoring DNS Query With Netdata&lt;/h3>
&lt;p>Netdata offers a robust DNS Query monitoring tool that allows users to track the round-trip time (RTT) of DNS queries. Utilizing Netdata, you can efficiently monitor DNS query performance, ensuring that your applications and services maintain optimal connectivity.&lt;/p></description></item><item><title>DNSBL</title><link>https://www.netdata.cloud/integrations/data-collection/networking/dnsbl/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/dnsbl/</guid><description/></item><item><title>DNSBL Monitoring</title><link>https://www.netdata.cloud/monitoring-101/dnsbl-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/dnsbl-monitoring/</guid><description>&lt;h2 id="dnsbl-monitoring">DNSBL Monitoring&lt;/h2>
&lt;h3 id="what-is-dnsbl">What Is DNSBL?&lt;/h3>
&lt;p>DNSBL, or Domain Name System-based Blackhole List, is a mechanism that helps identify and block email messages from known spammers by using a real-time database of blacklisted IP addresses often employed by email servers to filter out unwanted messages and reduce spam.&lt;/p>
&lt;h3 id="monitoring-dnsbl-with-netdata">Monitoring DNSBL With Netdata&lt;/h3>
&lt;p>To monitor DNSBL, Netdata utilizes an openmetrics (Prometheus) exporter. This approach ensures that users can gather comprehensive metrics related to DNSBL effectively. With Netdata, there is no necessity for a dedicated Prometheus server or Grafana setup as Netdata integrates seamlessly with any Prometheus exporter, showcasing automated dashboards, alerts, and in-depth visibility into DNSBL metrics such as domain reputation and security management.&lt;/p></description></item><item><title>DNSdist</title><link>https://www.netdata.cloud/integrations/data-collection/networking/dnsdist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/dnsdist/</guid><description/></item><item><title>DNSdist Monitoring</title><link>https://www.netdata.cloud/monitoring-101/dnsdist-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/dnsdist-monitoring/</guid><description>&lt;h2 id="dnsdist-monitoring">DNSdist Monitoring&lt;/h2>
&lt;h3 id="what-is-dnsdist">What Is DNSdist?&lt;/h3>
&lt;p>DNSdist is a highly flexible and powerful DNS proxy server designed to handle a wide range of queries efficiently. It helps balance and protect your DNS infrastructure with advanced features like query filtering, load balancing, and telemetry.&lt;/p>
&lt;h3 id="monitoring-dnsdist-with-netdata">Monitoring DNSdist With Netdata&lt;/h3>
&lt;p>Netdata provides unparalleled visibility into DNSdist operations through its comprehensive dashboard. By using the &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/dnsdist/">DNSdist monitoring tool&lt;/a>, you can quickly understand traffic patterns, identify bottlenecks, and ensure DNS reliability.&lt;/p></description></item><item><title>Dnsmasq</title><link>https://www.netdata.cloud/integrations/data-collection/networking/dnsmasq/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/dnsmasq/</guid><description/></item><item><title>Dnsmasq DHCP</title><link>https://www.netdata.cloud/integrations/data-collection/networking/dnsmasq-dhcp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/dnsmasq-dhcp/</guid><description/></item><item><title>Dnsmasq DHCP Monitoring</title><link>https://www.netdata.cloud/monitoring-101/dnsmasq-dhcp-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/dnsmasq-dhcp-monitoring/</guid><description>&lt;h2 id="what-is-dnsmasq-for-dhcp">What is Dnsmasq for DHCP?&lt;/h2>
&lt;p>&lt;a href="https://thekelleys.org.uk/dnsmasq/doc.html">Dnsmasq&lt;/a> is an open-source, lightweight, DNS caching and forwarding server. It is designed to provide DNS resolution for small and home networks. Dnsmasq provides local DNS caching, forwarding, and recursive lookups, as well as DHCP, TFTP, and other related services. It also has support for DNS and DHCPv6, as well as various other features such as DNS-over-TLS and IPv6 privacy extensions. Dnsmasq is a versatile and highly configurable tool that is simple to use.
DNSMasq_DHCP is a feature of DNSMasq that provides a combined server to serve both DNS (Domain Name System) and DHCP (Dynamic Host Configuration Protocol) requests. It is a fast and lightweight DHCP server with support for both IPv4 and IPv6, and can be used to serve IP addresses to hosts on a LAN. DNSMasq_DHCP also offers features such as DNS and DHCP performance tuning, DHCP address range management, and support for multiple DNS domains.&lt;/p></description></item><item><title>Dnsmasq DHCP Monitoring</title><link>https://www.netdata.cloud/monitoring-101/dnsmasq_dhcp-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/dnsmasq_dhcp-monitoring/</guid><description>&lt;h2 id="dnsmasq-dhcp-monitoring">Dnsmasq DHCP Monitoring&lt;/h2>
&lt;p>Dnsmasq is a lightweight, easy-to-configure, and highly efficient DHCP server that serves small- to medium-sized networks. Monitoring Dnsmasq DHCP effectively ensures the reliability and performance of your network configuration services.&lt;/p>
&lt;h3 id="what-is-dnsmasq-dhcp">What Is Dnsmasq DHCP?&lt;/h3>
&lt;p>Dnsmasq provides network infrastructure for small networks: DNS, DHCP, router advertisement, and network boot. These server functionalities are crucial for network management and troubleshooting.&lt;/p>
&lt;h3 id="monitoring-dnsmasq-dhcp-with-netdata">Monitoring Dnsmasq DHCP With Netdata&lt;/h3>
&lt;p>Netdata is the perfect solution for monitoring Dnsmasq DHCP as it provides real-time monitoring, enabling swift identification and resolution of potential network issues. With its turnkey installation, Netdata automatically detects and visualizes metrics needed for effective Dnsmasq DHCP monitoring.&lt;/p></description></item><item><title>Dnsmasq Monitoring</title><link>https://www.netdata.cloud/monitoring-101/dnsmasq-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/dnsmasq-monitoring/</guid><description>&lt;h2 id="dnsmasq-monitoring">Dnsmasq Monitoring&lt;/h2>
&lt;h3 id="what-is-dnsmasq">What Is Dnsmasq?&lt;/h3>
&lt;p>Dnsmasq is a lightweight and robust DNS forwarder and DHCP server, perfect for networking tools used in small to medium scale networks. It&amp;rsquo;s designed to be easy to configure while maintaining a wide variety of functionalities. Dnsmasq can cache DNS queries for improved performance of local networks. You can learn more about Dnsmasq on &lt;a href="https://thekelleys.org.uk/dnsmasq/doc.html">its official documentation&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-dnsmasq-with-netdata">Monitoring Dnsmasq With Netdata&lt;/h3>
&lt;p>Netdata provides an unparalleled real-time monitoring solution for Dnsmasq. By leveraging Netdata&amp;rsquo;s &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/dnsmasq/?utm_source=website&amp;amp;utm_content=monitoring101">Dnsmasq monitoring tool&lt;/a>, administrators and developers gain seamless insights into the DNS and DHCP services. Netdata&amp;rsquo;s easy-to-navigate UI, with real-time visualizations, and &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">live demo&lt;/a>, brings valuable metrics to the forefront, aiding proactive diagnosis and trend recognition.&lt;/p></description></item><item><title>Docker</title><link>https://www.netdata.cloud/integrations/all/docker/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/all/docker/</guid><description/></item><item><title>Docker</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/docker/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/docker/</guid><description/></item><item><title>Docker</title><link>https://www.netdata.cloud/integrations/deploy/docker-kubernetes/docker/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/docker-kubernetes/docker/</guid><description/></item><item><title>Docker commands hang: docker ps, inspect, and exec freezes</title><link>https://www.netdata.cloud/guides/docker/docker-commands-hang/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-commands-hang/</guid><description>&lt;h1 id="docker-commands-hang-docker-ps-inspect--exec-freezes">Docker Commands Hang: docker ps, inspect &amp;amp; exec Freezes&lt;/h1>
&lt;p>&lt;code>docker ps&lt;/code>, &lt;code>docker inspect&lt;/code>, and &lt;code>docker exec&lt;/code> hang while containers continue serving traffic. This is a Docker daemon hang: the management plane is dead while the data plane survives. Standard process monitors show &lt;code>dockerd&lt;/code> as alive, yet you cannot manage, inspect, or evacuate workloads.&lt;/p>
&lt;p>This guide covers how to distinguish a hang from a crash, identify whether the root cause is storage, a stuck shim, a plugin, or an internal deadlock, and recover without unnecessary host reboots or container kills.&lt;/p></description></item><item><title>Docker container cannot connect to another container</title><link>https://www.netdata.cloud/guides/docker/docker-container-cannot-connect-to-container/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-container-cannot-connect-to-container/</guid><description>&lt;h1 id="docker-container-cannot-connect-to-another-container">Docker Container Cannot Connect To Another Container&lt;/h1>
&lt;p>Connection timeouts or &amp;ldquo;connection refused&amp;rdquo; between containers on the same host point to a failure in one of four layers: the bridge network, Docker&amp;rsquo;s embedded DNS, iptables, or the application bind address. Inter-container networking relies on Linux bridges, veth pairs, iptables rules, and an embedded DNS resolver at 127.0.0.11. This guide isolates the faulty layer and fixes it.&lt;/p>
&lt;h2 id="what-this-means">What This Means&lt;/h2>
&lt;p>&lt;a href="https://www.netdata.cloud/guides/docker/">Docker&lt;/a> isolates each container in its own network namespace. Containers on the same custom bridge communicate through a Linux bridge and resolve each other by name via Docker&amp;rsquo;s embedded DNS. The default &lt;code>bridge&lt;/code> network has no embedded DNS. If containers are on different networks, if the target application binds to 127.0.0.1, or if iptables rules were wiped by an external firewall reload, traffic stops. The symptom appears at the application layer, but the root cause can be Layer 2, Layer 3, or the application itself.&lt;/p></description></item><item><title>Docker container cannot connect to the internet: diagnosis and fixes</title><link>https://www.netdata.cloud/guides/docker/docker-container-cannot-connect-to-internet/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-container-cannot-connect-to-internet/</guid><description>&lt;h1 id="docker-container-cannot-connect-to-the-internet">Docker container cannot connect to the internet&lt;/h1>
&lt;p>The host may be online while Docker&amp;rsquo;s bridge, iptables rules, DNS proxy, or connection tracking table is in a broken state. Applications fail with connection timeouts, package managers stall, and external health checks return unhealthy. Isolate whether the failure is DNS, routing, packet filtering, or daemon state before fixing it.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>Outbound connectivity from a container traverses the container&amp;rsquo;s network namespace, a veth pair attached to a bridge (docker0 or user-defined), iptables NAT and filter rules managed by Docker&amp;rsquo;s libnetwork, the host&amp;rsquo;s routing table, and the upstream physical interface. On user-defined networks, Docker&amp;rsquo;s embedded DNS resolver at 127.0.0.11 proxies queries to the host&amp;rsquo;s configured resolvers. On the default bridge network, there is no embedded DNS; containers inherit the host&amp;rsquo;s /etc/resolv.conf directly. A failure at any layer produces the same symptom: requests time out.&lt;/p></description></item><item><title>Docker container exits immediately: how to diagnose it</title><link>https://www.netdata.cloud/guides/docker/docker-container-exits-immediately/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-container-exits-immediately/</guid><description>&lt;h1 id="docker-container-exits-immediately-how-to-diagnose-it">Docker container exits immediately: how to diagnose it&lt;/h1>
&lt;p>A container that exits immediately prints its ID during &lt;code>docker run&lt;/code>, appears briefly in &lt;code>docker ps&lt;/code>, then vanishes. &lt;code>docker ps -a&lt;/code> lists it as &lt;code>Exited&lt;/code> with an uptime measured in seconds. Unlike a restart loop, it stays stopped.&lt;/p>
&lt;p>A Docker container is a Linux process wrapper, not a virtual machine. Its lifecycle is bound to PID 1 in the container&amp;rsquo;s namespace. When that process completes, crashes, or is killed, the container exits. An immediate exit means the process finished its work, failed to start, or was terminated during initialization. This guide gives you a diagnostic flow to find out why PID 1 terminated, interpret the exit code, and fix the root cause.&lt;/p></description></item><item><title>Docker container high CPU usage: causes and fixes</title><link>https://www.netdata.cloud/guides/docker/docker-container-high-cpu-usage/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-container-high-cpu-usage/</guid><description>&lt;h1 id="docker-container-high-cpu-usage-causes-and-fixes">Docker container high CPU usage: causes and fixes&lt;/h1>
&lt;p>A container alert for high CPU is easy to generate but hard to interpret. Docker reports CPU as a percentage of total host cores, so a value of 340% on an 8-core machine simply means the container is using 3.4 cores. That might be expected for a multi-threaded workload, or it might signal a runaway process, CFS throttling, or a cryptominer. This guide shows how to distinguish legitimate compute from pathological behavior, and how to fix the root cause without restarting the container blindly.&lt;/p></description></item><item><title>Docker container high memory usage: how to diagnose it</title><link>https://www.netdata.cloud/guides/docker/docker-container-high-memory-usage/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-container-high-memory-usage/</guid><description>&lt;h1 id="docker-container-high-memory-usage-how-to-diagnose-it">Docker Container High Memory Usage: How To Diagnose It&lt;/h1>
&lt;p>Your container is sitting at 90% of its memory limit but has not been OOMKilled. Or it is being killed repeatedly and you cannot tell whether the limit is too low or the application is leaking. docker stats shows a single percentage, but that number mixes reclaimable page cache with anonymous memory that the kernel cannot reclaim. To diagnose this correctly, you need to decompose cgroup memory.stat, map it to your runtime&amp;rsquo;s actual allocations, and decide whether the problem is cache pressure, a runtime mismatch, or a true leak.&lt;/p></description></item><item><title>Docker container keeps restarting: causes, checks, and fixes</title><link>https://www.netdata.cloud/guides/docker/docker-container-keeps-restarting/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-container-keeps-restarting/</guid><description>&lt;h1 id="docker-container-keeps-restarting-causes-checks--fixes">Docker Container Keeps Restarting: Causes, Checks &amp;amp; Fixes&lt;/h1>
&lt;p>A container that restarts every few seconds is not self-healing. It is a crash loop that wastes CPU, floods logs, and masks the real failure. Docker&amp;rsquo;s restart policy can hide whether the application is OOM-killed, segfaulting, or waiting for a dependency that never arrives. This guide shows how to read the signals, map exit codes to causes, and stop the loop before it degrades the host.&lt;/p></description></item><item><title>Docker container memory leak: how to find one and prove it</title><link>https://www.netdata.cloud/guides/docker/docker-container-memory-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-container-memory-leak/</guid><description>&lt;h1 id="docker-container-memory-leak-how-to-find-one-and-prove-it">Docker Container Memory Leak: How To Find One And Prove It&lt;/h1>
&lt;p>Memory that only ever climbs is easy to spot. The harder problem is proving whether the growth is a leak, unbounded caching, or a limit set below the working set. During an incident, operators need to decide in minutes whether to page an on-call developer or bump a cgroup limit. This guide shows how to use cgroup memory.stat, process-level RSS, and container restart patterns to build a defensible diagnosis. You will be able to separate anonymous memory growth from &lt;a href="https://www.netdata.cloud/guides/docker/docker-memory-usage-explained/">reclaimable cache&lt;/a>, identify whether the leak lives in application heap or runtime overhead, and present evidence that justifies either a code fix or a capacity change.&lt;/p></description></item><item><title>Docker container running but unhealthy: how to diagnose health check failures</title><link>https://www.netdata.cloud/guides/docker/docker-container-running-but-unhealthy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-container-running-but-unhealthy/</guid><description>&lt;h1 id="docker-container-running-but-unhealthy-how-to-diagnose-health-check-failures">Docker Container Running But Unhealthy: How To Diagnose Health Check Failures&lt;/h1>
&lt;p>The container is up. &lt;code>docker ps&lt;/code> shows it running. But the status is &lt;code>(unhealthy)&lt;/code> and your load balancer or orchestrator has stopped sending traffic. This means Docker&amp;rsquo;s health check command is returning a non-zero exit code, or timing out, while the container&amp;rsquo;s main process stays alive. The container is running, but Docker does not consider it ready.&lt;/p>
&lt;p>An unhealthy state is not a crash. In Swarm, it triggers replacement. In Compose, &lt;code>depends_on&lt;/code> with &lt;code>condition: service_healthy&lt;/code> blocks downstream services. Even on a single host, an unhealthy mark often precedes a restart loop that buries the real error in noise. You need to distinguish between a broken application, a broken probe, and a broken runtime configuration.&lt;/p></description></item><item><title>Docker Containers</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/docker-containers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/docker-containers/</guid><description/></item><item><title>Docker CPU throttling: the hidden cause of container latency</title><link>https://www.netdata.cloud/guides/docker/docker-cpu-throttling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-cpu-throttling/</guid><description>&lt;h1 id="docker-cpu-throttling-the-hidden-cause-of-container-latency">Docker CPU Throttling: The Hidden Cause Of Container Latency&lt;/h1>
&lt;p>Your application latency just spiked. p99 response times doubled or tripled. CPU dashboards show the container at 40% utilization. Memory is fine. Network is quiet. You restart the container, redeploy, or blame the code, but the pattern repeats.&lt;/p>
&lt;p>The culprit is often CPU throttling. &lt;a href="https://www.netdata.cloud/guides/docker/">Docker&lt;/a> uses Linux CFS bandwidth control to enforce CPU limits in discrete 100ms periods. A container can exhaust its quota in a burst, spend the rest of each period throttled by the kernel, and still report a modest average CPU over a longer window. This guide shows how to confirm throttling from cgroup metrics, calculate its severity, and fix it without guessing.&lt;/p></description></item><item><title>Docker daemon not responding: how to troubleshoot a hung dockerd</title><link>https://www.netdata.cloud/guides/docker/docker-daemon-not-responding/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-daemon-not-responding/</guid><description>&lt;h1 id="docker-daemon-not-responding-how-to-troubleshoot-a-hung-dockerd">Docker Daemon Not Responding: How To Troubleshoot A Hung Dockerd&lt;/h1>
&lt;p>A hung Docker daemon is one of the more disorienting production failures you can face. The process is still running, your containers are still serving traffic, but every &lt;code>docker&lt;/code> command hangs. You cannot inspect, stop, or create containers. Deployments stall. Automation times out.&lt;/p>
&lt;p>This guide covers the diagnostic ladder from socket probe to storage driver investigation, explains when to wait versus when to restart, and describes what to avoid when you are not sure what is wrong.&lt;/p></description></item><item><title>Docker disk space full: how to troubleshoot /var/lib/docker</title><link>https://www.netdata.cloud/guides/docker/docker-disk-space-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-disk-space-full/</guid><description>&lt;h1 id="docker-disk-space-full-how-to-troubleshoot-varlibdocker">Docker Disk Space Full: How To Troubleshoot /var/lib/docker&lt;/h1>
&lt;p>You notice deployments failing with &amp;ldquo;no space left on device,&amp;rdquo; image pulls hanging, or the Docker daemon becoming sluggish. On a &lt;a href="https://www.netdata.cloud/guides/docker/">Docker&lt;/a> host, everything lives under /var/lib/docker: image layers, container writable layers and logs, named volumes, and build cache. When this filesystem fills, the failure is cascading and abrupt. New containers cannot start, running containers may fail on writes, and daemon operations deadlock.&lt;/p></description></item><item><title>Docker DNS not working inside containers</title><link>https://www.netdata.cloud/guides/docker/docker-dns-not-working/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-dns-not-working/</guid><description>&lt;h1 id="docker-dns-not-working-inside-containers">Docker DNS not working inside containers&lt;/h1>
&lt;p>Your application logs show connection timeouts. Health checks against dependency names are failing. &lt;code>curl&lt;/code> from inside a container returns &amp;ldquo;Could not resolve host&amp;rdquo; while the same name resolves fine on the host. DNS inside Docker is not a simple passthrough to the host resolver. It is a stack of namespace-specific forwarders, embedded resolvers, and inherited configuration that breaks in specific, repeatable ways.&lt;/p>
&lt;p>This guide will help you determine whether the failure is a missing embedded DNS, a poisoned resolv.conf, an upstream forwarding issue, or a version regression. You will be able to distinguish between inter-container name resolution failures and external lookup failures, identify the root cause with safe read-only checks, and apply the correct fix without guessing.&lt;/p></description></item><item><title>Docker Engine</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/docker-engine/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/docker-engine/</guid><description/></item><item><title>Docker exit code 1: application errors and how to find them</title><link>https://www.netdata.cloud/guides/docker/docker-exit-code-1/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-exit-code-1/</guid><description>&lt;h1 id="docker-exit-code-1-application-errors-and-how-to-find-them">Docker Exit Code 1: Application Errors And How To Find Them&lt;/h1>
&lt;p>A container exiting immediately with code 1 is a common production incident. Unlike exit code 137 (OOM killer) or code 125 (Docker daemon error), code 1 means the application inside the container called &lt;code>exit(1)&lt;/code>. &lt;a href="https://www.netdata.cloud/guides/docker/">Docker&lt;/a> is only reporting what PID 1 did.&lt;/p>
&lt;p>The challenge is that code 1 is a catch-all. It can mask an unhandled JavaScript exception, a Python import error, a missing configuration file, a shell script failing under &lt;code>set -e&lt;/code>, or a Go binary that cannot reach its database. Time spent checking Docker daemon health or host memory is wasted when the real failure is an application-level error written to stdout or stderr.&lt;/p></description></item><item><title>Docker exit code 137: OOMKilled or SIGKILL?</title><link>https://www.netdata.cloud/guides/docker/docker-exit-code-137/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-exit-code-137/</guid><description>&lt;h1 id="docker-exit-code-137-oomkilled-or-sigkill">Docker Exit Code 137: OOMKilled Or SIGKILL?&lt;/h1>
&lt;p>A container exits with code 137. Docker restarts it, or it stays down, and you need to know why. The number itself only tells you that the process received SIGKILL. What matters for your next step is whether the kernel&amp;rsquo;s cgroup OOM killer fired because the container exceeded its memory limit, or whether an external actor sent the signal. The remediation for an undersized memory limit is completely different from fixing a misconfigured stop timeout or an orchestrator sending a premature kill. This guide shows how to classify the cause in under a minute using only the Docker CLI and cgroup files.&lt;/p></description></item><item><title>Docker exit code 143: SIGTERM and graceful shutdown failures</title><link>https://www.netdata.cloud/guides/docker/docker-exit-code-143/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-exit-code-143/</guid><description>&lt;h1 id="docker-exit-code-143-sigterm-and-graceful-shutdown-failures">Docker exit code 143: SIGTERM and graceful shutdown failures&lt;/h1>
&lt;p>You are reviewing container exit codes after a deployment or node drain and see 143. If your monitoring alerts on it, you might think something failed. Exit code 143 is not an error. It is 128 plus signal 15 (SIGTERM), and it means the container&amp;rsquo;s PID 1 process received SIGTERM and exited voluntarily. This is exactly what &lt;code>docker stop&lt;/code> is designed to do.&lt;/p></description></item><item><title>Docker Hub Monitoring</title><link>https://www.netdata.cloud/monitoring-101/dockerhub-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/dockerhub-monitoring/</guid><description>&lt;h2 id="docker-hub-monitoring">Docker Hub Monitoring&lt;/h2>
&lt;h3 id="what-is-docker-hub">What Is Docker Hub?&lt;/h3>
&lt;p>Docker Hub is a cloud-based repository where Docker users and partners create, test, store, and distribute container images. It is the centralize hub for approximately 100,000 container images; providing the ability to automate workflows, build images from GitHub or Bitbucket, and distribute them to your users.&lt;/p>
&lt;h3 id="monitoring-docker-hub-with-netdata">Monitoring Docker Hub With Netdata&lt;/h3>
&lt;p>To effectively monitor Docker Hub, leverage the insightful metrics retrieved by Netdata’s robust monitoring agent—the go-to Docker Hub monitoring tool. Netdata&amp;rsquo;s Docker Hub collector keeps track of your repository statistics, providing essential visibility into its health and performance.&lt;/p></description></item><item><title>Docker Hub repository</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/docker-hub-repository/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/docker-hub-repository/</guid><description/></item><item><title>Docker image cleanup: safe pruning strategies for production hosts</title><link>https://www.netdata.cloud/guides/docker/docker-image-cleanup/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-image-cleanup/</guid><description>&lt;h1 id="docker-image-cleanup-safe-pruning-strategies-for-production-hosts">Docker image cleanup: safe pruning strategies for production hosts&lt;/h1>
&lt;p>When &lt;code>df -h /var/lib/docker&lt;/code> shows 87% utilization, the urge to run &lt;code>docker system prune -a&lt;/code> is strong. On production hosts, that is a mistake. Cleanup is not about finding the single command that reclaims the most space. It is about knowing exactly what each flag deletes, what it leaves behind, and which filters prevent a 3 a.m. image re-pull because a base layer was removed.&lt;/p></description></item><item><title>Docker image pull failures: registry, network, and auth diagnosis</title><link>https://www.netdata.cloud/guides/docker/docker-image-pull-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-image-pull-failures/</guid><description>&lt;h1 id="docker-image-pull-failures-registry-network--auth-diagnosis">Docker Image Pull Failures: Registry, Network &amp;amp; Auth Diagnosis&lt;/h1>
&lt;h2 id="what-this-means">What This Means&lt;/h2>
&lt;p>When you run &lt;code>docker pull&lt;/code>, the daemon negotiates a TLS connection to the registry, authenticates if required, resolves the manifest for the requested tag and architecture, then downloads missing layers. If any step fails, the pull aborts.&lt;/p>
&lt;p>Registry errors surface as HTTP 429 or 401/403 responses. Network errors appear as timeouts, connection resets, or TLS handshake failures. Local problems such as a full disk or a hung daemon can also abort a pull even when the registry is healthy. In orchestrated environments, a single node&amp;rsquo;s pull failure can trigger the scheduler to retry on other nodes, turning a localized auth error into a cluster-wide rate limit storm. Distinguishing these layers quickly is the difference between a five-minute fix and a prolonged outage.&lt;/p></description></item><item><title>Docker JVM memory tuning: heap, off-heap, and the cgroup mismatch</title><link>https://www.netdata.cloud/guides/docker/docker-jvm-memory-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-jvm-memory-tuning/</guid><description>&lt;h1 id="docker-jvm-memory-tuning-heap-off-heap-and-the-cgroup-mismatch">Docker JVM Memory Tuning: Heap, Off-Heap, And The cgroup Mismatch&lt;/h1>
&lt;p>Your Java container is &lt;a href="https://www.netdata.cloud/guides/docker/docker-oomkilled/">OOMKilled&lt;/a> at 02:00. Heap usage is 60%. Docker reports exit code 137 and &lt;code>OOMKilled: true&lt;/code>. The JVM never threw an &lt;code>OutOfMemoryError&lt;/code>.&lt;/p>
&lt;p>In a container, the kernel enforces memory limits through cgroups, but the JVM heap is only one component of process RSS. Off-heap memory, metaspace, thread stacks, direct byte buffers, and GC overhead all count against the same cgroup limit. When total RSS crosses that limit, the kernel kills the container without a JVM-level error. In some JDK and kernel combinations, the JVM fails to detect the cgroup limit entirely and sizes the heap against host RAM, which guarantees an OOM kill.&lt;/p></description></item><item><title>Docker log rotation: preventing json-file logs from filling disk</title><link>https://www.netdata.cloud/guides/docker/docker-log-rotation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-log-rotation/</guid><description>&lt;h1 id="docker-log-rotation-preventing-json-file-logs-from-filling-disk">Docker Log Rotation: Preventing json-file Logs From Filling Disk&lt;/h1>
&lt;p>Docker&amp;rsquo;s default &lt;code>json-file&lt;/code> log driver appends every line of container stdout and stderr to a JSON file on the host under &lt;code>/var/lib/docker/containers/&amp;lt;id&amp;gt;/&lt;/code>. Without size limits, that file grows monotonically. A single verbose container can consume tens of gigabytes, and because the driver is the default, this often happens silently until &lt;code>/var/lib/docker&lt;/code> fills. At that point image pulls fail, container creates are rejected, and the daemon may hang on storage operations. This guide covers how to cap log files with &lt;code>daemon.json&lt;/code> and per-container &lt;code>log-opt&lt;/code> overrides, verify the caps are working, and choose a different driver when json-file is not appropriate.&lt;/p></description></item><item><title>Docker logs taking too much disk space: how to fix log growth</title><link>https://www.netdata.cloud/guides/docker/docker-logs-disk-space/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-logs-disk-space/</guid><description>&lt;h1 id="docker-logs-taking-too-much-disk-space-how-to-fix-log-growth">Docker Logs Taking Too Much Disk Space: How To Fix Log Growth&lt;/h1>
&lt;p>The default json-file logging driver captures every line a container writes to stdout and stderr and appends it to a file under &lt;code>/var/lib/docker/containers/&lt;/code>. There is no size limit out of the box. A container with verbose logging, a stuck retry loop, or a crash spiral can grow its log file by gigabytes per day until the filesystem is full.&lt;/p></description></item><item><title>Docker memory limits: how to set them and what happens when they hit</title><link>https://www.netdata.cloud/guides/docker/docker-memory-limits/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-memory-limits/</guid><description>&lt;h1 id="docker-memory-limits-how-to-set-them-and-what-happens-when-they-hit">Docker memory limits: how to set them and what happens when they hit&lt;/h1>
&lt;p>A container without a memory limit can consume all available RAM, force the kernel to reclaim page cache, push the system into swap, and eventually trigger a host-level OOM kill that takes down the Docker daemon or other critical processes. Setting &lt;code>--memory&lt;/code> is not enough: limits can be silently ignored, misread by runtimes, or masked by swap behavior that turns a clean failure into a slow crawl.&lt;/p></description></item><item><title>Docker memory usage explained: anonymous, file, slab, and what counts</title><link>https://www.netdata.cloud/guides/docker/docker-memory-usage-explained/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-memory-usage-explained/</guid><description>&lt;h1 id="docker-memory-usage-explained-anonymous-file-slab-and-what-counts">Docker memory usage explained: anonymous, file, slab, and what counts&lt;/h1>
&lt;p>You look at &lt;code>docker stats&lt;/code>, see a container sitting at 1.2 GB of a 1.5 GB limit, and assume it is about to explode. It might be fine. That total includes page cache the kernel will reclaim the second another process asks for memory. Meanwhile, a different container reports 600 MB with no limit set, yet its anonymous memory grows 50 MB per hour and will force a host-level OOM kill before lunch.&lt;/p></description></item><item><title>Docker Monitoring</title><link>https://www.netdata.cloud/monitoring-101/docker-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/docker-monitoring/</guid><description>&lt;h2 id="docker-monitoring">Docker Monitoring&lt;/h2>
&lt;h3 id="what-is-docker">What Is Docker?&lt;/h3>
&lt;p>Docker is an open platform that packages applications and their dependencies into lightweight, portable units called &lt;strong>containers&lt;/strong>. A container bundles everything the software needs to run — code, runtime, system tools, libraries, and settings — so it behaves the same way on a developer&amp;rsquo;s laptop, in staging, and in production. Because containers share the host operating system&amp;rsquo;s kernel instead of shipping a full guest OS, they start in milliseconds and use far fewer resources than virtual machines.&lt;/p></description></item><item><title>Docker monitoring checklist: the signals every production host needs</title><link>https://www.netdata.cloud/guides/docker/docker-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-monitoring-checklist/</guid><description>&lt;h1 id="docker-monitoring-checklist-the-signals-every-production-host-needs">Docker Monitoring Checklist: The Signals Every Production Host Needs&lt;/h1>
&lt;p>Production Docker incidents rarely look like Docker problems at first. They show up as application latency, deployment failures, or hosts that suddenly refuse to schedule containers. By the time you notice, the daemon may be hung, a log file has filled the disk, or a container has been silently throttled into unusable latency. This checklist groups the essential production signals into three priority tiers: must-have alerts that keep the host alive, should-have metrics that expose resource pressure before it becomes an outage, and nice-to-have security and internal signals for mature environments. Every signal includes where to read it from the raw cgroup filesystem or the Docker API so you can instrument hosts without guessing paths.&lt;/p></description></item><item><title>Docker OOMKilled: causes, detection, and prevention</title><link>https://www.netdata.cloud/guides/docker/docker-oomkilled/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-oomkilled/</guid><description>&lt;h1 id="docker-oomkilled-causes-detection-and-prevention">Docker OOMKilled: causes, detection, and prevention&lt;/h1>
&lt;p>A container exits with code 137 and restarts. The application loses in-memory state. Dependent services start failing. The restart loop begins. This is the OOMKilled pattern, and it is one of the most common and most misdiagnosed failure modes in Docker environments.&lt;/p>
&lt;p>This article covers how to confirm an OOM kill, distinguish it from an external SIGKILL, understand why it happened, and prevent recurrence. It also covers the JVM-in-container memory mismatch, which is responsible for a large share of OOM kills in Java workloads.&lt;/p></description></item><item><title>Docker port binding: address already in use</title><link>https://www.netdata.cloud/guides/docker/docker-port-binding-address-already-in-use/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-port-binding-address-already-in-use/</guid><description>&lt;h1 id="docker-port-binding-address-already-in-use">Docker port binding: address already in use&lt;/h1>
&lt;p>A &lt;code>docker run -p 8080:80&lt;/code> or &lt;code>docker compose up&lt;/code> fails with &lt;code>bind: address already in use&lt;/code>. The error is clear, but the owner is not. &lt;code>ss&lt;/code> may show a system service. &lt;code>docker ps&lt;/code> may show nothing, yet Docker still refuses. The port can appear free while an orphaned DNAT rule or a split firewall backend blocks the bind.&lt;/p>
&lt;p>Distinguish a genuine socket conflict from an orphaned DNAT rule, a rootless Docker regression, and a WSL2 iptables/nftables split brain. Then reclaim the port and prevent recurrence.&lt;/p></description></item><item><title>Docker published port not reachable: troubleshooting -p and EXPOSE</title><link>https://www.netdata.cloud/guides/docker/docker-published-port-not-reachable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-published-port-not-reachable/</guid><description>&lt;h1 id="docker-published-port-not-reachable-troubleshooting--p-and-expose">Docker published port not reachable: troubleshooting -p and EXPOSE&lt;/h1>
&lt;p>You mapped a port with &lt;code>-p 8080:80&lt;/code>, but &lt;code>curl&lt;/code> against the host IP returns connection refused. &lt;code>docker ps&lt;/code> shows the mapping, the container is running, and the port still appears closed.&lt;/p>
&lt;p>A published port depends on three layers: a runtime mapping rule (&lt;code>-p&lt;/code>), a host forwarding path (iptables DNAT and FORWARD policy), and an application listener inside the container bound to an interface that receives the forwarded packet. EXPOSE in a Dockerfile is metadata. It does not publish ports, create firewall rules, or set bind addresses.&lt;/p></description></item><item><title>Docker socket security: why /var/run/docker.sock is root access</title><link>https://www.netdata.cloud/guides/docker/docker-socket-security/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-socket-security/</guid><description>&lt;h1 id="docker-socket-security-why-varrundockersock-is-root-access">Docker socket security: why /var/run/docker.sock is root access&lt;/h1>
&lt;p>Mounting &lt;code>/var/run/docker.sock&lt;/code> into a container grants that container a root-equivalent privilege boundary. This is not a Docker vulnerability. The daemon runs as root, listens on a filesystem socket, and trusts any client that can write to it.&lt;/p>
&lt;p>Symptoms of abuse look like container escape or host compromise, but the container never escaped. It asked the root-owned daemon to perform privileged operations on its behalf. This guide explains the mechanism, how to audit exposure, and what constraints you can apply without rebuilding your pipeline.&lt;/p></description></item><item><title>Docker volume cleanup: finding and removing orphaned volumes</title><link>https://www.netdata.cloud/guides/docker/docker-volume-cleanup/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-volume-cleanup/</guid><description>&lt;h1 id="docker-volume-cleanup-finding-and-removing-orphaned-volumes">Docker volume cleanup: finding and removing orphaned volumes&lt;/h1>
&lt;p>Run &lt;code>docker system df&lt;/code> and the &lt;code>Local Volumes&lt;/code> line keeps growing. Run &lt;code>docker system prune&lt;/code> and you reclaim images and build cache, but the volume count barely drops. A few weeks later the disk alert fires again.&lt;/p>
&lt;p>Data persists in &lt;code>/var/lib/docker/volumes/&lt;/code> after its consumer is long gone. These volumes waste disk, complicate capacity planning, and can hide sensitive data in forgotten corners of the filesystem.&lt;/p></description></item><item><title>Domain expiration date</title><link>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/domain-expiration-date/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/domain-expiration-date/</guid><description/></item><item><title>Domain Monitoring</title><link>https://www.netdata.cloud/monitoring-101/whoisquery-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/whoisquery-monitoring/</guid><description>&lt;h2 id="domain-monitoring">Domain Monitoring&lt;/h2>
&lt;h3 id="what-is-domain-monitoring">What Is Domain Monitoring?&lt;/h3>
&lt;p>Domain monitoring involves tracking various domain-related metrics such as the time until a domain expires. This ensures that your domain remains operational without unexpected downtime due to expiry. By using tools designed for monitoring domains, businesses can avoid disruptions to their online presence.&lt;/p>
&lt;h3 id="monitoring-domain-expiration-with-netdata">Monitoring Domain Expiration With Netdata&lt;/h3>
&lt;p>Netdata offers a comprehensive domain monitoring tool through its WhoisQuery module. This tool helps in tracking domain expiration dates effectively. By integrating this into your monitoring strategy, you can maintain continuous oversight of your domains’ expiration status and take proactive measures to renew them on time.&lt;/p></description></item><item><title>Dovecot</title><link>https://www.netdata.cloud/integrations/data-collection/applications/dovecot/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/dovecot/</guid><description/></item><item><title>Dovecot Monitoring</title><link>https://www.netdata.cloud/monitoring-101/dovecot-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/dovecot-monitoring/</guid><description>&lt;h2 id="dovecot-monitoring">Dovecot Monitoring&lt;/h2>
&lt;h3 id="what-is-dovecot">What Is Dovecot?&lt;/h3>
&lt;p>Dovecot is a widely-used open-source IMAP and POP3 server for Unix-like operating systems. Its focus on security, ease of use, and performance makes it a preferred choice for many system administrators when it comes to mail server deployments. For more details on Dovecot, visit the &lt;a href="https://www.dovecot.org/">official Dovecot website&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-dovecot-with-netdata">Monitoring Dovecot With Netdata&lt;/h3>
&lt;p>Netdata provides a comprehensive $name monitoring tool designed to capture and visualize a wide array of metrics from your Dovecot server. By monitoring Dovecot with Netdata, you gain real-time insights into your server&amp;rsquo;s performance, understand traffic patterns, and detect anomalies that could indicate underlying issues.&lt;/p></description></item><item><title>Dps Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dps-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dps-inc-snmp-traps/</guid><description/></item><item><title>Dragonwave SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dragonwave-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dragonwave-snmp-traps/</guid><description/></item><item><title>Drive disappeared from the bus: sudden controller or electronics death</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-drive-disappeared-from-bus/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-drive-disappeared-from-bus/</guid><description>&lt;h1 id="drive-disappeared-from-the-bus-sudden-controller-or-electronics-death">Drive disappeared from the bus: sudden controller or electronics death&lt;/h1>
&lt;p>A device node that was present minutes or hours ago is now gone. &lt;code>lsblk&lt;/code> no longer shows it. &lt;code>smartctl&lt;/code> returns &amp;ldquo;No such device.&amp;rdquo; The drive did not warn you through SMART because the component that failed is the one that would have reported the problem. This is one of the few storage failure modes that produces zero SMART telemetry before it happens.&lt;/p></description></item><item><title>Drive temperature too high: HDD, SATA SSD, and NVMe thresholds</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-drive-overheating/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-drive-overheating/</guid><description>&lt;h1 id="drive-temperature-too-high-hdd-sata-ssd-and-nvme-thresholds">Drive temperature too high: HDD, SATA SSD, and NVMe thresholds&lt;/h1>
&lt;p>A temperature reading of 68C means very different things depending on what is in the slot. For an enterprise HDD, it is past the danger threshold and the drive is likely sustaining damage. For an NVMe SSD, it may be within normal operating range and the drive is not even throttling yet. Alerting on a single global threshold across mixed drive types produces two failure modes simultaneously: false pages on normal NVMe temperatures and missed alerts on dangerously hot HDDs.&lt;/p></description></item><item><title>Dutch Electricity Smart Meter</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/dutch-electricity-smart-meter/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/dutch-electricity-smart-meter/</guid><description/></item><item><title>Dutch Electricity Smart Meter Monitoring</title><link>https://www.netdata.cloud/monitoring-101/dutch_electricity_smart_meter-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/dutch_electricity_smart_meter-monitoring/</guid><description>&lt;h2 id="dutch-electricity-smart-meter-monitoring">Dutch Electricity Smart Meter Monitoring&lt;/h2>
&lt;h3 id="what-is-dutch-electricity-smart-meter">What Is Dutch Electricity Smart Meter?&lt;/h3>
&lt;p>Dutch Electricity Smart Meters are advanced IoT devices used for tracking energy consumption in households. They provide granular data through the P1 port, which is beneficial for energy management and efficient monitoring. To gather insights, these meters can be hooked up with various monitoring tools, such as the Prometheus P1 Exporter.&lt;/p>
&lt;h3 id="monitoring-dutch-electricity-smart-meters-with-netdata">Monitoring Dutch Electricity Smart Meters With Netdata&lt;/h3>
&lt;p>To monitor Dutch Electricity Smart Meters, Netdata utilizes an openmetrics (Prometheus) exporter. This integration enables Netdata to ingest data from any Prometheus exporter. With this setup, users benefit from Netdata’s automated dashboards, alerts, and other features without the need for a Prometheus server or Grafana. This means you can have a comprehensive and real-time view of your smart meter data effortlessly.&lt;/p></description></item><item><title>Dynatech Communications SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dynatech-communications-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/dynatech-communications-snmp-traps/</guid><description/></item><item><title>Dynatrace</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/dynatrace/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/dynatrace/</guid><description/></item><item><title>Dynatrace</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/dynatrace/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/dynatrace/</guid><description/></item><item><title>Dynatrace Monitoring</title><link>https://www.netdata.cloud/monitoring-101/dynatrace-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/dynatrace-monitoring/</guid><description>&lt;h2 id="dynatrace-monitoring">Dynatrace Monitoring&lt;/h2>
&lt;h3 id="what-is-dynatrace">What Is Dynatrace?&lt;/h3>
&lt;p>Dynatrace is a leading application performance management (APM) solution that provides comprehensive visibility into your applications, infrastructure, and user experiences. It is designed to help organizations monitor, optimize, and scale their environments effectively. Whether it&amp;rsquo;s cloud-native applications or more traditional server-client architectures, Dynatrace offers automated and intelligent observability across the stack.&lt;/p>
&lt;h3 id="monitoring-dynatrace-with-netdata">Monitoring Dynatrace With Netdata&lt;/h3>
&lt;p>Monitoring Dynatrace with Netdata offers seamless integration via the openmetrics (Prometheus) exporter. Netdata can ingest metrics from any Prometheus exporter, allowing it to create automated dashboards, alerts, and more without the need for a Prometheus server or Grafana. By utilizing the &lt;a href="https://github.com/Apside-TOP/dynatrace_exporter">Dynatrace Exporter&lt;/a>, you can easily monitor Dynatrace metrics in real-time with Netdata, enabling better application performance management.&lt;/p></description></item><item><title>E Dynamics Org SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/e-dynamics-org-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/e-dynamics-org-snmp-traps/</guid><description/></item><item><title>E T A Elektrotechnische Apparate GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/e-t-a-elektrotechnische-apparate-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/e-t-a-elektrotechnische-apparate-gmbh-snmp-traps/</guid><description/></item><item><title>Eastern Research Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/eastern-research-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/eastern-research-inc-snmp-traps/</guid><description/></item><item><title>Eaton Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/eaton-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/eaton-corporation-snmp-traps/</guid><description/></item><item><title>Eaton Epdu</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/eaton-epdu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/eaton-epdu/</guid><description/></item><item><title>Eaton UPS</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/eaton-ups/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/eaton-ups/</guid><description/></item><item><title>eBPF Cachestat</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-cachestat/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-cachestat/</guid><description/></item><item><title>eBPF DCstat</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-dcstat/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-dcstat/</guid><description/></item><item><title>eBPF Disk</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ebpf-disk/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ebpf-disk/</guid><description/></item><item><title>eBPF Filedescriptor</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-filedescriptor/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-filedescriptor/</guid><description/></item><item><title>eBPF Filesystem</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ebpf-filesystem/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ebpf-filesystem/</guid><description/></item><item><title>eBPF Hardirq</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-hardirq/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-hardirq/</guid><description/></item><item><title>eBPF MDflush</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ebpf-mdflush/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ebpf-mdflush/</guid><description/></item><item><title>eBPF Mount</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ebpf-mount/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ebpf-mount/</guid><description/></item><item><title>eBPF OOMkill</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-oomkill/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-oomkill/</guid><description/></item><item><title>eBPF Process</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-process/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-process/</guid><description/></item><item><title>eBPF Processes</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-processes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-processes/</guid><description/></item><item><title>eBPF SHM</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-shm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-shm/</guid><description/></item><item><title>eBPF Socket</title><link>https://www.netdata.cloud/integrations/data-collection/networking/ebpf-socket/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/ebpf-socket/</guid><description/></item><item><title>eBPF SoftIRQ</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-softirq/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-softirq/</guid><description/></item><item><title>eBPF SWAP</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-swap/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/ebpf-swap/</guid><description/></item><item><title>eBPF Sync</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ebpf-sync/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ebpf-sync/</guid><description/></item><item><title>eBPF VFS</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ebpf-vfs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ebpf-vfs/</guid><description/></item><item><title>Ecreso SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ecreso-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ecreso-snmp-traps/</guid><description/></item><item><title>Edial Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/edial-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/edial-inc-snmp-traps/</guid><description/></item><item><title>Egenera Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/egenera-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/egenera-inc-snmp-traps/</guid><description/></item><item><title>Egnite GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/egnite-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/egnite-gmbh-snmp-traps/</guid><description/></item><item><title>Eicon SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/eicon-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/eicon-snmp-traps/</guid><description/></item><item><title>Ekinops Sas SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ekinops-sas-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ekinops-sas-snmp-traps/</guid><description/></item><item><title>Elasticsearch</title><link>https://www.netdata.cloud/integrations/data-collection/databases/elasticsearch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/elasticsearch/</guid><description/></item><item><title>ElasticSearch</title><link>https://www.netdata.cloud/integrations/exporters/elasticsearch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/elasticsearch/</guid><description/></item><item><title>Elasticsearch all shards failed: diagnosing search_phase_execution_exception</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-all-shards-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-all-shards-failed/</guid><description>&lt;h1 id="elasticsearch-all-shards-failed-diagnosing-search_phase_execution_exception">Elasticsearch all shards failed: diagnosing search_phase_execution_exception&lt;/h1>
&lt;p>You run a search and Elasticsearch returns &lt;code>search_phase_execution_exception&lt;/code> with reason &lt;code>all shards failed&lt;/code>. Every shard copy involved in the query returned a failure to the coordinating node. The outer error is a container; the actual root cause lives in the per-shard failure reasons inside the response body. Do not assume the cluster is down. This error fires on clusters with green health and stable nodes when a query is malformed, a mapping is incompatible, or a resource limit is breached uniformly across every target shard.&lt;/p></description></item><item><title>Elasticsearch ALLOCATION_FAILED after max retries: reroute and corrupt shard recovery</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-shard-allocation-failed-max-retries/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-shard-allocation-failed-max-retries/</guid><description>&lt;h1 id="elasticsearch-allocation_failed-after-max-retries-reroute-and-corrupt-shard-recovery">Elasticsearch ALLOCATION_FAILED after max retries: reroute and corrupt shard recovery&lt;/h1>
&lt;p>A shard that repeatedly fails allocation exhausts &lt;code>index.allocation.max_retries&lt;/code> (default 5) and becomes permanently UNASSIGNED. Elasticsearch stops automatic placement. Cluster health is RED if the shard is a primary, YELLOW if it is a replica. Indexing to the affected index is blocked until the primary is assigned.&lt;/p>
&lt;p>This state typically follows transient node restarts, disk pressure events, or translog corruption. After the fifth failed attempt, the allocator stops retrying. The shard will not move without explicit operator action: &lt;code>POST /_cluster/reroute?retry_failed=true&lt;/code> for transient failures, or forced allocation with accepted data loss when no valid copy remains.&lt;/p></description></item><item><title>Elasticsearch authentication failures: audit logs, brute force, and credential drift</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-authentication-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-authentication-failures/</guid><description>&lt;h1 id="elasticsearch-authentication-failures-audit-logs-brute-force-and-credential-drift">Elasticsearch authentication failures: audit logs, brute force, and credential drift&lt;/h1>
&lt;p>Elasticsearch does not expose authentication failure counts through &lt;code>_nodes/stats&lt;/code> or &lt;code>_cluster/health&lt;/code>. The security audit log is the only structured source for &lt;code>authentication_failed&lt;/code>, &lt;code>access_denied&lt;/code>, and &lt;code>run_as_denied&lt;/code> events. Without it, brute force attempts, credential stuffing, and expiring service tokens are invisible until they cause an outage.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>When &lt;code>xpack.security.audit.enabled&lt;/code> is &lt;code>true&lt;/code>, each node writes security events to a local audit log file. The events that matter for auth issues are:&lt;/p></description></item><item><title>Elasticsearch CircuitBreakingException: [parent] Data too large - causes and fixes</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-circuitbreakingexception-parent-data-too-large/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-circuitbreakingexception-parent-data-too-large/</guid><description>&lt;h1 id="elasticsearch-circuitbreakingexception-parent-data-too-large---causes-and-fixes">Elasticsearch CircuitBreakingException: [parent] Data too large - causes and fixes&lt;/h1>
&lt;p>When a search or indexing request returns HTTP 429 with &lt;code>CircuitBreakingException: [parent] Data too large, data for [&amp;lt;http_request&amp;gt;] would be [X], which is larger than the limit of [Y]&lt;/code>, the parent circuit breaker has rejected the operation. This is Elasticsearch protecting the JVM from an out-of-memory kill, not a client-side rate limit.&lt;/p>
&lt;p>Since version 7.0, the parent breaker tracks real memory usage by default. It can trip even when individual child breakers are within limits. The node is under genuine heap pressure. Determine quickly whether the cause is a single abusive query or structural memory exhaustion.&lt;/p></description></item><item><title>Elasticsearch cluster health red: unassigned primaries and how to recover</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-cluster-health-red/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-cluster-health-red/</guid><description>&lt;h1 id="elasticsearch-cluster-health-red-unassigned-primaries-and-how-to-recover">Elasticsearch cluster health red: unassigned primaries and how to recover&lt;/h1>
&lt;p>Cluster health &lt;code>red&lt;/code> means at least one primary shard is unassigned. Queries against affected indices return partial results or fail; writes are blocked. &lt;code>yellow&lt;/code> only signals missing replicas, but &lt;code>red&lt;/code> signals active data unavailability.&lt;/p>
&lt;p>Cluster health is a lagging indicator. A red status sustained longer than two minutes after the cluster has formed is a real fault; a brief flash during startup is normal. By the time the status turns red, a node has likely departed, a disk has crossed a watermark, or a shard copy has been rejected as corrupt.&lt;/p></description></item><item><title>Elasticsearch cluster health yellow: unassigned replicas vs real allocation blocks</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-cluster-health-yellow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-cluster-health-yellow/</guid><description>&lt;h1 id="elasticsearch-cluster-health-yellow-unassigned-replicas-vs-real-allocation-blocks">Elasticsearch cluster health yellow: unassigned replicas vs real allocation blocks&lt;/h1>
&lt;p>&lt;code>GET /_cluster/health&lt;/code> returning &lt;code>status: yellow&lt;/code> means all primaries are assigned but at least one replica is not. Some clusters are yellow by design. Others are yellow because a disk cascade, stuck allocator, or failed shard is blocking recovery. Benign yellow resolves. Structural yellow persists. If your cluster has been yellow for more than thirty minutes, something is actively blocking allocation. Teams that treat yellow as cosmetic often miss the transition from transient recovery to a real incident, leaving them one node failure away from red.&lt;/p></description></item><item><title>Elasticsearch cluster state too large: field count, index count, and per-node heap</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-cluster-state-too-large/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-cluster-state-too-large/</guid><description>&lt;h1 id="elasticsearch-cluster-state-too-large-field-count-index-count-and-per-node-heap">Elasticsearch cluster state too large: field count, index count, and per-node heap&lt;/h1>
&lt;p>Every node holds a copy of the cluster state in heap. When it grows large, the cost is paid everywhere: 200 MB of state consumes 200 MB on every node, and the elected master burns additional CPU and heap serializing and publishing updates. Symptoms show up indirectly: the master feels sluggish, pending tasks queue for minutes, heap pressure climbs on nodes that should be idle, and master elections stall indexing and shard allocation. The usual drivers are too many indices from per-minute or per-hour time-series patterns; a mapping explosion from uncontrolled dynamic fields; excessive aliases; or churn from frequent template and setting changes. Raw size measured via &lt;code>/_cluster/state&lt;/code> is a rough proxy. The indicators that matter are field-count growth, cluster state version churn, pending-task age, and master node heap pressure.&lt;/p></description></item><item><title>Elasticsearch cluster_block_exception: blocked by, the read-only blocks explained</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-cluster-block-exception/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-cluster-block-exception/</guid><description>&lt;h1 id="elasticsearch-cluster_block_exception-blocked-by-the-read-only-blocks-explained">Elasticsearch cluster_block_exception: blocked by, the read-only blocks explained&lt;/h1>
&lt;p>Elasticsearch returns &lt;code>cluster_block_exception&lt;/code> when a write or metadata operation hits an active index or cluster-level block. The &lt;code>blocked by&lt;/code> array contains a &lt;code>FORBIDDEN&lt;/code> or &lt;code>TOO_MANY_REQUESTS&lt;/code> string, a numeric code, and an API label. Writes stop. Depending on the code, deletes and metadata changes may also fail.&lt;/p>
&lt;p>This article maps the four numeric codes most common in production, explains how to confirm which block is active, and gives the exact commands to clear it. Most incidents involve the &lt;code>/12&lt;/code> flood-stage disk watermark, but manual index blocks and cluster-level overrides produce the same exception and require different fixes.&lt;/p></description></item><item><title>Elasticsearch coordinating node overload: aggregation merge, heap spikes, and 429s</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-coordinating-node-overload/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-coordinating-node-overload/</guid><description>&lt;h1 id="elasticsearch-coordinating-node-overload-aggregation-merge-heap-spikes-and-429s">Elasticsearch coordinating node overload: aggregation merge, heap spikes, and 429s&lt;/h1>
&lt;p>HTTP 429 or 503 responses appear on search requests while data nodes look healthy. Heap spikes on one node while others stay flat. The slow log shows heavy aggregation queries. That node is the coordinator, and it is running out of heap during the reduce phase.&lt;/p>
&lt;p>Every node can act as a coordinating node. For each search, the coordinator broadcasts the query to relevant shards, collects partial results, and merges them. Aggregations compute locally per shard and reduce in memory on the coordinator. High-cardinality terms aggregations, deep pagination, or large fetch sizes force the coordinator to hold massive intermediate structures in heap. If the estimate exceeds the circuit breaker limit, the request is rejected. If the breaker is too slow, the node may suffer long GC pauses or disconnect from the cluster.&lt;/p></description></item><item><title>Elasticsearch CPU saturation: search, merges, GC, and hot-spotting</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-cpu-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-cpu-saturation/</guid><description>&lt;h1 id="elasticsearch-cpu-saturation-search-merges-gc-and-hot-spotting">Elasticsearch CPU saturation: search, merges, GC, and hot-spotting&lt;/h1>
&lt;p>Search latency climbs, bulk indexing slows, and clients see &lt;code>EsRejectedExecutionException&lt;/code> or HTTP 429. On data nodes, CPU is pinned above 90% and load average rises. Elasticsearch burns CPU in query execution, text analysis, segment merges, and JVM garbage collection. If saturation persists, &lt;code>search&lt;/code> and &lt;code>write&lt;/code> thread pool queues grow until the cluster rejects work.&lt;/p>
&lt;p>The first trap is assuming all CPU belongs to Elasticsearch. &lt;code>os.cpu.percent&lt;/code> includes every process on the host, co-located containers, and kernel work. In containers this diverges sharply from &lt;code>process.cpu.percent&lt;/code>, which tracks only the Elasticsearch process. A node may show 89% OS CPU and 67% process CPU, while the container runtime reports a third figure. Attribute CPU correctly before tuning thread pools or adding nodes.&lt;/p></description></item><item><title>Elasticsearch disk full: emergency recovery and freeing space safely</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-disk-full-emergency-recovery/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-disk-full-emergency-recovery/</guid><description>&lt;h1 id="elasticsearch-disk-full-emergency-recovery-and-freeing-space-safely">Elasticsearch disk full: emergency recovery and freeing space safely&lt;/h1>
&lt;p>Writes fail with &lt;code>TOO_MANY_REQUESTS/12/index read-only / allow delete (api)&lt;/code>. Cluster health is red or yellow and data nodes are pinned above 90 percent disk. Indexers buffer or drop data while the cluster attempts shard relocations onto already-full disks. Recover without corrupting metadata or amplifying pressure.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>Elasticsearch uses three disk watermarks. Low (85 percent) stops new shard allocation. High (90 percent) starts relocating shards off the node. Flood stage (95 percent) forces every index with a shard on that node into &lt;code>index.blocks.read_only_allow_delete: true&lt;/code>, which stops writes.&lt;/p></description></item><item><title>Elasticsearch disk I/O saturation: merges, fsync, and page-cache starvation</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-disk-io-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-disk-io-saturation/</guid><description>&lt;h1 id="elasticsearch-disk-io-saturation-merges-fsync-and-page-cache-starvation">Elasticsearch disk I/O saturation: merges, fsync, and page-cache starvation&lt;/h1>
&lt;p>When Elasticsearch data nodes show climbing I/O wait while indexing and search latency rise, but CPU is not the bottleneck, the cluster stays green while throughput falls and thread pool queues grow. This pattern usually traces to one of three disk pressures: background segment merges rewriting data faster than storage can absorb, translog fsync overhead from durability guarantees, or OS page-cache eviction forcing every search to read from disk. This guide shows how to tell them apart, confirm the bottleneck with safe read-only checks, and relieve pressure.&lt;/p></description></item><item><title>Elasticsearch disk watermark cascade: from low watermark to cluster-wide read-only</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-disk-watermark-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-disk-watermark-cascade/</guid><description>&lt;h1 id="elasticsearch-disk-watermark-cascade-from-low-watermark-to-cluster-wide-read-only">Elasticsearch disk watermark cascade: from low watermark to cluster-wide read-only&lt;/h1>
&lt;p>Writes fail with &lt;code>cluster_block_exception&lt;/code> or &lt;code>FORBIDDEN/12/index read-only / allow delete (api)&lt;/code>. Kibana becomes unreachable; Logstash and Beats buffer or drop data. The cluster did not fail at once. It crossed a sequence of thresholds that turned single-node disk pressure into a cluster-wide write outage. With homogeneous disk sizes, every data node likely hit the thresholds within minutes, leaving no relocation target and no relief valve.&lt;/p></description></item><item><title>Elasticsearch disk watermark tuning: thresholds, max_headroom, and multiple data paths</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-watermark-tuning-and-max-headroom/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-watermark-tuning-and-max-headroom/</guid><description>&lt;h1 id="elasticsearch-disk-watermark-tuning-thresholds-max_headroom-and-multiple-data-paths">Elasticsearch disk watermark tuning: thresholds, max_headroom, and multiple data paths&lt;/h1>
&lt;p>Elasticsearch uses disk watermarks to decide when a node is too full to accept shards. The defaults (85%, 90%, 95%) were designed for small disks. On modern nodes with multi-terabyte volumes, those percentages leave hundreds of gigabytes free while still blocking allocation and triggering expensive rebalancing. Elasticsearch 8.x introduced &lt;code>max_headroom&lt;/code> to cap the free-space requirement on large disks, but the interaction between percentages, absolute bytes, and &lt;code>max_headroom&lt;/code> is not obvious. If you run multiple data paths, watermarks apply per path rather than per node, so a partially full disk can cause node-wide allocation restrictions. This article explains how the allocator evaluates disk space, how to tune thresholds without causing relocation storms, and why multiple data paths complicate the picture.&lt;/p></description></item><item><title>Elasticsearch document indexing failures: index_failed, bulk item errors, and version conflicts</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-document-indexing-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-document-indexing-failures/</guid><description>&lt;h1 id="elasticsearch-document-indexing-failures-index_failed-bulk-item-errors-and-version-conflicts">Elasticsearch document indexing failures: index_failed, bulk item errors, and version conflicts&lt;/h1>
&lt;p>A bulk request can return HTTP 200 while rejecting individual documents inside it. &lt;code>indices.indexing.index_failed&lt;/code> climbs, but pipelines that only check HTTP status miss the rejections and documents disappear. This guide covers three failure classes: mapper parsing and type conflicts, per-item bulk errors hidden in HTTP 200 responses, and &lt;code>version_conflict_engine_exception&lt;/code> under concurrent updates. It also distinguishes these from node-level write rejections and circuit breaker trips, which produce different symptoms and require different fixes.&lt;/p></description></item><item><title>Elasticsearch EsRejectedExecutionException: write thread pool rejections and HTTP 429</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-esrejectedexecutionexception-write-queue/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-esrejectedexecutionexception-write-queue/</guid><description>&lt;h1 id="elasticsearch-esrejectedexecutionexception-write-thread-pool-rejections-and-http-429">Elasticsearch EsRejectedExecutionException: write thread pool rejections and HTTP 429&lt;/h1>
&lt;p>HTTP 429 responses from Elasticsearch, or stack traces containing &lt;code>EsRejectedExecutionException&lt;/code>, mean the &lt;code>write&lt;/code> thread pool queue is full. The &lt;code>write&lt;/code> thread pool (named &lt;code>bulk&lt;/code> before Elasticsearch 6.3) executes indexing, bulk, update, and delete operations on each data node using a bounded queue. When all threads are busy and the queue fills, the node rejects the operation. This guide covers how to confirm the diagnosis, distinguish it from other rejection paths, and fix the root cause instead of masking the symptom.&lt;/p></description></item><item><title>Elasticsearch expensive queries: leading wildcards, regex, deep pagination, and scripts</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-expensive-queries/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-expensive-queries/</guid><description>&lt;h1 id="elasticsearch-expensive-queries-leading-wildcards-regex-deep-pagination-and-scripts">Elasticsearch expensive queries: leading wildcards, regex, deep pagination, and scripts&lt;/h1>
&lt;p>Search latency spikes and data-node CPU saturation usually trace to one of four query patterns: leading wildcards or unbounded regex, deep &lt;code>from+size&lt;/code> pagination, runtime Painless scripts, and deep aggregations. Elasticsearch uses a scatter-gather read path: the coordinating node broadcasts every query to one copy of every target shard. A single expensive query fans out, and the slowest shard sets overall latency. Coordinating-node heap pressure rises when it merges large intermediate result sets during the fetch phase or aggregation reduce phase.&lt;/p></description></item><item><title>Elasticsearch exposed without authentication: open clusters and snapshot exfiltration</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-exposed-without-authentication/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-exposed-without-authentication/</guid><description>&lt;h1 id="elasticsearch-exposed-without-authentication-open-clusters-and-snapshot-exfiltration">Elasticsearch exposed without authentication: open clusters and snapshot exfiltration&lt;/h1>
&lt;p>TCP/9200 is externally reachable and responds without credentials. When Elasticsearch binds to a public interface with security disabled, anyone who can reach the HTTP port can query indices, modify cluster state, and register snapshot repositories. The fastest exfiltration path is not reading documents individually. It is registering an attacker-controlled snapshot repository and copying entire indices out in a single background operation.&lt;/p></description></item><item><title>Elasticsearch fielddata circuit breaker tripped: text-field aggregations and the keyword fix</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-fielddata-circuit-breaker-tripped/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-fielddata-circuit-breaker-tripped/</guid><description>&lt;h1 id="elasticsearch-fielddata-circuit-breaker-tripped-text-field-aggregations-and-the-keyword-fix">Elasticsearch fielddata circuit breaker tripped: text-field aggregations and the keyword fix&lt;/h1>
&lt;p>Queries return &lt;code>CircuitBreakingException: [fielddata] Data too large...&lt;/code> and HTTP 429s while JVM heap on one or more data nodes climbs toward the breaker limit. The node rejects queries to protect itself before OOM. This almost always means a query is aggregating, sorting, or scripting against an analyzed &lt;code>text&lt;/code> field that lacks a &lt;code>keyword&lt;/code> sub-field, forcing Elasticsearch to load an expensive fielddata cache into heap.&lt;/p></description></item><item><title>Elasticsearch FORBIDDEN/12/index read-only / allow delete (api) — flood stage recovery</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-forbidden-12-index-read-only-allow-delete/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-forbidden-12-index-read-only-allow-delete/</guid><description>&lt;h1 id="elasticsearch-forbidden12index-read-only--allow-delete-api--flood-stage-recovery">Elasticsearch FORBIDDEN/12/index read-only / allow delete (api) — flood stage recovery&lt;/h1>
&lt;p>When Elasticsearch returns &lt;code>cluster_block_exception&lt;/code> with &lt;code>FORBIDDEN/12/index read-only / allow delete (api)&lt;/code>, every write, update, and index creation fails with HTTP 403. Search continues to work.&lt;/p>
&lt;p>This happens when a data node crosses the flood-stage disk watermark (95% by default). Elasticsearch auto-applies &lt;code>index.blocks.read_only_allow_delete&lt;/code> to every index with a shard on that node, preventing writes to avoid Lucene segment corruption from a full disk.&lt;/p></description></item><item><title>Elasticsearch heap pressure death spiral: GC, node removal, and the cascade</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-heap-pressure-death-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-heap-pressure-death-spiral/</guid><description>&lt;h1 id="elasticsearch-heap-pressure-death-spiral-gc-node-removal-and-the-cascade">Elasticsearch heap pressure death spiral: GC, node removal, and the cascade&lt;/h1>
&lt;p>When a node drops from &lt;code>_cat/nodes&lt;/code> and network tests pass, check its JVM GC logs. Stop-the-world pauses over 10 seconds cause the master to remove the node. Survivors then absorb recovery traffic, their heap climbs, and they begin missing fault-detection checks too. This feedback loop is the heap pressure death spiral. It masquerades as network instability because operators check connectivity while the real problem is memory saturation.&lt;/p></description></item><item><title>Elasticsearch high disk watermark [90%] exceeded: shard relocation and the cascade</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-high-disk-watermark-exceeded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-high-disk-watermark-exceeded/</guid><description>&lt;h1 id="elasticsearch-high-disk-watermark-90-exceeded-shard-relocation-and-the-cascade">Elasticsearch high disk watermark [90%] exceeded: shard relocation and the cascade&lt;/h1>
&lt;p>When Elasticsearch logs &lt;code>high disk watermark [90%] exceeded on [node] ... shards will be relocated away&lt;/code>, the allocator immediately begins moving shards off the affected node. Relocation generates disk I/O and network traffic on both source and target. If targets were already close to their own watermarks, incoming shards can push them past 90%. This is the disk watermark cascade: a self-reinforcing loop where one full node triggers relocations that make other nodes full, eventually leaving the cluster with no legal allocation target and, if flood stage is hit, read-only indices.&lt;/p></description></item><item><title>Elasticsearch ILM stuck: indices not rolling over, shrinking, or deleting</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-ilm-stuck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-ilm-stuck/</guid><description>&lt;h1 id="elasticsearch-ilm-stuck-indices-not-rolling-over-shrinking-or-deleting">Elasticsearch ILM stuck: indices not rolling over, shrinking, or deleting&lt;/h1>
&lt;p>Disk usage climbs steadily. Old indices that should have been deleted remain. Shard count grows, and the cluster approaches &lt;code>cluster.max_shards_per_node&lt;/code>. In ILM, indices are stuck in one phase for hours or days. This is the ILM stuck pattern: silent accumulation that becomes a disk watermark crisis, heap pressure, or unassigned shard storm when the cluster runs out of room.&lt;/p>
&lt;p>ILM polls every ten minutes by default. When an index cannot advance, it sits. Because the failure is gradual, it rarely pages until a secondary limit is breached. Detect the stuck state early and fix the root cause before accumulation triggers cascading failures.&lt;/p></description></item><item><title>Elasticsearch indexing pressure rejections: memory backpressure before heap failure</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-indexing-pressure-rejections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-indexing-pressure-rejections/</guid><description>&lt;h1 id="elasticsearch-indexing-pressure-rejections-memory-backpressure-before-heap-failure">Elasticsearch indexing pressure rejections: memory backpressure before heap failure&lt;/h1>
&lt;p>Bulk indexing clients report rejections while cluster health is green, disks are below the high watermark, and write thread pool queues are not saturated. Yet the nodes are pushing back. Pull &lt;code>_nodes/stats/indexing_pressure&lt;/code> and you will see climbing &lt;code>coordinating_rejections&lt;/code>, &lt;code>primary_rejections&lt;/code>, or &lt;code>replica_rejections&lt;/code>. This is the indexing pressure framework, introduced in Elasticsearch 7.9&lt;!-- TODO: verify exact version -->, enforcing memory-based backpressure. It tracks in-flight indexing bytes at the coordinating, primary, and replica stages. The default limit is 10% of the JVM heap for coordinating and primary work, and 1.5 times that limit for replica operations. It fires before write thread pool rejections, indicating that in-flight write memory is too high rather than disk or CPU.&lt;/p></description></item><item><title>Elasticsearch indexing rate dropped to zero: where the write path stalls</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-indexing-rate-dropped/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-indexing-rate-dropped/</guid><description>&lt;h1 id="elasticsearch-indexing-rate-dropped-to-zero-where-the-write-path-stalls">Elasticsearch indexing rate dropped to zero: where the write path stalls&lt;/h1>
&lt;p>Your ingestion pipeline reports healthy connections, but Elasticsearch stopped accepting writes. The &lt;code>index_total&lt;/code> counter is flat, upstream queues are building, and documents are erroring or disappearing. Because &lt;code>index_total&lt;/code> increments for every document, update, delete, and individual bulk item, a sustained rate of zero means the write path is stalled. The cluster may still report green health, nodes may still respond to pings, and search may still work, but the pipeline is backed up.&lt;/p></description></item><item><title>Elasticsearch ingest pipeline bottleneck: grok, enrich, and per-processor time</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-ingest-pipeline-bottleneck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-ingest-pipeline-bottleneck/</guid><description>&lt;h1 id="elasticsearch-ingest-pipeline-bottleneck-grok-enrich-and-per-processor-time">Elasticsearch ingest pipeline bottleneck: grok, enrich, and per-processor time&lt;/h1>
&lt;p>Logstash or Beats instances log HTTP 429 errors. Elasticsearch indexing rate drops while upstream volume is flat. The write thread pool queue grows, but cluster health is green and there is no disk pressure or heap pressure. The culprit is often a single slow processor inside an ingest pipeline. A grok pattern backtracking on malformed logs, an enrich processor doing synchronous lookups, or a heavy Painless script throttles every document before it reaches Lucene. Because ingest processing runs on the write path, the backlog overflows the write thread pool queue and returns bulk rejections to upstream clients.&lt;/p></description></item><item><title>Elasticsearch IOException: Too many open files -- file descriptors, segments, and ulimit</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-too-many-open-files/</guid><description>&lt;h1 id="elasticsearch-ioexception-too-many-open-files----file-descriptors-segments-and-ulimit">Elasticsearch IOException: Too many open files &amp;ndash; file descriptors, segments, and ulimit&lt;/h1>
&lt;p>&lt;code>IOException: Too many open files&lt;/code> means the Elasticsearch process has reached its per-process &lt;code>RLIMIT_NOFILE&lt;/code>. Once &lt;code>open_file_descriptors&lt;/code> reaches &lt;code>max_file_descriptors&lt;/code>, the kernel returns &lt;code>EMFILE&lt;/code> on every new &lt;code>open()&lt;/code> call. The node cannot create Lucene segments, accept transport or HTTP connections, or maintain cluster membership. Each shard is a Lucene index composed of multiple segment files, and every network connection consumes a descriptor.&lt;/p></description></item><item><title>Elasticsearch JVM heap usage high: reading the sawtooth and the post-GC floor</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-jvm-heap-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-jvm-heap-high/</guid><description>&lt;h1 id="elasticsearch-jvm-heap-usage-high-reading-the-sawtooth-and-the-post-gc-floor">Elasticsearch JVM heap usage high: reading the sawtooth and the post-GC floor&lt;/h1>
&lt;p>Your Elasticsearch alert fires: &lt;code>jvm.mem.heap_used_percent&lt;/code> has crossed 75 percent and is holding there. You pull up the graph and see a jagged sawtooth climbing toward the ceiling. The first instinct is to add heap or restart the node. Both are usually wrong.&lt;/p>
&lt;p>The sawtooth is normal. Elasticsearch runs on the JVM with a young generation that fills with short-lived objects and empties on young garbage collections. The peak of the tooth is noise. The signal that matters is the post-GC floor: the minimum heap used immediately after a collection. In a healthy node, the floor stays between roughly 30 and 50 percent of max heap, and young GC dominates. When the floor trends upward, old generation objects are accumulating. Old GC pauses stop the world, and once a pause exceeds the cluster fault detection timeout, the master removes the node and triggers shard reallocation. That reallocation places more heap pressure on the survivors, beginning a death spiral.&lt;/p></description></item><item><title>Elasticsearch Limit of total fields [1000] in index has been exceeded — mapping explosion</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-limit-of-total-fields-exceeded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-limit-of-total-fields-exceeded/</guid><description>&lt;h1 id="elasticsearch-limit-of-total-fields-1000-in-index-has-been-exceeded--mapping-explosion">Elasticsearch Limit of total fields [1000] in index has been exceeded — mapping explosion&lt;/h1>
&lt;p>Every write suddenly returns &lt;code>illegal_argument_exception: Limit of total fields [1000] in index [X] has been exceeded&lt;/code>. Indexing stops. The temptation is to raise &lt;code>index.mapping.total_fields.limit&lt;/code> and move on. Do not. This is a mapping explosion: dynamic mapping creates a new field for every unique key in your documents. The 1000-field limit is a guardrail, not the root cause. Runaway mappings bloat the cluster state, inflate heap on every node, and eventually destabilize the master. Diagnose the source, relieve pressure safely, and fix the data shape so it does not recur.&lt;/p></description></item><item><title>Elasticsearch long GC pauses: old-generation stop-the-world and node drops</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-old-gc-long-pauses/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-old-gc-long-pauses/</guid><description>&lt;h1 id="elasticsearch-long-gc-pauses-old-generation-stop-the-world-and-node-drops">Elasticsearch long GC pauses: old-generation stop-the-world and node drops&lt;/h1>
&lt;p>In Elasticsearch 8.x, nodes can drop out of the cluster without logging errors. The master logs &lt;code>node-left&lt;/code> with reason &lt;code>disconnected&lt;/code>, while the departed node shows no ERROR entries because its JVM was frozen in an old-generation stop-the-world GC pause. A single pause longer than 10 seconds fails a fault-detection check; roughly 30 seconds of total unresponsiveness triggers removal. Once the master reallocates shards, remaining nodes face additional heap pressure and the cascade continues.&lt;/p></description></item><item><title>Elasticsearch mapper_parsing_exception: type conflicts and failed document indexing</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-mapper-parsing-exception/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-mapper-parsing-exception/</guid><description>&lt;h1 id="elasticsearch-mapper_parsing_exception-type-conflicts-and-failed-document-indexing">Elasticsearch mapper_parsing_exception: type conflicts and failed document indexing&lt;/h1>
&lt;p>Document counts do not match what your pipeline sent. The &lt;code>_bulk&lt;/code> endpoint returns HTTP 200, yet documents are missing from queries. In the Elasticsearch logs you see &lt;code>mapper_parsing_exception&lt;/code> with messages like &lt;code>failed to parse field [fieldname] of type [typename] in document with id [id]&lt;/code>. The &lt;code>indices.indexing.index_failed&lt;/code> counter is climbing. This is a per-document schema rejection, not a cluster outage, and it silently drops data.&lt;/p></description></item><item><title>Elasticsearch mapping explosion: dynamic mapping, cluster state bloat, and master pressure</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-mapping-explosion-dynamic-mapping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-mapping-explosion-dynamic-mapping/</guid><description>&lt;h1 id="elasticsearch-mapping-explosion-dynamic-mapping-cluster-state-bloat-and-master-pressure">Elasticsearch mapping explosion: dynamic mapping, cluster state bloat, and master pressure&lt;/h1>
&lt;p>Intermittent master elections, climbing heap on every node, and spiking indexing latency are hallmarks of a mapping explosion. Queries slow down. Administrative operations such as index creation or snapshot management crawl. If your data sources include unstructured JSON, such as application logs with variable keys, Kubernetes labels, or user-generated metadata, check mapping growth first.&lt;/p>
&lt;p>Dynamic mapping creates a new field for every unique key it encounters. Over hours, an index can accumulate tens of thousands of fields. Every mapping is part of the cluster state, which every node holds in its JVM heap. The master serializes and publishes the updated state on every mapping change. As the state grows, publication slows, pending tasks queue up, and the master becomes unstable. Unchecked, this leads to heap pressure death spirals and cascading node removals.&lt;/p></description></item><item><title>Elasticsearch master instability: frequent elections and metadata overload</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-master-instability-flapping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-master-instability-flapping/</guid><description>&lt;h1 id="elasticsearch-master-instability-frequent-elections-and-metadata-overload">Elasticsearch master instability: frequent elections and metadata overload&lt;/h1>
&lt;p>Index creation requests time out. &lt;code>_cluster/health&lt;/code> hangs or returns timeouts. The node listed by &lt;code>_cat/master&lt;/code> changes every few minutes outside planned maintenance. Shard allocation stalls, and new indices stay red or unassigned even though all data nodes are reachable. These symptoms indicate a master node that cannot keep up with cluster state updates, triggering repeated elections and leaving the cluster without stable coordination.&lt;/p></description></item><item><title>Elasticsearch master_not_discovered_exception: no elected master and stalled writes</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-no-master-not-discovered/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-no-master-not-discovered/</guid><description>&lt;h1 id="elasticsearch-master_not_discovered_exception-no-elected-master-and-stalled-writes">Elasticsearch master_not_discovered_exception: no elected master and stalled writes&lt;/h1>
&lt;p>HTTP 503 and &lt;code>master_not_discovered_exception&lt;/code> mean bulk indexing, index creation, mapping updates, and shard allocation checks are being rejected. Search requests that do not need fresh cluster state may return cached results briefly, but the cluster cannot process writes or administrative work. The data nodes may be healthy, but without an elected master, the cluster cannot update shard routing, publish state changes, or acknowledge document writes. The root cause is usually one of four problems inside the master-eligible cohort.&lt;/p></description></item><item><title>Elasticsearch merge storms: segment explosion, I/O saturation, and refresh tuning</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-merge-storms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-merge-storms/</guid><description>&lt;h1 id="elasticsearch-merge-storms-segment-explosion-io-saturation-and-refresh-tuning">Elasticsearch merge storms: segment explosion, I/O saturation, and refresh tuning&lt;/h1>
&lt;p>Search latency climbs, indexing slows, and heap usage rises while cluster health stays green. The cause is often a merge storm. Background Lucene segment consolidation has fallen behind, leaving nodes with hundreds or thousands of small segments. Each extra segment adds search overhead, consumes file descriptors, and increases memory pressure. This guide covers how merge storms develop, how to confirm the diagnosis, and how to fix them without making things worse.&lt;/p></description></item><item><title>Elasticsearch Monitoring</title><link>https://www.netdata.cloud/monitoring-101/elasticsearch-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/elasticsearch-monitoring/</guid><description>&lt;h2 id="elasticsearch-monitoring">Elasticsearch Monitoring&lt;/h2>
&lt;h3 id="what-is-elasticsearch">What Is Elasticsearch?&lt;/h3>
&lt;p>&lt;a href="https://www.elastic.co/elasticsearch/">Elasticsearch&lt;/a> is a powerful search and analytics engine designed to quickly query large volumes of data, supporting use cases such as log and event data analytics, full-text searches, and more.&lt;/p>
&lt;h3 id="monitoring-elasticsearch-with-netdata">Monitoring Elasticsearch With Netdata&lt;/h3>
&lt;p>To effectively monitor Elasticsearch, leveraging a comprehensive and dynamic monitoring tool is essential. Netdata offers a robust &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/elasticsearch/?utm_source=website&amp;amp;utm_content=monitoring101">Elasticsearch monitoring tool&lt;/a>, providing real-time insights into Elasticsearch&amp;rsquo;s performance and health. Netdata&amp;rsquo;s agent continuously collects key metrics, presenting them in interactive dashboards, which helps diagnose issues efficiently.&lt;/p></description></item><item><title>Elasticsearch monitoring checklist: the signals every production cluster needs</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-monitoring-checklist/</guid><description>&lt;h1 id="elasticsearch-monitoring-checklist-the-signals-every-production-cluster-needs">Elasticsearch monitoring checklist: the signals every production cluster needs&lt;/h1>
&lt;p>Elasticsearch failures cascade. A long GC pause on one node causes it to miss fault detection checks; the master removes it. Shards relocate to survivors, increasing heap pressure and thread pool load. If disk is near the high watermark, relocation I/O pushes other nodes toward flood stage, which sets indices to read-only and blocks writes. By the time &lt;code>GET /_cluster/health&lt;/code> returns red, the leading indicators fired minutes ago.&lt;/p></description></item><item><title>Elasticsearch monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-monitoring-maturity-model/</guid><description>&lt;h1 id="elasticsearch-monitoring-maturity-model-from-survival-to-expert">Elasticsearch monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Incidents escalate when teams monitor only the survival layer and miss leading indicators that predict cascades. A rising heap floor, a growing segment count, or a master task backlog surface long before cluster health turns red.&lt;/p>
&lt;p>This guide organizes Elasticsearch monitoring into four levels: survival, operational, mature, and expert. Each level adds signals that reduce mean time to detection and prevent composite failure patterns. Build level 1 before going live, level 2 before handling production traffic, and levels 3 and 4 after your first serious incident.&lt;/p></description></item><item><title>Elasticsearch node left the cluster: fault detection, reallocation, and recovery</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-node-left-cluster/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-node-left-cluster/</guid><description>&lt;h1 id="elasticsearch-node-left-the-cluster-fault-detection-reallocation-and-recovery">Elasticsearch node left the cluster: fault detection, reallocation, and recovery&lt;/h1>
&lt;p>Your cluster health turned yellow and &lt;code>number_of_nodes&lt;/code> dropped by one. The master logs a &lt;code>NODE_LEFT&lt;/code> event, shards are unassigned, and the remaining nodes absorb extra load. In the next minute, the allocator decides whether to move data. Misread the cause and a transient restart becomes an expensive reallocation storm, or a genuine hardware failure goes unaddressed while replicas rebalance.&lt;/p>
&lt;p>This guide covers how Elasticsearch decides a node is gone, what happens to its shards, and how to recover without deepening the incident.&lt;/p></description></item><item><title>Elasticsearch node OOM-killed: heap ceiling, page cache, and container limits</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-out-of-memory-oom-killed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-out-of-memory-oom-killed/</guid><description>&lt;h1 id="elasticsearch-node-oom-killed-heap-ceiling-page-cache-and-container-limits">Elasticsearch node OOM-killed: heap ceiling, page cache, and container limits&lt;/h1>
&lt;p>An Elasticsearch node leaves the cluster, restarts seconds later via systemd or a supervisor, and is killed again. Kernel logs show the OOM-killer terminated the Java process. &lt;code>heap.percent&lt;/code> often looks reasonable right up until the kill.&lt;/p>
&lt;p>The JVM heap is only one component of resident set size. Off-heap allocations, memory-mapped Lucene segments, and co-located processes all compete for the same memory budget. In containers, the cgroup limit is the hard boundary, not the host&amp;rsquo;s physical RAM.&lt;/p></description></item><item><title>Elasticsearch pending cluster tasks backlog: the master can't keep up</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-pending-tasks-backlog/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-pending-tasks-backlog/</guid><description>&lt;h1 id="elasticsearch-pending-cluster-tasks-backlog-the-master-cant-keep-up">Elasticsearch pending cluster tasks backlog: the master can&amp;rsquo;t keep up&lt;/h1>
&lt;p>Index creation hangs, shards stop allocating, and administrative settings updates time out. &lt;code>GET /_cluster/pending_tasks&lt;/code> shows a queue that grows instead of draining, with some tasks waiting for minutes. The elected master cannot keep up with cluster state updates, and every metadata-dependent operation stalls.&lt;/p>
&lt;p>Elasticsearch processes cluster state changes serially on the master. Index creation, mapping updates, shard allocation decisions, and settings changes all enter a single priority-ordered queue. When tasks arrive faster than the master can compute, serialize, and publish the updated state to the rest of the cluster, the backlog grows. A healthy cluster typically holds zero to five pending tasks, each resolving in under a second. Once the queue exceeds a few hundred tasks, or any individual task ages past several minutes, the cluster is approaching an operational outage.&lt;/p></description></item><item><title>Elasticsearch refresh_interval tuning: search visibility vs indexing throughput</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-refresh-interval-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-refresh-interval-tuning/</guid><description>&lt;h1 id="elasticsearch-refresh_interval-tuning-search-visibility-vs-indexing-throughput">Elasticsearch refresh_interval tuning: search visibility vs indexing throughput&lt;/h1>
&lt;p>Elasticsearch makes newly indexed documents searchable within one second by default. Every refresh creates a new Lucene segment, invalidates per-segment query caches, and adds merge debt. During bulk loads, reindexing, or high-throughput indexing, the default interval becomes a throughput bottleneck. Tuning refresh_interval requires understanding its interaction with the write path, the merge scheduler, and the query cache.&lt;/p>
&lt;p>In self-managed clusters, the default index.refresh_interval is one second, though recent versions implement lazy refresh: an index refreshes every second only if it received a search request in the last 30 seconds. If idle, auto-refresh stops until the next search arrives. Explicitly setting refresh_interval overrides lazy refresh and forces refreshes on the interval regardless of search traffic. In Elastic Cloud Serverless, the default is five seconds with a hard floor of five seconds (or -1). Setting the interval to -1 disables automatic refresh.&lt;/p></description></item><item><title>Elasticsearch replica lag: sequence-number gaps and in-sync set removal</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-replica-lag-sequence-numbers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-replica-lag-sequence-numbers/</guid><description>&lt;h1 id="elasticsearch-replica-lag-sequence-number-gaps-and-in-sync-set-removal">Elasticsearch replica lag: sequence-number gaps and in-sync set removal&lt;/h1>
&lt;p>A replica shard is falling behind its primary. The local checkpoint on the replica is lower than the global checkpoint on the primary, and the gap is growing across consecutive samples. If the replica cannot acknowledge writes fast enough, Elasticsearch removes it from the in-sync set. At that point, &lt;code>wait_for_active_shards&lt;/code> semantics change for the replication group and data redundancy is reduced. The replica must undergo peer recovery to rejoin. Read sequence-number checkpoints, find the bottleneck, and stop the lag before the master drops the replica.&lt;/p></description></item><item><title>Elasticsearch search latency high: query phase, fetch phase, and the slow shard</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-search-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-search-latency-high/</guid><description>&lt;h1 id="elasticsearch-search-latency-high-query-phase-fetch-phase-and-the-slow-shard">Elasticsearch search latency high: query phase, fetch phase, and the slow shard&lt;/h1>
&lt;p>Elasticsearch P95 search latency jumped from 50 ms to multiple seconds. Cluster health is green, search thread pool queues are growing, and users report timeouts. Elasticsearch executes every search as a two-phase scatter-gather across targeted shards, so the slowest shard sets the floor for response time. The &lt;code>_nodes/stats&lt;/code> API reports cumulative counters, not per-shard or per-query latencies, meaning one bad shard or one expensive query can be invisible to cluster-wide averages while destroying tail latency. Determine whether time is lost in the query phase, the fetch phase, or the coordinating node merge step, then isolate the outlier.&lt;/p></description></item><item><title>Elasticsearch search thread pool rejections: failed queries and small search queues</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-search-thread-pool-rejections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-search-thread-pool-rejections/</guid><description>&lt;h1 id="elasticsearch-search-thread-pool-rejections-failed-queries-and-small-search-queues">Elasticsearch search thread pool rejections: failed queries and small search queues&lt;/h1>
&lt;p>Your application logs show HTTP 429 responses from Elasticsearch. The error is &lt;code>es_rejected_execution_exception&lt;/code> and the message points to the &lt;code>search&lt;/code> thread pool. User-facing queries are failing, not just slowing down.&lt;/p>
&lt;p>The search thread pool uses a fixed number of threads and a bounded queue. The default queue size is 1000. Because search is a scatter-gather operation across shards, a single expensive query can hold a thread for seconds while the coordinating node waits for every shard to respond. The queue drains slowly, and under burst or pathological load it fills fast. Once full, Elasticsearch rejects new search requests immediately.&lt;/p></description></item><item><title>Elasticsearch shard recovery stuck: throttling, translog replay, and concurrent limits</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-shard-recovery-stuck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-shard-recovery-stuck/</guid><description>&lt;h1 id="elasticsearch-shard-recovery-stuck-throttling-translog-replay-and-concurrent-limits">Elasticsearch shard recovery stuck: throttling, translog replay, and concurrent limits&lt;/h1>
&lt;p>You restart a data node during a rollout. Ten minutes later the cluster is still yellow. You check &lt;code>_cat/recovery&lt;/code> and see a replica pinned at 38 percent for the last hour. The network is not saturated, the source node is healthy, and there are no obvious errors in the logs, yet the shard refuses to reach STARTED state.&lt;/p>
&lt;p>Stuck recovery usually comes down to four constraints: a bandwidth throttle capping transfer speed, a large translog that must replay sequentially, a concurrent recovery limit serializing work, or a disk watermark on the target node blocking allocation entirely.&lt;/p></description></item><item><title>Elasticsearch slow search after restart: cold OS page cache and warmup</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-cold-page-cache-after-restart/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-cold-page-cache-after-restart/</guid><description>&lt;h1 id="elasticsearch-slow-search-after-restart-cold-os-page-cache-and-warmup">Elasticsearch slow search after restart: cold OS page cache and warmup&lt;/h1>
&lt;p>Restart an Elasticsearch node and searches that normally return in tens of milliseconds now take seconds. CPU stays low, disk read throughput spikes, and &lt;code>iowait&lt;/code> climbs. This is not a failing disk, a runaway query, or a JVM heap problem. It is a cold OS page cache.&lt;/p>
&lt;p>Elasticsearch relies on the OS filesystem cache to serve search requests from Lucene segment files. After any restart, that cache is empty. The kernel must read segments from disk into memory on demand. Until the working set is resident, queries incur disk I/O that should have been cache hits. Depending on the ratio of dataset size to available RAM, warmup can last minutes to hours.&lt;/p></description></item><item><title>Elasticsearch snapshot failed or partial: backups that silently stop working</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-snapshot-failed-or-partial/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-snapshot-failed-or-partial/</guid><description>&lt;h1 id="elasticsearch-snapshot-failed-or-partial-backups-that-silently-stop-working">Elasticsearch snapshot failed or partial: backups that silently stop working&lt;/h1>
&lt;p>You check backups and the last successful snapshot is three days old. Or you see PARTIAL snapshots where only some indices were captured. Cluster health is green, indexing and search are fine, but your recovery point is slipping. Elasticsearch does not fail the cluster when snapshots break, so the problem surfaces only when someone asks, &amp;ldquo;When did we last test a restore?&amp;rdquo;&lt;/p></description></item><item><title>Elasticsearch this action would add too many shards: max_shards_per_node limit</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-max-shards-per-node-exceeded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-max-shards-per-node-exceeded/</guid><description>&lt;h1 id="elasticsearch-this-action-would-add-too-many-shards-max_shards_per_node-limit">Elasticsearch this action would add too many shards: max_shards_per_node limit&lt;/h1>
&lt;p>Creating an index or rolling over a data stream returns HTTP 400 &lt;code>validation_exception&lt;/code>: &amp;ldquo;this action would add [N] shards, but this cluster currently has [X]/[Y] maximum normal shards open&amp;rdquo;. The cluster has hit &lt;code>cluster.max_shards_per_node&lt;/code>, which defaults to 1000 open shards per non-frozen data node. Raising the limit via &lt;code>_cluster/settings&lt;/code> unblocks writes but postpones the outage. The durable fix is consolidation.&lt;/p></description></item><item><title>Elasticsearch thread pool queue growing: the precursor to rejection</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-thread-pool-queue-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-thread-pool-queue-growing/</guid><description>&lt;h1 id="elasticsearch-thread-pool-queue-growing-the-precursor-to-rejection">Elasticsearch thread pool queue growing: the precursor to rejection&lt;/h1>
&lt;p>A climbing &lt;code>write&lt;/code> or &lt;code>search&lt;/code> queue in &lt;code>_cat/thread_pool&lt;/code> means a node is receiving work faster than it can complete it. Rejections are the lagging indicator. By the time clients see &lt;code>EsRejectedExecutionException&lt;/code>, the cluster is already degraded.&lt;/p>
&lt;p>For the &lt;code>write&lt;/code> pool, the default queue size is 10000 (ES 7.x+). For &lt;code>search&lt;/code>, it is 1000. Sustained write queues above 1000 or search queues above 100 warrant investigation. The &lt;code>management&lt;/code> pool is different: even small amounts of sustained queuing mean the master is falling behind on cluster state operations, which blocks allocation, mapping updates, and recovery.&lt;/p></description></item><item><title>Elasticsearch TLS certificate expiry: cluster fragmentation and client lockout</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-tls-certificate-expiry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-tls-certificate-expiry/</guid><description>&lt;h1 id="elasticsearch-tls-certificate-expiry-cluster-fragmentation-and-client-lockout">Elasticsearch TLS certificate expiry: cluster fragmentation and client lockout&lt;/h1>
&lt;p>A node reboots and never rejoins the cluster. Kibana shows connection errors while your data pipeline buffers events. curl returns a TLS handshake failure even though the Elasticsearch process is still listening on port 9200. In Elasticsearch 8.x, security is enabled by default: every node and client relies on TLS certificates. When they expire, failure is abrupt and total. Transport-layer expiry fragments the cluster by rejecting inter-node handshakes. HTTP-layer expiry locks out clients while the cluster internals may still operate. There is no built-in grace period. At the expiry timestamp, connections fail immediately, often with no prior warning in the application logs.&lt;/p></description></item><item><title>Elasticsearch too many segments per shard: search slowdown and force-merge</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-too-many-segments/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-too-many-segments/</guid><description>&lt;h1 id="elasticsearch-too-many-segments-per-shard-search-slowdown-and-force-merge">Elasticsearch too many segments per shard: search slowdown and force-merge&lt;/h1>
&lt;p>Searches that used to complete in tens of milliseconds are now taking hundreds. CPU on data nodes climbs without a traffic spike. Cluster health stays green, but query latency degrades and timeouts increase. The problem is invisible at the cluster level: shards have accumulated too many Lucene segments, and every query scans every segment in every relevant shard. When a shard holds more than roughly 100 segments, the background merge policy has fallen behind. Search slows, file descriptor usage rises, and merge I/O competes with indexing and queries. Since Elasticsearch 7.7, most segment metadata moved off-heap&lt;!-- TODO: verify 7.7 is the correct version for off-heap segment metadata change -->, so the node may not run out of JVM heap, but search performance degrades all the same.&lt;/p></description></item><item><title>Elasticsearch too many shards per node: overallocation and the heap tax</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-shard-overallocation-too-many-shards/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-shard-overallocation-too-many-shards/</guid><description>&lt;h1 id="elasticsearch-too-many-shards-per-node-overallocation-and-the-heap-tax">Elasticsearch too many shards per node: overallocation and the heap tax&lt;/h1>
&lt;p>Searches slow from milliseconds to seconds. Nodes drop out after long GC pauses, triggering shard relocations that overload the remaining nodes. Cluster health stays green while &lt;code>_cat/nodes&lt;/code> shows climbing heap and &lt;code>segments.memory&lt;/code> growing steadily. The cause is usually shard overallocation: thousands of open shards with individual nodes carrying more than their heap can sustain.&lt;/p>
&lt;p>Each shard is a self-contained Lucene index. Elasticsearch keeps per-segment metadata in heap memory, and that overhead is charged per field, not per segment. A node with 1,000 shards may be spending multiple gigabytes of heap on metadata alone, leaving less room for caches, aggregations, and in-flight requests. &lt;!-- TODO: verify aggregate heap-to-disk ratio for segment metadata is approximately 1:4000 since ES 7.7 -->&lt;/p></description></item><item><title>Elasticsearch translog growing: flush problems, durability, and slow recovery</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-translog-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-translog-growing/</guid><description>&lt;h1 id="elasticsearch-translog-growing-flush-problems-durability-and-slow-recovery">Elasticsearch translog growing: flush problems, durability, and slow recovery&lt;/h1>
&lt;p>Shard recovery estimates climb from minutes to hours. A node restart stalls during translog replay. Disk usage creeps up on data nodes, and uncommitted translog size per shard is past the flush threshold. The translog is growing faster than Elasticsearch can flush it to Lucene commit points.&lt;/p>
&lt;p>Every indexing operation appends to the per-shard write-ahead log before the document is searchable. A flush commits those operations to a Lucene segment and truncates the log. When flush cannot keep pace with the write path, uncommitted operations accumulate. During recovery, Elasticsearch replays every operation sequentially. A multi-gigabyte translog translates directly into a multi-hour recovery window, extending time-to-recovery and increasing cascade risk if another node fails during replay.&lt;/p></description></item><item><title>Elasticsearch unassigned shards: reading allocation explain and fixing each reason</title><link>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-unassigned-shards/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/elasticsearch-unassigned-shards/</guid><description>&lt;h1 id="elasticsearch-unassigned-shards-reading-allocation-explain-and-fixing-each-reason">Elasticsearch unassigned shards: reading allocation explain and fixing each reason&lt;/h1>
&lt;p>Yellow or red cluster health with &lt;code>unassigned_shards &amp;gt; 0&lt;/code> means the allocator cannot place one or more shard copies on any node. Missing primaries block queries and risk data loss; missing replicas only cost redundancy. Do not guess from the cluster color. The allocator already knows why it rejected every node. Ask it.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>Unassigned primaries make their data unreachable. Affected indices return partial results or fail. Unassigned replicas remove redundancy; a second failure on those primaries drops the data. The master allocator evaluates every node through a chain of deciders: disk watermarks, allocation filters, awareness attributes, the same-shard rule, and retry limits. When every node is rejected, the shard stays &lt;code>UNASSIGNED&lt;/code> until the blocking condition clears or you intervene.&lt;/p></description></item><item><title>Electroline Equipment Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/electroline-equipment-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/electroline-equipment-inc-snmp-traps/</guid><description/></item><item><title>Elfiq Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/elfiq-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/elfiq-inc-snmp-traps/</guid><description/></item><item><title>Elgato Key Light devices.</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/elgato-key-light-devices./</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/elgato-key-light-devices./</guid><description/></item><item><title>Elgato Key Light Monitoring</title><link>https://www.netdata.cloud/monitoring-101/elgato_keylight-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/elgato_keylight-monitoring/</guid><description>&lt;h2 id="elgato-key-light-monitoring">Elgato Key Light Monitoring&lt;/h2>
&lt;h3 id="what-is-elgato-key-light">What Is Elgato Key Light?&lt;/h3>
&lt;p>Elgato Key Light devices are high-quality lighting solutions that offer advanced controls for content creators, videographers, and professionals. These devices ensure optimal lighting conditions and are integral to achieving the perfect visual setup with ease. By intelligently managing your lighting environment, Elgato Key Light ensures consistent quality outputs in your video productions or live streams.&lt;/p>
&lt;h3 id="monitoring-elgato-key-light-with-netdata">Monitoring Elgato Key Light With Netdata&lt;/h3>
&lt;p>To monitor Elgato Key Light devices, Netdata utilizes an OpenMetrics (Prometheus) exporter. This allows users to gather valuable metrics by periodically sending HTTP requests to the &lt;a href="https://github.com/mdlayher/keylight_exporter">Elgato Key Light exporter&lt;/a>. Netdata excels in ingesting data from any Prometheus exporter, providing automated dashboards, alerts, and more, without requiring a Prometheus server or Grafana. This seamless integration enables you to keep track of lighting metrics in real-time, enhancing control and management of your environment.&lt;/p></description></item><item><title>Elitecore Technologies Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/elitecore-technologies-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/elitecore-technologies-ltd-snmp-traps/</guid><description/></item><item><title>Eltek Energy AS SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/eltek-energy-as-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/eltek-energy-as-snmp-traps/</guid><description/></item><item><title>Eltek Valere Inc Formerly Valere Power Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/eltek-valere-inc-formerly-valere-power-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/eltek-valere-inc-formerly-valere-power-inc-snmp-traps/</guid><description/></item><item><title>Eltex Enterprise Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/eltex-enterprise-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/eltex-enterprise-ltd-snmp-traps/</guid><description/></item><item><title>Email</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/email/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/email/</guid><description/></item><item><title>Emc Clariion Advanced Storage Solutions SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/emc-clariion-advanced-storage-solutions-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/emc-clariion-advanced-storage-solutions-snmp-traps/</guid><description/></item><item><title>Emc Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/emc-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/emc-corp-snmp-traps/</guid><description/></item><item><title>Emc Data General Division SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/emc-data-general-division-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/emc-data-general-division-snmp-traps/</guid><description/></item><item><title>Empire Technologies Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/empire-technologies-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/empire-technologies-inc-snmp-traps/</guid><description/></item><item><title>Endace Technology SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/endace-technology-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/endace-technology-snmp-traps/</guid><description/></item><item><title>Endrun Technologies LLC SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/endrun-technologies-llc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/endrun-technologies-llc-snmp-traps/</guid><description/></item><item><title>Energi Core Wallet Monitoring</title><link>https://www.netdata.cloud/monitoring-101/energid-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/energid-monitoring/</guid><description>&lt;h2 id="what-is-energi-core-wallet">What is Energi Core Wallet?&lt;/h2>
&lt;p>Energi is a cryptocurrency that is built on the Ethereum blockchain. It is designed to be a self-funding and self-governing cryptocurrency that uses a hybrid Proof-of-Stake (PoS) and Proof-of-Work (PoW) consensus model to improve scalability and security. The Energi Core Wallet is the official wallet software for Energi. It is a software application that allows users to securely store, send and receive Energi coins. The wallet also includes advanced features such as staking, Masternode setup, and coin control.&lt;/p></description></item><item><title>Energomera smart power meters</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/energomera-smart-power-meters/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/energomera-smart-power-meters/</guid><description/></item><item><title>Energomera Smart Power Meters Monitoring</title><link>https://www.netdata.cloud/monitoring-101/energomera-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/energomera-monitoring/</guid><description>&lt;h2 id="energomera-smart-power-meters-monitoring">Energomera Smart Power Meters Monitoring&lt;/h2>
&lt;h3 id="what-is-energomera">What Is Energomera?&lt;/h3>
&lt;p>Energomera smart power meters are innovative devices designed to efficiently measure and manage electricity usage. These meters offer high accuracy and reliability, providing essential data for effective energy management. Used widely in the industrial and consumer sectors, Energomera helps in monitoring power consumption patterns, optimizing energy usage, and reducing costs.&lt;/p>
&lt;h3 id="monitoring-energomera-with-netdata">Monitoring Energomera With Netdata&lt;/h3>
&lt;p>To effectively monitor Energomera smart power meters, Netdata leverages an openmetrics (Prometheus) exporter. One of Netdata&amp;rsquo;s strengths lies in its ability to ingest data from any Prometheus exporter, presenting users with automated dashboards, alerts, and more—all without the need for a Prometheus server or Grafana. This ensures that monitoring solutions remain lightweight, easy to deploy, and comprehensive in their data collection.&lt;/p></description></item><item><title>Engenio Information Technologies Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/engenio-information-technologies-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/engenio-information-technologies-inc-snmp-traps/</guid><description/></item><item><title>Enlogic Systems LLC SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/enlogic-systems-llc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/enlogic-systems-llc-snmp-traps/</guid><description/></item><item><title>Enter our waitlist to trial Netdata's AI features</title><link>https://www.netdata.cloud/netdata-ai-waitlist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/netdata-ai-waitlist/</guid><description/></item><item><title>Enterasys Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/enterasys-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/enterasys-networks-inc-snmp-traps/</guid><description/></item><item><title>Enterasys Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/enterasys-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/enterasys-networks-snmp-traps/</guid><description/></item><item><title>Enterprise 1004849 SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/enterprise-1004849-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/enterprise-1004849-snmp-traps/</guid><description/></item><item><title>Entropy</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/entropy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/entropy/</guid><description/></item><item><title>Envoy</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/envoy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/envoy/</guid><description/></item><item><title>Envoy 502 and upstream resets: rx_reset, tx_reset, and mid-response failures</title><link>https://www.netdata.cloud/guides/envoy/envoy-502-upstream-connection-termination/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-502-upstream-connection-termination/</guid><description>&lt;h1 id="envoy-502-and-upstream-resets-rx_reset-tx_reset-and-mid-response-failures">Envoy 502 and upstream resets: rx_reset, tx_reset, and mid-response failures&lt;/h1>
&lt;p>A client request made it past Envoy&amp;rsquo;s routing, established a TCP connection to an upstream host, and then the connection died before the response completed. The access log shows a &lt;code>UC&lt;/code> or &lt;code>UPE&lt;/code> response flag, &lt;code>upstream_rq_rx_reset&lt;/code> is climbing, and the client received a 502 or 503.&lt;/p>
&lt;p>This is a different failure class from &amp;ldquo;no healthy upstream&amp;rdquo; (503, &lt;code>NR&lt;/code> flag) or &amp;ldquo;connection refused&amp;rdquo; (503, &lt;code>UF&lt;/code> flag). Those happen before a connection is established. Resets happen after the handshake succeeds, which means the upstream was reachable and then failed during request processing. The diagnosis and fix are completely different.&lt;/p></description></item><item><title>Envoy 503 with response flag UO: a tripped circuit breaker, not a dead backend</title><link>https://www.netdata.cloud/guides/envoy/envoy-503-uo-circuit-breaker/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-503-uo-circuit-breaker/</guid><description>&lt;h1 id="envoy-503-with-response-flag-uo-a-tripped-circuit-breaker-not-a-dead-backend">Envoy 503 with response flag UO: a tripped circuit breaker, not a dead backend&lt;/h1>
&lt;p>You see 503 responses in access logs with the &lt;code>UO&lt;/code> response flag. The flag is what pins the cause: &lt;code>UO&lt;/code> means upstream overflow. A circuit breaker fast-failed the request locally before forwarding. The backend did not refuse the connection, did not time out, and may not be aware the request existed. Envoy decided the cluster was at capacity and returned an immediate 503.&lt;/p></description></item><item><title>Envoy 504 upstream timeout: upstream_rq_timeout, per-try timeouts, and the UT flag</title><link>https://www.netdata.cloud/guides/envoy/envoy-504-upstream-timeout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-504-upstream-timeout/</guid><description>&lt;h1 id="envoy-504-upstream-timeout-upstream_rq_timeout-per-try-timeouts-and-the-ut-flag">Envoy 504 upstream timeout: upstream_rq_timeout, per-try timeouts, and the UT flag&lt;/h1>
&lt;p>A 504 from Envoy with response flag &lt;code>UT&lt;/code> means the upstream request timeout fired: Envoy sent the request to a backend and did not receive a complete response within the configured budget. The backend may be genuinely slow, the timeout may be too tight for the workload, or a filter may be interfering with the timeout clock. Each case needs a different fix.&lt;/p></description></item><item><title>Envoy circuit breaker open: cx_open, rq_pending_open, and fast-failed requests</title><link>https://www.netdata.cloud/guides/envoy/envoy-circuit-breaker-cx-open/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-circuit-breaker-cx-open/</guid><description>&lt;h1 id="envoy-circuit-breaker-open-cx_open-rq_pending_open-and-fast-failed-requests">Envoy circuit breaker open: cx_open, rq_pending_open, and fast-failed requests&lt;/h1>
&lt;p>You see &lt;code>cx_open=1&lt;/code> or &lt;code>rq_pending_open=1&lt;/code> on a production cluster. Access logs show 503 responses tagged with the &lt;code>UO&lt;/code> response flag. Clients receive fast-failed requests, sometimes with an &lt;code>x-envoy-overloaded&lt;/code> header. The circuit breaker gauges are binary: 0 means the breaker has headroom and can admit more work, 1 means it is at capacity and rejecting. Each gauge is scoped per-cluster and per-priority (default or high), so you need the right cluster and priority combination to read the signal correctly.&lt;/p></description></item><item><title>Envoy clusters stuck warming: warming_clusters non-zero and routes returning 503</title><link>https://www.netdata.cloud/guides/envoy/envoy-cluster-warming-stuck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-cluster-warming-stuck/</guid><description>&lt;p>&lt;code>cluster_manager.warming_clusters&lt;/code> reading non-zero during steady state is a config convergence problem. Envoy has accepted a new cluster (or a cluster update) but cannot activate it because a dependency has not resolved: DNS, an SDS secret, an EDS endpoint set, or an active health-check initialization. Until warming completes, the cluster is invisible to routing for new additions, or held in swap-for-update for modifications. Routes targeting a freshly added but still-warming cluster return 503 with response flag NC (no cluster), or 404 with NR (no route), depending on whether the route table entry has been pushed yet.&lt;/p></description></item><item><title>Envoy connection churn: a low reuse ratio and keepalive misconfiguration</title><link>https://www.netdata.cloud/guides/envoy/envoy-upstream-connection-reuse-churn/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-upstream-connection-reuse-churn/</guid><description>&lt;h1 id="envoy-connection-churn-a-low-reuse-ratio-and-keepalive-misconfiguration">Envoy connection churn: a low reuse ratio and keepalive misconfiguration&lt;/h1>
&lt;p>Envoy&amp;rsquo;s upstream connection pool amortizes TCP (and TLS) handshakes across many requests. When it stops doing that, you have connection churn: the pool establishes a fresh connection for nearly every request, and the reuse ratio collapses toward 1.0.&lt;/p>
&lt;p>The headline signal is the ratio &lt;code>upstream_rq_total / upstream_cx_total&lt;/code>. For HTTP/1.1 with keepalive working, expect a number well above 1, often in the tens. For HTTP/2 with stream multiplexing, expect far higher still. When the ratio sits near 1.0, every request pays the full handshake tax.&lt;/p></description></item><item><title>Envoy connection pool exhaustion: a slow upstream that fills the pool</title><link>https://www.netdata.cloud/guides/envoy/envoy-connection-pool-exhaustion-slow-upstream/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-connection-pool-exhaustion-slow-upstream/</guid><description>&lt;h1 id="envoy-connection-pool-exhaustion-a-slow-upstream-that-fills-the-pool">Envoy connection pool exhaustion: a slow upstream that fills the pool&lt;/h1>
&lt;p>Clients are seeing 503 responses. Access logs show response flag UO on those requests. But upstream health looks fine: &lt;code>membership_healthy&lt;/code> is stable and hosts are passing active health checks. The upstream is not down, yet Envoy is refusing to forward new requests to it.&lt;/p>
&lt;p>This is connection pool exhaustion. An upstream that was previously fast has become slow. Each request now holds a connection longer, so the same request rate fills more connection slots. The per-worker connection pool saturates, new requests queue in the pending buffer, the buffer overflows, and Envoy fast-fails those requests with 503 UO rather than piling on more load.&lt;/p></description></item><item><title>Envoy control_plane.connected_state = 0: running on stale xDS config</title><link>https://www.netdata.cloud/guides/envoy/envoy-control-plane-connected-state/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-control-plane-connected-state/</guid><description>&lt;p>&lt;code>control_plane.connected_state&lt;/code> is a gauge that flips to 0 the moment Envoy&amp;rsquo;s gRPC stream to its xDS management server drops. The flip produces no user-visible symptom: existing traffic keeps flowing, listener sockets stay open, and Envoy keeps serving its last-known-good configuration. There is no default expiry on that cached config.&lt;/p>
&lt;p>That silence is why this signal is under-monitored. By the time anyone notices, new endpoints have been invisible for hours, removed endpoints have been sending traffic to dead or reassigned IPs, routes never updated, and SDS-managed certificates stopped rotating. The disconnect happened long before the visible incident, which makes root-cause correlation non-obvious.&lt;/p></description></item><item><title>Envoy DNS resolution failures: STRICT_DNS clusters serving stale endpoints</title><link>https://www.netdata.cloud/guides/envoy/envoy-dns-resolution-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-dns-resolution-failure/</guid><description>&lt;h1 id="envoy-dns-resolution-failures-strict_dns-clusters-serving-stale-endpoints">Envoy DNS resolution failures: STRICT_DNS clusters serving stale endpoints&lt;/h1>
&lt;p>A STRICT_DNS cluster stops picking up new endpoints, and endpoints you removed from DNS hours ago are still receiving traffic. Envoy&amp;rsquo;s own dashboard shows the cluster as healthy, error rates look flat, and the upstream service reports traffic to IPs that no longer exist. The first concrete evidence is often a wave of 502s or connection failures when a stale IP gets reassigned to an unrelated workload.&lt;/p></description></item><item><title>Envoy downstream 4xx spike: 401s, 403s, and 404s from the client side</title><link>https://www.netdata.cloud/guides/envoy/envoy-downstream-4xx-auth-spike/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-downstream-4xx-auth-spike/</guid><description>&lt;h1 id="envoy-downstream-4xx-spike-401s-403s-and-404s-from-the-client-side">Envoy downstream 4xx spike: 401s, 403s, and 404s from the client side&lt;/h1>
&lt;p>A spike in &lt;code>http.&amp;lt;stat_prefix&amp;gt;.downstream_rq_4xx&lt;/code> is usually a client-side story, not an Envoy story. The proxy is reporting that clients sent bad, unauthorized, or unroutable requests. The single counter lumps 400s, 401s, 403s, and 404s together, and Envoy does not expose per-status-code downstream counters. You will not find &lt;code>downstream_rq_401&lt;/code> or &lt;code>downstream_rq_403&lt;/code> in the stats dump. &lt;!-- TODO: verify the per-code downstream counter gap is still present in current Envoy builds as of mid-2026 -->&lt;/p></description></item><item><title>Envoy downstream connection flood: slowloris, the cx-to-rq ratio, and oversized requests</title><link>https://www.netdata.cloud/guides/envoy/envoy-downstream-connection-flood-slowloris/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-downstream-connection-flood-slowloris/</guid><description>&lt;h1 id="envoy-downstream-connection-flood-slowloris-the-cx-to-rq-ratio-and-oversized-requests">Envoy downstream connection flood: slowloris, the cx-to-rq ratio, and oversized requests&lt;/h1>
&lt;p>The proxy is rejecting new connections, latency is climbing, and upstreams look healthy. The problem is at the front door. A downstream connection flood exhausts Envoy&amp;rsquo;s resources before a request reaches a filter chain. The classic signal: a high new-connection rate paired with a low request rate. Many TCP handshakes, few HTTP requests. That ratio is the cx-to-rq ratio, and when it inverts you are looking at a slowloris-style attack or a misconfigured client pool.&lt;/p></description></item><item><title>Envoy downstream request rate dropping to zero: the traffic blackhole</title><link>https://www.netdata.cloud/guides/envoy/envoy-traffic-blackhole-downstream-rq-drop/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-traffic-blackhole-downstream-rq-drop/</guid><description>&lt;h1 id="envoy-downstream-request-rate-dropping-to-zero-the-traffic-blackhole">Envoy downstream request rate dropping to zero: the traffic blackhole&lt;/h1>
&lt;p>A sudden collapse of &lt;code>http.&amp;lt;stat_prefix&amp;gt;.downstream_rq_total&lt;/code> to near zero is one of the loudest availability signals Envoy can emit, and one of the easiest to misread. Operators trained to chase 5xx spikes often treat a falling request rate as &amp;ldquo;the system is calm.&amp;rdquo; It is not. A drop below roughly 10% of baseline on a listener that should be receiving traffic is a traffic blackhole: clients have stopped reaching Envoy, Envoy has stopped accepting requests, or traffic is reaching Envoy but being answered by local replies before any upstream work happens.&lt;/p></description></item><item><title>Envoy downstream_cx_active growing: connection leaks and idle-timeout gaps</title><link>https://www.netdata.cloud/guides/envoy/envoy-downstream-cx-active-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-downstream-cx-active-leak/</guid><description>&lt;h1 id="envoy-downstream_cx_active-growing-connection-leaks-and-idle-timeout-gaps">Envoy downstream_cx_active growing: connection leaks and idle-timeout gaps&lt;/h1>
&lt;p>&lt;code>listener.&amp;lt;address&amp;gt;.downstream_cx_active&lt;/code> is a gauge counting the downstream (client-facing) connections currently held open by a listener. When it climbs without a matching rise in request rate, the usual first call is &amp;ldquo;connection leak&amp;rdquo;. That diagnosis is sometimes correct and sometimes wrong, and the fixes are completely different.&lt;/p>
&lt;p>The gauge reflects three independent inputs: how fast new connections arrive (&lt;code>downstream_cx_total&lt;/code>), how fast old connections close (&lt;code>downstream_cx_destroy&lt;/code>), and how long each connection is allowed to live once idle (idle, stream, drain, and connection-duration timers). A leak is only one explanation. Others include clients that legitimately hold connections open (HTTP/2 and gRPC multiplexing), loopback sidecar traffic, an idle timeout set far longer than the workload needs, or a slow upstream that holds streams, and therefore connections, open past the point the idle timer could reclaim them.&lt;/p></description></item><item><title>Envoy downstream_cx_overflow and overload_reject: connections turned away at the door</title><link>https://www.netdata.cloud/guides/envoy/envoy-downstream-cx-overflow-overload-reject/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-downstream-cx-overflow-overload-reject/</guid><description>&lt;h1 id="envoy-downstream_cx_overflow-and-overload_reject-connections-turned-away-at-the-door">Envoy downstream_cx_overflow and overload_reject: connections turned away at the door&lt;/h1>
&lt;p>When Envoy rejects a new TCP connection at the listener, the client gets no HTTP error or retry hint: the SYN is answered with a RST or dropped, and the client sees connection refused or a connect timeout. The only evidence is three listener-level counters that should always be zero: &lt;code>downstream_cx_overflow&lt;/code>, &lt;code>downstream_cx_overload_reject&lt;/code>, and &lt;code>downstream_global_cx_overflow&lt;/code>.&lt;/p>
&lt;p>Each counter corresponds to a distinct rejection mechanism with a different root cause and fix. Confusing them wastes time tuning the wrong limit while the real problem is memory pressure, file descriptor exhaustion, or a missing overload manager.&lt;/p></description></item><item><title>Envoy downstream_rq_time high: client-observed latency and proxy overhead</title><link>https://www.netdata.cloud/guides/envoy/envoy-downstream-rq-time-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-downstream-rq-time-high/</guid><description>&lt;h1 id="envoy-downstream_rq_time-high-client-observed-latency-and-proxy-overhead">Envoy downstream_rq_time high: client-observed latency and proxy overhead&lt;/h1>
&lt;p>High &lt;code>downstream_rq_time&lt;/code> means client-observed latency through the proxy is climbing. The histogram lives under &lt;code>http.&amp;lt;stat_prefix&amp;gt;.downstream_rq_time&lt;/code>, measured in milliseconds, and it maps most directly to user-perceived slowness. When it breaches SLO, you are in an incident whether the upstreams are healthy or not.&lt;/p>
&lt;p>The common reflex is to subtract &lt;code>upstream_rq_time&lt;/code> from &lt;code>downstream_rq_time&lt;/code> and label the remainder &amp;ldquo;Envoy overhead.&amp;rdquo; That delta is useful as a trend, but it is not a clean measurement of proxy processing time. It folds in filter execution, downstream upload and download behavior, retry time across multiple attempts, buffering delays, streaming duration, and how fast the client reads the response. A high delta can mean Envoy is working hard, a client is slow, a request was retried, or a long-lived SSE stream is open. The histogram alone will not tell you which.&lt;/p></description></item><item><title>Envoy ext_authz failure_mode_allowed: unauthenticated traffic when auth is down</title><link>https://www.netdata.cloud/guides/envoy/envoy-ext-authz-failure-mode-allowed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-ext-authz-failure-mode-allowed/</guid><description>&lt;h1 id="envoy-ext_authz-failure_mode_allowed-unauthenticated-traffic-when-auth-is-down">Envoy ext_authz failure_mode_allowed: unauthenticated traffic when auth is down&lt;/h1>
&lt;p>The Envoy &lt;code>ext_authz&lt;/code> filter calls an external authorization service on every request that matches the filter chain. That call is on the request critical path: nothing is forwarded upstream until the auth service returns a decision. When the auth service is unreachable, returns an HTTP 5xx, or exceeds its timeout, Envoy either lets the request through without an auth decision or rejects it. That choice is controlled by &lt;code>failure_mode_allow&lt;/code>, and every time the fail-open branch fires, Envoy increments &lt;code>http.&amp;lt;stat_prefix&amp;gt;.ext_authz.failure_mode_allowed&lt;/code>.&lt;/p></description></item><item><title>Envoy file descriptor exhaustion: the FD cliff that refuses every new connection</title><link>https://www.netdata.cloud/guides/envoy/envoy-file-descriptor-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-file-descriptor-exhaustion/</guid><description>&lt;h1 id="envoy-file-descriptor-exhaustion-the-fd-cliff-that-refuses-every-new-connection">Envoy file descriptor exhaustion: the FD cliff that refuses every new connection&lt;/h1>
&lt;p>When Envoy runs out of file descriptors, it is not a slow degradation. At the &lt;code>RLIMIT_NOFILE&lt;/code> ceiling every new downstream &lt;code>accept()&lt;/code> and every new upstream &lt;code>connect()&lt;/code> fails at once. Clients see resets, timeouts, or 503s, and the proxy looks dead even though established streams keep flowing. Unlike memory or CPU pressure, there is no graceful ramp and rarely an overload-manager warning before the cliff.&lt;/p></description></item><item><title>Envoy health checks vs outlier detection: two systems that eject hosts differently</title><link>https://www.netdata.cloud/guides/envoy/envoy-health-check-vs-outlier-detection/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-health-check-vs-outlier-detection/</guid><description>&lt;h1 id="envoy-health-checks-vs-outlier-detection-two-systems-that-eject-hosts-differently">Envoy health checks vs outlier detection: two systems that eject hosts differently&lt;/h1>
&lt;p>Envoy has two independent mechanisms for removing unhealthy upstream hosts from load balancing. Active health checks send synthetic probes on a schedule. Outlier detection watches real request traffic and ejects hosts based on observed errors. They are not redundant, and they do not always agree.&lt;/p>
&lt;p>A host can pass every health check and still be ejected by outlier detection. The two systems run on different threads, observe different signals, and act on different timelines. Operators who watch only &lt;code>membership_healthy&lt;/code> or only &lt;code>outlier_detection.ejections_active&lt;/code> miss half the picture.&lt;/p></description></item><item><title>Envoy hot restart races: draining, epoch churn, and dropped connections</title><link>https://www.netdata.cloud/guides/envoy/envoy-hot-restart-race-draining/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-hot-restart-race-draining/</guid><description>&lt;h1 id="envoy-hot-restart-races-draining-epoch-churn-and-dropped-connections">Envoy hot restart races: draining, epoch churn, and dropped connections&lt;/h1>
&lt;p>You deployed a new Envoy binary or triggered a reload that invoked a hot restart. Seconds later, clients report connection resets, your 5xx rate ticks up, and &lt;code>server.hot_restart_epoch&lt;/code> is climbing on your dashboard. The new process was supposed to inherit listen sockets gracefully while the old one drained.&lt;/p>
&lt;p>Hot restart is Envoy&amp;rsquo;s mechanism for zero-downtime binary upgrades and certain reloads. A new process launches, coordinates with the old one over a Unix domain socket, takes over the listen sockets, and the old process enters a drain sequence. Both processes run simultaneously during the handoff. That coexistence is where the races live.&lt;/p></description></item><item><title>Envoy listener_create_failure: a listener config Envoy could not apply</title><link>https://www.netdata.cloud/guides/envoy/envoy-listener-create-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-listener-create-failure/</guid><description>&lt;h1 id="envoy-listener_create_failure-a-listener-config-envoy-could-not-apply">Envoy listener_create_failure: a listener config Envoy could not apply&lt;/h1>
&lt;p>The &lt;code>listener_manager.listener_create_failure&lt;/code> counter increments when Envoy cannot add a listener object to its workers. In a healthy deployment this counter is zero. Any non-zero value means Envoy received a listener configuration it could not apply: a failed bind (port conflict, permission denied) or invalid configuration.&lt;/p>
&lt;p>The signal is easy to miss because Envoy does not crash. It keeps serving the previous listener configuration, traffic on existing routes flows normally, and the listener the operator intended to add or modify never becomes active. Deployment pipelines report success. The control plane believes the config was accepted. Only the process log and this counter reveal the failure.&lt;/p></description></item><item><title>Envoy membership_healthy dropping: reading the single most important cluster signal</title><link>https://www.netdata.cloud/guides/envoy/envoy-membership-healthy-dropping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-membership-healthy-dropping/</guid><description>&lt;h1 id="envoy-membership_healthy-dropping-reading-the-single-most-important-cluster-signal">Envoy membership_healthy dropping: reading the single most important cluster signal&lt;/h1>
&lt;p>If you watch only one availability metric per Envoy cluster, watch the ratio of &lt;code>cluster.&amp;lt;name&amp;gt;.membership_healthy&lt;/code> to &lt;code>cluster.&amp;lt;name&amp;gt;.membership_total&lt;/code>. When the ratio collapses, the remaining healthy hosts take proportionally more load, and the cluster is one or two failures away from panic mode. Upstream 5xx rate, latency, circuit breaker state, and retries are all downstream of host availability.&lt;/p>
&lt;p>The hard part is reading the gauge correctly, not collecting it. A drop can mean the upstream is genuinely broken, the control plane removed endpoints, or outlier detection ejected hosts that are still passing active health checks. Each case has a different response.&lt;/p></description></item><item><title>Envoy memory pressure spiral: server.memory_allocated climbing toward OOM</title><link>https://www.netdata.cloud/guides/envoy/envoy-memory-pressure-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-memory-pressure-spiral/</guid><description>&lt;h1 id="envoy-memory-pressure-spiral-servermemory_allocated-climbing-toward-oom">Envoy memory pressure spiral: server.memory_allocated climbing toward OOM&lt;/h1>
&lt;p>&lt;code>server.memory_allocated&lt;/code> is climbing, traffic is flat, and you are hours or minutes from an OOM kill. In a sidecar deployment, the pod dies and the application gets the blame. In an edge or gateway deployment, Envoy vanishes and clients see connection resets across the board.&lt;/p>
&lt;p>Envoy has a built-in protection mechanism (the overload manager), but many deployments never configure it. Without it, there is no graceful degradation: Envoy goes straight from &amp;ldquo;memory looks fine&amp;rdquo; to &amp;ldquo;OOM killed&amp;rdquo; with nothing in between.&lt;/p></description></item><item><title>Envoy memory_heap_size vs memory_allocated: tcmalloc fragmentation that fools dashboards</title><link>https://www.netdata.cloud/guides/envoy/envoy-memory-heap-fragmentation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-memory-heap-fragmentation/</guid><description>&lt;h1 id="envoy-memory_heap_size-vs-memory_allocated-tcmalloc-fragmentation-that-fools-dashboards">Envoy memory_heap_size vs memory_allocated: tcmalloc fragmentation that fools dashboards&lt;/h1>
&lt;p>Envoy exposes three memory gauges that operators naturally try to map onto process RSS: &lt;code>server.memory_allocated&lt;/code>, &lt;code>server.memory_heap_size&lt;/code>, and &lt;code>server.memory_physical_size&lt;/code>. On a busy proxy the first two diverge by 2x-4x routinely, and dashboards that graph &lt;code>heap_size&lt;/code> look like a slow leak even when nothing is wrong.&lt;/p>
&lt;p>The gap is not a bug. Envoy ships with tcmalloc as its default allocator, and tcmalloc is built for allocation latency, not for prompt return of freed memory to the OS. It parks freed pages in per-thread caches, central caches, and the page heap free lists so the next allocation is fast. Those pages stay mapped into the process, counted in &lt;code>memory_heap_size&lt;/code>, and very often still resident in RSS.&lt;/p></description></item><item><title>Envoy Monitoring</title><link>https://www.netdata.cloud/monitoring-101/envoy-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/envoy-monitoring/</guid><description>&lt;h2 id="envoy-monitoring">Envoy Monitoring&lt;/h2>
&lt;h3 id="what-is-envoy">What Is Envoy?&lt;/h3>
&lt;p>Envoy is an open-source edge and service proxy, designed for cloud-native applications and microservices architectures. It plays a pivotal role in managing ingress and egress traffic between microservices, thus ensuring seamless communications within distributed systems. As a Layer 7 proxy, it provides advanced load balancing, traffic management, and observability features, crucial for high-availability applications.&lt;/p>
&lt;h3 id="monitoring-envoy-with-netdata">Monitoring Envoy With Netdata&lt;/h3>
&lt;p>When it comes to monitoring Envoy proxies, the Netdata monitoring tool provides comprehensive insights into various metrics that are key to maintaining optimal performance. Netdata&amp;rsquo;s real-time monitoring capabilities allow you to track the health of your Envoy instances, detect anomalies, and troubleshoot issues promptly. With its ease of configuration and powerful visualization features, Netdata is an ideal solution for monitoring Envoy.&lt;/p></description></item><item><title>Envoy monitoring checklist: the signals every production proxy needs</title><link>https://www.netdata.cloud/guides/envoy/envoy-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-monitoring-checklist/</guid><description>&lt;h1 id="envoy-monitoring-checklist-the-signals-every-production-proxy-needs">Envoy monitoring checklist: the signals every production proxy needs&lt;/h1>
&lt;p>This checklist organizes Envoy monitoring into four maturity levels: survival, operational, mature, and expert. Each level catches failure modes the previous level misses. The levels are cumulative. Level 2 assumes Level 1 is in place. A team alerting on outlier detection ejections without basic server liveness has gaps in the wrong direction.&lt;/p>
&lt;p>Use this as an audit tool. Walk each level, confirm each signal is collected and alerted on (or deliberately omitted), and note the gaps. Most production Envoy deployments plateau around Level 2 with a few Level 3 additions.&lt;/p></description></item><item><title>Envoy monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/envoy/envoy-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-monitoring-maturity-model/</guid><description>&lt;h1 id="envoy-monitoring-maturity-model-from-survival-to-expert">Envoy monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Envoy exposes thousands of stats and dozens of admin endpoints. The trap is not lack of data; it is lack of the right data at the right depth for your operational maturity. A team monitoring &lt;code>/ready&lt;/code> and &lt;code>membership_healthy&lt;/code> will survive most outages but will be blind to the failure modes that cause multi-hour incidents: silent xDS NACKs, retry amplification, connection pool exhaustion, CFS throttling.&lt;/p></description></item><item><title>Envoy NC no cluster: the cluster vanished mid-request</title><link>https://www.netdata.cloud/guides/envoy/envoy-nc-no-cluster/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-nc-no-cluster/</guid><description>&lt;h1 id="envoy-nc-no-cluster-the-cluster-vanished-mid-request">Envoy NC no cluster: the cluster vanished mid-request&lt;/h1>
&lt;p>A request lands in Envoy, the route matches, the router filter looks up the cluster the route points at, and the cluster manager has no cluster by that name. Envoy fast-fails the request with the &lt;code>NC&lt;/code> (No Cluster) response flag. This is not an upstream health problem or a transient network issue. Any non-zero rate of &lt;code>NC&lt;/code> in production is a configuration or timing bug.&lt;/p></description></item><item><title>Envoy no healthy upstream: the 503 when a cluster has no host to route to</title><link>https://www.netdata.cloud/guides/envoy/envoy-no-healthy-upstream/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-no-healthy-upstream/</guid><description>&lt;h1 id="envoy-no-healthy-upstream-the-503-when-a-cluster-has-no-host-to-route-to">Envoy no healthy upstream: the 503 when a cluster has no host to route to&lt;/h1>
&lt;p>When Envoy returns a 503 with body &amp;ldquo;no healthy upstream&amp;rdquo; and access log flag &lt;code>UH&lt;/code>, the cluster has zero hosts available for load balancing. The cluster exists, routes point to it, traffic is flowing, but the load balancer cannot select a host.&lt;/p>
&lt;p>The response body is literal. The flag &lt;code>UH&lt;/code> appears in the &lt;code>%RESPONSE_FLAGS%&lt;/code> access log field and means &amp;ldquo;No healthy upstream hosts in upstream cluster in addition to 503 response code.&amp;rdquo; &lt;!-- TODO: verify exact Envoy version in which UH was introduced; draft cited v1.5.0 --> It is distinct from &lt;code>UO&lt;/code> (circuit breaker overflow), &lt;code>UF&lt;/code> (upstream connection failure), and &lt;code>NR&lt;/code> (no route).&lt;/p></description></item><item><title>Envoy NR no route: 404s and 503s after a bad xDS route push</title><link>https://www.netdata.cloud/guides/envoy/envoy-nr-no-route/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-nr-no-route/</guid><description>&lt;h1 id="envoy-nr-no-route-404s-and-503s-after-a-bad-xds-route-push">Envoy NR no route: 404s and 503s after a bad xDS route push&lt;/h1>
&lt;p>The NR response flag in Envoy access logs means no route matched the request. The HTTP connection manager walked its route table (delivered via RDS or static config) and found no match for the incoming Host header, path, or other match criteria. In production, sustained NR is a configuration error. A spike right after an xDS route push points to a bad or rejected config update.&lt;/p></description></item><item><title>Envoy outlier detection mass ejection: when passive health checks empty a cluster</title><link>https://www.netdata.cloud/guides/envoy/envoy-outlier-detection-mass-ejection/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-outlier-detection-mass-ejection/</guid><description>&lt;h1 id="envoy-outlier-detection-mass-ejection-when-passive-health-checks-empty-a-cluster">Envoy outlier detection mass ejection: when passive health checks empty a cluster&lt;/h1>
&lt;p>&lt;code>membership_healthy&lt;/code> is collapsing. &lt;code>outlier_detection.ejections_active&lt;/code> is climbing toward &lt;code>membership_total&lt;/code>, and &lt;code>ejections_overflow&lt;/code> is ticking up because Envoy wanted to eject more hosts than &lt;code>max_ejection_percent&lt;/code> allows. This is the outlier detection mass ejection pattern: one of the few Envoy failure modes that can amplify a partial upstream degradation into a cluster-wide outage.&lt;/p>
&lt;p>The mechanism is subtle because outlier detection is doing exactly what it was configured to do. It is a passive health check: it ejects hosts based on the actual traffic they are serving, not on synthetic probes. When one host starts returning 5xx or its success rate drops, ejecting it is correct. The problem is what happens next. Load concentrates on the survivors, they get slower, their success rates drop, and they get ejected too. The cascade feeds itself.&lt;/p></description></item><item><title>Envoy overload manager actions: stop_accepting_connections and the last line before OOM</title><link>https://www.netdata.cloud/guides/envoy/envoy-overload-manager-actions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-overload-manager-actions/</guid><description>&lt;h1 id="envoy-overload-manager-actions-stop_accepting_connections-and-the-last-line-before-oom">Envoy overload manager actions: stop_accepting_connections and the last line before OOM&lt;/h1>
&lt;p>When &lt;code>server.overload_manager.envoy.overload_actions.stop_accepting_connections.active&lt;/code> flips to 1, Envoy has stopped accepting new TCP connections on its listeners. When &lt;code>stop_accepting_requests&lt;/code> is active, Envoy returns 503 to new HTTP requests before they reach an upstream. These are not bugs. They are Envoy&amp;rsquo;s last intentional actions before the kernel OOM-kills the process.&lt;/p>
&lt;p>The overload manager connects resource pressure (heap size, connection counts) to a cascade of protective actions. When it fires, you are looking at both a root cause (memory or connection pressure) and a symptom (traffic being refused). Many deployments never configure it, skipping every graceful degradation step and going straight from &amp;ldquo;fine&amp;rdquo; to OOM-killed with nothing in between.&lt;/p></description></item><item><title>Envoy panic threshold: why traffic routes to unhealthy hosts at 50%</title><link>https://www.netdata.cloud/guides/envoy/envoy-panic-threshold-routing-all-hosts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-panic-threshold-routing-all-hosts/</guid><description>&lt;h1 id="envoy-panic-threshold-why-traffic-routes-to-unhealthy-hosts-at-50">Envoy panic threshold: why traffic routes to unhealthy hosts at 50%&lt;/h1>
&lt;p>You are looking at an Envoy cluster where &lt;code>membership_healthy&lt;/code> has dropped below half of &lt;code>membership_total&lt;/code>. The error rate has climbed. The first instinct is that something new has broken: the health checks are wrong, outlier detection is misfiring, or the upstream has a second fault. In most cases none of that is true. Envoy has entered panic mode, and the elevated error rate is a direct and expected consequence of the design.&lt;/p></description></item><item><title>Envoy rate limiting: over_limit, 429s, and fail-open vs fail-closed</title><link>https://www.netdata.cloud/guides/envoy/envoy-rate-limiting-over-limit-429/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-rate-limiting-over-limit-429/</guid><description>&lt;h1 id="envoy-rate-limiting-over_limit-429s-and-fail-open-vs-fail-closed">Envoy rate limiting: over_limit, 429s, and fail-open vs fail-closed&lt;/h1>
&lt;p>Envoy has two HTTP filters for rate limiting with very different failure characteristics. The global rate limit filter delegates every applicable request to an external rate limit service (RLS) over gRPC. The local rate limit filter applies an in-process token bucket with no external dependency. Both can produce 429 responses, but the signals that tell you what happened live in different stat namespaces and mean different things.&lt;/p></description></item><item><title>Envoy RBAC: access denied: 403s from the RBAC filter and shadow-mode tuning</title><link>https://www.netdata.cloud/guides/envoy/envoy-rbac-access-denied/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-rbac-access-denied/</guid><description>&lt;h1 id="envoy-rbac-access-denied-403s-from-the-rbac-filter-and-shadow-mode-tuning">Envoy RBAC: access denied: 403s from the RBAC filter and shadow-mode tuning&lt;/h1>
&lt;p>A 403 with the body &lt;code>RBAC: access denied&lt;/code> is the fingerprint of Envoy&amp;rsquo;s HTTP RBAC filter blocking a request. It is not the upstream service refusing the request, and it is not the network RBAC filter, which closes the TCP connection instead of returning an HTTP status. When this counter rises you have two questions to answer: which policy matched, and was the match correct.&lt;/p></description></item><item><title>Envoy response flags: decoding UO, NR, UF, UT, UC and the rest</title><link>https://www.netdata.cloud/guides/envoy/envoy-response-flags/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-response-flags/</guid><description>&lt;h1 id="envoy-response-flags-decoding-uo-nr-uf-ut-uc-and-the-rest">Envoy response flags: decoding UO, NR, UF, UT, UC and the rest&lt;/h1>
&lt;p>A 503 from Envoy can mean a dozen different things. The HTTP status code tells you what the client saw; it does not tell you why Envoy generated that response. The &lt;code>%RESPONSE_FLAGS%&lt;/code> access log field separates a circuit breaker trip from a missing route, a dead upstream, or a client that hung up.&lt;/p>
&lt;p>Response flags are the most precise debugging signal Envoy emits, and the most operationally misunderstood. They are not aggregate stats; they appear only in access logs. Most teams discover this the first time they try to alert on &lt;code>UO&lt;/code> and find no Prometheus counter for it.&lt;/p></description></item><item><title>Envoy retry policy tuning: retry_on, budgets, idempotency, and hedging</title><link>https://www.netdata.cloud/guides/envoy/envoy-retry-policy-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-retry-policy-tuning/</guid><description>&lt;h1 id="envoy-retry-policy-tuning-retry_on-budgets-idempotency-and-hedging">Envoy retry policy tuning: retry_on, budgets, idempotency, and hedging&lt;/h1>
&lt;p>Retries are one of the few Envoy features that can make an outage measurably worse. A correctly tuned retry policy absorbs transient failures cleanly. A poorly tuned one multiplies upstream load by 2x-3x during a partial failure and accelerates the collapse it was supposed to mask. This article covers the knobs that matter operationally: &lt;code>retry_on&lt;/code> conditions, retry budgets, &lt;code>per_try_timeout&lt;/code>, and hedging. It also covers the one thing Envoy cannot decide for you: whether a request is safe to repeat.&lt;/p></description></item><item><title>Envoy retry storm: when the upstream-to-downstream request ratio climbs</title><link>https://www.netdata.cloud/guides/envoy/envoy-retry-storm-amplification/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-retry-storm-amplification/</guid><description>&lt;h1 id="envoy-retry-storm-when-the-upstream-to-downstream-request-ratio-climbs">Envoy retry storm: when the upstream-to-downstream request ratio climbs&lt;/h1>
&lt;p>A retry storm is one of the most dangerous failure modes in an Envoy-based service mesh or edge proxy. An upstream service starts failing a fraction of requests. Envoy&amp;rsquo;s retry policy, configured to mask transient errors, fires retries on the failures. The extra load lands on an already-degraded upstream. More requests fail under the additional load. More retries fire. Within minutes, the upstream receives two to three times its normal traffic, mostly retries, and collapses under amplification it cannot escape.&lt;/p></description></item><item><title>Envoy server.state not LIVE: draining, initializing, and the /ready probe</title><link>https://www.netdata.cloud/guides/envoy/envoy-server-state-not-live/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-server-state-not-live/</guid><description>&lt;h1 id="envoy-serverstate-not-live-draining-initializing-and-the-ready-probe">Envoy server.state not LIVE: draining, initializing, and the /ready probe&lt;/h1>
&lt;p>A Kubernetes readiness probe starts failing. The pod shows NotReady, the service stops sending traffic, and a rolling deploy stalls. You exec in and hit Envoy&amp;rsquo;s &lt;code>/ready&lt;/code> endpoint:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-bash" data-lang="bash">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># Check readiness state on the admin port&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>curl -s -o /dev/null -w &lt;span style="color:#e6db74">&amp;#34;%{http_code}\n&amp;#34;&lt;/span> http://localhost:9901/ready
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#ae81ff">503&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;ol start="503">
&lt;li>The probe is doing its job: Envoy is reporting &lt;code>server.state != LIVE&lt;/code>. The real question is whether this is expected (rolling deploy, hot restart, normal warm-up) or genuine capacity loss (stuck init, orphaned drain, crashed child). Non-LIVE is normal during deploys, so alerting on &lt;code>server.state != LIVE&lt;/code> alone is noisy. You need to combine it with listener state and uptime before paging.&lt;/li>
&lt;/ol>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>Envoy exposes its lifecycle as a single gauge, &lt;code>server.state&lt;/code>:&lt;/p></description></item><item><title>Envoy ssl.fail_verify_error: certificate verification failures on the TLS path</title><link>https://www.netdata.cloud/guides/envoy/envoy-tls-fail-verify-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-tls-fail-verify-error/</guid><description>&lt;h1 id="envoy-sslfail_verify_error-certificate-verification-failures-on-the-tls-path">Envoy ssl.fail_verify_error: certificate verification failures on the TLS path&lt;/h1>
&lt;p>When &lt;code>ssl.fail_verify_error&lt;/code> starts climbing on an Envoy proxy, TLS handshakes are failing peer certificate verification. A small burst during a planned rotation is operational noise. A sustained climb past roughly 1% of handshakes means the trust relationship between Envoy and its peers has broken.&lt;/p>
&lt;p>This counter is one of several SSL stats Envoy tracks, and the others help narrow the failure mode. &lt;code>ssl.connection_error&lt;/code> covers protocol-level issues such as TLS version or cipher mismatch. &lt;code>ssl.no_certificate&lt;/code> fires when a client presented no certificate where one was required. &lt;code>ssl.fail_verify_error&lt;/code> means a certificate was presented and parsed, but Envoy could not validate it against its configured trust context, SAN matchers, or pin hashes.&lt;/p></description></item><item><title>Envoy stats cardinality explosion: the stats region fills and new metrics vanish</title><link>https://www.netdata.cloud/guides/envoy/envoy-stats-cardinality-explosion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-stats-cardinality-explosion/</guid><description>&lt;h1 id="envoy-stats-cardinality-explosion-the-stats-region-fills-and-new-metrics-vanish">Envoy stats cardinality explosion: the stats region fills and new metrics vanish&lt;/h1>
&lt;p>Envoy builds each stat name from the resource names in your config: clusters, listeners, HTTP connection managers, routes. When those names are stable, cardinality is bounded. When they are dynamic, every distinct value mints a new metric. The budget that backs those stats fills, and new stats stop appearing with no error and no log line.&lt;/p>
&lt;p>The failure is quiet by design. Envoy treats stat registration as best-effort relative to its memory budget. A missing metric raises no alert, increments no hot-path error counter, and never appears in access logs. You usually discover it weeks later, when an incident sends you hunting for a per-cluster or per-route metric that was never recorded.&lt;/p></description></item><item><title>Envoy stuck initializing: waiting on xDS while the pod never turns ready</title><link>https://www.netdata.cloud/guides/envoy/envoy-stuck-initializing-warming/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-stuck-initializing-warming/</guid><description>&lt;h1 id="envoy-stuck-initializing-waiting-on-xds-while-the-pod-never-turns-ready">Envoy stuck initializing: waiting on xDS while the pod never turns ready&lt;/h1>
&lt;p>The pod sits at 0/1 Running indefinitely. &lt;code>kubectl describe&lt;/code> shows failing readiness probes, the Envoy sidecar never reports ready, and the application container may be healthy but unreachable because the sidecar gate is closed. Envoy&amp;rsquo;s &lt;code>/ready&lt;/code> admin endpoint returns HTTP 503. &lt;code>/server_info&lt;/code> reports &lt;code>PRE_INITIALIZING&lt;/code> or &lt;code>INITIALIZING&lt;/code>. The process is alive, but it never received its initial xDS configuration, so the listener sockets exist but no listeners are active.&lt;/p></description></item><item><title>Envoy symptoms from the kernel: conntrack exhaustion and CFS throttling</title><link>https://www.netdata.cloud/guides/envoy/envoy-conntrack-cfs-throttling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-conntrack-cfs-throttling/</guid><description>&lt;h1 id="envoy-symptoms-from-the-kernel-conntrack-exhaustion-and-cfs-throttling">Envoy symptoms from the kernel: conntrack exhaustion and CFS throttling&lt;/h1>
&lt;p>Envoy looks healthy. The admin endpoint returns 200, cluster membership is stable, &lt;code>upstream_rq_time&lt;/code> is in its normal band. But every few minutes a small fraction of requests fail with &lt;code>cluster.&amp;lt;name&amp;gt;.upstream_cx_connect_fail&lt;/code> ticking up, and nothing inside Envoy explains it. Upstream hosts are up, the network path looks clean, the response flag is &lt;code>UF&lt;/code> with no further detail.&lt;/p>
&lt;p>Or: &lt;code>server.watchdog_miss&lt;/code> starts incrementing during a traffic burst, tail latency spikes, and yet the cgroup CPU chart shows Envoy using well under its declared limit. No hot worker, no expensive filter, no lock contention. Envoy&amp;rsquo;s own telemetry says &amp;ldquo;I am not the problem.&amp;rdquo;&lt;/p></description></item><item><title>Envoy TLS certificate expiry: expired certs, broken SDS rotation, and total outage</title><link>https://www.netdata.cloud/guides/envoy/envoy-tls-certificate-expiry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-tls-certificate-expiry/</guid><description>&lt;h1 id="envoy-tls-certificate-expiry-expired-certs-broken-sds-rotation-and-total-outage">Envoy TLS certificate expiry: expired certs, broken SDS rotation, and total outage&lt;/h1>
&lt;p>TLS certificate expiry in Envoy is one of the few failure modes that can take down an entire service mesh at once. A single expired server certificate breaks new handshakes on one listener. An expired CA root breaks every mTLS connection in the data plane, every Envoy-to-control-plane link, and every health check that uses TLS.&lt;/p>
&lt;p>Expiry is silent until handshakes start failing. Envoy does not expose certificate expiration as a standard metric. The only authoritative source is the &lt;code>/certs&lt;/code> admin endpoint, which must be polled externally. Teams running automated rotation via SDS or cert-manager often assume rotation is working and discover otherwise when TLS breaks across the fleet.&lt;/p></description></item><item><title>Envoy update_rejected (xDS NACK): the config the control plane pushed and Envoy quietly refused</title><link>https://www.netdata.cloud/guides/envoy/envoy-update-rejected-nack/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-update-rejected-nack/</guid><description>&lt;h1 id="envoy-update_rejected-xds-nack-the-config-the-control-plane-pushed-and-envoy-quietly-refused">Envoy update_rejected (xDS NACK): the config the control plane pushed and Envoy quietly refused&lt;/h1>
&lt;p>You deploy a config change through your control plane. The deployment reports success. Traffic flows, error rates are flat, &lt;code>control_plane.connected_state&lt;/code> stays at 1. But the change never took effect: the new route is missing, the timeout override is gone, the certificate rotation did not happen.&lt;/p>
&lt;p>Envoy received the new config, validated it, found it invalid, and rejected it. It kept the previous config and kept routing. The rejection is recorded in &lt;code>update_rejected&lt;/code> and, for listeners, in &lt;code>listener_manager.listener_create_failure&lt;/code>. The reason itself is only in the Envoy process log, not in stats. The control plane reports success because the gRPC stream is healthy, not because Envoy accepted the config. Nothing fails visibly. The only symptoms are the things that should have changed but did not.&lt;/p></description></item><item><title>Envoy upstream connect error or disconnect/reset before headers: reading the reset reason</title><link>https://www.netdata.cloud/guides/envoy/envoy-upstream-connect-error-disconnect-reset-before-headers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-upstream-connect-error-disconnect-reset-before-headers/</guid><description>&lt;p>The 503 body &lt;code>upstream connect error or disconnect/reset before headers. reset reason: &amp;lt;reason&amp;gt;&lt;/code> means Envoy selected an upstream host, tried to open a request stream on a pooled connection, and the stream was reset before any response headers came back. The body looks generic, but the trailing &lt;code>reset reason:&lt;/code> field is the diagnostic payload. It is one of the values in Envoy&amp;rsquo;s &lt;code>StreamResetReason&lt;/code> enum &lt;!-- TODO: verify the current count (the draft says nine); recent Envoy versions may add OverloadManager or others -->, and each value points at a different failure mechanism.&lt;/p></description></item><item><title>Envoy upstream_cx_active near max_connections: the pool filling up</title><link>https://www.netdata.cloud/guides/envoy/envoy-upstream-cx-active-max-connections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-upstream-cx-active-max-connections/</guid><description>&lt;h1 id="envoy-upstream_cx_active-near-max_connections-the-pool-filling-up">Envoy upstream_cx_active near max_connections: the pool filling up&lt;/h1>
&lt;p>The &lt;code>cluster.&amp;lt;name&amp;gt;.upstream_cx_active&lt;/code> gauge reports active connections across all hosts in a cluster. When it trends toward the circuit-breaker &lt;code>max_connections&lt;/code> limit, the cluster is approaching saturation. Past the limit, Envoy stops opening new connections, queues new requests in &lt;code>upstream_rq_pending_active&lt;/code>, and once &lt;code>max_pending_requests&lt;/code> is hit, returns 503s with response flag &lt;code>UO&lt;/code>.&lt;/p>
&lt;p>Envoy is fast-failing locally to protect an upstream that cannot absorb more concurrent work. Raising &lt;code>max_connections&lt;/code> without addressing the upstream removes that protection. The operator&amp;rsquo;s job is to find what is shrinking effective pool capacity.&lt;/p></description></item><item><title>Envoy upstream_cx_connect_fail: failed TCP connections to upstream hosts</title><link>https://www.netdata.cloud/guides/envoy/envoy-upstream-cx-connect-fail/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-upstream-cx-connect-fail/</guid><description>&lt;h1 id="envoy-upstream_cx_connect_fail-failed-tcp-connections-to-upstream-hosts">Envoy upstream_cx_connect_fail: failed TCP connections to upstream hosts&lt;/h1>
&lt;p>&lt;code>cluster.&amp;lt;name&amp;gt;.upstream_cx_connect_fail&lt;/code> is a per-cluster counter that increments each time Envoy fails to establish a TCP connection to an upstream host. On a healthy cluster the rate is flat. The threshold for concern is &lt;code>connect_fail / connect_total &amp;gt; 0.05&lt;/code> &lt;!-- TODO: verify whether upstream_cx_total counts only established connections or all attempts; if the former, denominator should be connect_total + connect_fail -->, meaning more than 5% of connection attempts are failing. Sustained nonzero rates point at one of four root causes: the upstream process is down, the upstream is out of file descriptors, a firewall or ACL changed, or the upstream listen backlog is overflowing.&lt;/p></description></item><item><title>Envoy upstream_cx_connect_ms high: slow TCP connects to upstream hosts</title><link>https://www.netdata.cloud/guides/envoy/envoy-upstream-cx-connect-ms-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-upstream-cx-connect-ms-high/</guid><description>&lt;h1 id="envoy-upstream_cx_connect_ms-high-slow-tcp-connects-to-upstream-hosts">Envoy upstream_cx_connect_ms high: slow TCP connects to upstream hosts&lt;/h1>
&lt;p>When &lt;code>cluster.&amp;lt;name&amp;gt;.upstream_cx_connect_ms&lt;/code> starts climbing, understand what this histogram actually measures: the time Envoy spent establishing a TCP connection to an upstream host. When upstream TLS is configured, the TLS handshake is rolled into the same number. It is not request latency, not application processing time, and not pure network RTT once TLS is in the path.&lt;/p>
&lt;p>Same-zone connections typically sit under 2ms at P99. The alert threshold is a sudden 5x rise over rolling baseline. Anything beyond that means the network path has changed, the upstream kernel cannot accept connections fast enough, or something is forcing Envoy to open far more new connections than usual.&lt;/p></description></item><item><title>Envoy upstream_rq_pending_overflow: the pending queue fills and 503s begin</title><link>https://www.netdata.cloud/guides/envoy/envoy-upstream-rq-pending-overflow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-upstream-rq-pending-overflow/</guid><description>&lt;h1 id="envoy-upstream_rq_pending_overflow-the-pending-queue-fills-and-503s-begin">Envoy upstream_rq_pending_overflow: the pending queue fills and 503s begin&lt;/h1>
&lt;p>You see &lt;code>cluster.&amp;lt;name&amp;gt;.upstream_rq_pending_overflow&lt;/code> incrementing on one or more clusters. Access logs show 503 with response flag &lt;code>UO&lt;/code>. Clients get fast-fail 503s, not timeouts. The curve is cliff-edge: requests were flowing fine, and now a slice of them are rejected with no upstream attempt at all.&lt;/p>
&lt;p>This is Envoy&amp;rsquo;s pending request circuit breaker doing exactly what it was configured to do. Requests arrive, there is no available upstream connection to attach them to, so they wait in &lt;code>upstream_rq_pending_active&lt;/code>. When that queue hits &lt;code>max_pending_requests&lt;/code>, Envoy stops queueing and starts rejecting. The overflow counter is the rejected count.&lt;/p></description></item><item><title>Envoy upstream_rq_retry_overflow: the retry budget exhausted and retries dropped</title><link>https://www.netdata.cloud/guides/envoy/envoy-upstream-rq-retry-overflow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-upstream-rq-retry-overflow/</guid><description>&lt;h1 id="envoy-upstream_rq_retry_overflow-the-retry-budget-exhausted-and-retries-dropped">Envoy upstream_rq_retry_overflow: the retry budget exhausted and retries dropped&lt;/h1>
&lt;p>&lt;code>upstream_rq_retry_overflow&lt;/code> is climbing on a cluster. Error rates are elevated, and retries should absorb the failures. But &lt;code>upstream_rq_retry&lt;/code> is not growing proportionally, and the system looks like it stopped retrying. It did: the retry circuit breaker or retry budget is full, and Envoy is dropping retries to protect the upstream.&lt;/p>
&lt;p>The retry system is working as designed, but legitimate retries are being dropped at the moment they are needed most. A budget that is too small loses the resilience retries provide. A budget that is too large risks 2x-3x traffic amplification during partial failures.&lt;/p></description></item><item><title>Envoy upstream_rq_time high: upstream latency as the proxy sees it</title><link>https://www.netdata.cloud/guides/envoy/envoy-upstream-rq-time-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-upstream-rq-time-high/</guid><description>&lt;h1 id="envoy-upstream_rq_time-high-upstream-latency-as-the-proxy-sees-it">Envoy upstream_rq_time high: upstream latency as the proxy sees it&lt;/h1>
&lt;p>&lt;code>cluster.&amp;lt;name&amp;gt;.upstream_rq_time&lt;/code> is the histogram that answers &amp;ldquo;how long did the upstream interaction take, as Envoy observed it?&amp;rdquo; When it climbs, the instinct is to page the backend team. That instinct is often wrong, or at least incomplete. The metric is wall-clock time measured at the HTTP router filter. It bundles several distinct latencies: TCP connect, upstream TLS handshake (for new connections), request transmission, upstream processing, and response transfer. It is not pure backend service time.&lt;/p></description></item><item><title>Envoy URX upstream retry limit exceeded: all retries used, last error returned</title><link>https://www.netdata.cloud/guides/envoy/envoy-urx-upstream-retry-limit-exceeded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-urx-upstream-retry-limit-exceeded/</guid><description>&lt;h1 id="envoy-urx-upstream-retry-limit-exceeded-all-retries-used-last-error-returned">Envoy URX upstream retry limit exceeded: all retries used, last error returned&lt;/h1>
&lt;p>When you see &lt;code>URX&lt;/code> in Envoy access logs, every configured retry attempt has fired and failed, and the last upstream error is what the client receives. This is distinct from &lt;code>retry_overflow&lt;/code>, where retries were never attempted because the retry budget or circuit breaker was full. The distinction matters because the two conditions have opposite root causes and opposite fixes.&lt;/p></description></item><item><title>Envoy watchdog_miss and the hot worker: one saturated event loop hiding in the average</title><link>https://www.netdata.cloud/guides/envoy/envoy-watchdog-miss-hot-worker/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-watchdog-miss-hot-worker/</guid><description>&lt;p>&lt;code>server.watchdog_miss&lt;/code> just incremented. Aggregate CPU shows Envoy at 40%, comfortably under capacity. But a slice of requests are spiking P99, P50 is fine, and latency variance is high.&lt;/p>
&lt;p>This is the hot worker pattern. Envoy is single-threaded per worker: each worker owns its connections for their entire lifetime and runs its own event loop. When one worker&amp;rsquo;s event loop blocks past the watchdog timeout (200ms by default), the main thread increments &lt;code>watchdog_miss&lt;/code>. A single saturated worker is invisible in process-level CPU averages because the other workers idle along.&lt;/p></description></item><item><title>EOS</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/eos/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/eos/</guid><description/></item><item><title>EOS Monitoring</title><link>https://www.netdata.cloud/monitoring-101/eos_web-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/eos_web-monitoring/</guid><description>&lt;h2 id="eos-monitoring">EOS Monitoring&lt;/h2>
&lt;h3 id="what-is-eos">What Is EOS?&lt;/h3>
&lt;p>EOS is a high-performance, scalable storage service developed by CERN. It is primarily used for managing massive volumes of scientific data, providing a reliable and efficient storage solution that ensures data integrity and accessibility.&lt;/p>
&lt;h3 id="monitoring-eos-with-netdata">Monitoring EOS With Netdata&lt;/h3>
&lt;p>To ensure optimal performance and reliability of EOS, monitoring it with Netdata can provide real-time insights into system metrics and help troubleshoot issues proactively. Netdata uses an openmetrics (Prometheus) exporter to collect EOS metrics efficiently. With Netdata, you can ingest data from any Prometheus exporter, offering you automated dashboards and alerts without the need for a Prometheus server or Grafana. This makes Netdata an ideal EOS monitoring tool that simplifies the monitoring process while delivering comprehensive metrics insights.&lt;/p></description></item><item><title>Epicenter Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/epicenter-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/epicenter-inc-snmp-traps/</guid><description/></item><item><title>Equallogic SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/equallogic-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/equallogic-snmp-traps/</guid><description/></item><item><title>Equinox Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/equinox-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/equinox-systems-inc-snmp-traps/</guid><description/></item><item><title>Equipe Communications Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/equipe-communications-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/equipe-communications-corporation-snmp-traps/</guid><description/></item><item><title>Era A S SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/era-a-s-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/era-a-s-snmp-traps/</guid><description/></item><item><title>Ericsson AB Packet Core Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ericsson-ab-packet-core-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ericsson-ab-packet-core-networks-snmp-traps/</guid><description/></item><item><title>Ericsson AB SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ericsson-ab-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ericsson-ab-snmp-traps/</guid><description/></item><item><title>Ericsson Inc Formerly Redback Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ericsson-inc-formerly-redback-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ericsson-inc-formerly-redback-networks-snmp-traps/</guid><description/></item><item><title>Essential Communications SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/essential-communications-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/essential-communications-snmp-traps/</guid><description/></item><item><title>etcd</title><link>https://www.netdata.cloud/integrations/data-collection/applications/etcd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/etcd/</guid><description/></item><item><title>etcd Monitoring</title><link>https://www.netdata.cloud/monitoring-101/etcd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/etcd-monitoring/</guid><description>&lt;h2 id="etcd-monitoring">etcd Monitoring&lt;/h2>
&lt;h3 id="what-is-etcd">What Is etcd?&lt;/h3>
&lt;p>etcd, a distributed key-value store, is a critical component for service discovery and storing all of a cluster&amp;rsquo;s data reliably. It ensures that the systems in production environments are available and consistent. etcd is often implemented in environments requiring fault tolerance, distributed networking, or consensus building. &lt;a href="https://etcd.io/">Learn more about etcd.&lt;/a>&lt;/p>
&lt;h3 id="monitoring-etcd-with-netdata">Monitoring etcd With Netdata&lt;/h3>
&lt;p>When it comes to monitoring etcd, Netdata stands out as a powerful tool. Netdata leverages an OpenMetrics (Prometheus) exporter to monitor etcd, facilitating effortless integration. This capability permits Netdata to ingest metrics from any Prometheus exporter, delivering dynamic dashboards, alerts, and comprehensive insights that require neither a dedicated Prometheus server nor Grafana setup.&lt;/p></description></item><item><title>Etherwan Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/etherwan-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/etherwan-systems-inc-snmp-traps/</guid><description/></item><item><title>Eurologic Systems Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/eurologic-systems-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/eurologic-systems-ltd-snmp-traps/</guid><description/></item><item><title>Exablaze SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/exablaze-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/exablaze-snmp-traps/</guid><description/></item><item><title>Exabyte Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/exabyte-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/exabyte-corporation-snmp-traps/</guid><description/></item><item><title>Exagrid</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/exagrid/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/exagrid/</guid><description/></item><item><title>Exalt Communications SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/exalt-communications-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/exalt-communications-snmp-traps/</guid><description/></item><item><title>Exanet SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/exanet-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/exanet-snmp-traps/</guid><description/></item><item><title>Exceliance SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/exceliance-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/exceliance-snmp-traps/</guid><description/></item><item><title>Exim</title><link>https://www.netdata.cloud/integrations/data-collection/applications/exim/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/exim/</guid><description/></item><item><title>Exim Monitoring</title><link>https://www.netdata.cloud/monitoring-101/exim-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/exim-monitoring/</guid><description>&lt;h2 id="exim-monitoring">Exim Monitoring&lt;/h2>
&lt;h3 id="what-is-exim">What Is Exim?&lt;/h3>
&lt;p>Exim is a mail transfer agent (MTA) used on Unix-like operating systems. It is highly configurable and can handle email routing, spam control, and address verification. For more details, you can visit the &lt;a href="https://www.exim.org/">Exim website&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-exim-with-netdata">Monitoring Exim With Netdata&lt;/h3>
&lt;p>Monitoring Exim&amp;rsquo;s performance is crucial to ensure the optimal operation of the mail server. With Netdata, you can monitor Exim effortlessly using the go.d.plugin, which tracks the Exim mail queue. This integration provides real-time insights, helping you troubleshoot and maintain reliable email service. Learn more in the &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/exim/?utm_source=website&amp;amp;utm_content=monitoring101">Exim collector documentation&lt;/a>.&lt;/p></description></item><item><title>Exinda Networks Pty Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/exinda-networks-pty-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/exinda-networks-pty-ltd-snmp-traps/</guid><description/></item><item><title>Expand Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/expand-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/expand-networks-inc-snmp-traps/</guid><description/></item><item><title>Extended Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/extended-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/extended-systems-inc-snmp-traps/</guid><description/></item><item><title>Extrahop Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/extrahop-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/extrahop-networks-inc-snmp-traps/</guid><description/></item><item><title>Extreme Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/extreme-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/extreme-networks-snmp-traps/</guid><description/></item><item><title>Extreme Switching</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/extreme-switching/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/extreme-switching/</guid><description/></item><item><title>Extricomltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/extricomltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/extricomltd-snmp-traps/</guid><description/></item><item><title>F5 BGP</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/f5-bgp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/f5-bgp/</guid><description/></item><item><title>F5 BIG IP</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/f5-big-ip/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/f5-big-ip/</guid><description/></item><item><title>F5 Labs Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/f5-labs-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/f5-labs-inc-snmp-traps/</guid><description/></item><item><title>F5 Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/f5-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/f5-networks-inc-snmp-traps/</guid><description/></item><item><title>Fail2ban</title><link>https://www.netdata.cloud/integrations/data-collection/applications/fail2ban/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/fail2ban/</guid><description/></item><item><title>Fail2ban Monitoring</title><link>https://www.netdata.cloud/monitoring-101/fail2ban-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/fail2ban-monitoring/</guid><description>&lt;h2 id="fail2ban-monitoring">Fail2ban Monitoring&lt;/h2>
&lt;h3 id="what-is-fail2ban">What Is Fail2ban?&lt;/h3>
&lt;p>Fail2ban is an open-source intrusion prevention software framework that protects servers from brute-force attacks. It monitors log files and bans IPs that exhibit malicious behavior, such as too many failed login attempts, by modifying firewall rules.&lt;/p>
&lt;h3 id="monitoring-fail2ban-with-netdata">Monitoring Fail2ban With Netdata&lt;/h3>
&lt;p>Netdata is a powerful monitoring solution that offers real-time insights into the performance and security of your systems. With Netdata, you can monitor Fail2ban, ensuring your servers remain secured from unwanted access. &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">Check out the Live Demo&lt;/a> to see Netdata&amp;rsquo;s capabilities in action.&lt;/p></description></item><item><title>Fair Usage Policy</title><link>https://www.netdata.cloud/fair-usage-policy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/fair-usage-policy/</guid><description/></item><item><title>Fastd</title><link>https://www.netdata.cloud/integrations/data-collection/networking/fastd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/fastd/</guid><description/></item><item><title>Fastd Monitoring</title><link>https://www.netdata.cloud/monitoring-101/fastd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/fastd-monitoring/</guid><description>&lt;h2 id="fastd-monitoring">Fastd Monitoring&lt;/h2>
&lt;h3 id="what-is-fastd">What Is Fastd?&lt;/h3>
&lt;p>Fastd, or Fast and Secure Tunneling Daemon, is a VPN solution renowned for its simplicity and flexibility. It&amp;rsquo;s widely used in various network environments, particularly in community wireless networks. Fastd enables encrypted internet connections and is known for its efficient resource usage and adaptability across different platforms.&lt;/p>
&lt;h3 id="monitoring-fastd-with-netdata">Monitoring Fastd With Netdata&lt;/h3>
&lt;p>To monitor Fastd, Netdata employs an &lt;a href="https://github.com/freifunk-darmstadt/fastd-exporter">openmetrics (Prometheus) exporter&lt;/a>. With Netdata, you can ingest data from any Prometheus exporter, offering automated dashboards, real-time alerts, and more, without the need for a Prometheus server or Grafana. This seamless integration makes it a powerful Fastd monitoring tool, providing comprehensive insights into your VPN&amp;rsquo;s performance and health.&lt;/p></description></item><item><title>FDB / MAC Forwarding Topology</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/fdb---mac-forwarding-topology/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/fdb---mac-forwarding-topology/</guid><description/></item><item><title>Fedora</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/fedora/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/fedora/</guid><description/></item><item><title>Fial Computer Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fial-computer-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fial-computer-inc-snmp-traps/</guid><description/></item><item><title>Fiberhome Telecommunication Technologies Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fiberhome-telecommunication-technologies-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fiberhome-telecommunication-technologies-co-ltd-snmp-traps/</guid><description/></item><item><title>Fibernet International SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fibernet-international-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fibernet-international-snmp-traps/</guid><description/></item><item><title>Fibrolan SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fibrolan-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fibrolan-snmp-traps/</guid><description/></item><item><title>Fibronics SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fibronics-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fibronics-snmp-traps/</guid><description/></item><item><title>Files and directories</title><link>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/files-and-directories/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/files-and-directories/</guid><description/></item><item><title>Files and Directories Monitoring</title><link>https://www.netdata.cloud/monitoring-101/filecheck-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/filecheck-monitoring/</guid><description>&lt;h2 id="files-and-directories-monitoring">Files and Directories Monitoring&lt;/h2>
&lt;h3 id="what-is-files-and-directories-monitoring">What Is Files and Directories Monitoring?&lt;/h3>
&lt;p>Files and directories are integral components of any computing system. Monitoring these elements involves keeping track of their existence, modification times, sizes, and other changes. Effective monitoring ensures data integrity, improves security, and supports efficient system maintenance.&lt;/p>
&lt;h3 id="monitoring-files-and-directories-with-netdata">Monitoring Files and Directories With Netdata&lt;/h3>
&lt;p>Netdata provides an intuitive and powerful files monitoring tool leveraging its &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/filecheck/?utm_source=website&amp;amp;utm_content=monitoring101">Filecheck module&lt;/a>. This tool simplifies the process of gathering key metrics in real time. With Netdata, you can visualize data from multiple files and directories across different servers, offering a comprehensive view of your infrastructure.&lt;/p></description></item><item><title>Fireeye</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/fireeye/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/fireeye/</guid><description/></item><item><title>Fireeye Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fireeye-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fireeye-inc-snmp-traps/</guid><description/></item><item><title>First-observation baselining: don't page on lifetime counters at rollout</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-first-observation-baseline/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-first-observation-baseline/</guid><description>&lt;h1 id="first-observation-baselining-dont-page-on-lifetime-counters-at-rollout">First-observation baselining: don&amp;rsquo;t page on lifetime counters at rollout&lt;/h1>
&lt;p>When you deploy SMART monitoring on an existing fleet for the first time, every in-service drive carries accumulated history in its lifetime counters. Offline Uncorrectable sectors, NVMe Media and Data Integrity Errors, Power-On Hours, Unsafe Shutdowns, Reallocated Sector Count, UDMA CRC Error Count. These counters started incrementing the moment the drive left the factory and never reset.&lt;/p>
&lt;p>If your alerting fires on any non-zero absolute value, every drive with any history pages within minutes of enabling monitoring. A drive with 3 reallocated sectors from factory QA, a drive with 200 uncorrectable errors from a past thermal event three years ago, a drive with 5 unsafe shutdowns from a UPS failure. All produce identical alerts to a naive greater-than-zero rule.&lt;/p></description></item><item><title>Flarion Technologies SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/flarion-technologies-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/flarion-technologies-snmp-traps/</guid><description/></item><item><title>Flock</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/flock/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/flock/</guid><description/></item><item><title>Flow export-to-ingest latency: why your NetFlow data is minutes behind</title><link>https://www.netdata.cloud/guides/network/network-flow-export-ingest-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-flow-export-ingest-latency/</guid><description>&lt;h1 id="flow-export-to-ingest-latency-why-your-netflow-data-is-minutes-behind">Flow export-to-ingest latency: why your NetFlow data is minutes behind&lt;/h1>
&lt;p>Flow export-to-ingest latency accumulates across a pipeline: the exporter&amp;rsquo;s active timeout, the UDP transport path, the kernel socket buffer, the collector&amp;rsquo;s parser, and the storage write queue. Each stage can add seconds or minutes, and each has a different fix.&lt;/p>
&lt;p>The most common cause is the active timeout default on most network devices: 30 minutes. Long-lived flows (VPN tunnels, database connections, bulk transfers) are not exported until the timer expires. The collector is not slow and the network is not congested. The device is behaving as configured. But if you need near-real-time visibility, a 30-minute export delay is indistinguishable from broken telemetry.&lt;/p></description></item><item><title>Fluentd</title><link>https://www.netdata.cloud/integrations/data-collection/applications/fluentd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/fluentd/</guid><description/></item><item><title>Fluentd average flush time rising: the earliest sign of destination slowdown</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-flush-time-rising/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-flush-time-rising/</guid><description>&lt;h1 id="fluentd-average-flush-time-rising-the-earliest-sign-of-destination-slowdown">Fluentd average flush time rising: the earliest sign of destination slowdown&lt;/h1>
&lt;p>Your Fluentd process is alive. &lt;code>retry_count&lt;/code> is zero. The buffer queue looks flat. And yet, if you are computing average flush time from &lt;code>flush_time_count&lt;/code> and &lt;code>write_count&lt;/code>, you can see the destination getting slower, sometimes hours before the first retry fires. This is the earliest signal of destination or network degradation in a Fluentd pipeline, and most teams never look at it because the monitor_agent API does not expose it as a ready-made field.&lt;/p></description></item><item><title>Fluentd broken pipe / connection reset: dropped output connections and LB timeouts</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-broken-pipe-connection-reset/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-broken-pipe-connection-reset/</guid><description>&lt;h1 id="fluentd-broken-pipe--connection-reset-dropped-output-connections-and-lb-timeouts">Fluentd broken pipe / connection reset: dropped output connections and LB timeouts&lt;/h1>
&lt;p>Your Fluentd logs show &lt;code>broken pipe&lt;/code> or &lt;code>Connection reset by peer&lt;/code> during flush, retries start climbing, and chunks roll back into the buffer queue. The destination is not down. It answers health checks, other clients reach it fine, and the errors come and go in a pattern that looks almost random.&lt;/p>
&lt;p>The usual explanation: Fluentd&amp;rsquo;s output plugin is holding a long-lived TCP connection to the destination, something in the middle (a load balancer, NAT gateway, firewall, or the destination itself) has an idle timeout, and it silently drops the connection after a period of inactivity. Fluentd does not find out until the next flush writes into the dead socket. The kernel returns EPIPE (&amp;ldquo;broken pipe&amp;rdquo;) or the peer returns RST (&amp;ldquo;connection reset by peer&amp;rdquo;), and the chunk goes back for retry.&lt;/p></description></item><item><title>Fluentd buffer available space low: computing time-to-overflow before it fires</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-buffer-available-space-low/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-buffer-available-space-low/</guid><description>&lt;h1 id="fluentd-buffer-available-space-low-computing-time-to-overflow-before-it-fires">Fluentd buffer available space low: computing time-to-overflow before it fires&lt;/h1>
&lt;p>The monitor_agent metric &lt;code>buffer_available_buffer_space_ratios&lt;/code> tells you what percentage of an output plugin&amp;rsquo;s configured buffer capacity is still free. When it drops, the buffer is filling. When it hits zero, the &lt;code>overflow_action&lt;/code> fires: with the default &lt;code>throw_exception&lt;/code>, new events are rejected at the input with a BufferOverflowError; with &lt;code>block&lt;/code>, input threads stall; with &lt;code>drop_oldest_chunk&lt;/code>, your oldest undelivered data is thrown away. None of those are outcomes you want to discover after the fact.&lt;/p></description></item><item><title>Fluentd buffer queue length growing: the output cannot keep pace with the input</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-buffer-queue-length-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-buffer-queue-length-growing/</guid><description>&lt;h1 id="fluentd-buffer-queue-length-growing-the-output-cannot-keep-pace-with-the-input">Fluentd buffer queue length growing: the output cannot keep pace with the input&lt;/h1>
&lt;p>You are looking at a graph of &lt;code>buffer_queue_length&lt;/code> for one of your Fluentd outputs and it has been climbing for twenty minutes. No alerts have fired, nothing has crashed, but the trend only goes one direction. This is the earliest visible stage of the most common Fluentd failure mode: the output destination is falling behind the input, and the buffer is absorbing the difference.&lt;/p></description></item><item><title>Fluentd buffer stage vs queue: telling healthy batching from backpressure</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-stage-vs-queue/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-stage-vs-queue/</guid><description>&lt;h1 id="fluentd-buffer-stage-vs-queue-telling-healthy-batching-from-backpressure">Fluentd buffer stage vs queue: telling healthy batching from backpressure&lt;/h1>
&lt;p>A common mistake in Fluentd operations is treating &amp;ldquo;the buffer&amp;rdquo; as a single number. Teams graph one buffer metric, see it climb, and either panic over normal batching or ignore a real backpressure signal because &amp;ldquo;the buffer number looks like it always does.&amp;rdquo; The buffer is not one thing. It is two distinct states with two distinct gauges, and they mean opposite things.&lt;/p></description></item><item><title>Fluentd buffer_oldest_timekey lag: how far behind the oldest buffered data is</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-buffer-oldest-timekey-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-buffer-oldest-timekey-lag/</guid><description>&lt;h1 id="fluentd-buffer_oldest_timekey-lag-how-far-behind-the-oldest-buffered-data-is">Fluentd buffer_oldest_timekey lag: how far behind the oldest buffered data is&lt;/h1>
&lt;p>&lt;code>buffer_oldest_timekey&lt;/code> is a timestamp: the timekey of the oldest chunk still sitting in an output plugin&amp;rsquo;s buffer, waiting to be delivered. Subtract it from the current time and you get a number that answers the question operators care about during an incident: how old is the oldest data that has not reached its destination yet?&lt;/p>
&lt;p>That number is more honest than the alternatives. Comparing input and output &lt;code>emit_records&lt;/code> rates works for steady streams, but it misleads you for time-sliced outputs (which legitimately hold chunks until the slice expires) and for bursty workloads (where rates oscillate and any short window looks alarming). The oldest timekey does not care about rates. It tells you directly how stale the data is.&lt;/p></description></item><item><title>Fluentd BufferOverflowError: buffer space has too many data</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-buffer-overflow-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-buffer-overflow-error/</guid><description>&lt;h1 id="fluentd-bufferoverflowerror-buffer-space-has-too-many-data">Fluentd BufferOverflowError: buffer space has too many data&lt;/h1>
&lt;p>Fluentd just started logging &lt;code>BufferOverflowError: buffer space has too many data&lt;/code> and your destination stopped receiving events. The buffer has reached &lt;code>total_limit_size&lt;/code> (512MB by default for memory buffers, 64GB for file buffers) and the default &lt;code>overflow_action&lt;/code>, &lt;code>throw_exception&lt;/code>, is rejecting every new event at the input. Those events are gone. They do not retry, they do not queue, and no error counter reliably tracks them.&lt;/p></description></item><item><title>Fluentd config integrity: leaked credentials and silent tampering</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-config-integrity/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-config-integrity/</guid><description>&lt;h1 id="fluentd-config-integrity-leaked-credentials-and-silent-tampering">Fluentd config integrity: leaked credentials and silent tampering&lt;/h1>
&lt;p>A Fluentd config file is not just routing logic. In most production deployments it is also a credential store: Elasticsearch passwords, S3 access keys, Kafka SASL secrets, TLS client certificates, and forward shared keys all live in the same file that tells Fluentd where your logs go. If that file is world-readable, every local user and every compromised process on the host can read those credentials. If it is writable outside the deploy pipeline, an attacker can redirect your log stream to their own endpoint, quietly drop collection for the services they are touching, or add an output you never approved.&lt;/p></description></item><item><title>Fluentd config reload failed: SIGHUP that partially applies</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-config-reload-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-config-reload-failed/</guid><description>&lt;h1 id="fluentd-config-reload-failed-sighup-that-partially-applies">Fluentd config reload failed: SIGHUP that partially applies&lt;/h1>
&lt;p>You edited the Fluentd config, sent SIGHUP, watched the log line saying the config reloaded, and moved on. Hours later you notice a new output never started receiving data, or an old filter is still dropping events you told it to keep. The process never crashed. No error fired. But the pipeline running in memory is not the pipeline in the config file.&lt;/p></description></item><item><title>Fluentd CPU bottleneck: the Ruby GVL caps a single worker at one core</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-cpu-gvl-bottleneck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-cpu-gvl-bottleneck/</guid><description>&lt;h1 id="fluentd-cpu-bottleneck-the-ruby-gvl-caps-a-single-worker-at-one-core">Fluentd CPU bottleneck: the Ruby GVL caps a single worker at one core&lt;/h1>
&lt;p>Fluentd throughput has flatlined. The host has idle cores, the destination is healthy, &lt;code>retry_count&lt;/code> is zero, and yet the buffer is growing and &lt;code>in_tail&lt;/code> is falling behind the files it watches. &lt;code>top&lt;/code> shows the Fluentd process pinned at 100% CPU, but 100% of exactly one core.&lt;/p>
&lt;p>This is the CRuby Global VM Lock (GVL) doing what it is designed to do. Within a single Fluentd worker process, only one Ruby thread can execute Ruby code at a time. The event router, parsers, filters, and flush threads all compete for that lock. The moment one thread does sustained CPU-bound work, such as regex parsing or JSON serialization, everything else in the process waits.&lt;/p></description></item><item><title>Fluentd CrashLoopBackOff: rapid restart cycling in Kubernetes</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-crashloopbackoff/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-crashloopbackoff/</guid><description>&lt;h1 id="fluentd-crashloopbackoff-rapid-restart-cycling-in-kubernetes">Fluentd CrashLoopBackOff: rapid restart cycling in Kubernetes&lt;/h1>
&lt;p>A Fluentd DaemonSet pod in CrashLoopBackOff is not &amp;ldquo;down.&amp;rdquo; It is oscillating: the container starts, runs for seconds or minutes, dies, and kubelet restarts it with an exponentially growing backoff. &lt;code>kubectl get pods&lt;/code> shows the pod flapping between Running, Error, and CrashLoopBackOff while the restart count climbs.&lt;/p>
&lt;p>Because kubelet keeps resurrecting the process, you rarely see one long outage. Instead you see repeated brief absences: buffers drain partially and refill, and downstream destinations see gaps and duplicate windows as file-backed buffers replay on each restart. The auto-restart that keeps the node &amp;ldquo;mostly covered&amp;rdquo; is also what masks the root cause.&lt;/p></description></item><item><title>Fluentd drop_oldest_chunk_count incrementing: confirmed buffer data loss</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-drop-oldest-chunk/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-drop-oldest-chunk/</guid><description>&lt;h1 id="fluentd-drop_oldest_chunk_count-incrementing-confirmed-buffer-data-loss">Fluentd drop_oldest_chunk_count incrementing: confirmed buffer data loss&lt;/h1>
&lt;p>&lt;code>drop_oldest_chunk_count&lt;/code> is not a warning signal. It is a receipt for data that no longer exists. Every increment means Fluentd discarded the oldest buffered chunk to make room for incoming events, and those events are permanently gone. No retry, no secondary output, no replay will bring them back.&lt;/p>
&lt;p>This counter only moves when you have explicitly configured &lt;code>overflow_action drop_oldest_chunk&lt;/code> in an output&amp;rsquo;s &lt;code>&amp;lt;buffer&amp;gt;&lt;/code> section. That setting is a deliberate trade: keep the pipeline alive under sustained output failure by sacrificing the oldest data first. The failure mode is that nothing else in the system complains when it fires. Fluentd keeps running, input keeps flowing, output keeps writing. The only evidence of loss is this counter and a warning line in Fluentd&amp;rsquo;s own log, which most pipelines never scrape.&lt;/p></description></item><item><title>Fluentd duplicate events: why the same log shows up twice downstream</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-duplicate-events/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-duplicate-events/</guid><description>&lt;h1 id="fluentd-duplicate-events-why-the-same-log-shows-up-twice-downstream">Fluentd duplicate events: why the same log shows up twice downstream&lt;/h1>
&lt;p>You query Elasticsearch, S3, or your log backend and the same event appears twice. Sometimes the duplication is a burst that lines up exactly with a Fluentd restart or a pod reschedule. Sometimes it is a slow, persistent trickle that inflates dashboards and breaks counts. The two situations have different root causes and different fixes.&lt;/p>
&lt;p>Fluentd does not guarantee exactly-once delivery. It guarantees at-least-once for file-backed buffers, and the brief duplication window after a restart is by design. Persistent or large-scale duplication almost always traces back to one of a small set of mechanisms: position tracking resets at the input, or full-chunk retries at the output after a partial write.&lt;/p></description></item><item><title>Fluentd emit_error_count: the number-one under-monitored data-loss signal</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-emit-error-count/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-emit-error-count/</guid><description>&lt;h1 id="fluentd-emit_error_count-the-number-one-under-monitored-data-loss-signal">Fluentd emit_error_count: the number-one under-monitored data-loss signal&lt;/h1>
&lt;p>Most Fluentd monitoring setups watch &lt;code>retry_count&lt;/code>, &lt;code>buffer_queue_length&lt;/code>, and process liveness. Almost none watch &lt;code>emit_error_count&lt;/code>. The usual discovery is postmortem: hours of logs are missing, and the missing window is the exact window the incident needed.&lt;/p>
&lt;p>&lt;code>emit_error_count&lt;/code> counts emit transactions that failed inside the pipeline: events that could not be handed off and will never be delivered. Any nonzero rate means data loss is happening right now. This is not a degradation signal or a leading indicator. It is confirmation that events are gone.&lt;/p></description></item><item><title>Fluentd end-to-end pipeline latency: stale logs during an incident</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-end-to-end-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-end-to-end-latency/</guid><description>&lt;h1 id="fluentd-end-to-end-pipeline-latency-stale-logs-during-an-incident">Fluentd end-to-end pipeline latency: stale logs during an incident&lt;/h1>
&lt;p>Your dashboards show errors spiking at 02:00, but when you query the log store for the same window, the newest events are 40 minutes old. Fluentd is running, the process is healthy, no alerts fired. The pipeline is alive but the data is stale, and you are debugging an incident on a delay you did not know existed.&lt;/p>
&lt;p>Fluentd end-to-end pipeline latency is the time from event generation to arrival at the destination. Fluentd does not expose this as a native metric. There is no &lt;code>pipeline_latency_seconds&lt;/code> field in the monitor_agent API. You have to derive it: compare event timestamps to arrival time at the destination, or inject synthetic events with known timestamps and measure round-trip time.&lt;/p></description></item><item><title>Fluentd failed to flush the buffer: the output cannot deliver and retries begin</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-failed-to-flush-the-buffer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-failed-to-flush-the-buffer/</guid><description>&lt;h1 id="fluentd-failed-to-flush-the-buffer-the-output-cannot-deliver-and-retries-begin">Fluentd failed to flush the buffer: the output cannot deliver and retries begin&lt;/h1>
&lt;p>Your Fluentd log starts emitting lines like this, over and over:&lt;/p>
&lt;pre tabindex="0">&lt;code>[warn]: #0 failed to flush the buffer. retry_time=3 next_retry_seconds=2026-07-21 22:50:11 +0000 chunk=&amp;#34;5e1a2b...&amp;#34; error_class=Net::OpenTimeout error=&amp;#34;execution expired&amp;#34;
&lt;/code>&lt;/pre>&lt;p>An output plugin tried to flush a buffer chunk to its destination and failed. Fluentd did not lose the chunk; it rolled the chunk back into the queue and scheduled a retry with exponential backoff. The warning repeats once per failed attempt, with &lt;code>retry_time&lt;/code> climbing and &lt;code>next_retry_seconds&lt;/code> drifting further into the future.&lt;/p></description></item><item><title>Fluentd file buffer filling the disk: when the buffer partition runs out</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-buffer-disk-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-buffer-disk-full/</guid><description>&lt;h1 id="fluentd-file-buffer-filling-the-disk-when-the-buffer-partition-runs-out">Fluentd file buffer filling the disk: when the buffer partition runs out&lt;/h1>
&lt;p>A file-backed buffer is supposed to be the durable option: the output stalls, chunks pile up on disk, the destination recovers, the backlog drains. That story only holds while the partition underneath the buffer directory has free space. When it does not, the failure is abrupt. At 100% full, Fluentd cannot stage new chunks, incoming events are lost, and anything else writing to the same partition fails at the same time.&lt;/p></description></item><item><title>Fluentd in_tail not reading: the file is growing but no events are emitted</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-in-tail-not-reading/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-in-tail-not-reading/</guid><description>&lt;h1 id="fluentd-in_tail-not-reading-the-file-is-growing-but-no-events-are-emitted">Fluentd in_tail not reading: the file is growing but no events are emitted&lt;/h1>
&lt;p>The application is writing logs. &lt;code>ls -la&lt;/code> shows the file growing. But Fluentd&amp;rsquo;s input &lt;code>emit_records&lt;/code> for that &lt;code>in_tail&lt;/code> source has been flat for minutes or hours, and nothing is arriving downstream. The daemon runs, the buffers stay empty, and the only evidence is missing data. This is an input-side stall, and it is one of the quieter Fluentd failures because nothing page-worthy breaks.&lt;/p></description></item><item><title>Fluentd in_tail with multi-worker: pinning to a single worker</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-in-tail-multi-worker/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-in-tail-multi-worker/</guid><description>&lt;h1 id="fluentd-in_tail-with-multi-worker-pinning-to-a-single-worker">Fluentd in_tail with multi-worker: pinning to a single worker&lt;/h1>
&lt;p>You added &lt;code>workers 4&lt;/code> to &lt;code>&amp;lt;system&amp;gt;&lt;/code> to get past the single-core ceiling the Ruby GVL imposes on one Fluentd process, and now Fluentd refuses to start. The log says: &lt;code>Plugin 'tail' does not support multi workers configuration (Fluent::Plugin::TailInput)&lt;/code>. Or worse: it starts, and you discover weeks later that some log files were read twice by different workers while others were never read at all.&lt;/p></description></item><item><title>Fluentd input emit_records dropped to zero: ingestion has stopped</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-input-emit-records-dropped-to-zero/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-input-emit-records-dropped-to-zero/</guid><description>&lt;h1 id="fluentd-input-emit_records-dropped-to-zero-ingestion-has-stopped">Fluentd input emit_records dropped to zero: ingestion has stopped&lt;/h1>
&lt;p>The &lt;code>emit_records&lt;/code> counter for an input plugin is the number of records that plugin has handed to the Fluentd router since process start. It is cumulative and monotonically increasing. If its computed rate drops to zero and holds there while the upstream source is still writing logs, ingestion is broken: events are being generated somewhere and silently going nowhere.&lt;/p>
&lt;p>This is one of the few Fluentd conditions that justifies paging immediately. A dead process is obvious. A live process that has stopped ingesting is a silent observability blackout, and it will not announce itself anywhere except this counter.&lt;/p></description></item><item><title>Fluentd input emit_records stuck at zero: enable_input_metrics on older versions</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-enable-input-metrics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-enable-input-metrics/</guid><description>&lt;h1 id="fluentd-input-emit_records-stuck-at-zero-enable_input_metrics-on-older-versions">Fluentd input emit_records stuck at zero: enable_input_metrics on older versions&lt;/h1>
&lt;p>You open the monitor_agent API (or a dashboard built on top of it) and every input plugin reports &lt;code>emit_records: 0&lt;/code>. First read: ingestion has stopped. Then you notice output plugins are delivering records, buffers are cycling, and downstream log storage is receiving fresh data. The pipeline is fine. The metric is blind.&lt;/p>
&lt;p>On Fluentd versions before v1.19.0, input plugin metrics are disabled by default. Unless you set &lt;code>enable_input_metrics true&lt;/code> in the &lt;code>&amp;lt;system&amp;gt;&lt;/code> block, every input plugin&amp;rsquo;s &lt;code>emit_records&lt;/code> counter stays at 0 forever, no matter how much data flows through it. This is a common false alarm, and it also cuts the other way: teams running real ingestion stalls on older versions have no input-side visibility at all because they never turned the metrics on.&lt;/p></description></item><item><title>Fluentd input spike: a log storm that overwhelms buffers and outputs</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-log-storm-spike/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-log-storm-spike/</guid><description>&lt;h1 id="fluentd-input-spike-a-log-storm-that-overwhelms-buffers-and-outputs">Fluentd input spike: a log storm that overwhelms buffers and outputs&lt;/h1>
&lt;p>Your dashboards show input &lt;code>emit_records&lt;/code> suddenly at 5x or 50x baseline. Within minutes, &lt;code>buffer_queue_length&lt;/code> starts climbing, average flush time rises, and you are watching the buffer fill in real time. This is a log storm: an input spike large enough that the output side of the pipeline cannot drain it.&lt;/p>
&lt;p>The danger is not the spike itself; Fluentd is built to absorb bursts, that is what the buffer is for. The danger is the cascade: the buffer fills to &lt;code>total_limit_size&lt;/code>, &lt;code>overflow_action&lt;/code> fires, and depending on configuration you either drop new events (&lt;code>throw_exception&lt;/code>, the default), stall all input threads (&lt;code>block&lt;/code>), or discard your oldest buffered data (&lt;code>drop_oldest_chunk&lt;/code>). None of these outcomes is obvious unless you know which counters to watch.&lt;/p></description></item><item><title>Fluentd log rotation loss: copytruncate, rotate_wait, and missed lines</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-log-rotation-loss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-log-rotation-loss/</guid><description>&lt;h1 id="fluentd-log-rotation-loss-copytruncate-rotate_wait-and-missed-lines">Fluentd log rotation loss: copytruncate, rotate_wait, and missed lines&lt;/h1>
&lt;p>Log rotation is the most common source of quiet data loss in a Fluentd deployment. Pipeline metrics look healthy, the destination keeps receiving data, and yet there is a gap in the log stream at exactly 00:00 every night, or a burst of duplicate records right after logrotate runs. The cause is almost always the interaction between the rotation method used by logrotate (or the container runtime) and the assumptions Fluentd&amp;rsquo;s &lt;code>in_tail&lt;/code> plugin makes about how files change on disk.&lt;/p></description></item><item><title>Fluentd memory growing: leak versus the normal Ruby fragmentation plateau</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-memory-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-memory-growing/</guid><description>&lt;h1 id="fluentd-memory-growing-leak-versus-the-normal-ruby-fragmentation-plateau">Fluentd memory growing: leak versus the normal Ruby fragmentation plateau&lt;/h1>
&lt;p>Fluentd RSS has been climbing for hours or days and you are trying to decide whether you have a memory leak or whether this is just what Ruby does. The answer matters because the two cases have completely different responses: one requires no action at all, the other ends in an OOM kill and, if your buffers are memory-backed, permanent loss of buffered log data.&lt;/p></description></item><item><title>Fluentd memory vs file buffer: why the default buffer loses data on restart</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-memory-vs-file-buffer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-memory-vs-file-buffer/</guid><description>&lt;h1 id="fluentd-memory-vs-file-buffer-why-the-default-buffer-loses-data-on-restart">Fluentd memory vs file buffer: why the default buffer loses data on restart&lt;/h1>
&lt;p>You restarted Fluentd for a config change, or the OOM killer restarted it for you, or Kubernetes rescheduled the pod. The process came back healthy. Every metric looks normal. But downstream there is a gap in the logs covering the minutes before the restart, and no error anywhere explains it.&lt;/p>
&lt;p>The explanation is almost always the same: the output was using the memory buffer, and every chunk that had not been flushed at the moment the process died was deleted with it. This is not a bug. It is the documented behavior of the memory buffer. The Fluentd troubleshooting documentation itself lists &amp;ldquo;change buffer type from memory to file&amp;rdquo; as a standard remediation for exactly this symptom.&lt;/p></description></item><item><title>Fluentd monitor_agent not responding: a process that is up but hung</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-monitor-agent-not-responding/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-monitor-agent-not-responding/</guid><description>&lt;h1 id="fluentd-monitor_agent-not-responding-a-process-that-is-up-but-hung">Fluentd monitor_agent not responding: a process that is up but hung&lt;/h1>
&lt;p>Your liveness check says Fluentd is fine. The PID exists, &lt;code>systemctl status&lt;/code> shows active, but &lt;code>curl http://localhost:24220/api/plugins.json&lt;/code> hangs until it times out, or returns something other than 200, and no logs are moving.&lt;/p>
&lt;p>This is the zombie state: a process that passes every cheap liveness check but is functionally dead. The monitor_agent endpoint is the cheapest honest probe you have for this. If it does not answer, the internal runtime is not making progress, regardless of what the process table says.&lt;/p></description></item><item><title>Fluentd Monitoring</title><link>https://www.netdata.cloud/monitoring-101/fluentd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/fluentd-monitoring/</guid><description>&lt;h2 id="fluentd-monitoring">Fluentd Monitoring&lt;/h2>
&lt;h3 id="what-is-fluentd">What Is Fluentd?&lt;/h3>
&lt;p>&lt;a href="https://www.fluentd.org/">Fluentd&lt;/a> is a powerful open-source data collector that allows you to unify the collection and consumption of data streams. It helps streamline log data management and enables real-time data processing.&lt;/p>
&lt;h3 id="monitoring-fluentd-with-netdata">Monitoring Fluentd With Netdata&lt;/h3>
&lt;p>Using Netdata as your Fluentd monitoring tool provides an efficient, real-time overview of the performance and health of your Fluentd instances. Netdata&amp;rsquo;s lightweight but comprehensive monitoring capabilities make it a preferred choice for DevOps and IT professionals.&lt;/p></description></item><item><title>Fluentd monitoring checklist: the signals every production log pipeline needs</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-monitoring-checklist/</guid><description>&lt;h1 id="fluentd-monitoring-checklist-the-signals-every-production-log-pipeline-needs">Fluentd monitoring checklist: the signals every production log pipeline needs&lt;/h1>
&lt;p>Most Fluentd monitoring setups answer one question: is the process running? That is necessary and nowhere near sufficient. A Fluentd process can be alive, responsive, and completely idle because the buffer is full and the default &lt;code>overflow_action&lt;/code> (&lt;code>throw_exception&lt;/code>) is dropping every new event at the input. No error counter reliably increments for that path. The process looks healthy while the pipeline loses data.&lt;/p></description></item><item><title>Fluentd monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-monitoring-maturity-model/</guid><description>&lt;h1 id="fluentd-monitoring-maturity-model-from-survival-to-expert">Fluentd monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most teams monitoring Fluentd stop at &amp;ldquo;is the process running?&amp;rdquo; and discover the gap during an incident, when the SIEM is missing the exact logs they need for a postmortem. Fluentd can be alive, responsive, and green on every dashboard while silently dropping events, accumulating a buffer that will overflow in forty minutes, or retrying into a backoff so deep the pipeline is effectively dead.&lt;/p></description></item><item><title>Fluentd OOM killed: memory bloat, the OOM killer, and lost buffers</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-oom-killed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-oom-killed/</guid><description>&lt;h1 id="fluentd-oom-killed-memory-bloat-the-oom-killer-and-lost-buffers">Fluentd OOM killed: memory bloat, the OOM killer, and lost buffers&lt;/h1>
&lt;p>Fluentd disappeared. The process is gone, systemd or Kubernetes restarted it, and there is a gap in your log storage covering the last few minutes or hours. A check of &lt;code>dmesg&lt;/code> shows the OOM killer picked Fluentd as its victim. Memory grew for hours, GC fought harder and harder, and then the kernel ended it.&lt;/p>
&lt;p>This failure mode is expensive because of what dies with the process. If your outputs use memory-backed buffers, every staged and queued chunk that had not yet been flushed is gone. The OOM kill is not a graceful shutdown. Ruby cleanup code does not run, so &lt;code>flush_at_shutdown&lt;/code> never gets a chance to fire. In Kubernetes the same mechanism shows up as CrashLoopBackOff: the pod restarts, memory climbs back to the cgroup limit, and the kubelet kills it again.&lt;/p></description></item><item><title>Fluentd output authentication errors: 401/403 and rejected credentials</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-auth-errors-output/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-auth-errors-output/</guid><description>&lt;h1 id="fluentd-output-authentication-errors-401403-and-rejected-credentials">Fluentd output authentication errors: 401/403 and rejected credentials&lt;/h1>
&lt;p>A &lt;code>401&lt;/code>, &lt;code>403&lt;/code>, &lt;code>unauthorized&lt;/code>, &lt;code>forbidden&lt;/code>, or TLS certificate error in the Fluentd log means the destination is actively rejecting the output plugin&amp;rsquo;s connection. The pipeline does not crash. Fluentd keeps accepting input, buffering events, and attempting flushes that keep failing, so from the outside the agent looks alive while no data reaches that destination.&lt;/p>
&lt;p>The operational risk is the backlog that builds behind the failure. Every rejected flush pushes chunks back into the buffer, and if credentials stay broken long enough the buffer fills and the configured &lt;code>overflow_action&lt;/code> decides what data you lose. With the default &lt;code>throw_exception&lt;/code>, new events are silently discarded. With &lt;code>drop_oldest_chunk&lt;/code>, the oldest buffered chunks are destroyed. Neither emits an obvious alert unless you are watching the right counters.&lt;/p></description></item><item><title>Fluentd output rate lower than input rate: the deficit that is quietly dropping logs</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-output-emit-lower-than-input/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-output-emit-lower-than-input/</guid><description>&lt;h1 id="fluentd-output-rate-lower-than-input-rate-the-deficit-that-is-quietly-dropping-logs">Fluentd output rate lower than input rate: the deficit that is quietly dropping logs&lt;/h1>
&lt;p>Your Fluentd process is up. The monitor agent responds. &lt;code>retry_count&lt;/code> is zero. And yet, when you query the destination, whole time windows of logs are missing. The cause is usually the same: output &lt;code>emit_records&lt;/code> has been running below input &lt;code>emit_records&lt;/code> for hours or days, and nobody was comparing them.&lt;/p>
&lt;p>Most teams chart input rate and output rate independently and never compute the ratio. A sustained 5% deficit on a host doing 200 events per second is over 6 million events lost per week. The two rates should converge over any reasonable window. When they do not, data is either being dropped or piling up in a buffer that will eventually overflow and drop it anyway.&lt;/p></description></item><item><title>Fluentd overflow_action: throw_exception, block, and drop_oldest_chunk</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-overflow-action/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-overflow-action/</guid><description>&lt;h1 id="fluentd-overflow_action-throw_exception-block-and-drop_oldest_chunk">Fluentd overflow_action: throw_exception, block, and drop_oldest_chunk&lt;/h1>
&lt;p>When a Fluentd output buffer reaches &lt;code>total_limit_size&lt;/code>, the pipeline does not degrade gradually. It hits a binary transition: the buffer was accepting events, and now it is not. What happens in that moment is decided by a single parameter in the &lt;code>&amp;lt;buffer&amp;gt;&lt;/code> section: &lt;code>overflow_action&lt;/code>. It accepts three values: &lt;code>throw_exception&lt;/code> (the default), &lt;code>block&lt;/code>, and &lt;code>drop_oldest_chunk&lt;/code>.&lt;/p>
&lt;p>Most teams assume &amp;ldquo;buffer full = pipeline blocks and waits.&amp;rdquo; That is only true if you explicitly configure &lt;code>block&lt;/code>. The default, &lt;code>throw_exception&lt;/code>, means new events are rejected when the buffer is full, and depending on the input plugin, those events are simply never ingested. There is no reliable metric that counts these losses.&lt;/p></description></item><item><title>Fluentd pattern not matched: the parser is silently dropping log lines</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-pattern-not-matched/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-pattern-not-matched/</guid><description>&lt;h1 id="fluentd-pattern-not-matched-the-parser-is-silently-dropping-log-lines">Fluentd pattern not matched: the parser is silently dropping log lines&lt;/h1>
&lt;p>You found this in the Fluentd log:&lt;/p>
&lt;pre tabindex="0">&lt;code>[warn]: #0 pattern not match: &amp;#34;2026-07-21T22:58:11.123456789Z stdout F some application log line&amp;#34;
&lt;/code>&lt;/pre>&lt;p>One warning line per dropped record. Or worse: someone turned the warnings off to silence the noise, and now log lines are vanishing with no signal at all. This is input-level data loss. The events never enter the filter chain, never reach the buffer, and never appear in your destination. Pipeline metrics can look healthy because most counters only count what the parser accepted.&lt;/p></description></item><item><title>Fluentd per-worker imbalance: one struggling worker hidden by aggregate metrics</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-per-worker-imbalance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-per-worker-imbalance/</guid><description>&lt;h1 id="fluentd-per-worker-imbalance-one-struggling-worker-hidden-by-aggregate-metrics">Fluentd per-worker imbalance: one struggling worker hidden by aggregate metrics&lt;/h1>
&lt;p>Your Fluentd dashboards look fine. Input rate is steady, summed output rate roughly matches, and average buffer usage across the instance is comfortably below the limit. Then someone notices a three-hour gap in the destination for a subset of logs, or the node starts flapping on disk pressure, and the aggregate metrics still insist nothing is wrong.&lt;/p>
&lt;p>This is the defining trap of multi-worker Fluentd. When you set &lt;code>workers N&lt;/code> in &lt;code>&amp;lt;system&amp;gt;&lt;/code>, Fluentd spawns N independent Ruby processes. Each worker has its own event router, its own buffers, its own flush threads, its own retry state, and its own memory footprint. Workers do not share buffer state. If worker 1&amp;rsquo;s output is stuck in a retry storm while workers 0, 2, and 3 are draining normally, any metric you sum or average across workers will report &amp;ldquo;mostly healthy&amp;rdquo; right up until worker 1 starts losing data.&lt;/p></description></item><item><title>Fluentd plugin load error at startup: LoadError and missing gems</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-plugin-load-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-plugin-load-error/</guid><description>&lt;h1 id="fluentd-plugin-load-error-at-startup-loaderror-and-missing-gems">Fluentd plugin load error at startup: LoadError and missing gems&lt;/h1>
&lt;p>Fluentd fails to start, or starts with part of the pipeline missing, and the log shows &lt;code>LoadError&lt;/code>, &lt;code>cannot load such file&lt;/code>, or a &lt;code>load_plugin&lt;/code> failure. If the failed plugin is an output, that destination silently receives no data while everything else looks healthy. If the plugin is critical to the config, the process refuses to start entirely and you find out via your process-alive check, or via a gap in downstream logs.&lt;/p></description></item><item><title>Fluentd poison pill crash loop: one bad log line that kills the process on every restart</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-poison-pill-crash-loop/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-poison-pill-crash-loop/</guid><description>&lt;h1 id="fluentd-poison-pill-crash-loop-one-bad-log-line-that-kills-the-process-on-every-restart">Fluentd poison pill crash loop: one bad log line that kills the process on every restart&lt;/h1>
&lt;p>Fluentd is restarting every few seconds. The service comes up, CPU spikes to 100 percent almost immediately, the process dies, systemd (or Kubernetes) restarts it, and the cycle repeats. Input throughput is effectively zero, but the process is &amp;ldquo;running&amp;rdquo; most of the time, so a naive process-alive check flaps between green and red without telling you anything.&lt;/p></description></item><item><title>Fluentd pos_file corruption: duplicates and gaps after a restart</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-pos-file-corruption/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-pos-file-corruption/</guid><description>&lt;h1 id="fluentd-pos_file-corruption-duplicates-and-gaps-after-a-restart">Fluentd pos_file corruption: duplicates and gaps after a restart&lt;/h1>
&lt;p>You restarted Fluentd (a deploy, an OOM kill, a pod reschedule) and now one of two things is wrong downstream: the same log lines appear twice, or a window of logs never arrived. Both symptoms point at the same component: the &lt;code>in_tail&lt;/code> position file.&lt;/p>
&lt;p>The pos_file is how &lt;code>in_tail&lt;/code> remembers where it stopped reading. It is a plain text file with one line per tailed file, recording the file path, a byte offset in hexadecimal, and an inode number in hexadecimal. There are no checksums and no integrity metadata. On startup, Fluentd reads it and trusts it completely.&lt;/p></description></item><item><title>Fluentd process not running: the log pipeline is dead and the host has gone dark</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-process-not-running/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-process-not-running/</guid><description>&lt;h1 id="fluentd-process-not-running-the-log-pipeline-is-dead-and-the-host-has-gone-dark">Fluentd process not running: the log pipeline is dead and the host has gone dark&lt;/h1>
&lt;p>The Fluentd process is gone. No logs are being collected, parsed, buffered, or forwarded from this host. Your downstream systems (Elasticsearch, S3, a SIEM, an aggregator tier) are now receiving nothing from here, and most of them will not tell you that. Log pipelines fail silently at the consumer side: the absence of data looks identical to a quiet host.&lt;/p></description></item><item><title>Fluentd read_from_head replay: a memory and CPU spike that re-reads whole files</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-read-from-head-replay/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-read-from-head-replay/</guid><description>&lt;h1 id="fluentd-read_from_head-replay-a-memory-and-cpu-spike-that-re-reads-whole-files">Fluentd read_from_head replay: a memory and CPU spike that re-reads whole files&lt;/h1>
&lt;p>Fluentd has just started (or restarted) and the process is immediately pinned: CPU at or near 100% of one core, RSS climbing fast, and the pipeline emitting a burst of events with old timestamps. Alerts on process CPU, memory growth, or &amp;ldquo;events older than X arriving at the destination&amp;rdquo; are probably firing.&lt;/p>
&lt;p>In most cases this is &lt;code>read_from_head true&lt;/code> doing exactly what you configured. The &lt;code>in_tail&lt;/code> plugin is reading each watched file from byte zero, parsing every historical line, and pushing those events through the filter chain, buffer, and outputs. The spike is real resource consumption, but it is one-time per file per position entry, and it stops when the read catches up.&lt;/p></description></item><item><title>Fluentd retry backoff: a pipeline that is 'retrying' but effectively dead</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-retry-backoff-stalled/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-retry-backoff-stalled/</guid><description>&lt;h1 id="fluentd-retry-backoff-a-pipeline-that-is-retrying-but-effectively-dead">Fluentd retry backoff: a pipeline that is &amp;lsquo;retrying&amp;rsquo; but effectively dead&lt;/h1>
&lt;p>The Fluentd process is up. The monitor agent responds. &lt;code>retry_count&lt;/code> is non-zero, which you already knew, because the destination had a bad night. What &lt;code>retry_count&lt;/code> does not tell you is that the next retry attempt is scheduled 4 hours from now. No data is flowing, none will flow for hours, and every dashboard that only tracks the retry counter shows the same flat number it showed an hour ago. The pipeline is technically retrying and operationally dead.&lt;/p></description></item><item><title>Fluentd retry storm: thundering-herd resonance that keeps a destination down</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-retry-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-retry-storm/</guid><description>&lt;h1 id="fluentd-retry-storm-thundering-herd-resonance-that-keeps-a-destination-down">Fluentd retry storm: thundering-herd resonance that keeps a destination down&lt;/h1>
&lt;p>Your log destination had a hiccup. It recovered. Then it fell over again. And again. Each time it comes back, it survives for a minute or two, gets hit by a wall of buffered log traffic, and falls over. Fluentd&amp;rsquo;s own metrics look confusing: retries intermittently succeed, the buffer queue drains a little, then grows again. The destination team insists their service &amp;ldquo;works fine when we test it.&amp;rdquo;&lt;/p></description></item><item><title>Fluentd retry_count climbing: the destination is rejecting or unreachable</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-retry-count-climbing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-retry-count-climbing/</guid><description>&lt;h1 id="fluentd-retry_count-climbing-the-destination-is-rejecting-or-unreachable">Fluentd retry_count climbing: the destination is rejecting or unreachable&lt;/h1>
&lt;p>You pulled &lt;code>/api/plugins.json&lt;/code> from the monitor_agent and one of your output plugins shows a nonzero &lt;code>retry_count&lt;/code>, and it keeps going up. At the same time, &lt;code>write_count&lt;/code> has stopped incrementing and &lt;code>buffer_queue_length&lt;/code> is growing. Fluentd itself is alive and inputs are still collecting.&lt;/p>
&lt;p>This is the destination-unavailable failure pattern: the output plugin cannot deliver chunks, the retry engine has taken over, and Fluentd is now in exponential backoff against a destination that is rejecting connections, rejecting data, or simply gone. Every event that arrives from now on accumulates in the buffer. The clock you are racing is buffer capacity, not the retry count itself.&lt;/p></description></item><item><title>Fluentd rollback_count: chunks recycling back into the queue after failed flushes</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-rollback-count/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-rollback-count/</guid><description>&lt;h1 id="fluentd-rollback_count-chunks-recycling-back-into-the-queue-after-failed-flushes">Fluentd rollback_count: chunks recycling back into the queue after failed flushes&lt;/h1>
&lt;p>&lt;code>rollback_count&lt;/code> is one of the output-plugin counters exposed by Fluentd&amp;rsquo;s monitor_agent. It answers a specific question: how many times has a buffer chunk been taken out of the queue for flushing, failed, and been put back? If &lt;code>buffer_queue_length&lt;/code> sits stubbornly non-zero while &lt;code>write_count&lt;/code> refuses to move, &lt;code>rollback_count&lt;/code> tells you chunks are actively cycling through the failure path rather than sitting idle.&lt;/p></description></item><item><title>Fluentd Ruby GC pressure: garbage-collection pauses that stall event processing</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-ruby-gc-pressure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-ruby-gc-pressure/</guid><description>&lt;h1 id="fluentd-ruby-gc-pressure-garbage-collection-pauses-that-stall-event-processing">Fluentd Ruby GC pressure: garbage-collection pauses that stall event processing&lt;/h1>
&lt;p>Fluentd runs on CRuby. Every parsed record, buffered chunk, and serialized payload creates Ruby objects that the garbage collector eventually has to reclaim. Under memory pressure, GC stops being background work and starts blocking the pipeline: inputs stop reading, flush threads miss their windows, chunks roll back into the queue, and retries allocate even more objects.&lt;/p>
&lt;p>The outside symptom is easy to misread. The process is alive, CPU is high, and throughput dips in pulses or sags steadily. Buffer metrics can make it look like a slow destination; CPU metrics can make it look like parser load. The distinguishing pattern is Ruby GC activity rising with memory pressure while destination health remains otherwise explainable.&lt;/p></description></item><item><title>Fluentd silent data loss: logs missing downstream with no error at all</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-silent-data-loss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-silent-data-loss/</guid><description>&lt;h1 id="fluentd-silent-data-loss-logs-missing-downstream-with-no-error-at-all">Fluentd silent data loss: logs missing downstream with no error at all&lt;/h1>
&lt;p>Someone queries the log store for an incident window and the logs are not there. Not delayed, not misindexed: missing. You check Fluentd. The process is running, retry_count is zero, no error logs, buffer metrics look unremarkable. Everything is green, and the data is gone.&lt;/p>
&lt;p>This is the hardest Fluentd failure to detect because it is designed into the defaults. When the buffer fills and &lt;code>overflow_action&lt;/code> is &lt;code>throw_exception&lt;/code> (the default), new events are rejected at the input and lost. When &lt;code>overflow_action&lt;/code> is &lt;code>drop_oldest_chunk&lt;/code>, chunks are discarded with only a log warning and one counter that almost nobody alerts on. Parse failures can drop records with no visible trace if the error stream is not routed anywhere. In all three cases the pipeline looks healthy from every angle except the one that matters: the destination.&lt;/p></description></item><item><title>Fluentd slow flush: buffer flush took longer than slow_flush_log_threshold</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-slow-flush/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-slow-flush/</guid><description>&lt;h1 id="fluentd-slow-flush-buffer-flush-took-longer-than-slow_flush_log_threshold">Fluentd slow flush: buffer flush took longer than slow_flush_log_threshold&lt;/h1>
&lt;p>You see this in the Fluentd log:&lt;/p>
&lt;pre tabindex="0">&lt;code>[warn]: buffer flush took longer time than slow_flush_log_threshold: elapsed_time=... slow_flush_log_threshold=20.0 plugin_id=&amp;#34;...&amp;#34;
&lt;/code>&lt;/pre>&lt;p>This is a performance warning, not a data-loss signal. Fluentd delivered (or is still retrying) a buffer chunk to the destination, and the write took longer than &lt;code>slow_flush_log_threshold&lt;/code>, which defaults to 20 seconds. Each occurrence also increments the &lt;code>slow_flush_count&lt;/code> counter on that output plugin.&lt;/p></description></item><item><title>Fluentd throttled_log_count: in_tail rate limiting is dropping lines at the source</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-throttled-log-count/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-throttled-log-count/</guid><description>&lt;h1 id="fluentd-throttled_log_count-in_tail-rate-limiting-is-dropping-lines-at-the-source">Fluentd throttled_log_count: in_tail rate limiting is dropping lines at the source&lt;/h1>
&lt;p>&lt;code>throttled_log_count&lt;/code> is a cumulative counter exposed by the &lt;code>in_tail&lt;/code> input plugin. Every increment means Fluentd deferred log lines at the source because a configured rate limit was exceeded. Unlike buffer-side loss, this happens before the event enters the pipeline: no filter sees it, no buffer holds it, and no output will ever deliver it.&lt;/p>
&lt;p>The counter only moves when you have configured throttling in &lt;code>in_tail&lt;/code>, either via the &lt;code>&amp;lt;group&amp;gt;&lt;/code> section with rate rules or via byte-rate limiting on reads. If you never configured throttling, this counter stays at zero forever, and a zero value tells you nothing. If you did configure it, any non-zero increment rate is a decision point: either the throttle is doing what you designed it to do, or it is silently eating log volume you expected to keep.&lt;/p></description></item><item><title>Fluentd TLS certificate expiry: a cliff-edge that stops every output at once</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-tls-certificate-expiry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-tls-certificate-expiry/</guid><description>&lt;h1 id="fluentd-tls-certificate-expiry-a-cliff-edge-that-stops-every-output-at-once">Fluentd TLS certificate expiry: a cliff-edge that stops every output at once&lt;/h1>
&lt;p>Every TLS output on a Fluentd node was working. Then, at one exact second, all of them started failing with the same handshake error. &lt;code>retry_count&lt;/code> is climbing on every output that uses TLS, &lt;code>write_count&lt;/code> has flatlined, and the buffer is filling. Nothing was deployed. Nothing changed on the network. The only thing that changed is the wall clock: the certificate crossed its &lt;code>notAfter&lt;/code> timestamp.&lt;/p></description></item><item><title>Fluentd too many open files: file descriptor exhaustion stalls the pipeline</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-too-many-open-files/</guid><description>&lt;h1 id="fluentd-too-many-open-files-file-descriptor-exhaustion-stalls-the-pipeline">Fluentd too many open files: file descriptor exhaustion stalls the pipeline&lt;/h1>
&lt;p>Fluentd logs &lt;code>Errno::EMFILE: Too many open files&lt;/code> and the pipeline stops making progress. Log files keep growing on disk, the destination sees nothing new, and the Fluentd process is still alive. It is not crashed; it is wedged against its file descriptor limit, and almost every operation it needs to do next requires opening something.&lt;/p>
&lt;p>What makes this failure nasty is the cliff edge. Everything works until the limit is reached, and then several things fail at once: &lt;code>in_tail&lt;/code> cannot open newly created log files (and may stop watching them without a loud error), the buffer cannot create new chunk files, and outputs cannot establish new connections. A single resource limit takes out input, buffer, and output simultaneously.&lt;/p></description></item><item><title>Fluentd unauthorized in_forward connections: log injection on the aggregator tier</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-unauthorized-forward/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-unauthorized-forward/</guid><description>&lt;h1 id="fluentd-unauthorized-in_forward-connections-log-injection-on-the-aggregator-tier">Fluentd unauthorized in_forward connections: log injection on the aggregator tier&lt;/h1>
&lt;p>You ran &lt;code>ss -tn&lt;/code> on your Fluentd aggregator and saw connections to port 24224 from IP addresses you do not recognize. Or a security review turned up the fact that your aggregator-tier Fluentd accepts forwarded events from anything that can reach it. Either way: your &lt;code>in_forward&lt;/code> input is unauthenticated, and any network-reachable host can inject events directly into your log pipeline.&lt;/p></description></item><item><title>Fluentd worker died: a partial outage the supervisor hides</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-worker-died/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-worker-died/</guid><description>&lt;h1 id="fluentd-worker-died-a-partial-outage-the-supervisor-hides">Fluentd worker died: a partial outage the supervisor hides&lt;/h1>
&lt;p>&lt;code>systemctl status td-agent&lt;/code> says &lt;code>active (running)&lt;/code>. The supervisor PID exists. Host-level checks are green. Meanwhile, one of your four Fluentd workers died twenty minutes ago, and a quarter of the log pipeline is gone or degraded.&lt;/p>
&lt;p>This is the trap of Fluentd multi-worker mode. When you set &lt;code>workers N&lt;/code> in &lt;code>&amp;lt;system&amp;gt;&lt;/code>, Fluentd starts a supervisor process and N independent Ruby workers. Each worker has its own event loop, buffers, output threads, and, if configured, monitor_agent endpoint. Workers do not share in-memory state.&lt;/p></description></item><item><title>Fluentd write_secondary_count: the primary output has failed to its backup</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-write-secondary/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-write-secondary/</guid><description>&lt;h1 id="fluentd-write_secondary_count-the-primary-output-has-failed-to-its-backup">Fluentd write_secondary_count: the primary output has failed to its backup&lt;/h1>
&lt;p>&lt;code>write_secondary_count&lt;/code> is a per-output counter in Fluentd&amp;rsquo;s monitor_agent API. A nonzero value means the primary output plugin exhausted its retries for at least one chunk, and Fluentd wrote that chunk to the configured &lt;code>&amp;lt;secondary&amp;gt;&lt;/code> backup destination instead. The primary pipeline for that output is broken.&lt;/p>
&lt;p>This is a ticket-level signal. Data is not lost yet (that is the point of the secondary), but it is no longer flowing where downstream systems expect it. Dashboards, SIEM rules, and alerts that read from the primary destination are now working from a gap.&lt;/p></description></item><item><title>Force10 Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/force10-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/force10-networks-inc-snmp-traps/</guid><description/></item><item><title>Forcepoint LLC SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/forcepoint-llc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/forcepoint-llc-snmp-traps/</guid><description/></item><item><title>Fore Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fore-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fore-systems-inc-snmp-traps/</guid><description/></item><item><title>Fort Telecom SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fort-telecom-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fort-telecom-snmp-traps/</guid><description/></item><item><title>Forte Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/forte-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/forte-networks-inc-snmp-traps/</guid><description/></item><item><title>Fortinet</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/fortinet/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/fortinet/</guid><description/></item><item><title>Fortinet Appliance</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/fortinet-appliance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/fortinet-appliance/</guid><description/></item><item><title>Fortinet Fortigate</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/fortinet-fortigate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/fortinet-fortigate/</guid><description/></item><item><title>Fortinet Fortiswitch</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/fortinet-fortiswitch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/fortinet-fortiswitch/</guid><description/></item><item><title>Fortinet Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fortinet-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fortinet-inc-snmp-traps/</guid><description/></item><item><title>Fortinet Licensing</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/licensing-monitoring/fortinet-licensing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/licensing-monitoring/fortinet-licensing/</guid><description/></item><item><title>Fraunhofer Fokus SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fraunhofer-fokus-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fraunhofer-fokus-snmp-traps/</guid><description/></item><item><title>FreeBSD</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/freebsd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/freebsd/</guid><description/></item><item><title>FreeBSD Monitoring</title><link>https://www.netdata.cloud/monitoring-101/freebsd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/freebsd-monitoring/</guid><description>&lt;h2 id="why-freebsd">Why FreeBSD?&lt;/h2>
&lt;p>&lt;a href="https://www.freebsd.org/">FreeBSD&lt;/a> is a free and open source Unix-like operating system descended from the Berkeley Software Distribution (BSD). FreeBSD is the most widely used open source BSD distribution. It is used by companies such as Netflix, Baidu, the United States Department of Defense, and the European Organization for Nuclear Research (CERN).&lt;/p>
&lt;p>FreeBSD is a high-quality, stable, and secure operating system used in a wide variety of applications. FreeBSD is developed by a large and passionate community of volunteers and is an excellent choice if you are looking for a high-quality, stable, and secure operating system.&lt;/p></description></item><item><title>FreeBSD NFS</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/freebsd-nfs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/freebsd-nfs/</guid><description/></item><item><title>FreeBSD NFS Monitoring</title><link>https://www.netdata.cloud/monitoring-101/freebsd_nfs-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/freebsd_nfs-monitoring/</guid><description>&lt;h2 id="freebsd-nfs-monitoring">FreeBSD NFS Monitoring&lt;/h2>
&lt;h3 id="what-is-freebsd-nfs">What Is FreeBSD NFS?&lt;/h3>
&lt;p>FreeBSD Network File System (NFS) is a protocol that allows file and directory sharing across different systems over a network. Designed for efficient, streamlined connections, FreeBSD NFS is a reliable choice for environments requiring seamless file sharing capabilities.&lt;/p>
&lt;h3 id="monitoring-freebsd-nfs-with-netdata">Monitoring FreeBSD NFS With Netdata&lt;/h3>
&lt;p>Monitoring FreeBSD NFS efficiently is crucial to ensure optimal performance and reliability across networked systems. Netdata offers an effective FreeBSD NFS monitoring tool that leverages an OpenMetrics (Prometheus) exporter. Netdata can ingest data from any Prometheus exporter, meaning users can enjoy automated dashboards, alerts, and more without the need for running a full Prometheus server or Grafana setup. Simply connect to the &lt;a href="https://github.com/Axcient/freebsd-nfs-exporter">FreeBSD NFS Exporter&lt;/a>, and Netdata takes care of the rest, from data collection to visualization and alerting.&lt;/p></description></item><item><title>FreeBSD RCTL-RACCT</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/freebsd-rctl-racct/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/freebsd-rctl-racct/</guid><description/></item><item><title>FreeBSD RCTL-RACCT Monitoring</title><link>https://www.netdata.cloud/monitoring-101/freebsd_rctl-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/freebsd_rctl-monitoring/</guid><description>&lt;h2 id="freebsd-rctl-racct-monitoring">FreeBSD RCTL-RACCT Monitoring&lt;/h2>
&lt;h3 id="what-is-freebsd-rctl-racct">What Is FreeBSD RCTL-RACCT?&lt;/h3>
&lt;p>FreeBSD RCTL-RACCT provides an intricate framework for resource consumption accounting and control on machines running the FreeBSD operating system. It enables precise monitoring and limiting of resource usage, which is vital for maintaining optimal system performance. Key tasks include controlling CPU, memory, and other resource utilizations across different processes and users.&lt;/p>
&lt;h3 id="monitoring-freebsd-rctl-racct-with-netdata">Monitoring FreeBSD RCTL-RACCT With Netdata&lt;/h3>
&lt;p>To monitor FreeBSD RCTL-RACCT, Netdata employs an openmetrics (Prometheus) exporter. Users can rely on the &lt;a href="https://github.com/yo000/rctl_exporter">FreeBSD RCTL Exporter&lt;/a> to gather crucial metrics seamlessly. Netdata excels in ingesting data from any Prometheus exporter, generating automated dashboards, real-time alerts, and more without necessitating a separate Prometheus server or Grafana setup. This streamlined approach leverages robust open-source technology to ensure comprehensive observability of your FreeBSD resources.&lt;/p></description></item><item><title>FreeRADIUS</title><link>https://www.netdata.cloud/integrations/data-collection/applications/freeradius/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/freeradius/</guid><description/></item><item><title>FreeRADIUS Monitoring</title><link>https://www.netdata.cloud/monitoring-101/freeradius-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/freeradius-monitoring/</guid><description>&lt;h2 id="freeradius-monitoring">FreeRADIUS Monitoring&lt;/h2>
&lt;h3 id="what-is-freeradius">What Is FreeRADIUS?&lt;/h3>
&lt;p>&lt;a href="https://freeradius.org/">FreeRADIUS&lt;/a> is an open-source implementation of the RADIUS protocol aimed at providing centralized Authentication, Authorization, and Accounting (AAA) services for networked applications. Its broad adoption and flexible configuration options make it a preferred choice for ISPs and enterprises looking to manage network access efficiently.&lt;/p>
&lt;h3 id="monitoring-freeradius-with-netdata">Monitoring FreeRADIUS With Netdata&lt;/h3>
&lt;p>Monitoring FreeRADIUS is crucial for maintaining the reliability and performance of your network services. With Netdata&amp;rsquo;s FreeRADIUS monitoring tool, you can achieve real-time observability, receive timely alerts, and access detailed visualizations of your FreeRADIUS servers&amp;rsquo; performance. Netdata&amp;rsquo;s collector gathers comprehensive metrics, allowing you to keep a keen eye on the health of your RADIUS server in a single dashboard.&lt;/p></description></item><item><title>Freifunk network</title><link>https://www.netdata.cloud/integrations/data-collection/networking/freifunk-network/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/freifunk-network/</guid><description/></item><item><title>Freifunk Network Monitoring</title><link>https://www.netdata.cloud/monitoring-101/freifunk-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/freifunk-monitoring/</guid><description>&lt;h2 id="freifunk-network-monitoring">Freifunk Network Monitoring&lt;/h2>
&lt;h3 id="what-is-freifunk">What Is Freifunk?&lt;/h3>
&lt;p>Freifunk is a grassroots initiative to create free and open wireless community networks. It enables various communities to contribute and connect using decentralized technology. The Freifunk network operates on the principle of crowd-sourced internet, where each node in the network can communicate with others, facilitating a global mesh of connections.&lt;/p>
&lt;h3 id="monitoring-freifunk-networks-with-netdata">Monitoring Freifunk Networks With Netdata&lt;/h3>
&lt;p>Monitoring the Freifunk network is crucial for maintaining seamless connectivity and optimizing network performance. With Netdata, this task becomes efficient and manageable. By leveraging an openmetrics (Prometheus) exporter such as the &lt;a href="https://github.com/xperimental/freifunk-exporter">Freifunk Exporter&lt;/a>, Netdata can ingest valuable data to provide real-time insights. Notably, Netdata does not require a Prometheus server or Grafana to operate. The platform offers automated dashboards and alerts, significantly easing the monitoring process. Learn more about how &lt;a href="https://learn.netdata.cloud/docs/collecting-metrics/generic-collecting-metrics/prometheus-endpoint/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata can ingest data from any Prometheus exporter&lt;/a>.&lt;/p></description></item><item><title>FRRouting</title><link>https://www.netdata.cloud/integrations/data-collection/networking/frrouting/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/frrouting/</guid><description/></item><item><title>FRRouting Monitoring</title><link>https://www.netdata.cloud/monitoring-101/frrouting-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/frrouting-monitoring/</guid><description>&lt;h2 id="frrouting-monitoring">FRRouting Monitoring&lt;/h2>
&lt;h3 id="what-is-frrouting">What Is FRRouting?&lt;/h3>
&lt;p>FRRouting (FRR) is a robust, high-performance suite for software-based routing on Linux and Unix systems. It offers a comprehensive set of tools aimed at providing scalability and flexibility, allowing networks to handle both small and large network infrastructures effectively. With FRR, network administrators can ensure optimal routing protocols are in place for efficient network traffic management.&lt;/p>
&lt;h3 id="monitoring-frrouting-with-netdata">Monitoring FRRouting With Netdata&lt;/h3>
&lt;p>Monitoring FRRouting with Netdata is simplified by utilizing the openmetrics (Prometheus) exporter. Netdata has streamlined the process, allowing users to easily ingest data from any Prometheus exporter. This functionality ensures automated dashboards, alerts, and insights without necessitating a standalone Prometheus server or Grafana. You can &lt;a href="https://github.com/tynany/frr_exporter">get the community exporter here&lt;/a>.&lt;/p></description></item><item><title>Fs Com Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fs-com-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fs-com-inc-snmp-traps/</guid><description/></item><item><title>Fujitsu Access Limited SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fujitsu-access-limited-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fujitsu-access-limited-snmp-traps/</guid><description/></item><item><title>Fujitsu Germany GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fujitsu-germany-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fujitsu-germany-gmbh-snmp-traps/</guid><description/></item><item><title>Fujitsu Limited SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fujitsu-limited-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fujitsu-limited-snmp-traps/</guid><description/></item><item><title>Fujitsu Network Communications Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fujitsu-network-communications-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fujitsu-network-communications-inc-snmp-traps/</guid><description/></item><item><title>Fusionio SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fusionio-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/fusionio-snmp-traps/</guid><description/></item><item><title>Future Software SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/future-software-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/future-software-snmp-traps/</guid><description/></item><item><title>G-Sense_Error_Rate rising: shock and vibration reaching the drive</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-g-sense-error-rate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-g-sense-error-rate/</guid><description>&lt;h1 id="g-sense_error_rate-rising-shock-and-vibration-reaching-the-drive">G-Sense_Error_Rate rising: shock and vibration reaching the drive&lt;/h1>
&lt;p>G-Sense_Error_Rate (SMART attribute ID 221 on most drives, ID 191 on some) is a cumulative counter of shock and vibration events that exceeded the drive&amp;rsquo;s internal threshold, as detected by its built-in accelerometer. The attribute is HDD-only. SSDs and NVMe drives do not report it. Some enterprise HDD models do not expose the attribute at all, which does not mean vibration is absent, only that the drive does not instrument it.&lt;/p></description></item><item><title>Gadzoox Microsystems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/gadzoox-microsystems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/gadzoox-microsystems-inc-snmp-traps/</guid><description/></item><item><title>Gamatronic Electronic Industries Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/gamatronic-electronic-industries-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/gamatronic-electronic-industries-ltd-snmp-traps/</guid><description/></item><item><title>Gandalf SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/gandalf-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/gandalf-snmp-traps/</guid><description/></item><item><title>Garderos GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/garderos-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/garderos-gmbh-snmp-traps/</guid><description/></item><item><title>Gcom Technologies Co Ltd Formerly Greennet Technology Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/gcom-technologies-co-ltd-formerly-greennet-technology-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/gcom-technologies-co-ltd-formerly-greennet-technology-co-ltd-snmp-traps/</guid><description/></item><item><title>GCP GCE</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/gcp-gce/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/gcp-gce/</guid><description/></item><item><title>GCP GCE Monitoring</title><link>https://www.netdata.cloud/monitoring-101/gcp_gce-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/gcp_gce-monitoring/</guid><description>&lt;h2 id="gcp-gce-monitoring">GCP GCE Monitoring&lt;/h2>
&lt;h3 id="what-is-gcp-gce">What Is GCP GCE?&lt;/h3>
&lt;p>Google Cloud Platform&amp;rsquo;s Compute Engine (GCE) is a robust cloud service that provides scalable virtual machines on demand. GCE is designed to be flexible, offering many choices in terms of operating systems, resource configurations, and networking, making it an ideal choice for developers and enterprises looking to leverage Google&amp;rsquo;s infrastructure for building, testing, and deploying applications.&lt;/p>
&lt;h3 id="monitoring-gcp-gce-with-netdata">Monitoring GCP GCE With Netdata&lt;/h3>
&lt;p>To effectively monitor GCP GCE, Netdata leverages a Prometheus-based approach through an openmetrics exporter specifically designed for Google Cloud. The &lt;a href="https://github.com/O1ahmad/gcp-gce-exporter">GCP GCE Exporter&lt;/a> collects rich and detailed metrics, enabling seamless cloud resource management and enhanced performance monitoring. Netdata can ingest data from any Prometheus exporter, providing automated dashboards and alerts without the need for a standalone Prometheus server or Grafana. This capability ensures that monitoring your cloud infrastructure is as efficient and real-time as possible.&lt;/p></description></item><item><title>GCP IP Ranges</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/gcp-ip-ranges/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/gcp-ip-ranges/</guid><description/></item><item><title>Gearman</title><link>https://www.netdata.cloud/integrations/data-collection/applications/gearman/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/gearman/</guid><description/></item><item><title>Gearman Monitoring</title><link>https://www.netdata.cloud/monitoring-101/gearman-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/gearman-monitoring/</guid><description>&lt;h2 id="gearman-monitoring">Gearman Monitoring&lt;/h2>
&lt;h3 id="what-is-gearman">What is Gearman?&lt;/h3>
&lt;p>&lt;a href="https://gearman.org/">Gearman&lt;/a> is a distributed job processing system that allows applications to farm out work to other machines or processes that are better suited to do it. It consists of clients, workers, and a Gearman server to queue tasks. This decouples the workload from the tasks, helping in distributing various tasks in a wide set of machines or processes.&lt;/p>
&lt;h3 id="monitoring-gearman-with-netdata">Monitoring Gearman With Netdata&lt;/h3>
&lt;p>Monitor Gearman effectively using the Netdata monitoring tool, designed to keep an eye on your Gearman instances in real-time. Netdata provides significant insight into jobs&amp;rsquo; activity, priority, and available workers through its advanced monitoring solutions. This helps in maintaining an optimal work balance and swiftly addressing potential issues.&lt;/p></description></item><item><title>Gemtek Systems Holding B.V. SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/gemtek-systems-holding-b.v.-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/gemtek-systems-holding-b.v.-snmp-traps/</guid><description/></item><item><title>General Datacomm Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/general-datacomm-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/general-datacomm-inc-snmp-traps/</guid><description/></item><item><title>General Instrument SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/general-instrument-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/general-instrument-snmp-traps/</guid><description/></item><item><title>Generex Systems GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/generex-systems-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/generex-systems-gmbh-snmp-traps/</guid><description/></item><item><title>Generic BGP (BGP4-MIB)</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/generic-bgp-bgp4-mib/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/generic-bgp-bgp4-mib/</guid><description/></item><item><title>Generic JSON-over-HTTP IPAM</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/generic-json-over-http-ipam/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/generic-json-over-http-ipam/</guid><description/></item><item><title>Generic SNMP Device</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/generic-snmp-device/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/generic-snmp-device/</guid><description/></item><item><title>Generic storage enclosure tool</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/generic-storage-enclosure-tool/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/generic-storage-enclosure-tool/</guid><description/></item><item><title>Generic storage enclosure tool Monitoring</title><link>https://www.netdata.cloud/monitoring-101/enclosure-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/enclosure-monitoring/</guid><description>&lt;h2 id="generic-storage-enclosure-tool-monitoring">Generic storage enclosure tool Monitoring&lt;/h2>
&lt;h3 id="what-is-generic-storage-enclosure-tool">What Is Generic storage enclosure tool?&lt;/h3>
&lt;p>The Generic storage enclosure tool, accessible from its &lt;a href="https://github.com/Gandi/jbod-rs">GitHub repository&lt;/a>, is a pivotal utility for managing storage enclosure metrics. It aids in effective management and performance evaluation of storage devices. Specifically created for IT infrastructures heavily relying on storage, it ensures your IT assets are running smoothly and efficiently by tracking critical metrics.&lt;/p>
&lt;h3 id="monitoring-generic-storage-enclosure-tool-with-netdata">Monitoring Generic storage enclosure tool With Netdata&lt;/h3>
&lt;p>Monitoring your storage environment is crucial and Netdata offers a comprehensive solution for it. Using Netdata, you can monitor the Generic storage enclosure tool without needing complex setups like a full Prometheus server or Grafana. Netdata uses an openmetrics (Prometheus) exporter to pull metrics, automatically generating dashboards and alerts for a seamless monitoring experience. Netdata can ingest data effortlessly from any Prometheus exporter, offering profound insights and easy setup.&lt;/p></description></item><item><title>Generic UPS (UPS-MIB)</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/generic-ups-ups-mib/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/generic-ups-ups-mib/</guid><description/></item><item><title>Genie Network Resource Management SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/genie-network-resource-management-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/genie-network-resource-management-snmp-traps/</guid><description/></item><item><title>Get support!</title><link>https://www.netdata.cloud/support/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/support/</guid><description/></item><item><title>getifaddrs</title><link>https://www.netdata.cloud/integrations/data-collection/networking/getifaddrs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/getifaddrs/</guid><description/></item><item><title>getmntinfo</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/getmntinfo/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/getmntinfo/</guid><description/></item><item><title>Gigamon</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/gigamon/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/gigamon/</guid><description/></item><item><title>Gigamon Systems LLC SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/gigamon-systems-llc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/gigamon-systems-llc-snmp-traps/</guid><description/></item><item><title>Giganet Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/giganet-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/giganet-ltd-snmp-traps/</guid><description/></item><item><title>GitHub API rate limit</title><link>https://www.netdata.cloud/integrations/data-collection/applications/github-api-rate-limit/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/github-api-rate-limit/</guid><description/></item><item><title>GitHub API Rate Limit Monitoring</title><link>https://www.netdata.cloud/monitoring-101/github_ratelimit-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/github_ratelimit-monitoring/</guid><description>&lt;h2 id="github-api-rate-limit-monitoring">GitHub API Rate Limit Monitoring&lt;/h2>
&lt;h3 id="what-is-github-api-rate-limit">What Is GitHub API Rate Limit?&lt;/h3>
&lt;p>GitHub API rate limits are set thresholds that guide the number of requests a user can make within a certain time frame using the GitHub API. Monitoring these limits ensures that your applications and services make efficient use of API calls without exceeding the permitted usage, preventing disruptions and maintaining performance reliability.&lt;/p>
&lt;h3 id="monitoring-github-api-rate-limit-with-netdata">Monitoring GitHub API Rate Limit With Netdata&lt;/h3>
&lt;p>Netdata provides a seamless approach to monitor GitHub API rate limits using an openmetrics (Prometheus) exporter. The &lt;a href="https://github.com/lunarway/github-ratelimit-exporter">GitHub API rate limit Exporter&lt;/a> efficiently gathers metrics on rate limits, allowing you to gain insights into your API usage. Netdata can ingest data from any Prometheus exporter, delivering automated dashboards and alerts without the need for a Prometheus server or Grafana setup. This unique capability sets Netdata apart, simplifying the process of tracking API consumption in real-time and enhancing your operational insights.&lt;/p></description></item><item><title>GitHub repository</title><link>https://www.netdata.cloud/integrations/data-collection/applications/github-repository/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/github-repository/</guid><description/></item><item><title>GitHub Repository Monitoring</title><link>https://www.netdata.cloud/monitoring-101/github_repo-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/github_repo-monitoring/</guid><description>&lt;h2 id="github-repository-monitoring">GitHub Repository Monitoring&lt;/h2>
&lt;h3 id="what-is-github-repository-monitoring">What Is GitHub Repository Monitoring?&lt;/h3>
&lt;p>GitHub repository monitoring involves tracking a plethora of metrics essential for maintaining robust and high-performing code repositories. As various developers contribute code, issues may arise that affect performance and reliability. Monitoring these aspects is crucial for ensuring smooth DevOps processes. A well-rounded approach to monitoring examines pull requests, commits, issues, and broader analytics related to repository activity.&lt;/p>
&lt;h3 id="monitoring-github-repository-with-netdata">Monitoring GitHub Repository With Netdata&lt;/h3>
&lt;p>Netdata provides a seamless path to monitor GitHub repositories. To monitor GitHub repositories, Netdata uses an openmetrics (Prometheus) exporter, the &lt;a href="https://github.com/githubexporter/github-exporter">GitHub Exporter&lt;/a>. Netdata is designed to ingest data from any Prometheus exporter, allowing you to benefit from automated dashboards, alerts, and real-time insights without requiring standalone Prometheus servers or Grafana dashboards. This streamlined approach facilitates effective monitoring while saving overhead on setting up intricate monitoring systems.&lt;/p></description></item><item><title>GitLab Runner</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/gitlab-runner/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/gitlab-runner/</guid><description/></item><item><title>GitLab Runner Monitoring</title><link>https://www.netdata.cloud/monitoring-101/gitlab_runner-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/gitlab_runner-monitoring/</guid><description>&lt;h2 id="gitlab-runner-monitoring">GitLab Runner Monitoring&lt;/h2>
&lt;h3 id="what-is-gitlab-runner">What Is GitLab Runner?&lt;/h3>
&lt;p>GitLab Runner is an open-source project that is used to run your jobs and send the results back to GitLab. It is a key component in GitLab&amp;rsquo;s CI/CD pipeline, responsible for executing code and providing continuous feedback. GitLab Runner supports multiple platforms, cloud-native applications, and many advanced configuration options, making it versatile and essential for developers looking to automate their CI/CD processes.&lt;/p>
&lt;h3 id="monitoring-gitlab-runner-with-netdata">Monitoring GitLab Runner With Netdata&lt;/h3>
&lt;p>Netdata is a powerful tool to monitor GitLab Runner. Using a Prometheus exporter, Netdata collects detailed metrics from GitLab Runner without the need for a Prometheus server or Grafana installation. This setup enables users to access automated dashboards and real-time alerts which significantly ease the process of monitoring your CI/CD environment. Moreover, Netdata is capable of ingesting data from any Prometheus exporter, providing flexibility and scalability for growing operations.&lt;/p></description></item><item><title>Gnocchi</title><link>https://www.netdata.cloud/integrations/exporters/gnocchi/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/gnocchi/</guid><description/></item><item><title>Go applications (EXPVAR)</title><link>https://www.netdata.cloud/integrations/data-collection/applications/go-applications-expvar/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/go-applications-expvar/</guid><description/></item><item><title>Go Beyond 'Good Enough' Monitoring</title><link>https://www.netdata.cloud/ebook/ai-observability-new/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/ebook/ai-observability-new/</guid><description/></item><item><title>Go Beyond 'Good Enough' Monitoring</title><link>https://www.netdata.cloud/ebook/ai-observability/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/ebook/ai-observability/</guid><description/></item><item><title>Go-ethereum</title><link>https://www.netdata.cloud/integrations/data-collection/applications/go-ethereum/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/go-ethereum/</guid><description/></item><item><title>Go-ethereum Monitoring</title><link>https://www.netdata.cloud/monitoring-101/geth-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/geth-monitoring/</guid><description>&lt;h2 id="go-ethereum-monitoring">Go-ethereum Monitoring&lt;/h2>
&lt;h3 id="what-is-go-ethereum">What Is Go-ethereum?&lt;/h3>
&lt;p>Go-ethereum, often abbreviated as Geth, is one of the most popular implementations of the Ethereum protocol. Written in Go, it enables developers and enthusiasts to run a full Ethereum node, providing functionalities such as mining, creating contracts, and deploying DApps on the Ethereum blockchain. Go-ethereum is a fundamental tool within the blockchain ecosystem and is highly valued for its robustness and performance.&lt;/p>
&lt;h3 id="monitoring-go-ethereum-with-netdata">Monitoring Go-ethereum With Netdata&lt;/h3>
&lt;p>For anyone invested in the blockchain landscape, ensuring optimal performance and reliability of Go-ethereum nodes is crucial. With Netdata, a highly efficient monitoring solution, you can keep an eye on the performance and health of your Go-ethereum instances in real-time. As a powerful Go-ethereum monitoring tool, Netdata offers deep insights into blockchain operations, facilitating rapid troubleshooting and ensuring seamless operations.&lt;/p></description></item><item><title>Gobetween</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/gobetween/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/gobetween/</guid><description/></item><item><title>Gobetween Monitoring</title><link>https://www.netdata.cloud/monitoring-101/gobetween-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/gobetween-monitoring/</guid><description>&lt;h2 id="gobetween-monitoring">Gobetween Monitoring&lt;/h2>
&lt;h3 id="what-is-gobetween">What Is Gobetween?&lt;/h3>
&lt;p>Gobetween is a high-performance, protocol-agnostic load balancer designed for distributed network traffic management. As an open-source solution, it allows for the efficient routing of requests between multiple servers, offering improved performance, scalability, and reliability for modern web applications.&lt;/p>
&lt;h3 id="monitoring-gobetween-with-netdata">Monitoring Gobetween With Netdata&lt;/h3>
&lt;p>Monitor Gobetween effectively using Netdata&amp;rsquo;s robust monitoring capabilities. Netdata enhances the $name monitoring tool by utilizing an openmetrics (prometheus) exporter, which allows for seamless data collection without the need for a Prometheus server or Grafana setup. With Netdata, users can ingest data from any Prometheus exporter and gain access to automated dashboards, real-time alerts, and a wealth of metrics that facilitate optimized network traffic management and performance analysis.&lt;/p></description></item><item><title>Google BigQuery</title><link>https://www.netdata.cloud/integrations/exporters/google-bigquery/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/google-bigquery/</guid><description/></item><item><title>Google Cloud Platform</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/google-cloud-platform/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/google-cloud-platform/</guid><description/></item><item><title>Google Cloud Platform Monitoring</title><link>https://www.netdata.cloud/monitoring-101/gcp-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/gcp-monitoring/</guid><description>&lt;h2 id="google-cloud-platform-monitoring">Google Cloud Platform Monitoring&lt;/h2>
&lt;h3 id="what-is-google-cloud-platform">What Is Google Cloud Platform?&lt;/h3>
&lt;p>Google Cloud Platform (GCP) is a suite of cloud computing services provided by Google. It offers a wide range of services including computing power, storage, and machine learning capabilities, all of which are scalable to meet business needs. Whether you are deploying applications in Kubernetes, analyzing big data, or developing AI/ML models, GCP provides the infrastructure and tools to do so efficiently.&lt;/p></description></item><item><title>Google Cloud Pub Sub</title><link>https://www.netdata.cloud/integrations/exporters/google-cloud-pub-sub/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/google-cloud-pub-sub/</guid><description/></item><item><title>Google Colab Monitoring</title><link>https://www.netdata.cloud/monitoring-101/colab-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/colab-monitoring/</guid><description>&lt;p>Google Colab offers an exceptional platform for running Notebooks, developing machine learning models, and conducting various data science and analytics tasks. To optimize your experience, it is crucial to understand the performance of your Colab instance.&lt;/p>
&lt;h2 id="why-monitor-colab">Why monitor Colab?&lt;/h2>
&lt;p>Monitoring your Google Colab instance provides vital insights into its performance, enabling you to optimize your work and make the most of available resources. The benefits of monitoring your Colab instance include:&lt;/p></description></item><item><title>Google Pagespeed</title><link>https://www.netdata.cloud/integrations/data-collection/applications/google-pagespeed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/google-pagespeed/</guid><description/></item><item><title>Google Pagespeed Monitoring</title><link>https://www.netdata.cloud/monitoring-101/google_pagespeed-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/google_pagespeed-monitoring/</guid><description>&lt;h2 id="google-pagespeed-monitoring">Google Pagespeed Monitoring&lt;/h2>
&lt;h3 id="what-is-google-pagespeed">What Is Google Pagespeed?&lt;/h3>
&lt;p>Google Pagespeed is a widely-utilized performance tool that offers valuable insights into the speed and optimization of web pages. It&amp;rsquo;s an essential service in the realm of cloud computing, providing metrics that help ensure efficient website performance and user experience.&lt;/p>
&lt;h3 id="monitoring-google-pagespeed-with-netdata">Monitoring Google Pagespeed With Netdata&lt;/h3>
&lt;p>Monitor Google Pagespeed effectively with Netdata, using the robust capabilities of Netdata&amp;rsquo;s openmetrics (Prometheus) exporter. Netdata allows you to seamlessly ingest data from any Prometheus exporter, including Google Pagespeed, providing automated dashboards, alerts, and more. This toolkit eliminates the need for a Prometheus server or Grafana, streamlining your monitoring setup and enabling real-time insights.&lt;/p></description></item><item><title>Google Secret Manager</title><link>https://www.netdata.cloud/integrations/all/google-secret-manager/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/all/google-secret-manager/</guid><description/></item><item><title>Google Stackdriver</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/google-stackdriver/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/google-stackdriver/</guid><description/></item><item><title>Google Stackdriver Monitoring</title><link>https://www.netdata.cloud/monitoring-101/gcp_stackdriver-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/gcp_stackdriver-monitoring/</guid><description>&lt;h2 id="google-stackdriver-monitoring">Google Stackdriver Monitoring&lt;/h2>
&lt;h3 id="what-is-google-stackdriver">What Is Google Stackdriver?&lt;/h3>
&lt;p>Google Stackdriver, now part of Google Cloud&amp;rsquo;s operations suite, provides powerful monitoring, logging, and diagnostics services for applications hosted on Google Cloud Platform (GCP) and Amazon Web Services (AWS). With Stackdriver, you can proactively identify and troubleshoot issues across cloud services to ensure optimal performance and reliability.&lt;/p>
&lt;h3 id="monitoring-google-stackdriver-with-netdata">Monitoring Google Stackdriver With Netdata&lt;/h3>
&lt;p>To monitor Google Stackdriver, Netdata employs an openmetrics (prometheus) exporter, allowing users to gather diverse metrics without needing a Prometheus server or Grafana. The integration seamlessly ingests data, providing automated dashboards, alerts, and insights for real-time monitoring. You can explore the community-supported &lt;a href="https://github.com/prometheus-community/stackdriver_exporter">Google Stackdriver exporter&lt;/a> to start monitoring with accuracy and efficiency.&lt;/p></description></item><item><title>Gotify</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/gotify/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/gotify/</guid><description/></item><item><title>gpsd</title><link>https://www.netdata.cloud/integrations/data-collection/applications/gpsd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/gpsd/</guid><description/></item><item><title>GPSD Monitoring</title><link>https://www.netdata.cloud/monitoring-101/gpsd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/gpsd-monitoring/</guid><description>&lt;h2 id="gpsd-monitoring">GPSD Monitoring&lt;/h2>
&lt;h3 id="what-is-gpsd">What Is GPSD?&lt;/h3>
&lt;p>GPSD (GPS Daemon) is a service daemon that monitors one or more GPS or AIS receivers, which share data with various applications. It manages the communication between GPS devices and client applications, ensuring that all connected applications receive accurate and timely positional and navigational data.&lt;/p>
&lt;h3 id="monitoring-gpsd-with-netdata">Monitoring GPSD With Netdata&lt;/h3>
&lt;p>Netdata offers robust solutions for monitoring GPSD by leveraging the &lt;a href="https://github.com/natesales/gpsd-exporter">gpsd-exporter&lt;/a>. This tool uses openmetrics (Prometheus) to collect data seamlessly. With Netdata, you can ingest this data without the need for a dedicated Prometheus server or custom Grafana dashboards. Netdata automates dashboards, alerts, and allows you to keep a finger on the pulse of your GPS data systems effortlessly.&lt;/p></description></item><item><title>Grafana</title><link>https://www.netdata.cloud/integrations/data-collection/applications/grafana/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/grafana/</guid><description/></item><item><title>Grafana Monitoring</title><link>https://www.netdata.cloud/monitoring-101/grafana-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/grafana-monitoring/</guid><description>&lt;h2 id="grafana-monitoring">Grafana Monitoring&lt;/h2>
&lt;h3 id="what-is-grafana">What Is Grafana?&lt;/h3>
&lt;p>Grafana is a powerful open-source platform for monitoring and observability that allows you to query, visualize, and alert on your metrics. By enabling diverse data visualization options, Grafana helps DevOps, SREs, and IT admins understand complex datasets in real-time, which makes it an invaluable tool in modern IT environments.&lt;/p>
&lt;h3 id="monitoring-grafana-with-netdata">Monitoring Grafana With Netdata&lt;/h3>
&lt;p>Monitoring Grafana effectively is crucial for leveraging its full capabilities in visualizing data. Netdata offers a comprehensive solution to monitor Grafana using its OpenMetrics (Prometheus) exporter. Unlike traditional setups where a Prometheus server and Grafana dashboard need to be maintained separately, Netdata simplifies the process by allowing direct ingestion of Prometheus metrics, offering automated dashboards, alerts, and more without the need for additional server infrastructure.&lt;/p></description></item><item><title>Grand Junction Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/grand-junction-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/grand-junction-networks-snmp-traps/</guid><description/></item><item><title>Grandstream Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/grandstream-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/grandstream-networks-inc-snmp-traps/</guid><description/></item><item><title>Graphite</title><link>https://www.netdata.cloud/integrations/exporters/graphite/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/graphite/</guid><description/></item><item><title>Graylog Server</title><link>https://www.netdata.cloud/integrations/data-collection/applications/graylog-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/graylog-server/</guid><description/></item><item><title>Graylog Server Monitoring</title><link>https://www.netdata.cloud/monitoring-101/graylog-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/graylog-monitoring/</guid><description>&lt;h2 id="graylog-server-monitoring">Graylog Server Monitoring&lt;/h2>
&lt;h3 id="what-is-graylog-server">What Is Graylog Server?&lt;/h3>
&lt;p>Graylog Server is a powerful open-source log management platform designed to aggregate and analyze log data effectively. It centralizes log data from across an infrastructure, providing DevOps, SREs, developers, IT admins, and engineers with valuable insights for operational intelligence and security.&lt;/p>
&lt;h3 id="monitoring-graylog-server-with-netdata">Monitoring Graylog Server With Netdata&lt;/h3>
&lt;p>To monitor Graylog Server, Netdata uses an openmetrics exporter, commonly known as a Prometheus exporter. Netdata seamlessly integrates with any Prometheus exporter, allowing you to get instantaneous access to automated dashboards and alerts, without the need for a complete Prometheus setup or Grafana. This ease of use makes Netdata an excellent choice for tools for monitoring Graylog Server, ensuring efficient log management and analysis.&lt;/p></description></item><item><title>GreptimeDB</title><link>https://www.netdata.cloud/integrations/exporters/greptimedb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/greptimedb/</guid><description/></item><item><title>Gude Analog Und Digitalsysteme GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/gude-analog-und-digitalsysteme-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/gude-analog-und-digitalsysteme-gmbh-snmp-traps/</guid><description/></item><item><title>Gw Technologies Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/gw-technologies-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/gw-technologies-co-ltd-snmp-traps/</guid><description/></item><item><title>H3C SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/h3c-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/h3c-snmp-traps/</guid><description/></item><item><title>Hadoop Distributed File System (HDFS)</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/hadoop-distributed-file-system-hdfs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/hadoop-distributed-file-system-hdfs/</guid><description/></item><item><title>Hadoop Distributed File System (HDFS) Monitoring</title><link>https://www.netdata.cloud/monitoring-101/hdfs-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/hdfs-monitoring/</guid><description>&lt;h2 id="hadoop-distributed-file-system-hdfs-monitoring">Hadoop Distributed File System (HDFS) Monitoring&lt;/h2>
&lt;h3 id="what-is-hdfs">What Is HDFS?&lt;/h3>
&lt;p>The &lt;a href="https://hadoop.apache.org/docs/r1.2.1/hdfs_design.html">Hadoop Distributed File System (HDFS)&lt;/a> is the primary data storage system used by Hadoop applications. It&amp;rsquo;s designed to store very large data sets reliably and to stream those data sets at high bandwidth to user applications.&lt;/p>
&lt;h3 id="monitoring-hdfs-with-netdata">Monitoring HDFS With Netdata&lt;/h3>
&lt;p>Monitoring HDFS is pivotal for ensuring data reliability and system performance. Netdata&amp;rsquo;s HDFS monitoring tool provides deep insights by efficiently gathering metrics via JMX from HDFS daemons. This ensures that any anomaly within HDFS is quickly detected and addressed.&lt;/p></description></item><item><title>Halcyon Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/halcyon-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/halcyon-inc-snmp-traps/</guid><description/></item><item><title>Halon</title><link>https://www.netdata.cloud/integrations/data-collection/applications/halon/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/halon/</guid><description/></item><item><title>Halon Monitoring</title><link>https://www.netdata.cloud/monitoring-101/halon-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/halon-monitoring/</guid><description>&lt;h2 id="halon-monitoring">Halon Monitoring&lt;/h2>
&lt;h3 id="what-is-halon">What Is Halon?&lt;/h3>
&lt;p>Halon is a high-performance email security and delivery platform designed to streamline email management and enhance protection against spam and malicious content. Halon offers a modular environment whereby organizations can tailor configurations to address specific email handling and security needs effectively.&lt;/p>
&lt;h3 id="monitoring-halon-with-netdata">Monitoring Halon With Netdata&lt;/h3>
&lt;p>Monitoring Halon is crucial to ensure seamless email operations and security. With Netdata, you can monitor Halon using the powerful openmetrics (Prometheus) exporter. Netdata can ingest data from any Prometheus exporter, enabling automated dashboards, alerts, and more without the need for setting up a separate Prometheus server or Grafana. This integration streamlines the monitoring process and delivers real-time insights on Halon&amp;rsquo;s performance and health.&lt;/p></description></item><item><title>HANA</title><link>https://www.netdata.cloud/integrations/data-collection/databases/hana/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/hana/</guid><description/></item><item><title>HANA Monitoring</title><link>https://www.netdata.cloud/monitoring-101/hana-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/hana-monitoring/</guid><description>&lt;h2 id="hana-monitoring">HANA Monitoring&lt;/h2>
&lt;h3 id="what-is-hana">What Is HANA?&lt;/h3>
&lt;p>SAP HANA is an in-memory, column-oriented, relational database management system designed to handle both high transaction rates and complex query processing. As a powerful database server, it&amp;rsquo;s essential to monitor HANA to ensure efficient data storage and quick query performance.&lt;/p>
&lt;h3 id="monitoring-hana-with-netdata">Monitoring HANA With Netdata&lt;/h3>
&lt;p>To monitor HANA, Netdata leverages an openmetrics (Prometheus) exporter. By integrating with &lt;a href="https://github.com/jenningsloy318/hana_exporter">HANA Exporter&lt;/a>, Netdata collects and visualizes key metrics from your HANA database. One of the standout benefits of using Netdata is that it can ingest data from any Prometheus exporter, providing automated dashboards, alerts, and insights into your system&amp;rsquo;s performance—without the need for a Prometheus server or Grafana setup.&lt;/p></description></item><item><title>HAProxy</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/haproxy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/haproxy/</guid><description/></item><item><title>HAProxy 400 and 408 spikes: bad requests, request timeouts, and Slowloris</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-400-408-request-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-400-408-request-errors/</guid><description>&lt;h1 id="haproxy-400-and-408-spikes-bad-requests-request-timeouts-and-slowloris">HAProxy 400 and 408 spikes: bad requests, request timeouts, and Slowloris&lt;/h1>
&lt;p>Your frontend &lt;code>hrsp_4xx&lt;/code> counter just doubled, and it is not 404s. The spike is 400s and 408s, which means HAProxy itself is rejecting or timing out requests before they reach a backend. Unlike backend-generated 4xx, these errors say something about the bytes arriving at your frontend: malformed requests, protocol mismatches, oversized headers, or clients that open connections and never finish the request.&lt;/p></description></item><item><title>HAProxy 502 Bad Gateway: backend connection and response failures</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-502-bad-gateway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-502-bad-gateway/</guid><description>&lt;h1 id="haproxy-502-bad-gateway-backend-connection-and-response-failures">HAProxy 502 Bad Gateway: backend connection and response failures&lt;/h1>
&lt;p>Clients are getting 502 Bad Gateway from HAProxy. The backend may look healthy, health checks may be green, and yet some fraction of requests comes back with a 502 the application never logged. That is the signature: a 502 from HAProxy is usually generated by HAProxy itself, not passed through from the backend.&lt;/p>
&lt;p>HAProxy emits a 502 in two situations. First, it could not establish or use a connection to the backend server at all. Second, it connected fine, but the server sent an invalid, truncated, or incomplete response, or closed the connection mid-response. These are different failure modes with different fixes, and the fastest path to resolution is separating them before touching anything.&lt;/p></description></item><item><title>HAProxy 503 Service Unavailable: no server is available to handle this request</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-503-service-unavailable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-503-service-unavailable/</guid><description>&lt;h1 id="haproxy-503-service-unavailable-no-server-is-available-to-handle-this-request">HAProxy 503 Service Unavailable: no server is available to handle this request&lt;/h1>
&lt;p>Clients are getting &lt;code>503 Service Unavailable&lt;/code> with the message &lt;code>no server is available to handle this request&lt;/code>. HAProxy generated this error. Unlike a 500 from your application or a 502 from a broken backend response, a 503 with this message never touched a backend server. HAProxy looked at its own state, concluded it had nowhere to send the request, and answered on its own.&lt;/p></description></item><item><title>HAProxy 504 Gateway Timeout: timeout server and the backend timeout cascade</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-504-gateway-timeout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-504-gateway-timeout/</guid><description>&lt;h1 id="haproxy-504-gateway-timeout-timeout-server-and-the-backend-timeout-cascade">HAProxy 504 Gateway Timeout: timeout server and the backend timeout cascade&lt;/h1>
&lt;p>Your frontend error rate jumps. Clients report &amp;ldquo;504 Gateway Timeout.&amp;rdquo; HAProxy looks alive: the process is running, servers are UP, health checks are green. Yet requests are dying on a timer.&lt;/p>
&lt;p>This is the backend timeout cascade, and the 504 is its final symptom. The servers are reachable (TCP accepts succeed), but the application behind them has slowed past the &lt;code>timeout server&lt;/code> budget. HAProxy holds the client connection, waits, gives up, and generates a 504 itself.&lt;/p></description></item><item><title>HAProxy 5xx delta: telling HAProxy-generated errors from backend errors</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-frontend-backend-5xx-delta/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-frontend-backend-5xx-delta/</guid><description>&lt;h1 id="haproxy-5xx-delta-telling-haproxy-generated-errors-from-backend-errors">HAProxy 5xx delta: telling HAProxy-generated errors from backend errors&lt;/h1>
&lt;p>Your frontend &lt;code>hrsp_5xx&lt;/code> counter is climbing and users are seeing errors. The first question that decides the next hour of your incident response: is HAProxy generating these errors, or is it passing through errors your backend servers produced?&lt;/p>
&lt;p>HAProxy counts 5xx responses in both places. Frontend rows count every 5xx the client saw, including responses HAProxy fabricated itself. Backend and server rows count only 5xx responses that actually came back from a server. The difference between the two answers &amp;ldquo;is the backend failing or is HAProxy failing?&amp;rdquo; without walking every backend one by one.&lt;/p></description></item><item><title>HAProxy backend losing servers: active server count and cascade risk</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-backend-losing-servers-capacity/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-backend-losing-servers-capacity/</guid><description>&lt;h1 id="haproxy-backend-losing-servers-active-server-count-and-cascade-risk">HAProxy backend losing servers: active server count and cascade risk&lt;/h1>
&lt;p>A backend can lose half its servers and HAProxy will still report the BACKEND aggregate row as UP. As long as one server is available, the rollup status says traffic is being served, and status-based dashboards stay green. Meanwhile the surviving servers are absorbing redistributed load, their session counts are climbing, and you are one health check failure away from a full 503 outage for that backend.&lt;/p></description></item><item><title>HAProxy backend queue building (qcur): requests waiting for a free server slot</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-backend-queue-building-qcur/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-backend-queue-building-qcur/</guid><description>&lt;h1 id="haproxy-backend-queue-building-qcur-requests-waiting-for-a-free-server-slot">HAProxy backend queue building (qcur): requests waiting for a free server slot&lt;/h1>
&lt;p>You are looking at HAProxy stats and &lt;code>qcur&lt;/code> is nonzero on a backend, or worse, climbing. Clients have not seen errors yet, but latency is drifting up. That instinct that something is wrong is correct: &lt;code>qcur&lt;/code> is the earliest signal that backend capacity is exhausted, and it shows up before any error counter moves.&lt;/p>
&lt;p>&lt;code>qcur&lt;/code> is the instantaneous count of connections sitting in a queue because every server that could take them is already at its &lt;code>maxconn&lt;/code> limit. Those requests are not being served. They are parked, consuming their client&amp;rsquo;s timeout budget, waiting for a slot to free up. If the queue grows faster than slots free up, the outcome is queue timeouts and 503s.&lt;/p></description></item><item><title>HAProxy bandwidth (bin/bout): bytes in, bytes out, and capacity planning</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-bandwidth-bytes-in-out/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-bandwidth-bytes-in-out/</guid><description>&lt;h1 id="haproxy-bandwidth-binbout-bytes-in-bytes-out-and-capacity-planning">HAProxy bandwidth (bin/bout): bytes in, bytes out, and capacity planning&lt;/h1>
&lt;p>Every row in HAProxy&amp;rsquo;s stats output carries two cumulative byte counters: &lt;code>bin&lt;/code> (bytes in) and &lt;code>bout&lt;/code> (bytes out). They are the ground truth for how much data actually moves through the proxy, per frontend, per backend, and per server. Most teams only look at them when a network link saturates. Used properly, they are a baseline signal, an anomaly detector, and the raw material for bandwidth capacity planning.&lt;/p></description></item><item><title>HAProxy compression overhead: comp_byp, CPU cost, and when compression backfires</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-compression-cpu-overhead/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-compression-cpu-overhead/</guid><description>&lt;h1 id="haproxy-compression-overhead-comp_byp-cpu-cost-and-when-compression-backfires">HAProxy compression overhead: comp_byp, CPU cost, and when compression backfires&lt;/h1>
&lt;p>HTTP compression in HAProxy is a trade: CPU cycles on the proxy in exchange for fewer bytes on the wire. When it works, text responses shrink by 50 to 80 percent and clients on slow links see faster page loads. When it backfires, the proxy burns event-loop CPU compressing content that does not compress, or it sheds compression work under load and you lose the bandwidth savings exactly when traffic is highest.&lt;/p></description></item><item><title>HAProxy connect time (ctime) high: network latency and backend accept-queue overflow</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-connect-time-ctime-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-connect-time-ctime-high/</guid><description>&lt;h1 id="haproxy-connect-time-ctime-high-network-latency-and-backend-accept-queue-overflow">HAProxy connect time (ctime) high: network latency and backend accept-queue overflow&lt;/h1>
&lt;p>Your HAProxy backend shows &lt;code>ctime&lt;/code> climbing from a steady 1-2ms to tens or hundreds of milliseconds. Total request latency (&lt;code>ttime&lt;/code>) rises with it, and if the trend continues, connect timeouts (&lt;code>econ&lt;/code> increments, 502/504s on the frontend) are next. The connect phase of the HAProxy-to-backend path is degrading, and it almost always means one of three things: the network path is congested or losing packets, the backend&amp;rsquo;s kernel accept queue is overflowing, or the backend is too CPU-starved to answer the handshake promptly.&lt;/p></description></item><item><title>HAProxy connection errors (econ): backend connections refused, timed out, or reset</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-connection-errors-econ/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-connection-errors-econ/</guid><description>&lt;h1 id="haproxy-connection-errors-econ-backend-connections-refused-timed-out-or-reset">HAProxy connection errors (econ): backend connections refused, timed out, or reset&lt;/h1>
&lt;p>Your HAProxy backend shows a rising &lt;code>econ&lt;/code> counter. Clients may not see errors yet, because HAProxy retries failed connections and redispatches to other servers, but the counter means HAProxy is trying to open TCP connections to a backend server and failing. The connection is being refused, timing out, or getting reset.&lt;/p>
&lt;p>&lt;code>econ&lt;/code> is one of the highest-signal error counters HAProxy exposes, and one of the most misread. It is cumulative, it only increments on new connection attempts (so connection reuse can mask a real problem for hours), and it lumps three very different failure modes into one number. Telling &amp;ldquo;the server process is dead&amp;rdquo; apart from &amp;ldquo;a firewall is silently dropping SYNs&amp;rdquo; requires correlating &lt;code>econ&lt;/code> with &lt;code>ctime&lt;/code> and a few other signals.&lt;/p></description></item><item><title>HAProxy connection reuse dropping: http-reuse, pooling, and handshake overhead</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-connection-reuse-dropping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-connection-reuse-dropping/</guid><description>&lt;h1 id="haproxy-connection-reuse-dropping-http-reuse-pooling-and-handshake-overhead">HAProxy connection reuse dropping: http-reuse, pooling, and handshake overhead&lt;/h1>
&lt;p>Backend latency crept up, CPU on the HAProxy host is climbing, and &lt;code>ctime&lt;/code> went from zero to measurable milliseconds, but the backends are healthy and the network looks fine. Before you blame the application or the network, check the one ratio almost nobody graphs: connection reuse, &lt;code>reuse / (connect + reuse)&lt;/code> on the backend. If it fell off a cliff, HAProxy stopped reusing pooled backend connections and is now paying a full TCP (and possibly TLS) handshake for every request.&lt;/p></description></item><item><title>HAProxy counter resets on reload: phantom drops and spikes in your metrics</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-counter-resets-on-reload/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-counter-resets-on-reload/</guid><description>&lt;h1 id="haproxy-counter-resets-on-reload-phantom-drops-and-spikes-in-your-metrics">HAProxy counter resets on reload: phantom drops and spikes in your metrics&lt;/h1>
&lt;p>Your HAProxy 5xx dashboard just showed a cliff-edge drop to zero, then a vertical spike an hour later. No user complaints. No backend incidents. No deploys. If this pattern repeats on a schedule, or right after config changes, you are almost certainly looking at counter resets from HAProxy reloads, not real traffic events.&lt;/p>
&lt;p>Every time HAProxy reloads its configuration, a new worker process starts. That new process has its own in-memory statistics, and every cumulative counter starts at zero: &lt;code>hrsp_5xx&lt;/code>, &lt;code>econ&lt;/code>, &lt;code>eresp&lt;/code>, &lt;code>bin&lt;/code>, &lt;code>bout&lt;/code>, &lt;code>wretr&lt;/code>, &lt;code>wredis&lt;/code>, &lt;code>chkfail&lt;/code>, the SSL counters, all of them. A reload is a new process for stats purposes, not a continuation of the old one.&lt;/p></description></item><item><title>HAProxy DNS resolver failure: stale IPs and silently misrouted traffic</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-dns-resolver-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-dns-resolver-failure/</guid><description>&lt;h1 id="haproxy-dns-resolver-failure-stale-ips-and-silently-misrouted-traffic">HAProxy DNS resolver failure: stale IPs and silently misrouted traffic&lt;/h1>
&lt;p>Your backend servers report UP. Health checks pass. The stats page is green. And traffic is going to IP addresses that belong to pods deleted an hour ago, or to a node that now runs something else entirely.&lt;/p>
&lt;p>This is the HAProxy DNS resolver failure mode, and it only bites deployments that resolve backends dynamically: &lt;code>server-template&lt;/code> with Consul DNS, Kubernetes headless services, SRV records, or any setup where server addresses come from DNS instead of static config lines. When the resolver times out or a TTL expires without a successful re-resolution, HAProxy keeps routing to the last IPs it knew about. There is no stats CSV counter for this. &lt;code>show resolvers&lt;/code> is the only direct view into it, and most teams never look.&lt;/p></description></item><item><title>HAProxy DroppedLogs: losing log lines exactly when you need them</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-dropped-logs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-dropped-logs/</guid><description>&lt;h1 id="haproxy-droppedlogs-losing-log-lines-exactly-when-you-need-them">HAProxy DroppedLogs: losing log lines exactly when you need them&lt;/h1>
&lt;p>You are mid-incident. Traffic is spiking, 5xx rates are climbing, and you go to the HAProxy logs for per-request timing and termination codes. The logs are not there. Whole seconds of traffic are missing, right at the peak. Then you notice &lt;code>DroppedLogs&lt;/code> in &lt;code>show info&lt;/code> has been incrementing for hours and nobody was watching it.&lt;/p>
&lt;p>&lt;code>DroppedLogs&lt;/code> counts log messages HAProxy tried to emit but could not deliver. The cause is almost never HAProxy itself: it is the log pipeline downstream. A syslog daemon that cannot keep up, a socket buffer that overflowed, a log target that stalled. There is no error in the HAProxy log, because the line that would have told you is one of the ones that got dropped.&lt;/p></description></item><item><title>HAProxy ephemeral port exhaustion: TIME_WAIT and backend connection churn</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-ephemeral-port-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-ephemeral-port-exhaustion/</guid><description>&lt;h1 id="haproxy-ephemeral-port-exhaustion-time_wait-and-backend-connection-churn">HAProxy ephemeral port exhaustion: TIME_WAIT and backend connection churn&lt;/h1>
&lt;p>Your backends are healthy. Every server is UP, health checks pass, the application team sees nothing wrong on their side. Yet HAProxy is logging failed backend connections, &lt;code>econ&lt;/code> counters are climbing, and clients see intermittent 502s or stalls. On the HAProxy host, tens of thousands of sockets sit in TIME_WAIT.&lt;/p>
&lt;p>This is ephemeral port exhaustion on the backend-facing side of HAProxy. Every new TCP connection from HAProxy to a backend consumes one local ephemeral port. When the connection closes, that port is held in TIME_WAIT for 60 seconds before it can be reused. If the rate of new backend connections is high enough, the port range empties and &lt;code>connect()&lt;/code> starts failing with &lt;code>EADDRNOTAVAIL&lt;/code>. HAProxy reports these as connection errors even though nothing is wrong with the backend.&lt;/p></description></item><item><title>HAProxy health check L4TOUT and L4CON: server marked DOWN and unreachable</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-health-check-l4tout-l4con/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-health-check-l4tout-l4con/</guid><description>&lt;h1 id="haproxy-health-check-l4tout-and-l4con-server-marked-down-and-unreachable">HAProxy health check L4TOUT and L4CON: server marked DOWN and unreachable&lt;/h1>
&lt;p>You look at the HAProxy stats page (or your monitoring) and a server row is red. The &lt;code>last_chk&lt;/code> field says &lt;code>L4TOUT&lt;/code> or &lt;code>L4CON&lt;/code>, &lt;code>status&lt;/code> says &lt;code>DOWN&lt;/code>, and HAProxy has stopped routing traffic to that server. Both codes mean the failure happened before any HTTP, before any application code, at the plain TCP layer.&lt;/p>
&lt;p>&lt;code>L4TOUT&lt;/code> means HAProxy opened a socket and the TCP connect did not complete within the check timeout. &lt;code>L4CON&lt;/code> means the connect attempt failed immediately with an error, typically &amp;ldquo;Connection refused&amp;rdquo; (TCP RST) or &amp;ldquo;No route to host&amp;rdquo; (ICMP). The server is unreachable at the TCP layer, and after &lt;code>fall&lt;/code> consecutive failures (default 3), HAProxy marks it DOWN and redistributes its traffic.&lt;/p></description></item><item><title>HAProxy health check L7STS: server DOWN on the wrong HTTP status</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-health-check-l7sts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-health-check-l7sts/</guid><description>&lt;h1 id="haproxy-health-check-l7sts-server-down-on-the-wrong-http-status">HAProxy health check L7STS: server DOWN on the wrong HTTP status&lt;/h1>
&lt;p>A backend server flips to DOWN in the stats page, &lt;code>last_chk&lt;/code> shows &lt;code>L7STS&lt;/code>, and the log line says something like &lt;code>Layer7 wrong status, code: 503, info: &amp;quot;Service Unavailable&amp;quot;&lt;/code>. The server process is running. The port is open. You can curl it by hand and get a response. But HAProxy has pulled it out of rotation and traffic is concentrating on the survivors.&lt;/p></description></item><item><title>HAProxy health checks green but the application is broken: when UP does not mean healthy</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-health-check-green-app-broken/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-health-check-green-app-broken/</guid><description>&lt;h1 id="haproxy-health-checks-green-but-the-application-is-broken-when-up-does-not-mean-healthy">HAProxy health checks green but the application is broken: when UP does not mean healthy&lt;/h1>
&lt;p>Every server in the backend shows UP. Health checks are passing. And users are getting 500s. This is one of the most common HAProxy postmortem findings: &amp;ldquo;health checks said everything was fine.&amp;rdquo;&lt;/p>
&lt;p>The failure is conceptual, not a bug. HAProxy&amp;rsquo;s health check engine runs independently of traffic and only proves that the specific probe you configured succeeds. A TCP check proves the port accepts connections. An HTTP check to &lt;code>/health&lt;/code> proves that one URL returns the expected status. Neither proves that real requests, which exercise the database, auth, and downstream dependencies, actually work.&lt;/p></description></item><item><title>HAProxy hot-thread skew: when average Idle_pct hides a saturated thread</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-thread-skew-show-activity/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-thread-skew-show-activity/</guid><description>&lt;h1 id="haproxy-hot-thread-skew-when-average-idle_pct-hides-a-saturated-thread">HAProxy hot-thread skew: when average Idle_pct hides a saturated thread&lt;/h1>
&lt;p>Latency is climbing, throughput has plateaued below what the hardware should deliver, and &lt;code>show info&lt;/code> reports &lt;code>Idle_pct&lt;/code> at a comfortable 60 or 70 percent. System CPU shows one core pegged while the others coast. The bottleneck is not HAProxy as a whole. It is one thread.&lt;/p>
&lt;p>With &lt;code>nbthread &amp;gt; 1&lt;/code> (the default since HAProxy 2.0, where &lt;code>nbthread&lt;/code> is set to the number of available CPUs), work is spread across per-thread event loops. &lt;code>Idle_pct&lt;/code> is a process-level average across all of them. One thread pinned at 100% while the rest sit at 80% idle averages out to &amp;ldquo;looks fine,&amp;rdquo; but that hot thread caps overall throughput: every connection it owns gets processed late, and no other thread can take them over.&lt;/p></description></item><item><title>HAProxy Idle_pct dropping: event-loop CPU saturation and the busy-polling trap</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-idle-pct-low/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-idle-pct-low/</guid><description>&lt;h1 id="haproxy-idle_pct-dropping-event-loop-cpu-saturation-and-the-busy-polling-trap">HAProxy Idle_pct dropping: event-loop CPU saturation and the busy-polling trap&lt;/h1>
&lt;p>Your HAProxy latency is climbing, nothing is queuing on the backends, and the one metric that looks wrong is &lt;code>Idle_pct&lt;/code> sliding toward zero. Or worse: &lt;code>Idle_pct&lt;/code> has been pinned at zero for weeks and you only just noticed because someone finally asked what it meant. Both situations are common, and they have very different explanations.&lt;/p>
&lt;p>&lt;code>Idle_pct&lt;/code> is HAProxy&amp;rsquo;s self-reported measure of event-loop headroom: the share of time the loop spends waiting in &lt;code>poll()&lt;/code> versus processing events. It measures HAProxy CPU headroom, not system CPU. A host at 50% system CPU with &lt;code>Idle_pct&lt;/code> at 80% is fine. A host at 50% system CPU with &lt;code>Idle_pct&lt;/code> at 10% means HAProxy itself is saturated. That distinction is the entire point of the metric, and it is also where the traps live.&lt;/p></description></item><item><title>HAProxy in TCP mode: the HTTP signals you lose and what to watch instead</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-tcp-mode-monitoring-blind-spots/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-tcp-mode-monitoring-blind-spots/</guid><description>&lt;h1 id="haproxy-in-tcp-mode-the-http-signals-you-lose-and-what-to-watch-instead">HAProxy in TCP mode: the HTTP signals you lose and what to watch instead&lt;/h1>
&lt;p>When a frontend and its backend run with &lt;code>mode tcp&lt;/code>, HAProxy stops parsing the protocol inside the connection. It forwards bytes and nothing more. That is often correct: databases, gRPC, Redis, TLS passthrough, or protocols HAProxy does not understand. But it amputates a large part of the monitoring surface most HAProxy alerting assumes exists.&lt;/p>
&lt;p>The failure mode this creates is not an outage; it is a blind spot. Dashboards that worked in HTTP mode start reporting zeros that look like &amp;ldquo;no traffic&amp;rdquo; or &amp;ldquo;no errors.&amp;rdquo; Health checks stay green while the application behind them is broken. Latency alerts keyed on &lt;code>rtime&lt;/code> never fire again because &lt;code>rtime&lt;/code> is permanently zero. The proxy is fine; your visibility into what flows through it is not.&lt;/p></description></item><item><title>HAProxy kernel accept-queue overflow: silent SYN drops and somaxconn</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-listen-queue-overflow-somaxconn/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-listen-queue-overflow-somaxconn/</guid><description>&lt;h1 id="haproxy-kernel-accept-queue-overflow-silent-syn-drops-and-somaxconn">HAProxy kernel accept-queue overflow: silent SYN drops and somaxconn&lt;/h1>
&lt;p>Clients report connection timeouts. You open the HAProxy stats page and everything looks fine: session rate is normal, no 5xx spike, servers are UP, Idle_pct is healthy. The network team sees nothing. The clients insist the load balancer is dropping them.&lt;/p>
&lt;p>Both sides are right. The kernel is silently dropping SYN packets in the accept queue before they ever reach HAProxy&amp;rsquo;s event loop. From HAProxy&amp;rsquo;s perspective, those connections never existed. There is no log line, no error counter, no stat field that increments. The only evidence lives in two kernel counters, &lt;code>ListenOverflows&lt;/code> and &lt;code>ListenDrops&lt;/code>, that almost nobody graphs.&lt;/p></description></item><item><title>HAProxy maxconn hierarchy: global, frontend, backend, and per-server limits</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-maxconn-hierarchy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-maxconn-hierarchy/</guid><description>&lt;h1 id="haproxy-maxconn-hierarchy-global-frontend-backend-and-per-server-limits">HAProxy maxconn hierarchy: global, frontend, backend, and per-server limits&lt;/h1>
&lt;p>HAProxy does not have one connection limit. It has four, enforced independently at the global, frontend, backend, and per-server levels. Any one of them can throttle traffic while the other three show headroom. This is why &amp;ldquo;global maxconn is at 30%&amp;rdquo; tells you almost nothing during an incident: the request path can be blocked at a per-server limit you never set, or queuing behind a backend constraint you did not know existed.&lt;/p></description></item><item><title>HAProxy maxconn reached: new connections queued and rejected</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-maxconn-reached/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-maxconn-reached/</guid><description>&lt;h1 id="haproxy-maxconn-reached-new-connections-queued-and-rejected">HAProxy maxconn reached: new connections queued and rejected&lt;/h1>
&lt;p>HAProxy is refusing work. Clients see connection timeouts or resets, load tests fail at a suspiciously round concurrency number, and yet the HAProxy process looks healthy: CPU is fine, memory is fine, the process is up, health checks pass. The stats tell the real story: &lt;code>CurrConns&lt;/code> is flat at &lt;code>Maxconn&lt;/code>, or &lt;code>scur&lt;/code> is pinned at &lt;code>slim&lt;/code> on a frontend, and it does not move.&lt;/p></description></item><item><title>HAProxy memory growth and OOM: connection buffers, pools, and memory ceilings</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-memory-growth-oom/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-memory-growth-oom/</guid><description>&lt;h1 id="haproxy-memory-growth-and-oom-connection-buffers-pools-and-memory-ceilings">HAProxy memory growth and OOM: connection buffers, pools, and memory ceilings&lt;/h1>
&lt;p>The symptom usually arrives as a dead process, not a slow one. HAProxy was proxying fine, then the PID changed, connections dropped for a few seconds, and &lt;code>dmesg&lt;/code> shows the OOM killer picked haproxy as the victim. Or, before it gets that far, RSS climbs day over day and nobody knows whether it is a leak or just load.&lt;/p></description></item><item><title>HAProxy Monitoring</title><link>https://www.netdata.cloud/monitoring-101/haproxy-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/haproxy-monitoring/</guid><description>&lt;h2 id="haproxy-monitoring">HAProxy Monitoring&lt;/h2>
&lt;h3 id="what-is-haproxy">What Is HAProxy?&lt;/h3>
&lt;p>HAProxy, short for High Availability Proxy, is a popular open-source software widely used for load balancing and proxying for TCP and HTTP-based applications. It is known for its reliability, performance, and feature-rich nature, making it a staple in server infrastructure for distributing workloads efficiently.&lt;/p>
&lt;h3 id="monitoring-haproxy-with-netdata">Monitoring HAProxy With Netdata&lt;/h3>
&lt;p>Monitoring HAProxy is crucial for ensuring your application infrastructure performs optimally and remains highly reliable. &lt;a href="https://www.netdata.cloud/">Netdata&lt;/a> offers comprehensive and real-time monitoring of HAProxy, delivering rich visualization of health and performance metrics. By leveraging such a tool, you can observe metrics like response times, session rates, and network throughput in a live, interactive interface. You can get started by &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">checking out the Live Demo&lt;/a> or &lt;a href="https://app.netdata.cloud/?utm_source=website&amp;amp;utm_content=monitoring101">signing up for a free trial&lt;/a>.&lt;/p></description></item><item><title>HAProxy monitoring checklist: the signals every production proxy needs</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-monitoring-checklist/</guid><description>&lt;h1 id="haproxy-monitoring-checklist-the-signals-every-production-proxy-needs">HAProxy monitoring checklist: the signals every production proxy needs&lt;/h1>
&lt;p>Most HAProxy outages are visible in its own stats long before users notice. The problem is that the stats socket exposes dozens of fields, and teams usually wire up the five metrics their dashboard template shipped with and stop there. Then a reload storm, a retry storm, or a silent stick-table overflow teaches them what they were missing.&lt;/p>
&lt;p>This checklist organizes the signals that catch production failures into four maturity levels: survival, operational, mature, and expert. Each level builds on the previous one. You do not need to reach expert on day one, but you should know which level you are at, because that determines which failure modes are invisible to you.&lt;/p></description></item><item><title>HAProxy monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-monitoring-maturity-model/</guid><description>&lt;h1 id="haproxy-monitoring-maturity-model-from-survival-to-expert">HAProxy monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most HAProxy monitoring setups fall into one of two states: a process check and a prayer, or a wall of charts nobody reads. Neither answers the question that matters during an incident: is HAProxy the problem, the messenger, or the victim.&lt;/p>
&lt;p>This article defines four maturity levels for HAProxy monitoring: Survival, Operational, Mature, and Expert. Each level adds signals that catch a class of failures the previous level cannot see. The goal is not to collect everything. The goal is to know, at any moment, which questions your monitoring can answer and which it cannot.&lt;/p></description></item><item><title>HAProxy peers sync broken: divergent stick tables across instances</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-peers-sync-broken/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-peers-sync-broken/</guid><description>&lt;h1 id="haproxy-peers-sync-broken-divergent-stick-tables-across-instances">HAProxy peers sync broken: divergent stick tables across instances&lt;/h1>
&lt;p>You have two or more HAProxy instances fronting the same service, and their behavior no longer matches. Rate limiting triggers on one node but lets the same client through on another. A failover happens and every sticky session evaporates, even though you configured stick tables and a peers section specifically to survive that event. Or you compare &lt;code>show table&lt;/code> output across instances and the entry counts are wildly different when they should be near-identical.&lt;/p></description></item><item><title>HAProxy per-source-IP rate limiting: stick tables, conn_rate, and http_req_rate</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-rate-limiting-per-source-ip/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-rate-limiting-per-source-ip/</guid><description>&lt;h1 id="haproxy-per-source-ip-rate-limiting-stick-tables-conn_rate-and-http_req_rate">HAProxy per-source-IP rate limiting: stick tables, conn_rate, and http_req_rate&lt;/h1>
&lt;p>Per-source-IP rate limiting in HAProxy is built on stick tables: in-memory key-value stores that track counters per client key (usually the source IP) and expose those counters to ACLs. When it breaks, it breaks silently. A full table, a mismatched expire, or an IPv6 client population can disable your limiting without a single error in the logs, and most teams find out only after an abuse incident.&lt;/p></description></item><item><title>HAProxy PoolFailed: buffer starvation and memory-allocator failures</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-poolfailed-buffer-starvation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-poolfailed-buffer-starvation/</guid><description>&lt;h1 id="haproxy-poolfailed-buffer-starvation-and-memory-allocator-failures">HAProxy PoolFailed: buffer starvation and memory-allocator failures&lt;/h1>
&lt;p>&lt;code>PoolFailed&lt;/code> in &lt;code>show info&lt;/code> counts failed internal memory-pool allocations since process start. It is normally zero. Any nonzero value means HAProxy tried to allocate a buffer, connection, or session object and could not get one. Work was dropped or stalled because of it.&lt;/p>
&lt;p>The tricky part is how this surfaces. Buffer starvation rarely looks like a memory problem from the outside. It looks like mysterious latency: &lt;code>rtime&lt;/code> and backend health are fine, CPU has headroom, the network is clean, and yet clients see stalls and sporadic failures. &lt;code>PoolFailed&lt;/code> is one of the few direct signals that HAProxy itself is refusing work because it ran out of an internal resource.&lt;/p></description></item><item><title>HAProxy queue time (qtime) rising: the earliest backend-saturation signal</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-queue-time-qtime-rising/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-queue-time-qtime-rising/</guid><description>&lt;h1 id="haproxy-queue-time-qtime-rising-the-earliest-backend-saturation-signal">HAProxy queue time (qtime) rising: the earliest backend-saturation signal&lt;/h1>
&lt;p>You are looking at a backend whose &lt;code>qtime&lt;/code> has crept from zero to a few hundred milliseconds, or a few seconds, and nothing is &amp;ldquo;down&amp;rdquo; yet. Health checks are green, error rates are flat, and no alert has fired. That is the point of qtime: it moves before the failure signals.&lt;/p>
&lt;p>Nonzero qtime means requests spent time in HAProxy&amp;rsquo;s queue because every usable server connection slot was busy when they arrived. It is usually the first measurable symptom of backend saturation: ahead of 503s, ahead of &lt;code>timeout server&lt;/code> expiries, ahead of user complaints.&lt;/p></description></item><item><title>HAProxy reload storm: lingering processes leaking file descriptors and memory</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-reload-storm-process-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-reload-storm-process-leak/</guid><description>&lt;h1 id="haproxy-reload-storm-lingering-processes-leaking-file-descriptors-and-memory">HAProxy reload storm: lingering processes leaking file descriptors and memory&lt;/h1>
&lt;p>You log into the HAProxy host because memory and file descriptor usage have been climbing for days, and you find a dozen haproxy processes running. Only one is serving traffic. The rest are old workers stuck in soft-stop, each holding open connections, connection buffers, and file descriptors it will never release. Traffic looks fine, per-process metrics look fine, and the host is still slowly running out of resources.&lt;/p></description></item><item><title>HAProxy response errors (eresp): truncated and invalid backend responses</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-response-errors-eresp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-response-errors-eresp/</guid><description>&lt;h1 id="haproxy-response-errors-eresp-truncated-and-invalid-backend-responses">HAProxy response errors (eresp): truncated and invalid backend responses&lt;/h1>
&lt;p>The &lt;code>eresp&lt;/code> counter is climbing on a backend, and clients are seeing broken pages, truncated downloads, or intermittent 502s. The backend servers report healthy. Health checks are green. The 5xx rate, if it moved at all, does not explain the volume of user complaints.&lt;/p>
&lt;p>This is the failure mode &lt;code>eresp&lt;/code> exists to capture: HAProxy connected to the backend successfully, sent the request, and then received a response it could not treat as valid. Either the bytes were not parseable HTTP, or the server closed the connection before the response finished. That is a protocol-level failure, not an application returning an error status code, and it needs a different investigation path than a 5xx spike.&lt;/p></description></item><item><title>HAProxy response time (rtime) climbing: slow backends behind the proxy</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-response-time-rtime-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-response-time-rtime-high/</guid><description>&lt;h1 id="haproxy-response-time-rtime-climbing-slow-backends-behind-the-proxy">HAProxy response time (rtime) climbing: slow backends behind the proxy&lt;/h1>
&lt;p>Your HAProxy dashboards show &lt;code>rtime&lt;/code> creeping up on one or more backends. Clients have not started timing out yet, or maybe they just have. HAProxy itself looks fine: CPU is reasonable, no servers are DOWN, error rates are flat or only slightly elevated. But the average time between HAProxy forwarding a request and the backend&amp;rsquo;s first response byte is growing, and it is not coming back down.&lt;/p></description></item><item><title>HAProxy retries and redispatches (wretr/wredis): the backend instability nobody watches</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-retries-redispatches-wretr-wredis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-retries-redispatches-wretr-wredis/</guid><description>&lt;h1 id="haproxy-retries-and-redispatches-wretrwredis-the-backend-instability-nobody-watches">HAProxy retries and redispatches (wretr/wredis): the backend instability nobody watches&lt;/h1>
&lt;p>Your dashboards are clean. No 5xx spike, latency averages look normal, every backend server is UP. Then the backend falls over in what feels like seconds, and the postmortem shows the application was flapping for six hours before the outage. The warning was there the whole time, sitting in two counters almost nobody charts: &lt;code>wretr&lt;/code> and &lt;code>wredis&lt;/code>.&lt;/p>
&lt;p>These counters record HAProxy&amp;rsquo;s built-in compensation mechanism at work. When a connection to a backend server fails, HAProxy does not immediately return an error to the client. It retries. If &lt;code>option redispatch&lt;/code> is set, it can also give up on the original server and dispatch the request to a different one. From the client&amp;rsquo;s perspective the request succeeds, possibly a bit slower. From the proxy&amp;rsquo;s perspective, something in your backend just failed and got papered over.&lt;/p></description></item><item><title>HAProxy run queue rising: task backlog and scheduler pressure</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-run-queue-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-run-queue-saturation/</guid><description>&lt;h1 id="haproxy-run-queue-rising-task-backlog-and-scheduler-pressure">HAProxy run queue rising: task backlog and scheduler pressure&lt;/h1>
&lt;p>HAProxy&amp;rsquo;s event loop processes every connection, request, health check, and timer as tasks scheduled across its threads. When tasks arrive faster than threads can drain them, they pile up in the run queue. That backlog is the earliest internal signal that HAProxy is CPU-saturated: it appears before latency spikes, before backend queuing, and before client timeouts.&lt;/p>
&lt;p>It is also the signal to use when &lt;code>Idle_pct&lt;/code> is meaningless. With &lt;code>busy-polling&lt;/code> enabled, the event loop never sleeps, so &lt;code>Idle_pct&lt;/code> sits near zero by design. The &lt;code>Run_queue&lt;/code> field from &lt;code>show info&lt;/code> and the per-thread counters from &lt;code>show activity&lt;/code> are the replacement.&lt;/p></description></item><item><title>HAProxy scur approaching slim: the concurrent-session saturation signal</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-current-sessions-near-maxconn/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-current-sessions-near-maxconn/</guid><description>&lt;h1 id="haproxy-scur-approaching-slim-the-concurrent-session-saturation-signal">HAProxy scur approaching slim: the concurrent-session saturation signal&lt;/h1>
&lt;p>Your alert says &lt;code>scur&lt;/code> is at 85% of &lt;code>slim&lt;/code> on a frontend, or &lt;code>CurrConns&lt;/code> is closing on &lt;code>Maxconn&lt;/code> in &lt;code>show info&lt;/code>. HAProxy is still up. Traffic is still flowing. But you are running out of concurrent-session headroom, and what happens next depends entirely on which of the four independent limits you are about to hit.&lt;/p>
&lt;p>The danger is that the system still appears responsive: HAProxy is alive and healthy, but clients are about to see queuing delays or 503 rejections. The scur/slim ratio is the warning, if you read it correctly.&lt;/p></description></item><item><title>HAProxy server flapping: rise, fall, and health checks oscillating UP and DOWN</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-server-flapping-up-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-server-flapping-up-down/</guid><description>&lt;h1 id="haproxy-server-flapping-rise-fall-and-health-checks-oscillating-up-and-down">HAProxy server flapping: rise, fall, and health checks oscillating UP and DOWN&lt;/h1>
&lt;p>A backend server that alternates between UP and DOWN is worse than one that is cleanly down. Every transition redistributes traffic: when the server goes DOWN its connections move to the survivors, when it comes back UP it absorbs a fresh share of load before it has warmed up, and every cycle costs retries, redispatches, and connection churn. If the flap rate is high enough, the oscillation itself becomes the incident.&lt;/p></description></item><item><title>HAProxy server-template and dynamic backends: monitoring service-discovery routing</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-server-template-dynamic-backends/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-server-template-dynamic-backends/</guid><description>&lt;h1 id="haproxy-server-template-and-dynamic-backends-monitoring-service-discovery-routing">HAProxy server-template and dynamic backends: monitoring service-discovery routing&lt;/h1>
&lt;p>With a static HAProxy config, the server list lives in the config file. Capacity review is config review: count the &lt;code>server&lt;/code> lines and you know what the backend can absorb. With &lt;code>server-template&lt;/code>, the server list lives in DNS. HAProxy pre-provisions a pool of empty server slots, and its internal resolver fills and empties them as A or SRV records change. This is the mechanism behind HAProxy fronting Kubernetes Ingress, Consul, and most autoscaled service-discovery setups.&lt;/p></description></item><item><title>HAProxy show pools: finding which internal pool is consuming memory</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-show-pools-memory-breakdown/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-show-pools-memory-breakdown/</guid><description>&lt;h1 id="haproxy-show-pools-finding-which-internal-pool-is-consuming-memory">HAProxy show pools: finding which internal pool is consuming memory&lt;/h1>
&lt;p>HAProxy&amp;rsquo;s memory footprint is not one number. The process holds dozens of internal object pools: buffers, connections, sessions, tasks, stick-table entries, SSL state. When RSS climbs, when &lt;code>PoolFailed&lt;/code> goes nonzero, or when you are sizing a memory budget for a new instance, the question is never &amp;ldquo;how much memory is HAProxy using&amp;rdquo; but &amp;ldquo;which pool is using it&amp;rdquo;. &lt;code>show pools&lt;/code> is the runtime API command that answers that.&lt;/p></description></item><item><title>HAProxy Slowloris and idle-connection pile-up: high scur, low request rate</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-slowloris-idle-connections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-slowloris-idle-connections/</guid><description>&lt;h1 id="haproxy-slowloris-and-idle-connection-pile-up-high-scur-low-request-rate">HAProxy Slowloris and idle-connection pile-up: high scur, low request rate&lt;/h1>
&lt;p>Your HAProxy frontend shows &lt;code>scur&lt;/code> pinned at or near &lt;code>slim&lt;/code>, but &lt;code>req_rate&lt;/code> is a fraction of what those connections should be producing. Throughput (&lt;code>bin&lt;/code>/&lt;code>bout&lt;/code>) is near zero, backend queue depth (&lt;code>qcur&lt;/code>) is zero because backends are doing nothing, and &lt;code>ereq&lt;/code> with 408 responses is climbing. HAProxy is not down. It is not slow. It is full: thousands of connections are holding session slots open and sending data so slowly they may as well be dead.&lt;/p></description></item><item><title>HAProxy SSL certificate expired: total TLS failure and how to catch it first</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-ssl-certificate-expired/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-ssl-certificate-expired/</guid><description>&lt;h1 id="haproxy-ssl-certificate-expired-total-tls-failure-and-how-to-catch-it-first">HAProxy SSL certificate expired: total TLS failure and how to catch it first&lt;/h1>
&lt;p>Every client connecting to the affected frontend gets a TLS handshake failure. Not a percentage of clients, not a slow degradation: all of them, immediately, from the moment the certificate&amp;rsquo;s &lt;code>notAfter&lt;/code> timestamp passes. Browsers show &lt;code>NET::ERR_CERT_DATE_INVALID&lt;/code>, API clients throw certificate validation errors, and HAProxy itself is perfectly healthy, running, and passing every process-liveness check you have.&lt;/p>
&lt;p>The mechanism is trivial, the blast radius is total, and HAProxy gives you no built-in metric, counter, or log warning that a certificate is about to expire. If you are reading this during an incident, skip to Quick checks and Fixes. If you are reading it afterward, the Prevention section is the part that matters.&lt;/p></description></item><item><title>HAProxy SSL session cache misses: full handshakes and lost resumption</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-ssl-session-cache-misses/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-ssl-session-cache-misses/</guid><description>&lt;h1 id="haproxy-ssl-session-cache-misses-full-handshakes-and-lost-resumption">HAProxy SSL session cache misses: full handshakes and lost resumption&lt;/h1>
&lt;p>Your HAProxy CPU is climbing but connection counts look normal. &lt;code>Idle_pct&lt;/code> is dropping, latency on new connections is rising, and existing sessions seem fine. When you pull &lt;code>show info&lt;/code>, &lt;code>SslFrontendKeyRate&lt;/code> is far above baseline and the SSL session cache numbers tell the story: almost every lookup is a miss.&lt;/p>
&lt;p>A high cache miss rate means most TLS clients are paying for a full handshake instead of resuming a previous session. Full handshakes are the most CPU-intensive operation HAProxy performs; resumed handshakes skip most of that cost. When resumption breaks, the load balancer loses a large fraction of its capacity without any change in traffic volume.&lt;/p></description></item><item><title>HAProxy SslFrontendKeyRate: the handshake rate that drives CPU cost</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-ssl-frontend-key-rate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-ssl-frontend-key-rate/</guid><description>&lt;h1 id="haproxy-sslfrontendkeyrate-the-handshake-rate-that-drives-cpu-cost">HAProxy SslFrontendKeyRate: the handshake rate that drives CPU cost&lt;/h1>
&lt;p>SslFrontendKeyRate is the number of new (non-resumed) TLS handshakes per second across your HAProxy frontends. Because full TLS handshakes are the most CPU-intensive operation HAProxy performs, this metric is effectively the CPU cost of your TLS traffic expressed as a rate.&lt;/p>
&lt;p>Operators usually look for this metric in two situations: HAProxy CPU is climbing and they need to know whether TLS is the cause, or they are doing capacity planning and need a defensible number for &amp;ldquo;how many handshakes per second can this box absorb.&amp;rdquo; Both depend on understanding what the metric counts, where it comes from, and what breaks it.&lt;/p></description></item><item><title>HAProxy stick-table overflow: rate limiting that silently stops working</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-stick-table-overflow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-stick-table-overflow/</guid><description>&lt;h1 id="haproxy-stick-table-overflow-rate-limiting-that-silently-stops-working">HAProxy stick-table overflow: rate limiting that silently stops working&lt;/h1>
&lt;p>Your HAProxy config has rate limiting. It has session persistence. Both are enforced by stick tables, and both can stop working without an error counter, a log line, or a page.&lt;/p>
&lt;p>Stick tables are fixed-size, in-memory key-value stores. When a table reaches its configured &lt;code>size&lt;/code>, HAProxy has two possible behaviors and neither one alerts you. With the default configuration, HAProxy flushes expired entries to make room for new ones. If the table is full of entries that have not yet expired, new entries are refused. With the &lt;code>nopurge&lt;/code> option, eviction is disabled entirely and new entries are silently rejected. In both paths, whatever the table was doing (rate limiting, sticky sessions, abuse tracking) degrades or stops for affected clients.&lt;/p></description></item><item><title>HAProxy stuck in soft-stop: draining connections and hard-stop-after</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-soft-stop-draining-stuck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-soft-stop-draining-stuck/</guid><description>&lt;h1 id="haproxy-stuck-in-soft-stop-draining-connections-and-hard-stop-after">HAProxy stuck in soft-stop: draining connections and hard-stop-after&lt;/h1>
&lt;p>You reloaded HAProxy minutes ago, but &lt;code>pgrep haproxy&lt;/code> still shows two, three, maybe a dozen PIDs. The old process is not serving new traffic, but it refuses to die. Memory and file descriptor usage on the host keep climbing with every reload, and your monitoring loses counter history each time it happens.&lt;/p>
&lt;p>The old process is in soft-stop: it has unbound from its listeners and is waiting for its existing connections to close before exiting. A drain of a few seconds to a couple of minutes is normal. A process draining for more than 10 minutes is stuck, and the cause is almost always long-lived connections (WebSocket, Server-Sent Events, gRPC streaming, or plain TCP keepalive sessions) that have no reason to close on their own.&lt;/p></description></item><item><title>HAProxy TLS handshake failures: broken clients after a cert or cipher change</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-tls-handshake-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-tls-handshake-failures/</guid><description>&lt;h1 id="haproxy-tls-handshake-failures-broken-clients-after-a-cert-or-cipher-change">HAProxy TLS handshake failures: broken clients after a cert or cipher change&lt;/h1>
&lt;p>A certificate rotation or TLS policy change goes out, HAProxy reloads cleanly, health checks pass, traffic graphs look normal, and then the tickets start: a subset of clients can no longer connect. The load balancer itself is fine. The handshake rate even looks normal. What changed is the failure rate &lt;em>within&lt;/em> handshakes, and HAProxy has no dedicated stats counter that shows it to you directly.&lt;/p></description></item><item><title>HAProxy TLS handshake storm: new-connection floods that saturate the event loop</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-tls-handshake-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-tls-handshake-storm/</guid><description>&lt;h1 id="haproxy-tls-handshake-storm-new-connection-floods-that-saturate-the-event-loop">HAProxy TLS handshake storm: new-connection floods that saturate the event loop&lt;/h1>
&lt;p>Traffic looks normal at the TCP layer. Connections are arriving, the frontend is accepting them, but latency is climbing, some clients time out before ever sending a request, and backend servers look healthy. When you check &lt;code>show info&lt;/code>, &lt;code>Idle_pct&lt;/code> is near zero and &lt;code>SslFrontendKeyRate&lt;/code> is far above anything you have tested. You are in a TLS handshake storm: HAProxy is burning every CPU cycle on cryptographic handshake work and the event loop has nothing left for the rest of the pipeline.&lt;/p></description></item><item><title>HAProxy Too many open files: file descriptor exhaustion and refused connections</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-too-many-open-files/</guid><description>&lt;h1 id="haproxy-too-many-open-files-file-descriptor-exhaustion-and-refused-connections">HAProxy Too many open files: file descriptor exhaustion and refused connections&lt;/h1>
&lt;p>HAProxy is refusing new connections. Clients see connection timeouts or resets, health checks start failing for no application reason, and the logs show &lt;code>Too many open files&lt;/code>. The process is alive and CPU looks fine, which makes this confusing the first time you hit it: nothing is overloaded in the usual sense. The proxy has run out of file descriptors.&lt;/p></description></item><item><title>HAProxy tune.bufsize and 400 errors: buffers too small for large headers</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-tune-bufsize-large-headers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-tune-bufsize-large-headers/</guid><description>&lt;h1 id="haproxy-tunebufsize-and-400-errors-buffers-too-small-for-large-headers">HAProxy tune.bufsize and 400 errors: buffers too small for large headers&lt;/h1>
&lt;p>Your frontend 400 rate just spiked. The backends insist they never saw the requests, and they are telling the truth: HAProxy generated these 400s itself, during header parsing, before any backend was selected. The usual trigger is a request header block that no longer fits in the per-connection buffer, and the usual suspects are large cookies, JWT bearer tokens, or a newly added header that pushed total header size past the limit.&lt;/p></description></item><item><title>Hardware information collected from kernel ring.</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/hardware-information-collected-from-kernel-ring./</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/hardware-information-collected-from-kernel-ring./</guid><description/></item><item><title>Harmonic Lightwaves SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/harmonic-lightwaves-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/harmonic-lightwaves-snmp-traps/</guid><description/></item><item><title>HDD temperature</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/hdd-temperature/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/hdd-temperature/</guid><description/></item><item><title>HDD Temperature Monitoring</title><link>https://www.netdata.cloud/monitoring-101/hddtemp-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/hddtemp-monitoring/</guid><description>&lt;h2 id="hdd-temperature-monitoring">HDD Temperature Monitoring&lt;/h2>
&lt;h3 id="what-is-hdd-temperature">What Is HDD Temperature?&lt;/h3>
&lt;p>HDD temperature refers to the heat emitted by your hard disk drive during its operation. Monitoring the HDD temperature can prevent overheating, which could compromise data integrity and hardware stability. Using an efficient HDD temperature monitoring tool is crucial for maintaining your system’s overall health.&lt;/p>
&lt;h3 id="monitoring-hdd-temperature-with-netdata">Monitoring HDD Temperature With Netdata&lt;/h3>
&lt;p>Netdata provides a comprehensive HDD temperature monitoring solution. This open-source platform leverages the hddtemp daemon to collect and display real-time temperature data across all your disks, ensuring you stay informed about any potential overheating issues. With &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata&amp;rsquo;s live demo&lt;/a>, you can see Netdata&amp;rsquo;s monitoring capabilities in action.&lt;/p></description></item><item><title>Hewlett Packard Enterprise SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/hewlett-packard-enterprise-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/hewlett-packard-enterprise-snmp-traps/</guid><description/></item><item><title>Hewlett Packard SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/hewlett-packard-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/hewlett-packard-snmp-traps/</guid><description/></item><item><title>Hitachi Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/hitachi-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/hitachi-ltd-snmp-traps/</guid><description/></item><item><title>Hitron CODA Cable Modem</title><link>https://www.netdata.cloud/integrations/data-collection/networking/hitron-coda-cable-modem/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/hitron-coda-cable-modem/</guid><description/></item><item><title>Hitron CODA Cable Modem Monitoring</title><link>https://www.netdata.cloud/monitoring-101/hitron_coda-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/hitron_coda-monitoring/</guid><description>&lt;h2 id="hitron-coda-cable-modem-monitoring">Hitron CODA Cable Modem Monitoring&lt;/h2>
&lt;h3 id="what-is-hitron-coda-cable-modem">What Is Hitron CODA Cable Modem?&lt;/h3>
&lt;p>The Hitron CODA Cable Modem is a device that facilitates high-speed internet connectivity, often used in homes and businesses. This modem is designed to enhance internet performance through efficient data management and superior networking capabilities. Monitoring is crucial for maintaining optimal performance and troubleshooting any connectivity issues that may arise.&lt;/p>
&lt;h3 id="monitoring-hitron-coda-cable-modem-with-netdata">Monitoring Hitron CODA Cable Modem With Netdata&lt;/h3>
&lt;p>To monitor the Hitron CODA Cable Modem, Netdata utilizes an openmetrics (Prometheus) exporter &lt;a href="https://github.com/hairyhenderson/hitron_coda_exporter">developed by the community&lt;/a>. This integration allows Netdata to ingest data from any Prometheus exporter, offering users automated dashboards, alerts, and more—without the need for setting up a complex Prometheus server or Grafana. The seamless integration with Netdata means you get actionable insights with minimal setup.&lt;/p></description></item><item><title>Homebridge</title><link>https://www.netdata.cloud/integrations/data-collection/applications/homebridge/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/homebridge/</guid><description/></item><item><title>Homebridge Monitoring</title><link>https://www.netdata.cloud/monitoring-101/homebridge-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/homebridge-monitoring/</guid><description>&lt;h2 id="homebridge-monitoring">Homebridge Monitoring&lt;/h2>
&lt;h3 id="what-is-homebridge">What Is Homebridge?&lt;/h3>
&lt;p>Homebridge is a lightweight Node.js server that emulates the iOS HomeKit API. It allows you to connect all your smart home devices, which may not natively support HomeKit, to the Apple ecosystem. By integrating a wide array of home automation devices from various brands into a single cohesive environment, Homebridge plays a pivotal role in efficient home automation management.&lt;/p>
&lt;h3 id="monitoring-homebridge-with-netdata">Monitoring Homebridge With Netdata&lt;/h3>
&lt;p>To monitor Homebridge effectively, utilizing the &lt;a href="https://github.com/lstrojny/homebridge-prometheus-exporter">Homebridge Prometheus Exporter&lt;/a>, Netdata provides a comprehensive solution. Netdata uses an openmetrics (Prometheus) exporter approach, capable of ingesting data from any Prometheus exporter, offering you automated dashboards, alerts, and more—all without the need for a separate Prometheus server or Grafana. This streamlined, powerful integration makes Homebridge monitoring seamless and highly efficient.&lt;/p></description></item><item><title>Homey</title><link>https://www.netdata.cloud/integrations/data-collection/applications/homey/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/homey/</guid><description/></item><item><title>Homey Monitoring</title><link>https://www.netdata.cloud/monitoring-101/homey-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/homey-monitoring/</guid><description>&lt;h2 id="homey-monitoring">Homey Monitoring&lt;/h2>
&lt;h3 id="what-is-homey">What Is Homey?&lt;/h3>
&lt;p>Homey is a smart home controller that allows users to manage and automate their home appliances and devices, offering a seamless and integrated smart living experience. Its user-friendly interface and compatibility with a wide range of smart devices make it a popular choice for home automation enthusiasts.&lt;/p>
&lt;h3 id="monitoring-homey-with-netdata">Monitoring Homey With Netdata&lt;/h3>
&lt;p>Netdata provides a comprehensive and efficient way to monitor Homey using the Homey Prometheus Exporter. By leveraging an openmetrics (Prometheus) exporter, Netdata enables the ingestion of data from any Prometheus exporter. This means you can monitor Homey efficiently without the need for a Prometheus server or Grafana. Netdata offers automated dashboards, alerts, and detailed real-time monitoring, making it an essential tool for maintaining the optimal performance of your smart home setup.&lt;/p></description></item><item><title>Honeypot</title><link>https://www.netdata.cloud/integrations/data-collection/applications/honeypot/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/honeypot/</guid><description/></item><item><title>Honeypot Monitoring</title><link>https://www.netdata.cloud/monitoring-101/honeypot-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/honeypot-monitoring/</guid><description>&lt;h2 id="honeypot-monitoring">Honeypot Monitoring&lt;/h2>
&lt;h3 id="what-is-honeypot">What Is Honeypot?&lt;/h3>
&lt;p>A honeypot is a security mechanism that creates a virtual trap to entice cyber attackers. It essentially offers an isolated and controlled environment where malicious activities can be detected and studied without compromising real systems. Used prominently in cybersecurity, a honeypot mimics a legitimate system to lure hackers, allowing organizations to analyze attack techniques and enhance system defenses.&lt;/p>
&lt;h3 id="monitoring-honeypot-with-netdata">Monitoring Honeypot With Netdata&lt;/h3>
&lt;p>Monitoring Honeypot with Netdata leverages the &lt;a href="https://github.com/Intrinsec/honeypot_exporter">Intrinsec honeypot_exporter&lt;/a> using an openmetrics (Prometheus) exporter. Netdata excels by ingesting data from any Prometheus exporter, providing users automated dashboards, alerts, and insights without the need for a Prometheus server or Grafana. This seamless integration enables real-time visibility into Honeypot metrics, facilitating efficient threat detection and management.&lt;/p></description></item><item><title>How ActiveMQ Classic actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/activemq/activemq-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/activemq/activemq-how-it-works-in-production/</guid><description>&lt;h1 id="how-activemq-classic-actually-works-in-production-a-mental-model-for-operators">How ActiveMQ Classic actually works in production: a mental model for operators&lt;/h1>
&lt;p>Most ActiveMQ incidents are misdiagnosed in the first hour because the operator is looking at the wrong subsystem. A &amp;ldquo;hung producer&amp;rdquo; is flow control doing its job. A &amp;ldquo;slow broker&amp;rdquo; is fsync latency on the KahaDB device. A &amp;ldquo;healthy&amp;rdquo; broker that is losing messages is a full temp store with flow control disabled. Each of these looks like a different problem until you know which subsystem is actually involved.&lt;/p></description></item><item><title>How an NVIDIA GPU actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-how-gpus-work-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-how-gpus-work-in-production/</guid><description>&lt;h1 id="how-an-nvidia-gpu-actually-works-in-production-a-mental-model-for-operators">How an NVIDIA GPU actually works in production: a mental model for operators&lt;/h1>
&lt;p>GPU runbooks fail when they treat the GPU as a black box that emits a utilization percentage. &amp;ldquo;100% utilization but slow&amp;rdquo; makes sense only when you know what utilization measures. &amp;ldquo;nvidia-smi hangs&amp;rdquo; makes sense only when you know where nvidia-smi gets its data.&lt;/p>
&lt;p>This guide covers the two pieces underneath GPU troubleshooting: the subsystems that produce production failures, and the telemetry plane through which you observe them. It applies primarily to datacenter GPUs such as V100, A100, and H100, with differences noted for consumer cards.&lt;/p></description></item><item><title>How Apache HTTPD actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-httpd/apache-httpd-how-it-works-in-production/</guid><description>&lt;h1 id="how-apache-httpd-actually-works-in-production-a-mental-model-for-operators">How Apache HTTPD actually works in production: a mental model for operators&lt;/h1>
&lt;p>Most Apache incidents are misdiagnosed for the same reason: the operator is reading signals without knowing which execution model produced them. A scoreboard full of &lt;code>K&lt;/code> states is a crisis on prefork and a rounding error on event. &amp;ldquo;Apache is down&amp;rdquo; and &amp;ldquo;the backend is down&amp;rdquo; look identical from outside a reverse proxy. MaxRequestWorkers means something different depending on whether your concurrency unit is a process or a thread.&lt;/p></description></item><item><title>How Apache Pulsar actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/apache-pulsar/apache-pulsar-how-it-works-in-production/</guid><description>&lt;h1 id="how-apache-pulsar-actually-works-in-production-a-mental-model-for-operators">How Apache Pulsar actually works in production: a mental model for operators&lt;/h1>
&lt;p>Most Pulsar incidents are misdiagnosed in the first thirty minutes because the operator is looking at the wrong layer. A publish latency spike gets chased on the broker when the real bottleneck is a bookie journal disk. A &amp;ldquo;consumer problem&amp;rdquo; turns out to be a cursor that stopped advancing. A broker that looks healthy on its HTTP endpoint is fenced off from ZooKeeper and losing topic ownership in a loop.&lt;/p></description></item><item><title>How BIND actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-how-it-works-in-production/</guid><description>&lt;h1 id="how-bind-actually-works-in-production-a-mental-model-for-operators">How BIND actually works in production: a mental model for operators&lt;/h1>
&lt;p>During an incident, &lt;code>named&lt;/code> can feel like a black box. A single process handles packet reception, DNS parsing, ACL evaluation, cache lookup, zone lookup, recursive fetch, DNSSEC validation, RPZ policy enforcement, and response serialization. There is no separate resolver process, no separate authoritative process, no separate cache daemon. Every query flows through the same pipeline, and every subsystem competes for the same pool of CPU, memory, file descriptors, and kernel network buffers.&lt;/p></description></item><item><title>How Cassandra actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/cassandra/how-cassandra-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cassandra/how-cassandra-works-in-production/</guid><description>&lt;h1 id="how-cassandra-actually-works-in-production-a-mental-model-for-operators">How Cassandra actually works in production: a mental model for operators&lt;/h1>
&lt;p>If you operate Cassandra in production, you are managing a distributed, partitioned, replicated log-structured merge tree where every node is a peer. There is no master to restart, no central query planner to tune, and no automatic load balancing that will save you from a hot partition. Every write is a sequential append to a commitlog and an in-memory update on multiple replicas. Every read is a merge of memtables and immutable SSTables, filtered by probabilistic bloom filters and reconciled by timestamp.&lt;/p></description></item><item><title>How Ceph actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/ceph/ceph-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/ceph/ceph-how-it-works-in-production/</guid><description>&lt;h1 id="how-ceph-actually-works-in-production-a-mental-model-for-operators">How Ceph actually works in production: a mental model for operators&lt;/h1>
&lt;p>Most Ceph incidents become legible the moment you stop reasoning about RBD, CephFS, and RGW as separate products and start reasoning about one system: RADOS placing objects on OSDs. The three client interfaces are thin translations. Underneath them, every write is an object, every object lives in a placement group, and every placement group is mapped to a set of OSDs by a deterministic algorithm.&lt;/p></description></item><item><title>How ClickHouse actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/clickhouse/how-clickhouse-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/clickhouse/how-clickhouse-works-in-production/</guid><description>&lt;h1 id="how-clickhouse-actually-works-in-production-a-mental-model-for-operators">How ClickHouse actually works in production: a mental model for operators&lt;/h1>
&lt;p>ClickHouse is a column-oriented OLAP database built around the MergeTree family of table engines. It does not fail like a transactional database. Its failure modes are distinctive and almost always originate in the storage layer rather than the query layer. In production, most incidents begin as storage-structure debt: immutable data parts accumulate faster than background processes can consolidate them. By the time queries slow down or inserts fail, the underlying merge crisis has been developing for hours.&lt;/p></description></item><item><title>How CockroachDB actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/cockroachdb/how-cockroachdb-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/cockroachdb/how-cockroachdb-works-in-production/</guid><description>&lt;h1 id="how-cockroachdb-actually-works-in-production-a-mental-model-for-operators">How CockroachDB actually works in production: a mental model for operators&lt;/h1>
&lt;p>CockroachDB is a distributed, strongly-consistent SQL database built on a replicated key-value store. It layers SQL execution on top of a transactional KV engine that uses Raft consensus for replication and MVCC for concurrency control. To reason about its failures, you must hold several interacting subsystems in your head simultaneously.&lt;/p>
&lt;p>This is not a tutorial. It is the mental model that experienced operators use when diagnosing slow queries, unavailable ranges, clock skew, and compaction death spirals. Each subsystem has its own failure modes, but the most damaging incidents happen when failures cascade across layers.&lt;/p></description></item><item><title>How Consul actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/consul/consul-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/consul/consul-how-it-works-in-production/</guid><description>&lt;h1 id="how-consul-actually-works-in-production-a-mental-model-for-operators">How Consul actually works in production: a mental model for operators&lt;/h1>
&lt;p>Most Consul incidents are debugged with the wrong mental model. Operators check &lt;code>consul members&lt;/code>, see all nodes alive, conclude the cluster is healthy, and miss that Raft has no leader or that the catalog is stale by minutes. Consul runs three concurrent subsystems with different failure modes, signals, and consistency guarantees. Most monitoring collapses them into one &amp;ldquo;is Consul up?&amp;rdquo; check.&lt;/p></description></item><item><title>How CoreDNS actually works in production: the plugin chain mental model</title><link>https://www.netdata.cloud/guides/coredns/coredns-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/coredns/coredns-how-it-works-in-production/</guid><description>&lt;h1 id="how-coredns-actually-works-in-production-the-plugin-chain-mental-model">How CoreDNS actually works in production: the plugin chain mental model&lt;/h1>
&lt;p>Most CoreDNS incidents are pipeline incidents, not DNS incidents. An upstream resolver gets slow, goroutines pile up, memory climbs, and the pod gets OOMKilled. Or the Kubernetes API watch drops silently and CoreDNS keeps answering every query correctly except the answers are three hours stale. Each failure mode is a direct consequence of how CoreDNS is built internally, and most CoreDNS runbooks assume you already hold that internal model in your head.&lt;/p></description></item><item><title>How Elasticsearch actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/elasticsearch/how-elasticsearch-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/elasticsearch/how-elasticsearch-works-in-production/</guid><description>&lt;h1 id="how-elasticsearch-actually-works-in-production-a-mental-model-for-operators">How Elasticsearch actually works in production: a mental model for operators&lt;/h1>
&lt;p>Elasticsearch is a distributed search and analytics engine built on Apache Lucene. Every node runs interacting subsystems that compete for the same resources. If you treat it as a black box that stores JSON, you will miss the failure modes that kill clusters at 3 a.m. The interactions between subsystems determine whether a cluster survives a traffic spike or collapses into a heap-pressure death spiral.&lt;/p></description></item><item><title>How Envoy actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/envoy/envoy-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/envoy/envoy-how-it-works-in-production/</guid><description>&lt;h1 id="how-envoy-actually-works-in-production-a-mental-model-for-operators">How Envoy actually works in production: a mental model for operators&lt;/h1>
&lt;p>Envoy is a multi-threaded, event-driven L4/L7 proxy written in C++. Its architecture directly shapes what you see in metrics, access logs, and user-visible behavior during incidents. If you do not know that worker threads share nothing in the hot path, aggregate CPU utilization will mislead you. If you do not know that circuit breaker 503s are indistinguishable from upstream-generated 503s at the counter level, you will blame the wrong component. If you do not know that Envoy keeps serving traffic on stale configuration after an xDS disconnect, you will miss the slow-burn failure that surfaces hours later.&lt;/p></description></item><item><title>How Fluentd actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/fluentd/fluentd-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/fluentd/fluentd-how-it-works-in-production/</guid><description>&lt;h1 id="how-fluentd-actually-works-in-production-a-mental-model-for-operators">How Fluentd actually works in production: a mental model for operators&lt;/h1>
&lt;p>Most Fluentd incidents are not random. They are the predictable output of a small set of internal mechanisms: a single-threaded event router, a chunked buffer with a hard limit, Ruby threads fighting over one GVL, and a retry engine with exponential backoff. If you understand those four, you can predict almost every Fluentd failure mode before you open a dashboard.&lt;/p></description></item><item><title>How HAProxy actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/haproxy/haproxy-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/haproxy/haproxy-how-it-works-in-production/</guid><description>&lt;h1 id="how-haproxy-actually-works-in-production-a-mental-model-for-operators">How HAProxy actually works in production: a mental model for operators&lt;/h1>
&lt;p>Most HAProxy incidents come from an operator applying the wrong mental model: treating HAProxy like a web server, like nginx, or like a black box that either works or does not. HAProxy is an event-driven state machine that multiplexes hundreds of thousands of connections across a small number of threads, and almost every confusing symptom traces back to one of a handful of internal mechanisms: the accept pipeline, the maxconn hierarchy, the buffer and connection pools, the health check engine, or a subsystem you forgot was running.&lt;/p></description></item><item><title>How Kafka actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/kafka/how-kafka-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/how-kafka-works-in-production/</guid><description>&lt;h1 id="how-kafka-actually-works-in-production-a-mental-model-for-operators">How Kafka actually works in production: a mental model for operators&lt;/h1>
&lt;h2 id="what-it-is-and-why-it-matters">What it is and why it matters&lt;/h2>
&lt;p>Each broker is a concurrent system of subsystems sharing disk, network, page cache, and threads. Tail latency spikes and offline partitions usually stem from interactions between these subsystems, not single component failures.&lt;/p>
&lt;p>Underneath the distributed commit log abstraction, each broker runs a reactor-pattern server with three saturation layers: network threads, a bounded request queue, and I/O handler threads. Between I/O threads and the log sits a delayed-operation timer wheel called the purgatory. A separate controller maintains cluster metadata in a single-threaded event queue. Consumer groups are managed by a Group Coordinator that hashes group IDs to the &lt;code>__consumer_offsets&lt;/code> topic. Log data lives in segment files that rely on OS page cache and zero-copy sendfile. These subsystems compete for the same disk I/O, network bandwidth, file descriptors, and cache. Treating them as independent pipelines leads to unnecessary restarts and missed bottlenecks.&lt;/p></description></item><item><title>How Logstash actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/logstash/logstash-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-how-it-works-in-production/</guid><description>&lt;h1 id="how-logstash-actually-works-in-production-a-mental-model-for-operators">How Logstash actually works in production: a mental model for operators&lt;/h1>
&lt;p>Most Logstash incidents are diagnosed badly because the operator is reasoning about the wrong system. They see a &amp;ldquo;log shipper&amp;rdquo; and check whether the process is up. The process is up. The API returns 200. Nothing is being delivered. Or they see high CPU and assume the host is undersized, when the real problem is one grok pattern with catastrophic backtracking. Or they watch queue depth stay flat for hours and conclude all is well, while the persistent queue quietly absorbs a downstream outage that will page someone at 4 a.m. when it fills.&lt;/p></description></item><item><title>How LVM actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/lvm/lvm-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-how-it-works-in-production/</guid><description>&lt;h1 id="how-lvm-actually-works-in-production-a-mental-model-for-operators">How LVM actually works in production: a mental model for operators&lt;/h1>
&lt;p>Most LVM incidents are caused by operators holding the wrong mental model: thinking of LVM as &amp;ldquo;partitioning with extra steps&amp;rdquo; rather than what it actually is, a userspace management layer that programs the kernel&amp;rsquo;s device-mapper subsystem. Once you internalize that, the confusing behaviors stop being confusing. A full thin pool freezing every process in D-state, a snapshot silently invalidating, an &lt;code>lvs&lt;/code> command hanging during the exact incident you need it for: all of these follow directly from the architecture.&lt;/p></description></item><item><title>How Memcached actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/memcached/memcached-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-how-it-works-in-production/</guid><description>&lt;h1 id="how-memcached-actually-works-in-production-a-mental-model-for-operators">How Memcached actually works in production: a mental model for operators&lt;/h1>
&lt;p>Memcached is a multi-threaded, in-memory key-value cache daemon built on libevent. No persistence, no replication, no clustering. Clients handle sharding. A restart means total data loss. These are the design, not limitations to work around.&lt;/p>
&lt;p>Before you can debug a memcached incident, you need three abstractions: how memory is partitioned (the slab allocator), how items age within each partition (the segmented LRU), and how connections and threads interact under load. Without these, the stats output is noise. With them, the same numbers tell you exactly which resource is saturated.&lt;/p></description></item><item><title>How Microsoft SQL Server actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-how-it-works-in-production/</guid><description>&lt;h1 id="how-microsoft-sql-server-actually-works-in-production-a-mental-model-for-operators">How Microsoft SQL Server actually works in production: a mental model for operators&lt;/h1>
&lt;p>Most SQL Server incidents look confusing because operators bring a Linux-process mental model to a system that is not one. &lt;code>sqlservr&lt;/code> (&lt;code>sqlservr.exe&lt;/code> on Windows) is a single multi-threaded process, but inside it runs SQLOS, a user-mode operating system with its own scheduler, memory manager, and I/O completion handling. The host OS does not schedule your queries, does not cache your data pages, and does not decide which transaction gets a lock. SQL Server does all of that itself.&lt;/p></description></item><item><title>How MongoDB actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/mongodb/how-mongodb-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/how-mongodb-works-in-production/</guid><description>&lt;h1 id="how-mongodb-actually-works-in-production-a-mental-model-for-operators">How MongoDB actually works in production: a mental model for operators&lt;/h1>
&lt;p>A write latency spike in MongoDB is rarely a query problem. It is more often a storage journal stall, cache eviction cascade, ticket exhaustion, or replication flow-control throttle. Operators who treat MongoDB as a monolithic document database often tune indexes while the real bottleneck is a WiredTiger checkpoint falling behind.&lt;/p>
&lt;p>This guide is a runtime mental model for the subsystems that matter in production: the WiredTiger storage engine, replication state machine, thread-per-connection execution model, and sharding coordination layer. Understanding how these compete for memory, disk, tickets, and connections lets you read signals correctly instead of guessing.&lt;/p></description></item><item><title>How MySQL actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/mysql/how-mysql-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/how-mysql-works-in-production/</guid><description>&lt;h1 id="how-mysql-actually-works-in-production-a-mental-model-for-operators">How MySQL actually works in production: a mental model for operators&lt;/h1>
&lt;p>MySQL is a connection-oriented request processor with a pluggable storage engine. In production, that engine is almost always InnoDB. Operating MySQL at scale means managing memory pressure, write-ahead log capacity, version-chain traversal, and fsync latency. The optimizer matters, but the stalls that wake you up start inside InnoDB.&lt;/p>
&lt;p>This guide maps the subsystems that fail in production. It assumes you know how to run &lt;code>SHOW PROCESSLIST&lt;/code> and need to move from symptoms to root cause: why the server stalls, why replication drifts, or why a fast query suddenly saturates disk I/O.&lt;/p></description></item><item><title>How NATS actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/nats/nats-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-how-it-works-in-production/</guid><description>&lt;h1 id="how-nats-actually-works-in-production-a-mental-model-for-operators">How NATS actually works in production: a mental model for operators&lt;/h1>
&lt;p>Most NATS incidents come from a mismatch between how operators think NATS works and how it actually works. Teams coming from Kafka or RabbitMQ carry a queue-centric mental model: messages go somewhere, wait there, and can be retrieved later. Core NATS does not work like that. The mismatch shows up at 3 a.m. as silent message loss, slow consumer cascades, or a Raft election storm nobody saw coming.&lt;/p></description></item><item><title>How NGINX actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/nginx/how-nginx-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/how-nginx-works-in-production/</guid><description>&lt;h1 id="how-nginx-actually-works-in-production-a-mental-model-for-operators">How NGINX actually works in production: a mental model for operators&lt;/h1>
&lt;p>NGINX is not a multi-threaded server that spawns a thread per connection. It is an event-driven, non-blocking, single-threaded-per-worker process architecture. If you are debugging a production incident where connections are timing out, memory is climbing, or CPU is pinned, this architecture is the lens through which every symptom must be interpreted.&lt;/p>
&lt;p>Most production issues involving NGINX are not NGINX bugs. They are resource accounting problems: file descriptors, connection slots, buffer boundaries, or event loop latency. An operator who understands the internal mechanics can read the signals correctly instead of chasing phantom upstream problems or adding hardware that hits the same limit.&lt;/p></description></item><item><title>How NVMe actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/nvme/nvme-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-how-it-works-in-production/</guid><description>&lt;h1 id="how-nvme-actually-works-in-production-a-mental-model-for-operators">How NVMe actually works in production: a mental model for operators&lt;/h1>
&lt;p>Most storage incidents on NVMe are misdiagnosed for the same reason: operators debug them as if NVMe were a faster SATA disk. It is not. NVMe is a host-to-controller communication protocol that exposes flash storage over PCIe, and the failure modes live in places generic disk monitoring never looks: the PCIe link, the controller firmware, and the flash translation layer sitting between your filesystem and the NAND.&lt;/p></description></item><item><title>How Oracle Database actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/oracle-database/how-oracle-database-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/how-oracle-database-works-in-production/</guid><description>&lt;h1 id="how-oracle-database-actually-works-in-production-a-mental-model-for-operators">How Oracle Database actually works in production: a mental model for operators&lt;/h1>
&lt;p>Oracle is dense, but four abstractions make most production incidents legible: the shared memory region (SGA), per-process private memory (PGA), the background processes that move data to disk, and the wait-event model that tells you where time is going. When you see &lt;code>log file sync&lt;/code> climbing, or &lt;code>free buffer waits&lt;/code> appearing, or the database hanging while basic health checks still pass, you should know immediately which subsystem is involved and what it competes for.&lt;/p></description></item><item><title>How PgBouncer actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-how-it-works-in-production/</guid><description>&lt;h1 id="how-pgbouncer-actually-works-in-production-a-mental-model-for-operators">How PgBouncer actually works in production: a mental model for operators&lt;/h1>
&lt;p>Most PgBouncer incidents are the predictable consequence of a handful of internal mechanisms interacting under load: a single-threaded event loop, per-pool FIFO wait queues, fixed-size socket buffers, a small set of server connection states, and a pool mode that decides when server connections change hands. Hold these in your head and the alert thresholds stop being arbitrary numbers.&lt;/p>
&lt;p>This article is the model, not the triage. It explains what PgBouncer is doing internally so that when &lt;code>cl_waiting&lt;/code> spikes or &lt;code>sv_login&lt;/code> won&amp;rsquo;t drain, you already know which part of the machine is hurting and why.&lt;/p></description></item><item><title>How PHP-FPM actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-how-it-works-in-production/</guid><description>&lt;h1 id="how-php-fpm-actually-works-in-production-a-mental-model-for-operators">How PHP-FPM actually works in production: a mental model for operators&lt;/h1>
&lt;p>PHP-FPM looks like a single service from the outside: a socket, a process, a stream of responses. Inside, it is a process-based concurrency system with a small number of moving parts, each of which fails in a specific, predictable way. Most production incidents trace back to a misunderstanding of one of those parts.&lt;/p>
&lt;p>The single most important fact about PHP-FPM: each worker process handles exactly one request at a time. There is no in-process concurrency. Your maximum concurrent request capacity equals your number of active worker processes. Everything else in the system exists to feed, queue, recycle, or protect those workers.&lt;/p></description></item><item><title>How Postfix actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/postfix/postfix-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-how-it-works-in-production/</guid><description>&lt;h1 id="how-postfix-actually-works-in-production-a-mental-model-for-operators">How Postfix actually works in production: a mental model for operators&lt;/h1>
&lt;p>Postfix is not a monolith. It is a collection of small, specialized programs that a supervisor process spawns on demand, coordinated by a single-threaded queue manager, with all durable state living as files on disk. Every operational failure in Postfix, from queue gridlock to inode exhaustion to silent DNS degradation, traces back to how these pieces interact.&lt;/p>
&lt;p>The specific symptom matters less than the architecture underneath it. The same model that explains why one slow destination stalls all outbound mail also explains why a burst of inbound traffic degrades delivery performance, and why a restart on a large queue looks like recovery before it looks like failure again.&lt;/p></description></item><item><title>How PostgreSQL actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/postgres/how-postgres-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/how-postgres-works-in-production/</guid><description>&lt;h1 id="how-postgresql-actually-works-in-production-a-mental-model-for-operators">How PostgreSQL actually works in production: a mental model for operators&lt;/h1>
&lt;p>PostgreSQL is a reliable relational database, but that reliability is not magic. It is the emergent behavior of five concrete subsystems that interact in specific, observable ways: the Write-Ahead Log, Multi-Version Concurrency Control, checkpoints, autovacuum, and the process-per-connection model. If you do not understand how these interact, you will misdiagnose table bloat as a missing index, interpret checkpoint I/O spikes as disk failures, and respond to connection exhaustion by raising &lt;code>max_connections&lt;/code> until the OOM killer arrives.&lt;/p></description></item><item><title>How ProxySQL actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-how-it-works-in-production/</guid><description>&lt;h1 id="how-proxysql-actually-works-in-production-a-mental-model-for-operators">How ProxySQL actually works in production: a mental model for operators&lt;/h1>
&lt;p>When ProxySQL fails in production, the layer where it fails determines what you see: connection exhaustion at the frontend pool, mis-routing in the query processor, pool starvation from multiplexing collapse, or false health decisions in the monitor module. Understanding these layers is prerequisite to debugging any of them.&lt;/p>
&lt;p>This article covers the request path from client connection to backend response, the components that make routing and pooling decisions, and the failure modes characteristic of each. The ProxySQL runbooks in this section build on the terminology and component relationships described here.&lt;/p></description></item><item><title>How Redis actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/redis/how-redis-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/how-redis-works-in-production/</guid><description>&lt;h1 id="how-redis-actually-works-in-production-a-mental-model-for-operators">How Redis actually works in production: a mental model for operators&lt;/h1>
&lt;p>Production incidents do not come from the Redis API. They come from invisible internals: a fork that doubles memory usage, a replication backlog that wraps around, a single slow command that freezes every client. You cannot debug a cascading failure at 3 a.m. without knowing which abstractions compete for which resources.&lt;/p>
&lt;h2 id="what-it-is-and-why-it-matters">What it is and why it matters&lt;/h2>
&lt;p>Redis is a single-threaded event loop around an in-memory dataset. Every incident traces back to resource competition inside one process: memory consumed by the dataset, client buffers, replication backlogs, and allocator fragmentation; CPU consumed by command execution, active expiry, and defragmentation; disk I/O consumed by AOF fsync and RDB snapshots; and network bandwidth consumed by replication and Pub/Sub fan-out.&lt;/p></description></item><item><title>How S.M.A.R.T. actually works: a mental model for operators</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-how-smart-works/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-how-smart-works/</guid><description>&lt;h1 id="how-smart-actually-works-a-mental-model-for-operators">How S.M.A.R.T. actually works: a mental model for operators&lt;/h1>
&lt;p>S.M.A.R.T. (Self-Monitoring, Analysis, and Reporting Technology) is not a tool, an agent, or a daemon. It is firmware-level self-instrumentation embedded in every modern HDD, SSD, and NVMe drive. The drive itself continuously monitors its internal health and reports what it finds. The &lt;code>smartctl&lt;/code> utility from &lt;code>smartmontools&lt;/code> is simply a reader: it queries the drive and prints what the firmware already knows. It has no independent intelligence about drive health.&lt;/p></description></item><item><title>How the Kubernetes control plane works: a mental model for operators</title><link>https://www.netdata.cloud/guides/kubernetes/how-kubernetes-control-plane-works/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/how-kubernetes-control-plane-works/</guid><description>&lt;h1 id="how-the-kubernetes-control-plane-works-a-mental-model-for-operators">How the Kubernetes control plane works: a mental model for operators&lt;/h1>
&lt;p>The Kubernetes control plane is often drawn as a single box labeled &amp;ldquo;Master.&amp;rdquo; In production, that abstraction fails the moment kubectl hangs, a node disappears, or a Deployment stays pending. The control plane is not a monolith. It is a pipeline of specialized components that hand off state through a single API server, and understanding those handoffs is what lets you triage a cluster outage without guessing.&lt;/p></description></item><item><title>How Tomcat actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-how-it-works-in-production/</guid><description>&lt;h1 id="how-tomcat-actually-works-in-production-a-mental-model-for-operators">How Tomcat actually works in production: a mental model for operators&lt;/h1>
&lt;p>Tomcat looks like a single process serving HTTP, but inside it is a nested container hierarchy with a three-tier buffering model in front of your servlet code. Most production incidents are not bugs in Tomcat; they are mismatches between what the operator thought was happening and what the connector and thread pool were doing. The thread pool exhausts while CPU sits at 5%. The JVM is alive but every request hangs. Sessions fill the heap while request rate looks normal. Clients see connection timeouts while Tomcat logs nothing.&lt;/p></description></item><item><title>How Traefik actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/traefik/traefik-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-how-it-works-in-production/</guid><description>&lt;h1 id="how-traefik-actually-works-in-production-a-mental-model-for-operators">How Traefik actually works in production: a mental model for operators&lt;/h1>
&lt;p>Most Traefik incidents are confusing because operators reason about it like nginx: a static proxy with a config file that gets reloaded on change. Traefik is not that. It is two systems sharing one process: a data-plane reverse proxy that terminates and forwards connections, and a control-plane reconciler that continuously watches configuration providers and rebuilds the routing table without a restart.&lt;/p></description></item><item><title>How uWSGI actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-how-it-works-in-production/</guid><description>&lt;h1 id="how-uwsgi-actually-works-in-production-a-mental-model-for-operators">How uWSGI actually works in production: a mental model for operators&lt;/h1>
&lt;p>uWSGI is a pre-fork application server. A single master process spawns a pool of worker processes, each running a full copy of your application. The master never serves requests. It manages worker lifecycle, enforces timeouts, and coordinates graceful reloads. Every request flows through the same path: a client connects to a socket, the kernel queues that connection in a listen backlog, and a worker calls &lt;code>accept()&lt;/code> to pull it out and process it synchronously.&lt;/p></description></item><item><title>How Varnish actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/varnish/varnish-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-how-it-works-in-production/</guid><description>&lt;h1 id="how-varnish-actually-works-in-production-a-mental-model-for-operators">How Varnish actually works in production: a mental model for operators&lt;/h1>
&lt;p>Varnish Cache is a reverse HTTP proxy that serves cached content from memory. Behind that description sits a set of interacting subsystems, each with its own failure modes and saturation points. Before you can interpret a counter or diagnose an incident, you need to understand how a request moves through the process, where worker threads are consumed, how the cache store is managed, and what runs in the background.&lt;/p></description></item><item><title>How vSphere actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-how-it-works-in-production/</guid><description>&lt;h1 id="how-vsphere-actually-works-in-production-a-mental-model-for-operators">How vSphere actually works in production: a mental model for operators&lt;/h1>
&lt;p>vSphere is a layered virtualization stack with three interdependent planes, and most production incidents cross plane boundaries. A guest that &amp;ldquo;feels slow&amp;rdquo; may be starved at the hypervisor scheduler, throttled by a forgotten CPU limit, fighting for IOPS at the storage layer, or sitting behind a vCenter whose database has bloated to the point that DRS stopped rebalancing the cluster.&lt;/p></description></item><item><title>How ZFS actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/zfs/zfs-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-how-it-works-in-production/</guid><description>&lt;h1 id="how-zfs-actually-works-in-production-a-mental-model-for-operators">How ZFS actually works in production: a mental model for operators&lt;/h1>
&lt;p>Every ZFS incident traces back to a small set of internal mechanisms. A write freeze at 3 a.m. is the transaction group pipeline stalling. A server that &amp;ldquo;runs out of memory&amp;rdquo; with gigabytes free is the ARC doing exactly what it was designed to do. A pool that was fine at 78% full and unusable at 88% is the metaslab allocator crossing a threshold, not a disk dying.&lt;/p></description></item><item><title>How ZooKeeper actually works in production: a mental model for operators</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-how-it-works-in-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-how-it-works-in-production/</guid><description>&lt;h1 id="how-zookeeper-actually-works-in-production-a-mental-model-for-operators">How ZooKeeper actually works in production: a mental model for operators&lt;/h1>
&lt;p>ZooKeeper is a distributed coordination service built on a replicated state machine. It maintains a hierarchical namespace of data nodes (znodes) entirely in memory and replicates every mutation across an ensemble of servers using the ZAB (ZooKeeper Atomic Broadcast) protocol.&lt;/p>
&lt;p>This is the mental model the rest of the ZooKeeper runbooks assume. It is the set of abstractions an on-call engineer needs to reason about why writes stall, sessions expire, a &amp;ldquo;healthy&amp;rdquo; node can return stale reads, and why the leader matters more than any other node in the cluster.&lt;/p></description></item><item><title>HP H3C Switch</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/hp-h3c-switch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/hp-h3c-switch/</guid><description/></item><item><title>HP ICF Switch</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/hp-icf-switch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/hp-icf-switch/</guid><description/></item><item><title>HP ILO</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/hp-ilo/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/hp-ilo/</guid><description/></item><item><title>HP Ilo4</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/hp-ilo4/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/hp-ilo4/</guid><description/></item><item><title>HPE Bladesystem Enclosure</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/hpe-bladesystem-enclosure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/hpe-bladesystem-enclosure/</guid><description/></item><item><title>HPE MSA</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/hpe-msa/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/hpe-msa/</guid><description/></item><item><title>HPE Nimble</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/hpe-nimble/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/hpe-nimble/</guid><description/></item><item><title>HPE Proliant</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/hpe-proliant/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/hpe-proliant/</guid><description/></item><item><title>HPE Smart Arrays</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/hpe-smart-arrays/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/hpe-smart-arrays/</guid><description/></item><item><title>HPE Smart Arrays Monitoring</title><link>https://www.netdata.cloud/monitoring-101/hpssa-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/hpssa-monitoring/</guid><description>&lt;h2 id="hpe-smart-arrays-monitoring">HPE Smart Arrays Monitoring&lt;/h2>
&lt;h3 id="what-is-hpe-smart-arrays">What Is HPE Smart Arrays?&lt;/h3>
&lt;p>HPE Smart Arrays are hardware RAID controllers from Hewlett Packard Enterprise designed to manage data storage that is highly reliable and enhances system performance. These controllers are integral for businesses that require robust and scalable storage systems, often used in server environments to support critical applications.&lt;/p>
&lt;h3 id="monitoring-hpe-smart-arrays-with-netdata">Monitoring HPE Smart Arrays With Netdata&lt;/h3>
&lt;p>Netdata provides a comprehensive monitoring tool specifically designed for HPE Smart Arrays, which allows real-time insights into the health and performance of your storage systems. By leveraging the HPE Smart Arrays monitoring tool, you can stay informed about any potential issues or performance bottlenecks before they impact your operations.&lt;/p></description></item><item><title>HTTP endpoint</title><link>https://www.netdata.cloud/integrations/all/http-endpoint/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/all/http-endpoint/</guid><description/></item><item><title>HTTP Endpoints</title><link>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/http-endpoints/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/http-endpoints/</guid><description/></item><item><title>HTTP Endpoints Monitoring</title><link>https://www.netdata.cloud/monitoring-101/http-endpoints-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/http-endpoints-monitoring/</guid><description>&lt;p>The &lt;a href="https://en.wikipedia.org/wiki/Hypertext_Transfer_Protocol">HTTP protocol&lt;/a> has become the de facto standard application layer protocol of the internet. From publicly available web sites and APIs to “inter-process” communications in REST based microservice architectures or large &lt;a href="https://en.wikipedia.org/wiki/Service-oriented_architecture">Service Oriented Architectures&lt;/a> based on &lt;a href="https://en.wikipedia.org/wiki/SOAP">SOAP&lt;/a>, you find HTTP being used again and again, due to its simplicity and our familiarity with it. How many protocols can you name that have &lt;a href="https://imgur.com/gallery/4KqWq">memes&lt;/a> for their status codes? Of course, such a popular protocol has endless pages written about how to properly monitor the services that rely on it, with many options specific to every use case. What you will learn here is how to get your basics done in monitoring HTTP endpoints, so you can be up and running in a few minutes, monitoring all HTTP services in your entire infrastructure.&lt;/p></description></item><item><title>HTTP Endpoints Monitoring</title><link>https://www.netdata.cloud/monitoring-101/httpcheck-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/httpcheck-monitoring/</guid><description>&lt;h2 id="http-endpoints-monitoring">HTTP Endpoints Monitoring&lt;/h2>
&lt;h3 id="what-is-http-endpoints">What Is HTTP Endpoints?&lt;/h3>
&lt;p>HTTP Endpoints are specific URLs or web services that applications or users interact with. Monitoring these endpoints involves tracking their availability, response times, and reliability to ensure smooth and efficient operation of web services.&lt;/p>
&lt;h3 id="monitoring-http-endpoints-with-netdata">Monitoring HTTP Endpoints With Netdata&lt;/h3>
&lt;p>Netdata offers a comprehensive HTTP Endpoints monitoring tool, designed to provide real-time insights into your web services. With Netdata, you can monitor HTTP Endpoints to identify issues like increased response times or server connectivity problems before they impact user experience. You can &lt;a href="https://app.netdata.cloud/?utm_source=website&amp;amp;utm_content=monitoring101">sign up for a free trial&lt;/a> or check out a &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">live demo&lt;/a> to see how it works.&lt;/p></description></item><item><title>HTTPD</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/httpd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/httpd/</guid><description/></item><item><title>Huawei</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/huawei/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/huawei/</guid><description/></item><item><title>Huawei Access Controllers</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/huawei-access-controllers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/huawei-access-controllers/</guid><description/></item><item><title>Huawei BGP</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/huawei-bgp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/huawei-bgp/</guid><description/></item><item><title>Huawei Routers</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/huawei-routers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/huawei-routers/</guid><description/></item><item><title>Huawei Switches</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/huawei-switches/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/huawei-switches/</guid><description/></item><item><title>Huawei Symantec Technologies Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/huawei-symantec-technologies-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/huawei-symantec-technologies-co-ltd-snmp-traps/</guid><description/></item><item><title>Huawei Technology Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/huawei-technology-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/huawei-technology-co-ltd-snmp-traps/</guid><description/></item><item><title>Hubble</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/hubble/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/hubble/</guid><description/></item><item><title>Hubble Monitoring</title><link>https://www.netdata.cloud/monitoring-101/hubble-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/hubble-monitoring/</guid><description>&lt;h2 id="hubble-monitoring">Hubble Monitoring&lt;/h2>
&lt;h3 id="what-is-hubble">What Is Hubble?&lt;/h3>
&lt;p>Hubble is an advanced, open-source network observability tool that provides detailed telemetry and security observability for cloud-native environments. By leveraging Hubble, DevOps engineers and IT professionals can gain deep insights into the network traffic and infrastructure within their Kubernetes clusters, powered by the Cilium network plugin.&lt;/p>
&lt;h3 id="monitoring-hubble-with-netdata">Monitoring Hubble With Netdata&lt;/h3>
&lt;p>Monitoring Hubble with Netdata allows you to leverage its openmetrics (Prometheus) exporter capabilities. Netdata seamlessly ingests data from any Prometheus exporter, allowing you to monitor Hubble efficiently. With Netdata, users benefit from automated dashboards and smart alerting systems without the need for setting up a Prometheus server or Grafana. This streamlined approach facilitates a faster and more efficient workflow, enabling SREs and IT admins to focus on responding to the alerts and insights rather than setting up and maintaining complex systems.&lt;/p></description></item><item><title>Huber Suhner Bktel GmbH Hfc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/huber-suhner-bktel-gmbh-hfc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/huber-suhner-bktel-gmbh-hfc-snmp-traps/</guid><description/></item><item><title>Hw Group S R O SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/hw-group-s-r-o-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/hw-group-s-r-o-snmp-traps/</guid><description/></item><item><title>hw.intrcnt</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/hw.intrcnt/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/hw.intrcnt/</guid><description/></item><item><title>Hyper-V</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/hyper-v/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/hyper-v/</guid><description/></item><item><title>Hypercom Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/hypercom-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/hypercom-inc-snmp-traps/</guid><description/></item><item><title>I/O errors in dmesg with clean SMART: the failure the drive can't see</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-host-io-errors-clean-smart/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-host-io-errors-clean-smart/</guid><description>&lt;h1 id="io-errors-in-dmesg-with-clean-smart-the-failure-the-drive-cant-see">I/O errors in dmesg with clean SMART: the failure the drive can&amp;rsquo;t see&lt;/h1>
&lt;p>&lt;code>I/O error, dev sda&lt;/code> scrolling through dmesg. The application is timing out on disk reads. You run &lt;code>smartctl -H /dev/sda&lt;/code> and it returns &lt;code>PASSED&lt;/code>. Reallocated sectors, pending sectors, offline uncorrectable: all zero. The ATA error log is empty.&lt;/p>
&lt;p>SMART reports what the drive firmware can observe about itself: media integrity, mechanical health, thermal state, NAND endurance. It cannot see failures in the transport layer between the host and the drive. When the failure lives in the SATA cable, the backplane connector, the HBA firmware, a SCSI error recovery loop, or a PCIe link, the drive firmware has no way to observe it. The host kernel sees every failed transaction. The drive does not.&lt;/p></description></item><item><title>IBM</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ibm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ibm/</guid><description/></item><item><title>IBM AIX systems Njmon</title><link>https://www.netdata.cloud/integrations/data-collection/applications/ibm-aix-systems-njmon/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/ibm-aix-systems-njmon/</guid><description/></item><item><title>IBM AIX Systems Njmon Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ibm_aix_njmon-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ibm_aix_njmon-monitoring/</guid><description>&lt;h2 id="ibm-aix-systems-njmon-monitoring">IBM AIX Systems Njmon Monitoring&lt;/h2>
&lt;h3 id="what-is-ibm-aix-systems-njmon">What Is IBM AIX Systems Njmon?&lt;/h3>
&lt;p>IBM AIX systems Njmon is a state-of-the-art tool designed for performance monitoring and management of IBM AIX operating environments. It effectively provides insights into system health and performance metrics, ensuring smooth IT infrastructure operation. With Njmon, you can seamlessly monitor data collection in real time without impacting system performance.&lt;/p>
&lt;h3 id="monitoring-ibm-aix-systems-njmon-with-netdata">Monitoring IBM AIX Systems Njmon With Netdata&lt;/h3>
&lt;p>To monitor IBM AIX Systems Njmon efficiently, Netdata provides a robust solution utilizing an openmetrics (Prometheus) exporter. This integration allows users to ingest data directly from any Prometheus endpoint without needing a dedicated Prometheus server or Grafana setup. With Netdata, you gain access to automated dashboards, alerts, and insights, drastically simplifying the process of monitoring complex systems like IBM AIX.&lt;/p></description></item><item><title>IBM CryptoExpress (CEX) cards</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/ibm-cryptoexpress-cex-cards/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/ibm-cryptoexpress-cex-cards/</guid><description/></item><item><title>IBM CryptoExpress (CEX) Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ibm_cex-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ibm_cex-monitoring/</guid><description>&lt;h2 id="ibm-cryptoexpress-cex-monitoring">IBM CryptoExpress (CEX) Monitoring&lt;/h2>
&lt;h3 id="what-is-ibm-cryptoexpress-cex">What Is IBM CryptoExpress (CEX)?&lt;/h3>
&lt;p>IBM CryptoExpress (CEX) cards are specialized hardware devices used in IBM systems for cryptographic operations, providing secure encryption and decryption services. These cards are crucial for businesses that require high-security standards for their data transactions and management.&lt;/p>
&lt;h3 id="monitoring-ibm-cryptoexpress-cex-with-netdata">Monitoring IBM CryptoExpress (CEX) With Netdata&lt;/h3>
&lt;p>To efficiently monitor IBM CryptoExpress (CEX), Netdata utilizes the openmetrics (Prometheus) exporter available for these cards. This allows Netdata to seamlessly integrate with any existing Prometheus exporter to collect vital metrics that give insights into cryptographic performance and management. With Netdata, automated dashboards and alerts are readily available without the need for setting up a separate Prometheus server or Grafana, simplifying the monitoring process.&lt;/p></description></item><item><title>IBM Datapower Gateway</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ibm-datapower-gateway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ibm-datapower-gateway/</guid><description/></item><item><title>IBM DB2</title><link>https://www.netdata.cloud/integrations/data-collection/databases/ibm-db2/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/ibm-db2/</guid><description/></item><item><title>Ibm Eserver X SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ibm-eserver-x-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ibm-eserver-x-snmp-traps/</guid><description/></item><item><title>Ibm Https W3 Ibm Com Standards SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ibm-https-w3-ibm-com-standards-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ibm-https-w3-ibm-com-standards-snmp-traps/</guid><description/></item><item><title>IBM i (AS/400)</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/ibm-i-as-400/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/ibm-i-as-400/</guid><description/></item><item><title>IBM Lenovo Server</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ibm-lenovo-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ibm-lenovo-server/</guid><description/></item><item><title>IBM MQ</title><link>https://www.netdata.cloud/integrations/data-collection/applications/ibm-mq/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/ibm-mq/</guid><description/></item><item><title>IBM MQ</title><link>https://www.netdata.cloud/integrations/data-collection/databases/ibm-mq/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/ibm-mq/</guid><description/></item><item><title>IBM MQ Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ibm_mq-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ibm_mq-monitoring/</guid><description>&lt;h2 id="ibm-mq-monitoring">IBM MQ Monitoring&lt;/h2>
&lt;h3 id="what-is-ibm-mq">What Is IBM MQ?&lt;/h3>
&lt;p>IBM MQ is a powerful messaging middleware designed to simplify, accelerate, and facilitate the transport of messages between different platforms and applications. It allows for secure and reliable message transfer, ensuring that businesses can integrate multiple complex systems seamlessly.&lt;/p>
&lt;h3 id="monitoring-ibm-mq-with-netdata">Monitoring IBM MQ With Netdata&lt;/h3>
&lt;p>To monitor IBM MQ effectively, Netdata provides a comprehensive solution using the &lt;a href="https://github.com/agebhar1/mq_exporter">MQ Exporter&lt;/a>, which is an openmetrics (prometheus) exporter. With Netdata, you can effortlessly ingest data from any Prometheus exporter without the need for a Prometheus server or Grafana for visualizations. This means you can access automated dashboards, receive intelligent alerts, and enjoy rich visual insights into your IBM MQ metrics—all integrated seamlessly within the Netdata platform.&lt;/p></description></item><item><title>IBM Spectrum</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ibm-spectrum/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ibm-spectrum/</guid><description/></item><item><title>IBM Spectrum Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ibm_spectrum-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ibm_spectrum-monitoring/</guid><description>&lt;h2 id="ibm-spectrum-monitoring">IBM Spectrum Monitoring&lt;/h2>
&lt;h3 id="what-is-ibm-spectrum">What Is IBM Spectrum?&lt;/h3>
&lt;p>IBM Spectrum is a suite of data management software from IBM, designed to manage storage resources as efficiently and cost-effectively as possible. With solutions that optimize analytics, data protection, and storage, IBM Spectrum is essential for organizations that rely on large-scale data environments.&lt;/p>
&lt;h3 id="monitoring-ibm-spectrum-with-netdata">Monitoring IBM Spectrum With Netdata&lt;/h3>
&lt;p>To monitor IBM Spectrum, Netdata leverages an openmetrics, specifically a Prometheus exporter, known as the &lt;a href="https://github.com/topine/ibm-spectrum-exporter">IBM Spectrum Exporter&lt;/a>. Netdata’s integration provides dynamic monitoring capabilities, turning data collected from IBM Spectrum into insightful analytics. Netdata is unique in its ability to ingest data from any Prometheus exporter, and users benefit from automated dashboards, real-time alerts, and more—all without needing a separate Prometheus server or Grafana.&lt;/p></description></item><item><title>IBM Spectrum Virtualize</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ibm-spectrum-virtualize/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ibm-spectrum-virtualize/</guid><description/></item><item><title>IBM Spectrum Virtualize Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ibm_spectrum_virtualize-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ibm_spectrum_virtualize-monitoring/</guid><description>&lt;h2 id="ibm-spectrum-virtualize-monitoring">IBM Spectrum Virtualize Monitoring&lt;/h2>
&lt;h3 id="what-is-ibm-spectrum-virtualize">What Is IBM Spectrum Virtualize?&lt;/h3>
&lt;p>IBM Spectrum Virtualize is a sophisticated software that enables efficient storage virtualization, streamlining storage management and enhancing performance. It is widely utilized in various IT environments to manage storage resources, boosting performance and scalability across different platforms.&lt;/p>
&lt;h3 id="monitoring-ibm-spectrum-virtualize-with-netdata">Monitoring IBM Spectrum Virtualize With Netdata&lt;/h3>
&lt;p>Monitoring IBM Spectrum Virtualize becomes seamless and highly efficient with Netdata. By using an openmetrics (Prometheus) exporter, specifically the &lt;a href="https://github.com/bluecmd/spectrum_virtualize_exporter">spectrum_virtualize_exporter&lt;/a>, Netdata can effortlessly ingest data from any Prometheus exporter. This functionality allows users to access automated dashboards, receive real-time alerts, and gain deep insights without requiring a Prometheus server or Grafana. The integration facilitates comprehensive visibility into your storage virtualization metrics, ensuring optimal performance and resource utilization.&lt;/p></description></item><item><title>IBM WebSphere JMX</title><link>https://www.netdata.cloud/integrations/data-collection/applications/ibm-websphere-jmx/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/ibm-websphere-jmx/</guid><description/></item><item><title>IBM WebSphere MicroProfile</title><link>https://www.netdata.cloud/integrations/data-collection/applications/ibm-websphere-microprofile/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/ibm-websphere-microprofile/</guid><description/></item><item><title>IBM WebSphere PMI</title><link>https://www.netdata.cloud/integrations/data-collection/applications/ibm-websphere-pmi/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/ibm-websphere-pmi/</guid><description/></item><item><title>IBM Z Hardware Management Console</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/ibm-z-hardware-management-console/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/ibm-z-hardware-management-console/</guid><description/></item><item><title>IBM Z Hardware Management Console Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ibm_zhmc-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ibm_zhmc-monitoring/</guid><description>&lt;h2 id="ibm-z-hardware-management-console-monitoring">IBM Z Hardware Management Console Monitoring&lt;/h2>
&lt;h3 id="what-is-ibm-z-hardware-management-console">What Is IBM Z Hardware Management Console?&lt;/h3>
&lt;p>The IBM Z Hardware Management Console (HMC) is a critical component for managing enterprise-class mainframe systems. It provides an interface for controlling hardware, managing partitions, and configuring features of IBM Z servers, enabling efficient mainframe management.&lt;/p>
&lt;h3 id="monitoring-ibm-z-hardware-management-console-with-netdata">Monitoring IBM Z Hardware Management Console With Netdata&lt;/h3>
&lt;p>To optimize performance and ensure reliable operation, monitoring the IBM Z HMC is crucial. Netdata provides a robust solution for this task by using an openmetrics (Prometheus) exporter, specifically the &lt;a href="https://github.com/zhmcclient/zhmc-prometheus-exporter">IBM Z HMC Exporter&lt;/a>. This enables seamless integration with Netdata&amp;rsquo;s monitoring platform, allowing users to ingest metrics without needing a standalone Prometheus server or Grafana setup.&lt;/p></description></item><item><title>Ibrix Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ibrix-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ibrix-corp-snmp-traps/</guid><description/></item><item><title>Icecast</title><link>https://www.netdata.cloud/integrations/data-collection/applications/icecast/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/icecast/</guid><description/></item><item><title>Icecast Monitoring</title><link>https://www.netdata.cloud/monitoring-101/icecast-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/icecast-monitoring/</guid><description>&lt;h2 id="icecast-monitoring">Icecast Monitoring&lt;/h2>
&lt;h3 id="what-is-icecast">What Is Icecast?&lt;/h3>
&lt;p>Icecast is a free and open-source streaming media server which supports various streaming formats, including MP3. It&amp;rsquo;s widely used for setting up online radio stations and creating or distributing online audio content. For more information, visit the &lt;a href="https://icecast.org/">Icecast official website&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-icecast-with-netdata">Monitoring Icecast With Netdata&lt;/h3>
&lt;p>Netdata offers an unparalleled solution for monitoring Icecast, providing real-time insights and detailed metrics to optimize performance. As a comprehensive Icecast monitoring tool, Netdata allows you to track listener counts, server performance, and more, ensuring your streaming server operates at its best.&lt;/p></description></item><item><title>Idirect SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/idirect-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/idirect-snmp-traps/</guid><description/></item><item><title>Idle OS Jitter</title><link>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/idle-os-jitter/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/idle-os-jitter/</guid><description/></item><item><title>IDRAC</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/idrac/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/idrac/</guid><description/></item><item><title>Ieee 802 SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ieee-802-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ieee-802-snmp-traps/</guid><description/></item><item><title>Ieee Lldp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ieee-lldp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ieee-lldp-snmp-traps/</guid><description/></item><item><title>IIS</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/iis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/iis/</guid><description/></item><item><title>ilert</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/ilert/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/ilert/</guid><description/></item><item><title>ilert</title><link>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/ilert/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/ilert/</guid><description/></item><item><title>Image Processing Techniques Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/image-processing-techniques-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/image-processing-techniques-ltd-snmp-traps/</guid><description/></item><item><title>Image Project Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/image-project-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/image-project-inc-snmp-traps/</guid><description/></item><item><title>Imv Victron B.V. SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/imv-victron-b.v.-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/imv-victron-b.v.-snmp-traps/</guid><description/></item><item><title>Inca Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/inca-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/inca-networks-inc-snmp-traps/</guid><description/></item><item><title>Incognito Software Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/incognito-software-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/incognito-software-systems-inc-snmp-traps/</guid><description/></item><item><title>Independence Technologies Inc Iti SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/independence-technologies-inc-iti-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/independence-technologies-inc-iti-snmp-traps/</guid><description/></item><item><title>Infinera Coriant Groove</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/infinera-coriant-groove/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/infinera-coriant-groove/</guid><description/></item><item><title>Infinet LLC Formerly Aqua Project Group SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/infinet-llc-formerly-aqua-project-group-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/infinet-llc-formerly-aqua-project-group-snmp-traps/</guid><description/></item><item><title>InfiniBand</title><link>https://www.netdata.cloud/integrations/data-collection/networking/infiniband/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/infiniband/</guid><description/></item><item><title>Inflection Systems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/inflection-systems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/inflection-systems-snmp-traps/</guid><description/></item><item><title>InfluxDB</title><link>https://www.netdata.cloud/integrations/data-collection/databases/influxdb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/influxdb/</guid><description/></item><item><title>InfluxDB Monitoring</title><link>https://www.netdata.cloud/monitoring-101/influxdb-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/influxdb-monitoring/</guid><description>&lt;h2 id="influxdb-monitoring">InfluxDB Monitoring&lt;/h2>
&lt;h3 id="what-is-influxdb">What Is InfluxDB?&lt;/h3>
&lt;p>InfluxDB is a high-performance time series database designed for storing and analyzing time-stamped data. Used extensively for monitoring, it offers powerful data retention, querying, and analytics features that make it a popular choice for real-time applications. If you&amp;rsquo;re working with metrics, logs, or any form of timestamped data, understanding how to effectively monitor InfluxDB can significantly enhance your infrastructure&amp;rsquo;s performance.&lt;/p>
&lt;h3 id="monitoring-influxdb-with-netdata">Monitoring InfluxDB With Netdata&lt;/h3>
&lt;p>To monitor InfluxDB seamlessly, Netdata utilizes an openmetrics (Prometheus) exporter. The flexibility of Netdata allows it to ingest data from any Prometheus exporter. This means you can access automated dashboards, receive alerts, and more, without the need for a dedicated Prometheus server or Grafana. By leveraging the &lt;a href="https://github.com/prometheus/influxdb_exporter">InfluxDB exporter&lt;/a>, Netdata provides comprehensive visibility into your InfluxDB instances, enabling you to optimize performance efficiently.&lt;/p></description></item><item><title>Infoblox Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/infoblox-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/infoblox-inc-snmp-traps/</guid><description/></item><item><title>Infoblox Ipam</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/infoblox-ipam/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/infoblox-ipam/</guid><description/></item><item><title>Infortrend Technology Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/infortrend-technology-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/infortrend-technology-inc-snmp-traps/</guid><description/></item><item><title>Infosim SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/infosim-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/infosim-snmp-traps/</guid><description/></item><item><title>Infratec Plus GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/infratec-plus-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/infratec-plus-gmbh-snmp-traps/</guid><description/></item><item><title>Ingrasys SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ingrasys-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ingrasys-snmp-traps/</guid><description/></item><item><title>Inktomi Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/inktomi-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/inktomi-corporation-snmp-traps/</guid><description/></item><item><title>Innominate Security Technologies AG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/innominate-security-technologies-ag-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/innominate-security-technologies-ag-snmp-traps/</guid><description/></item><item><title>Innovaphone GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/innovaphone-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/innovaphone-gmbh-snmp-traps/</guid><description/></item><item><title>Innovative Circuit Technology Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/innovative-circuit-technology-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/innovative-circuit-technology-ltd-snmp-traps/</guid><description/></item><item><title>Inoc Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/inoc-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/inoc-inc-snmp-traps/</guid><description/></item><item><title>Inova Dc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/inova-dc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/inova-dc-snmp-traps/</guid><description/></item><item><title>Inovonics Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/inovonics-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/inovonics-inc-snmp-traps/</guid><description/></item><item><title>Integrations</title><link>https://www.netdata.cloud/integrations/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/</guid><description/></item><item><title>Intel Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/intel-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/intel-corporation-snmp-traps/</guid><description/></item><item><title>Intel GPU</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/intel-gpu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/intel-gpu/</guid><description/></item><item><title>Intel GPU Monitoring</title><link>https://www.netdata.cloud/monitoring-101/intelgpu-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/intelgpu-monitoring/</guid><description>&lt;h2 id="intel-gpu-monitoring">Intel GPU Monitoring&lt;/h2>
&lt;h3 id="what-is-intel-gpu">What Is Intel GPU?&lt;/h3>
&lt;p>Intel Graphics Processing Units (GPUs) are integrated into Intel CPUs, offering powerful graphics capabilities for a variety of applications. These GPUs are particularly important for managing graphics-intensive workloads and improving the overall performance of systems using Intel-based hardware.&lt;/p>
&lt;h3 id="monitoring-intel-gpu-with-netdata">Monitoring Intel GPU With Netdata&lt;/h3>
&lt;p>Netdata&amp;rsquo;s Intel GPU monitoring tool provides real-time, insightful data that allows you to track the performance and health of your Intel GPU seamlessly. With Netdata, you can monitor Intel GPU metrics like frequency, power consumption, and engine utilization without compromising system security or performance.&lt;/p></description></item><item><title>Intelligent Platform Management Interface (IPMI)</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/intelligent-platform-management-interface-ipmi/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/intelligent-platform-management-interface-ipmi/</guid><description/></item><item><title>Inter Process Communication</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/inter-process-communication/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/inter-process-communication/</guid><description/></item><item><title>Interface discards with low utilization: diagnosing ifInDiscards/ifOutDiscards</title><link>https://www.netdata.cloud/guides/network/network-interface-discards/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-interface-discards/</guid><description>&lt;h1 id="interface-discards-with-low-utilization-diagnosing-ifindiscardsifoutdiscards">Interface discards with low utilization: diagnosing ifInDiscards/ifOutDiscards&lt;/h1>
&lt;p>ifOutDiscards is climbing on a critical uplink. Utilization sits at 35%. No CRC errors, no input errors, no physical-layer alarms. The link is up and passing traffic, but something is silently dropping packets, and your averaged utilization metrics are not telling you why.&lt;/p>
&lt;p>The gap between what the counters show and what the silicon is doing comes down to two things: the averaging window on utilization, and the fact that discards happen at buffer-queue granularity, not at link-rate granularity.&lt;/p></description></item><item><title>Interface flapping: link up/down storms and their blast radius</title><link>https://www.netdata.cloud/guides/network/network-interface-flapping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-interface-flapping/</guid><description>&lt;h1 id="interface-flapping-link-updown-storms-and-their-blast-radius">Interface flapping: link up/down storms and their blast radius&lt;/h1>
&lt;p>Interface flapping is when a network interface oscillates rapidly between up and down states. Each transition generates a linkDown/linkUp trap pair, syslog entries, and an STP topology change notification. At low rates this is operational noise. At high rates it becomes a multi-layer failure: the trap receiver overflows, the syslog pipeline saturates, STP reconvergence flushes MAC tables across the VLAN, and the monitoring platform reports misleading availability because the poll interval is slower than the flap cadence.&lt;/p></description></item><item><title>Interface input/output errors: finding the bad link with ifInErrors/ifOutErrors</title><link>https://www.netdata.cloud/guides/network/network-interface-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-interface-errors/</guid><description>&lt;h1 id="interface-inputoutput-errors-finding-the-bad-link-with-ifinerrorsifouterrors">Interface input/output errors: finding the bad link with ifInErrors/ifOutErrors&lt;/h1>
&lt;p>A switch port reports 47,000 input errors and climbing. The interface is up and traffic is flowing. Your monitoring fired an alert on ifInErrors crossing threshold. Now you need to determine whether this is a dirty fiber, a dying SFP, a duplex mismatch, buffer exhaustion, or a counter artifact from an interface flap.&lt;/p>
&lt;p>ifInErrors (.1.3.6.1.2.1.2.2.1.14) and ifOutErrors (.1.3.6.1.2.1.2.2.1.20) are aggregate counters. They tell you something is wrong, but not what. The counter is a sum of multiple error types: CRC, alignment, runts, giants, overruns, frame errors, and on some platforms, input drops. A frame that arrives with both a CRC error and a runt condition increments ifInErrors by exactly 1, not 2. You cannot reconcile the sub-counters against the aggregate by simple addition.&lt;/p></description></item><item><title>Interface saturation: measuring utilization against ifHighSpeed correctly</title><link>https://www.netdata.cloud/guides/network/network-interface-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-interface-saturation/</guid><description>&lt;h1 id="interface-saturation-measuring-utilization-against-ifhighspeed-correctly">Interface saturation: measuring utilization against ifHighSpeed correctly&lt;/h1>
&lt;p>Interface utilization is one of the most frequently miscomputed network metrics. The formula is simple: divide throughput by capacity, multiply by 100. But the SNMP objects for throughput and capacity have different precision, units, and failure modes. An interface that is genuinely saturated can report 0% utilization. A healthy link can report 300%.&lt;/p>
&lt;p>The root cause is almost always the denominator. &lt;code>ifSpeed&lt;/code> (IF-MIB &lt;code>.1.3.6.1.2.1.2.2.1.5&lt;/code>) is a 32-bit Gauge that caps at 4,294,967,295 bps, approximately 4.29 Gbps. For any link faster than that, &lt;code>ifSpeed&lt;/code> saturates at its maximum value and utilization computed against it is wrong. The correct denominator is &lt;code>ifHighSpeed&lt;/code> (&lt;code>.1.3.6.1.2.1.31.1.1.1.15&lt;/code>), which reports speed in units of 1,000,000 bps (Mbps) and has no practical upper bound. A value of 10000 means 10 Gbps; 100000 means 100 Gbps.&lt;/p></description></item><item><title>Internet Information Services (IIS) Monitoring</title><link>https://www.netdata.cloud/monitoring-101/iis-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/iis-monitoring/</guid><description>&lt;h2 id="iis-internet-information-services">IIS (Internet Information Services)&lt;/h2>
&lt;p>IIS stands for Internet Information Services, which is a web server software package designed for Windows Server. IIS can host static or dynamic websites, and serve content such as HTML, JavaScript, or media files to a user’s browser. IIS is versatile and stable, and it has been widely used in production for many years.&lt;/p>
&lt;p>IIS is composed of the following basic components:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Web Server&lt;/strong>
&lt;ul>
&lt;li>The most common use-case IIS is used for is to host websites and ASP.NET web applications.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;strong>Security&lt;/strong>
&lt;ul>
&lt;li>IIS provides robust security through Windows authentication, and is capable of filtering requests, and management of TLS certificates, so that you can enable HTTPS or SFTP on your web server.&lt;/li>
&lt;/ul>
&lt;/li>
&lt;li>&lt;strong>Management&lt;/strong>
&lt;ul>
&lt;li>IIS can be managed locally or remotely through the Console, CLI, PowerShell and other tools&lt;/li>
&lt;/ul>
&lt;/li>
&lt;/ul>
&lt;h2 id="iis-monitoring">IIS Monitoring&lt;/h2>
&lt;p>To holistically monitor IIS, you need to track various aspects of its performance, such as:&lt;/p></description></item><item><title>Interrupts</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/interrupts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/interrupts/</guid><description/></item><item><title>Intersystems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/intersystems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/intersystems-snmp-traps/</guid><description/></item><item><title>Invidi Technologies SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/invidi-technologies-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/invidi-technologies-snmp-traps/</guid><description/></item><item><title>IOPing</title><link>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/ioping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/ioping/</guid><description/></item><item><title>Ip Infusion Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ip-infusion-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ip-infusion-inc-snmp-traps/</guid><description/></item><item><title>IP Virtual Server</title><link>https://www.netdata.cloud/integrations/data-collection/networking/ip-virtual-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/ip-virtual-server/</guid><description/></item><item><title>IP2Location LITE IP-Country</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/ip2location-lite-ip-country/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/ip2location-lite-ip-country/</guid><description/></item><item><title>IPDeny Country Zones</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/ipdeny-country-zones/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/ipdeny-country-zones/</guid><description/></item><item><title>Ipf Technology Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ipf-technology-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ipf-technology-ltd-snmp-traps/</guid><description/></item><item><title>IPFIX</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/flow-protocols/ipfix/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/flow-protocols/ipfix/</guid><description/></item><item><title>IPFS</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ipfs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/ipfs/</guid><description/></item><item><title>IPFS Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ipfs-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ipfs-monitoring/</guid><description>&lt;h2 id="ipfs-monitoring">IPFS Monitoring&lt;/h2>
&lt;h3 id="what-is-ipfs">What Is IPFS?&lt;/h3>
&lt;p>&lt;a href="https://ipfs.tech/">IPFS (InterPlanetary File System)&lt;/a> is a protocol and peer-to-peer network for storing and sharing data in a distributed file system. IPFS allows users to host and receive content in a decentralized manner, making it well-suited for distributed and resilient applications.&lt;/p>
&lt;h3 id="monitoring-ipfs-with-netdata">Monitoring IPFS With Netdata&lt;/h3>
&lt;p>To effectively monitor IPFS, you need a robust tool that provides real-time insights into network activity, file storage, and other critical metrics. Netdata&amp;rsquo;s IPFS monitoring tool offers comprehensive visibility, enabling you to ensure the stability and performance of your IPFS nodes.&lt;/p></description></item><item><title>ipfw</title><link>https://www.netdata.cloud/integrations/data-collection/networking/ipfw/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/ipfw/</guid><description/></item><item><title>IPIP Country Database</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/ipip-country-database/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/ipip-country-database/</guid><description/></item><item><title>IPtoASN</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/iptoasn/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/iptoasn/</guid><description/></item><item><title>IPv6 Socket Statistics</title><link>https://www.netdata.cloud/integrations/data-collection/networking/ipv6-socket-statistics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/ipv6-socket-statistics/</guid><description/></item><item><title>IRC</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/irc/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/irc/</guid><description/></item><item><title>IRONdb</title><link>https://www.netdata.cloud/integrations/exporters/irondb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/irondb/</guid><description/></item><item><title>Ironport Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ironport-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ironport-systems-inc-snmp-traps/</guid><description/></item><item><title>ISC DHCP</title><link>https://www.netdata.cloud/integrations/data-collection/networking/isc-dhcp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/isc-dhcp/</guid><description/></item><item><title>ISC DHCP Monitoring</title><link>https://www.netdata.cloud/monitoring-101/isc_dhcpd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/isc_dhcpd-monitoring/</guid><description>&lt;h2 id="isc-dhcp-monitoring">ISC DHCP Monitoring&lt;/h2>
&lt;h3 id="what-is-isc-dhcp">What Is ISC DHCP?&lt;/h3>
&lt;p>&lt;a href="https://www.isc.org/dhcp/">ISC DHCP&lt;/a> is a widely used open-source DHCP server software that automates the assignment and management of IP addresses, enabling efficient network management. It is capable of serving both IPv4 and IPv6 clients and supports various network topologies.&lt;/p>
&lt;h3 id="monitoring-isc-dhcp-with-netdata">Monitoring ISC DHCP With Netdata&lt;/h3>
&lt;p>Monitoring ISC DHCP efficiently is crucial for ensuring network reliability and performance. With Netdata, a leading ISC DHCP monitoring tool, you gain real-time insights into DHCP lease usage and pool utilization. Netdata provides comprehensive data visualization to help you track key metrics and quickly diagnose any issues in your DHCP environment.&lt;/p></description></item><item><title>ISC DHCPd Monitoring</title><link>https://www.netdata.cloud/monitoring-101/isc-dhcpd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/isc-dhcpd-monitoring/</guid><description>&lt;h2 id="what-is-isc-dhcpd">What is ISC DHCPd?&lt;/h2>
&lt;p>ISC DHCP is a DHCP server that supports both IPv4 and IPv6, and is suitable for use in high-volume and high-reliability applications.&lt;/p>
&lt;h2 id="monitoring-isc-dhcpd-with-netdata">Monitoring ISC DHCPd with Netdata&lt;/h2>
&lt;p>The prerequisites for monitoring ISC DHCPd with Netdata are to have ISC DHCPd and &lt;a href="https://learn.netdata.cloud/docs/cloud/get-started/">Netdata installed&lt;/a> on your system.&lt;/p>
&lt;p>Netdata auto discovers hundreds of services, and for those it doesn&amp;rsquo;t turning on manual discovery is a one line configuration. For more information on configuring Netdata for ISC DHCPd monitoring please read the collector &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/isc_dhcpd/">documentation&lt;/a>.&lt;/p></description></item><item><title>ISCBind (RNDC) Monitoring</title><link>https://www.netdata.cloud/monitoring-101/iscbind-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/iscbind-monitoring/</guid><description>&lt;h2 id="what-is-iscbind-rndc">What is ISCBind (RNDC)?&lt;/h2>
&lt;h2 id="monitoring-iscbind-rndc-with-netdata">Monitoring ISCBind (RNDC) with Netdata&lt;/h2>
&lt;p>The prerequisites for monitoring ISCBind with Netdata are to have ISCBind and &lt;a href="https://learn.netdata.cloud/docs/cloud/get-started/">Netdata installed&lt;/a> on your system.&lt;/p>
&lt;p>Netdata auto discovers hundreds of services, and for those it doesn&amp;rsquo;t turning on manual discovery is a one line configuration. For more information on configuring Netdata for ISCBind monitoring please read the collector &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/python.d.plugin/bind_rndc/">documentation&lt;/a>.&lt;/p>
&lt;p>You should now see the ISCBind section on the Overview tab in Netdata Cloud already populated with charts about all the metrics you care about.&lt;/p></description></item><item><title>Isilon</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/isilon/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/isilon/</guid><description/></item><item><title>Isilon Ststems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/isilon-ststems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/isilon-ststems-snmp-traps/</guid><description/></item><item><title>It Watchdogs Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/it-watchdogs-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/it-watchdogs-inc-snmp-traps/</guid><description/></item><item><title>Ixsystems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ixsystems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ixsystems-snmp-traps/</guid><description/></item><item><title>Ixsystems Truenas</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ixsystems-truenas/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ixsystems-truenas/</guid><description/></item><item><title>Jacarta Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/jacarta-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/jacarta-ltd-snmp-traps/</guid><description/></item><item><title>Jacobs University Bremen SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/jacobs-university-bremen-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/jacobs-university-bremen-snmp-traps/</guid><description/></item><item><title>Jacques Technologies SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/jacques-technologies-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/jacques-technologies-snmp-traps/</guid><description/></item><item><title>Janitza Electronics GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/janitza-electronics-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/janitza-electronics-gmbh-snmp-traps/</guid><description/></item><item><title>Jarvis Standing Desk</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/jarvis-standing-desk/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/jarvis-standing-desk/</guid><description/></item><item><title>Jarvis Standing Desk Monitoring</title><link>https://www.netdata.cloud/monitoring-101/jarvis-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/jarvis-monitoring/</guid><description>&lt;h2 id="jarvis-standing-desk-monitoring">Jarvis Standing Desk Monitoring&lt;/h2>
&lt;h3 id="what-is-jarvis-standing-desk">What Is Jarvis Standing Desk?&lt;/h3>
&lt;p>The Jarvis Standing Desk is an ergonomic workspace solution designed to promote better posture and flexibility in the office or home environment. Its ability to adjust height with ease helps users to maintain productive and healthy working conditions. With the increasing need for workspace ergonomics, monitoring the usage of standing desks like Jarvis becomes crucial to ensure optimal usage and benefit from the ergonomic advantages they offer.&lt;/p></description></item><item><title>Java Springboot Applications Monitoring</title><link>https://www.netdata.cloud/monitoring-101/springboot2-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/springboot2-monitoring/</guid><description>&lt;h2 id="what-is-java-springboot2">What is Java Springboot2?&lt;/h2>
&lt;p>Java Springboot2 is an open source Java-based framework used to create web applications. It is based on the Spring framework and provides a range of features such as enhanced dependency injection, abstract data access layers, and a powerful configuration system. Springboot2 also enables developers to quickly create and deploy applications and services. It can be used to rapidly develop applications and services that are cloud-native and highly secure.&lt;/p></description></item><item><title>Jds Uniphase Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/jds-uniphase-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/jds-uniphase-corporation-snmp-traps/</guid><description/></item><item><title>Jenkins</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/jenkins/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/jenkins/</guid><description/></item><item><title>Jenkins Monitoring</title><link>https://www.netdata.cloud/monitoring-101/jenkins-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/jenkins-monitoring/</guid><description>&lt;h2 id="jenkins-monitoring">Jenkins Monitoring&lt;/h2>
&lt;h3 id="what-is-jenkins">What Is Jenkins?&lt;/h3>
&lt;p>&lt;a href="https://www.jenkins.io/">Jenkins&lt;/a> is an open-source automation server that is used to automate the process of building, testing, and deploying software. It is a continuous integration tool that enables developers to continually test and integrate their codebase. It supports a range of plugins for building and deploying applications.&lt;/p>
&lt;h3 id="monitoring-jenkins-with-netdata">Monitoring Jenkins With Netdata&lt;/h3>
&lt;p>To effectively monitor Jenkins, Netdata employs a robust methodology using an openmetrics (Prometheus) exporter. This integration is seamless as Netdata can ingest data from any Prometheus exporter. It saves users from having to maintain separate Prometheus server and Grafana dashboard setups. With Netdata, automated dashboards and alerts become instantly accessible, providing real-time insights and alerting on critical metrics.&lt;/p></description></item><item><title>JMX</title><link>https://www.netdata.cloud/integrations/data-collection/applications/jmx/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/jmx/</guid><description/></item><item><title>JMX Monitoring</title><link>https://www.netdata.cloud/monitoring-101/jmx-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/jmx-monitoring/</guid><description>&lt;h2 id="jmx-monitoring">JMX Monitoring&lt;/h2>
&lt;h3 id="what-is-jmx">What Is JMX?&lt;/h3>
&lt;p>Java Management Extensions (JMX) is a technology that facilitates the management and monitoring of Java applications. It provides crucial insights into application performance through MBeans (Managed Beans), which represent resources such as applications or any device that needs to be managed.&lt;/p>
&lt;h3 id="monitoring-jmx-with-netdata">Monitoring JMX With Netdata&lt;/h3>
&lt;p>Leveraging Netdata for JMX monitoring adds a powerful dimension to your DevOps toolkit. With Netdata, you can monitor JMX seamlessly by using an OpenMetrics (Prometheus) exporter. Netdata can ingest data from any Prometheus exporter, offering automated dashboards, alerts, and more, without requiring a Prometheus server or Grafana. This means you can have real-time visibility into JMX metrics with minimal setup.&lt;/p></description></item><item><title>Johnson Controls Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/johnson-controls-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/johnson-controls-inc-snmp-traps/</guid><description/></item><item><title>journald</title><link>https://www.netdata.cloud/integrations/data-collection/applications/journald/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/journald/</guid><description/></item><item><title>Journald Monitoring</title><link>https://www.netdata.cloud/monitoring-101/journald-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/journald-monitoring/</guid><description>&lt;h2 id="journald-monitoring">Journald Monitoring&lt;/h2>
&lt;h3 id="what-is-journald">What Is Journald?&lt;/h3>
&lt;p>Journald is a component of the systemd suite that manages and facilitates logging within Linux systems. It collects and stores logs in a structured manner, offering a unified logging system to better manage system events and logs. Understanding the intricacies of journald is essential for system administrators and developers who rely on consistent and efficient log management.&lt;/p>
&lt;h3 id="monitoring-journald-with-netdata">Monitoring Journald With Netdata&lt;/h3>
&lt;p>To monitor journald, Netdata employs an openmetrics (Prometheus) exporter, specifically the &lt;a href="https://github.com/dead-claudia/journald-exporter">journald-exporter&lt;/a>. This approach allows Netdata to ingest data from any Prometheus exporter, providing users with automated dashboards and alerts without the need for a Prometheus server or Grafana. This capability makes Netdata an excellent journald monitoring tool due to its simplicity and efficiency in onboarding and dashboard automation.&lt;/p></description></item><item><title>JSON</title><link>https://www.netdata.cloud/integrations/exporters/json/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/json/</guid><description/></item><item><title>Juniper</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/juniper/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/juniper/</guid><description/></item><item><title>Juniper BGP</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/juniper-bgp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/juniper-bgp/</guid><description/></item><item><title>Juniper EX</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/juniper-ex/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/juniper-ex/</guid><description/></item><item><title>Juniper MX</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/juniper-mx/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/juniper-mx/</guid><description/></item><item><title>Juniper Networks Funk Software SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/juniper-networks-funk-software-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/juniper-networks-funk-software-snmp-traps/</guid><description/></item><item><title>Juniper Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/juniper-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/juniper-networks-inc-snmp-traps/</guid><description/></item><item><title>Juniper Networks Unisphere SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/juniper-networks-unisphere-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/juniper-networks-unisphere-snmp-traps/</guid><description/></item><item><title>Juniper Pulse Secure</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/juniper-pulse-secure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/juniper-pulse-secure/</guid><description/></item><item><title>Juniper QFX</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/juniper-qfx/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/juniper-qfx/</guid><description/></item><item><title>Juniper SRX</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/juniper-srx/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/juniper-srx/</guid><description/></item><item><title>K2Net SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/k2net-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/k2net-snmp-traps/</guid><description/></item><item><title>K8s Kubelet Monitoring</title><link>https://www.netdata.cloud/monitoring-101/kubelet-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/kubelet-monitoring/</guid><description>&lt;h2 id="what-is-k8s-kubelet">What is K8s Kubelet?&lt;/h2>
&lt;p>&lt;a href="https://kubernetes.io/docs/concepts/overview/components/#kubelet">&lt;code>Kubelet&lt;/code>&lt;/a> is an agent that runs on each node in the cluster. It makes sure that containers are running in a pod.&lt;/p>
&lt;h2 id="monitoring-k8s-kubelet-with-netdata">Monitoring K8s Kubelet with Netdata&lt;/h2>
&lt;p>The prerequisite for monitoring K8s Kubelet with Netdata is to &lt;a href="https://learn.netdata.cloud/docs/cloud/get-started/">Netdata installed&lt;/a> on your system.&lt;/p>
&lt;p>Netdata auto discovers hundreds of services, and for those it doesn&amp;rsquo;t turning on manual discovery is a one line configuration. For more information on configuring Netdata for K8s Kubelet monitoring please read the collector &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/k8s_kubelet/">documentation&lt;/a>.&lt;/p></description></item><item><title>Kafka</title><link>https://www.netdata.cloud/integrations/data-collection/databases/kafka/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/kafka/</guid><description/></item><item><title>Kafka</title><link>https://www.netdata.cloud/integrations/exporters/kafka/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/kafka/</guid><description/></item><item><title>Kafka __consumer_offsets growing huge: compaction failure on the offsets topic</title><link>https://www.netdata.cloud/guides/kafka/kafka-consumer-offsets-topic-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-consumer-offsets-topic-growing/</guid><description>&lt;h1 id="kafka-__consumer_offsets-growing-huge-compaction-failure-on-the-offsets-topic">Kafka __consumer_offsets growing huge: compaction failure on the offsets topic&lt;/h1>
&lt;p>One broker&amp;rsquo;s disk is climbing faster than its peers, or &lt;code>__consumer_offsets&lt;/code> has grown to multiple gigabytes while producer traffic is flat. This internal topic is compacted by default; its size should stay roughly proportional to active consumer groups and partitions. Unbounded growth means the log cleaner has stalled or crashed. The failure is silent: producers and consumers keep working, but every offset commit appends a record that compaction will never remove. Growth is not evenly distributed: &lt;code>__consumer_offsets&lt;/code> is partitioned by &lt;code>group.id&lt;/code> hash, so a stalled cleaner on one broker affects only the partitions in that broker&amp;rsquo;s &lt;code>log.dirs&lt;/code>. Eventually the log directory fills, and the broker may mark it offline.&lt;/p></description></item><item><title>Kafka 'Broker may not be available': clients that can't connect or stay connected</title><link>https://www.netdata.cloud/guides/kafka/kafka-broker-may-not-be-available/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-broker-may-not-be-available/</guid><description>&lt;h1 id="kafka-broker-may-not-be-available-clients-that-cant-connect-or-stay-connected">Kafka &amp;lsquo;Broker may not be available&amp;rsquo;: clients that can&amp;rsquo;t connect or stay connected&lt;/h1>
&lt;p>Client logs show a warning like this:&lt;/p>
&lt;pre tabindex="0">&lt;code>WARN [Producer clientId=...] Connection to node -1 could not be established. Broker may not be available.
&lt;/code>&lt;/pre>&lt;p>Python clients may raise &lt;code>kafka.errors.NoBrokersAvailable&lt;/code>, while librdkafka-based clients report &lt;code>Connection refused&lt;/code> against the broker address. The bootstrap server is often reachable; ping and telnet succeed, yet the client still fails. This happens because Kafka clients use the bootstrap connection only to fetch metadata. After that, they disconnect and try to open fresh TCP connections to the host:port pairs advertised in the metadata response. If those endpoints are unreachable, misconfigured, or secured differently than the client expects, the connection fails even though the bootstrap succeeded.&lt;/p></description></item><item><title>Kafka ActiveControllerCount not equal to 1: no controller or split brain</title><link>https://www.netdata.cloud/guides/kafka/kafka-no-active-controller/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-no-active-controller/</guid><description>&lt;h1 id="kafka-activecontrollercount-not-equal-to-1-no-controller-or-split-brain">Kafka ActiveControllerCount not equal to 1: no controller or split brain&lt;/h1>
&lt;p>When the cluster-wide sum of &lt;code>kafka.controller:type=KafkaController,name=ActiveControllerCount&lt;/code> is not 1, the cluster has either no active controller or multiple active controllers. A sum of 0 means no broker is steering metadata. A sum greater than 1 means split brain. Either way, this is a control-plane failure. Existing partition leaders usually keep serving produce and fetch requests, so the data plane may look healthy at first. Any broker failure, partition reassignment, or topic operation will stall because there is no controller to process it, or multiple controllers are processing conflicting operations.&lt;/p></description></item><item><title>Kafka authentication failures: SASL/mTLS errors, credential rotation, and brute force</title><link>https://www.netdata.cloud/guides/kafka/kafka-authentication-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-authentication-failures/</guid><description>&lt;h1 id="kafka-authentication-failures-saslmtls-errors-credential-rotation-and-brute-force">Kafka authentication failures: SASL/mTLS errors, credential rotation, and brute force&lt;/h1>
&lt;p>&lt;code>AuthenticationException&lt;/code> lines in broker logs are spiking and &lt;code>failed-authentication-total&lt;/code> is climbing on one or more listeners. Producers or consumers are failing to connect, and every failed attempt consumes a network thread and a file descriptor before the broker closes the connection. The same metric pattern can mean routine credential rotation, a client misconfiguration, or an active brute-force attempt. Telling the difference determines whether you need a config fix, a secret rotation, or a security incident response.&lt;/p></description></item><item><title>Kafka authorization failures: ACL denials, wrong-topic clients, and audit trails</title><link>https://www.netdata.cloud/guides/kafka/kafka-authorization-acl-denials/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-authorization-acl-denials/</guid><description>&lt;h1 id="kafka-authorization-failures-acl-denials-wrong-topic-clients-and-audit-trails">Kafka authorization failures: ACL denials, wrong-topic clients, and audit trails&lt;/h1>
&lt;p>Your consumers or producers throw &lt;code>TOPIC_AUTHORIZATION_FAILED&lt;/code> or &lt;code>CLUSTER_AUTHORIZATION_FAILED&lt;/code>. Your security team spots repeated log lines in the broker logs. You need to know within minutes whether this is a deployment misconfiguration, a client pointing at the wrong topic, or a security incident.&lt;/p>
&lt;p>In clusters where &lt;code>authorizer.class.name&lt;/code> is configured, Kafka writes authorization decisions to &lt;code>kafka-authorizer.log&lt;/code>. Denials log at INFO level; allowed operations are silent unless you enable DEBUG logging for the authorizer.&lt;/p></description></item><item><title>Kafka broker out of disk: log.dirs full, the cliff-edge shutdown, and recovery</title><link>https://www.netdata.cloud/guides/kafka/kafka-broker-out-of-disk/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-broker-out-of-disk/</guid><description>&lt;h1 id="kafka-broker-out-of-disk-logdirs-full-the-cliff-edge-shutdown-and-recovery">Kafka broker out of disk: log.dirs full, the cliff-edge shutdown, and recovery&lt;/h1>
&lt;p>When a Kafka broker exhausts disk space on a log directory, it does not throttle. It marks that directory offline or crashes entirely. Partitions on the failed directory become unavailable immediately. If all &lt;code>log.dirs&lt;/code> fail, the broker exits. You may see under-replicated partitions or producer timeouts seconds before failure, but often the first sign is a hard process exit with &lt;code>No space left on device&lt;/code>.&lt;/p></description></item><item><title>Kafka CommitFailedException: rebalanced-out consumers and poll loop timeouts</title><link>https://www.netdata.cloud/guides/kafka/kafka-commit-failed-exception/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-commit-failed-exception/</guid><description>&lt;h1 id="kafka-commitfailedexception-rebalanced-out-consumers-and-poll-loop-timeouts">Kafka CommitFailedException: rebalanced-out consumers and poll loop timeouts&lt;/h1>
&lt;p>&lt;code>CommitFailedException&lt;/code> with the message that the group has already rebalanced and assigned the partitions to another member means the time between &lt;code>poll()&lt;/code> calls exceeded &lt;code>max.poll.interval.ms&lt;/code>. The coordinator evicted the consumer and rejected the in-flight offset commit.&lt;/p>
&lt;p>When one consumer is evicted, the group rebalances. If other consumers are also slow, or if the rebalance itself takes long enough that healthy consumers miss the same deadline, the group enters a rebalance storm: it oscillates between &lt;code>JoinGroup&lt;/code> and &lt;code>SyncGroup&lt;/code> without stabilizing, and lag grows without bound.&lt;/p></description></item><item><title>Kafka connection storms: connection-count spikes, FD pressure, and network threads</title><link>https://www.netdata.cloud/guides/kafka/kafka-connection-count-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-connection-count-storm/</guid><description>&lt;h1 id="kafka-connection-storms-connection-count-spikes-fd-pressure-and-network-threads">Kafka connection storms: connection-count spikes, FD pressure, and network threads&lt;/h1>
&lt;p>Connection-count jumps on one or more brokers. An alert fires for open file descriptors, or consumers report timeouts while broker logs show nothing dramatic. In Kafka, a connection storm is insidious: every TCP connection costs one file descriptor and network-thread time, and the reactor architecture means network-thread saturation blocks everything, including metadata requests.&lt;/p>
&lt;p>Incremental fetch sessions keep consumer connections open even when idle. This is normal, but it means brokers run with a high steady-state connection count. The danger is the rate of change, not the absolute number. When connections spike above twice the established baseline, network threads drop below 30% idle, and FD usage approaches the process limit, the broker is in a connection storm.&lt;/p></description></item><item><title>Kafka consumer group lag growing: detection, lag-as-time, and root causes</title><link>https://www.netdata.cloud/guides/kafka/kafka-consumer-group-lag-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-consumer-group-lag-growing/</guid><description>&lt;h1 id="kafka-consumer-group-lag-growing-detection-lag-as-time-and-root-causes">Kafka consumer group lag growing: detection, lag-as-time, and root causes&lt;/h1>
&lt;p>When a consumer group falls behind, lag climbs monotonically. If the committed offset crosses the retention boundary, consumers hit &lt;code>OffsetOutOfRangeException&lt;/code> and must reset to earliest or latest, reprocessing or skipping data. Restarting the consumer is a common first reaction, but broker-side fetch latency, page cache eviction, and network saturation are equally common culprits. Detect lag accurately, convert it to time, and trace the root cause without guessing.&lt;/p></description></item><item><title>Kafka consumer group rebalancing too often: heartbeats, session timeout, and assignors</title><link>https://www.netdata.cloud/guides/kafka/kafka-consumer-group-rebalancing-frequently/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-consumer-group-rebalancing-frequently/</guid><description>&lt;h1 id="kafka-consumer-group-rebalancing-too-often-heartbeats-session-timeout-and-assignors">Kafka consumer group rebalancing too often: heartbeats, session timeout, and assignors&lt;/h1>
&lt;p>Consumer group lag is growing, application logs are full of &lt;code>JoinGroup&lt;/code> and &lt;code>SyncGroup&lt;/code> messages, and &lt;code>kafka-consumer-groups.sh&lt;/code> shows the group flipping between &lt;code>Stable&lt;/code> and &lt;code>PreparingRebalance&lt;/code>. Healthy groups rebalance only during membership changes and planned deployments. More than two or three rebalances per hour for a stable group signals instability. The usual cause is a mismatch between processing latency and one of three timeouts: &lt;code>session.timeout.ms&lt;/code>, &lt;code>heartbeat.interval.ms&lt;/code>, or &lt;code>max.poll.interval.ms&lt;/code>. The assignor strategy and static membership configuration determine how painful each rebalance is.&lt;/p></description></item><item><title>Kafka consumer group stuck Empty or Dead: no members consuming</title><link>https://www.netdata.cloud/guides/kafka/kafka-consumer-group-empty-stuck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-consumer-group-empty-stuck/</guid><description>&lt;h1 id="kafka-consumer-group-stuck-empty-or-dead-no-members-consuming">Kafka consumer group stuck Empty or Dead: no members consuming&lt;/h1>
&lt;p>A consumer group that should be active shows &lt;code>Empty&lt;/code> or &lt;code>Dead&lt;/code> state with no members. For a real-time pipeline, every new message piles up unread. For a batch job, &lt;code>Empty&lt;/code> between runs is expected. The difference between an intentional idle state and a silent outage is simple: is lag growing, and was the group supposed to be active?&lt;/p>
&lt;p>An &lt;code>Empty&lt;/code> group still exists in the coordinator. Committed offsets are retained until they expire, so consumers can resume on restart. A &lt;code>Dead&lt;/code> group has been removed, usually because offsets expired while no members were active. If the group is &lt;code>Dead&lt;/code> and data is still being produced, new consumers join a fresh group state and default to &lt;code>auto.offset.reset&lt;/code>, which can skip or reprocess data.&lt;/p></description></item><item><title>Kafka Consumer Lag</title><link>https://www.netdata.cloud/integrations/data-collection/databases/kafka-consumer-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/kafka-consumer-lag/</guid><description/></item><item><title>Kafka Consumer Lag Monitoring</title><link>https://www.netdata.cloud/monitoring-101/kafka_consumer_lag-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/kafka_consumer_lag-monitoring/</guid><description>&lt;h2 id="kafka-consumer-lag-monitoring">Kafka Consumer Lag Monitoring&lt;/h2>
&lt;h3 id="what-is-kafka-consumer-lag">What Is Kafka Consumer Lag?&lt;/h3>
&lt;p>Kafka Consumer Lag represents the difference between the latest offset of a partition and the offset of the consumer group. It is a crucial metric within Kafka&amp;rsquo;s architecture as it indicates the latency for messages consumed by consumers in a Kafka topic. Properly monitoring Kafka Consumer Lag ensures efficient message queue management and prevents bottlenecks in your streaming applications.&lt;/p>
&lt;h3 id="monitoring-kafka-consumer-lag-with-netdata">Monitoring Kafka Consumer Lag With Netdata&lt;/h3>
&lt;p>To monitor Kafka Consumer Lag, Netdata uses an openmetrics (Prometheus) exporter. Netdata&amp;rsquo;s $name monitoring tool allows you to ingest data from any Prometheus exporter seamlessly, offering automatic dashboards, alerts, and comprehensive insights without requiring a Prometheus server or Grafana. By leveraging Netdata&amp;rsquo;s dashboards, DevOps and IT teams can effectively track Kafka performance, ensuring real-time data processing stays smooth and efficient.&lt;/p></description></item><item><title>Kafka consumer rebalance storm: stuck in PreparingRebalance and max.poll.interval.ms</title><link>https://www.netdata.cloud/guides/kafka/kafka-consumer-rebalance-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-consumer-rebalance-storm/</guid><description>&lt;h1 id="kafka-consumer-rebalance-storm-stuck-in-preparingrebalance-and-maxpollintervalms">Kafka consumer rebalance storm: stuck in PreparingRebalance and max.poll.interval.ms&lt;/h1>
&lt;p>Consumer group lag climbs while brokers report zero under-replicated partitions, normal produce and fetch latency, and normal request handler idle percent. The consumer group oscillates between Stable, PreparingRebalance, and CompletingRebalance without settling long enough to make progress. Every rebalance cycle pauses consumption; lag grows because time spent rebalancing dwarfs time spent fetching. This is a consumer rebalance storm. It is almost always a client-side timeout or processing issue. Look for CommitFailedException or max.poll.interval.ms exceeded in consumer logs.&lt;/p></description></item><item><title>Kafka controller event queue backing up: overwhelmed controller and stalled metadata</title><link>https://www.netdata.cloud/guides/kafka/kafka-controller-event-queue-backup/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-controller-event-queue-backup/</guid><description>&lt;h1 id="kafka-controller-event-queue-backing-up-overwhelmed-controller-and-stalled-metadata">Kafka controller event queue backing up: overwhelmed controller and stalled metadata&lt;/h1>
&lt;p>You see &lt;code>NOT_LEADER_FOR_PARTITION&lt;/code> errors spike in client logs. Leader elections stop completing. The active controller&amp;rsquo;s event queue grows and does not drain. Partitions that need new leaders stay offline. The single thread processing partition state changes cannot keep up, and metadata operations stall.&lt;/p>
&lt;p>In ZooKeeper mode, each event writes to ZooKeeper. In KRaft mode, each event appends to the Raft metadata log. The active controller processes them sequentially. When the queue backs up, the control plane slows. The data plane may continue serving existing leaders, but any failure requiring a state change gets stuck.&lt;/p></description></item><item><title>Kafka disk I/O latency high: await, LocalTimeMs, and the slow-disk broker</title><link>https://www.netdata.cloud/guides/kafka/kafka-disk-io-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-disk-io-latency-high/</guid><description>&lt;h1 id="kafka-disk-io-latency-high-await-localtimems-and-the-slow-disk-broker">Kafka disk I/O latency high: await, LocalTimeMs, and the slow-disk broker&lt;/h1>
&lt;p>&lt;code>iostat&lt;/code> shows &lt;code>await&lt;/code> climbing, maybe a disk alert fired. But &lt;code>await&lt;/code> alone is not a pageable event. On SSDs and RAID arrays, &lt;code>%util&lt;/code> hits 100% under modest load because it measures device busy time, not saturation. What matters is &lt;code>await&lt;/code>, the average time for I/O requests to be served. At the Kafka layer, the mirror image is &lt;code>LocalTimeMs&lt;/code> in the request latency breakdown. Both spike during normal operations &amp;ndash; broker restart with cold page cache, log compaction, and partition reassignment all drive up disk latency without indicating hardware fault. This guide shows how to distinguish a transient spike from a slow disk that will shrink your ISR and block produce requests.&lt;/p></description></item><item><title>Kafka disk space planning: retention, replication, and runway estimation</title><link>https://www.netdata.cloud/guides/kafka/kafka-disk-space-runway-planning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-disk-space-runway-planning/</guid><description>&lt;h1 id="kafka-disk-space-planning-retention-replication-and-runway-estimation">Kafka disk space planning: retention, replication, and runway estimation&lt;/h1>
&lt;p>When a Kafka broker exhausts a &lt;code>log.dirs&lt;/code> volume, the directory goes offline. If it is the only volume, the broker halts. Partitions on that volume become unavailable and the broker cannot rejoin until space is freed. Disk planning is a continuous operational calculation, not a one-time purchase. You need a working estimate of steady-state usage, a runway model against an operational full threshold, and an explicit map of the retention rules that actually delete data.&lt;/p></description></item><item><title>Kafka enable.auto.commit data loss: committed offsets that outrun processing</title><link>https://www.netdata.cloud/guides/kafka/kafka-auto-commit-silent-data-loss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-auto-commit-silent-data-loss/</guid><description>&lt;h1 id="kafka-enableautocommit-data-loss-committed-offsets-that-outrun-processing">Kafka enable.auto.commit data loss: committed offsets that outrun processing&lt;/h1>
&lt;p>Downstream systems are missing events, but your Kafka consumer group reports lag near zero and the group state is Stable. There are no broker errors, no rebalances, and no visible backpressure. The application crashed or deployed minutes ago, then caught up instantly on recovery. Messages are lost silently. This is the signature of &lt;code>enable.auto.commit=true&lt;/code>: offsets are committed to &lt;code>__consumer_offsets&lt;/code> on a fixed schedule, not after processing completes.&lt;/p></description></item><item><title>Kafka FailedProduceRequestsPerSec rising: the single best 'producers are hurting' signal</title><link>https://www.netdata.cloud/guides/kafka/kafka-failed-produce-requests/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-failed-produce-requests/</guid><description>&lt;h1 id="kafka-failedproducerequestspersec-rising-the-single-best-producers-are-hurting-signal">Kafka FailedProduceRequestsPerSec rising: the single best &amp;lsquo;producers are hurting&amp;rsquo; signal&lt;/h1>
&lt;p>&lt;code>FailedProduceRequestsPerSec&lt;/code> rising is broker-side confirmation that producers are actively being rejected. Unlike client-side timeouts from network blips or producer memory pressure, it only increments when the broker receives a produce request and fails to complete it. The MBean &lt;code>kafka.server:type=BrokerTopicMetrics,name=FailedProduceRequestsPerSec&lt;/code> rolls up server-side produce failures into a single rate: &lt;code>NotEnoughReplicasException&lt;/code>, &lt;code>NotLeaderOrFollowerException&lt;/code>, &lt;code>CorruptRecordException&lt;/code>, and others. A sustained nonzero rate means data is not being accepted, and the window for data loss or unavailability is open.&lt;/p></description></item><item><title>Kafka fetch request latency high: FetchConsumer vs FetchFollower and page cache misses</title><link>https://www.netdata.cloud/guides/kafka/kafka-fetch-request-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-fetch-request-latency-high/</guid><description>&lt;h1 id="kafka-fetch-request-latency-high-fetchconsumer-vs-fetchfollower-and-page-cache-misses">Kafka fetch request latency high: FetchConsumer vs FetchFollower and page cache misses&lt;/h1>
&lt;p>Your tail consumers are lagging, or your under-replicated partition count is climbing. On the broker, &lt;code>kafka.network:type=RequestMetrics,name=TotalTimeMs,request=FetchConsumer&lt;/code> or &lt;code>FetchFollower&lt;/code> is elevated. Raw fetch latency is a poor signal: consumer long-polling means &lt;code>TotalTimeMs&lt;/code> routinely includes the full &lt;code>fetch.max.wait.ms&lt;/code> wait even on an idle topic, and follower fetches are paced by the leader&amp;rsquo;s ability to serve segments. The actionable metric is &lt;code>LocalTimeMs&lt;/code>, the time the leader spends reading the log. When &lt;code>FetchConsumer&lt;/code> &lt;code>LocalTimeMs&lt;/code> spikes, the data was not in the OS page cache. When &lt;code>FetchFollower&lt;/code> &lt;code>LocalTimeMs&lt;/code> spikes, the leader is slow to serve replication reads and ISR shrinks will follow. The sections below show how to tell the two apart, confirm page cache misses, and fix the root cause.&lt;/p></description></item><item><title>Kafka ISR shrinking: IsrShrinksPerSec, flapping, and the cascade to offline</title><link>https://www.netdata.cloud/guides/kafka/kafka-isr-shrink-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-isr-shrink-storm/</guid><description>&lt;h1 id="kafka-isr-shrinking-isrshrinkspersec-flapping-and-the-cascade-to-offline">Kafka ISR shrinking: IsrShrinksPerSec, flapping, and the cascade to offline&lt;/h1>
&lt;p>&lt;code>IsrShrinksPerSec&lt;/code> is climbing on your leaders and &lt;code>UnderReplicatedPartitions&lt;/code> is no longer zero. If it is flapping &amp;ndash; shrinks followed by expands every few minutes &amp;ndash; the path ends with &lt;code>OfflinePartitionsCount&lt;/code> rising and &lt;code>acks=all&lt;/code> producers throwing &lt;code>NotEnoughReplicasException&lt;/code>. This guide covers that path: how a lagging follower becomes a cluster-wide problem, how to separate flapping from one-way degradation, and how to stop the cascade before partitions go offline.&lt;/p></description></item><item><title>Kafka JVM heap and Full GC pauses: ISR drops, session timeouts, and right-sizing the heap</title><link>https://www.netdata.cloud/guides/kafka/kafka-jvm-heap-full-gc-pauses/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-jvm-heap-full-gc-pauses/</guid><description>&lt;h1 id="kafka-jvm-heap-and-full-gc-pauses-isr-drops-session-timeouts-and-right-sizing-the-heap">Kafka JVM heap and Full GC pauses: ISR drops, session timeouts, and right-sizing the heap&lt;/h1>
&lt;p>Sporadic &lt;code>UnderReplicatedPartitions&lt;/code> and ISR shrinks that do not correlate with disk I/O or network faults, combined with consumer rebalances and &lt;code>NotEnoughReplicasException&lt;/code> from producers using &lt;code>acks=all&lt;/code>, point to broker JVM heap pressure. Check broker logs for GC pauses in the Old Generation lasting several seconds.&lt;/p>
&lt;p>Brokers use the JVM heap for metadata, request buffers, and message format conversion. They do not store messages on the heap; the OS page cache handles that. When the heap is misconfigured or under pressure, garbage collection pauses can freeze a broker long enough to trigger ZooKeeper session expirations, follower lag, and cascading availability issues. Full GC pauses exceeding five seconds are the common threshold where these symptoms begin.&lt;/p></description></item><item><title>Kafka KRaft metadata log lag: standby controllers and brokers falling behind</title><link>https://www.netdata.cloud/guides/kafka/kafka-kraft-metadata-log-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-kraft-metadata-log-lag/</guid><description>&lt;h1 id="kafka-kraft-metadata-log-lag-standby-controllers-and-brokers-falling-behind">Kafka KRaft metadata log lag: standby controllers and brokers falling behind&lt;/h1>
&lt;p>When client errors like &lt;code>NOT_LEADER_FOR_PARTITION&lt;/code> or &lt;code>LeaderNotAvailableException&lt;/code> hit a stable-looking cluster, check whether standby controllers and broker metadata lag are climbing. In KRaft mode, the active controller is the Raft quorum leader. It appends metadata changes to &lt;code>__cluster_metadata&lt;/code>. Standby controllers and brokers apply that log asynchronously. When they fall behind, they act on stale partition leadership, outdated ISR memberships, and old broker registrations. This guide shows how to distinguish a slow follower, a sick leader, and a network partition.&lt;/p></description></item><item><title>Kafka KRaft quorum has no leader: current-leader = -1 and frozen metadata</title><link>https://www.netdata.cloud/guides/kafka/kafka-kraft-quorum-no-leader/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-kraft-quorum-no-leader/</guid><description>&lt;h1 id="kafka-kraft-quorum-has-no-leader-current-leader---1-and-frozen-metadata">Kafka KRaft quorum has no leader: current-leader = -1 and frozen metadata&lt;/h1>
&lt;p>Topic creation hangs. Partition reassignments stall. Broker logs show metadata operations timing out. On controller nodes, JMX reports &lt;code>kafka.server:type=raft-metrics,attribute=current-leader&lt;/code> with value &lt;code>-1&lt;/code>, and the quorum state is frozen. Existing producers and consumers continue to read and write, but the control plane is stuck.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>The quorum leader acts as the active controller. When &lt;code>current-leader&lt;/code> is &lt;code>-1&lt;/code>, the local controller has not discovered a leader, which means no node in the quorum can commit new entries to the metadata log. Topic creation, deletion, configuration updates, reassignments, and ISR changes are blocked. Leader elections for partitions that lose their broker also cannot proceed.&lt;/p></description></item><item><title>Kafka LEADER_NOT_AVAILABLE: causes during elections, restarts, and topic creation</title><link>https://www.netdata.cloud/guides/kafka/kafka-leader-not-available/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-leader-not-available/</guid><description>&lt;h1 id="kafka-leader_not_available-causes-during-elections-restarts-and-topic-creation">Kafka LEADER_NOT_AVAILABLE: causes during elections, restarts, and topic creation&lt;/h1>
&lt;p>&lt;code>LEADER_NOT_AVAILABLE&lt;/code> means a client asked a broker to produce or fetch from a partition that has no assigned leader. In healthy clusters this is brief during rolling restarts, controller elections, or topic creation. Persistent errors correlate with &lt;code>OfflinePartitionsCount &amp;gt; 0&lt;/code> and indicate the data plane is broken for those partitions. Distinguish this from &lt;code>NOT_LEADER_FOR_PARTITION&lt;/code>, which means a leader exists but the client contacted the wrong broker and needs a metadata refresh.&lt;/p></description></item><item><title>Kafka LeaderElectionRateAndTimeMs spiking: election storms and slow elections</title><link>https://www.netdata.cloud/guides/kafka/kafka-leader-election-rate-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-leader-election-rate-high/</guid><description>&lt;h1 id="kafka-leaderelectionrateandtimems-spiking-election-storms-and-slow-elections">Kafka LeaderElectionRateAndTimeMs spiking: election storms and slow elections&lt;/h1>
&lt;p>&lt;code>LeaderElectionRateAndTimeMs&lt;/code> climbing outside maintenance windows means the controller is struggling to keep the partition map consistent. Producers and consumers may see &lt;code>NOT_LEADER_FOR_PARTITION&lt;/code> or request timeouts. Partitions can hang under-replicated or offline while the controller works through a backlog.&lt;/p>
&lt;p>This metric has two dimensions. The rate is how often leadership changes. The time is how long the metadata store takes to commit each change. High rate with low time means broker flapping or repeated administrative actions. Normal rate with high time means the controller event queue is backed up, ZooKeeper writes are slow, or the KRaft quorum is lagging. Determine which mode you are in before the cluster degrades further.&lt;/p></description></item><item><title>Kafka leadership imbalance: LeaderCount skew and preferred replica election</title><link>https://www.netdata.cloud/guides/kafka/kafka-leadership-imbalance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-leadership-imbalance/</guid><description>&lt;h1 id="kafka-leadership-imbalance-leadercount-skew-and-preferred-replica-election">Kafka leadership imbalance: LeaderCount skew and preferred replica election&lt;/h1>
&lt;p>One broker in your cluster is handling 40 percent of all produce requests. Its &lt;code>RequestHandlerAvgIdlePercent&lt;/code> is dropping toward 0.3 while the rest of the cluster is nearly idle. Cluster-level throughput dashboards look healthy because aggregate ingress is within capacity, but tail latency on the hot broker is climbing and consumers connected to it are falling behind. The root cause is leadership skew: a disproportionate number of partition leaders have landed on one broker, and every leader carries the full request load for its partitions. This happens silently after rolling restarts, controlled shutdowns, or broker decommissions. Even a perfectly even &lt;code>PartitionCount&lt;/code> across brokers can hide severe leadership imbalance. You will not see it without monitoring &lt;code>LeaderCount&lt;/code> per broker.&lt;/p></description></item><item><title>Kafka log compaction falling behind: the dead cleaner thread and unbounded disk growth</title><link>https://www.netdata.cloud/guides/kafka/kafka-log-compaction-falling-behind/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-log-compaction-falling-behind/</guid><description>&lt;h1 id="kafka-log-compaction-falling-behind-the-dead-cleaner-thread-and-unbounded-disk-growth">Kafka log compaction falling behind: the dead cleaner thread and unbounded disk growth&lt;/h1>
&lt;p>Disk utilization climbs steadily on one or more brokers while producer traffic stays flat. &lt;code>__consumer_offsets&lt;/code> balloons from a few gigabytes to hundreds. Under-replicated partitions are at zero, producers show no errors, and standard Kafka alerts are silent. The usual culprit is a dead log cleaner thread: it hit a corrupt record or an OOM during compaction, crashed, and never restarted. Every compacted topic now accumulates records without bound.&lt;/p></description></item><item><title>Kafka Log directory failed / OfflineLogDirectoryCount > 0: disk errors and JBOD recovery</title><link>https://www.netdata.cloud/guides/kafka/kafka-log-directory-failed-offline/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-log-directory-failed-offline/</guid><description>&lt;h1 id="kafka-log-directory-failed--offlinelogdirectorycount--0-disk-errors-and-jbod-recovery">Kafka Log directory failed / OfflineLogDirectoryCount &amp;gt; 0: disk errors and JBOD recovery&lt;/h1>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>When Kafka catches an IOException on a &lt;code>log.dirs&lt;/code> path, it marks that log directory offline. The broker increments &lt;code>kafka.log:type=LogManager,name=OfflineLogDirectoryCount&lt;/code> and logs the failure. Partitions with replicas on the failed directory lose those replicas. If the partition leader was on that directory and &lt;code>unclean.leader.election.enable=false&lt;/code>, the partition becomes unavailable until the controller elects a new leader from the remaining ISR. Producers with &lt;code>acks=all&lt;/code> see &lt;code>NotEnoughReplicasException&lt;/code> when the surviving ISR drops below &lt;code>min.insync.replicas&lt;/code>.&lt;/p></description></item><item><title>Kafka LogFlushRateAndTimeMs high: fsync latency and a failing disk</title><link>https://www.netdata.cloud/guides/kafka/kafka-log-flush-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-log-flush-latency-high/</guid><description>&lt;h1 id="kafka-logflushrateandtimems-high-fsync-latency-and-a-failing-disk">Kafka LogFlushRateAndTimeMs high: fsync latency and a failing disk&lt;/h1>
&lt;p>When &lt;code>kafka.log:type=LogFlushStats,name=LogFlushRateAndTimeMs&lt;/code> p99 stays above 200 ms, the broker&amp;rsquo;s fsync path is slow. On SSD-backed clusters, p99 should stay below 50 ms; sustained values above 2 s point to disk degradation. Because most deployments leave &lt;code>log.flush.interval.messages&lt;/code> and &lt;code>log.flush.interval.ms&lt;/code> unset and rely on replication plus OS lazy flush, this metric reflects kernel-driven flushes or explicit fsyncs. A slow flush raises produce &lt;code>LocalTimeMs&lt;/code>, blocks request handler threads, and can push followers out of ISR. This guide shows how to tell whether the cause is a failing disk, a bad flush policy, or transient I/O contention.&lt;/p></description></item><item><title>Kafka min.insync.replicas and acks: configuring durability you actually have</title><link>https://www.netdata.cloud/guides/kafka/kafka-min-insync-replicas-misconfigured/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-min-insync-replicas-misconfigured/</guid><description>&lt;h1 id="kafka-mininsyncreplicas-and-acks-configuring-durability-you-actually-have">Kafka min.insync.replicas and acks: configuring durability you actually have&lt;/h1>
&lt;p>Most operators set producers to &lt;code>acks=all&lt;/code> and assume the cluster acks only when every replica has the message. It does not. With &lt;code>acks=all&lt;/code>, the broker waits only for the current in-sync replica set (ISR). Because the ISR shrinks dynamically when followers lag, a partition with replication factor three can have an ISR of one &amp;ndash; the leader itself. Without raising &lt;code>min.insync.replicas&lt;/code> from its default, the leader acks with zero followers caught up. Your durability guarantee collapses to leader-only persistence, and you only find out when the leader dies and data is missing.&lt;/p></description></item><item><title>Kafka Monitoring</title><link>https://www.netdata.cloud/monitoring-101/kafka-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/kafka-monitoring/</guid><description>&lt;h2 id="kafka-monitoring">Kafka Monitoring&lt;/h2>
&lt;h3 id="what-is-kafka">What Is Kafka?&lt;/h3>
&lt;p>Apache Kafka is a distributed event streaming platform capable of handling trillions of events a day. It is used by thousands of companies for streaming analytics, data integration, and data pipelines. As a &lt;strong>message broker&lt;/strong>, Kafka is crucial for managing big data and ensuring real-time data streaming operations. Monitoring Kafka is vital to ensure that it operates efficiently and to prevent costly downtime or data loss.&lt;/p></description></item><item><title>Kafka monitoring checklist: the signals every production cluster needs</title><link>https://www.netdata.cloud/guides/kafka/kafka-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-monitoring-checklist/</guid><description>&lt;h1 id="kafka-monitoring-checklist-the-signals-every-production-cluster-needs">Kafka monitoring checklist: the signals every production cluster needs&lt;/h1>
&lt;p>Kafka failures follow predictable paths: an ISR shrinks, a controller queue backs up, a disk fills while the cleaner thread hangs, or a consumer rebalance storm hides behind healthy broker metrics. You need to know which signals matter and when they justify a 3 AM page.&lt;/p>
&lt;p>This checklist organizes broker-side signals into four levels. Each builds on the last: Level 1 prevents data loss. Level 2 prevents surprises. Level 3 exposes leading indicators. Level 4 catches silent killers. Use it to audit dashboards, tune alert severity, and justify instrumentation.&lt;/p></description></item><item><title>Kafka monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/kafka/kafka-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-monitoring-maturity-model/</guid><description>&lt;h1 id="kafka-monitoring-maturity-model-from-survival-to-expert">Kafka monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Broker logs and a single health check are not enough for production Kafka. Monitoring everything at once creates noise and hides the signals that prevent outages. This model gives you a prioritized path: start with telemetry that prevents total data loss, then add layers that catch degradation before it becomes an incident.&lt;/p>
&lt;p>Use this as an onboarding checklist for new clusters, an incident reference, and a roadmap when you have instrumentation budget. Signals come from Kafka JMX, OS metrics, and admin APIs. In KRaft mode (mandatory in Kafka 4.0+), substitute ZooKeeper-specific signals with the corresponding Raft quorum metrics.&lt;/p></description></item><item><title>Kafka network egress saturation: BytesOutPerSec, replication amplification, and fan-out</title><link>https://www.netdata.cloud/guides/kafka/kafka-bytes-out-network-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-bytes-out-network-saturation/</guid><description>&lt;h1 id="kafka-network-egress-saturation-bytesoutpersec-replication-amplification-and-fan-out">Kafka network egress saturation: BytesOutPerSec, replication amplification, and fan-out&lt;/h1>
&lt;p>Kafka is usually sized for producer ingress because that is what the business reports. In practice, the first resource to saturate is almost always network egress. Every byte written to a partition leader is read at least once by each follower replica and once by every active consumer group keeping up with the log. A broker can hit its NIC ceiling while producer throughput looks comfortable and disk I/O is barely warmed up.&lt;/p></description></item><item><title>Kafka NetworkProcessorAvgIdlePercent low: network thread saturation and TLS overhead</title><link>https://www.netdata.cloud/guides/kafka/kafka-network-processor-idle-percent-low/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-network-processor-idle-percent-low/</guid><description>&lt;h1 id="kafka-networkprocessoravgidlepercent-low-network-thread-saturation-and-tls-overhead">Kafka NetworkProcessorAvgIdlePercent low: network thread saturation and TLS overhead&lt;/h1>
&lt;p>NetworkProcessorAvgIdlePercent drops on one or more brokers. A ticket may fire when the value falls below 0.3, or producers and consumers may timeout while broker CPU looks fine and I/O threads are not obviously saturated. Network thread saturation blocks all socket activity, including metadata requests, so every client suffers. With the default num.network.threads set to 3, brokers running TLS and dense client pools are especially vulnerable. Sustained values below 0.1 make the broker effectively unreachable even though the process is still running.&lt;/p></description></item><item><title>Kafka NOT_LEADER_FOR_PARTITION: stale metadata, controller lag, and client retries</title><link>https://www.netdata.cloud/guides/kafka/kafka-not-leader-for-partition/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-not-leader-for-partition/</guid><description>&lt;h1 id="kafka-not_leader_for_partition-stale-metadata-controller-lag-and-client-retries">Kafka NOT_LEADER_FOR_PARTITION: stale metadata, controller lag, and client retries&lt;/h1>
&lt;p>Producers and consumers log &lt;code>NOT_LEADER_FOR_PARTITION&lt;/code>. Broker response metrics show spikes in failed produce or fetch requests. The cluster usually self-heals within seconds as clients refresh metadata. When the error persists for minutes, or flaps across many partitions, the root cause is typically a controller that cannot keep up with leadership changes. Distinguishing a routine leader election from a controller queue backup that blocks metadata propagation is the first step.&lt;/p></description></item><item><title>Kafka NotEnoughReplicasException: acks=all writes rejected below min.insync.replicas</title><link>https://www.netdata.cloud/guides/kafka/kafka-not-enough-replicas-exception/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-not-enough-replicas-exception/</guid><description>&lt;h1 id="kafka-notenoughreplicasexception-acksall-writes-rejected-below-mininsyncreplicas">Kafka NotEnoughReplicasException: acks=all writes rejected below min.insync.replicas&lt;/h1>
&lt;p>Producers are throwing &lt;code>org.apache.kafka.common.errors.NotEnoughReplicasException&lt;/code> or &lt;code>NotEnoughReplicasAfterAppendException&lt;/code>, and &lt;code>acks=all&lt;/code> writes are failing while &lt;code>acks=1&lt;/code> or &lt;code>acks=0&lt;/code> writes may still succeed. The affected partitions no longer have enough in-sync replicas to satisfy &lt;code>min.insync.replicas&lt;/code>. The immediate operational question is whether the ISR shrink is a transient recovery blip or a sustained degradation that will block writes until you fix the follower.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>The leader tracks followers caught up within &lt;code>replica.lag.time.max.ms&lt;/code> &lt;!-- TODO: verify default value and version --> in the In-Sync Replica set (ISR). For &lt;code>acks=all&lt;/code>, the leader waits for all current ISR members before acknowledging the producer. If the ISR size drops below &lt;code>min.insync.replicas&lt;/code>, the leader rejects the produce request. The broker-level default for &lt;code>min.insync.replicas&lt;/code> is 1, so a lone leader can acknowledge alone. In practice, with &lt;code>replication.factor=3&lt;/code> and &lt;code>acks=all&lt;/code>, set &lt;code>min.insync.replicas=2&lt;/code> so a single follower loss blocks writes instead of silently weakening durability.&lt;/p></description></item><item><title>Kafka OfflinePartitionsCount > 0: partitions with no leader and how to recover</title><link>https://www.netdata.cloud/guides/kafka/kafka-offline-partitions-count/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-offline-partitions-count/</guid><description>&lt;h1 id="kafka-offlinepartitionscount--0-partitions-with-no-leader-and-how-to-recover">Kafka OfflinePartitionsCount &amp;gt; 0: partitions with no leader and how to recover&lt;/h1>
&lt;p>When &lt;code>kafka.controller:type=KafkaController,name=OfflinePartitionsCount&lt;/code> is nonzero, at least one partition has no active leader. Those partitions are completely unavailable: producers receive errors, consumers stall, and no data is written or read until a leader is elected. This is a data-plane outage.&lt;/p>
&lt;p>This metric is only meaningful on the active controller. Non-controller brokers always report zero. If the controller itself is down, the metric may be stale or unreachable at the exact moment you need it. Brief spikes can occur during controller re-election or ungraceful broker shutdown, but any sustained nonzero value past 60 seconds is an active incident that requires immediate intervention.&lt;/p></description></item><item><title>Kafka OffsetOutOfRangeException: when retention deletes data before the consumer reads it</title><link>https://www.netdata.cloud/guides/kafka/kafka-offset-out-of-range-exception/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-offset-out-of-range-exception/</guid><description>&lt;h1 id="kafka-offsetoutofrangeexception-when-retention-deletes-data-before-the-consumer-reads-it">Kafka OffsetOutOfRangeException: when retention deletes data before the consumer reads it&lt;/h1>
&lt;p>&lt;code>OffsetOutOfRangeException&lt;/code> means the consumer requested an offset the broker no longer holds. The log segment containing the consumer&amp;rsquo;s committed position was deleted by retention before the consumer caught up. This is not a transient fetch error; it is data loss, and the outcome depends entirely on &lt;code>auto.offset.reset&lt;/code>. Many clients default to &lt;code>latest&lt;/code>, which turns this exception into silent skipping.&lt;/p></description></item><item><title>Kafka page cache thrashing: the backfill consumer that 100x's tail latency</title><link>https://www.netdata.cloud/guides/kafka/kafka-page-cache-thrashing-latency-cliff/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-page-cache-thrashing-latency-cliff/</guid><description>&lt;h1 id="kafka-page-cache-thrashing-the-backfill-consumer-that-100xs-tail-latency">Kafka page cache thrashing: the backfill consumer that 100x&amp;rsquo;s tail latency&lt;/h1>
&lt;p>Tail latency jumps from milliseconds to seconds while producer throughput and replication stay flat. CPU is normal. Every fetch request from tail consumers starts hitting disk. The culprit is usually a single backfill consumer reading historical data, evicting the hot working set from the OS page cache and turning a memory-speed system into a disk-bound one.&lt;/p>
&lt;p>The write path stays green. Under-replicated partitions do not increase. No broker has crashed. The only visible signs are elevated disk read latency and slow consumers. If the cluster slowed suddenly with no broker fault, suspect page cache thrashing from a backfill consumer.&lt;/p></description></item><item><title>Kafka produce request latency high: reading the TotalTimeMs breakdown</title><link>https://www.netdata.cloud/guides/kafka/kafka-produce-request-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-produce-request-latency-high/</guid><description>&lt;h1 id="kafka-produce-request-latency-high-reading-the-totaltimems-breakdown">Kafka produce request latency high: reading the TotalTimeMs breakdown&lt;/h1>
&lt;p>Producers are timing out or retrying. Client-side &lt;code>request-latency-avg&lt;/code> is elevated and broker &lt;code>TotalTimeMs&lt;/code> p99 is spiking. The total is unactionable by itself. Kafka breaks it into five sub-components, each implicating a different subsystem. You need all five via JMX or an equivalent metrics collector to know whether to fix disk, scale a thread pool, or replace a follower.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>&lt;code>TotalTimeMs&lt;/code> is the wall-clock time from when the broker&amp;rsquo;s network thread receives a produce request until the response is fully sent. It is the arithmetic sum of five phases:&lt;/p></description></item><item><title>Kafka producer timeout cascade: when retries pile load onto a slow broker</title><link>https://www.netdata.cloud/guides/kafka/kafka-produce-timeout-retry-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-produce-timeout-retry-cascade/</guid><description>&lt;h1 id="kafka-producer-timeout-cascade-when-retries-pile-load-onto-a-slow-broker">Kafka producer timeout cascade: when retries pile load onto a slow broker&lt;/h1>
&lt;p>Your producers are timing out. P99 produce latency is climbing. You see more requests hitting the brokers, yet actual throughput of new messages is flat or falling. This is not a traffic spike. It is a producer timeout cascade: one slow broker causes clients to retry, and those retries add load to the same overloaded broker, closing the loop until the cluster is pinned.&lt;/p></description></item><item><title>Kafka purgatory size growing: delayed produce and fetch operations</title><link>https://www.netdata.cloud/guides/kafka/kafka-purgatory-size-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-purgatory-size-growing/</guid><description>&lt;h1 id="kafka-purgatory-size-growing-delayed-produce-and-fetch-operations">Kafka purgatory size growing: delayed produce and fetch operations&lt;/h1>
&lt;p>JMX shows &lt;code>kafka.server:type=DelayedOperationPurgatory,name=PurgatorySize&lt;/code> climbing on one or more brokers. Produce purgatory holds &lt;code>acks=all&lt;/code> requests waiting for ISR completion. Fetch purgatory holds consumer and follower requests waiting for &lt;code>fetch.min.bytes&lt;/code>. A growing queue means requests spend more time inside the broker than clients expected. Producers time out and retry. Consumers sit idle. The cluster is not dead, but it is backing up at a precise choke point. This guide distinguishes normal long-polling from replication crisis.&lt;/p></description></item><item><title>Kafka quota throttling: throttle-time, runaway clients, and protecting the cluster</title><link>https://www.netdata.cloud/guides/kafka/kafka-quota-throttling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-quota-throttling/</guid><description>&lt;h1 id="kafka-quota-throttling-throttle-time-runaway-clients-and-protecting-the-cluster">Kafka quota throttling: throttle-time, runaway clients, and protecting the cluster&lt;/h1>
&lt;p>Your Kafka producers are timing out. Consumer lag is growing. Client dashboards show elevated request latency, but the brokers are healthy: UnderReplicatedPartitions is zero and disk I/O looks fine. The culprit is often quota throttling: the cluster is enforcing per-client or per-user byte-rate limits, and one or more clients have hit the wall. The broker exposes this as throttle-time in JMX. That backpressure slows the client, but if the client is a runaway service or a backfill consumer, the throttling can cascade into widespread latency.&lt;/p></description></item><item><title>Kafka RecordTooLargeException / MESSAGE_TOO_LARGE: message size limits across the path</title><link>https://www.netdata.cloud/guides/kafka/kafka-message-too-large-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-message-too-large-error/</guid><description>&lt;h1 id="kafka-recordtoolargeexception--message_too_large-message-size-limits-across-the-path">Kafka RecordTooLargeException / MESSAGE_TOO_LARGE: message size limits across the path&lt;/h1>
&lt;p>Two errors surface: the producer client throws &lt;code>RecordTooLargeException&lt;/code> before the request hits the wire, or the broker returns &lt;code>MESSAGE_TOO_LARGE&lt;/code> in the produce response. Both mean a record batch exceeds a limit somewhere in the path, but the fix depends on exactly which limit and where the size is measured. These settings are not a single knob: &lt;code>max.request.size&lt;/code>, &lt;code>message.max.bytes&lt;/code>, &lt;code>max.message.bytes&lt;/code>, &lt;code>replica.fetch.max.bytes&lt;/code>, &lt;code>fetch.max.bytes&lt;/code>, &lt;code>max.partition.fetch.bytes&lt;/code>, and &lt;code>socket.request.max.bytes&lt;/code> must all align. Raise one limit without raising the others and you will wedge replication, strand consumers, or silently lose data on the next large message.&lt;/p></description></item><item><title>Kafka replica MaxLag growing: slow followers and replica fetcher health</title><link>https://www.netdata.cloud/guides/kafka/kafka-replica-fetcher-max-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-replica-fetcher-max-lag/</guid><description>&lt;h1 id="kafka-replica-maxlag-growing-slow-followers-and-replica-fetcher-health">Kafka replica MaxLag growing: slow followers and replica fetcher health&lt;/h1>
&lt;p>When &lt;code>kafka.server:type=ReplicaFetcherManager,name=MaxLag,clientId=Replica&lt;/code> climbs on a broker, the worst follower is failing to replicate fast enough. This metric is the maximum offset distance between a leader and its most lagging follower. In a healthy cluster it stays near zero. If the gap persists longer than &lt;code>replica.lag.time.max.ms&lt;/code>, the leader removes the follower from the ISR. Once enough replicas drop, partitions can fall below &lt;code>min.insync.replicas&lt;/code>, and producers using &lt;code>acks=all&lt;/code> hit &lt;code>NotEnoughReplicasException&lt;/code>.&lt;/p></description></item><item><title>Kafka request queue filling up: RequestQueueSize, queued.max.requests, and backpressure</title><link>https://www.netdata.cloud/guides/kafka/kafka-request-queue-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-request-queue-full/</guid><description>&lt;h1 id="kafka-request-queue-filling-up-requestqueuesize-queuedmaxrequests-and-backpressure">Kafka request queue filling up: RequestQueueSize, queued.max.requests, and backpressure&lt;/h1>
&lt;p>Your Kafka producers are timing out. Broker logs show no errors, but client-side metrics reveal growing latency and retries. On the broker, &lt;code>RequestQueueSize&lt;/code> climbs toward &lt;code>queued.max.requests&lt;/code> (default 500). Once the queue fills, network threads block on enqueue and stop reading from their sockets. Clients see TCP backpressure, time out, and retry. Retries add load, deepening the queue. This feedback loop correlates with falling &lt;code>RequestHandlerAvgIdlePercent&lt;/code>.&lt;/p></description></item><item><title>Kafka REQUEST_TIMED_OUT: produce requests that expire before replication completes</title><link>https://www.netdata.cloud/guides/kafka/kafka-request-timed-out-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-request-timed-out-error/</guid><description>&lt;h1 id="kafka-request_timed_out-produce-requests-that-expire-before-replication-completes">Kafka REQUEST_TIMED_OUT: produce requests that expire before replication completes&lt;/h1>
&lt;p>Producers using &lt;code>acks=all&lt;/code> throw &lt;code>TimeoutException&lt;/code> (error code &lt;code>REQUEST_TIMED_OUT&lt;/code>) when the broker accepts a produce request but cannot complete replication before &lt;code>request.timeout.ms&lt;/code> expires. The leader appends the record to its local log and waits in purgatory for acknowledgments from all in-sync replicas. If the ISR ack does not arrive before the client deadline, the producer disconnects and surfaces the error. The broker may still finish the write, but the producer has already moved on, creating a hidden replication backlog and a potential retry storm.&lt;/p></description></item><item><title>Kafka RequestHandlerAvgIdlePercent low: I/O thread saturation and overload</title><link>https://www.netdata.cloud/guides/kafka/kafka-request-handler-idle-percent-low/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-request-handler-idle-percent-low/</guid><description>&lt;h1 id="kafka-requesthandleravgidlepercent-low-io-thread-saturation-and-overload">Kafka RequestHandlerAvgIdlePercent low: I/O thread saturation and overload&lt;/h1>
&lt;p>You get paged because &lt;code>RequestHandlerAvgIdlePercent&lt;/code> is below 0.3 and falling. This is an exponentially weighted moving average, not a spike metric. A low value means the broker&amp;rsquo;s I/O handler threads have been saturated long enough that the request queue is backing up and clients are timing out.&lt;/p>
&lt;p>Treat above 0.5 as healthy, below 0.3 as critical, and below 0.1 as active overload. Even a stable 0.45 is dangerous: one broker failure can shift enough load to collapse the survivors.&lt;/p></description></item><item><title>Kafka retention not deleting old segments: retention.ms, retention.bytes, and the active segment</title><link>https://www.netdata.cloud/guides/kafka/kafka-retention-not-deleting-segments/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-retention-not-deleting-segments/</guid><description>&lt;h1 id="kafka-retention-not-deleting-old-segments-retentionms-retentionbytes-and-the-active-segment">Kafka retention not deleting old segments: retention.ms, retention.bytes, and the active segment&lt;/h1>
&lt;p>You set &lt;code>retention.ms&lt;/code> to 24 hours, but broker disk keeps climbing. A partition shows segment files older than the threshold, or the active segment has grown so large it consumes most of the volume. A topic with both &lt;code>retention.ms&lt;/code> and &lt;code>retention.bytes&lt;/code> may still appear to ignore them.&lt;/p>
&lt;p>Kafka retention is not a continuous sweep. &lt;code>retention.ms&lt;/code> and &lt;code>retention.bytes&lt;/code> apply independently to closed segments, and the check interval adds latency. The active segment is never deleted by retention alone. On compacted topics, &lt;code>retention.bytes&lt;/code> is ignored. Per-topic overrides shadow broker defaults.&lt;/p></description></item><item><title>Kafka Too many open files: file descriptor exhaustion from segments and connections</title><link>https://www.netdata.cloud/guides/kafka/kafka-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-too-many-open-files/</guid><description>&lt;h1 id="kafka-too-many-open-files-file-descriptor-exhaustion-from-segments-and-connections">Kafka Too many open files: file descriptor exhaustion from segments and connections&lt;/h1>
&lt;p>Producers time out, consumers disconnect, and the broker log shows &lt;code>java.io.IOException: Too many open files&lt;/code>. The broker may still run, but it cannot open new log segments or accept additional TCP connections. File descriptor exhaustion is a cliff-edge failure: the broker operates normally until it hits the hard limit, then the data path stops.&lt;/p>
&lt;p>Kafka brokers hold a file descriptor for every log segment and every network connection. Each partition maintains an active segment and older retained segments. A broker with thousands of partitions and tens of segments each, plus hundreds of client connections, can hold tens of thousands of open file descriptors. The default Linux per-process limit of 1024 is far below production requirements.&lt;/p></description></item><item><title>Kafka too many partitions per broker: controller load, recovery time, and the 4000 guideline</title><link>https://www.netdata.cloud/guides/kafka/kafka-too-many-partitions-per-broker/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-too-many-partitions-per-broker/</guid><description>&lt;h1 id="kafka-too-many-partitions-per-broker-controller-load-recovery-time-and-the-4000-guideline">Kafka too many partitions per broker: controller load, recovery time, and the 4000 guideline&lt;/h1>
&lt;p>Every replica on a Kafka broker carries fixed overhead: file descriptors, memory-mapped index files, replication fetcher load, and controller metadata. The guideline of roughly 4000 partitions per broker is not a hard architectural limit. It is an operational warning based on a serial bottleneck: the controller processes partition state changes one at a time. When a broker with 5000 partitions fails, the active controller queues thousands of events. Until the queue drains, leader elections stall, metadata propagation slows, and clients see &lt;code>NOT_LEADER_FOR_PARTITION&lt;/code>. Restarting the broker forces every replica to recover its logs and catch up from leaders before rejoining the ISR. Recovery time is a step function; a routine restart can become a multi-hour incident. Partition count is easy to ignore in steady state and catastrophic to discover during an outage.&lt;/p></description></item><item><title>Kafka UncleanLeaderElectionsPerSec > 0: confirmed silent data loss</title><link>https://www.netdata.cloud/guides/kafka/kafka-unclean-leader-election/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-unclean-leader-election/</guid><description>&lt;h1 id="kafka-uncleanleaderelectionspersec--0-confirmed-silent-data-loss">Kafka UncleanLeaderElectionsPerSec &amp;gt; 0: confirmed silent data loss&lt;/h1>
&lt;p>Your &lt;code>UncleanLeaderElectionsPerSec&lt;/code> alert fired. The JMX metric &lt;code>kafka.controller:type=ControllerStats,name=UncleanLeaderElectionsPerSec&lt;/code> shows &lt;code>OneMinuteRate &amp;gt; 0&lt;/code>, or &lt;code>Count&lt;/code> has incremented since your last check. This is confirmed data loss: a partition leader was elected from outside the ISR, and acknowledged records that the new leader does not possess are silently truncated.&lt;/p>
&lt;p>Producers that received acks for those records were given a durability guarantee the cluster just broke. Stop additional loss, understand scope, and fix the conditions that allowed it.&lt;/p></description></item><item><title>Kafka UnderMinIsrPartitionCount: confirming the write path is blocked</title><link>https://www.netdata.cloud/guides/kafka/kafka-under-min-isr-partition-count/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-under-min-isr-partition-count/</guid><description>&lt;h1 id="kafka-underminisrpartitioncount-confirming-the-write-path-is-blocked">Kafka UnderMinIsrPartitionCount: confirming the write path is blocked&lt;/h1>
&lt;p>&lt;code>UnderMinIsrPartitionCount&lt;/code> is non-zero and &lt;code>acks=all&lt;/code> producers are failing. Unlike &lt;code>UnderReplicatedPartitions&lt;/code>, which only opens a durability window, this metric confirms the broker is actively rejecting writes. It counts leader partitions where the in-sync replica set has shrunk below &lt;code>min.insync.replicas&lt;/code>. In steady state it must be zero.&lt;/p>
&lt;p>This article explains the metric, when to PAGE, and the exact checks to run when it fires.&lt;/p></description></item><item><title>Kafka UnderReplicatedPartitions > 0: the most important metric and how to clear it</title><link>https://www.netdata.cloud/guides/kafka/kafka-under-replicated-partitions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-under-replicated-partitions/</guid><description>&lt;h1 id="kafka-underreplicatedpartitions--0-the-most-important-metric-and-how-to-clear-it">Kafka UnderReplicatedPartitions &amp;gt; 0: the most important metric and how to clear it&lt;/h1>
&lt;p>UnderReplicatedPartitions climbing above zero means a follower has fallen behind, the ISR has shrunk, and the cluster&amp;rsquo;s durability guarantee is degraded. One more failure on the wrong broker could make partitions unavailable or cause &lt;code>acks=all&lt;/code> writes to be rejected. Determine whether this is a transient blip from maintenance or the start of a cascading replication failure.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>The JMX MBean &lt;code>kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions&lt;/code> reports the number of partitions where this broker is the leader and the ISR count is below the configured replication factor. It is a per-broker gauge. A broker without leadership always reports zero, even when the cluster is degraded, so aggregate across all brokers to see the full picture.&lt;/p></description></item><item><title>Kafka ZooKeeper</title><link>https://www.netdata.cloud/integrations/data-collection/databases/kafka-zookeeper/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/kafka-zookeeper/</guid><description/></item><item><title>Kafka ZooKeeper Monitoring</title><link>https://www.netdata.cloud/monitoring-101/kafka_zookeeper-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/kafka_zookeeper-monitoring/</guid><description>&lt;h2 id="kafka-zookeeper-monitoring">Kafka ZooKeeper Monitoring&lt;/h2>
&lt;h3 id="what-is-kafka-zookeeper">What Is Kafka ZooKeeper?&lt;/h3>
&lt;p>Kafka ZooKeeper is a centralized service for maintaining configuration information, naming, providing distributed synchronization, and providing group services. It is a critical component in the architecture of distributed systems, especially for managing and coordinating services like Apache Kafka.&lt;/p>
&lt;h3 id="monitoring-kafka-zookeeper-with-netdata">Monitoring Kafka ZooKeeper With Netdata&lt;/h3>
&lt;p>Monitoring Kafka ZooKeeper with Netdata offers a comprehensive view of the system’s performance and health. To monitor Kafka ZooKeeper, Netdata uses an openmetrics (Prometheus) exporter called the &lt;a href="https://github.com/cloudflare/kafka_zookeeper_exporter">Kafka ZooKeeper Exporter&lt;/a>. This allows Netdata to ingest data from any Prometheus exporter efficiently. With Netdata, you get automated dashboards, alerts, and more without the need for a Prometheus server or Grafana. This streamlined approach with the Netdata $name monitoring tool ensures ease of use and rapid insights into your Kafka ZooKeeper instances.&lt;/p></description></item><item><title>Kafka ZooKeeper request latency high: metadata slowdowns in ZK mode</title><link>https://www.netdata.cloud/guides/kafka/kafka-zookeeper-request-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-zookeeper-request-latency-high/</guid><description>&lt;h1 id="kafka-zookeeper-request-latency-high-metadata-slowdowns-in-zk-mode">Kafka ZooKeeper request latency high: metadata slowdowns in ZK mode&lt;/h1>
&lt;p>If you run Kafka in ZooKeeper mode and &lt;code>kafka.server:type=ZooKeeperClientMetrics,name=ZooKeeperRequestLatencyMs&lt;/code> is climbing, a sustained p99 above 100 ms means the controller event queue is already backing up. Above 1 s, you approach the &lt;code>zookeeper.session.timeout.ms&lt;/code> boundary and brokers risk session expiry. In ZK mode, every leader election, ISR change, and topic update flows through the ensemble. When ZK stalls, the controller stalls, metadata propagation freezes, and the cluster degrades in a way that looks like a broker problem but originates in the metadata plane.&lt;/p></description></item><item><title>Kafka ZooKeeper session expired: GC pauses, ISR drops, and controller loss</title><link>https://www.netdata.cloud/guides/kafka/kafka-zookeeper-session-expired/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kafka/kafka-zookeeper-session-expired/</guid><description>&lt;h1 id="kafka-zookeeper-session-expired-gc-pauses-isr-drops-and-controller-loss">Kafka ZooKeeper session expired: GC pauses, ISR drops, and controller loss&lt;/h1>
&lt;p>In ZooKeeper mode, a session expiry means ZK declared a broker dead. The Java process may still be running, but the cluster treats it as gone. If the broker was the controller, the metadata plane re-elects. If it was a leader, every partition it led starts a new leader election. If it was a follower, leaders remove it from the ISR. The default &lt;code>zookeeper.session.timeout.ms&lt;/code> is 18000 ms, and the most common trigger is a Full GC pause longer than that window. This guide applies to Kafka clusters running in ZooKeeper mode; KRaft mode does not use ZK sessions.&lt;/p></description></item><item><title>KairosDB</title><link>https://www.netdata.cloud/integrations/exporters/kairosdb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/kairosdb/</guid><description/></item><item><title>Kannel</title><link>https://www.netdata.cloud/integrations/data-collection/applications/kannel/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/kannel/</guid><description/></item><item><title>Kannel Monitoring</title><link>https://www.netdata.cloud/monitoring-101/kannel-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/kannel-monitoring/</guid><description>&lt;h2 id="kannel-monitoring">Kannel Monitoring&lt;/h2>
&lt;h3 id="what-is-kannel">What Is Kannel?&lt;/h3>
&lt;p>Kannel is an open-source WAP and SMS gateway used widely for mobile communication gateways. It enables applications to send and receive SMS messages and connect to mobile network carriers. Given its crucial role in mobile communication, oversight of its operational metrics is paramount.&lt;/p>
&lt;h3 id="monitoring-kannel-with-netdata">Monitoring Kannel With Netdata&lt;/h3>
&lt;p>To effectively monitor Kannel, Netdata offers an intuitive solution that employs an openmetrics (Prometheus) exporter. This integration does not require a standalone Prometheus server or Grafana setup. Netdata strives as a comprehensive Kannel monitoring tool by allowing users to visualize data in real time with automated dashboards and alerts across various metrics provided by any compatible Prometheus exporter. To get started, you can use the &lt;a href="https://github.com/apostvav/kannel_exporter">Kannel Exporter&lt;/a>.&lt;/p></description></item><item><title>Kashya SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/kashya-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/kashya-snmp-traps/</guid><description/></item><item><title>Kaspersky Lab Zao SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/kaspersky-lab-zao-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/kaspersky-lab-zao-snmp-traps/</guid><description/></item><item><title>Katron Technologies Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/katron-technologies-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/katron-technologies-inc-snmp-traps/</guid><description/></item><item><title>Kavenegar</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/kavenegar/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/kavenegar/</guid><description/></item><item><title>Kcp Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/kcp-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/kcp-inc-snmp-traps/</guid><description/></item><item><title>Keepalived</title><link>https://www.netdata.cloud/integrations/data-collection/networking/keepalived/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/keepalived/</guid><description/></item><item><title>Keepalived Monitoring</title><link>https://www.netdata.cloud/monitoring-101/keepalived-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/keepalived-monitoring/</guid><description>&lt;h2 id="keepalived-monitoring">Keepalived Monitoring&lt;/h2>
&lt;h3 id="what-is-keepalived">What Is Keepalived?&lt;/h3>
&lt;p>Keepalived is a robust software used primarily for high-availability and load-balancing on Linux systems. It leverages VRRP (Virtual Router Redundancy Protocol) to increase network uptime by offering failover protocols. By monitoring network interfaces, Keepalived helps in ensuring that servers are capable of handling traffic efficiently.&lt;/p>
&lt;h3 id="monitoring-keepalived-with-netdata">Monitoring Keepalived With Netdata&lt;/h3>
&lt;p>Monitoring Keepalived with Netdata is a seamless process that ensures you have real-time insights into the performance and availability of your network infrastructure. Netdata uses an &lt;strong>openmetrics (Prometheus) exporter&lt;/strong> to retrieve Keepalived metrics. This means that, to monitor Keepalived, Netdata can accept data from any Prometheus exporter, allowing you to benefit from automated dashboards and alerts without needing a dedicated Prometheus server or Grafana. This integration allows for comprehensive monitoring of Keepalived metrics, ensuring your high-availability setup runs smoothly.&lt;/p></description></item><item><title>Kentix GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/kentix-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/kentix-gmbh-snmp-traps/</guid><description/></item><item><title>Kentrox SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/kentrox-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/kentrox-snmp-traps/</guid><description/></item><item><title>kern.cp_time</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/kern.cp_time/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/kern.cp_time/</guid><description/></item><item><title>kern.ipc.msq</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/kern.ipc.msq/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/kern.ipc.msq/</guid><description/></item><item><title>kern.ipc.sem</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/kern.ipc.sem/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/kern.ipc.sem/</guid><description/></item><item><title>kern.ipc.shm</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/kern.ipc.shm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/kern.ipc.shm/</guid><description/></item><item><title>Kernel Same-Page Merging</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/kernel-same-page-merging/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/kernel-same-page-merging/</guid><description/></item><item><title>Kernel Same-page Merging (KSM) Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ksm-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ksm-monitoring/</guid><description>&lt;h2 id="kernel-same-page-merging-ksm">Kernel Same-page Merging (KSM)?&lt;/h2>
&lt;p>Linux kernels store memory in &lt;strong>pages&lt;/strong> which are moved in and out of memory as a single block. On most Linux architectures pages are 4096 bytes. &lt;strong>KSM&lt;/strong> (Kernel Same-page Merging) is a kernel feature that scans memory looking for pages with identical content, and then de-duplicates them. The most common use-case where such duplicate pages occur is on hosts running multiple virtual machines (VMs).&lt;/p>
&lt;p>KSM can greatly reduce the amount of memory used by VMs. When it finds two or more identical pages, it replaces them with a single page that is shared by all VMs that are using it.&lt;/p></description></item><item><title>Kevin Ether Boulain SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/kevin-ether-boulain-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/kevin-ether-boulain-snmp-traps/</guid><description/></item><item><title>Knuerr AG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/knuerr-ag-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/knuerr-ag-snmp-traps/</guid><description/></item><item><title>Kube-proxy Monitoring</title><link>https://www.netdata.cloud/monitoring-101/kubeproxy-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/kubeproxy-monitoring/</guid><description>&lt;h2 id="what-is-kube-proxy">What is Kube-proxy?&lt;/h2>
&lt;p>&lt;a href="https://kubernetes.io/docs/concepts/overview/components/#kube-proxy">&lt;code>Kube-proxy&lt;/code>&lt;/a> is a network proxy that runs on each node in your cluster, implementing part of the Kubernetes Service.&lt;/p>
&lt;h2 id="monitoring-kube-proxy-with-netdata">Monitoring Kube-proxy with Netdata&lt;/h2>
&lt;p>The prerequisites for monitoring Kube-proxy with Netdata are to have Kube-proxy and &lt;a href="https://learn.netdata.cloud/docs/cloud/get-started/">Netdata installed&lt;/a> on your system.&lt;/p>
&lt;p>Netdata auto discovers hundreds of services, and for those it doesn&amp;rsquo;t turning on manual discovery is a one line configuration. For more information on configuring Netdata for Kube-proxy monitoring please read the collector &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/k8s_kubeproxy/">documentation&lt;/a>.&lt;/p></description></item><item><title>Kubelet</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/kubelet/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/kubelet/</guid><description/></item><item><title>Kubelet Monitoring</title><link>https://www.netdata.cloud/monitoring-101/k8s_kubelet-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/k8s_kubelet-monitoring/</guid><description>&lt;h2 id="kubelet-monitoring">Kubelet Monitoring&lt;/h2>
&lt;p>In today&amp;rsquo;s fast-paced DevOps environments, monitoring the &lt;a href="https://kubernetes.io/docs/concepts/overview/components/#kubelet">Kubelet&lt;/a> effectively is crucial for maintaining healthy Kubernetes clusters. Understanding its metrics allows teams to ensure high availability and performance efficiency for their applications.&lt;/p>
&lt;h3 id="what-is-kubelet">What Is Kubelet?&lt;/h3>
&lt;p>Kubelet is a core component of Kubernetes that runs on each node in the cluster. It ensures that containers are running as expected, watching for changes in Pod specifications and reporting back to the Kubernetes API server.&lt;/p></description></item><item><title>Kubeproxy</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/kubeproxy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/kubeproxy/</guid><description/></item><item><title>Kubeproxy Monitoring</title><link>https://www.netdata.cloud/monitoring-101/k8s_kubeproxy-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/k8s_kubeproxy-monitoring/</guid><description>&lt;h2 id="kubeproxy-monitoring">Kubeproxy Monitoring&lt;/h2>
&lt;h3 id="what-is-kubeproxy">What Is Kubeproxy?&lt;/h3>
&lt;p>Kubeproxy is a critical component of the Kubernetes ecosystem, acting as a network proxy that runs on each node in your cluster. It manages the IP table rules and forwards traffic to correct Pod IP addresses, playing a vital role in facilitating seamless communication within your Kubernetes deployment. Its efficient operation is crucial to maintaining healthy network traffic in Kubernetes environments.&lt;/p>
&lt;h3 id="monitoring-kubeproxy-with-netdata">Monitoring Kubeproxy With Netdata&lt;/h3>
&lt;p>Netdata provides a comprehensive Kubernetes monitoring tool that allows you to monitor Kubeproxy in real-time. By harnessing the &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/k8s_kubeproxy/">Kubeproxy collector&lt;/a>, you can gather pivotal metrics that give insights into your cluster&amp;rsquo;s network performance. These metrics help in diagnosing issues swiftly and ensuring optimal Kubeproxy operations. You can even &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">check out the live demo here&lt;/a> to see these monitoring capabilities in action.&lt;/p></description></item><item><title>Kubernetes</title><link>https://www.netdata.cloud/integrations/all/kubernetes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/all/kubernetes/</guid><description/></item><item><title>Kubernetes (Helm)</title><link>https://www.netdata.cloud/integrations/deploy/docker-kubernetes/kubernetes-helm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/docker-kubernetes/kubernetes-helm/</guid><description/></item><item><title>Kubernetes admission webhook death spiral: detection and recovery</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-webhook-death-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-webhook-death-spiral/</guid><description>&lt;h1 id="kubernetes-admission-webhook-death-spiral-detection-and-recovery">Kubernetes admission webhook death spiral: detection and recovery&lt;/h1>
&lt;p>You deploy a mutating webhook to enforce policy. Later, a node drains and the webhook pod evicts. Now no pods can start, including the webhook&amp;rsquo;s own replacement. The cluster does not crash, but it stops moving. Every &lt;code>kubectl apply&lt;/code> hangs. Horizontal autoscalers freeze. Rolling updates stall. This is the admission webhook death spiral: a circular dependency where the webhook must admit a pod it itself requires.&lt;/p></description></item><item><title>Kubernetes anonymous API access: detection, audit, and lockdown</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-anonymous-access-detection/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-anonymous-access-detection/</guid><description>&lt;h1 id="kubernetes-anonymous-api-access-detection-audit-and-lockdown">Kubernetes anonymous API access: detection, audit, and lockdown&lt;/h1>
&lt;p>Anonymous requests to the Kubernetes API server are authenticated as &lt;code>system:anonymous&lt;/code> and evaluated as part of the &lt;code>system:unauthenticated&lt;/code> group. On many self-managed clusters this behavior is enabled by default, which means requests that arrive without a valid client certificate, bearer token, or other credential are passed to the authorization layer instead of being rejected immediately. Public discovery endpoints and health checks are common legitimate uses, but anonymous access to namespaced resources, secrets, or RBAC objects is a direct security exposure.&lt;/p></description></item><item><title>Kubernetes API Server</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/kubernetes-api-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/kubernetes-api-server/</guid><description/></item><item><title>Kubernetes API server audit logging: policy, backends, and forensics</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-audit-logging/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-audit-logging/</guid><description>&lt;h1 id="kubernetes-api-server-audit-logging-policy-backends-and-forensics">Kubernetes API server audit logging: policy, backends, and forensics&lt;/h1>
&lt;p>Kubernetes API server audit logging is the authoritative record of every request that reaches the control plane. It captures the identity of the caller, the resource and verb, the timestamp, the stage, and the outcome. Without it, a security investigation into unauthorized access, a compliance audit, or a postmortem into a failed certificate rotation is built on inference rather than evidence.&lt;/p></description></item><item><title>Kubernetes API server certificate rotation: detection and grace handling</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-certificate-rotation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-certificate-rotation/</guid><description>&lt;h1 id="kubernetes-api-server-certificate-rotation-detection-and-grace-handling">Kubernetes API server certificate rotation: detection and grace handling&lt;/h1>
&lt;p>Kubernetes control plane certificates created by kubeadm expire after one year. The API server does not auto-rotate its serving certificate. When it expires, etcd rejects control plane connections, kubelets cannot authenticate, and the cluster becomes unreachable. The failure is sudden and total.&lt;/p>
&lt;p>This guide covers kubeadm-managed clusters where you own the control plane. Distinguish between the API server serving certificate and the broader control plane bundle, detect expiration before the outage, and renew with minimal disruption.&lt;/p></description></item><item><title>Kubernetes API server etcd latency: detection and cascading failures</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-etcd-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-etcd-latency/</guid><description>&lt;h1 id="kubernetes-api-server-etcd-latency-detection-and-cascading-failures">Kubernetes API server etcd latency: detection and cascading failures&lt;/h1>
&lt;p>When etcd slows down, the entire control plane slows with it. A few extra milliseconds on disk fsync turns into hung kubectl commands, backed-up controller queues, and eventually a cluster that cannot schedule pods or update endpoints. Detect the etcd latency cascade, confirm whether storage is the root cause, and break the feedback loop before the cluster becomes effectively read-only.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>etcd serializes every Kubernetes mutation. Every API server write becomes a Raft proposal that must fsync to the WAL before etcd acknowledges it. When the disk under etcd is slow, every fsync waits longer. The API server holds mutating requests open until etcd responds. Requests pile up in the inflight queue. Once the queue hits the limit, the API server returns 429 Too Many Requests. Controllers that depend on writes (scheduler, replica set controller, and others) fall behind and retry. Retries generate more write load. The result is a feedback loop: slow disk -&amp;gt; slow etcd -&amp;gt; slow API server -&amp;gt; retry storm -&amp;gt; amplified etcd load.&lt;/p></description></item><item><title>Kubernetes API server FlowSchemas and PriorityLevels: design and tuning</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-flow-schemas/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-flow-schemas/</guid><description>&lt;h1 id="kubernetes-api-server-flowschemas-and-prioritylevels-design-and-tuning">Kubernetes API server FlowSchemas and PriorityLevels: design and tuning&lt;/h1>
&lt;p>Before Kubernetes 1.20, the API server protected itself with two global hard limits: &lt;code>--max-requests-inflight&lt;/code> and &lt;code>--max-mutating-requests-inflight&lt;/code>. Every request, whether a kubelet heartbeat or a runaway controller LIST, competed for the same pool. API Priority and Fairness (APF), enabled by default since 1.20, replaces that coarse model with a two-stage classification and fair-queuing system. It separates requests into priority levels, isolates flows within each level, and rejects or queues traffic before it can starve critical control plane operations.&lt;/p></description></item><item><title>Kubernetes API server memory pressure: OOM cycle and tuning</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-memory-pressure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-memory-pressure/</guid><description>&lt;h1 id="kubernetes-api-server-memory-pressure-oom-cycle-and-tuning">Kubernetes API server memory pressure: OOM cycle and tuning&lt;/h1>
&lt;p>Your control plane is crash-looping. The kube-apiserver process climbs toward its container memory limit, the Go garbage collector stalls trying to reclaim space, and the kernel OOM killer terminates it. The replacement pod starts with cold caches; every connected client immediately re-lists watched resources, and memory spikes again before caches warm. The root cause is usually a mismatch between how the API server uses memory and how much headroom you have given it.&lt;/p></description></item><item><title>Kubernetes API server rate limiting: APF priority levels and starvation</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-rate-limited/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-rate-limited/</guid><description>&lt;h1 id="kubernetes-api-server-rate-limiting-apf-priority-levels-and-starvation">Kubernetes API server rate limiting: APF priority levels and starvation&lt;/h1>
&lt;p>Your API server is running. &lt;code>/healthz&lt;/code> returns 200. &lt;code>/readyz&lt;/code> passes. Yet nodes drop to &lt;code>NotReady&lt;/code>, the scheduler stops placing pods, and controller logs fill with &lt;code>context deadline exceeded&lt;/code>. The cluster is not down, but it is frozen. This pattern often points to API Priority and Fairness (APF) starvation: low-priority traffic consumes the API server&amp;rsquo;s concurrency budget, and critical control plane requests queue or get rejected.&lt;/p></description></item><item><title>Kubernetes API server slow or unresponsive: causes and fixes</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-slow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-slow/</guid><description>&lt;h1 id="kubernetes-api-server-slow-or-unresponsive-causes-and-fixes">Kubernetes API server slow or unresponsive: causes and fixes&lt;/h1>
&lt;p>When &lt;code>kubectl&lt;/code> hangs, controllers log &lt;code>context deadline exceeded&lt;/code>, and deployments stall, the Kubernetes API server is usually the bottleneck. It is the single funnel for every read and write to cluster state. Slowness propagates to scheduling, pod lifecycle, service discovery, and external automation.&lt;/p>
&lt;p>This article covers operational causes and gives a step-by-step diagnostic flow to run during an incident. Use it to distinguish etcd latency, admission webhook stalls, request saturation, and memory pressure.&lt;/p></description></item><item><title>Kubernetes API server watch storm: re-list cascades and connection floods</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-watch-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-watch-storm/</guid><description>&lt;h1 id="kubernetes-api-server-watch-storm-re-list-cascades-and-connection-floods">Kubernetes API server watch storm: re-list cascades and connection floods&lt;/h1>
&lt;p>A sudden wall of LIST requests pins API server CPU, climbs memory, spikes etcd read latency, and lags controllers. The culprit is usually a watch storm: hundreds or thousands of clients simultaneously re-listing because their watch connections failed or fell behind. Each re-list triggers an expensive etcd range scan and serializes all matching objects. Under load, this saturates CPU, fills network bandwidth, and can trigger APF throttling or memory pressure. If the API server restarts before the storm subsides, the cycle repeats.&lt;/p></description></item><item><title>Kubernetes bound service account tokens: rotation, audience, and expiry</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-bound-service-account-tokens/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-bound-service-account-tokens/</guid><description>&lt;h1 id="kubernetes-bound-service-account-tokens-rotation-audience-and-expiry">Kubernetes bound service account tokens: rotation, audience, and expiry&lt;/h1>
&lt;p>Pods fail with 401 Unauthorized, CSI volume mounts hang with token errors, or security audits surface long-lived credentials that never rotate. Bound tokens are projected, audience-scoped, and short-lived. Legacy tokens are static Secrets that persist forever. After Kubernetes 1.24, both coexist in most upgraded clusters. This guide covers the TokenRequest API lifecycle, kubelet rotation behavior, audience binding, and version-specific changes from 1.24 through 1.33 to help you diagnose auth failures and remove stale credentials safely.&lt;/p></description></item><item><title>Kubernetes Cluster State</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/kubernetes-cluster-state/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/kubernetes-cluster-state/</guid><description/></item><item><title>Kubernetes Cluster State Monitoring</title><link>https://www.netdata.cloud/monitoring-101/k8s_state-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/k8s_state-monitoring/</guid><description>&lt;h2 id="kubernetes-cluster-state-monitoring">Kubernetes Cluster State Monitoring&lt;/h2>
&lt;h3 id="what-is-kubernetes-cluster-state">What Is Kubernetes Cluster State?&lt;/h3>
&lt;p>Kubernetes Cluster State refers to the real-time condition or the status of your Kubernetes cluster, including nodes, pods, and containers. Monitoring these elements ensures that your clusters are operating at optimal capacity and helps in rapidly identifying any discrepancies or issues.&lt;/p>
&lt;h3 id="monitoring-kubernetes-cluster-state-with-netdata">Monitoring Kubernetes Cluster State With Netdata&lt;/h3>
&lt;p>Netdata provides a comprehensive Kubernetes Cluster State monitoring tool that allows you to effortlessly oversee node and pod metrics. The integration with the go.d.plugin module makes it simple for you to garner insights into the dynamic environment of Kubernetes.&lt;/p></description></item><item><title>Kubernetes conntrack exhaustion: dropped connections under load</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-conntrack-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-conntrack-exhaustion/</guid><description>&lt;h1 id="kubernetes-conntrack-exhaustion-dropped-connections-under-load">Kubernetes conntrack exhaustion: dropped connections under load&lt;/h1>
&lt;p>Intermittent connection timeouts under load in Kubernetes often trace to a full nf_conntrack table on the node. Existing TCP sessions stay open, but new connections fail silently. DNS resolution becomes unreliable. Application logs show timeouts to healthy dependencies. The root cause is usually not the application, network policy, or CNI, but kernel connection tracking exhaustion.&lt;/p>
&lt;p>Every connection that traverses kube-proxy NAT rules creates an entry in the node&amp;rsquo;s nf_conntrack table. This finite, node-level table is shared by all workloads and invisible to most application monitoring. When it fills, the kernel drops new connection attempts without sending a TCP reset or ICMP error. The application sees a timeout.&lt;/p></description></item><item><title>Kubernetes container runtime shim failures: containerd, CRI-O troubleshooting</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-runtime-shim-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-runtime-shim-failures/</guid><description>&lt;h1 id="kubernetes-container-runtime-shim-failures-containerd-cri-o-troubleshooting">Kubernetes container runtime shim failures: containerd, CRI-O troubleshooting&lt;/h1>
&lt;p>Pods stuck in &lt;code>ContainerCreating&lt;/code>, nodes flapping &lt;code>NotReady&lt;/code>, and PLEG timeouts that clear only after a node reboot usually point to the container runtime shim layer, not the kubelet or network. The shim sits between the kubelet and the low-level runtime. When it hangs, crashes, or leaks, the kubelet cannot enumerate containers, start sandboxes, or reap terminated pods. Existing containers may keep running, but the node stops accepting new work.&lt;/p></description></item><item><title>Kubernetes Containers</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/kubernetes-containers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/kubernetes-containers/</guid><description/></item><item><title>Kubernetes controller-manager leader election failures</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-controller-manager-leader-election/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-controller-manager-leader-election/</guid><description>&lt;h1 id="kubernetes-controller-manager-leader-election-failures">Kubernetes controller-manager leader election failures&lt;/h1>
&lt;p>Your Deployment has stopped scaling. Nodes cordoned hours ago are still draining. Garbage collection is paused, and orphaned volumes are not being cleaned up. The kube-controller-manager runs these reconciliation loops, and in an HA cluster only the leader performs work. When leader election fails, the controller-manager exits, and the control plane stops acting on desired state. Existing workloads keep running, but nothing new is managed.&lt;/p></description></item><item><title>Kubernetes CSI driver failures: detection, recovery, and version skew</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-csi-driver-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-csi-driver-failures/</guid><description>&lt;h1 id="kubernetes-csi-driver-failures-detection-recovery-and-version-skew">Kubernetes CSI driver failures: detection, recovery, and version skew&lt;/h1>
&lt;p>When a workload pod hangs in &lt;code>ContainerCreating&lt;/code> with &lt;code>FailedMount&lt;/code> or &lt;code>FailedAttach&lt;/code> events, the root cause is often a CSI driver pod that has crashed, a node plugin that is missing on the target node, or a version mismatch between the driver and its sidecars. Unlike application pods, CSI drivers sit on the critical path for every volume operation. Their failure is not isolated; it blocks scheduling, provisioning, and recovery.&lt;/p></description></item><item><title>Kubernetes DaemonSet pods Pending: scheduling and tolerations</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-daemonset-pods-pending/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-daemonset-pods-pending/</guid><description>&lt;h1 id="kubernetes-daemonset-pods-pending-scheduling-and-tolerations">Kubernetes DaemonSet pods Pending: scheduling and tolerations&lt;/h1>
&lt;p>A DaemonSet pod in &lt;code>Pending&lt;/code> on a node where it should run is a scheduling failure, not a workload crash. Since Kubernetes 1.18, DaemonSet pods pass through the default scheduler like any other pod. The DaemonSet controller creates the pod and pins it to a target node via &lt;code>nodeAffinity&lt;/code>, but the scheduler still evaluates predicates: taints, tolerations, resource requests, and node state. If the scheduler rejects the pod, it stays &lt;code>Pending&lt;/code>. The controller will not create a replacement; it waits for the scheduler to succeed.&lt;/p></description></item><item><title>Kubernetes Deployment rollout stuck: stalled rollouts and ready replicas</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-deployment-rollout-stuck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-deployment-rollout-stuck/</guid><description>&lt;h1 id="kubernetes-deployment-rollout-stuck-stalled-rollouts-and-ready-replicas">Kubernetes Deployment rollout stuck: stalled rollouts and ready replicas&lt;/h1>
&lt;p>A Deployment rollout that stalls is a silent capacity leak. The old ReplicaSet scales down, the new ReplicaSet stops halfway, and &lt;code>kubectl rollout status&lt;/code> blocks indefinitely. Kubernetes does not automatically recover. The controller sets &lt;code>ProgressDeadlineExceeded&lt;/code> only after &lt;code>progressDeadlineSeconds&lt;/code> elapses, and takes no corrective action. You need to distinguish between a Pod lifecycle blockage, a readiness probe or gate failure, and a rare controller bug that freezes the rollout entirely.&lt;/p></description></item><item><title>Kubernetes DNS resolution failures inside pods</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-dns-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-dns-failures/</guid><description>&lt;h1 id="kubernetes-dns-resolution-failures-inside-pods">Kubernetes DNS resolution failures inside pods&lt;/h1>
&lt;p>DNS failures inside pods break service discovery. A single overloaded CoreDNS replica or saturated conntrack table on one node can look like a multi-service outage. Before fixing, determine whether the failure is cluster-wide, node-specific, or workload-specific.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>Kubernetes injects an &lt;code>/etc/resolv.conf&lt;/code> into every pod that points to the cluster DNS service, typically CoreDNS. CoreDNS resolves cluster-internal names via the kubernetes plugin and forwards external queries to an upstream resolver. A failure at any point produces the same symptom: the name cannot be resolved.&lt;/p></description></item><item><title>Kubernetes etcd defragmentation: when, how, and what breaks</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-etcd-defragmentation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-etcd-defragmentation/</guid><description>&lt;h1 id="kubernetes-etcd-defragmentation-when-how-and-what-breaks">Kubernetes etcd defragmentation: when, how, and what breaks&lt;/h1>
&lt;p>etcd&amp;rsquo;s database file grows over time even when you are not adding objects. Compaction removes old revisions logically, but the bbolt backend does not shrink the file on disk. Without defragmentation, the gap between physical file size and actual data footprint widens until the cluster hits a &lt;code>NOSPACE&lt;/code> alarm and rejects all writes. This guide covers measuring that gap, sequencing defragmentation across an HA cluster without triggering leader elections, and the monitoring signals that predict when you need to act.&lt;/p></description></item><item><title>Kubernetes etcd disk fsync latency: detection and tuning</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-etcd-disk-fsync/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-etcd-disk-fsync/</guid><description>&lt;h1 id="kubernetes-etcd-disk-fsync-latency-detection-and-tuning">Kubernetes etcd disk fsync latency: detection and tuning&lt;/h1>
&lt;p>etcd serializes every Kubernetes write. When disk fsync latency rises, Raft heartbeats stall, leaders step down, and the API server returns 500s and 429s while controllers retry into a death spiral. Unlike CPU or memory pressure, etcd disk latency is invisible until it is catastrophic. This guide shows how to detect it early, isolate the root cause between storage hardware, database size, and configuration, and tune the cluster to survive production load.&lt;/p></description></item><item><title>Kubernetes etcd snapshot failures: backup, restore, and verification</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-etcd-snapshot-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-etcd-snapshot-failures/</guid><description>&lt;h1 id="kubernetes-etcd-snapshot-failures-backup-restore-and-verification">Kubernetes etcd snapshot failures: backup, restore, and verification&lt;/h1>
&lt;p>An etcd snapshot failure usually surfaces during an incident, not during backup. A snapshot from an unhealthy member, a corrupted transfer to object storage, or a restore that writes no data renders disaster recovery useless. This guide gives the checks, commands, and decision logic to verify snapshot integrity, fix backup failures, and perform clean restores.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>An etcd snapshot captures the entire key-value store at a point in time. Because committed Raft log entries exist on a majority of members, a snapshot from any healthy member contains the full cluster state. The snapshot file includes a SHA-256 hash computed at save time. If that hash does not match after transfer, or if the snapshot is taken while etcd is under NOSPACE alarm or leader instability, the file may be inconsistent.&lt;/p></description></item><item><title>Kubernetes eviction cascade: when one node failure takes down the cluster</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-eviction-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-eviction-cascade/</guid><description>&lt;h1 id="kubernetes-eviction-cascade-when-one-node-failure-takes-down-the-cluster">Kubernetes eviction cascade: when one node failure takes down the cluster&lt;/h1>
&lt;p>You see pods entering Evicted status across multiple nodes. Nodes flap between Ready and MemoryPressure or DiskPressure. The scheduler keeps placing replacements, but the new pods are evicted again before they become ready. Workloads never stabilize, and every remediation attempt seems to make the cluster more volatile.&lt;/p>
&lt;p>This is a node-pressure eviction cascade. It happens when the scheduler&amp;rsquo;s view of capacity diverges from the kubelet&amp;rsquo;s view. One node under pressure evicts pods; those pods land on other nodes that are also overcommitted; those nodes tip into pressure and evict more pods. The result is a cluster-wide feedback loop that looks like a resource shortage but is often a scheduling and configuration problem.&lt;/p></description></item><item><title>Kubernetes headless service resolution: SRV records and pod discovery</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-headless-service-resolution/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-headless-service-resolution/</guid><description>&lt;h1 id="kubernetes-headless-service-resolution-srv-records-and-pod-discovery">Kubernetes headless service resolution: SRV records and pod discovery&lt;/h1>
&lt;p>You deployed a StatefulSet with a headless Service so peers can discover each other, but nslookup returns NXDOMAIN or only a single IP when several pods are running. Your application might rely on SRV records for port discovery and the lookup returns nothing. Or a pod rescheduled onto a new node and clients kept trying the old IP for minutes because the TTL behavior surprised you.&lt;/p></description></item><item><title>Kubernetes imagePullSecrets: configuration, propagation, and rotation</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-image-pull-secrets/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-image-pull-secrets/</guid><description>&lt;h1 id="kubernetes-imagepullsecrets-configuration-propagation-and-rotation">Kubernetes imagePullSecrets: configuration, propagation, and rotation&lt;/h1>
&lt;p>Pods stuck in &lt;code>ImagePullBackOff&lt;/code> are rarely caused by a missing or incorrect image tag. More often, the kubelet lacks valid registry credentials. Kubernetes uses &lt;code>imagePullSecrets&lt;/code>, namespaced Secrets of type &lt;code>kubernetes.io/dockerconfigjson&lt;/code>, to inject registry auth into a pod. These can be attached directly to the pod spec or propagated through a ServiceAccount.&lt;/p>
&lt;p>This guide covers propagation from registry to kubelet, verification at each link, and rotation without forcing a rolling restart of every workload.&lt;/p></description></item><item><title>Kubernetes init container fails: blocking main container start</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-init-container-fails/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-init-container-fails/</guid><description>&lt;h1 id="kubernetes-init-container-fails-blocking-main-container-start">Kubernetes init container fails: blocking main container start&lt;/h1>
&lt;p>An init container failure blocks the entire pod. Until every init container exits zero, no main container starts. In automated clusters, one failing init container can stall a deployment while pods sit in &lt;code>PodInitializing&lt;/code>. This guide covers init container retry mechanics, how to read pod status to find the failing step, and how to distinguish application bugs, resource limits, and kubelet issues that leave pods permanently stuck.&lt;/p></description></item><item><title>Kubernetes Job and CronJob troubleshooting: history, backoff, and missed runs</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-job-cronjob-troubleshooting/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-job-cronjob-troubleshooting/</guid><description>&lt;h1 id="kubernetes-job-and-cronjob-troubleshooting-history-backoff-and-missed-runs">Kubernetes Job and CronJob troubleshooting: history, backoff, and missed runs&lt;/h1>
&lt;p>You deployed a CronJob to run every minute, but the last successful run was three hours ago. Or a critical data-processing Job failed with BackoffLimitExceeded after six silent retries, leaving a trail of failed Pods and no clear signal about what broke. Batch workloads fail differently from long-running services: they are time-bound, retry-sensitive, and leave debris in etcd if you do not clean them up. Read the failure signals, distinguish retry storms from control plane delays, and fix the root cause without guessing.&lt;/p></description></item><item><title>Kubernetes kube-proxy and CNI rule conflicts: detection and fix</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-kube-proxy-cni-conflict/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-kube-proxy-cni-conflict/</guid><description>&lt;h1 id="kubernetes-kube-proxy-and-cni-rule-conflicts-detection-and-fix">Kubernetes kube-proxy and CNI rule conflicts: detection and fix&lt;/h1>
&lt;p>Pods stuck in ContainerCreating while the node reports Ready. Services time out despite existing endpoints. Intermittent connection resets during rolling updates. These symptoms usually indicate a conflict between kube-proxy and the container network interface (CNI) plugin over netfilter rules. Both components program the same kernel tables, compete for the same locks, and can corrupt each other&amp;rsquo;s chains. The data plane degrades while the control plane stays healthy.&lt;/p></description></item><item><title>Kubernetes kube-proxy iptables sync stall: causes and recovery</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-iptables-sync-stall/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-iptables-sync-stall/</guid><description>&lt;h1 id="kubernetes-kube-proxy-iptables-sync-stall-causes-and-recovery">Kubernetes kube-proxy iptables sync stall: causes and recovery&lt;/h1>
&lt;p>Pods fail to start. Services intermittently route traffic to dead endpoints. The kube-proxy health endpoint still returns HTTP 200, so the DaemonSet looks healthy, yet rules drift further behind with every sync cycle.&lt;/p>
&lt;p>An iptables sync stall is not a crash. It is a slowdown or blockage in the control loop that translates Service and EndpointSlice state into kernel NAT rules. When kube-proxy cannot acquire the global xtables lock, when &lt;code>iptables-restore&lt;/code> hangs, or when the rule set grows too large to reconcile within the sync period, the node forwards packets using stale rules. New endpoints are invisible. Terminated pods still receive connections. CNI plugins that also need the xtables lock time out, and pod sandbox creation fails.&lt;/p></description></item><item><title>Kubernetes kube-proxy IPVS: stale rules and session affinity issues</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-ipvs-stale-rules/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-ipvs-stale-rules/</guid><description>&lt;h1 id="kubernetes-kube-proxy-ipvs-stale-rules-and-session-affinity-issues">Kubernetes kube-proxy IPVS: stale rules and session affinity issues&lt;/h1>
&lt;p>DNS queries start timing out from one node after a CoreDNS rolling update. A UDP Service returns timeouts for some clients but not others. New Services are unreachable from a specific node while older Services continue to work.&lt;/p>
&lt;p>In IPVS mode, kube-proxy programs the kernel&amp;rsquo;s IPVS table with virtual servers and real servers. The IPVS connection table lives outside kube-proxy&amp;rsquo;s direct control and outside nf_conntrack. That separation creates two IPVS-specific failure modes: stale rules that diverge from EndpointSlice state, and UDP session affinity that sticks to dead backends long after a pod terminates.&lt;/p></description></item><item><title>Kubernetes kubelet certificate expired: detection, rotation, and recovery</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-kubelet-certificate-expired/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-kubelet-certificate-expired/</guid><description>&lt;h1 id="kubernetes-kubelet-certificate-expired-detection-rotation-and-recovery">Kubernetes kubelet certificate expired: detection, rotation, and recovery&lt;/h1>
&lt;p>A healthy node suddenly shows NotReady. Pods keep running, but the kubelet stops reporting status. &lt;code>kubectl logs&lt;/code> and &lt;code>kubectl exec&lt;/code> fail with TLS errors. The cluster event stream is quiet. This is usually an expired kubelet client certificate that failed to rotate.&lt;/p>
&lt;p>Every kubelet maintains two independent TLS credentials: a client certificate that authenticates it to the kube-apiserver, and a serving certificate that secures the kubelet&amp;rsquo;s own HTTPS endpoints. Both typically have a one-year validity. When the client certificate expires, the kubelet cannot authenticate to the API server. The node goes NotReady. Workloads may continue running, but they are unmanaged: no evictions, no probe execution, no status updates, and no new pod scheduling. When the serving certificate expires, metrics-server, &lt;code>kubectl exec&lt;/code>, and &lt;code>kubectl logs&lt;/code> break even if the node is otherwise Ready.&lt;/p></description></item><item><title>Kubernetes kubelet goroutine leaks: detection and bisection</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-kubelet-goroutine-leaks/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-kubelet-goroutine-leaks/</guid><description>&lt;h1 id="kubernetes-kubelet-goroutine-leaks-detection-and-bisection">Kubernetes kubelet goroutine leaks: detection and bisection&lt;/h1>
&lt;p>A kubelet goroutine leak is a slow-burn failure. The process stays up, the node often remains Ready for hours, and then PLEG timeouts start, sync loops lag, and the kubelet is eventually OOM-killed or unresponsive. By the time the node flips to NotReady, the original leak signature is usually obscured by secondary symptoms.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>The kubelet spawns goroutines for pod workers, probes, API server watches, volume operations, PLEG relists, and CRI calls. In a healthy node, goroutine count correlates with pod density and returns to a stable floor when churn stops. A leak means goroutines are created but never exit, causing three cascading effects:&lt;/p></description></item><item><title>Kubernetes kubelet memory leak: detection and OOM cycle</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-kubelet-memory-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-kubelet-memory-leak/</guid><description>&lt;h1 id="kubernetes-kubelet-memory-leak-detection-and-oom-cycle">Kubernetes kubelet memory leak: detection and OOM cycle&lt;/h1>
&lt;p>Kubelet memory growth ends one of two ways: the process hits its cgroup limit or the node runs out of memory. The kernel OOM killer sends SIGKILL. Systemd restarts kubelet, but the new process has cold caches and immediately runs a full reconciliation pass: relisting all containers, re-syncing every pod status, and re-attaching every volume. On a busy node, that burst spikes CPU and memory, which can push the fresh kubelet back over the edge and create a Ready/NotReady flap cycle.&lt;/p></description></item><item><title>Kubernetes kubelet not responding: PLEG, runtime, and certificate issues</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-kubelet-not-responding/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-kubelet-not-responding/</guid><description>&lt;h1 id="kubernetes-kubelet-not-responding-pleg-runtime--certificate-issues">Kubernetes Kubelet Not Responding: PLEG, Runtime &amp;amp; Certificate Issues&lt;/h1>
&lt;p>&lt;a href="https://www.netdata.cloud/guides/kubernetes/">A Kubernetes node&lt;/a> flipping to NotReady while containers keep running is one of the most confusing production failure modes. The kubelet is the node agent that reconciles API server intent with running containers. When it stops responding or reports unhealthy subsystems, the control plane marks the node NotReady and reschedules workloads, even though the data plane may still serve traffic.&lt;/p>
&lt;p>This guide covers three failure domains: Pod Lifecycle Event Generator (PLEG) stalls, container runtime disconnections, and kubelet certificate expiration or rotation failures. Distinguish these symptoms, run safe targeted diagnostics, and apply fixes without blind node reboots.&lt;/p></description></item><item><title>Kubernetes kubelet pod CIDR changes: detection and rolling fix</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-kubelet-pod-cidr-changes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-kubelet-pod-cidr-changes/</guid><description>&lt;h1 id="kubernetes-kubelet-pod-cidr-changes-detection-and-rolling-fix">Kubernetes kubelet pod CIDR changes: detection and rolling fix&lt;/h1>
&lt;p>Pod sandbox creation errors, nodes registering without an IP range, and cross-node traffic failures are symptoms of pod CIDR drift. The kubelet does not assign its own pod CIDR; the kube-controller-manager node IPAM controller writes the range into &lt;code>node.spec.podCIDR&lt;/code> and &lt;code>node.spec.podCIDRs&lt;/code>. When that assignment fails or diverges, the result is usually a slow fracture: some nodes host pods while others cannot, or pods on different nodes lose connectivity.&lt;/p></description></item><item><title>Kubernetes kubelet pprof troubleshooting: capturing heap and goroutine profiles</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-kubelet-pprof-troubleshooting/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-kubelet-pprof-troubleshooting/</guid><description>&lt;h1 id="kubernetes-kubelet-pprof-troubleshooting-capturing-heap-and-goroutine-profiles">Kubernetes kubelet pprof troubleshooting: capturing heap and goroutine profiles&lt;/h1>
&lt;p>When a node starts flapping between Ready and NotReady, or kubelet memory climbs steadily while pod count stays flat, node-level metrics like &lt;code>kubelet_goroutines&lt;/code> and &lt;code>process_resident_memory_bytes&lt;/code> will tell you that the kubelet process is struggling. They will not tell you whether the leak is in the PLEG relist path, the volume manager, or a probe goroutine pool. For that, you need a profile.&lt;/p></description></item><item><title>Kubernetes kubelet volume deadlock: detection and recovery</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-kubelet-volume-deadlock/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-kubelet-volume-deadlock/</guid><description>&lt;h1 id="kubernetes-kubelet-volume-deadlock-detection-and-recovery">Kubernetes kubelet volume deadlock: detection and recovery&lt;/h1>
&lt;p>When pods hang in &lt;code>ContainerCreating&lt;/code> or &lt;code>Terminating&lt;/code> while the node stays &lt;code>Ready=True&lt;/code>, the kubelet volume manager is often the culprit. A blocked mount, unmount, attach, or detach operation consumes a goroutine in the volume manager&amp;rsquo;s finite pool. Once enough operations hang, the queue saturates. Subsequent pods needing volumes stall indefinitely, while pods without volumes start normally. The scheduler, seeing a healthy node, may keep placing volume-bound pods onto the saturated node, deepening the backlog. This guide covers how to confirm a volume deadlock, distinguish it from slow storage, and recover safely.&lt;/p></description></item><item><title>Kubernetes monitoring checklist: the signals every production cluster needs</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-monitoring-checklist/</guid><description>&lt;h1 id="kubernetes-monitoring-checklist-the-signals-every-production-cluster-needs">Kubernetes Monitoring Checklist: The Signals Every Production Cluster Needs&lt;/h1>
&lt;p>This article is a reference checklist for senior engineers who are wiring up, auditing, or hardening monitoring for a production Kubernetes cluster. It assumes you already understand the control plane architecture and focuses on what to collect, where to find it, and which symptoms matter. Use it during greenfield instrumentation, post-incident gap analysis, or routine health audits.&lt;/p>
&lt;p>The signals are grouped by domain. Each entry leads with a short noun phrase, followed by one sentence explaining why it matters, and a concrete warning sign to alert on. Thresholds are drawn from upstream SLOs, kubelet defaults, and etcd operational limits documented in the Kubernetes source and production playbooks. If you run a managed service such as EKS, GKE, or AKS, treat control-plane metrics as provider-mediated; many etcd and API server internals are opaque in those environments.&lt;/p></description></item><item><title>Kubernetes monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-monitoring-maturity-model/</guid><description>&lt;h1 id="kubernetes-monitoring-maturity-model-from-survival-to-expert">Kubernetes monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Kubernetes failures rarely announce themselves. A slow etcd disk cascades into API latency, which backs up controller workqueues, which delays pod scheduling, which triggers autoscaling, which amplifies the load that caused the original latency spike. Without the right signals, you will debug the autoscaling event while the real problem is a WAL fsync that crossed 100ms ten minutes earlier.&lt;/p>
&lt;p>A monitoring maturity model is not a tooling checklist. It is a coverage framework that tells you which signals you are missing and what those gaps cost you during an incident. The four levels below map the progression from &amp;ldquo;Is the cluster on fire?&amp;rdquo; to &amp;ldquo;Why did that single pod take three milliseconds longer to start on node twelve?&amp;rdquo; Each level assumes the previous and adds signals that change your mean time to detection and your mean time to root cause.&lt;/p></description></item><item><title>Kubernetes NetworkPolicy debugging: when traffic is denied silently</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-network-policy-debugging/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-network-policy-debugging/</guid><description>&lt;h1 id="kubernetes-networkpolicy-debugging-when-traffic-is-denied-silently">Kubernetes NetworkPolicy debugging: when traffic is denied silently&lt;/h1>
&lt;p>A pod that could reach its dependency yesterday now times out today. There is no TCP RST, no ICMP unreachable, and often no application log. The packet is dropped in the CNI data plane. If a policy change, namespace reorganization, or cluster upgrade preceded the outage, you are likely dealing with silent NetworkPolicy denial. This guide shows how to confirm it, find the rule or semantic gap responsible, and restore connectivity without opening the cluster.&lt;/p></description></item><item><title>Kubernetes node CPU saturation: load, throttling, and runqueue depth</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-node-cpu-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-node-cpu-saturation/</guid><description>&lt;h1 id="kubernetes-node-cpu-saturation-load-throttling-and-runqueue-depth">Kubernetes node CPU saturation: load, throttling, and runqueue depth&lt;/h1>
&lt;p>Application latency climbs and pods slow down. &lt;code>kubectl top nodes&lt;/code> reports 70 percent CPU, so you assume headroom exists. It does not. CPU percent is a time-average that masks micro-bursts, runqueue backlog, and CFS throttling. A container can throttle to a crawl while node utilization looks comfortable, and a node can show 50 percent utilization with every runnable thread queued behind a noisy neighbor. Distinguish node-level CPU contention from limit-induced throttling using runqueue depth, CFS bandwidth metrics, and Pressure Stall Information (PSI).&lt;/p></description></item><item><title>Kubernetes node DiskPressure: detection, eviction, and recovery</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-node-disk-pressure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-node-disk-pressure/</guid><description>&lt;h1 id="kubernetes-node-diskpressure-detection-eviction--recovery">Kubernetes Node DiskPressure: Detection, Eviction &amp;amp; Recovery&lt;/h1>
&lt;p>A node reporting DiskPressure is actively shedding workloads. The kubelet has detected that nodefs or imagefs has crossed an eviction threshold. It is garbage collecting images, terminating pods, and applying the &lt;code>node.kubernetes.io/disk-pressure&lt;/code> taint to block new scheduling. Existing pods may continue running, but any pod requiring disk for logs, emptyDir volumes, or image pulls is at risk.&lt;/p>
&lt;p>Disk pressure builds predictably, unlike memory pressure. This guide covers how the kubelet evaluates disk pressure, how to distinguish nodefs from imagefs exhaustion, how to find the specific consumer, and how to recover without causing a cascading eviction loop.&lt;/p></description></item><item><title>Kubernetes node MemoryPressure: detection, eviction order, and prevention</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-node-memory-pressure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-node-memory-pressure/</guid><description>&lt;h1 id="kubernetes-node-memorypressure-detection-eviction-order-and-prevention">Kubernetes node MemoryPressure: detection, eviction order, and prevention&lt;/h1>
&lt;p>Before adding RAM, determine whether kubelet is evicting because workloads are genuinely starving or because memory requests are misaligned with reality.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>Kubelet evaluates &lt;code>memory.available&lt;/code> against an eviction threshold. On Linux the default hard threshold is &lt;code>memory.available &amp;lt; 100Mi&lt;/code>. Kubelet derives this from cgroup stats, not &lt;code>free -m&lt;/code>. It measures working-set memory (RSS plus active file-backed pages) and subtracts that from total capacity. When the threshold is crossed, kubelet sets the node condition &lt;code>MemoryPressure=True&lt;/code> and adds the taint &lt;code>node.kubernetes.io/memory-pressure:NoSchedule&lt;/code>. New pods are blocked from scheduling until the condition clears.&lt;/p></description></item><item><title>Kubernetes node NotReady: kubelet, runtime, and network diagnosis</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-node-not-ready/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-node-not-ready/</guid><description>&lt;h1 id="kubernetes-node-notready-kubelet-runtime--network-diagnosis">Kubernetes Node NotReady: Kubelet, Runtime &amp;amp; Network Diagnosis&lt;/h1>
&lt;p>When &lt;a href="https://www.netdata.cloud/guides/kubernetes/">a Kubernetes node&lt;/a> becomes NotReady, existing containers usually keep running, but the cluster stops scheduling new pods, removes endpoints from Services, and eventually evicts workloads after the pod eviction timeout. Root causes fall into three domains: kubelet health, container runtime responsiveness, and CNI or control plane connectivity.&lt;/p>
&lt;h2 id="what-this-means">What This Means&lt;/h2>
&lt;p>Kubernetes marks a node NotReady when the kubelet Ready condition is False, or when the node controller has not received a heartbeat within &lt;code>--node-monitor-grace-period&lt;/code> (default 40 seconds). The node receives the &lt;code>node.kubernetes.io/not-ready:NoSchedule&lt;/code> taint. If the condition persists longer than the pod eviction timeout (default 5 minutes), the controller manager marks pods on the node for rescheduling.&lt;/p></description></item><item><title>Kubernetes node PIDPressure: detection and remediation</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-node-pid-pressure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-node-pid-pressure/</guid><description>&lt;h1 id="kubernetes-node-pidpressure-detection-and-remediation">Kubernetes node PIDPressure: detection and remediation&lt;/h1>
&lt;p>PID exhaustion is a cliff-edge failure: once the kernel cannot fork, containers fail to start, health checks fail, and ssh to the node may hang. Kubernetes surfaces this through the PIDPressure node condition, but many clusters ship without PID-based eviction thresholds. Without them, the first symptom is usually &lt;code>EAGAIN&lt;/code> or &lt;code>ENOMEM&lt;/code> from fork failures, not a kubelet eviction.&lt;/p>
&lt;p>This guide shows how to detect PIDPressure before it triggers an outage, distinguish between application leaks, runtime shim accumulation, and kernel limits, and remediate the root cause. You will correlate node-level PID utilization with specific pods, validate kubelet cgroup enforcement, and configure thresholds that provide lead time.&lt;/p></description></item><item><title>Kubernetes PLEG is not healthy: runtime stalls and node degradation</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pleg-not-healthy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pleg-not-healthy/</guid><description>&lt;h1 id="kubernetes-pleg-is-not-healthy-runtime-stalls-and-node-degradation">Kubernetes PLEG is not healthy: runtime stalls and node degradation&lt;/h1>
&lt;p>A node suddenly flips to NotReady with the message &amp;ldquo;PLEG is not healthy.&amp;rdquo; Containers on the node keep running, but the control plane evicts workloads and reschedules them elsewhere. New pods cannot start, and existing pods run without health checks or status updates. This is one of the most common kubelet failure modes in production.&lt;/p>
&lt;p>The Pod Lifecycle Event Generator (PLEG) is the kubelet subsystem that polls the container runtime every second and emits events when containers start, stop, or change state. When the runtime becomes slow or unresponsive, the PLEG relist loop stalls. If the elapsed time since the last successful relist exceeds three minutes, kubelet declares PLEG unhealthy, marks the node NotReady, and skips pod synchronization.&lt;/p></description></item><item><title>Kubernetes pod CrashLoopBackOff: causes, diagnosis, and fixes</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-crashloopbackoff/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-crashloopbackoff/</guid><description>&lt;h1 id="kubernetes-pod-crashloopbackoff-causes-diagnosis--fixes">Kubernetes Pod CrashLoopBackOff: Causes, Diagnosis &amp;amp; Fixes&lt;/h1>
&lt;p>CrashLoopBackOff means a container in a Pod has terminated after starting, and the kubelet is delaying the next restart with exponential backoff. The status describes behavior, not root cause. Underlying failures include application panics, OOM kills, misconfigured liveness probes, missing secrets, or node-level resource pressure.&lt;/p>
&lt;p>Use pod status, previous container logs, node conditions, and kubelet events to narrow the cause. Monitor restart rate, node pressure, and probe failures to catch loops before they degrade capacity.&lt;/p></description></item><item><title>Kubernetes pod creation fails: admission, quota, and CRI errors</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-creation-fails/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-creation-fails/</guid><description>&lt;h1 id="kubernetes-pod-creation-fails-admission-quota-and-cri-errors">Kubernetes pod creation fails: admission, quota, and CRI errors&lt;/h1>
&lt;p>Pre-scheduling failures happen when the API server or container runtime rejects a Pod before the scheduler assigns it. You apply a Deployment, but &lt;code>kubectl get pods&lt;/code> returns nothing. Or a Pod hangs in &lt;code>ImagePullBackOff&lt;/code> before &lt;code>ContainerCreating&lt;/code>. These cases surface as missing Pods, &lt;code>FailedCreate&lt;/code> events on ReplicaSets or Jobs, or explicit API rejections. This guide covers admission control, quota and policy limits, and CRI-level image pull and sandbox failures.&lt;/p></description></item><item><title>Kubernetes pod Evicted: detection, root cause, and prevention</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-evicted/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-evicted/</guid><description>&lt;h1 id="kubernetes-pod-evicted-detection-root-cause-and-prevention">Kubernetes pod Evicted: detection, root cause, and prevention&lt;/h1>
&lt;p>Pods with status &lt;code>Evicted&lt;/code> are not application crashes. They are the kubelet&amp;rsquo;s emergency response to node-level resource pressure. When memory, disk, inodes, or PIDs approach exhaustion, the kubelet terminates pods to reclaim resources and protect node availability. The pod phase changes to &lt;code>Failed&lt;/code> with reason &lt;code>Evicted&lt;/code>, and the node reports conditions such as &lt;code>MemoryPressure&lt;/code> or &lt;code>DiskPressure&lt;/code>.&lt;/p>
&lt;p>This guide covers node-pressure eviction triggered by the kubelet, not voluntary disruption from &lt;code>kubectl drain&lt;/code> or PodDisruptionBudget enforcement.&lt;/p></description></item><item><title>Kubernetes pod exits immediately: how to diagnose it</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-exits-immediately/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-exits-immediately/</guid><description>&lt;h1 id="kubernetes-pod-exits-immediately-how-to-diagnose-it">Kubernetes pod exits immediately: how to diagnose it&lt;/h1>
&lt;p>When a pod shows &lt;code>Completed&lt;/code> or &lt;code>Error&lt;/code> with zero restarts, the container exited on its first run. The diagnostic evidence lives in termination metadata, not in a growing restart count. This is distinct from &lt;code>CrashLoopBackOff&lt;/code>, where the kubelet has already applied exponential backoff after multiple restarts.&lt;/p>
&lt;p>This guide covers how to distinguish a clean exit, an OOM kill, an application crash, and a configuration error using only the kubelet&amp;rsquo;s reported state and the previous container logs, plus which node-level and control-plane signals to check when the container produced no logs.&lt;/p></description></item><item><title>Kubernetes pod ImagePullBackOff: registry, auth, and network diagnosis</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-imagepullbackoff/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-imagepullbackoff/</guid><description>&lt;h1 id="kubernetes-pod-imagepullbackoff-registry-auth-and-network-diagnosis">Kubernetes pod ImagePullBackOff: registry, auth, and network diagnosis&lt;/h1>
&lt;p>&lt;code>ImagePullBackOff&lt;/code> means the kubelet cannot pull a required image. After each &lt;code>ErrImagePull&lt;/code> failure, the kubelet retries with exponential backoff capped at five minutes. When &lt;code>serializeImagePulls&lt;/code> is true, a single slow pull blocks every subsequent pull on that node. Read the exact error from the CRI in pod events, test the registry directly from the node, and fix the root cause without blindly recreating pods.&lt;/p></description></item><item><title>Kubernetes pod liveness probe killing healthy containers</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-liveness-probe-killing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-liveness-probe-killing/</guid><description>&lt;h1 id="kubernetes-pod-liveness-probe-killing-healthy-containers">Kubernetes pod liveness probe killing healthy containers&lt;/h1>
&lt;p>A container that is processing requests, not OOMKilled, and not crashed can still be restarted repeatedly by the kubelet because a liveness probe failed. The application is alive, but the probe says it is not. This usually shows up as a pod stuck in &lt;code>CrashLoopBackOff&lt;/code> with &lt;code>Liveness probe failed&lt;/code> events, even though application logs show no fatal error. The restarts waste resources, break active connections, and can trigger cascading load on the cluster as other pods absorb the shifted traffic.&lt;/p></description></item><item><title>Kubernetes pod network isolation: when one node loses pod connectivity</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-network-isolation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-network-isolation/</guid><description>&lt;h1 id="kubernetes-pod-network-isolation-when-one-node-loses-pod-connectivity">Kubernetes pod network isolation: when one node loses pod connectivity&lt;/h1>
&lt;p>One node in your cluster drops off the pod network. Pods still show Running, but they cannot reach Services, other pods, or external endpoints. The rest of the cluster is unaffected. This is single-node pod network isolation, and it is almost always a node-local CNI or data path failure.&lt;/p>
&lt;p>Symptoms are subtle compared to a full cluster outage. Applications log connection timeouts. Health checks fail. New pods stick in ContainerCreating. Because the node often stays Ready, operators may blame the application rather than the network layer.&lt;/p></description></item><item><title>Kubernetes pod OOMKilled: cgroup limits, evictions, and fixes</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-oomkilled/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-oomkilled/</guid><description>&lt;h1 id="kubernetes-pod-oomkilled-cgroup-limits-evictions-and-fixes">Kubernetes pod OOMKilled: cgroup limits, evictions, and fixes&lt;/h1>
&lt;p>A pod status of &lt;code>OOMKilled&lt;/code> means the container restarted after the kernel sent SIGKILL because it could not satisfy a memory allocation. There is no graceful shutdown.&lt;/p>
&lt;p>Distinguish whether the kill happened at the container cgroup level (a limit you set) or at the node level (a system-wide shortage). Then separate kernel OOM kills from kubelet evictions, identify the correct fix, and prevent recurrence without guessing at memory limits.&lt;/p></description></item><item><title>Kubernetes pod readiness probe failures: traffic exclusion and debug</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-readiness-probe-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-readiness-probe-failures/</guid><description>&lt;h1 id="kubernetes-pod-readiness-probe-failures-traffic-exclusion-and-debug">Kubernetes pod readiness probe failures: traffic exclusion and debug&lt;/h1>
&lt;p>You check a Deployment and see pods in &lt;code>Running&lt;/code> phase but not &lt;code>Ready&lt;/code>. Traffic to the Service drops. The containers have not restarted. This is a readiness probe failure. Unlike liveness, readiness does not restart the container. It removes the pod from Service endpoints. The container may be starting slowly, temporarily overloaded, or waiting on a dependency that should not be in the probe path. This guide explains how Kubernetes excludes traffic, how to distinguish readiness from liveness failures, and how to debug the root cause.&lt;/p></description></item><item><title>Kubernetes Pod Security Standards violations: detection and remediation</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-security-violations/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-security-violations/</guid><description>&lt;h1 id="kubernetes-pod-security-standards-violations-detection-and-remediation">Kubernetes Pod Security Standards violations: detection and remediation&lt;/h1>
&lt;p>You apply a Deployment and the ReplicaSet creates zero pods. You roll out a DaemonSet update and new pods vanish without events. You add a namespace label to improve security posture, and CI pipelines fail with opaque admission errors. Pod Security Admission (PSA) blocks non-compliant workloads silently when &lt;code>enforce&lt;/code> mode is active. To fix these violations, correlate namespace labels with enforcement outcomes, parse audit log annotations into exact spec changes, distinguish admission-time rejections from runtime failures, and prevent enforcement from silently blocking future rollouts.&lt;/p></description></item><item><title>Kubernetes pod stuck ContainerCreating: volume, network, and image issues</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-containercreating/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-containercreating/</guid><description>&lt;h1 id="kubernetes-pod-stuck-containercreating-volume-network-and-image-issues">Kubernetes pod stuck ContainerCreating: volume, network, and image issues&lt;/h1>
&lt;p>A pod stuck in &lt;code>ContainerCreating&lt;/code> never produces logs or readiness events. The kubelet accepted the spec but blocked during initialization after scheduling and before the container runtime starts the user process. The dominant failure domains are volume mount deadlocks, CNI sandbox creation failures, and image pull problems. They all surface the same status but need different fixes. This guide shows how to identify the stuck subsystem and resolve it.&lt;/p></description></item><item><title>Kubernetes pod stuck on volume mount: CSI, permissions, and timeouts</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-volume-mount-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-volume-mount-failures/</guid><description>&lt;h1 id="kubernetes-pod-stuck-on-volume-mount-csi-permissions-and-timeouts">Kubernetes pod stuck on volume mount: CSI, permissions, and timeouts&lt;/h1>
&lt;p>You scale a StatefulSet and the new pods sit in &lt;code>ContainerCreating&lt;/code> for ten minutes. The node is &lt;code>Ready&lt;/code>. The CSI driver pods are running. &lt;code>kubectl describe&lt;/code> shows no &lt;code>FailedMount&lt;/code> events, yet the containers never start. The absence of volume events often misleads operators into checking image registries or resource quotas instead of the storage path. The kubelet volume manager is blocked somewhere between attach and mount, and Kubernetes will not retry fast enough to hide the problem.&lt;/p></description></item><item><title>Kubernetes pod stuck Pending: scheduling failures explained</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-pending/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-pending/</guid><description>&lt;h1 id="kubernetes-pod-stuck-pending-scheduling-failures-explained">Kubernetes pod stuck Pending: scheduling failures explained&lt;/h1>
&lt;p>A Deployment scales up and the new replicas stay Pending. No containers start, no logs appear, and &lt;code>kubectl logs&lt;/code> returns nothing because no node is assigned yet. When a pod is stuck in Pending, the scheduler has either not yet evaluated it or has rejected every candidate node. Containers cannot start until the pod is assigned, so this blocks rollouts, autoscaling, and recovery.&lt;/p></description></item><item><title>Kubernetes pod stuck Terminating: finalizers, grace periods, and force delete</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-stuck-terminating/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pod-stuck-terminating/</guid><description>&lt;h1 id="kubernetes-pod-stuck-terminating-finalizers-grace-periods--force-delete">Kubernetes Pod Stuck Terminating: Finalizers, Grace Periods &amp;amp; Force Delete&lt;/h1>
&lt;p>A pod stuck in &lt;code>Terminating&lt;/code> stays visible in the &lt;a href="https://www.netdata.cloud/guides/kubernetes/kubernetes-api-server-slow/">API server&lt;/a> after &lt;code>kubectl delete&lt;/code>, sometimes for minutes or hours. Usually a finalizer blocks removal, a CSI volume is still attached, or the container is ignoring SIGTERM. Force deleting without diagnosis orphans containers and can violate StatefulSet guarantees. Check the signals first, then decide whether to wait, patch a finalizer, or force delete.&lt;/p></description></item><item><title>Kubernetes priority class evictions: how preemption actually works</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-priority-class-eviction/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-priority-class-eviction/</guid><description>&lt;h1 id="kubernetes-priority-class-evictions-how-preemption-actually-works">Kubernetes priority class evictions: how preemption actually works&lt;/h1>
&lt;p>Scheduler preemption and kubelet node-pressure eviction both use Pod Priority, but they follow different rules, interact differently with PodDisruptionBudgets, and produce different failure signatures. Misunderstanding the two mechanisms leads to critical workloads being terminated unexpectedly, high-priority pods staying pending while lower-priority pods run, or system pods being preempted after a PriorityClass change.&lt;/p>
&lt;p>The scheduler preempts lower-priority pods to make room for a higher-priority pending pod. The kubelet evicts pods when a node is under resource pressure. A PodDisruptionBudget protects against voluntary disruptions, but the scheduler treats PDB violations during preemption as best-effort, not guaranteed.&lt;/p></description></item><item><title>Kubernetes PV reclaim policies: Retain, Delete, Recycle in practice</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pv-reclaim-policy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pv-reclaim-policy/</guid><description>&lt;h1 id="kubernetes-pv-reclaim-policies-retain-delete-recycle-in-practice">Kubernetes PV reclaim policies: Retain, Delete, Recycle in practice&lt;/h1>
&lt;p>Deleting a PVC does not always delete the underlying storage. If the PV enters Released state and the cloud bill keeps growing, or the application fails after redeployment because the PV is still bound to a deleted claim, the reclaim policy is misaligned with operator intent. This guide explains what happens when a bound PVC is deleted under each reclaim policy, how to recover a Retained volume, why Delete can orphan cloud resources, and why Recycle should not be used in modern clusters. The focus is on practical checks, CSI-specific gotchas, and preventing data loss.&lt;/p></description></item><item><title>Kubernetes PV volumeBindingMode WaitForFirstConsumer: when and why</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pv-binding-mode-late/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pv-binding-mode-late/</guid><description>&lt;h1 id="kubernetes-pv-volumebindingmode-waitforfirstconsumer-when-and-why">Kubernetes PV volumeBindingMode WaitForFirstConsumer: when and why&lt;/h1>
&lt;p>A PersistentVolumeClaim that stays Pending after creation often triggers an incident response. When the StorageClass uses &lt;code>volumeBindingMode: WaitForFirstConsumer&lt;/code>, that Pending state is usually intentional. The cluster defers provisioning and binding until a Pod that references the PVC is created and the scheduler tentatively selects a node. This article explains what Immediate and WaitForFirstConsumer mean, why Kubernetes defers binding, where the mechanism appears in production, and the operational tradeoffs you accept when you choose one mode over the other.&lt;/p></description></item><item><title>Kubernetes PVC stuck Pending: storage class, provisioner, and quota</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-pvc-stuck-pending/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-pvc-stuck-pending/</guid><description>&lt;h1 id="kubernetes-pvc-stuck-pending-storage-class-provisioner-and-quota">Kubernetes PVC stuck Pending: storage class, provisioner, and quota&lt;/h1>
&lt;p>A PersistentVolumeClaim stuck in &lt;code>Pending&lt;/code> is a storage-layer failure. Unlike a pod that is &lt;code>Pending&lt;/code> due to CPU or memory, a PVC in &lt;code>Pending&lt;/code> means the cluster cannot provision the volume. The cause usually sits in one of three layers: the StorageClass and its provisioner, namespace-level ResourceQuota limits, or topology and cloud constraints that prevent volume creation and attachment.&lt;/p>
&lt;p>When a PVC stays unbound, dependent pods cannot start. StatefulSets stall, rolling updates hang, and storage-dependent applications remain offline.&lt;/p></description></item><item><title>Kubernetes RBAC permission denied: detection and minimum-permission fix</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-rbac-permission-denied/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-rbac-permission-denied/</guid><description>&lt;h1 id="kubernetes-rbac-permission-denied-detection-and-minimum-permission-fix">Kubernetes RBAC permission denied: detection and minimum-permission fix&lt;/h1>
&lt;p>A 403 Forbidden from the Kubernetes API server means the caller was authenticated but RBAC refused the action. In production, this appears as a Deployment stuck creating pods, a CI pipeline failing to patch a ConfigMap, a controller logging repeated forbidden errors, or an operator unable to finalize a custom resource. One missing verb on one resource in one namespace blocks the entire workflow.&lt;/p></description></item><item><title>Kubernetes ResourceQuota exceeded: detection and remediation</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-resource-quota-exceeded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-resource-quota-exceeded/</guid><description>&lt;h1 id="kubernetes-resourcequota-exceeded-detection-and-remediation">Kubernetes ResourceQuota exceeded: detection and remediation&lt;/h1>
&lt;p>A Deployment looks healthy in &lt;code>kubectl get deployment&lt;/code> but the new ReplicaSet has zero pods. A Job is accepted but never creates a pod. A CI/CD pipeline fails with opaque 403 errors from the API server. These symptoms point to ResourceQuota exhaustion. Confirm the quota is the blocker, identify the exhausted resource, and fix it.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>ResourceQuota is a namespace-scoped admission controller that enforces hard aggregate limits on resource consumption. When a tracked resource hits its hard limit, the API server rejects subsequent create requests with HTTP 403 Forbidden and a message containing &lt;code>exceeded quota: &amp;lt;quota-name&amp;gt;&lt;/code>. Existing pods keep running; quota is enforced at admission time, not by terminating workloads.&lt;/p></description></item><item><title>Kubernetes scheduler not scheduling pods: queue depth and failure reasons</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-scheduler-not-scheduling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-scheduler-not-scheduling/</guid><description>&lt;h1 id="kubernetes-scheduler-not-scheduling-pods-queue-depth-and-failure-reasons">Kubernetes scheduler not scheduling pods: queue depth and failure reasons&lt;/h1>
&lt;p>Pods stay Pending for many reasons, but the scheduler process being down is rarely one. More often, pods accumulate in internal queues because the cluster is out of capacity, a control plane dependency stalls the binding cycle, or a filter plugin rejects every candidate node. Distinguishing &amp;ldquo;unschedulable&amp;rdquo; (no node fits) from &amp;ldquo;not scheduling&amp;rdquo; (the scheduler cannot keep up or the binding cycle is failing) prevents wasted node scaling when the real problem is an etcd latency spike or a volume affinity conflict.&lt;/p></description></item><item><title>Kubernetes secrets mount failures: detection and recovery</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-secrets-mount-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-secrets-mount-failures/</guid><description>&lt;h1 id="kubernetes-secrets-mount-failures-detection-and-recovery">Kubernetes secrets mount failures: detection and recovery&lt;/h1>
&lt;p>Pods that reference a Secret volume can hang in &lt;code>ContainerCreating&lt;/code> for minutes, crash on startup with missing files, or silently run with stale credentials. The kubelet mounts Secret data as a tmpfs-backed volume and validates the Secret object and its requested keys before any container starts. If the Secret does not exist, if a specific key is absent, or if the kubelet cannot reach the API server, the mount fails. Non-optional mounts block every container from starting. Optional mounts allow startup but leave the mount point empty, which can cause failures later.&lt;/p></description></item><item><title>Kubernetes Service Not Reachable: Kube-Proxy, Endpoints &amp; DNS</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-service-not-reachable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-service-not-reachable/</guid><description>&lt;p>A Service fails when the chain between the client and backend breaks. That chain depends on EndpointSlices to list healthy pods, kube-proxy to program kernel rules, and cluster DNS to resolve names to ClusterIPs. This guide covers the gap between healthy backend pods and an unreachable Service, focusing on kube-proxy data-plane programming, endpoint state, and DNS dependencies. It does not cover application-level bugs inside the pod.&lt;/p>
&lt;h2 id="what-this-means">What This Means&lt;/h2>
&lt;p>Reachability is a control-loop problem. kube-proxy watches Services and EndpointSlices, then programs iptables, IPVS, or nftables rules to DNAT traffic to healthy endpoints. These rules persist in the kernel if kube-proxy crashes, but updates stop until it resumes. DNS resolution targets the CoreDNS ClusterIP, so a kube-proxy failure often appears as a DNS outage before a direct Service timeout.&lt;/p></description></item><item><title>Kubernetes stale conntrack during rolling updates: intermittent connection resets</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-stale-conntrack-rolling-update/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-stale-conntrack-rolling-update/</guid><description>&lt;h1 id="kubernetes-stale-conntrack-during-rolling-updates-intermittent-connection-resets">Kubernetes stale conntrack during rolling updates: intermittent connection resets&lt;/h1>
&lt;p>You deploy a new version of a service. The rollout reports healthy. Error budgets look fine. Then you notice sporadic connection resets in application logs, a handful of timeout errors, or brief latency spikes that correlate exactly with pod terminations. The failures are intermittent, last only seconds, and never trigger a full outage. This is the stale conntrack race during rolling updates.&lt;/p></description></item><item><title>Kubernetes StatefulSet pod not ready: ordering, PVCs, and recovery</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-statefulset-pod-not-ready/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-statefulset-pod-not-ready/</guid><description>&lt;h1 id="kubernetes-statefulset-pod-not-ready-ordering-pvcs-and-recovery">Kubernetes StatefulSet pod not ready: ordering, PVCs, and recovery&lt;/h1>
&lt;p>A StatefulSet pod stuck at 0/1 Ready, Pending, or Unknown blocks the entire chain when the default OrderedReady policy is in effect. The controller creates pods sequentially from ordinal 0 to N-1, so one failure halts all higher ordinals. A volumeClaimTemplate adds dependencies on storage binding, node affinity, and CSI attachment that can outlive the pod. Force-deleting a pod risks violating at-most-one semantics and causing split-brain in quorum-sensitive workloads.&lt;/p></description></item><item><title>Kubernetes volume snapshot failures: VolumeSnapshot CRDs and CSI snapshotter</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-volume-snapshot-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-volume-snapshot-failures/</guid><description>&lt;h1 id="kubernetes-volume-snapshot-failures-volumesnapshot-crds-and-csi-snapshotter">Kubernetes volume snapshot failures: VolumeSnapshot CRDs and CSI snapshotter&lt;/h1>
&lt;p>A VolumeSnapshot that never becomes &lt;code>readyToUse&lt;/code>, a restore yielding an empty PVC, or a CSI snapshotter sidecar logging GRPC errors with no cluster progress all indicate the same problem: the snapshot pipeline broke between the API server and the storage backend.&lt;/p>
&lt;p>This guide covers failure modes between applying a VolumeSnapshot manifest and the storage backend completing the operation. It covers verifying CRDs, the snapshot-controller, the CSI snapshotter sidecar, driver names, capabilities, and secrets, and reading the error signatures each misalignment produces.&lt;/p></description></item><item><title>Kyocera Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/kyocera-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/kyocera-corporation-snmp-traps/</guid><description/></item><item><title>Kyocera Printer</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/kyocera-printer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/kyocera-printer/</guid><description/></item><item><title>Lanart Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lanart-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lanart-corp-snmp-traps/</guid><description/></item><item><title>Lancast Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lancast-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lancast-inc-snmp-traps/</guid><description/></item><item><title>Lancity Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lancity-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lancity-corporation-snmp-traps/</guid><description/></item><item><title>Lancom Systems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lancom-systems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lancom-systems-snmp-traps/</guid><description/></item><item><title>Lanex Sp Z O O SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lanex-sp-z-o-o-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lanex-sp-z-o-o-snmp-traps/</guid><description/></item><item><title>Lannair Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lannair-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lannair-ltd-snmp-traps/</guid><description/></item><item><title>Lannet Company SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lannet-company-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lannet-company-snmp-traps/</guid><description/></item><item><title>Lantronix SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lantronix-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lantronix-snmp-traps/</guid><description/></item><item><title>Last Mile Gear SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/last-mile-gear-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/last-mile-gear-snmp-traps/</guid><description/></item><item><title>Latitude Communications SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/latitude-communications-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/latitude-communications-snmp-traps/</guid><description/></item><item><title>Laurel Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/laurel-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/laurel-networks-inc-snmp-traps/</guid><description/></item><item><title>Lefthand Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lefthand-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lefthand-networks-snmp-traps/</guid><description/></item><item><title>Lenovo Enterprise Business Group SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lenovo-enterprise-business-group-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lenovo-enterprise-business-group-snmp-traps/</guid><description/></item><item><title>Lenovoemc Ltd Formerly Iomega SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lenovoemc-ltd-formerly-iomega-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lenovoemc-ltd-formerly-iomega-snmp-traps/</guid><description/></item><item><title>Lexmark International SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lexmark-international-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lexmark-international-snmp-traps/</guid><description/></item><item><title>Librenms SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/librenms-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/librenms-snmp-traps/</guid><description/></item><item><title>Libreswan</title><link>https://www.netdata.cloud/integrations/data-collection/networking/libreswan/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/libreswan/</guid><description/></item><item><title>Libvirt VMs and Containers</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/libvirt-vms-and-containers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/libvirt-vms-and-containers/</guid><description/></item><item><title>License expiry silently disabling features: monitor days-to-expiry</title><link>https://www.netdata.cloud/guides/network/network-license-expiry-silent-disable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-license-expiry-silent-disable/</guid><description>&lt;h1 id="license-expiry-silently-disabling-features-monitor-days-to-expiry">License expiry silently disabling features: monitor days-to-expiry&lt;/h1>
&lt;p>Your firewall dashboard shows green. Interfaces are up, CPU and memory are normal, traffic is flowing. But at 09:00, someone reports VPN connections failing, IPS no longer blocking threats, or URL filtering not enforcing policy. A feature license expired at midnight, and the device silently stopped performing the licensed function without raising a visible alarm.&lt;/p>
&lt;p>The device stays up, counters keep incrementing, throughput looks normal. The license-expiry message in syslog is low severity and gets buried under routine noise. By the time someone notices, the feature has been disabled for hours.&lt;/p></description></item><item><title>Lighttpd</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/lighttpd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/lighttpd/</guid><description/></item><item><title>Lighttpd Monitoring</title><link>https://www.netdata.cloud/monitoring-101/lighttpd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/lighttpd-monitoring/</guid><description>&lt;h2 id="lighttpd-monitoring">Lighttpd Monitoring&lt;/h2>
&lt;p>Welcome to the comprehensive guide on monitoring Lighttpd, a flexible and lightweight web server. This guide will take you through everything you need to know about monitoring Lighttpd using the Netdata platform. From understanding essential metrics to employ advanced performance monitoring techniques, Netdata provides a powerful Lighttpd monitoring tool for all your DevOps, SRE, and IT admin needs.&lt;/p>
&lt;h3 id="what-is-lighttpd">What Is Lighttpd?&lt;/h3>
&lt;p>&lt;a href="https://www.lighttpd.net/">Lighttpd&lt;/a> is an open-source web server optimized for speed-critical environments while maintaining a low memory footprint. Its feature set and performance make it a popular choice for many developers and engineers looking to implement efficient web solutions.&lt;/p></description></item><item><title>Ligowave SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ligowave-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ligowave-snmp-traps/</guid><description/></item><item><title>Linksys</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/linksys/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/linksys/</guid><description/></item><item><title>Linksys SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/linksys-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/linksys-snmp-traps/</guid><description/></item><item><title>Linode</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/linode/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/linode/</guid><description/></item><item><title>Linode Monitoring</title><link>https://www.netdata.cloud/monitoring-101/linode-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/linode-monitoring/</guid><description>&lt;h2 id="linode-monitoring">Linode Monitoring&lt;/h2>
&lt;h3 id="what-is-linode">What Is Linode?&lt;/h3>
&lt;p>Linode is a cloud hosting service that provides virtual servers to deploy applications, manage web hosting, and scale infrastructure as demands grow. Known for its simplicity and cost-effectiveness, Linode empowers developers, DevOps teams, and IT Administrators to manage cloud-based services with flexibility and efficiency.&lt;/p>
&lt;h3 id="monitoring-linode-with-netdata">Monitoring Linode With Netdata&lt;/h3>
&lt;p>Monitoring Linode is crucial to ensure optimal performance, cost management, and resource allocation. Netdata serves as an effective Linode monitoring tool, leveraging an openmetrics (Prometheus) exporter to collect real-time data without the need for standalone Prometheus or Grafana servers. This integration provides you with automated dashboards, alerts, and comprehensive insights into your Linode instances.&lt;/p></description></item><item><title>Linux</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/linux/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/linux/</guid><description/></item><item><title>Linux Audit Subsystem</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/linux-audit-subsystem/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/linux-audit-subsystem/</guid><description/></item><item><title>Linux Hardware Sensors (libsensors)</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/linux-hardware-sensors-libsensors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/linux-hardware-sensors-libsensors/</guid><description/></item><item><title>Linux kernel SLAB allocator statistics</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/linux-kernel-slab-allocator-statistics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/linux-kernel-slab-allocator-statistics/</guid><description/></item><item><title>Linux Sensors</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/linux-sensors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/linux-sensors/</guid><description/></item><item><title>Linux Sensors Monitoring</title><link>https://www.netdata.cloud/monitoring-101/sensors-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/sensors-monitoring/</guid><description>&lt;h2 id="linux-sensors-monitoring">Linux Sensors Monitoring&lt;/h2>
&lt;h3 id="what-is-linux-sensors">What Is Linux Sensors?&lt;/h3>
&lt;p>Linux Sensors is a comprehensive suite for monitoring hardware sensors in Linux systems, gathering real-time data about temperature, voltage, fan speed, energy consumption, and more. Leveraging the sysfs interface, it provides a robust way for systems and applications to access sensor data, crucial for maintaining optimal server performance and health.&lt;/p>
&lt;h3 id="monitoring-linux-sensors-with-netdata">Monitoring Linux Sensors With Netdata&lt;/h3>
&lt;p>Netdata is a powerful, real-time Linux sensors monitoring tool that allows you to effortlessly monitor Linux Sensors. By utilizing the &lt;a href="https://www.kernel.org/doc/Documentation/hwmon/sysfs-interface">sysfs interface&lt;/a>, Netdata automatically detects all available sensors on your system, providing you with instant visibility into sensor metrics such as temperature, voltage, current, and power usage.&lt;/p></description></item><item><title>Linux ZSwap</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/linux-zswap/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/linux-zswap/</guid><description/></item><item><title>Litespeed</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/litespeed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/litespeed/</guid><description/></item><item><title>Litespeed Monitoring</title><link>https://www.netdata.cloud/monitoring-101/litespeed-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/litespeed-monitoring/</guid><description>&lt;h2 id="litespeed-monitoring">Litespeed Monitoring&lt;/h2>
&lt;h3 id="what-is-litespeed">What Is Litespeed?&lt;/h3>
&lt;p>Litespeed is a powerful web server technology that boosts website speed and security. It is renowned for its high-performance capabilities, delivering superior HTTP/HTTPS content and optimizing traffic handling. Discover more about &lt;a href="https://www.litespeedtech.com/products/litespeed-web-server">Litespeed&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-litespeed-with-netdata">Monitoring Litespeed With Netdata&lt;/h3>
&lt;p>Utilizing Netdata to monitor Litespeed provides comprehensive insights into your web server&amp;rsquo;s performance. As a real-time, distributed monitoring tool, Netdata captures extensive metrics, ensuring your server operates efficiently and responds proactively to issues. You can begin monitoring Litespeed effortlessly with &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/litespeed/">Netdata&amp;rsquo;s Litespeed Monitoring Tool&lt;/a>.&lt;/p></description></item><item><title>Live Network Connections</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/live-network-connections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/live-network-connections/</guid><description/></item><item><title>Livingston Enterprises Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/livingston-enterprises-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/livingston-enterprises-inc-snmp-traps/</guid><description/></item><item><title>LLDP Topology</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/lldp-topology/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/lldp-topology/</guid><description/></item><item><title>Local listening processes</title><link>https://www.netdata.cloud/integrations/all/local-listening-processes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/all/local-listening-processes/</guid><description/></item><item><title>Locating endpoints behind NAT and wireless: the positioning problem</title><link>https://www.netdata.cloud/guides/network/network-endpoint-positioning-nat-wifi/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-endpoint-positioning-nat-wifi/</guid><description>&lt;h1 id="locating-endpoints-behind-nat-and-wireless-the-positioning-problem">Locating endpoints behind NAT and wireless: the positioning problem&lt;/h1>
&lt;p>Endpoint positioning maps a MAC address or IP to a specific switch port, access point, or VLAN. It underpins security investigations, access control enforcement, and day-to-day troubleshooting. When the endpoint sits behind a NAT boundary, the Layer 2 and Layer 3 signals that topology engines rely on (FDB entries, ARP tables, flow records) all report the NAT device&amp;rsquo;s identity, not the endpoint behind it. The endpoint becomes operationally invisible upstream.&lt;/p></description></item><item><title>Logstash</title><link>https://www.netdata.cloud/integrations/data-collection/applications/logstash/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/logstash/</guid><description/></item><item><title>Logstash _dateparsefailure: timestamp formats that stop parsing</title><link>https://www.netdata.cloud/guides/logstash/logstash-dateparsefailure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-dateparsefailure/</guid><description>&lt;h1 id="logstash-_dateparsefailure-timestamp-formats-that-stop-parsing">Logstash _dateparsefailure: timestamp formats that stop parsing&lt;/h1>
&lt;p>The &lt;code>_dateparsefailure&lt;/code> tag appears when Logstash&amp;rsquo;s date filter cannot parse a timestamp field against any of its configured &lt;code>match&lt;/code> patterns. The event still flows through the pipeline and reaches its output. &lt;code>@timestamp&lt;/code> stays at ingestion time instead of being set to the event&amp;rsquo;s actual time.&lt;/p>
&lt;p>The result: time-based dashboards show gaps or misplaced data points, index lifecycle management operates on ingestion time so events land in the wrong time bucket or expire at the wrong time, and the issue can persist for days because throughput metrics stay green. The date filter tries each format string in the &lt;code>match&lt;/code> array sequentially. The first successful parse wins. If none succeed, the filter appends its &lt;code>tag_on_failure&lt;/code> tag (default &lt;code>_dateparsefailure&lt;/code>) and leaves &lt;code>@timestamp&lt;/code> unchanged.&lt;/p></description></item><item><title>Logstash _grokparsefailure: why grok stops matching and how to fix it</title><link>https://www.netdata.cloud/guides/logstash/logstash-grokparsefailure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-grokparsefailure/</guid><description>&lt;h1 id="logstash-_grokparsefailure-why-grok-stops-matching-and-how-to-fix-it">Logstash _grokparsefailure: why grok stops matching and how to fix it&lt;/h1>
&lt;p>Events tagged with &lt;code>_grokparsefailure&lt;/code> pass through your pipeline unstructured. Grok could not match any of its configured patterns against the event&amp;rsquo;s input field, so it appended the failure tag and moved on. The event is not dropped. It is not counted as filtered. Unless you conditionally route it, it reaches your output destination carrying raw, unparsed data alongside the failure tag.&lt;/p></description></item><item><title>Logstash _jsonparsefailure: malformed JSON and codec mismatches</title><link>https://www.netdata.cloud/guides/logstash/logstash-jsonparsefailure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-jsonparsefailure/</guid><description>&lt;h1 id="logstash-_jsonparsefailure-malformed-json-and-codec-mismatches">Logstash _jsonparsefailure: malformed JSON and codec mismatches&lt;/h1>
&lt;p>When events arrive at your Elasticsearch indices carrying the &lt;code>_jsonparsefailure&lt;/code> tag, they contain raw, unparsed data instead of the structured fields your downstream consumers expect. Throughput metrics look healthy, events-out counts keep climbing, and no alerts fire. But the data is wrong.&lt;/p>
&lt;p>The &lt;code>_jsonparsefailure&lt;/code> tag is added by either the json codec or the json filter when it receives input it cannot parse as valid JSON. The event is not dropped. It flows through to the output carrying whatever raw data was received, plus the failure tag. This is a silent correctness failure, not an availability failure.&lt;/p></description></item><item><title>Logstash address already in use: input port conflicts on Beats, TCP, and HTTP</title><link>https://www.netdata.cloud/guides/logstash/logstash-address-already-in-use/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-address-already-in-use/</guid><description>&lt;h1 id="logstash-address-already-in-use-input-port-conflicts-on-beats-tcp-and-http">Logstash address already in use: input port conflicts on Beats, TCP, and HTTP&lt;/h1>
&lt;p>Logstash logs &lt;code>Address already in use&lt;/code> (a Java &lt;code>BindException&lt;/code>) when a listening input plugin, typically the Beats input on 5044, a tcp, http, or syslog input, tries to bind a port that is already taken. The bind fails, the input cannot start, and the pipeline that owns it stops ingesting.&lt;/p>
&lt;p>The blast radius is the problem. In a multi-pipeline deployment the JVM usually stays alive and the other pipelines keep running. Process-level checks stay green, the monitoring API answers, and one pipeline silently stops ingesting. Upstream Filebeat agents queue, syslog senders drop, and nobody notices until someone asks where a stream of logs went.&lt;/p></description></item><item><title>Logstash API unreachable on port 9600: crash, GC pause, or startup</title><link>https://www.netdata.cloud/guides/logstash/logstash-api-unreachable-9600/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-api-unreachable-9600/</guid><description>&lt;h1 id="logstash-api-unreachable-on-port-9600-crash-gc-pause-or-startup">Logstash API unreachable on port 9600: crash, GC pause, or startup&lt;/h1>
&lt;p>Your liveness check against &lt;code>http://127.0.0.1:9600/&lt;/code> just started timing out or returning connection refused. That single fact tells you almost nothing by itself. The Logstash monitoring API shares the JVM with the pipeline, so an unreachable API means one of four things: the process is dead, the JVM is frozen in a garbage collection pause, Logstash is still starting up, or something is wrong with how the API is bound to the network.&lt;/p></description></item><item><title>Logstash Beats input: Filebeat backpressure and connection health</title><link>https://www.netdata.cloud/guides/logstash/logstash-beats-input-backpressure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-beats-input-backpressure/</guid><description>&lt;h1 id="logstash-beats-input-filebeat-backpressure-and-connection-health">Logstash Beats input: Filebeat backpressure and connection health&lt;/h1>
&lt;p>Filebeat has stopped delivering logs. Its harvesters are running, files are being read, but the registry on the Filebeat host has not advanced in twenty minutes and downstream dashboards are going stale. Or the opposite variant: Filebeat&amp;rsquo;s log is a rotating wall of &amp;ldquo;Failed to publish events&amp;rdquo; and &amp;ldquo;connection reset by peer&amp;rdquo; while Logstash looks perfectly healthy on process checks.&lt;/p>
&lt;p>The Beats input is the front door to Logstash. Filebeat speaks the Lumberjack protocol over TCP, conventionally on port 5044, and that connection is where Logstash-side trouble shows up first. When Logstash cannot drain its queue fast enough, the backpressure does not stay inside Logstash. It propagates out the Beats input, across the TCP connection, and into Filebeat, where it appears as a stalled registry, growing disk spool, or publish errors.&lt;/p></description></item><item><title>Logstash certificate expiry: the silent, total outage no built-in metric shows</title><link>https://www.netdata.cloud/guides/logstash/logstash-certificate-expiry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-certificate-expiry/</guid><description>&lt;h1 id="logstash-certificate-expiry-the-silent-total-outage-no-built-in-metric-shows">Logstash certificate expiry: the silent, total outage no built-in metric shows&lt;/h1>
&lt;p>Every TLS connection into or out of Logstash depends on a certificate with a hard expiry date. When that date passes, the failure is not gradual: every new TLS handshake is rejected, every client disconnects, and every output stalls at the exact second the certificate lapses. Filebeat agents stop shipping. Elasticsearch outputs throw handshake errors. Nothing is delivered.&lt;/p>
&lt;p>The dangerous part is what Logstash does not tell you. The Node Stats API exposes JVM, pipeline, event, queue, and plugin metrics, but there is no certificate expiry date, no days-until-expiry gauge, and no TLS handshake failure counter anywhere in it. Your dashboards can be completely green at T-minus one hour and show a total ingestion outage at T-plus zero. Certificate expiry is a property of files on disk, so it needs an external check, not a Logstash metric.&lt;/p></description></item><item><title>Logstash config reload failed: reloads.failures and invisible configuration drift</title><link>https://www.netdata.cloud/guides/logstash/logstash-config-reload-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-config-reload-failed/</guid><description>&lt;h1 id="logstash-config-reload-failed-reloadsfailures-and-invisible-configuration-drift">Logstash config reload failed: reloads.failures and invisible configuration drift&lt;/h1>
&lt;p>You pushed a config change. The reload counter incremented, the process never restarted, and nothing paged. Two days later someone notices the new filter logic is not in the data. The deploy &amp;ldquo;worked&amp;rdquo; in every system that tracks deploys, but Logstash is still running the old pipeline.&lt;/p>
&lt;p>With &lt;code>config.reload.automatic&lt;/code> enabled, Logstash polls config files, validates changes, and hot-swaps the pipeline. When validation fails, it keeps the old pipeline running. That is the safe choice, but it means the config on disk and the config in memory diverge with no process-level symptom. The only evidence lives in three counters almost nobody graphs: &lt;code>reloads.failures&lt;/code>, &lt;code>reloads.last_error&lt;/code>, and &lt;code>reloads.successes&lt;/code>.&lt;/p></description></item><item><title>Logstash configuration drift: when the running config no longer matches the deployed one</title><link>https://www.netdata.cloud/guides/logstash/logstash-config-drift-after-reload/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-config-drift-after-reload/</guid><description>&lt;h1 id="logstash-configuration-drift-when-the-running-config-no-longer-matches-the-deployed-one">Logstash configuration drift: when the running config no longer matches the deployed one&lt;/h1>
&lt;p>You pushed a config fix to &lt;code>/etc/logstash/conf.d/&lt;/code> an hour ago. The pipeline is running, throughput is normal, no alert fired. But the fix never took effect: the reload failed validation, Logstash kept the old pipeline running, and nothing in your standard monitoring noticed. The running config and the deployed config are now two different things.&lt;/p>
&lt;p>This is Logstash configuration drift, and it is dangerous because it is invisible by default. Logstash has no built-in drift detection: there is no API endpoint that returns the currently active config text or a hash of it. The reload machinery is deliberately safe (a failed reload keeps the old pipeline alive rather than dropping it), but that safety creates a silent gap between what you think is deployed and what is actually processing your events.&lt;/p></description></item><item><title>Logstash configuration integrity: detecting unexpected changes to pipeline files</title><link>https://www.netdata.cloud/guides/logstash/logstash-config-integrity-changes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-config-integrity-changes/</guid><description>&lt;h1 id="logstash-configuration-integrity-detecting-unexpected-changes-to-pipeline-files">Logstash configuration integrity: detecting unexpected changes to pipeline files&lt;/h1>
&lt;p>Logstash has no built-in file integrity monitoring. It will load, reload, and run whatever it finds in &lt;code>/etc/logstash/conf.d/&lt;/code>, &lt;code>pipelines.yml&lt;/code>, and &lt;code>logstash.yml&lt;/code>, and it will not tell you that those files changed outside your deployment pipeline. A hand edit at 02:00, a config-management agent fighting your last deploy, or an unauthorized change all look identical to the process: new config, reload attempt, keep running.&lt;/p></description></item><item><title>Logstash could not be started: another instance is using the configured data.dir</title><link>https://www.netdata.cloud/guides/logstash/logstash-data-dir-lock-another-instance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-data-dir-lock-another-instance/</guid><description>&lt;h1 id="logstash-could-not-be-started-another-instance-is-using-the-configured-datadir">Logstash could not be started: another instance is using the configured data.dir&lt;/h1>
&lt;p>Logstash exits immediately at startup with an error like:&lt;/p>
&lt;pre tabindex="0">&lt;code>Logstash could not be started because there is already another instance
using the configured data.dir. Please change the value of path.data or
configure a different instance.
&lt;/code>&lt;/pre>&lt;p>The wording varies slightly by version, but the meaning is always the same: Logstash tried to take an exclusive lock on its data directory and failed. This is a hard startup block. No pipeline loads, no events flow, the process exits.&lt;/p></description></item><item><title>Logstash CPU-bound filters (grok hell): high CPU, saturated workers, growing queue</title><link>https://www.netdata.cloud/guides/logstash/logstash-cpu-bound-filters-grok-hell/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-cpu-bound-filters-grok-hell/</guid><description>&lt;h1 id="logstash-cpu-bound-filters-grok-hell-high-cpu-saturated-workers-growing-queue">Logstash CPU-bound filters (grok hell): high CPU, saturated workers, growing queue&lt;/h1>
&lt;p>Worker utilization is pinned above 90%. Host CPU is near its ceiling on allocated cores. The queue is growing, per-event processing duration is climbing, and output errors are absent. The pipeline is not blocked downstream. It is starved for compute because one or more filter plugins cannot keep up with the event rate.&lt;/p>
&lt;p>The usual culprit is complex grok regex evaluation, though Ruby filters, heavy JSON manipulation, and enrichment plugins can produce the same signature. The critical diagnostic fork: if CPU is high, you have a compute problem. If CPU is low or moderate with the same queue growth, you have backpressure or I/O blocking, and the fixes are entirely different. This article covers the compute-bound path.&lt;/p></description></item><item><title>Logstash dead letter queue growing: DLQ diversion, replay, and disabled-by-default risk</title><link>https://www.netdata.cloud/guides/logstash/logstash-dead-letter-queue-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-dead-letter-queue-growing/</guid><description>&lt;h1 id="logstash-dead-letter-queue-growing-dlq-diversion-replay-and-disabled-by-default-risk">Logstash dead letter queue growing: DLQ diversion, replay, and disabled-by-default risk&lt;/h1>
&lt;p>A growing &lt;code>dead_letter_queue.queue_size_in_bytes&lt;/code> means an output is permanently rejecting events and diverting them to the DLQ instead of delivering them downstream. This is a data-integrity failure, not just a metric ticking up. Every event in the DLQ is data your destination never received.&lt;/p>
&lt;p>Two design properties compound the risk. The DLQ is disabled by default in all Logstash versions; many production deployments run without it, meaning permanently failed events are logged and silently lost with no on-disk record. And the DLQ only captures specific output failure classes. Filter and parse failures (grok, json, date) do not go to the DLQ. They pass through the pipeline with &lt;code>_grokparsefailure&lt;/code> or similar tags and reach the output as malformed events.&lt;/p></description></item><item><title>Logstash disk full: PQ, DLQ, and log volumes competing for space</title><link>https://www.netdata.cloud/guides/logstash/logstash-disk-space-io-pressure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-disk-space-io-pressure/</guid><description>&lt;h1 id="logstash-disk-full-pq-dlq-and-log-volumes-competing-for-space">Logstash disk full: PQ, DLQ, and log volumes competing for space&lt;/h1>
&lt;p>Logstash is down, the logs show disk write failures, and the persistent queue metrics look innocent: &lt;code>queue_size_in_bytes&lt;/code> is well under &lt;code>max_queue_size_in_bytes&lt;/code>. The queue never filled. The disk did.&lt;/p>
&lt;p>This is a distinct outage path from queue fullness. &lt;code>queue.max_bytes&lt;/code> limits how much the persistent queue itself allocates. It says nothing about total disk consumption on the partition. When the PQ shares a filesystem with the dead letter queue, Logstash&amp;rsquo;s own log files, or the OS, any of those consumers can fill the partition while the PQ stays comfortably within its configured limit. Once the filesystem returns no-space errors, page writes fail, checkpoint writes fail, and the process dies.&lt;/p></description></item><item><title>Logstash downstream backpressure cascade: when a slow output stalls the whole pipeline</title><link>https://www.netdata.cloud/guides/logstash/logstash-downstream-backpressure-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-downstream-backpressure-cascade/</guid><description>&lt;h1 id="logstash-downstream-backpressure-cascade-when-a-slow-output-stalls-the-whole-pipeline">Logstash downstream backpressure cascade: when a slow output stalls the whole pipeline&lt;/h1>
&lt;p>Logstash is up, the API returns 200, and your dashboards are going dark. Beats agents are buffering, Kafka consumer lag is climbing, and throughput collapsed while CPU sits at 15%. The process looks healthy from the outside. It is not.&lt;/p>
&lt;p>A slow or failing output (Elasticsearch, Kafka, HTTP endpoint, syslog receiver) raises output duration. Worker threads block waiting for acknowledgment. The queue fills. The persistent queue absorbs the gap for a while, sometimes hours. When it hits &lt;code>max_bytes&lt;/code>, inputs are blocked. Upstream systems back up or drop events.&lt;/p></description></item><item><title>Logstash Elasticsearch 429: retrying failed action with response code 429</title><link>https://www.netdata.cloud/guides/logstash/logstash-elasticsearch-429-bulk-rejections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-elasticsearch-429-bulk-rejections/</guid><description>&lt;p>The log line &lt;code>retrying failed action with response code: 429&lt;/code> from &lt;code>logstash.outputs.elasticsearch&lt;/code> means Elasticsearch refused a bulk request because its bulk thread pool queue was full (&lt;code>es_rejected_execution_exception&lt;/code>). ES is saturated and cannot keep up with the incoming bulk request rate.&lt;/p>
&lt;p>This is the most common trigger of the backpressure wedge in Logstash-to-ES pipelines. The ES output plugin retries the rejected batch with exponential backoff. Worker threads block on those retries. The internal queue fills. Inputs are backpressured, and upstream systems like Filebeat buffer locally or begin dropping events.&lt;/p></description></item><item><title>Logstash Elasticsearch mapping conflict: field type mismatches and rejected documents</title><link>https://www.netdata.cloud/guides/logstash/logstash-elasticsearch-mapping-conflict/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-elasticsearch-mapping-conflict/</guid><description>&lt;h1 id="logstash-elasticsearch-mapping-conflict-field-type-mismatches-and-rejected-documents">Logstash Elasticsearch mapping conflict: field type mismatches and rejected documents&lt;/h1>
&lt;p>A mapping conflict occurs when Logstash sends a document whose field type does not match the index mapping. A field mapped as &lt;code>long&lt;/code> arrives as a string, or an object appears where a scalar was expected. Elasticsearch rejects the document at the bulk API level.&lt;/p>
&lt;p>The rejection is often invisible. Elasticsearch returns HTTP 200 for bulk requests that contain per-document failures. Logstash counts the batch as &lt;code>events.out&lt;/code> and moves on. If the Dead Letter Queue (DLQ) is enabled, rejected documents divert there. If it is not (the default), documents are logged at WARN and silently lost. Pipeline throughput stays green while data disappears.&lt;/p></description></item><item><title>Logstash Elasticsearch partial bulk failure: HTTP 200 with per-document errors and silent data loss</title><link>https://www.netdata.cloud/guides/logstash/logstash-elasticsearch-partial-bulk-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-elasticsearch-partial-bulk-failure/</guid><description>&lt;h1 id="logstash-elasticsearch-partial-bulk-failure-http-200-with-per-document-errors-and-silent-data-loss">Logstash Elasticsearch partial bulk failure: HTTP 200 with per-document errors and silent data loss&lt;/h1>
&lt;p>Your Logstash pipeline looks healthy. The process is up, the API responds on port 9600, &lt;code>events.out&lt;/code> is climbing steadily, and the queue is empty. But someone asks why last Tuesday&amp;rsquo;s application logs are missing from Elasticsearch. You run a count query against the index and the numbers do not add up. Logstash says it delivered 50 million events. Elasticsearch shows 48 million documents. Two million vanished.&lt;/p></description></item><item><title>Logstash event duplication: in/out ratio drift, clones, and re-read sources</title><link>https://www.netdata.cloud/guides/logstash/logstash-event-duplication/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-event-duplication/</guid><description>&lt;h1 id="logstash-event-duplication-inout-ratio-drift-clones-and-re-read-sources">Logstash event duplication: in/out ratio drift, clones, and re-read sources&lt;/h1>
&lt;p>Logstash exposes three cumulative event counters at the pipeline level: &lt;code>events.in&lt;/code>, &lt;code>events.out&lt;/code>, and &lt;code>events.filtered&lt;/code>. Their relationship reflects the pipeline&amp;rsquo;s intended transformation. A passthrough pipeline should see &lt;code>out&lt;/code> approximately equal to &lt;code>in&lt;/code> over any sustained window. A pipeline with &lt;code>drop {}&lt;/code> filters should see &lt;code>out&lt;/code> plus &lt;code>filtered&lt;/code> approximately equal to &lt;code>in&lt;/code>. Clone and split filters legitimately produce more output events than input events.&lt;/p></description></item><item><title>Logstash file descriptor pressure: leaks, tailed files, and reconnection churn</title><link>https://www.netdata.cloud/guides/logstash/logstash-file-descriptor-pressure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-file-descriptor-pressure/</guid><description>&lt;h1 id="logstash-file-descriptor-pressure-leaks-tailed-files-and-reconnection-churn">Logstash file descriptor pressure: leaks, tailed files, and reconnection churn&lt;/h1>
&lt;p>Logstash is up, throughput looks normal, and then new connections start failing, file opens error out, and the process dies with &amp;ldquo;too many open files.&amp;rdquo; The failure looks sudden. It usually is not. File descriptor pressure builds over days or weeks as &lt;code>process.open_file_descriptors&lt;/code> climbs toward &lt;code>process.max_file_descriptors&lt;/code>, and exhaustion is a cliff edge: everything works until the limit is hit, then new connections, new file opens, and sometimes logging itself fail at once.&lt;/p></description></item><item><title>Logstash file input and sincedb: re-read loops, duplicates, and FD pressure</title><link>https://www.netdata.cloud/guides/logstash/logstash-file-input-sincedb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-file-input-sincedb/</guid><description>&lt;h1 id="logstash-file-input-and-sincedb-re-read-loops-duplicates-and-fd-pressure">Logstash file input and sincedb: re-read loops, duplicates, and FD pressure&lt;/h1>
&lt;p>Your Logstash process looks healthy. The API answers, the pipeline is &lt;code>running&lt;/code>, workers are busy. Then someone downstream asks why the same log lines appear three times in Elasticsearch, or why &lt;code>events.in&lt;/code> just spiked to ten times its baseline, or why Logstash is logging &amp;ldquo;too many open files&amp;rdquo; while traffic looks normal.&lt;/p>
&lt;p>All three symptoms usually trace back to the same component: the file input plugin and its position-tracking file, the sincedb. The file input holds one file descriptor per tailed file and records how far it has read in a sincedb file on disk. When that tracking breaks (corruption, inode reuse after rotation, a remounted filesystem changing device numbers, a wildcard that suddenly matches thousands of files) the failure shows up as duplicate events, unexpected &lt;code>events.in&lt;/code> spikes, or climbing &lt;code>process.open_file_descriptors&lt;/code>.&lt;/p></description></item><item><title>Logstash flow.queue_backpressure: the input-throttling metric explained</title><link>https://www.netdata.cloud/guides/logstash/logstash-flow-queue-backpressure-metric/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-flow-queue-backpressure-metric/</guid><description>&lt;h1 id="logstash-flowqueue_backpressure-the-input-throttling-metric-explained">Logstash flow.queue_backpressure: the input-throttling metric explained&lt;/h1>
&lt;p>You opened the Logstash node stats API during an incident, found &lt;code>flow.queue_backpressure&lt;/code> sitting at 0.9, and now you need to know what that number is actually telling you. The short answer: it is the fraction of time your input threads spend blocked trying to push events into the pipeline queue. It measures ingestion throttling, not worker-side slowness, and that distinction drives the entire diagnosis.&lt;/p></description></item><item><title>Logstash GC death spiral: high garbage-collection overhead and collapsing throughput</title><link>https://www.netdata.cloud/guides/logstash/logstash-gc-death-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-gc-death-spiral/</guid><description>&lt;h1 id="logstash-gc-death-spiral-high-garbage-collection-overhead-and-collapsing-throughput">Logstash GC death spiral: high garbage-collection overhead and collapsing throughput&lt;/h1>
&lt;p>The Logstash process is up. systemd reports it as active. But event throughput has collapsed to near zero, and CPU is pinned at full utilization. This is a garbage-collection death spiral: the JVM spends most of its wall-clock time in GC, reclaiming almost nothing, while the pipeline starves for compute.&lt;/p>
&lt;p>The failure is misleading because the process appears alive. Process-liveness checks pass. The monitoring API on port 9600 may still respond, albeit slowly. But no useful work is happening. The JVM has entered a feedback loop: each GC cycle reclaims less than the last, so GC runs more frequently, steals CPU from event processing, causes events to accumulate, fills the heap faster, and triggers even more GC.&lt;/p></description></item><item><title>Logstash grok slow: catastrophic backtracking and per-event duration spikes</title><link>https://www.netdata.cloud/guides/logstash/logstash-grok-catastrophic-backtracking/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-grok-catastrophic-backtracking/</guid><description>&lt;h1 id="logstash-grok-slow-catastrophic-backtracking-and-per-event-duration-spikes">Logstash grok slow: catastrophic backtracking and per-event duration spikes&lt;/h1>
&lt;p>A Logstash pipeline that ran fine for months suddenly pegs CPU. Workers saturate, the persistent queue grows, and event throughput collapses. The input rate is unchanged, the output destination is healthy, and there are no error storms in the logs. The only visible anomaly is that one grok filter&amp;rsquo;s per-event processing duration has exploded.&lt;/p>
&lt;p>This is the signature of catastrophic regex backtracking, also known as ReDoS. A single malformed or near-miss log line hits a pattern with nested unbounded quantifiers or overlapping alternation, and the Oniguruma regex engine that grok uses takes exponential time to determine that the match fails. One bad event can occupy a pipeline worker thread for seconds or longer.&lt;/p></description></item><item><title>Logstash heap usage high: why the post-GC floor matters more than the peak</title><link>https://www.netdata.cloud/guides/logstash/logstash-heap-usage-high-post-gc-floor/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-heap-usage-high-post-gc-floor/</guid><description>&lt;h1 id="logstash-heap-usage-high-why-the-post-gc-floor-matters-more-than-the-peak">Logstash heap usage high: why the post-GC floor matters more than the peak&lt;/h1>
&lt;p>Logstash operators hit a recurring trap with JVM heap monitoring. They set an alert for &lt;code>heap_used_percent&lt;/code> above 80%, following standard JVM guidance. The alert fires constantly because the JVM heap naturally cycles through allocation and collection. They silence it or raise the threshold. Then a real memory leak or GC death spiral develops, heap stays elevated for hours, and nobody notices because the alert was muted.&lt;/p></description></item><item><title>Logstash input rate dropped to zero: upstream failure vs blocked inputs</title><link>https://www.netdata.cloud/guides/logstash/logstash-input-rate-dropped-to-zero/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-input-rate-dropped-to-zero/</guid><description>&lt;h1 id="logstash-input-rate-dropped-to-zero-upstream-failure-vs-blocked-inputs">Logstash input rate dropped to zero: upstream failure vs blocked inputs&lt;/h1>
&lt;p>The Logstash input rate metric (&lt;code>flow.input_throughput&lt;/code> or the &lt;code>events.in&lt;/code> counter delta) has dropped to zero. Before restarting anything, answer one question: is the queue empty or full?&lt;/p>
&lt;p>An input rate of zero with an empty queue means events are not arriving from upstream. An input rate of zero with a queue at capacity means Logstash has applied backpressure to its own inputs because it cannot drain events fast enough. These conditions require opposite responses, and treating one as the other wastes time during an incident.&lt;/p></description></item><item><title>Logstash JVM heap sizing: -Xms/-Xmx, the 1GB default, and why they should match</title><link>https://www.netdata.cloud/guides/logstash/logstash-jvm-heap-sizing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-jvm-heap-sizing/</guid><description>&lt;h1 id="logstash-jvm-heap-sizing--xms-xmx-the-1gb-default-and-why-they-should-match">Logstash JVM heap sizing: -Xms/-Xmx, the 1GB default, and why they should match&lt;/h1>
&lt;p>Logstash ships with a 1GB JVM heap (&lt;code>-Xms1g -Xmx1g&lt;/code> in &lt;code>config/jvm.options&lt;/code>), unchanged through 7.x, 8.x, and 9.x. That default works for trivial pipelines. It is catastrophically small for production and is the single most common cause of GC death spirals, throughput collapse, and OOM kills.&lt;/p>
&lt;p>The fix: increase the heap, set minimum and maximum to the same value, and account for off-heap memory. Set the heap too large and you starve the OS and risk longer full-GC pauses. Set it without understanding off-heap allocation and you get OOM kills with a heap that looks comfortably under capacity. Ignore container cgroup limits and the JVM sizes itself against host RAM, not the container limit.&lt;/p></description></item><item><title>Logstash Kafka input: consumer group lag and rebalances</title><link>https://www.netdata.cloud/guides/logstash/logstash-kafka-input-consumer-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-kafka-input-consumer-lag/</guid><description>&lt;h1 id="logstash-kafka-input-consumer-group-lag-and-rebalances">Logstash Kafka input: consumer group lag and rebalances&lt;/h1>
&lt;p>The consumer group lag for your &lt;code>logstash&lt;/code> group is climbing and not draining, or the Kafka input periodically drops to zero for seconds or minutes at a time while the brokers look fine. In both cases the evidence lives outside Logstash: consumer-group lag is a Kafka-side signal, and you will not find it in the Logstash monitoring API on port 9600.&lt;/p>
&lt;p>Rising lag means one thing: Logstash is consuming slower than producers are writing. Frequent rebalances mean the group cannot hold a stable assignment, and every rebalance is a window where consumption stops entirely. The two often appear together, because the most common rebalance trigger (slow event processing) is also the most common lag trigger.&lt;/p></description></item><item><title>Logstash memory queue vs persistent queue: durability, visibility, and failure modes</title><link>https://www.netdata.cloud/guides/logstash/logstash-memory-queue-vs-persistent-queue/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-memory-queue-vs-persistent-queue/</guid><description>&lt;h1 id="logstash-memory-queue-vs-persistent-queue-durability-visibility-and-failure-modes">Logstash memory queue vs persistent queue: durability, visibility, and failure modes&lt;/h1>
&lt;p>Every Logstash pipeline has a queue between its inputs and its worker threads. The default is the memory queue: a small, bounded, in-memory buffer with no durability. The alternative is the persistent queue (PQ): a page-based, checkpointed, on-disk buffer that survives restarts.&lt;/p>
&lt;p>This choice changes three things that matter operationally: what you lose when the process dies, what you can see in the metrics API, and how the pipeline fails when the queue fills. Teams enable PQ for durability and then discover it added disk I/O saturation, page corruption, and a false sense of safety to their failure catalogue. Teams running the memory queue often have no written answer for what a crash costs them.&lt;/p></description></item><item><title>Logstash Monitoring</title><link>https://www.netdata.cloud/monitoring-101/logstash-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/logstash-monitoring/</guid><description>&lt;h2 id="logstash-monitoring">Logstash Monitoring&lt;/h2>
&lt;h3 id="what-is-logstash">What Is Logstash?&lt;/h3>
&lt;p>Logstash is a powerful, open-source tool for managing events and logs. It plays a pivotal role in data collection within the &lt;a href="https://www.elastic.co/products/logstash">Elastic Stack&lt;/a>, allowing for seamless data transportation from a multitude of sources to your target destinations for further analysis or storage.&lt;/p>
&lt;h3 id="monitoring-logstash-with-netdata">Monitoring Logstash With Netdata&lt;/h3>
&lt;p>Netdata provides an effective Logstash monitoring tool that offers real-time insights into Logstash performance. Through the integration with &lt;a href="https://app.netdata.cloud/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata Cloud&lt;/a>, DevOps teams can oversee Logstash’s health and quickly identify any anomalies or inefficiencies that could affect their data pipelines. Check out our &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">Live Demo&lt;/a> to see Netdata in action.&lt;/p></description></item><item><title>Logstash monitoring API exposed: port 9600 on a routable interface without auth</title><link>https://www.netdata.cloud/guides/logstash/logstash-monitoring-api-exposed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-monitoring-api-exposed/</guid><description>&lt;h1 id="logstash-monitoring-api-exposed-port-9600-on-a-routable-interface-without-auth">Logstash monitoring API exposed: port 9600 on a routable interface without auth&lt;/h1>
&lt;p>Every Logstash node runs an HTTP monitoring API, by default on port 9600. It answers questions like &amp;ldquo;what pipelines are running&amp;rdquo;, &amp;ldquo;what plugins and versions are installed&amp;rdquo;, and &amp;ldquo;show me the hottest threads with stack traces&amp;rdquo;. By default it binds to 127.0.0.1, which is safe. Exposure happens when someone sets &lt;code>api.http.host&lt;/code> to 0.0.0.0 or a routable IP to make remote monitoring easier, and leaves it there.&lt;/p></description></item><item><title>Logstash monitoring checklist: the signals every production pipeline needs</title><link>https://www.netdata.cloud/guides/logstash/logstash-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-monitoring-checklist/</guid><description>&lt;h1 id="logstash-monitoring-checklist-the-signals-every-production-pipeline-needs">Logstash monitoring checklist: the signals every production pipeline needs&lt;/h1>
&lt;p>Most Logstash monitoring setups answer the wrong question. They answer &amp;ldquo;is the process running?&amp;rdquo; when the question that matters is &amp;ldquo;are events leaving the pipeline?&amp;rdquo; A Logstash JVM can be alive, healthy by systemd&amp;rsquo;s standards, and returning 200 from its monitoring API while the queue is full, workers are blocked on a dead Elasticsearch, and zero events have been delivered for an hour. Process liveness is necessary. It is nowhere near sufficient.&lt;/p></description></item><item><title>Logstash monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/logstash/logstash-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-monitoring-maturity-model/</guid><description>&lt;h1 id="logstash-monitoring-maturity-model-from-survival-to-expert">Logstash monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most Logstash deployments are monitored at the wrong level. Teams alert on process liveness and heap percentage, then get paged by users asking where Wednesday&amp;rsquo;s logs went. The process was up, the API returned 200, and throughput looked fine the whole time. The gap is not tooling. It is which signals the team decided to watch.&lt;/p>
&lt;p>This article defines a four-level maturity model for Logstash monitoring: Survival, Operational, Mature, and Expert. Each level names the specific signals to collect, why they matter, and what class of failure becomes visible that was invisible at the level below. Use it as an audit checklist against your current setup, and as a roadmap for what to add next. The levels are cumulative: every level assumes everything below it is already in place.&lt;/p></description></item><item><title>Logstash multi-pipeline monitoring: why aggregate stats hide a failed pipeline</title><link>https://www.netdata.cloud/guides/logstash/logstash-multi-pipeline-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-multi-pipeline-monitoring/</guid><description>&lt;h1 id="logstash-multi-pipeline-monitoring-why-aggregate-stats-hide-a-failed-pipeline">Logstash multi-pipeline monitoring: why aggregate stats hide a failed pipeline&lt;/h1>
&lt;p>You migrated to &lt;code>pipelines.yml&lt;/code> to isolate workloads: one pipeline per team, per source, or per destination. Each pipeline got its own queue, its own workers, its own failure domain. Right call for fault isolation. But if your monitoring still polls the node once and looks at global event counts, you have given back the visibility the isolation bought you.&lt;/p>
&lt;p>The concrete scenario: five pipelines, roughly equal traffic, one of them wedges. An output stalls, a config reload fails, a Kafka consumer group gets stuck rebalancing. That pipeline&amp;rsquo;s throughput goes to zero; the other four keep flowing. Aggregate node throughput drops by about 20 percent. If your alert is &amp;ldquo;throughput drops more than 50 percent&amp;rdquo; or &amp;ldquo;events per second below N&amp;rdquo;, nothing fires. The failed pipeline can sit dead for hours while node-level graphs show a normal dip.&lt;/p></description></item><item><title>Logstash one input stopped: per-input failures masked by aggregate metrics</title><link>https://www.netdata.cloud/guides/logstash/logstash-input-specific-failure-multi-input/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-input-specific-failure-multi-input/</guid><description>&lt;h1 id="logstash-one-input-stopped-per-input-failures-masked-by-aggregate-metrics">Logstash one input stopped: per-input failures masked by aggregate metrics&lt;/h1>
&lt;p>One of three inputs in your Logstash pipeline stops receiving events. The pipeline-level &lt;code>events.in&lt;/code> counter drops by a third. But because the other two inputs keep flowing, the aggregate rate never hits zero, and your threshold-based alert stays silent. The failed input&amp;rsquo;s upstream source starts accumulating: Kafka consumer lag grows, file tails fall behind, or Beats agents buffer locally. By the time someone notices, hours or days of data from that source are delayed or lost.&lt;/p></description></item><item><title>Logstash OutOfMemoryError: Java heap space and how to recover</title><link>https://www.netdata.cloud/guides/logstash/logstash-out-of-memory-java-heap-space/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-out-of-memory-java-heap-space/</guid><description>&lt;h1 id="logstash-outofmemoryerror-java-heap-space-and-how-to-recover">Logstash OutOfMemoryError: Java heap space and how to recover&lt;/h1>
&lt;p>&lt;code>java.lang.OutOfMemoryError: Java heap space&lt;/code> in the Logstash log means the JVM could not satisfy an allocation request and GC could not reclaim enough heap to proceed. The process exits or gets kernel OOM-killed, and every pipeline it was running stops with it.&lt;/p>
&lt;p>Before the hard crash, there is usually a warning period: the GC death spiral. Heap fills, garbage collection runs longer and more frequently, throughput collapses. This can last minutes or hours before the JVM finally fails to allocate. If you catch the spiral, you can intervene before data is lost. If you only catch the OOM, you are in recovery mode.&lt;/p></description></item><item><title>Logstash output errors and retries: reading downstream failure before the queue grows</title><link>https://www.netdata.cloud/guides/logstash/logstash-output-errors-retries/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-output-errors-retries/</guid><description>&lt;h1 id="logstash-output-errors-and-retries-reading-downstream-failure-before-the-queue-grows">Logstash output errors and retries: reading downstream failure before the queue grows&lt;/h1>
&lt;p>Output-side failures are the most common cause of Logstash queue growth and eventual outage. When the downstream destination (Elasticsearch, Kafka, an HTTP endpoint) starts rejecting, timing out, or slowing down, Logstash output plugins retry. Those retries preserve data temporarily but hide mounting delay. The pipeline appears alive while events accumulate.&lt;/p>
&lt;p>The critical window for diagnosis is between the first output errors and the moment the queue fills. Once the queue reaches capacity, inputs are blocked, upstream systems buffer or drop events, and the incident has cascaded beyond Logstash.&lt;/p></description></item><item><title>Logstash persistent queue full: max_bytes reached and inputs blocked</title><link>https://www.netdata.cloud/guides/logstash/logstash-persistent-queue-full-inputs-blocked/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-persistent-queue-full-inputs-blocked/</guid><description>&lt;h1 id="logstash-persistent-queue-full-max_bytes-reached-and-inputs-blocked">Logstash persistent queue full: max_bytes reached and inputs blocked&lt;/h1>
&lt;p>The persistent queue has hit &lt;code>queue.max_bytes&lt;/code>, input threads are blocked trying to push events in, and the pipeline has stopped accepting new data. A full PQ is almost never the root cause. It is the terminal stage of a downstream outage or capacity failure that the queue was absorbing. Fix the cause, drain the queue, prevent recurrence.&lt;/p>
&lt;p>This guide covers confirming the state, finding the downstream cause, calculating how much time you have, and recovering without making things worse. For the underlying throttling metric, see &lt;a href="https://www.netdata.cloud/guides/logstash/logstash-flow-queue-backpressure-metric/">Logstash flow.queue_backpressure: the input-throttling metric explained&lt;/a>.&lt;/p></description></item><item><title>Logstash persistent queue not draining: page-release lag after downstream recovery</title><link>https://www.netdata.cloud/guides/logstash/logstash-persistent-queue-not-draining/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-persistent-queue-not-draining/</guid><description>&lt;h1 id="logstash-persistent-queue-not-draining-page-release-lag-after-downstream-recovery">Logstash persistent queue not draining: page-release lag after downstream recovery&lt;/h1>
&lt;p>The downstream outage is over. Elasticsearch is green, output errors have stopped, and Logstash is delivering events again. But the persistent queue is shrinking far slower than it filled, disk I/O is pinned, and &lt;code>df&lt;/code> shows the same used space it did an hour ago even though &lt;code>queue_size_in_bytes&lt;/code> has clearly dropped.&lt;/p>
&lt;p>Most of the time, nothing is stuck. A persistent queue (PQ) draining after recovery looks unhealthy from almost every angle: high disk I/O, reduced effective throughput, and disk usage that refuses to go down. This is expected behavior driven by how PQ pages are released. The failure modes you actually need to rule out are narrower: page or checkpoint corruption, poison events pinning pages, and disk I/O saturation becoming the new bottleneck.&lt;/p></description></item><item><title>Logstash persistent queue runway: how long until the PQ fills</title><link>https://www.netdata.cloud/guides/logstash/logstash-persistent-queue-runway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-persistent-queue-runway/</guid><description>&lt;h1 id="logstash-persistent-queue-runway-how-long-until-the-pq-fills">Logstash persistent queue runway: how long until the PQ fills&lt;/h1>
&lt;p>A persistent queue changes the failure shape of a Logstash pipeline. When the output destination goes down, nothing looks broken for a while: inputs keep accepting events, throughput counters keep moving, the process stays green, and the queue absorbs the gap between arrival and delivery. The outage exists the whole time; the PQ just delays the moment anyone feels it.&lt;/p></description></item><item><title>Logstash pipeline not started: a healthy JVM with a missing pipeline</title><link>https://www.netdata.cloud/guides/logstash/logstash-pipeline-not-started/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-pipeline-not-started/</guid><description>&lt;h1 id="logstash-pipeline-not-started-a-healthy-jvm-with-a-missing-pipeline">Logstash pipeline not started: a healthy JVM with a missing pipeline&lt;/h1>
&lt;p>Systemd says Logstash is active. The monitoring API on port 9600 returns 200. Heap looks fine, GC is quiet, and yet one of your data sources has gone silent at the destination. When you query &lt;code>/_node/stats/pipelines&lt;/code>, the pipeline ID you expected is not there, or its event counters have not moved in hours.&lt;/p>
&lt;p>This is a partial outage that process-level checks cannot see. Since Logstash 7.11, a pipeline that crashes no longer takes the JVM down with it: the process stays alive, serving the API, with fewer pipelines running than you configured. &lt;!-- TODO: verify 7.11 as the exact version where a crashed pipeline stopped terminating the JVM --> A failed initial config load, a failed reload that stopped the old pipeline without starting the new one, or a worker error that terminated a single pipeline all produce the same shape: healthy JVM, missing pipeline.&lt;/p></description></item><item><title>Logstash pipeline stalled: output rate at zero while the process looks alive</title><link>https://www.netdata.cloud/guides/logstash/logstash-pipeline-stalled-output-zero/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-pipeline-stalled-output-zero/</guid><description>&lt;h1 id="logstash-pipeline-stalled-output-rate-at-zero-while-the-process-looks-alive">Logstash pipeline stalled: output rate at zero while the process looks alive&lt;/h1>
&lt;p>The Logstash JVM is running. &lt;code>systemctl status logstash&lt;/code> says active. The monitoring API on port 9600 returns 200. Every liveness check you have is green. But &lt;code>events.out&lt;/code> has not moved in twenty minutes, and data stopped arriving downstream at about the same time.&lt;/p>
&lt;p>This is the &amp;ldquo;living dead&amp;rdquo; state: the process exists, but the pipeline is functionally dead. It is one of the most common ways Logstash fails in production, and it is invisible to any monitoring that only checks process existence. Output throughput, not liveness, is the health signal.&lt;/p></description></item><item><title>Logstash pipeline.workers and batch_size: tuning the worker pool</title><link>https://www.netdata.cloud/guides/logstash/logstash-pipeline-workers-batch-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-pipeline-workers-batch-tuning/</guid><description>&lt;h1 id="logstash-pipelineworkers-and-batch_size-tuning-the-worker-pool">Logstash pipeline.workers and batch_size: tuning the worker pool&lt;/h1>
&lt;p>The worker pool is the most important throughput lever in Logstash. &lt;code>pipeline.workers&lt;/code> sets how many threads process events through the filter and output stages. &lt;code>pipeline.batch.size&lt;/code> sets how many events each worker handles per trip through the queue. Together, these define the maximum in-flight event count, the pipeline&amp;rsquo;s memory footprint, and how much CPU and downstream capacity it can consume.&lt;/p>
&lt;p>Tuning is not about finding a universal optimum. It is about matching the worker pool to three variables: CPU cost per event (dominated by filters), average event size, and the downstream system&amp;rsquo;s ability to absorb batches efficiently. Getting this wrong produces one of three outcomes: CPU saturation with queue growth, heap exhaustion from too many in-flight events, or output-blocking cascades that look like high worker utilization but produce no useful throughput.&lt;/p></description></item><item><title>Logstash queue events count growing: reading the in-flight backlog</title><link>https://www.netdata.cloud/guides/logstash/logstash-queue-events-count-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-queue-events-count-growing/</guid><description>&lt;h1 id="logstash-queue-events-count-growing-reading-the-in-flight-backlog">Logstash queue events count growing: reading the in-flight backlog&lt;/h1>
&lt;p>&lt;code>queue.events_count&lt;/code> trending upward on the pipeline stats API means one thing: events are arriving faster than workers can process and deliver them. It is the primary backpressure indicator, and the earliest warning you get before inputs block and upstream systems start backing up or dropping data.&lt;/p>
&lt;p>What the number alone does not tell you is how much trouble you are in. The same rising curve means different things for the memory queue versus a persistent queue, and the absolute value is workload-dependent enough that fixed thresholds are nearly useless. A memory queue climbing by a few hundred events can be more urgent than a persistent queue holding a million.&lt;/p></description></item><item><title>Logstash queue full: inputs blocked and the backpressure wedge</title><link>https://www.netdata.cloud/guides/logstash/logstash-queue-full-inputs-blocked/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-queue-full-inputs-blocked/</guid><description>&lt;h1 id="logstash-queue-full-inputs-blocked-and-the-backpressure-wedge">Logstash queue full: inputs blocked and the backpressure wedge&lt;/h1>
&lt;p>The symptom usually arrives from upstream first: Filebeat stops shipping, Kafka consumer lag climbs, or an HTTP input starts refusing connections. Logstash itself looks alive. The API answers on port 9600, the process is running, CPU is often unremarkable. But events are not moving, because the internal queue between inputs and workers is full, and every input thread is blocked trying to push into it.&lt;/p></description></item><item><title>Logstash silent data loss: events dropped with no DLQ and green dashboards</title><link>https://www.netdata.cloud/guides/logstash/logstash-silent-data-loss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-silent-data-loss/</guid><description>&lt;h1 id="logstash-silent-data-loss-events-dropped-with-no-dlq-and-green-dashboards">Logstash silent data loss: events dropped with no DLQ and green dashboards&lt;/h1>
&lt;p>The first sign is usually a question from downstream: &amp;ldquo;Where are last Wednesday&amp;rsquo;s logs?&amp;rdquo; Your Logstash dashboards show green. Process liveness is fine. Events-in and events-out curves track each other. The queue is empty. GC is quiet. Nothing paged.&lt;/p>
&lt;p>Silent data loss happens when events leave the Logstash accounting system as &amp;ldquo;delivered&amp;rdquo; but never reach their destination in usable form. The &lt;code>events.out&lt;/code> counter increments, throughput looks normal, and no alert fires. By the time anyone notices, the loss window may be hours or days old and the evidence is gone.&lt;/p></description></item><item><title>Logstash SSL/TLS handshake failures: expired certs and broken trust chains</title><link>https://www.netdata.cloud/guides/logstash/logstash-ssl-tls-handshake-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-ssl-tls-handshake-failures/</guid><description>&lt;h1 id="logstash-ssltls-handshake-failures-expired-certs-and-broken-trust-chains">Logstash SSL/TLS handshake failures: expired certs and broken trust chains&lt;/h1>
&lt;p>Your Logstash logs are filling with &lt;code>PKIX path building failed&lt;/code> or &lt;code>handshake_failure&lt;/code> errors. Depending on which side is failing, you either have a delivery outage in progress (an output that can no longer connect to its destination) or a security signal (clients failing to authenticate against an input). The triage path starts the same either way: find the failing plugin, read the actual certificate error, and determine whether this is an availability problem, a security problem, or a planned rotation you forgot about.&lt;/p></description></item><item><title>Logstash thread starvation: workers blocked, low CPU, and stalled throughput</title><link>https://www.netdata.cloud/guides/logstash/logstash-thread-starvation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-thread-starvation/</guid><description>&lt;h1 id="logstash-thread-starvation-workers-blocked-low-cpu-and-stalled-throughput">Logstash thread starvation: workers blocked, low CPU, and stalled throughput&lt;/h1>
&lt;p>Logstash throughput has collapsed. The queue is growing. Your pipeline workers are all occupied. But CPU is sitting at 20% with no obvious explanation. Restarting the process clears it temporarily, then the pattern repeats. This is thread starvation. The symptoms look contradictory, which makes it easy to misdiagnose.&lt;/p>
&lt;p>The mechanism: pipeline worker threads are stuck waiting on a shared resource rather than doing computation. Each worker pulls a batch from the queue and does not release it until every filter and output in the chain completes. When the thing that completes that chain is a slow network call, an exhausted connection pool, a mutex held by another worker, or a DNS resolver that hangs, the worker blocks. With enough workers blocked, no one drains the queue, backpressure propagates to inputs, and the pipeline stalls. CPU stays low because threads are parked in wait states, not burning cycles.&lt;/p></description></item><item><title>Logstash too many open files: file-descriptor exhaustion and refused connections</title><link>https://www.netdata.cloud/guides/logstash/logstash-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-too-many-open-files/</guid><description>&lt;h1 id="logstash-too-many-open-files-file-descriptor-exhaustion-and-refused-connections">Logstash too many open files: file-descriptor exhaustion and refused connections&lt;/h1>
&lt;p>Logstash is logging &lt;code>java.io.IOException: Too many open files&lt;/code> (EMFILE), Beats shippers report refused or reset connections, and new files are not being tailed. The JVM process is still alive, systemd still says &lt;code>active (running)&lt;/code>, and the monitoring API on port 9600 may still answer. Nothing looks crashed, yet the pipeline has stopped accepting new work.&lt;/p>
&lt;p>That is the defining shape of file-descriptor exhaustion: a partial failure that is routinely misread as a network or disk problem. The process cannot open new sockets, new persistent queue page files, or new files for the file input, because every open of those consumes one file descriptor and the per-process limit has been reached. Existing connections and already-open files keep working, which is why some traffic continues to flow while new connections are refused.&lt;/p></description></item><item><title>Logstash won't start after a crash: persistent queue corruption and checkpoint errors</title><link>https://www.netdata.cloud/guides/logstash/logstash-persistent-queue-corruption-wont-start/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-persistent-queue-corruption-wont-start/</guid><description>&lt;h1 id="logstash-wont-start-after-a-crash-persistent-queue-corruption-and-checkpoint-errors">Logstash won&amp;rsquo;t start after a crash: persistent queue corruption and checkpoint errors&lt;/h1>
&lt;p>Logstash was killed uncleanly (OOM kill, &lt;code>kill -9&lt;/code>, SIGKILL after a pod termination grace period, power loss) and now refuses to start. The log at &lt;code>/var/log/logstash/logstash-plain.log&lt;/code> shows &lt;code>java.io.IOException&lt;/code> and checkpoint-related errors during pipeline initialization, and the process exits before any events flow.&lt;/p>
&lt;p>The persistent queue (PQ) is a page-based on-disk queue: events live in page files (default 250MB each, up to &lt;code>queue.max_bytes&lt;/code> which defaults to 1GB), and checkpoint files record which events have been acknowledged as delivered. Both are updated continuously while the pipeline runs. An unclean shutdown can leave the checkpoint files and page files inconsistent with each other, and Logstash&amp;rsquo;s startup code refuses to open a queue it cannot reconcile.&lt;/p></description></item><item><title>Logstash worker utilization high: reading flow.worker_utilization and per-plugin skew</title><link>https://www.netdata.cloud/guides/logstash/logstash-worker-utilization-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/logstash/logstash-worker-utilization-high/</guid><description>&lt;h1 id="logstash-worker-utilization-high-reading-flowworker_utilization-and-per-plugin-skew">Logstash worker utilization high: reading flow.worker_utilization and per-plugin skew&lt;/h1>
&lt;p>High &lt;code>flow.worker_utilization&lt;/code> by itself means workers are busy. It becomes actionable only when the queue is growing or there is no headroom for traffic spikes. The first check is always queue depth, not utilization.&lt;/p>
&lt;p>&lt;code>flow.worker_utilization&lt;/code> is a percentage from 0 to 100, available in Logstash 8.x at the pipeline level. Plugin-level breakdowns live under &lt;code>plugins.filters[].flow.worker_utilization&lt;/code> and &lt;code>plugins.outputs[].flow.worker_utilization&lt;/code> in the Node Stats API. Pipeline-level tells you whether workers are saturated. Plugin-level tells you which filter or output is responsible.&lt;/p></description></item><item><title>loki</title><link>https://www.netdata.cloud/integrations/data-collection/applications/loki/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/loki/</guid><description/></item><item><title>Loki Monitoring</title><link>https://www.netdata.cloud/monitoring-101/loki-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/loki-monitoring/</guid><description>&lt;h2 id="loki-monitoring">Loki Monitoring&lt;/h2>
&lt;h3 id="what-is-loki">What Is Loki?&lt;/h3>
&lt;p>Loki is an open-source log aggregation system developed by Grafana Labs. It is designed to manage, aggregate, and analyze log data efficiently. Inspired by Prometheus, it focuses on performance and scalability for search across different log streams. Unlike traditional logging tools, Loki uses dynamic tagging to allow logs to be efficiently stored and accessed, promoting cost-effective log management.&lt;/p>
&lt;h3 id="monitoring-loki-with-netdata">Monitoring Loki With Netdata&lt;/h3>
&lt;p>Netdata offers an advanced monitoring solution specifically tailored for Loki. Utilizing an openmetrics Prometheus exporter, Netdata effectively gathers metrics from your Loki instances without the need for a Prometheus server or Grafana setup. This means users can benefit from automated dashboards, real-time alerts, and more. Netdata&amp;rsquo;s integration simplifies the process, allowing you to intensely monitor Loki&amp;rsquo;s performance and health with ease.&lt;/p></description></item><item><title>Loop Telecommunication International Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/loop-telecommunication-international-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/loop-telecommunication-international-inc-snmp-traps/</guid><description/></item><item><title>Loral Wdl SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/loral-wdl-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/loral-wdl-snmp-traps/</guid><description/></item><item><title>Lotus Development Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lotus-development-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lotus-development-corp-snmp-traps/</guid><description/></item><item><title>Lsi Logic SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lsi-logic-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lsi-logic-snmp-traps/</guid><description/></item><item><title>Lucent Technologies SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lucent-technologies-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/lucent-technologies-snmp-traps/</guid><description/></item><item><title>Luminous Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/luminous-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/luminous-networks-inc-snmp-traps/</guid><description/></item><item><title>Lustre metadata</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/lustre-metadata/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/lustre-metadata/</guid><description/></item><item><title>Lustre Metadata Monitoring</title><link>https://www.netdata.cloud/monitoring-101/lustre-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/lustre-monitoring/</guid><description>&lt;h2 id="lustre-metadata-monitoring">Lustre Metadata Monitoring&lt;/h2>
&lt;h3 id="what-is-lustre-metadata">What Is Lustre Metadata?&lt;/h3>
&lt;p>Lustre is a type of parallel distributed file system, widely used for large-scale cluster computing. Originally developed for research and enterprise sectors due to its capacity and high performance, Lustre metadata refers to the internal management data that keeps track of file location, size, and storage attributes. Monitoring Lustre metadata is crucial for maintaining optimal file system operations and ensuring efficient management.&lt;/p>
&lt;h3 id="monitoring-lustre-metadata-with-netdata">Monitoring Lustre Metadata With Netdata&lt;/h3>
&lt;p>Monitoring Lustre with Netdata provides seamless tracking of all critical metrics using an openmetrics (prometheus) exporter. To monitor Lustre metadata, Netdata taps into the &lt;a href="https://github.com/GSI-HPC/prometheus-cluster-exporter">Cluster Exporter&lt;/a> which is capable of gathering essential data. With Netdata, you can ingest data from any Prometheus exporter; unlocking automated dashboards and alerts without needing a Prometheus server or Grafana. This integration supports efficient Lustre metadata monitoring and ensures constant observability over your cluster&amp;rsquo;s performance.&lt;/p></description></item><item><title>Luxn Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/luxn-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/luxn-inc-snmp-traps/</guid><description/></item><item><title>LVM boot activation failure: emergency shell and missing mount points</title><link>https://www.netdata.cloud/guides/lvm/lvm-boot-activation-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-boot-activation-failure/</guid><description>&lt;h1 id="lvm-boot-activation-failure-emergency-shell-and-missing-mount-points">LVM boot activation failure: emergency shell and missing mount points&lt;/h1>
&lt;p>You boot a server and land in a dracut emergency shell or initramfs rescue prompt. Or the system boots but &lt;code>/mnt/data&lt;/code>, &lt;code>/var/lib/postgresql&lt;/code>, or other non-root filesystems are not mounted. The common thread: LVM volumes that should have activated during early boot did not.&lt;/p>
&lt;p>The initramfs activates LVM volumes before mounting root and before systemd mounts the rest of fstab. If the initramfs lacks the right LVM tools or configuration, or if a required PV is slow to appear (iSCSI not connected, multipath not configured, SAN LUN not presented), activation fails before you have a running system to diagnose from.&lt;/p></description></item><item><title>LVM cannot extend a logical volume: adding a PV when the VG is full</title><link>https://www.netdata.cloud/guides/lvm/lvm-cannot-extend-logical-volume/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-cannot-extend-logical-volume/</guid><description>&lt;h1 id="lvm-cannot-extend-a-logical-volume-adding-a-pv-when-the-vg-is-full">LVM cannot extend a logical volume: adding a PV when the VG is full&lt;/h1>
&lt;p>You ran &lt;code>lvextend&lt;/code> and it refused with some variation of &amp;ldquo;insufficient free space&amp;rdquo; or &amp;ldquo;insufficient free extents,&amp;rdquo; while the filesystem that prompted all this sits at 99% and application writes fail. The fix is mechanical once you know which of three independent constraints is actually blocking you.&lt;/p>
&lt;p>The trap is that &amp;ldquo;no space&amp;rdquo; means three different things in an LVM stack. The filesystem can be full while the LV has room. The VG can be out of free extents while every filesystem looks fine. A thin pool can be 100% full while the VG reports free space. Each constraint has a different remediation, and running the wrong one either does nothing or makes the incident worse.&lt;/p></description></item><item><title>LVM commands hang: when lvs, vgs, and pvs block on locks or dead devices</title><link>https://www.netdata.cloud/guides/lvm/lvm-commands-hang-lvs-vgs-pvs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-commands-hang-lvs-vgs-pvs/</guid><description>&lt;h1 id="lvm-commands-hang-when-lvs-vgs-and-pvs-block-on-locks-or-dead-devices">LVM commands hang: when lvs, vgs, and pvs block on locks or dead devices&lt;/h1>
&lt;p>You type &lt;code>lvs&lt;/code> on a host during an incident and the prompt never comes back. &lt;code>Ctrl-C&lt;/code> does nothing. &lt;code>kill -9&lt;/code> from another shell does nothing. The process is in D state, uninterruptible sleep, waiting on storage I/O or a metadata lock. Now &lt;code>vgs&lt;/code> and &lt;code>pvs&lt;/code> hang too, and your monitoring agent, which polls &lt;code>lvs&lt;/code> every 60 seconds, has gone silent on exactly the host that is on fire.&lt;/p></description></item><item><title>LVM Couldn't find device with uuid: a physical volume has gone missing</title><link>https://www.netdata.cloud/guides/lvm/lvm-couldnt-find-device-with-uuid/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-couldnt-find-device-with-uuid/</guid><description>&lt;h1 id="lvm-couldnt-find-device-with-uuid-a-physical-volume-has-gone-missing">LVM Couldn&amp;rsquo;t find device with uuid: a physical volume has gone missing&lt;/h1>
&lt;p>The error &amp;ldquo;Couldn&amp;rsquo;t find device with uuid&amp;rdquo; appears when LVM&amp;rsquo;s volume group metadata references a physical volume UUID that no block device on the system currently claims. Every LVM command that touches the affected VG prints the warning. The PV shows as &lt;code>[unknown]&lt;/code> in &lt;code>pvs&lt;/code> output, and the VG enters a partial state.&lt;/p>
&lt;p>The recovery path depends entirely on why the device disappeared and what LV layout sits on top of it. The wrong recovery command, applied too quickly, causes permanent data loss. If system uptime is more than a few minutes and a PV is gone, something has failed at the hardware, fabric, cloud, or operator layer.&lt;/p></description></item><item><title>LVM device node permissions: raw block access that bypasses the filesystem</title><link>https://www.netdata.cloud/guides/lvm/lvm-device-node-permissions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-device-node-permissions/</guid><description>&lt;h1 id="lvm-device-node-permissions-raw-block-access-that-bypasses-the-filesystem">LVM device node permissions: raw block access that bypasses the filesystem&lt;/h1>
&lt;p>Every active logical volume on a Linux host exposes device nodes at &lt;code>/dev/mapper/&amp;lt;VG&amp;gt;-&amp;lt;LV&amp;gt;&lt;/code> and &lt;code>/dev/&amp;lt;VG&amp;gt;/&amp;lt;LV&amp;gt;&lt;/code>. The default permissions on these nodes are &lt;code>brw-rw---- root:disk&lt;/code> (mode 0660). Anything more permissive than that, a mode like 0666, or a group assignment that includes unprivileged users, lets any local user read and write raw block data on the volume. Filesystem permissions, ACLs, and mount options do not apply at this layer. A user who can open the device node can read &lt;code>/etc/shadow&lt;/code> straight off the disk or overwrite filesystem metadata directly.&lt;/p></description></item><item><title>LVM dm-N device numbers change after reboot: use /dev/mapper, not /dev/dm-N</title><link>https://www.netdata.cloud/guides/lvm/lvm-dm-device-numbers-change-reboot/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-dm-device-numbers-change-reboot/</guid><description>&lt;h1 id="lvm-dm-n-device-numbers-change-after-reboot-use-devmapper-not-devdm-n">LVM dm-N device numbers change after reboot: use /dev/mapper, not /dev/dm-N&lt;/h1>
&lt;p>After a reboot, you run &lt;code>lsblk&lt;/code> or &lt;code>df -h&lt;/code> and notice your logical volumes have different &lt;code>/dev/dm-N&lt;/code> numbers. What was &lt;code>dm-0&lt;/code> before the reboot is now &lt;code>dm-2&lt;/code>, and what was &lt;code>dm-1&lt;/code> is now &lt;code>dm-0&lt;/code>. If anything on the system references &lt;code>/dev/dm-0&lt;/code> directly, it may now point at the wrong volume.&lt;/p>
&lt;p>This is not a bug and not a sign of corruption. Device-mapper minor numbers are assigned dynamically at activation time based on the order devices are discovered and activated. The numbers are not stored persistently anywhere. The device-mapper tables that back logical volumes exist only in kernel memory and are torn down on every shutdown, then recreated on every boot.&lt;/p></description></item><item><title>LVM dmeventd not running: the auto-extend safety net is offline</title><link>https://www.netdata.cloud/guides/lvm/lvm-dmeventd-not-running/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-dmeventd-not-running/</guid><description>&lt;h1 id="lvm-dmeventd-not-running-the-auto-extend-safety-net-is-offline">LVM dmeventd not running: the auto-extend safety net is offline&lt;/h1>
&lt;p>You found this page because &lt;code>systemctl is-active lvm2-monitor.service&lt;/code> returned something unexpected, or &lt;code>pgrep dmeventd&lt;/code> came back empty, or &lt;code>lvs -o+seg_monitor&lt;/code> showed &amp;ldquo;not monitored&amp;rdquo; on a thin pool you assumed was protected.&lt;/p>
&lt;p>Here is the uncomfortable part: nothing is broken right now. No errors in &lt;code>lvs&lt;/code>. No failed services your alerting noticed. No I/O hangs. The system looks healthy, and it is, right up until the moment a thin pool hits 100%, a mirror leg dies, or a snapshot overflows. Then you discover, during the incident, that the daemon which was supposed to fire the safety net has been dead for weeks.&lt;/p></description></item><item><title>LVM filesystem full while the volume group has space: the resize step everyone forgets</title><link>https://www.netdata.cloud/guides/lvm/lvm-filesystem-full-but-vg-has-space/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-filesystem-full-but-vg-has-space/</guid><description>&lt;h1 id="lvm-filesystem-full-while-the-volume-group-has-space-the-resize-step-everyone-forgets">LVM filesystem full while the volume group has space: the resize step everyone forgets&lt;/h1>
&lt;p>&lt;code>df -h&lt;/code> says the filesystem is 100% full. Applications are getting ENOSPC. &lt;code>vgs&lt;/code> shows plenty of free space. Maybe you already ran &lt;code>lvextend&lt;/code> and &lt;code>df&lt;/code> still shows the old size. Nothing here is broken. You are looking at three layers that measure three different things, and one of them was never told to grow.&lt;/p>
&lt;p>&lt;code>lvextend&lt;/code> grows the logical volume. The filesystem on top keeps its old size until you explicitly resize it. The LV is a bigger container; the filesystem inside has not expanded into the new space.&lt;/p></description></item><item><title>LVM Found duplicate PV: multipath devices and the lvm.conf filter</title><link>https://www.netdata.cloud/guides/lvm/lvm-found-duplicate-pv/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-found-duplicate-pv/</guid><description>&lt;h1 id="lvm-found-duplicate-pv-multipath-devices-and-the-lvmconf-filter">LVM Found duplicate PV: multipath devices and the lvm.conf filter&lt;/h1>
&lt;p>When LVM prints &amp;ldquo;Found duplicate PV&amp;rdquo; during boot or routine commands, it has found the same PV UUID on two or more block devices. On multipath-attached storage, this almost always means LVM is scanning both the raw SCSI paths (/dev/sdb, /dev/sdc) and the aggregated multipath device (/dev/mapper/mpatha). Each path carries the same PV metadata header.&lt;/p>
&lt;p>The warning is advisory at first: LVM picks one device and proceeds. The danger is which one. If LVM activates a logical volume through a raw path instead of the multipath device, you lose path redundancy and failover. I/O runs through a single HBA link until that link fails, then the volume goes dark. The system can boot fine for weeks until a udev timing race changes which device LVM selects, and the next path failure takes down production.&lt;/p></description></item><item><title>LVM has free space but striped or mirrored allocation still fails</title><link>https://www.netdata.cloud/guides/lvm/lvm-vg-free-space-fragmented-across-pvs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-vg-free-space-fragmented-across-pvs/</guid><description>&lt;h1 id="lvm-has-free-space-but-striped-or-mirrored-allocation-still-fails">LVM has free space but striped or mirrored allocation still fails&lt;/h1>
&lt;p>&lt;code>vgs&lt;/code> shows the volume group 30% free. &lt;code>lvcreate&lt;/code> for a striped or mirrored logical volume, or &lt;code>lvextend&lt;/code> on an existing one, fails anyway:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-text" data-lang="text">&lt;span style="display:flex;">&lt;span>Insufficient suitable allocatable extents for logical volume
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The VG has headroom and the operation needs less than that headroom. Both facts are true at once, and that is the trap: &lt;code>vg_free&lt;/code> is an aggregate across every physical volume in the group, but striped and mirrored allocations are not aggregate operations. They are per-PV placement problems. LVM must find extents on multiple distinct PVs at the same time, and if your free space is piled onto one PV, the allocation has nowhere legal to go.&lt;/p></description></item><item><title>LVM I/O hang: a suspended dm device and processes stuck in D state</title><link>https://www.netdata.cloud/guides/lvm/lvm-io-hang-suspended-device-d-state/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-io-hang-suspended-device-d-state/</guid><description>&lt;h1 id="lvm-io-hang-a-suspended-dm-device-and-processes-stuck-in-d-state">LVM I/O hang: a suspended dm device and processes stuck in D state&lt;/h1>
&lt;p>Processes are stuck in D state. &lt;code>kill -9&lt;/code> does nothing. &lt;code>lvs&lt;/code> hangs. Load average is climbing. The root cause is almost always a device-mapper device stuck in suspended state.&lt;/p>
&lt;p>A suspended dm device blocks all I/O to the underlying logical volume. Every process performing I/O to that volume enters uninterruptible sleep (D state) and cannot be killed, even with SIGKILL. The only way to release them is to resolve the underlying I/O blockage and let queued I/O drain.&lt;/p></description></item><item><title>LVM Insufficient free extents: the volume group is out of space</title><link>https://www.netdata.cloud/guides/lvm/lvm-insufficient-free-extents/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-insufficient-free-extents/</guid><description>&lt;h1 id="lvm-insufficient-free-extents-the-volume-group-is-out-of-space">LVM Insufficient free extents: the volume group is out of space&lt;/h1>
&lt;p>You ran &lt;code>lvcreate&lt;/code>, &lt;code>lvextend&lt;/code>, or &lt;code>lvresize&lt;/code> and got one of these:&lt;/p>
&lt;pre tabindex="0">&lt;code>Insufficient free extents
Insufficient free space: 12800 extents needed, but only 0 available
&lt;/code>&lt;/pre>&lt;p>Both messages mean the same thing: the volume group has no unallocated physical extents left for the allocation you requested. Every physical extent in the VG is already assigned to some logical volume, and the operation fails outright.&lt;/p></description></item><item><title>LVM lock contention: stale lock files and blocked commands</title><link>https://www.netdata.cloud/guides/lvm/lvm-lock-contention/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-lock-contention/</guid><description>&lt;h1 id="lvm-lock-contention-stale-lock-files-and-blocked-commands">LVM lock contention: stale lock files and blocked commands&lt;/h1>
&lt;p>You run &lt;code>lvs&lt;/code> and it sits there. You try &lt;code>vgs&lt;/code> in another terminal and it hangs too. Storage is fine, VMs are running, application I/O is flowing, but every LVM command blocks. Your monitoring, which polls &lt;code>lvs&lt;/code> every 10 seconds, goes stale at the exact moment you need it.&lt;/p>
&lt;p>This is LVM metadata lock contention. Every LVM command (&lt;code>pvs&lt;/code>, &lt;code>vgs&lt;/code>, &lt;code>lvs&lt;/code>, &lt;code>lvcreate&lt;/code>, &lt;code>lvextend&lt;/code>, &lt;code>pvmove&lt;/code>) acquires locks before reading or modifying VG metadata, using lock files in &lt;code>/run/lock/lvm/&lt;/code>. When one command holds a lock and stalls, everything else queues behind it. Because the monitoring commands are LVM commands too, the management plane goes blind under the same condition it is supposed to be observing.&lt;/p></description></item><item><title>LVM logical volume not active: lvchange -ay and why activation failed</title><link>https://www.netdata.cloud/guides/lvm/lvm-lv-not-active/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-lv-not-active/</guid><description>&lt;h1 id="lvm-logical-volume-not-active-lvchange--ay-and-why-activation-failed">LVM logical volume not active: lvchange -ay and why activation failed&lt;/h1>
&lt;p>A logical volume that should be serving data shows as inactive. In &lt;code>lvs&lt;/code> output, &lt;code>lv_attr&lt;/code> position 5 is &lt;code>-&lt;/code> where you expect &lt;code>a&lt;/code>. No device-mapper table is loaded, no &lt;code>/dev/dm-N&lt;/code> device exists, and nothing can mount, open, or write to that LV. Applications see a missing block device, and if this LV backs root or a critical service, the system may not have booted properly.&lt;/p></description></item><item><title>LVM logical volume partial (p) flag: which LVs the missing disk took down</title><link>https://www.netdata.cloud/guides/lvm/lvm-lv-partial-flag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-lv-partial-flag/</guid><description>&lt;h1 id="lvm-logical-volume-partial-p-flag-which-lvs-the-missing-disk-took-down">LVM logical volume partial (p) flag: which LVs the missing disk took down&lt;/h1>
&lt;p>You run &lt;code>lvs&lt;/code> and see a lowercase &lt;code>p&lt;/code> in the ninth character of &lt;code>lv_attr&lt;/code>. That means a physical volume (PV) backing extents in this logical volume has disappeared from the system. The VG is now in partial mode, and some LVs may be silently degraded or completely inaccessible.&lt;/p>
&lt;p>The &lt;code>p&lt;/code> flag is direct evidence of device loss: a disk failed, a SAN LUN was unpresented, a multipath device lost all paths, or a cloud volume was detached. The question is how bad the damage is and which volumes took the hit.&lt;/p></description></item><item><title>LVM logical volumes</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/lvm-logical-volumes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/lvm-logical-volumes/</guid><description/></item><item><title>LVM logical volumes Monitoring</title><link>https://www.netdata.cloud/monitoring-101/lvm-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/lvm-monitoring/</guid><description>&lt;h2 id="lvm-logical-volumes-monitoring">LVM logical volumes Monitoring&lt;/h2>
&lt;h3 id="what-is-lvm">What Is LVM?&lt;/h3>
&lt;p>LVM, or Logical Volume Manager, is a system for managing logical volumes, or filesystems, in Linux environments. It provides a high level of flexibility, allowing administrators to easily resize, extend, or shrink volumes according to their needs. This flexibility is particularly useful in dynamic and modern data environments.&lt;/p>
&lt;h3 id="monitoring-lvm-with-netdata">Monitoring LVM With Netdata&lt;/h3>
&lt;p>Monitoring LVM logical volumes efficiently is where the Netdata monitoring tool excels. With &lt;a href="https://github.com/netdata/netdata">Netdata&lt;/a>, you can monitor LVM to ensure each logical volume&amp;rsquo;s health and performance. Netdata leverages the &lt;code>lvs&lt;/code> CLI tool securely by using &lt;code>ndsudo&lt;/code> to facilitate robust and secure data collection without the need for &lt;code>sudo&lt;/code>.&lt;/p></description></item><item><title>LVM metadata backup stale: /etc/lvm/backup out of sync with the VG</title><link>https://www.netdata.cloud/guides/lvm/lvm-metadata-backup-stale/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-metadata-backup-stale/</guid><description>&lt;h1 id="lvm-metadata-backup-stale-etclvmbackup-out-of-sync-with-the-vg">LVM metadata backup stale: /etc/lvm/backup out of sync with the VG&lt;/h1>
&lt;p>Every LVM command that changes metadata (lvcreate, lvextend, lvremove, vgextend, snapshot creation) writes two things to the local filesystem: an archived copy of the new metadata in &lt;code>/etc/lvm/archive/&lt;/code> and a current-state backup in &lt;code>/etc/lvm/backup/&amp;lt;vg&amp;gt;&lt;/code>. That backup file is what &lt;code>vgcfgrestore&lt;/code> reads when you recover a VG after metadata corruption.&lt;/p>
&lt;p>The failure mode here is quiet: something blocks those writes, the backup file stops updating, and nothing in day-to-day operation tells you. LVM keeps working. Months later you hit metadata corruption, reach for &lt;code>vgcfgrestore&lt;/code>, and restore a VG layout from six months ago. LVs created since then are missing. LVs removed since then come back pointing at extents that have been reallocated. You have traded a recovery exercise for a data integrity problem.&lt;/p></description></item><item><title>LVM metadata corruption recovery: vgcfgrestore from /etc/lvm/archive</title><link>https://www.netdata.cloud/guides/lvm/lvm-metadata-corruption-vgcfgrestore/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-metadata-corruption-vgcfgrestore/</guid><description>&lt;p>LVM keeps the complete map of your logical volumes in a small metadata area at the start of each physical volume. If that metadata is corrupted, by a power loss during a metadata write or a bad sector in the first few megabytes of a PV, the entire volume group becomes unreadable. Every LV in the VG becomes inaccessible at once, regardless of how healthy the actual data extents are.&lt;/p></description></item><item><title>LVM mirror resync storm: multiple rebuilds saturating disk I/O</title><link>https://www.netdata.cloud/guides/lvm/lvm-mirror-resync-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-mirror-resync-storm/</guid><description>&lt;h1 id="lvm-mirror-resync-storm-multiple-rebuilds-saturating-disk-io">LVM mirror resync storm: multiple rebuilds saturating disk I/O&lt;/h1>
&lt;p>Production latency is climbing. Multiple applications are slow. Storage shows high I/O across all physical volumes, but there is no single failed disk and no obvious error in the logs. When you check LVM, several mirrored logical volumes show &lt;code>copy_percent&lt;/code> below 100%. They are all resyncing at the same time.&lt;/p>
&lt;p>This is a mirror resync storm. After a system crash, a disk replacement, or bulk mirror creation, multiple mirrored LVs may need to resynchronize simultaneously. Each resync generates substantial sequential and random I/O as the mirror copies dirty regions or entire leg contents to the recovering side. When several run in parallel, the combined I/O overwhelms storage bandwidth. Production workloads suffer elevated latency, and each resync takes longer than it would alone because they compete for the same physical devices.&lt;/p></description></item><item><title>LVM monitoring checklist: the signals every production volume manager needs</title><link>https://www.netdata.cloud/guides/lvm/lvm-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-monitoring-checklist/</guid><description>&lt;h1 id="lvm-monitoring-checklist-the-signals-every-production-volume-manager-needs">LVM monitoring checklist: the signals every production volume manager needs&lt;/h1>
&lt;p>Most LVM incidents are not exotic. They are a thin pool that hit 100% because nobody watched &lt;code>metadata_percent&lt;/code>, a mirror that ran degraded for three months because nobody alerted on &lt;code>copy_percent&lt;/code>, or a PV that went missing on a Friday and paged nobody because the check only ran &lt;code>df&lt;/code>. The failure modes are well known. What is usually missing is a deliberate, minimal set of signals with thresholds and severities attached.&lt;/p></description></item><item><title>LVM monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/lvm/lvm-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-monitoring-maturity-model/</guid><description>&lt;h1 id="lvm-monitoring-maturity-model-from-survival-to-expert">LVM monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most teams monitor LVM by accident. They alert on filesystem fullness, maybe on disk failure, and discover thin pool exhaustion or a silently degraded mirror only when applications start returning I/O errors. LVM has its own failure domains that filesystem metrics cannot see: a thin pool can be at 100% data usage while &lt;code>df&lt;/code> shows 50% free, and a mirror can run on one leg for months with no visible symptom.&lt;/p></description></item><item><title>LVM pvmove stuck or interrupted: temporary mirror state and how to unwind it</title><link>https://www.netdata.cloud/guides/lvm/lvm-pvmove-stuck-or-interrupted/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-pvmove-stuck-or-interrupted/</guid><description>&lt;h1 id="lvm-pvmove-stuck-or-interrupted-temporary-mirror-state-and-how-to-unwind-it">LVM pvmove stuck or interrupted: temporary mirror state and how to unwind it&lt;/h1>
&lt;p>You ran &lt;code>pvmove&lt;/code> to drain a disk, it has been running for hours, and now something is wrong. Maybe the SSH session died, maybe the host rebooted, maybe someone killed the process. Now &lt;code>lvs&lt;/code> shows a mirror LV you never created, and every LVM command against the volume group feels slow or blocked.&lt;/p>
&lt;p>This is the pvmove temporary mirror state. It is recoverable in almost every case, but only if you unwind it the way LVM expects: resume it with a bare &lt;code>pvmove&lt;/code>, or abort it with &lt;code>pvmove --abort&lt;/code>. The one move that turns a tedious situation into a bad one is &lt;code>kill -9&lt;/code> on the pvmove process, or worse, deleting the mirror LV by hand.&lt;/p></description></item><item><title>LVM RAID mismatch count: data integrity after a scrub</title><link>https://www.netdata.cloud/guides/lvm/lvm-raid-mismatch-count/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-raid-mismatch-count/</guid><description>&lt;h1 id="lvm-raid-mismatch-count-data-integrity-after-a-scrub">LVM RAID mismatch count: data integrity after a scrub&lt;/h1>
&lt;p>A non-zero &lt;code>raid_mismatch_count&lt;/code> means the scrub found blocks where data on one leg disagrees with data on the other leg(s). This surfaces as the &lt;code>m&lt;/code> flag in position 9 of &lt;code>lv_attr&lt;/code> and in the &lt;code>raid_mismatch_count&lt;/code> field of &lt;code>lvs&lt;/code>. It demands investigation, though some mismatches turn out to be benign.&lt;/p>
&lt;p>The operational question is not &amp;ldquo;do we have mismatches&amp;rdquo; but &amp;ldquo;is this real corruption or expected noise.&amp;rdquo; RAID1 and RAID10 arrays can report non-zero mismatch counts from documented edge cases in the kernel write path that produce alignment differences in transient data areas. Mismatches appearing after an unclean shutdown are expected until the scrub completes and should not trigger paging on initial boot.&lt;/p></description></item><item><title>LVM RAID or mirror degraded: a leg is dead and you are one failure from data loss</title><link>https://www.netdata.cloud/guides/lvm/lvm-raid-mirror-degraded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-raid-mirror-degraded/</guid><description>&lt;h1 id="lvm-raid-or-mirror-degraded-a-leg-is-dead-and-you-are-one-failure-from-data-loss">LVM RAID or mirror degraded: a leg is dead and you are one failure from data loss&lt;/h1>
&lt;p>Your LVM RAID1 or mirrored logical volume is still serving reads and writes. Applications see no errors. Filesystems are mounted. But one leg is dead and you are running on a single copy with zero redundancy. The next disk failure, cable disconnect, or SAN path loss is total data loss.&lt;/p>
&lt;p>The default &lt;code>raid_fault_policy&lt;/code> is &lt;code>warn&lt;/code>. When a leg fails, dmeventd logs a warning but takes no repair action. No resync is triggered, no alert fires unless your monitoring explicitly checks for degraded state. The system runs indefinitely on the surviving leg until the second failure removes all copies.&lt;/p></description></item><item><title>LVM RAID resync stuck: copy_percent not progressing</title><link>https://www.netdata.cloud/guides/lvm/lvm-raid-resync-stuck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-raid-resync-stuck/</guid><description>&lt;h1 id="lvm-raid-resync-stuck-copy_percent-not-progressing">LVM RAID resync stuck: copy_percent not progressing&lt;/h1>
&lt;p>You are watching &lt;code>lvs&lt;/code> output on a RAID logical volume, and &lt;code>copy_percent&lt;/code> has not moved for over an hour. The array is degraded, a resync or rebuild is supposed to be underway, and the percentage is frozen. Meanwhile, the volume is running on fewer healthy legs than intended, and every minute without full redundancy is a minute where a second disk failure means data loss.&lt;/p></description></item><item><title>LVM reached low water mark for data device: the thin pool warning before the freeze</title><link>https://www.netdata.cloud/guides/lvm/lvm-reached-low-water-mark-data-device/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-reached-low-water-mark-data-device/</guid><description>&lt;h1 id="lvm-reached-low-water-mark-for-data-device-the-thin-pool-warning-before-the-freeze">LVM reached low water mark for data device: the thin pool warning before the freeze&lt;/h1>
&lt;p>You found this line in dmesg or the journal:&lt;/p>
&lt;pre tabindex="0">&lt;code>device-mapper: thin: 253:4: reached low water mark for data device: sending event
&lt;/code>&lt;/pre>&lt;p>This is not an error. It is the dm-thin kernel target reporting that a thin pool&amp;rsquo;s data device has crossed its low water mark and that a device-mapper event has been sent to userspace. dmeventd listens for that event and can auto-extend the pool. On a correctly configured system, you may see this message once, dmeventd extends the pool, and usage drops back below the threshold.&lt;/p></description></item><item><title>LVM recovering a volume group after permanent PV loss: vgreduce --removemissing</title><link>https://www.netdata.cloud/guides/lvm/lvm-vgreduce-removemissing-recovery/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-vgreduce-removemissing-recovery/</guid><description>&lt;h1 id="lvm-recovering-a-volume-group-after-permanent-pv-loss-vgreduce---removemissing">LVM recovering a volume group after permanent PV loss: vgreduce &amp;ndash;removemissing&lt;/h1>
&lt;p>When a physical volume (PV) disappears from the system, the volume group (VG) enters a partial state. LVM commands start printing &amp;ldquo;Couldn&amp;rsquo;t find device with uuid&amp;rdquo; warnings, and any logical volume (LV) with extents on the missing PV becomes inaccessible (linear/striped) or degraded (mirror/RAID). Once you have confirmed the PV is permanently lost, &lt;code>vgreduce --removemissing&lt;/code> cleans the missing PV out of the VG metadata.&lt;/p></description></item><item><title>LVM snapshot COW usage climbing: extend or remove before it overflows</title><link>https://www.netdata.cloud/guides/lvm/lvm-snapshot-cow-usage-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-snapshot-cow-usage-growing/</guid><description>&lt;h1 id="lvm-snapshot-cow-usage-climbing-extend-or-remove-before-it-overflows">LVM snapshot COW usage climbing: extend or remove before it overflows&lt;/h1>
&lt;p>&lt;code>lvs&lt;/code> shows a traditional snapshot with &lt;code>snap_percent&lt;/code> at 78%. An hour ago it was 61%. You have a fixed-size COW exception store filling up from writes to the origin volume, and when it reaches 100% the snapshot is invalidated instantly and permanently. There is no warning state, no graceful degradation, and no recovery. If a backup is reading from that snapshot, the backup dies with it.&lt;/p></description></item><item><title>LVM snapshot invalid: the COW exception store filled and the snapshot is gone</title><link>https://www.netdata.cloud/guides/lvm/lvm-snapshot-invalid-cow-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-snapshot-invalid-cow-full/</guid><description>&lt;h1 id="lvm-snapshot-invalid-the-cow-exception-store-filled-and-the-snapshot-is-gone">LVM snapshot invalid: the COW exception store filled and the snapshot is gone&lt;/h1>
&lt;p>You run &lt;code>lvs&lt;/code> and see it: a snapshot with &lt;code>snap_percent&lt;/code> at 100.00 and an &lt;code>I&lt;/code> in the fifth position of &lt;code>lv_attr&lt;/code>. In the kernel log there is a line like &lt;code>Invalidating snapshot: Unable to allocate exception&lt;/code>. The snapshot is dead. It did not degrade, it did not warn you at the application layer, and it cannot be brought back.&lt;/p></description></item><item><title>LVM snapshot slowing the origin: copy-on-write write amplification</title><link>https://www.netdata.cloud/guides/lvm/lvm-snapshot-origin-latency-cow-overhead/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-snapshot-origin-latency-cow-overhead/</guid><description>&lt;h1 id="lvm-snapshot-slowing-the-origin-copy-on-write-write-amplification">LVM snapshot slowing the origin: copy-on-write write amplification&lt;/h1>
&lt;p>Write latency on a logical volume doubles or worse with no obvious cause. The disks are healthy, dmesg shows no I/O errors, no RAID resync is running, and the workload has not changed. The one thing that did change: someone created a snapshot, often hours or days ago, usually for a backup that has long since finished.&lt;/p>
&lt;p>This is a traditional (thick) LVM snapshot working as designed. While the snapshot exists, the first write to every chunk of the origin triggers a copy-on-write (COW) cycle: read the original chunk, write it to the snapshot&amp;rsquo;s exception store, then write the new data. One application write becomes a minimum of three physical I/Os. Under write-heavy load, origin latency doubles or worse, and it degrades further as the exception store fills.&lt;/p></description></item><item><title>LVM thin pool auto-extend not working: threshold 100 means disabled</title><link>https://www.netdata.cloud/guides/lvm/lvm-thin-pool-autoextend-not-working/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-thin-pool-autoextend-not-working/</guid><description>&lt;h1 id="lvm-thin-pool-auto-extend-not-working-threshold-100-means-disabled">LVM thin pool auto-extend not working: threshold 100 means disabled&lt;/h1>
&lt;p>Your thin pool hit 100% data usage, writes to every thin LV in the pool started failing, and the auto-extend you thought was protecting you never fired. Or you are reading this before that happens, auditing a pool that has been &amp;ldquo;protected&amp;rdquo; for months without anyone ever confirming the mechanism works.&lt;/p>
&lt;p>The most common root cause is a one-line surprise in &lt;code>/etc/lvm/lvm.conf&lt;/code>: the default value of &lt;code>thin_pool_autoextend_threshold&lt;/code> is &lt;strong>100&lt;/strong>, and a threshold of 100 does not mean &amp;ldquo;extend when the pool is full&amp;rdquo;. It means auto-extend is &lt;strong>disabled&lt;/strong>. The commented-out line &lt;code># thin_pool_autoextend_threshold = 100&lt;/code> in the shipped config looks like a sensible default. It is the off switch.&lt;/p></description></item><item><title>LVM thin pool check needed: thin_check and lvconvert --repair</title><link>https://www.netdata.cloud/guides/lvm/lvm-thin-pool-needs-repair/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-thin-pool-needs-repair/</guid><description>&lt;h1 id="lvm-thin-pool-check-needed-thin_check-and-lvconvert---repair">LVM thin pool check needed: thin_check and lvconvert &amp;ndash;repair&lt;/h1>
&lt;p>A thin pool that reports &amp;ldquo;check needed&amp;rdquo; is telling you the kernel no longer trusts the pool&amp;rsquo;s metadata B-tree. The pool may still be serving I/O, or it may already be erroring writes. Either way, the wrong move (resizing metadata, forcing activation, running the wrong repair path on an oversized metadata LV) can turn a recoverable state into permanent data loss.&lt;/p></description></item><item><title>LVM thin pool metadata full: the exhaustion that can corrupt the pool</title><link>https://www.netdata.cloud/guides/lvm/lvm-thin-pool-metadata-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-thin-pool-metadata-full/</guid><description>&lt;h1 id="lvm-thin-pool-metadata-full-the-exhaustion-that-can-corrupt-the-pool">LVM thin pool metadata full: the exhaustion that can corrupt the pool&lt;/h1>
&lt;p>Your thin pool&amp;rsquo;s &lt;code>metadata_percent&lt;/code> has hit 100. Writes to every thin LV in the pool are failing or hanging, and yet &lt;code>lvs&lt;/code> shows &lt;code>data_percent&lt;/code> at 40%. That is not a contradiction. It is the defining feature of thin pool metadata exhaustion, and it is the LVM failure mode most likely to cost you data.&lt;/p>
&lt;p>Metadata exhaustion is more dangerous than data exhaustion for one reason: a full data area stops new allocations, but a full metadata area can leave the pool structurally corrupted. The metadata LV holds the block mapping tables for every thin LV and snapshot in the pool. When the kernel can no longer allocate a metadata block mid-transaction, it aborts the transaction and switches the pool to read-only mode. From that point, recovery via &lt;code>lvconvert --repair&lt;/code> is possible but not guaranteed, and the &lt;code>lvmthin(7)&lt;/code> man page warns plainly that data from thin LVs may ultimately be unrecoverable.&lt;/p></description></item><item><title>LVM thin pool out of data space: every thin volume freezes at once</title><link>https://www.netdata.cloud/guides/lvm/lvm-thin-pool-full-out-of-data-space/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-thin-pool-full-out-of-data-space/</guid><description>&lt;h1 id="lvm-thin-pool-out-of-data-space-every-thin-volume-freezes-at-once">LVM thin pool out of data space: every thin volume freezes at once&lt;/h1>
&lt;p>The symptom usually arrives as &amp;ldquo;the server hung&amp;rdquo; rather than &amp;ldquo;the storage filled up.&amp;rdquo; Processes writing to any thin volume in the pool go into uninterruptible sleep (D state). Applications stop mid-write without returning errors at first. Databases stall, VMs pause, containers wedge. If root or swap sits on a thin LV in the pool, the whole system becomes unresponsive and SSH sessions freeze.&lt;/p></description></item><item><title>LVM thin pool overprovisioning: 60% full can still be dangerous</title><link>https://www.netdata.cloud/guides/lvm/lvm-thin-pool-overprovisioning-ratio/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-thin-pool-overprovisioning-ratio/</guid><description>&lt;h1 id="lvm-thin-pool-overprovisioning-60-full-can-still-be-dangerous">LVM thin pool overprovisioning: 60% full can still be dangerous&lt;/h1>
&lt;p>Your thin pool shows 60% data usage. Every dashboard is green, no alert has fired, and the pool has sat around that number for weeks. Then a single process on one thin volume starts writing aggressively, and eleven minutes later every thin LV in the pool is frozen: databases crash, VMs pause, and writes across the entire pool queue for 60 seconds and then fail with I/O errors.&lt;/p></description></item><item><title>LVM thin pool queue_if_no_space: why a full pool hangs instead of erroring</title><link>https://www.netdata.cloud/guides/lvm/lvm-thin-pool-queue-if-no-space-hang/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-thin-pool-queue-if-no-space-hang/</guid><description>&lt;h1 id="lvm-thin-pool-queue_if_no_space-why-a-full-pool-hangs-instead-of-erroring">LVM thin pool queue_if_no_space: why a full pool hangs instead of erroring&lt;/h1>
&lt;p>The server looks dead. Processes pile up in uninterruptible sleep, shells that touch certain filesystems never return, and even &lt;code>lvs&lt;/code> hangs. There are no I/O errors in the application logs, no kernel oops, nothing in &lt;code>dmesg&lt;/code> that screams hardware. Teams burn the first hour of this incident investigating a kernel or storage fault, when the actual cause is an LVM thin pool that ran out of data space.&lt;/p></description></item><item><title>LVM thin pool space not reclaimed: discard, TRIM, and fstrim</title><link>https://www.netdata.cloud/guides/lvm/lvm-reclaim-thin-pool-space-fstrim/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-reclaim-thin-pool-space-fstrim/</guid><description>&lt;h1 id="lvm-thin-pool-space-not-reclaimed-discard-trim-and-fstrim">LVM thin pool space not reclaimed: discard, TRIM, and fstrim&lt;/h1>
&lt;p>You deleted 200 GB of files from a filesystem on a thin LV. &lt;code>df&lt;/code> shows the space as free, but &lt;code>lvs&lt;/code> still reports the thin pool &lt;code>data_percent&lt;/code> exactly where it was. Nothing is broken in the reporting: this is how thin provisioning works. Deleting a file only updates filesystem metadata. The pool blocks that held the data stay allocated until a discard (TRIM) operation travels from the filesystem, through the thin LV, into the pool.&lt;/p></description></item><item><title>LVM thin snapshots vs old-style COW snapshots: which failure mode you inherit</title><link>https://www.netdata.cloud/guides/lvm/lvm-thin-vs-cow-snapshots/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-thin-vs-cow-snapshots/</guid><description>&lt;h1 id="lvm-thin-snapshots-vs-old-style-cow-snapshots-which-failure-mode-you-inherit">LVM thin snapshots vs old-style COW snapshots: which failure mode you inherit&lt;/h1>
&lt;p>LVM offers two snapshot mechanisms that look similar from the command line but fail in fundamentally different ways. The choice between them is about which failure mode you inherit when something goes wrong.&lt;/p>
&lt;p>Old-style COW (copy-on-write) snapshots allocate a fixed exception store at creation time. When that store fills to 100%, the snapshot is invalidated permanently and silently. The origin volume continues running. You lose the snapshot and whatever backup or rollback point it represented, but production keeps going.&lt;/p></description></item><item><title>LVM unexpected metadata changes: auditing pvcreate, vgcreate, and lvcreate</title><link>https://www.netdata.cloud/guides/lvm/lvm-unexpected-metadata-changes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-unexpected-metadata-changes/</guid><description>&lt;h1 id="lvm-unexpected-metadata-changes-auditing-pvcreate-vgcreate-and-lvcreate">LVM unexpected metadata changes: auditing pvcreate, vgcreate, and lvcreate&lt;/h1>
&lt;p>You found a logical volume nobody remembers creating. Or a volume group grew a new physical volume overnight. Or &lt;code>lvdisplay&lt;/code> output does not match your configuration management records. None of your change tickets explain it, and now you need to answer two questions fast: what changed, and who changed it.&lt;/p>
&lt;p>LVM metadata changes are a high-signal security event because every &lt;code>pvcreate&lt;/code>, &lt;code>vgcreate&lt;/code>, &lt;code>lvcreate&lt;/code>, &lt;code>lvremove&lt;/code>, and &lt;code>vgextend&lt;/code> requires root. An unexpected LVM change outside a change window means root ran that command, which narrows the possibilities to an authorized root process you forgot about (automation, a package script, a provisioning tool) or an unauthorized one (a compromised service that escalated, a rogue script, an attacker with a shell).&lt;/p></description></item><item><title>LVM vgck metadata inconsistency: PVs disagree about the volume group</title><link>https://www.netdata.cloud/guides/lvm/lvm-vgck-metadata-inconsistent/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-vgck-metadata-inconsistent/</guid><description>&lt;h1 id="lvm-vgck-metadata-inconsistency-pvs-disagree-about-the-volume-group">LVM vgck metadata inconsistency: PVs disagree about the volume group&lt;/h1>
&lt;p>&lt;code>vgck&lt;/code> returned non-zero on a production VG, or you are seeing warnings like &amp;ldquo;Inconsistent metadata found for VG&amp;rdquo; or &amp;ldquo;ignoring metadata seqno N on /dev/sdX for seqno M on /dev/sdY.&amp;rdquo; The physical volumes in the volume group hold different versions of the VG metadata. One or more PVs are stale: they missed a metadata update because they were offline, unreachable, or had bad blocks when the write happened.&lt;/p></description></item><item><title>LVM volume group is partial: operating a VG with a missing PV</title><link>https://www.netdata.cloud/guides/lvm/lvm-vg-partial-missing-pv/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-vg-partial-missing-pv/</guid><description>&lt;h1 id="lvm-volume-group-is-partial-operating-a-vg-with-a-missing-pv">LVM volume group is partial: operating a VG with a missing PV&lt;/h1>
&lt;p>You ran &lt;code>vgs&lt;/code> or &lt;code>lvs&lt;/code> and saw something wrong. The VG attributes show a &lt;code>p&lt;/code> in the fourth position (&lt;code>wz--pn-&lt;/code> instead of &lt;code>wz--n-&lt;/code>). One or more PVs show as &lt;code>[unknown]&lt;/code> in &lt;code>pvs&lt;/code> output. Commands that touch the VG now print warnings about a missing device and may refuse to run.&lt;/p>
&lt;p>The VG is in partial mode. LVM metadata still references a PV UUID that no block device on the system currently claims. The metadata itself is intact on surviving PVs, which is why you can still see the VG at all. But any LV that had physical extents on the missing PV now has a hole in its mapping table.&lt;/p></description></item><item><title>LVM volume group running low on free space: vg_free and runway estimation</title><link>https://www.netdata.cloud/guides/lvm/lvm-volume-group-free-space-low/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/lvm/lvm-volume-group-free-space-low/</guid><description>&lt;h1 id="lvm-volume-group-running-low-on-free-space-vg_free-and-runway-estimation">LVM volume group running low on free space: vg_free and runway estimation&lt;/h1>
&lt;p>Your volume group still works: LVs are active, filesystems are mounted, applications are writing. But &lt;code>vgs&lt;/code> shows &lt;code>vg_free&lt;/code> shrinking week over week, and the trend says you hit zero sometime next month. VG exhaustion is a cliff edge: there is no graceful degradation between 99% and 100% full. Everything works until an allocation request cannot be satisfied, and then it fails immediately. Watching &lt;code>vg_free&lt;/code> is the warning the cliff never gives you.&lt;/p></description></item><item><title>LXC Containers</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/lxc-containers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/lxc-containers/</guid><description/></item><item><title>Lynis audit reports</title><link>https://www.netdata.cloud/integrations/data-collection/applications/lynis-audit-reports/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/lynis-audit-reports/</guid><description/></item><item><title>Lynis Audit Reports Monitoring</title><link>https://www.netdata.cloud/monitoring-101/lynis-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/lynis-monitoring/</guid><description>&lt;h2 id="lynis-audit-reports-monitoring">Lynis Audit Reports Monitoring&lt;/h2>
&lt;h3 id="what-is-lynis">What Is Lynis?&lt;/h3>
&lt;p>Lynis is a renowned security auditing tool designed to perform comprehensive scans and audits on Unix-based systems. Its diverse range of checks helps ensure robust system security and compliance, making it a crucial element in any security-conscious IT infrastructure.&lt;/p>
&lt;h3 id="monitoring-lynis-with-netdata">Monitoring Lynis With Netdata&lt;/h3>
&lt;p>To effectively monitor Lynis audit reports, Netdata utilizes an openmetrics (prometheus) exporter. This setup allows Netdata to gather rich metrics by periodically sending HTTP requests to &lt;a href="https://github.com/MauveSoftware/lynis_exporter">lynis_exporter&lt;/a>. Netdata is adept at ingesting data from any Prometheus exporter, empowering users to visualize information with automated dashboards and receive instant alerts without needing a separate Prometheus server or Grafana installation.&lt;/p></description></item><item><title>M3DB</title><link>https://www.netdata.cloud/integrations/exporters/m3db/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/m3db/</guid><description/></item><item><title>macOS</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/macos/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/macos/</guid><description/></item><item><title>macOS</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/macos/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/macos/</guid><description/></item><item><title>macOS Unified Logs</title><link>https://www.netdata.cloud/integrations/logs/macos-unified-logs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/logs/macos-unified-logs/</guid><description/></item><item><title>Madge Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/madge-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/madge-networks-inc-snmp-traps/</guid><description/></item><item><title>Maipu Electric Industrial Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/maipu-electric-industrial-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/maipu-electric-industrial-co-ltd-snmp-traps/</guid><description/></item><item><title>Manjaro Linux</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/manjaro-linux/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/manjaro-linux/</guid><description/></item><item><title>Marc Hirsch SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/marc-hirsch-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/marc-hirsch-snmp-traps/</guid><description/></item><item><title>MariaDB</title><link>https://www.netdata.cloud/integrations/data-collection/databases/mariadb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/mariadb/</guid><description/></item><item><title>Matrix</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/matrix/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/matrix/</guid><description/></item><item><title>Mattermost</title><link>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/mattermost/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/mattermost/</guid><description/></item><item><title>MaxMind GeoIP / GeoLite2</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/maxmind-geoip---geolite2/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/maxmind-geoip---geolite2/</guid><description/></item><item><title>MaxScale</title><link>https://www.netdata.cloud/integrations/data-collection/databases/maxscale/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/maxscale/</guid><description/></item><item><title>MaxScale Monitoring</title><link>https://www.netdata.cloud/monitoring-101/maxscale-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/maxscale-monitoring/</guid><description>&lt;h2 id="maxscale-monitoring">MaxScale Monitoring&lt;/h2>
&lt;h3 id="what-is-maxscale">What Is MaxScale?&lt;/h3>
&lt;p>MaxScale is a database proxy that manages the traffic between client applications and a set of database servers. It is an integral part of the MariaDB ecosystem, coordinating database operations and optimizing performance and scalability for distributed database environments. &lt;a href="https://mariadb.com/kb/en/maxscale/">Learn more about MaxScale.&lt;/a>&lt;/p>
&lt;h3 id="monitoring-maxscale-with-netdata">Monitoring MaxScale With Netdata&lt;/h3>
&lt;p>Netdata provides a comprehensive MaxScale monitoring tool that helps you keep an eye on the performance and health of your MaxScale instances. The Netdata Agent collects real-time metrics from MaxScale, enabling you to monitor MaxScale servers efficiently. With Netdata’s extensive visualizations and intuitive interface, you can diagnose root causes of performance issues swiftly and effectively. For hands-on experience, check out our &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">Live Demo&lt;/a> or &lt;a href="https://app.netdata.cloud/?utm_source=website&amp;amp;utm_content=monitoring101">sign up for a free trial&lt;/a>.&lt;/p></description></item><item><title>Mcafee Associates Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mcafee-associates-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mcafee-associates-inc-snmp-traps/</guid><description/></item><item><title>Mcafee Formerly Secure Computing Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mcafee-formerly-secure-computing-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mcafee-formerly-secure-computing-corporation-snmp-traps/</guid><description/></item><item><title>Mcafee Inc Formerly Network Associates Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mcafee-inc-formerly-network-associates-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mcafee-inc-formerly-network-associates-inc-snmp-traps/</guid><description/></item><item><title>Mcafee WEB Gateway</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/mcafee-web-gateway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/mcafee-web-gateway/</guid><description/></item><item><title>MD RAID</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/md-raid/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/md-raid/</guid><description/></item><item><title>Media5 Corporation M5 Technologies SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/media5-corporation-m5-technologies-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/media5-corporation-m5-technologies-snmp-traps/</guid><description/></item><item><title>Meeting Netdata</title><link>https://www.netdata.cloud/meeting-netdata/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/meeting-netdata/</guid><description/></item><item><title>Meeting with Constantine</title><link>https://www.netdata.cloud/meeting-with-constantine/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/meeting-with-constantine/</guid><description/></item><item><title>Meeting with Netdata</title><link>https://www.netdata.cloud/meeting-with-netdata/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/meeting-with-netdata/</guid><description/></item><item><title>Meeting with Onboarding team</title><link>https://www.netdata.cloud/meeting-with-onboarding-team/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/meeting-with-onboarding-team/</guid><description/></item><item><title>Meeting With Onboarding Team</title><link>https://www.netdata.cloud/onboarding-meeting/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/onboarding-meeting/</guid><description/></item><item><title>Meeting with Satya</title><link>https://www.netdata.cloud/meeting-with-satya/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/meeting-with-satya/</guid><description/></item><item><title>Meeting with Shyam</title><link>https://www.netdata.cloud/meeting-with-shyam/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/meeting-with-shyam/</guid><description/></item><item><title>Meeting with Stuart</title><link>https://www.netdata.cloud/meeting-with-stuart/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/meeting-with-stuart/</guid><description/></item><item><title>Mega System Technologies Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mega-system-technologies-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mega-system-technologies-inc-snmp-traps/</guid><description/></item><item><title>MegaCLI MegaRAID</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/megacli-megaraid/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/megacli-megaraid/</guid><description/></item><item><title>MegaCLI MegaRAID Monitoring</title><link>https://www.netdata.cloud/monitoring-101/megacli-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/megacli-monitoring/</guid><description>&lt;h2 id="megacli-megaraid-monitoring">MegaCLI MegaRAID Monitoring&lt;/h2>
&lt;h3 id="what-is-megacli-megaraid">What Is MegaCLI MegaRAID?&lt;/h3>
&lt;p>MegaCLI is a powerful command-line utility from Broadcom, primarily used to manage and monitor MegaRAID storage controllers. &lt;a href="https://wikitech.wikimedia.org/wiki/MegaCli">Learn more&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-megacli-megaraid-with-netdata">Monitoring MegaCLI MegaRAID With Netdata&lt;/h3>
&lt;p>Using Netdata as your MegaCLI monitoring tool allows you to monitor the health and performance of your RAID adapters, physical drives, and backup batteries seamlessly. Netdata&amp;rsquo;s real-time monitoring capabilities make it one of the most effective tools for monitoring MegaCLI.&lt;/p></description></item><item><title>Megapac SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/megapac-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/megapac-snmp-traps/</guid><description/></item><item><title>Meilisearch</title><link>https://www.netdata.cloud/integrations/data-collection/databases/meilisearch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/meilisearch/</guid><description/></item><item><title>Meilisearch Monitoring</title><link>https://www.netdata.cloud/monitoring-101/meilisearch-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/meilisearch-monitoring/</guid><description>&lt;h2 id="meilisearch-monitoring">Meilisearch Monitoring&lt;/h2>
&lt;h3 id="what-is-meilisearch">What Is Meilisearch?&lt;/h3>
&lt;p>Meilisearch is a powerful, open-source, and real-time search engine designed for swift search interactions across large datasets. Its lightweight design and focus on performance and developer-friendliness make it an ideal choice for modern applications demanding fast and accurate search capabilities.&lt;/p>
&lt;h3 id="monitoring-meilisearch-with-netdata">Monitoring Meilisearch With Netdata&lt;/h3>
&lt;p>When it comes to tools for monitoring Meilisearch, Netdata stands out as a performance-efficient solution. Netdata uses an openmetrics (Prometheus) exporter to monitor Meilisearch, such as the &lt;a href="https://github.com/scottaglia/meilisearch_exporter">Meilisearch Exporter&lt;/a>, which efficiently gathers search engine metrics. With Netdata, you can ingest data from any Prometheus-compatible exporter, automatically generating on-demand dashboards and alerts without requiring a dedicated Prometheus server or Grafana setup. This streamlined monitoring process ensures you receive real-time insights without the need for extensive infrastructure.&lt;/p></description></item><item><title>Meinberg SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/meinberg-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/meinberg-snmp-traps/</guid><description/></item><item><title>Melco Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/melco-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/melco-inc-snmp-traps/</guid><description/></item><item><title>Mellanox Technologies Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mellanox-technologies-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mellanox-technologies-ltd-snmp-traps/</guid><description/></item><item><title>Memcached</title><link>https://www.netdata.cloud/integrations/data-collection/databases/memcached/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/memcached/</guid><description/></item><item><title>Memcached alive but not responding: the silent process hang</title><link>https://www.netdata.cloud/guides/memcached/memcached-silent-process-degradation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-silent-process-degradation/</guid><description>&lt;h1 id="memcached-alive-but-not-responding-the-silent-process-hang">Memcached alive but not responding: the silent process hang&lt;/h1>
&lt;p>The port check passes. The process shows up in &lt;code>ps&lt;/code>. The kernel still accepts TCP on 11211. But every &lt;code>stats&lt;/code> and &lt;code>version&lt;/code> probe times out, and dashboards show &lt;code>cmd_get&lt;/code> and &lt;code>cmd_set&lt;/code> falling off a cliff. This is the silent process hang: memcached is effectively down while appearing alive to every shallow health check.&lt;/p>
&lt;p>Most liveness checks stop at &amp;ldquo;is something listening on the port?&amp;rdquo; That question is answered by the kernel&amp;rsquo;s TCP stack, not by memcached&amp;rsquo;s worker threads. A frozen main thread, a stuck slab rebalancer, a process suspended by SIGSTOP, or an OOM kill in progress can all leave the listen backlog accepting connections while no command ever executes. Clients connect, send a command, and wait forever.&lt;/p></description></item><item><title>Memcached auth_errors: SASL failures, brute force, and the no-auth default</title><link>https://www.netdata.cloud/guides/memcached/memcached-auth-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-auth-failures/</guid><description>&lt;h1 id="memcached-auth_errors-sasl-failures-brute-force-and-the-no-auth-default">Memcached auth_errors: SASL failures, brute force, and the no-auth default&lt;/h1>
&lt;p>Two counters, present since memcached 1.4.4, are the only native signal the daemon gives you about authentication activity: &lt;code>auth_cmds&lt;/code> and &lt;code>auth_errors&lt;/code>. They count attempts and failures. They do not tell you who attempted, from where, or whether authentication is even enforced.&lt;/p>
&lt;p>The critical interpretation rule: if authentication is not enabled (no &lt;code>-S&lt;/code> for SASL, no &lt;code>-Y&lt;/code> for ASCII token auth&lt;/p></description></item><item><title>Memcached bound to 0.0.0.0: unauthenticated access on an open port</title><link>https://www.netdata.cloud/guides/memcached/memcached-network-exposure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-network-exposure/</guid><description>&lt;h1 id="memcached-bound-to-0000-unauthenticated-access-on-an-open-port">Memcached bound to 0.0.0.0: unauthenticated access on an open port&lt;/h1>
&lt;p>Memcached was designed for trusted internal networks. It has no authentication by default, no per-key access control, and no per-connection logging. Anyone who can open a TCP connection to the daemon can read every cached value, overwrite any key, enumerate key names and sizes, and issue &lt;code>flush_all&lt;/code> to wipe the entire cache in one command. Session tokens, PII, and application secrets cached in plaintext are all readable by that caller.&lt;/p></description></item><item><title>Memcached cache stampede: a hot key expires and the backend takes the hit</title><link>https://www.netdata.cloud/guides/memcached/memcached-cache-stampede/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-cache-stampede/</guid><description>&lt;h1 id="memcached-cache-stampede-a-hot-key-expires-and-the-backend-takes-the-hit">Memcached cache stampede: a hot key expires and the backend takes the hit&lt;/h1>
&lt;p>A cache stampede happens when a popular cached key expires or is evicted and many concurrent requests miss at the same instant. Each miss falls through to the backend independently. The backend, usually sized to handle only the small fraction of traffic that misses the cache, receives the full load for that key at once. Memcached itself is not malfunctioning. It is correctly reporting misses and returning them as fast as it can. The victim is the backend, and the fix lives in the application layer.&lt;/p></description></item><item><title>Memcached cas_badval climbing: check-and-set contention and lost updates</title><link>https://www.netdata.cloud/guides/memcached/memcached-cas-badval-conflicts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-cas-badval-conflicts/</guid><description>&lt;h1 id="memcached-cas_badval-climbing-check-and-set-contention-and-lost-updates">Memcached cas_badval climbing: check-and-set contention and lost updates&lt;/h1>
&lt;p>&lt;code>cas_badval&lt;/code> tracks check-and-set operations that failed because the CAS unique token changed between your &lt;code>gets&lt;/code> read and your &lt;code>cas&lt;/code> write. Each increment is a rejected write: another writer modified the key first, and memcached correctly refused the stale update to prevent a lost update.&lt;/p>
&lt;p>A sustained &lt;code>cas_badval&lt;/code> rate above 10% of total CAS attempts, computed as &lt;code>cas_badval / (cas_hits + cas_misses + cas_badval)&lt;/code>, indicates meaningful write contention. The application is spending CPU on writes that never land, and unbounded retry logic can amplify the problem: every failed CAS triggers an immediate re-read and re-write, increasing load on both the client and memcached without making progress.&lt;/p></description></item><item><title>Memcached command rate anomalies: cmd_get and cmd_set spikes and sudden drops</title><link>https://www.netdata.cloud/guides/memcached/memcached-command-rate-anomaly/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-command-rate-anomaly/</guid><description>&lt;p>&lt;code>cmd_get&lt;/code>, &lt;code>cmd_set&lt;/code>, and &lt;code>cmd_touch&lt;/code> are cumulative counters in memcached. They increase monotonically from process start and are meaningless as raw numbers. The signal is the derived rate: the delta between two samples divided by the interval. A sudden change in that rate is almost never the root cause. It is a proxy for something that changed upstream (clients stopped or started sending traffic), inside the cache (a hot key expired, the cache was flushed), or in the application logic (a write loop, a deploy with new key patterns).&lt;/p></description></item><item><title>Memcached conn_yields rising: one client's pipeline starving the others</title><link>https://www.netdata.cloud/guides/memcached/memcached-conn-yields/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-conn-yields/</guid><description>&lt;p>&lt;code>conn_yields&lt;/code> is climbing on a memcached instance. The name suggests thread contention or a locking bug. It is neither.&lt;/p>
&lt;p>&lt;code>conn_yields&lt;/code> is memcached&amp;rsquo;s per-connection fairness throttle. Each worker thread processes requests from its assigned connections in a libevent loop. The &lt;code>-R&lt;/code> flag (default 20, since memcached 1.4.0) caps how many sequential requests a worker pulls from a single connection in one event loop pass. When a connection hits that cap, the worker yields it: moves it to the back of the processing queue, serves other connections, then comes back. The &lt;code>conn_yields&lt;/code> stat increments on each yield.&lt;/p></description></item><item><title>Memcached connection churn: total_connections racing and TIME_WAIT buildup</title><link>https://www.netdata.cloud/guides/memcached/memcached-connection-churn/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-connection-churn/</guid><description>&lt;h1 id="memcached-connection-churn-total_connections-racing-and-time_wait-buildup">Memcached connection churn: total_connections racing and TIME_WAIT buildup&lt;/h1>
&lt;p>&lt;code>total_connections&lt;/code> is a cumulative counter that only moves up. When it climbs fast while &lt;code>curr_connections&lt;/code> stays flat, clients are opening a fresh TCP connection per request and closing it immediately instead of pooling. Each cycle costs a syscall pair on both ends, a socket on both ends, and on Linux the closed socket sits in &lt;code>TIME_WAIT&lt;/code> on the client host for roughly 60 seconds.&lt;/p></description></item><item><title>Memcached connection limit reached: accepting_conns=0 and clients being refused</title><link>https://www.netdata.cloud/guides/memcached/memcached-connection-limit-reached/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-connection-limit-reached/</guid><description>&lt;h1 id="memcached-connection-limit-reached-accepting_conns0-and-clients-being-refused">Memcached connection limit reached: accepting_conns=0 and clients being refused&lt;/h1>
&lt;p>Memcached enforces a hard ceiling on concurrent client connections via the &lt;code>-c&lt;/code> flag (default 1024). When &lt;code>curr_connections&lt;/code> reaches that limit, the daemon flips &lt;code>accepting_conns&lt;/code> to 0, disables the listen socket, and turns away every new TCP connection. Clients see connection-refused errors or timeouts, fall through to the backend, and the cache stops working for anyone not already connected.&lt;/p>
&lt;p>This is purely a connection-slot problem. Memory can be at 40% of &lt;code>limit_maxbytes&lt;/code>, CPU idle, evictions zero, and the process fully serving existing connections. The only thing wrong is that the connection table is full.&lt;/p></description></item><item><title>Memcached connection refused: telling a dead process from a hung or full one</title><link>https://www.netdata.cloud/guides/memcached/memcached-connection-refused/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-connection-refused/</guid><description>&lt;h1 id="memcached-connection-refused-telling-a-dead-process-from-a-hung-or-full-one">Memcached connection refused: telling a dead process from a hung or full one&lt;/h1>
&lt;p>&amp;ldquo;Connection refused&amp;rdquo; from a memcached client is commonly misdiagnosed. Operators run a TCP port check, see that the port is open, and act on that single data point. A port test tells you whether the kernel accepts a TCP SYN on 11211, not whether memcached is processing commands.&lt;/p>
&lt;p>Three distinct failure modes all surface to clients as &amp;ldquo;cannot reach memcached&amp;rdquo;:&lt;/p></description></item><item><title>Memcached CPU saturation: the per-thread ceiling that aggregate graphs hide</title><link>https://www.netdata.cloud/guides/memcached/memcached-cpu-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-cpu-saturation/</guid><description>&lt;h1 id="memcached-cpu-saturation-the-per-thread-ceiling-that-aggregate-graphs-hide">Memcached CPU saturation: the per-thread ceiling that aggregate graphs hide&lt;/h1>
&lt;p>Memcached is rarely CPU-bound. At normal request rates, worker threads spend most of their time in epoll waits. When CPU climbs on a memcached host, the reflex is to check aggregate &lt;code>rusage&lt;/code> and conclude the daemon is busy. That reflex misses the actual failure mode.&lt;/p>
&lt;p>Memcached dispatches each accepted connection round-robin to one of &lt;code>-t&lt;/code> worker threads (default 4). Once assigned, a connection lives on that worker for its lifetime. The worker runs its own libevent loop and serves its assigned connections exclusively. If one worker saturates at 100% of a core, every connection pinned to that worker sees elevated latency. The other workers may be at 20%. The aggregate graph reads roughly 40%. Three quarters of your clients are fine.&lt;/p></description></item><item><title>Memcached curr_connections climbing: connection leaks and missing pooling</title><link>https://www.netdata.cloud/guides/memcached/memcached-connection-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-connection-leak/</guid><description>&lt;h1 id="memcached-curr_connections-climbing-connection-leaks-and-missing-pooling">Memcached curr_connections climbing: connection leaks and missing pooling&lt;/h1>
&lt;p>&lt;code>curr_connections&lt;/code> trends upward with no corresponding traffic increase, edging toward the &lt;code>-c&lt;/code> ceiling (default 1024). When it hits the limit, the listen socket disables: &lt;code>accepting_conns&lt;/code> flips to 0 and new client TCP connects are refused. To the application it looks like memcached is down. To memcached it is idle and healthy on the connections it already holds.&lt;/p>
&lt;p>Most of the time this is not a memcached bug. It is a client-side problem: connections opened and never returned, a broken or absent pool, or a fleet of application instances each holding more sockets than you accounted for. The server is the victim, not the cause.&lt;/p></description></item><item><title>Memcached evicted_time low: distinguishing healthy turnover from cache thrash</title><link>https://www.netdata.cloud/guides/memcached/memcached-eviction-age-thrashing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-eviction-age-thrashing/</guid><description>&lt;h1 id="memcached-evicted_time-low-distinguishing-healthy-turnover-from-cache-thrash">Memcached evicted_time low: distinguishing healthy turnover from cache thrash&lt;/h1>
&lt;p>A high global eviction counter is one of the most over-alerted memcached signals. Operators see &lt;code>evictions&lt;/code> climbing and page someone at 3 a.m. or reflexively add memory. Both reactions are usually wrong. Raw eviction count says nothing about whether the evicted items mattered. The discriminator that turns &amp;ldquo;evictions are high&amp;rdquo; into &amp;ldquo;evictions are harmful&amp;rdquo; is &lt;code>evicted_time&lt;/code>, reported per slab class in the output of &lt;code>stats items&lt;/code>.&lt;/p></description></item><item><title>Memcached evicted_unfetched and expired_unfetched: caching data nobody ever reads</title><link>https://www.netdata.cloud/guides/memcached/memcached-cache-waste-unfetched/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-cache-waste-unfetched/</guid><description>&lt;h1 id="memcached-evicted_unfetched-and-expired_unfetched-caching-data-nobody-ever-reads">Memcached evicted_unfetched and expired_unfetched: caching data nobody ever reads&lt;/h1>
&lt;p>Memcached exposes two counters that most operators never look at, and both describe the same failure: the application is storing data that nobody ever reads. &lt;code>evicted_unfetched&lt;/code> counts items evicted from the LRU before they were ever touched by a read. &lt;code>expired_unfetched&lt;/code> counts items that lived their full TTL and expired without ever being read. Both are efficiency signals, not reliability signals. They will not tell you the cache is down. They tell you the cache is being used as a write-only buffer, and that write-only data is displacing items that might actually be read.&lt;/p></description></item><item><title>Memcached evicting with memory to spare: the slab-class imbalance trap</title><link>https://www.netdata.cloud/guides/memcached/memcached-slab-imbalance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-slab-imbalance/</guid><description>&lt;h1 id="memcached-evicting-with-memory-to-spare-the-slab-class-imbalance-trap">Memcached evicting with memory to spare: the slab-class imbalance trap&lt;/h1>
&lt;p>Evictions are climbing on your memcached instance. Global memory is at 55% of &lt;code>limit_maxbytes&lt;/code>. The natural response is to add memory, restart the daemon, or hunt for an eviction storm. None of those help here, because the cache is not out of memory in aggregate. One slab class is at 100% capacity and discarding recently-active items, while several other classes sit mostly empty.&lt;/p></description></item><item><title>Memcached eviction cascade: when a full cache overloads the backend</title><link>https://www.netdata.cloud/guides/memcached/memcached-eviction-cascade-backend-overload/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-eviction-cascade-backend-overload/</guid><description>&lt;h1 id="memcached-eviction-cascade-when-a-full-cache-overloads-the-backend">Memcached eviction cascade: when a full cache overloads the backend&lt;/h1>
&lt;p>The most common memcached-related production incident is not a memcached crash. The process stays up, serves hits at sub-millisecond latency, CPU is flat, and memory looks fine in aggregate. What breaks is the backend.&lt;/p>
&lt;p>An eviction cascade starts when the working set outgrows the cache. Memory fills, evictions accelerate, and every evicted item becomes a future miss. Those misses fall through to the database, which was sized for cache-assisted load, not raw traffic. The backend saturates, latency climbs stack-wide, and slow responses drive retries that add still more load. It is a positive feedback loop, and memcached is not where it closes.&lt;/p></description></item><item><title>Memcached evictions climbing: the cache is full and discarding live data</title><link>https://www.netdata.cloud/guides/memcached/memcached-evictions-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-evictions-high/</guid><description>&lt;h1 id="memcached-evictions-climbing-the-cache-is-full-and-discarding-live-data">Memcached evictions climbing: the cache is full and discarding live data&lt;/h1>
&lt;p>The &lt;code>evictions&lt;/code> counter is climbing. At least one slab class is saturated, and the daemon is removing valid, non-expired items to make room for new sets. The severity question is whether the cache is discarding cold data the application no longer needs, or live data that will be requested again within seconds.&lt;/p>
&lt;p>Adding memory is the right response only when the cache is globally undersized. If the problem is slab calcification (memory concentrated in the wrong size classes), adding RAM does nothing useful and may mask the real issue. First determine whether evictions are healthy LRU turnover or harmful thrash, and whether the pressure is global or per-slab.&lt;/p></description></item><item><title>Memcached flush_all: the accidental cache wipe and its cold-start blast radius</title><link>https://www.netdata.cloud/guides/memcached/memcached-flush-all-accidental/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-flush-all-accidental/</guid><description>&lt;h1 id="memcached-flush_all-the-accidental-cache-wipe-and-its-cold-start-blast-radius">Memcached flush_all: the accidental cache wipe and its cold-start blast radius&lt;/h1>
&lt;p>You are paged because hit ratio collapsed and the backend database is saturating. The memcached process is up, uptime is stable, memory looks allocated, and there are no evictions. The signature is an instant hit ratio cliff with no gradual decline. The cause is almost certainly a &lt;code>flush_all&lt;/code> command, and the &lt;code>cmd_flush&lt;/code> counter will confirm it.&lt;/p>
&lt;p>A single &lt;code>flush_all&lt;/code> invalidates every item in the cache. It does not lock the server, does not free memory immediately, and does not require authentication by default. Anyone with TCP access to port 11211 can issue it. The result is a cold-cache thundering herd: every subsequent GET misses and hits the backend simultaneously. If the backend is sized for cache-assisted load, it just received a multiple of its designed capacity.&lt;/p></description></item><item><title>Memcached hash table expansion: hash_is_expanding, memory spikes, and item churn</title><link>https://www.netdata.cloud/guides/memcached/memcached-hash-table-expansion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-hash-table-expansion/</guid><description>&lt;h1 id="memcached-hash-table-expansion-hash_is_expanding-memory-spikes-and-item-churn">Memcached hash table expansion: hash_is_expanding, memory spikes, and item churn&lt;/h1>
&lt;p>When you first see &lt;code>hash_is_expanding=1&lt;/code> in memcached stats, it looks like a warning. It is not. The daemon is doubling its internal key hash table to accommodate more items, a routine maintenance operation that runs on a dedicated background thread (since 1.4.0) and does not block request processing. It does temporarily spike memory and add CPU load on the maintenance thread, and in specific cases it can expose underlying problems: item-count oscillation, repeated expansion, or an overstretched memory budget.&lt;/p></description></item><item><title>Memcached high miss rate: separating cold start, new key patterns, and memory pressure</title><link>https://www.netdata.cloud/guides/memcached/memcached-high-miss-rate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-high-miss-rate/</guid><description>&lt;h1 id="memcached-high-miss-rate-separating-cold-start-new-key-patterns-and-memory-pressure">Memcached high miss rate: separating cold start, new key patterns, and memory pressure&lt;/h1>
&lt;p>A high miss rate on memcached is frequently misdiagnosed. The instinct is to assume the cache is too small and add memory. That instinct is usually wrong. Random or unique key access produces a near-0% hit rate no matter how healthy the cache is. A freshly restarted cache legitimately runs near 0% until it warms. A deploy that changed key naming produces instant misses on the new keys without any memory pressure.&lt;/p></description></item><item><title>Memcached hit ratio dropping: reading get_hits, get_misses, and cache effectiveness</title><link>https://www.netdata.cloud/guides/memcached/memcached-low-hit-ratio/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-low-hit-ratio/</guid><description>&lt;h1 id="memcached-hit-ratio-dropping-reading-get_hits-get_misses-and-cache-effectiveness">Memcached hit ratio dropping: reading get_hits, get_misses, and cache effectiveness&lt;/h1>
&lt;p>The cache hit ratio tells you whether memcached is earning its keep. When it drops, the backend database inherits the miss traffic, and a slow decline can cascade into a failure before anyone pages on &amp;ldquo;cache.&amp;rdquo; The ratio is also frequently read wrong: computing it from lifetime counters hides the exact degradation you are trying to catch.&lt;/p>
&lt;h2 id="what-the-hit-ratio-actually-measures">What the hit ratio actually measures&lt;/h2>
&lt;p>The cache hit ratio is the fraction of GET lookups that found the key in cache:&lt;/p></description></item><item><title>Memcached hot key: one key, one thread, and lopsided load</title><link>https://www.netdata.cloud/guides/memcached/memcached-hot-key/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-hot-key/</guid><description>&lt;h1 id="memcached-hot-key-one-key-one-thread-and-lopsided-load">Memcached hot key: one key, one thread, and lopsided load&lt;/h1>
&lt;p>Aggregate &lt;code>cmd_get&lt;/code> and &lt;code>cmd_set&lt;/code> rates look normal. Hit ratio is healthy. Evictions are zero. CPU across the memcached fleet is moderate. Yet a specific subset of clients reports p99 latency an order of magnitude above baseline, and in a consistent-hashing cluster one node runs noticeably hotter than the rest. Memory pressure and connection exhaustion are absent.&lt;/p>
&lt;p>The signature of a memcached hot key is concentration without saturation. A tiny number of keys, sometimes a single key, take a disproportionate share of traffic: a feature flag read on every request, a viral item page, a single bot&amp;rsquo;s session record, a shared rate-limit counter. All funnel traffic to one slab class on one node.&lt;/p></description></item><item><title>Memcached incr/decr misses: evicted counters that silently break rate limiters and locks</title><link>https://www.netdata.cloud/guides/memcached/memcached-incr-decr-misses/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-incr-decr-misses/</guid><description>&lt;h1 id="memcached-incrdecr-misses-evicted-counters-that-silently-break-rate-limiters-and-locks">Memcached incr/decr misses: evicted counters that silently break rate limiters and locks&lt;/h1>
&lt;p>Most memcached operators never look at &lt;code>incr_misses&lt;/code> and &lt;code>decr_misses&lt;/code>. They sit in the per-operation hit/miss breakdown and stay near zero. When they start climbing, the cache itself usually looks healthy: hit ratio is fine, evictions are modest, memory is not full. The damage is happening somewhere else.&lt;/p>
&lt;p>The pattern is specific. The application uses memcached as an atomic counter store. Rate limiters increment a per-client counter and reject when it crosses a threshold. Distributed locks hold a lease token with a TTL and decrement on release. Quota counters track usage per tenant. These counter keys are small, hot, and short-lived. When the slab class they live in comes under memory pressure, the LRU evicts them between operations.&lt;/p></description></item><item><title>Memcached latency high: diagnosing slow gets when the server exposes no latency metric</title><link>https://www.netdata.cloud/guides/memcached/memcached-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-latency-high/</guid><description>&lt;h1 id="memcached-latency-high-diagnosing-slow-gets-when-the-server-exposes-no-latency-metric">Memcached latency high: diagnosing slow gets when the server exposes no latency metric&lt;/h1>
&lt;p>Memcached does not expose a latency histogram, p50, or p99. The standard &lt;code>stats&lt;/code> output is all counters and gauges at the moment of query: no time-series, no distribution, no per-operation duration. If your application is reporting slow cache reads, the daemon itself cannot confirm or deny the problem. You must measure latency at the client, then use server-side and OS signals to localize the cause.&lt;/p></description></item><item><title>Memcached LRU crawler: crawler_reclaimed, disabled crawlers, and lazy expiry</title><link>https://www.netdata.cloud/guides/memcached/memcached-lru-crawler/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-lru-crawler/</guid><description>&lt;h1 id="memcached-lru-crawler-crawler_reclaimed-disabled-crawlers-and-lazy-expiry">Memcached LRU crawler: crawler_reclaimed, disabled crawlers, and lazy expiry&lt;/h1>
&lt;p>Memcached does not remove items when their TTL expires. An expired item holds its slab slot until something touches it: a client request, an eviction scan, or the LRU crawler. The crawler is a background thread that walks LRU chains, finds expired items, and frees their memory before eviction pressure forces the issue.&lt;/p>
&lt;p>A working crawler keeps slab slots available for new writes without evicting live data. Every expired item it reclaims is one fewer item evicted under pressure. Without it, expired items that are never accessed and never reach the eviction tail waste memory indefinitely.&lt;/p></description></item><item><title>Memcached maxconns and ulimit: why the default 1024 is too low for production</title><link>https://www.netdata.cloud/guides/memcached/memcached-max-connections-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-max-connections-tuning/</guid><description>&lt;h1 id="memcached-maxconns-and-ulimit-why-the-default-1024-is-too-low-for-production">Memcached maxconns and ulimit: why the default 1024 is too low for production&lt;/h1>
&lt;p>Memcached is up, the port is open, existing connections work fine. But new clients get connection refused. Their requests fall through to the backend, backend load climbs, and the cache looks healthy from the outside: CPU is low, memory is nowhere near the limit, and the slab allocator is not evicting anything.&lt;/p>
&lt;p>The most common cause is the default &lt;code>-c 1024&lt;/code> (maxconns) limit. It looks generous for a single application instance, but becomes a hard cliff-edge once you scale from 10 to 50 instances, each with its own connection pool. Monitoring agents, health checks, and load balancer probes also consume connections against this budget.&lt;/p></description></item><item><title>Memcached memory utilization: bytes vs limit_maxbytes and why the global number lies</title><link>https://www.netdata.cloud/guides/memcached/memcached-memory-utilization/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-memory-utilization/</guid><description>&lt;h1 id="memcached-memory-utilization-bytes-vs-limit_maxbytes-and-why-the-global-number-lies">Memcached memory utilization: bytes vs limit_maxbytes and why the global number lies&lt;/h1>
&lt;p>The &lt;code>bytes&lt;/code> and &lt;code>limit_maxbytes&lt;/code> fields from &lt;code>stats&lt;/code> look like a clean fill gauge. Divide one by the other and you know how full the cache is. Many dashboards and alerts are built on exactly that ratio.&lt;/p>
&lt;p>This article covers why the global ratio lies, what it actually measures, and which per-slab signals confirm real pressure. &lt;code>bytes&lt;/code> includes per-item overhead and never quite reaches the ceiling, allocation happens in 1MB page jumps rather than smoothly, and the pool is partitioned into fixed-size slab classes that saturate independently. A cache at 50% global utilization can be evicting live, actively-requested data.&lt;/p></description></item><item><title>Memcached monitoring checklist: the signals every production cache needs</title><link>https://www.netdata.cloud/guides/memcached/memcached-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-monitoring-checklist/</guid><description>&lt;h1 id="memcached-monitoring-checklist-the-signals-every-production-cache-needs">Memcached monitoring checklist: the signals every production cache needs&lt;/h1>
&lt;p>A production reference for memcached signals, organized by monitoring maturity. Use it to audit what you collect, identify gaps, and prioritize additions.&lt;/p>
&lt;p>Levels are cumulative: Level 2 includes Level 1. Level 1 detects crashes and restarts but misses slab calcification and eviction quality. Level 3 catches those before they hit the backend. Level 4 adds signals teams adopt after being burned by subtle failures.&lt;/p></description></item><item><title>Memcached monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/memcached/memcached-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-monitoring-maturity-model/</guid><description>&lt;h1 id="memcached-monitoring-maturity-model-from-survival-to-expert">Memcached monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Memcached exposes most of its internal state through plain text counters. Most teams stop at hit ratio, global memory, and eviction rate, then discover during an incident that the signal they needed was buried in &lt;code>stats items&lt;/code> or &lt;code>/proc/&amp;lt;pid&amp;gt;/status&lt;/code>. The four-level model below stages the path from &amp;ldquo;is the process alive&amp;rdquo; to &amp;ldquo;is the segmented LRU moving items between HOT and COLD the way the workload expects&amp;rdquo;.&lt;/p></description></item><item><title>Memcached network saturation: large values, multiget, and a maxed-out NIC</title><link>https://www.netdata.cloud/guides/memcached/memcached-network-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-network-saturation/</guid><description>&lt;h1 id="memcached-network-saturation-large-values-multiget-and-a-maxed-out-nic">Memcached network saturation: large values, multiget, and a maxed-out NIC&lt;/h1>
&lt;p>Memcached latency is climbing for every client, but the host looks healthy. CPU is well below saturation, memory utilization is far from the limit, evictions are zero, and &lt;code>curr_connections&lt;/code> is well below &lt;code>-c&lt;/code>. The daemon answers &lt;code>version&lt;/code> and &lt;code>stats&lt;/code> instantly. Yet every application reports slow gets, and the slowdown is uniform across keys and clients.&lt;/p>
&lt;p>When latency degrades uniformly and the usual suspects are clear, look at the wire. Memcached pulls values from RAM and writes them to a socket. If the response stream exceeds what the NIC can push, the kernel queues, TCP congestion control engages, and every client waits. The cache is fine; the link is the bottleneck.&lt;/p></description></item><item><title>Memcached reclaimed vs evictions: reading memory pressure through reclamation</title><link>https://www.netdata.cloud/guides/memcached/memcached-reclaimed-vs-evictions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-reclaimed-vs-evictions/</guid><description>&lt;h1 id="memcached-reclaimed-vs-evictions-reading-memory-pressure-through-reclamation">Memcached reclaimed vs evictions: reading memory pressure through reclamation&lt;/h1>
&lt;p>Most operators learn one signal first: &lt;code>evictions&lt;/code>. It goes up, the cache is full, hit ratio drops, someone pages. The counter next to it, &lt;code>reclaimed&lt;/code>, tells you how often memcached made room for a new item by reusing an expired item&amp;rsquo;s slot instead of throwing out live data. Read together, they describe the same process from two angles: how memcached finds space for new writes.&lt;/p></description></item><item><title>Memcached response_obj_oom: connections killed for lack of buffer memory</title><link>https://www.netdata.cloud/guides/memcached/memcached-response-obj-oom/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-response-obj-oom/</guid><description>&lt;h1 id="memcached-response_obj_oom-connections-killed-for-lack-of-buffer-memory">Memcached response_obj_oom: connections killed for lack of buffer memory&lt;/h1>
&lt;p>If &lt;code>response_obj_oom&lt;/code> is climbing in your &lt;code>stats&lt;/code> output, the daemon is closing client connections because it cannot allocate internal response buffers. This is not an item-storage problem and not an eviction problem. A cache can be otherwise healthy (zero evictions, free slab memory, low CPU) and still kill clients over this.&lt;/p>
&lt;p>The stat has existed since 1.6.0, when memcached reworked its connection buffer management to allocate buffers on demand instead of reserving them per connection. That cut idle connection overhead but introduced a new failure surface: when the pool of memory for response objects runs out, the daemon closes the connection rather than blocking or queueing. The close is immediate when the allocation fails. There is no retry, no backpressure, no graceful degradation. Applications see timeouts or resets against a daemon that still answers &lt;code>version&lt;/code> and &lt;code>stats&lt;/code> probes.&lt;/p></description></item><item><title>Memcached RSS above the -m limit: connection buffers, hash table, and fragmentation</title><link>https://www.netdata.cloud/guides/memcached/memcached-rss-memory-overhead/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-rss-memory-overhead/</guid><description>&lt;h1 id="memcached-rss-above-the--m-limit-connection-buffers-hash-table-and-fragmentation">Memcached RSS above the -m limit: connection buffers, hash table, and fragmentation&lt;/h1>
&lt;p>Operators set &lt;code>-m&lt;/code> to cap cache memory, then notice RSS is 40 percent or more above that number. This is not a leak. The &lt;code>-m&lt;/code> flag bounds only the slab allocator, the pool that stores cached items. RSS also includes the hash table, per-connection buffers, worker thread stacks, and allocator overhead.&lt;/p>
&lt;p>The &lt;code>bytes&lt;/code> / &lt;code>limit_maxbytes&lt;/code> gauge in memcached stats describes cache memory: the keys and values your application stores. RSS, reported as &lt;code>VmRSS&lt;/code> in &lt;code>/proc/&amp;lt;pid&amp;gt;/status&lt;/code>, describes the entire process. They measure different scopes and will almost never match.&lt;/p></description></item><item><title>Memcached segmented LRU: HOT/WARM/COLD tiers and what the move counters tell you</title><link>https://www.netdata.cloud/guides/memcached/memcached-lru-segments-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-lru-segments-tuning/</guid><description>&lt;h1 id="memcached-segmented-lru-hotwarmcold-tiers-and-what-the-move-counters-tell-you">Memcached segmented LRU: HOT/WARM/COLD tiers and what the move counters tell you&lt;/h1>
&lt;p>Since memcached 1.5.0, every slab class runs a segmented LRU instead of a single flat queue. Items land in HOT, WARM, COLD, and optionally TEMP tiers, and a background thread moves them between tiers based on access patterns. The goal is to protect frequently-accessed items from being evicted by one-off scans that touch cold data.&lt;/p>
&lt;p>The move counters under &lt;code>stats items&lt;/code> describe how well that sorting is working. &lt;code>moves_to_cold&lt;/code> counts items aging out of active use, &lt;code>moves_to_warm&lt;/code> counts items rescued from COLD by a re-access, and &lt;code>moves_within_lru&lt;/code> counts re-ranking within WARM. The ratios between these counters explain why a slab class evicts the way it does, even when global memory looks fine.&lt;/p></description></item><item><title>Memcached SERVER_ERROR object too large for cache: items over the max item size</title><link>https://www.netdata.cloud/guides/memcached/memcached-store-too-large/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-store-too-large/</guid><description>&lt;h1 id="memcached-server_error-object-too-large-for-cache-items-over-the-max-item-size">Memcached SERVER_ERROR object too large for cache: items over the max item size&lt;/h1>
&lt;p>&lt;code>SERVER_ERROR object too large for cache&lt;/code> is the protocol-level error memcached returns when a store operation targets a value larger than the daemon&amp;rsquo;s configured maximum item size. The default ceiling is 1 MB (1048576 bytes), set by the &lt;code>-I&lt;/code> flag. The store is rejected cleanly: nothing is written, and any pre-existing item under that key is left untouched. From the application&amp;rsquo;s side this is usually the first sign that something upstream is serializing more data than intended.&lt;/p></description></item><item><title>Memcached SERVER_ERROR out of memory storing object: stores rejected instead of evicting</title><link>https://www.netdata.cloud/guides/memcached/memcached-store-no-memory/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-store-no-memory/</guid><description>&lt;h1 id="memcached-server_error-out-of-memory-storing-object-stores-rejected-instead-of-evicting">Memcached SERVER_ERROR out of memory storing object: stores rejected instead of evicting&lt;/h1>
&lt;p>When a memcached client receives &lt;code>SERVER_ERROR out of memory storing object&lt;/code>, the daemon refused a SET because it could not allocate a chunk for the item. The &lt;code>store_no_memory&lt;/code> counter increments. Under default operation, memcached evicts LRU items from the relevant slab class to make room. This error means eviction was either disabled or could not produce a free chunk for that allocation.&lt;/p></description></item><item><title>Memcached slab calcification: pages locked to the wrong item sizes after a workload shift</title><link>https://www.netdata.cloud/guides/memcached/memcached-slab-calcification/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-slab-calcification/</guid><description>&lt;h1 id="memcached-slab-calcification-pages-locked-to-the-wrong-item-sizes-after-a-workload-shift">Memcached slab calcification: pages locked to the wrong item sizes after a workload shift&lt;/h1>
&lt;p>Hit ratio is dropping. Evictions are climbing. But global memory utilization sits at 50 to 70 percent, and adding memory does nothing. This is the signature of slab calcification: whole megabyte pages are locked into slab classes that serve item sizes the workload no longer produces, while the classes for the new sizes are starved and evicting actively used data.&lt;/p></description></item><item><title>Memcached slab_automove and slab_reassign: rebalancing pages between classes</title><link>https://www.netdata.cloud/guides/memcached/memcached-slab-automove/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-slab-automove/</guid><description>&lt;h1 id="memcached-slab_automove-and-slab_reassign-rebalancing-pages-between-classes">Memcached slab_automove and slab_reassign: rebalancing pages between classes&lt;/h1>
&lt;p>Memcached&amp;rsquo;s slab allocator divides its memory budget into 1MB pages and permanently assigns each page to a slab class based on item size. Once assigned, a page traditionally stayed locked to that class forever. If the workload&amp;rsquo;s item-size distribution shifted after deployment, you got slab calcification: one class full and evicting while others sat idle with free chunks. Global memory utilization looked healthy while cache effectiveness collapsed for the saturated size range.&lt;/p></description></item><item><title>Memcached swapping: why any VmSwap on an in-memory cache is an incident</title><link>https://www.netdata.cloud/guides/memcached/memcached-swap-usage/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-swap-usage/</guid><description>&lt;h1 id="memcached-swapping-why-any-vmswap-on-an-in-memory-cache-is-an-incident">Memcached swapping: why any VmSwap on an in-memory cache is an incident&lt;/h1>
&lt;p>A nonzero &lt;code>VmSwap&lt;/code> value in &lt;code>/proc/&amp;lt;pid&amp;gt;/status&lt;/code> for a memcached process is a production incident, not a tuning concern. Every access to a swapped page costs milliseconds instead of nanoseconds, so a fraction of your cache lookups runs at disk speed. The point of an in-memory cache is lost for those items, and the resulting latency outliers are random and intermittent, almost impossible to trace from the application layer without checking swap first.&lt;/p></description></item><item><title>Memcached UDP amplification: udpport, CVE-2018-0115, and internet exposure</title><link>https://www.netdata.cloud/guides/memcached/memcached-udp-amplification/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-udp-amplification/</guid><description>&lt;h1 id="memcached-udp-amplification-udpport-cve-2018-1000115-and-internet-exposure">Memcached UDP amplification: udpport, CVE-2018-1000115, and internet exposure&lt;/h1>
&lt;p>You are here because a third party reported your server in a DDoS attack, a security scan flagged UDP 11211, or you are auditing an older deployment. The condition is the same in every case: memcached has UDP enabled (&lt;code>udpport&lt;/code> is non-zero) and is reachable from a network you do not control.&lt;/p>
&lt;p>This is CVE-2018-1000115. In early 2018 it powered reflection attacks peaking at 1.3 to 1.7 Tbps. A single small UDP GET request with a spoofed source IP can trigger a response of hundreds of kilobytes from an unauthenticated memcached instance, with observed amplification factors around 51,000x. The upstream fix in 1.5.6 (Feb 2018) made &lt;code>udpport&lt;/code> default to 0, but that does not retroactively rewrite init scripts, container images, or distro configs.&lt;/p></description></item><item><title>Memcached unexpected restart: uptime reset, wiped cache, and the cold-start backend spike</title><link>https://www.netdata.cloud/guides/memcached/memcached-unexpected-restart/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/memcached/memcached-unexpected-restart/</guid><description>&lt;h1 id="memcached-unexpected-restart-uptime-reset-wiped-cache-and-the-cold-start-backend-spike">Memcached unexpected restart: uptime reset, wiped cache, and the cold-start backend spike&lt;/h1>
&lt;p>The &lt;code>uptime&lt;/code> stat in memcached&amp;rsquo;s &lt;code>stats&lt;/code> output is a seconds-since-process-start counter. When it drops from hours or days back to near-zero, the process restarted. Because memcached has no persistence, a restart is a full cache wipe: every item is gone, &lt;code>curr_items&lt;/code> falls to zero, and hit ratio collapses to 0%.&lt;/p>
&lt;p>The restart is rarely the incident. The incident is the cold-start backend spike that follows. Every cache miss now hits the origin database or service at full production volume. If the backend was sized assuming 90%+ cache offload, it now sees several times its normal read load.&lt;/p></description></item><item><title>Memory modules (DIMMs)</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/memory-modules-dimms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/memory-modules-dimms/</guid><description/></item><item><title>Memory Statistics</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/memory-statistics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/memory-statistics/</guid><description/></item><item><title>Memory Statistics (Win)</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/memory-statistics-win/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/memory-statistics-win/</guid><description/></item><item><title>Memory Usage</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/memory-usage/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/memory-usage/</guid><description/></item><item><title>Memotec Communications SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/memotec-communications-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/memotec-communications-snmp-traps/</guid><description/></item><item><title>Memotec Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/memotec-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/memotec-inc-snmp-traps/</guid><description/></item><item><title>Meraki</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/meraki/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/meraki/</guid><description/></item><item><title>Meraki Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/meraki-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/meraki-networks-inc-snmp-traps/</guid><description/></item><item><title>Merlin Gerin SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/merlin-gerin-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/merlin-gerin-snmp-traps/</guid><description/></item><item><title>Meru Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/meru-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/meru-networks-snmp-traps/</guid><description/></item><item><title>Mesos</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/mesos/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/mesos/</guid><description/></item><item><title>Mesos Monitoring</title><link>https://www.netdata.cloud/monitoring-101/mesos-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/mesos-monitoring/</guid><description>&lt;h2 id="mesos-monitoring">Mesos Monitoring&lt;/h2>
&lt;h3 id="what-is-mesos">What Is Mesos?&lt;/h3>
&lt;p>Apache Mesos is a powerful cluster manager that simplifies resource management and task scheduling across distributed systems. It&amp;rsquo;s widely used to organize and optimize processing tasks in cloud computing environments, taking the complexity out of managing large scale data centers.&lt;/p>
&lt;h3 id="monitoring-mesos-with-netdata">Monitoring Mesos With Netdata&lt;/h3>
&lt;p>To monitor Mesos, Netdata leverages an openmetrics (prometheus) exporter, specifically the &lt;a href="https://github.com/mesosphere/mesos_exporter">Mesos exporter&lt;/a>. Netdata stands out by ingesting data from any Prometheus exporter, providing automated dashboards and alerts without the need for a separate Prometheus server or Grafana setup. This integration with Netdata enables users to effortlessly track performance metrics and ensure efficient resource allocation across all running instances. Discover the potential by trying out our &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">Live Demo&lt;/a>.&lt;/p></description></item><item><title>MessageBird</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/messagebird/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/messagebird/</guid><description/></item><item><title>Metaswitch Networks Ltd Formerly Data Connection Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/metaswitch-networks-ltd-formerly-data-connection-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/metaswitch-networks-ltd-formerly-data-connection-ltd-snmp-traps/</guid><description/></item><item><title>MetricFire</title><link>https://www.netdata.cloud/integrations/exporters/metricfire/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/metricfire/</guid><description/></item><item><title>Metro Ethernet Forum SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/metro-ethernet-forum-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/metro-ethernet-forum-snmp-traps/</guid><description/></item><item><title>Micom Communication Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/micom-communication-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/micom-communication-corporation-snmp-traps/</guid><description/></item><item><title>Microbursts: catching sub-second congestion that minute averages hide</title><link>https://www.netdata.cloud/guides/network/network-microbursts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-microbursts/</guid><description>&lt;h1 id="microbursts-catching-sub-second-congestion-that-minute-averages-hide">Microbursts: catching sub-second congestion that minute averages hide&lt;/h1>
&lt;p>Your switches are dropping packets. The utilization charts say everything is fine. Interface counters show moderate load, error rates are clean, and no congestion alert has fired. But applications report retransmissions, latency spikes, and intermittent connectivity. The problem resolved between polls.&lt;/p>
&lt;p>A microburst is a short, intense spike of traffic that fills a switch egress queue faster than the queue can drain. The burst may last 50 milliseconds or less, but during that window the queue overflows and packets are tail-dropped. By the time your SNMP poller arrives 60 or 300 seconds later, the burst is over, the queue has drained, and interface utilization has been averaged down to an unremarkable number.&lt;/p></description></item><item><title>Microchip Technology Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/microchip-technology-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/microchip-technology-inc-snmp-traps/</guid><description/></item><item><title>Micromuse Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/micromuse-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/micromuse-inc-snmp-traps/</guid><description/></item><item><title>Microsens GmbH Co KG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/microsens-gmbh-co-kg-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/microsens-gmbh-co-kg-snmp-traps/</guid><description/></item><item><title>Microsoft Exchange Server Monitoring</title><link>https://www.netdata.cloud/monitoring-101/msexchange-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/msexchange-monitoring/</guid><description>&lt;h2 id="microsoft-exchange-server">Microsoft Exchange Server&lt;/h2>
&lt;p>Microsoft Exchange Server is a &lt;strong>mail server and calendaring server&lt;/strong> that runs on Windows Server operating systems. &lt;a href="https://www.microsoft.com/en-gb/microsoft-365/exchange/email">It enables users to send and receive email messages, schedule appointments, store contacts, and access shared mailboxes and calendars within an organization&lt;/a>.&lt;/p>
&lt;h2 id="exchange-server-monitoring">Exchange Server Monitoring&lt;/h2>
&lt;p>To monitor Microsoft Exchange Server, you need to &lt;strong>collect and analyze various performance metrics&lt;/strong> that reflect the health and performance of your server. Some of these metrics include:&lt;/p></description></item><item><title>Microsoft SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/microsoft-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/microsoft-snmp-traps/</guid><description/></item><item><title>Microsoft SQL Server</title><link>https://www.netdata.cloud/integrations/data-collection/databases/microsoft-sql-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/microsoft-sql-server/</guid><description/></item><item><title>Microsoft SQL Server (MSSQL) Monitoring</title><link>https://www.netdata.cloud/monitoring-101/mssql-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/mssql-monitoring/</guid><description>&lt;h2 id="microsoft-sql-server">Microsoft SQL Server&lt;/h2>
&lt;p>&lt;a href="https://en.wikipedia.org/wiki/Microsoft_SQL_Server">Microsoft SQL Server&lt;/a> is a relational database management system (RDBMS) developed by Microsoft. It is a software product that stores and retrieves data as requested by other software applications. &lt;a href="https://www.microsoft.com/en-us/sql-server/sql-server-downloads">It can run on various platforms, such as Windows, Linux, and Azure&lt;/a>.&lt;/p>
&lt;h2 id="monitoring-microsoft-sql-server">Monitoring Microsoft SQL Server&lt;/h2>
&lt;p>To holistically monitor Microsoft SQL Server, you need to track various aspects of its performance, health, and availability. Some of the key metrics to monitor include system level metrics about the impact of SQL Server such as:&lt;/p></description></item><item><title>Microsoft SQL Server monitoring checklist: the signals every production instance needs</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-monitoring-checklist/</guid><description>&lt;h1 id="microsoft-sql-server-monitoring-checklist-the-signals-every-production-instance-needs">Microsoft SQL Server monitoring checklist: the signals every production instance needs&lt;/h1>
&lt;p>Most SQL Server outages are not exotic. The transaction log fills because a backup job silently stopped. A sleeping session with an open transaction blocks forty other sessions until the worker pool runs dry. TempDB runs out of space and every database on the instance stalls at once. All of these are visible hours or days in advance if you collect the right signals. Most teams do not.&lt;/p></description></item><item><title>Microsoft SQL Server monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-monitoring-maturity-model/</guid><description>&lt;h1 id="microsoft-sql-server-monitoring-maturity-model-from-survival-to-expert">Microsoft SQL Server monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most SQL Server outages are not exotic. The transaction log fills because nobody noticed log backups stopped. A head blocker sits idle with an open transaction while the worker thread pool drains. Error 825 appears in the error log for weeks before the disk actually fails. In each case, the signal was available; the monitoring just was not looking at it.&lt;/p></description></item><item><title>Microsoft Teams</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/microsoft-teams/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/microsoft-teams/</guid><description/></item><item><title>Microsoft Teams</title><link>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/microsoft-teams/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/microsoft-teams/</guid><description/></item><item><title>Microwave Data Systems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/microwave-data-systems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/microwave-data-systems-snmp-traps/</guid><description/></item><item><title>Microwave Networks Incorporated SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/microwave-networks-incorporated-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/microwave-networks-incorporated-snmp-traps/</guid><description/></item><item><title>Mikom GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mikom-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mikom-gmbh-snmp-traps/</guid><description/></item><item><title>MikroTik Licensing</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/licensing-monitoring/mikrotik-licensing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/licensing-monitoring/mikrotik-licensing/</guid><description/></item><item><title>Mikrotik Router</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/mikrotik-router/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/mikrotik-router/</guid><description/></item><item><title>Mikrotik SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mikrotik-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mikrotik-snmp-traps/</guid><description/></item><item><title>Milestone Systems A S SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/milestone-systems-a-s-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/milestone-systems-a-s-snmp-traps/</guid><description/></item><item><title>Mimosa Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mimosa-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mimosa-networks-inc-snmp-traps/</guid><description/></item><item><title>Minecraft</title><link>https://www.netdata.cloud/integrations/data-collection/applications/minecraft/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/minecraft/</guid><description/></item><item><title>Minecraft Monitoring</title><link>https://www.netdata.cloud/monitoring-101/minecraft-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/minecraft-monitoring/</guid><description>&lt;h2 id="minecraft-monitoring">Minecraft Monitoring&lt;/h2>
&lt;h3 id="what-is-minecraft">What Is Minecraft?&lt;/h3>
&lt;p>Minecraft is a sandbox video game that gained worldwide popularity due to its open-ended nature. It allows players to explore and create in a block-based world, using tools to build structures, gather resources, and craft items. With hundreds of millions of copies sold, Minecraft&amp;rsquo;s appeal spans across all age groups and makes it not just a game but a platform for creativity and learning.&lt;/p>
&lt;h3 id="monitoring-minecraft-with-netdata">Monitoring Minecraft With Netdata&lt;/h3>
&lt;p>Monitoring your Minecraft server is essential to ensure smooth gameplay and optimal performance. Netdata provides a robust solution for monitoring Minecraft using the &lt;a href="https://github.com/sladkoff/minecraft-prometheus-exporter">Minecraft Exporter&lt;/a>. By leveraging openmetrics, Netdata integrates seamlessly with any Prometheus exporter. This means you get automated dashboards, customizable alerts, and advanced analytics without needing to set up a separate Prometheus server or Grafana. This streamlined approach simplifies the monitoring process, allowing you to focus on what matters most—keeping your players happy.&lt;/p></description></item><item><title>MISCONF Redis is configured to save RDB snapshots - what it means and how to fix it</title><link>https://www.netdata.cloud/guides/redis/redis-misconf-rdb-snapshots/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-misconf-rdb-snapshots/</guid><description>&lt;h1 id="misconf-redis-is-configured-to-save-rdb-snapshots---what-it-means-and-how-to-fix-it">MISCONF Redis is configured to save RDB snapshots - what it means and how to fix it&lt;/h1>
&lt;p>Applications see &lt;code>MISCONF Redis is configured to save RDB snapshots, but it is currently not able to persist on disk&lt;/code> on every write. Reads still work, but &lt;code>SET&lt;/code>, &lt;code>HSET&lt;/code>, &lt;code>LPUSH&lt;/code>, and all mutating commands are rejected.&lt;/p>
&lt;p>Redis makes itself read-only when the last background save failed and &lt;code>stop-writes-on-bgsave-error&lt;/code> is &lt;code>yes&lt;/code> (the default). The error persists until a subsequent &lt;code>BGSAVE&lt;/code> succeeds. Retrying writes will not help. The instance is protecting you from accepting writes that can never be persisted. To recover, fix the underlying persistence failure and clear the error state with a successful save.&lt;/p></description></item><item><title>Mission Critical Software Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mission-critical-software-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mission-critical-software-inc-snmp-traps/</guid><description/></item><item><title>Mitel Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mitel-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mitel-corp-snmp-traps/</guid><description/></item><item><title>Mitsubishi Electric Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mitsubishi-electric-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mitsubishi-electric-corporation-snmp-traps/</guid><description/></item><item><title>Moca Multimedia Over Coax Alliance SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/moca-multimedia-over-coax-alliance-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/moca-multimedia-over-coax-alliance-snmp-traps/</guid><description/></item><item><title>Modbus protocol</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/modbus-protocol/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/modbus-protocol/</guid><description/></item><item><title>Modbus Protocol Monitoring</title><link>https://www.netdata.cloud/monitoring-101/modbus_rtu-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/modbus_rtu-monitoring/</guid><description>&lt;h2 id="modbus-protocol-monitoring">Modbus Protocol Monitoring&lt;/h2>
&lt;h3 id="what-is-modbus-protocol">What Is Modbus Protocol?&lt;/h3>
&lt;p>The Modbus protocol is a messaging structure widely utilized in industrial automation systems to transmit data over serial lines between electronic devices. Originally developed in 1979 by Modicon (now Schneider Electric) to communicate with its PLCs, it has evolved to become a de facto standard in industries worldwide due to its simplicity and reliability.&lt;/p>
&lt;h3 id="monitoring-modbus-protocol-with-netdata">Monitoring Modbus Protocol With Netdata&lt;/h3>
&lt;p>Effectively monitoring the Modbus protocol is crucial for maintaining the performance and reliability of industrial control systems. Netdata offers an intuitive way to monitor Modbus by integrating with an openmetrics (Prometheus) exporter. Specifically, the &lt;a href="https://github.com/dernasherbrezon/modbusrtu_exporter">modbusrtu_exporter&lt;/a> can be leveraged to collect vital metrics seamlessly.&lt;/p></description></item><item><title>MogileFS</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/mogilefs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/mogilefs/</guid><description/></item><item><title>MogileFS Monitoring</title><link>https://www.netdata.cloud/monitoring-101/mogilefs-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/mogilefs-monitoring/</guid><description>&lt;h2 id="mogilefs-monitoring">MogileFS Monitoring&lt;/h2>
&lt;h3 id="what-is-mogilefs">What Is MogileFS?&lt;/h3>
&lt;p>MogileFS is a distributed file system designed for managing large volumes of files efficiently. It offers reliable file storage across multiple servers, enabling seamless scalability and redundancy. Ideal for infrastructures that require extensive data storage solutions, MogileFS ensures data replication and distribution, making it resilient against server failures.&lt;/p>
&lt;h3 id="monitoring-mogilefs-with-netdata">Monitoring MogileFS With Netdata&lt;/h3>
&lt;p>To monitor MogileFS, Netdata utilizes a powerful tool known as an openmetrics (Prometheus) exporter. This mechanism allows Netdata to ingest data from any Prometheus exporter, providing users with automated and dynamic dashboards, notifications, and alerts—without the need for a dedicated Prometheus server or Grafana. The &lt;a href="https://github.com/KKBOX/mogilefs-exporter">MogileFS Exporter&lt;/a> is designed specifically to collect metrics that provide insights into the performance and health of MogileFS systems. These metrics are crucial for real-time monitoring, ensuring prompt detection and resolution of any issues.&lt;/p></description></item><item><title>Monet Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/monet-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/monet-systems-inc-snmp-traps/</guid><description/></item><item><title>MongoDB</title><link>https://www.netdata.cloud/integrations/data-collection/databases/mongodb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/mongodb/</guid><description/></item><item><title>MongoDB</title><link>https://www.netdata.cloud/integrations/exporters/mongodb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/mongodb/</guid><description/></item><item><title>MongoDB and Transparent Huge Pages: why THP must be disabled</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-disable-transparent-huge-pages/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-disable-transparent-huge-pages/</guid><description>&lt;h1 id="mongodb-and-transparent-huge-pages-why-thp-must-be-disabled">MongoDB and Transparent Huge Pages: why THP must be disabled&lt;/h1>
&lt;p>If you see a mongod startup warning about Transparent Huge Pages, or you are chasing tail latency spikes and memory fragmentation that do not correlate with cache pressure, ticket exhaustion, or slow queries, the cause may be THP. The standard advice has always been to disable it, but that is now version-dependent. MongoDB 8.0 upgraded TCMalloc to use per-CPU caches, which changes the recommendation for most x86_64 Linux deployments. For MongoDB 7.0 and earlier, disabling THP is still correct. For 8.0 on x86_64 Linux, THP should be enabled. Running the wrong configuration silently degrades throughput and increases latency.&lt;/p></description></item><item><title>MongoDB Authentication failed: credential rotation, brute force, and the log signal</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-authentication-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-authentication-failed/</guid><description>&lt;h1 id="mongodb-authentication-failed-credential-rotation-brute-force-and-the-log-signal">MongoDB Authentication failed: credential rotation, brute force, and the log signal&lt;/h1>
&lt;p>A spike of &lt;code>Authentication failed&lt;/code> entries in the MongoDB log is abnormal: in a healthy deployment the baseline is near zero. The signal usually means one of four things: a brute-force attack, an application presenting stale credentials after rotation, an expired x.509 certificate, or an upstream LDAP or Kerberos directory issue. Distinguishing between these quickly determines whether you are facing a security incident or an impending application outage.&lt;/p></description></item><item><title>MongoDB balancer stuck and jumbo chunks: permanent imbalance and how to fix it</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-balancer-stuck-jumbo-chunks/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-balancer-stuck-jumbo-chunks/</guid><description>&lt;h1 id="mongodb-balancer-stuck-and-jumbo-chunks-permanent-imbalance-and-how-to-fix-it">MongoDB balancer stuck and jumbo chunks: permanent imbalance and how to fix it&lt;/h1>
&lt;p>One shard is hot while the others idle. &lt;code>sh.status()&lt;/code> shows chunk counts skewed more than 20%. The balancer is either stopped or running without closing the gap. Until you fix the root cause, the imbalance persists.&lt;/p>
&lt;p>Two failure modes cause this. The balancer itself can be disabled, restricted to a narrow window, or blocked by an unhealthy config server. Or the cluster has jumbo chunks: ranges that exceed the configured chunkSize but cannot split because too many documents share the exact same shard key value. MongoDB marks those chunks &lt;code>jumbo&lt;/code> in &lt;code>config.chunks&lt;/code> and the balancer skips them. That heavy chunk pins load and storage on a single shard and creates a floor on how balanced the cluster can become.&lt;/p></description></item><item><title>MongoDB cache too small: sizing the WiredTiger cache for your working set</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-cache-undersized-working-set/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-cache-undersized-working-set/</guid><description>&lt;h1 id="mongodb-cache-too-small-sizing-the-wiredtiger-cache-for-your-working-set">MongoDB cache too small: sizing the WiredTiger cache for your working set&lt;/h1>
&lt;p>When MongoDB latency doubles and disk read IOPS climb, operators usually check indexes and the query planner first. If &lt;code>db.currentOp()&lt;/code> shows no runaway query and the slow query log is quiet, the culprit is often the WiredTiger cache.&lt;/p>
&lt;p>WiredTiger maintains its own in-memory cache of uncompressed B-tree pages, separate from the OS page cache. MongoDB defaults the cache to &lt;code>max(0.5 * (RAM - 1 GB), 256 MB)&lt;/code>. That default works for a single mongod on a dedicated host, but it breaks down in containers, multi-tenant deployments, and during organic data growth. Once the working set exceeds the cache, reads fault to disk, pages are decompressed, and eviction threads compete with application threads for CPU.&lt;/p></description></item><item><title>MongoDB checkpoint duration climbing: diagnosing slow WiredTiger checkpoints</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-checkpoint-duration-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-checkpoint-duration-high/</guid><description>&lt;h1 id="mongodb-checkpoint-duration-climbing-diagnosing-slow-wiredtiger-checkpoints">MongoDB checkpoint duration climbing: diagnosing slow WiredTiger checkpoints&lt;/h1>
&lt;p>You notice &lt;code>transaction checkpoint most recent time (msecs)&lt;/code> climbing past 10 seconds, then 30, then 50. It is trending upward, check after check, approaching the 60-second default checkpoint interval. When checkpoint duration meets or exceeds the interval, WiredTiger has no margin left. The next checkpoint starts late, dirty pages accumulate faster than they flush, and the journal can fill to the point where all new writes block until the checkpoint finishes. This is a common production failure mode that starts as a slow climb and ends as a write freeze.&lt;/p></description></item><item><title>MongoDB checkpoint stall write freeze: when all writes stop with no error</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-checkpoint-stall-write-freeze/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-checkpoint-stall-write-freeze/</guid><description>&lt;h1 id="mongodb-checkpoint-stall-write-freeze-when-all-writes-stop-with-no-error">MongoDB checkpoint stall write freeze: when all writes stop with no error&lt;/h1>
&lt;p>Writes time out or hang while &lt;code>mongod&lt;/code> is running, TCP port 27017 is open, and reads still return results from cache. The MongoDB logs are quiet, but &lt;code>db.serverStatus().opcounters&lt;/code> shows write counts frozen. This is a WiredTiger checkpoint stall: the checkpoint process fell behind, dirty pages accumulated, and new writes blocked. The freeze lasts until the current checkpoint completes. If the I/O bottleneck remains, queued writes flood through and the next checkpoint stalls again.&lt;/p></description></item><item><title>MongoDB chunk migration storms: moveChunk I/O pressure and range locks</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-chunk-migration-storms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-chunk-migration-storms/</guid><description>&lt;h1 id="mongodb-chunk-migration-storms-movechunk-io-pressure-and-range-locks">MongoDB chunk migration storms: moveChunk I/O pressure and range locks&lt;/h1>
&lt;p>Write latency spikes across multiple shards simultaneously. Queries time out while &lt;code>mongos&lt;/code> logs show no election events. Donor shard primaries report growing queue depths, and &lt;code>config.changelog&lt;/code> shows a wall of &lt;code>moveChunk.error&lt;/code> &lt;!-- TODO: verify exact changelog event name for migration failures --> entries interleaved with retries. The balancer is hammering the cluster instead of helping.&lt;/p>
&lt;p>A single &lt;code>moveChunk&lt;/code> operation is expensive. It copies documents to the recipient, enters a critical section with a &lt;!-- TODO: verify exact lock mode during critical section --> shared range lock on the donor, then deletes orphaned documents. When the balancer triggers many migrations in quick succession, or individual migrations fail and retry in a tight loop, overlapping I/O pressure and lock contention compound into a cluster-wide latency event.&lt;/p></description></item><item><title>MongoDB connection churn: high totalCreated rate and thread creation overhead</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-connection-churn/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-connection-churn/</guid><description>&lt;h1 id="mongodb-connection-churn-high-totalcreated-rate-and-thread-creation-overhead">MongoDB connection churn: high totalCreated rate and thread creation overhead&lt;/h1>
&lt;p>&lt;code>db.serverStatus().connections&lt;/code> can show low &lt;code>current&lt;/code> and a rapidly climbing &lt;code>totalCreated&lt;/code>. That mismatch is connection churn: connections open and close rapidly instead of being reused. MongoDB uses a thread-per-connection model, so each cycle costs roughly a megabyte of thread stack, scheduling overhead, and file descriptor work. The result is rising RSS, CPU contention, and latency spikes that do not correlate with the active connection count.&lt;/p></description></item><item><title>MongoDB connection refused at maxIncomingConnections: hitting the connection ceiling</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-connection-limit-reached/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-connection-limit-reached/</guid><description>&lt;h1 id="mongodb-connection-refused-at-maxincomingconnections-hitting-the-connection-ceiling">MongoDB connection refused at maxIncomingConnections: hitting the connection ceiling&lt;/h1>
&lt;p>Application logs show connection timeouts. MongoDB logs show &lt;code>connection refused&lt;/code> or &lt;code>error accepting new connection&lt;/code>. &lt;code>db.serverStatus().connections&lt;/code> shows &lt;code>current&lt;/code> well below the configured maximum. This disconnect means you are hitting a hard ceiling at the TCP accept layer, not experiencing gradual degradation.&lt;/p>
&lt;p>MongoDB uses a one-thread-per-connection model. Each accepted connection consumes at least one file descriptor and a thread stack sized by the OS &lt;code>ulimit -s&lt;/code>. While &lt;code>maxIncomingConnections&lt;/code> sets the logical inbound cap, the OS file-descriptor limit (&lt;code>ulimit -n&lt;/code>) usually enforces the actual ceiling. Rejections happen before the connection handshake completes, so &lt;code>serverStatus().connections.current&lt;/code> never counts refused connections. Look at logs, ratios, and OS-level resource counts to find the real limit.&lt;/p></description></item><item><title>MongoDB connection storm spiral: reconnection floods after an election or deploy</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-connection-storm-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-connection-storm-spiral/</guid><description>&lt;h1 id="mongodb-connection-storm-spiral-reconnection-floods-after-an-election-or-deploy">MongoDB connection storm spiral: reconnection floods after an election or deploy&lt;/h1>
&lt;p>Connection count on a primary jumps from 200 to 4,000 in under a minute. Resident memory climbs, query latencies double, and application logs fill with timeout errors. The slow query log shows nothing unusual. Individual queries are not the problem. The database is drowning in threads.&lt;/p>
&lt;p>This is a connection storm spiral. A trigger event, usually a replica set election, application deploy, or network blip, invalidates existing connections across your application fleet. Every driver reconnects at once. Each new connection costs MongoDB a dedicated thread and roughly 1 MB of stack memory &lt;!-- TODO: verify per-connection committed memory on target platform; default thread stack reservation is often 8-10 MB -->. The resulting RSS spike and ticket contention slow down operations already in flight, causing more timeouts, which drives even more reconnections. The feedback loop ends in OOM kill or unresponsiveness.&lt;/p></description></item><item><title>MongoDB disk full: emergency recovery when mongod can't write the journal</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-disk-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-disk-full/</guid><description>&lt;h1 id="mongodb-disk-full-emergency-recovery-when-mongod-cant-write-the-journal">MongoDB disk full: emergency recovery when mongod can&amp;rsquo;t write the journal&lt;/h1>
&lt;p>When the filesystem backing the data or journal directory crosses a critical threshold, WiredTiger cannot allocate new journal extents. If mongod crashes or restarts, recovery replays journal files since the last checkpoint and requires free headroom to create or extend files during that replay. On a full disk, mongod hangs in recovery without binding to port 27017.&lt;/p>
&lt;p>If the node is a standalone, there is no replica to fail over to. If it is a secondary, cluster redundancy is reduced while the member is down. Recovery is complicated by a counterintuitive storage engine behavior: WiredTiger reclaims space internally after deletes, but does not automatically shrink data files or return bytes to the operating system. A volume that reads 99% full after a massive delete remains 99% full at the filesystem level.&lt;/p></description></item><item><title>MongoDB disk I/O saturation: correlating iostat with WiredTiger signals</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-disk-io-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-disk-io-saturation/</guid><description>&lt;h1 id="mongodb-disk-io-saturation-correlating-iostat-with-wiredtiger-signals">MongoDB disk I/O saturation: correlating iostat with WiredTiger signals&lt;/h1>
&lt;p>When &lt;code>opLatencies.writes&lt;/code> climbs and &lt;code>globalLock.currentQueue&lt;/code> grows, &lt;code>db.serverStatus().wiredTiger.transaction&lt;/code> often shows the most recent checkpoint took 45 seconds. WiredTiger metrics tell you &lt;em>what&lt;/em> is hurting, but they do not tell you &lt;em>why&lt;/em>. The next question is whether the disk is actually saturated.&lt;/p>
&lt;p>Disk I/O saturation surfaces as climbing journal sync latency, checkpoint duration exceeding the 60-second interval, application-thread evictions, and ticket exhaustion. The only way to separate a storage problem from a query problem is to correlate OS-level disk signals (&lt;code>iostat -x&lt;/code>) with WiredTiger internal signals in the same time window. This guide shows how to do that safely during an incident.&lt;/p></description></item><item><title>MongoDB exceeded memory limit for $group — aggregation spills and allowDiskUse</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-exceeded-memory-limit-group-sort/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-exceeded-memory-limit-group-sort/</guid><description>&lt;h1 id="mongodb-exceeded-memory-limit-for-group--aggregation-spills-and-allowdiskuse">MongoDB exceeded memory limit for $group — aggregation spills and allowDiskUse&lt;/h1>
&lt;p>Application logs show error 16945, or an aggregation pipeline slows by an order of magnitude. In MongoDB, every aggregation stage not backed by an index is limited to 100 megabytes of RAM. When a stage exceeds this limit and disk spilling is not enabled, the operation fails immediately. If spilling is enabled, MongoDB writes temporary files to disk, which keeps the pipeline alive but adds unpredictable latency and extra I/O load.&lt;/p></description></item><item><title>MongoDB exposed to the internet without authentication: bindIp and the breach scenario</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-exposed-without-auth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-exposed-without-auth/</guid><description>&lt;h1 id="mongodb-exposed-to-the-internet-without-authentication-bindip-and-the-breach-scenario">MongoDB exposed to the internet without authentication: bindIp and the breach scenario&lt;/h1>
&lt;p>A &lt;code>mongod&lt;/code> process bound to all interfaces with authentication disabled exposes every database to any host that can reach port 27017. If you are responding to a scan, PAGE, or audit, confirm the exposure, measure the blast radius, and eliminate the surface.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>MongoDB&amp;rsquo;s &lt;code>net.bindIp&lt;/code> controls which interfaces accept connections. Modern packages default &lt;code>bindIp&lt;/code> to &lt;code>127.0.0.1&lt;/code>; exposure usually follows an explicit override to a wildcard such as &lt;code>0.0.0.0&lt;/code> or &lt;code>::&lt;/code>. Without authentication, any reachable host can list databases, read or write documents, and execute administrative commands.&lt;/p></description></item><item><title>MongoDB flow control throttling writes: when the primary slows itself down</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-flow-control-throttling-writes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-flow-control-throttling-writes/</guid><description>&lt;h1 id="mongodb-flow-control-throttling-writes-when-the-primary-slows-itself-down">MongoDB flow control throttling writes: when the primary slows itself down&lt;/h1>
&lt;p>Write throughput on the primary has dropped by 30% or more. Application logs show intermittent write latency spikes, but the primary&amp;rsquo;s CPU, memory, and disk metrics look healthy. There are no elections, cache pressure warnings, or obvious errors. Check replication: secondaries are lagging. The primary is not sick; it is throttling itself.&lt;/p>
&lt;p>MongoDB 4.2 introduced flow control, a ticket-based admission mechanism that caps the primary&amp;rsquo;s write rate to keep secondaries from falling off the oplog. When &lt;code>isLagged&lt;/code> is true, the primary artificially limits throughput. The fix is rarely on the primary. Look at the replication pipeline: a slow secondary, an oplog window that is too small, or a topology change that distorts majority commit lag.&lt;/p></description></item><item><title>MongoDB globalLock.currentQueue growing: operations queuing for the storage engine</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-queue-depth-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-queue-depth-growing/</guid><description>&lt;h1 id="mongodb-globallockcurrentqueue-growing-operations-queuing-for-the-storage-engine">MongoDB globalLock.currentQueue growing: operations queuing for the storage engine&lt;/h1>
&lt;p>&lt;code>globalLock.currentQueue.total&lt;/code> is climbing and not stabilizing after a burst. Under WiredTiger, this metric reflects waits at the intent lock and ticket admission layers, not a single global mutex. The engine is receiving work faster than it can drain it.&lt;/p>
&lt;p>If the total queue sustains above 20 and grows, the system is saturated. Queued operations become latency spikes, then connection pileups as applications retry, then a degradation spiral. The root cause is usually ticket exhaustion, a single hot collection, a long-running operation, or storage I/O degradation inflating ticket hold times.&lt;/p></description></item><item><title>MongoDB journal sync latency high: the storage signal that warns 60 seconds early</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-journal-sync-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-journal-sync-latency-high/</guid><description>&lt;h1 id="mongodb-journal-sync-latency-high-the-storage-signal-that-warns-60-seconds-early">MongoDB journal sync latency high: the storage signal that warns 60 seconds early&lt;/h1>
&lt;p>Application write latency spikes. Connections pile up. Look back 60 seconds and WiredTiger journal sync latency was likely already climbing. Every write with &lt;code>j:true&lt;/code> or &lt;code>w:&amp;quot;majority&amp;quot;&lt;/code> blocks until the journal buffer is fsynced to disk. When storage struggles, journal sync is the first domino to fall.&lt;/p>
&lt;p>Journal sync latency is a storage subsystem signal, not a query or cache problem. The block device under &lt;code>mongod&lt;/code> cannot absorb small sequential writes fast enough. The result is head-of-line delay for all durable writes, which cascades into ticket exhaustion and connection backlog.&lt;/p></description></item><item><title>MongoDB lock wait times: collection and metadata lock contention during DDL</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-lock-wait-times/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-lock-wait-times/</guid><description>&lt;h1 id="mongodb-lock-wait-times-collection-and-metadata-lock-contention-during-ddl">MongoDB lock wait times: collection and metadata lock contention during DDL&lt;/h1>
&lt;p>When p99 latency jumps and &lt;code>globalLock.currentQueue&lt;/code> grows, check &lt;code>serverStatus().locks&lt;/code>. If &lt;code>timeAcquiringMicros&lt;/code> is climbing for &lt;code>Collection&lt;/code> or &lt;code>Metadata&lt;/code>, the cause is almost always DDL: &lt;code>createIndexes&lt;/code>, &lt;code>dropIndexes&lt;/code>, &lt;code>collMod&lt;/code>, &lt;code>renameCollection&lt;/code>, or similar commands that acquire exclusive collection, database, or metadata locks. WiredTiger uses document-level concurrency for ordinary reads and writes, so normal CRUD rarely blocks. A single schema change can serialize operations on a hot collection or across a database during peak traffic.&lt;/p></description></item><item><title>MongoDB long-running operations: finding and killing the query holding a ticket</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-long-running-operations/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-long-running-operations/</guid><description>&lt;h1 id="mongodb-long-running-operations-finding-and-killing-the-query-holding-a-ticket">MongoDB long-running operations: finding and killing the query holding a ticket&lt;/h1>
&lt;p>Your application latency just spiked. &lt;code>opLatencies&lt;/code> show reads and writes climbing. &lt;code>globalLock.currentQueue&lt;/code> is no longer zero. You check &lt;code>db.serverStatus().wiredTiger.concurrentTransactions&lt;/code>: available tickets are near zero, but throughput has not increased. An operation is holding a ticket without making progress.&lt;/p>
&lt;p>A collection scan, an unbounded aggregation, or a stalled write can hold a WiredTiger read or write ticket for minutes. The default is 128 read and 128 write tickets in most versions, so one long-running operation can cascade into system-wide queuing, connection pileup, and application timeouts. Find it and kill it, but killing the wrong operation can crash a node or leave data inconsistent.&lt;/p></description></item><item><title>MongoDB long-running transactions: pinned snapshots and silent cache pressure</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-long-running-transactions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-long-running-transactions/</guid><description>&lt;h1 id="mongodb-long-running-transactions-pinned-snapshots-and-silent-cache-pressure">MongoDB long-running transactions: pinned snapshots and silent cache pressure&lt;/h1>
&lt;p>MongoDB latency climbs while write throughput looks normal and the working set is unchanged. Check WiredTiger cache: the dirty ratio is rising and application threads have started evicting pages. The cause is often invisible in standard cache metrics &amp;ndash; a multi-document transaction that opened a snapshot and never let it go.&lt;/p>
&lt;p>Multi-document transactions pin WiredTiger snapshots until commit or abort. That snapshot prevents the storage engine from evicting old document versions. A single transaction left open for minutes can silently pin enough pages to push the cache into eviction pressure. An inactive-but-open transaction is worse than an active one because it holds the snapshot indefinitely with no forward progress.&lt;/p></description></item><item><title>MongoDB Monitoring</title><link>https://www.netdata.cloud/monitoring-101/mongodb-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/mongodb-monitoring/</guid><description>&lt;h2 id="mongodb-monitoring">MongoDB Monitoring&lt;/h2>
&lt;h3 id="what-is-mongodb">What Is MongoDB?&lt;/h3>
&lt;p>MongoDB is a leading NoSQL database platform designed for flexibility, scalability, and performance. It is used to store documents in a flexible, JSON-like format, which makes it perfect for handling large volumes of unstructured data. For more insights, check out &lt;a href="https://www.mongodb.com/">MongoDB&amp;rsquo;s official site&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-mongodb-with-netdata">Monitoring MongoDB With Netdata&lt;/h3>
&lt;p>Netdata provides comprehensive monitoring for MongoDB, allowing users to gain real-time insights into their MongoDB servers. By utilizing the &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/mongodb/">MongoDB monitoring tool from Netdata&lt;/a>, users can track critical metrics and enhance their troubleshooting capabilities.&lt;/p></description></item><item><title>MongoDB monitoring checklist: the signals every production cluster needs</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-monitoring-checklist/</guid><description>&lt;h1 id="mongodb-monitoring-checklist-the-signals-every-production-cluster-needs">MongoDB monitoring checklist: the signals every production cluster needs&lt;/h1>
&lt;p>Production MongoDB failures are preceded by signals that are visible but often unmonitored: climbing dirty cache ratio, shrinking oplog window, or ticket counts approaching zero. This guide organizes essential signals into four monitoring levels. Use them to audit instrumentation or triage gaps during an incident.&lt;/p>
&lt;p>Each level builds on the previous one. If you are missing a survival signal, instrument it before adding expert metrics. The thresholds below are drawn from the MongoDB &lt;code>serverStatus()&lt;/code> and &lt;code>rs.status()&lt;/code> contract and from operational patterns observed across WiredTiger deployments.&lt;/p></description></item><item><title>MongoDB monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-monitoring-maturity-model/</guid><description>&lt;h1 id="mongodb-monitoring-maturity-model-from-survival-to-expert">MongoDB monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>If you run MongoDB in production, you need a progression: what to watch first to know the database is alive, what to add next to know why it is slowing down, and what to track finally to predict a cascade before it becomes an outage.&lt;/p>
&lt;p>Treat the levels below as gates, not a shopping list. Automate and alert on Level 1 before you build Level 2 dashboards. If you are instrumenting Level 4 but lack a reliable page for disk space or member state, you have inverted your priorities.&lt;/p></description></item><item><title>MongoDB no primary / election storm: repeated elections and write outages</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-no-primary-election-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-no-primary-election-storm/</guid><description>&lt;h1 id="mongodb-no-primary--election-storm-repeated-elections-and-write-outages">MongoDB no primary / election storm: repeated elections and write outages&lt;/h1>
&lt;p>Applications log &amp;ldquo;not primary&amp;rdquo; errors. &lt;code>rs.status()&lt;/code> shows a different &lt;code>PRIMARY&lt;/code> than thirty seconds ago. MongoDB logs repeat &lt;code>&amp;quot;Starting an election&amp;quot;&lt;/code> and &lt;code>&amp;quot;Stepping down&amp;quot;&lt;/code>. Each election costs 2-12 seconds of write unavailability. More than two in ten minutes is an election storm.&lt;/p>
&lt;p>This pattern is more dangerous than a single failover because it creates rolling write outages that do not self-stabilize. Drivers reconnect, retry buffers fill, and application latency degrades even when a primary exists. Root causes usually fall into three categories: the primary is too slow to answer heartbeats, the network is dropping or delaying packets between members, or a misconfigured priority is forcing a healthy primary to step down.&lt;/p></description></item><item><title>MongoDB not master error: writes hitting a non-primary node after failover</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-not-master-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-not-master-error/</guid><description>&lt;h1 id="mongodb-not-master-error-writes-hitting-a-non-primary-node-after-failover">MongoDB not master error: writes hitting a non-primary node after failover&lt;/h1>
&lt;p>A node restart, network partition, or planned stepdown triggers a MongoDB election. Seconds later, application logs show &lt;code>NotWritablePrimary&lt;/code> (code 10107) or the legacy string &lt;code>not master and slaveOk=false&lt;/code>. Writes fail against a node that used to be PRIMARY, even though the cluster has elected a new one.&lt;/p>
&lt;p>This guide covers how to find the root cause and stop it from recurring.&lt;/p></description></item><item><title>MongoDB not primary and secondaryOk=false: reading from a secondary and how to fix it</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-not-primary-and-secondaryok-false/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-not-primary-and-secondaryok-false/</guid><description>&lt;h1 id="mongodb-not-primary-and-secondaryokfalse-reading-from-a-secondary-and-how-to-fix-it">MongoDB not primary and secondaryOk=false: reading from a secondary and how to fix it&lt;/h1>
&lt;p>Your application logs show &lt;code>NotPrimaryNoSecondaryOk&lt;/code> (code 13435) with the message &lt;code>&amp;quot;not master and slaveOk=false&amp;quot;&lt;/code>. Metrics show read failures against a specific host. The &lt;code>mongod&lt;/code> process is running, replica set heartbeats are clean, and replication lag looks normal. The cluster is not down. The error is a routing decision: a client sent a read to a replica set member that is not the primary, without declaring that reading from a non-primary is acceptable.&lt;/p></description></item><item><title>MongoDB noTimeout cursors causing cache pressure: pinned snapshots and silent eviction stalls</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-notimeout-cursors-cache-pressure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-notimeout-cursors-cache-pressure/</guid><description>&lt;h1 id="mongodb-notimeout-cursors-causing-cache-pressure-pinned-snapshots-and-silent-eviction-stalls">MongoDB noTimeout cursors causing cache pressure: pinned snapshots and silent eviction stalls&lt;/h1>
&lt;p>When &lt;code>wiredTiger.cache.bytes currently in the cache&lt;/code> climbs, the dirty ratio trends toward 20%, and read latencies spike without a single slow query in the log, check &lt;code>metrics.cursor.open.noTimeout&lt;/code>.&lt;/p>
&lt;p>Each noTimeout cursor pins a WiredTiger snapshot indefinitely. Old document versions cannot be evicted while that snapshot is open, so the cache fills with unreachable history until background eviction falls behind and application threads are forced to clean up. The result is a silent cache pressure cascade that looks like a capacity problem but is actually a cursor lifecycle problem.&lt;/p></description></item><item><title>MongoDB OOM-killed by the kernel: RSS, cache sizing, and oom_score_adj</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-oom-killed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-oom-killed/</guid><description>&lt;h1 id="mongodb-oom-killed-by-the-kernel-rss-cache-sizing-and-oom_score_adj">MongoDB OOM-killed by the kernel: RSS, cache sizing, and oom_score_adj&lt;/h1>
&lt;p>You find &lt;code>mongod&lt;/code> gone. The replica set has no primary. Applications time out. MongoDB logs show no graceful shutdown. Instead, &lt;code>dmesg&lt;/code> shows &lt;code>Out of memory: Killed process 12345 (mongod)&lt;/code>. The Linux OOM killer has reaped the process. MongoDB is a frequent target because its resident set size is usually the largest on the host.&lt;/p>
&lt;p>An OOM kill is not a MongoDB bug. It is the kernel freeing RAM by terminating the highest-scoring process. mongod&amp;rsquo;s RSS is dominated by the WiredTiger cache, plus roughly 1 MB per connection, plus roughly 500 MB to 1 GB of internal overhead for indexes, session buffers, and stack. When that sum comes within 1 GB of total RAM, the node is in the danger zone. The kill is abrupt: no stepdown, no replica set coordination, and after restart the cache must warm again.&lt;/p></description></item><item><title>MongoDB operation exceeded time limit (MaxTimeMSExpired): maxTimeMS and killed operations</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-operation-exceeded-time-limit/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-operation-exceeded-time-limit/</guid><description>&lt;h1 id="mongodb-operation-exceeded-time-limit-maxtimemsexpired-maxtimems-and-killed-operations">MongoDB operation exceeded time limit (MaxTimeMSExpired): maxTimeMS and killed operations&lt;/h1>
&lt;p>Error code 50, &lt;code>MaxTimeMSExpired&lt;/code>, means the server killed an operation that exceeded its processing budget. Raising the timeout without fixing the root cause turns acute failures into chronic resource exhaustion. The operation was already pathologically slow; &lt;code>maxTimeMS&lt;/code> ended it before it consumed more resources or held locks and tickets indefinitely.&lt;/p>
&lt;p>&lt;code>maxTimeMS&lt;/code> sets a cumulative processing budget in milliseconds. MongoDB enforces it using the same interrupt mechanism as &lt;code>killOp&lt;/code>, terminating the operation only at designated interrupt points. Idle time between cursor batches does not count toward the limit, and on direct connections network latency is excluded from the server-side clock. On sharded clusters, however, latency between &lt;code>mongos&lt;/code> and shard &lt;code>mongod&lt;/code> instances counts against the limit. Distinguish a true &lt;code>MaxTimeMSExpired&lt;/code> from a client-side socket timeout, where the client gives up before the server responds.&lt;/p></description></item><item><title>MongoDB oplog window collapse: secondaries falling off and forced full resync</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-oplog-window-collapse/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-oplog-window-collapse/</guid><description>&lt;h1 id="mongodb-oplog-window-collapse-secondaries-falling-off-and-forced-full-resync">MongoDB oplog window collapse: secondaries falling off and forced full resync&lt;/h1>
&lt;p>A secondary transitions to RECOVERING and logs &amp;ldquo;too stale to catch up.&amp;rdquo; The oplog window compresses from 48 hours to 90 minutes while replication lag on one secondary climbs steadily. These are the signatures of oplog window collapse: a write surge turns over the oplog faster than secondaries can consume it, and the safety margin between window and lag evaporates.&lt;/p></description></item><item><title>MongoDB oplog window too small: sizing the oplog for your write volume</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-oplog-window-too-small/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-oplog-window-too-small/</guid><description>&lt;h1 id="mongodb-oplog-window-too-small-sizing-the-oplog-for-your-write-volume">MongoDB oplog window too small: sizing the oplog for your write volume&lt;/h1>
&lt;p>The oplog window is the only thing standing between a routine secondary restart and a multi-hour full initial sync. It is a fixed-size capped collection that stores a variable amount of history. As your write volume grows, the window compresses. Most teams size the oplog once during initial deployment and never look at it again. Six months later, a routine maintenance window turns into an incident because the secondary fell off the oplog, entered RECOVERING, and forced a resync that saturated the remaining nodes.&lt;/p></description></item><item><title>MongoDB page faults high: working set exceeding memory after warmup</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-page-faults-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-page-faults-high/</guid><description>&lt;h1 id="mongodb-page-faults-high-working-set-exceeding-memory-after-warmup">MongoDB page faults high: working set exceeding memory after warmup&lt;/h1>
&lt;p>Hard page faults long after startup mean the active data set exceeds resident memory. On Linux, &lt;code>extra_info.page_faults&lt;/code> counts major faults: the OS read data from disk because the page was missing from both the WiredTiger cache and the OS page cache. A brief spike after restart is normal during warmup, but sustained faults mean the working set does not fit. On EBS gp3, 50 faults per second can degrade latency. On NVMe, hundreds per second may be tolerable, but neither is free. Confirm the cause, distinguish warmup from pressure, and reduce the fault rate without guessing.&lt;/p></description></item><item><title>MongoDB pages evicted by application threads: when eviction becomes user latency</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-application-thread-evictions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-application-thread-evictions/</guid><description>&lt;h1 id="mongodb-pages-evicted-by-application-threads-when-eviction-becomes-user-latency">MongoDB pages evicted by application threads: when eviction becomes user latency&lt;/h1>
&lt;p>Query p99 latency doubles or triples, but &lt;code>iostat&lt;/code> is not saturated and the slow query log shows no single offender. The signal is in &lt;code>db.serverStatus().wiredTiger.cache&lt;/code>: &lt;code>pages evicted by application threads&lt;/code> has moved from zero to a sustained nonzero rate.&lt;/p>
&lt;p>This metric marks the moment when WiredTiger&amp;rsquo;s dedicated eviction workers fall behind and application threads are drafted to do the work. Any sustained nonzero rate is abnormal. Once application threads evict, they perform page reconciliation and disk I/O inline with the request handler thread, directly inflating user-visible latency. The companion counter &lt;code>pages selected for eviction unable to be evicted&lt;/code> means eviction is stalled and the cache is effectively frozen.&lt;/p></description></item><item><title>MongoDB replica set member unhealthy: reading rs.status() states</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-replica-set-member-unhealthy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-replica-set-member-unhealthy/</guid><description>&lt;h1 id="mongodb-replica-set-member-unhealthy-reading-rsstatus-states">MongoDB replica set member unhealthy: reading rs.status() states&lt;/h1>
&lt;p>&lt;code>rs.status()&lt;/code> output is perspective-dependent and easy to misread. A member can show &lt;code>health: 1&lt;/code> while &lt;code>RECOVERING&lt;/code> and unable to serve reads, or appear &lt;code>UNKNOWN&lt;/code> from one node yet &lt;code>SECONDARY&lt;/code> from another because of an asymmetric firewall rule. &lt;code>health: 0&lt;/code> alone does not mean the process is dead, and &lt;code>health: 1&lt;/code> alone does not mean the node is healthy.&lt;/p>
&lt;p>This guide maps &lt;code>stateStr&lt;/code> values to failure modes and shows how to distinguish transient startup states, replication lag spirals, and network partitions.&lt;/p></description></item><item><title>MongoDB replication lag: detection, diagnosis, and fixes</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-replication-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-replication-lag/</guid><description>&lt;h1 id="mongodb-replication-lag-detection-diagnosis-and-fixes">MongoDB replication lag: detection, diagnosis, and fixes&lt;/h1>
&lt;p>Replication lag is the delay between the primary&amp;rsquo;s latest oplog entry and a secondary&amp;rsquo;s last applied entry. When lag grows faster than the oplog window, the secondary cannot catch up and requires a full initial sync, which can take hours to days. Until that happens, every second of lag erodes failover safety and read consistency.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>MongoDB replicates by having secondaries tail the primary&amp;rsquo;s oplog, a capped collection in &lt;code>local.oplog.rs&lt;/code>. The primary records every write with a timestamp; the secondary fetches entries and applies them locally. Lag is the delta between the primary&amp;rsquo;s &lt;code>optimeDate&lt;/code> and the secondary&amp;rsquo;s &lt;code>optimeDate&lt;/code> from &lt;code>rs.status()&lt;/code>.&lt;/p></description></item><item><title>MongoDB rollback after failover: silent data loss and the rollback directory</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-rollback-after-failover/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-rollback-after-failover/</guid><description>&lt;h1 id="mongodb-rollback-after-failover-silent-data-loss-and-the-rollback-directory">MongoDB rollback after failover: silent data loss and the rollback directory&lt;/h1>
&lt;p>A replica set member in &lt;code>ROLLBACK&lt;/code> state, or an application reporting vanished documents after failover, means a former primary held writes that never reached a majority. When that node rejoins, MongoDB erases the divergent history and writes the removed data to files under &lt;code>&amp;lt;dbPath&amp;gt;/rollback/&lt;/code>. The application may have received acknowledgment for those writes. With &lt;code>w:1&lt;/code>, acknowledgment meant only that the primary applied the write. It did not guarantee replication to a majority or survival through failover. That is silent data loss.&lt;/p></description></item><item><title>MongoDB RSS growing without cache growth: leaks, threads, and tcmalloc fragmentation</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-memory-rss-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-memory-rss-growing/</guid><description>&lt;h1 id="mongodb-rss-growing-without-cache-growth-leaks-threads-and-tcmalloc-fragmentation">MongoDB RSS growing without cache growth: leaks, threads, and tcmalloc fragmentation&lt;/h1>
&lt;p>&lt;code>db.serverStatus().mem.resident&lt;/code> climbs while WiredTiger cache utilization stays flat and the host is not swapping. Virtual memory is larger than RSS by design and is not an alert target. Only RSS reflects physical memory pressure. When RSS grows without cache growth, the problem lives outside the storage engine.&lt;/p>
&lt;p>This pattern points to one of three areas: tcmalloc heap retention and fragmentation, per-connection thread stack accumulation, or unbounded internal allocations from cursors, plan caches, or aggregation pipelines. Each connection reserves roughly 1 MB of stack space, so a connection storm can add gigabytes of RSS in minutes. TCMalloc caches freed memory in per-thread or per-CPU arenas, which inflates RSS independently of the WiredTiger cache.&lt;/p></description></item><item><title>MongoDB scanned objects ratio high: documents examined vs returned and wasted work</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-scanned-objects-ratio-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-scanned-objects-ratio-high/</guid><description>&lt;h1 id="mongodb-scanned-objects-ratio-high-documents-examined-vs-returned-and-wasted-work">MongoDB scanned objects ratio high: documents examined vs returned and wasted work&lt;/h1>
&lt;p>A high scanned-objects-to-returned-documents ratio means MongoDB examines far more documents than it returns. The alert usually fires as an Atlas query targeting event or a chart showing &lt;code>metrics.queryExecutor.scannedObjects&lt;/code> climbing faster than &lt;code>metrics.document.returned&lt;/code>. A ratio near 1:1 is efficient. A ratio above 1000:1 means the database examines roughly a thousand documents for every one it returns. Wasted work consumes read tickets, pushes unnecessary data through the WiredTiger cache, and often precedes latency spikes or ticket exhaustion.&lt;/p></description></item><item><title>MongoDB scatter-gather queries: when mongos fans out to every shard</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-scatter-gather-queries/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-scatter-gather-queries/</guid><description>&lt;h1 id="mongodb-scatter-gather-queries-when-mongos-fans-out-to-every-shard">MongoDB scatter-gather queries: when mongos fans out to every shard&lt;/h1>
&lt;p>In a sharded MongoDB cluster, mongos reads chunk-to-shard mappings from the config servers and routes each query to the smallest set of shards that can satisfy it. When the predicate includes the shard key, mongos targets one shard or a bounded subset. When it does not, mongos broadcasts the query to every shard that owns chunks for the collection and merges the results. This is a scatter-gather query.&lt;/p></description></item><item><title>MongoDB sharding hot shard: one shard saturating while others idle</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-sharding-hot-shard/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-sharding-hot-shard/</guid><description>&lt;h1 id="mongodb-sharding-hot-shard-one-shard-saturating-while-others-idle">MongoDB sharding hot shard: one shard saturating while others idle&lt;/h1>
&lt;p>Application latency climbs while aggregate CPU, I/O, and connection counts look healthy. Drill into individual shards and one node is pinned near 95% utilization while peers sit near 20%. &lt;code>sh.status()&lt;/code> shows an even chunk distribution. You have a hot shard.&lt;/p>
&lt;p>A hot shard occurs when one member of a sharded cluster receives a disproportionate share of read or write traffic. The cause is usually an access-pattern mismatch against the shard key, not a chunk-count imbalance. The overloaded shard saturates its WiredTiger cache and exhausts read or write tickets, pushing operations into queues. Because other shards are idle, aggregate cluster metrics hide the problem until the hot shard cascades into a latency spike or availability event.&lt;/p></description></item><item><title>MongoDB silent index regression: when a dropped index quietly becomes a collection scan</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-silent-index-regression/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-silent-index-regression/</guid><description>&lt;h1 id="mongodb-silent-index-regression-when-a-dropped-index-quietly-becomes-a-collection-scan">MongoDB silent index regression: when a dropped index quietly becomes a collection scan&lt;/h1>
&lt;p>Read latency on the primary doubles while connection counts and write throughput stay flat. There are no election events or cache pressure alerts. Traffic is unchanged. Yet p99 read latency climbs until operations time out.&lt;/p>
&lt;p>The slow query log shows queries that used to finish in milliseconds now taking seconds. The plans show &lt;code>COLLSCAN&lt;/code>. An index that existed last week is gone, or the query planner switched to a less efficient index after a cache invalidation. Because queries still return correct results, the regression is silent until it becomes an outage.&lt;/p></description></item><item><title>MongoDB slow query COLLSCAN: collection scans and the missing index</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-slow-query-collscan/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-slow-query-collscan/</guid><description>&lt;h1 id="mongodb-slow-query-collscan-collection-scans-and-the-missing-index">MongoDB slow query COLLSCAN: collection scans and the missing index&lt;/h1>
&lt;p>Queries that used to return in tens of milliseconds now breach application timeouts. Read latency climbs while write throughput stays flat. In the MongoDB slow query log, you see &lt;code>planSummary: &amp;quot;COLLSCAN&amp;quot;&lt;/code> attached to operations that should be indexed. A collection scan reads documents that will never be returned, wastes disk I/O, floods the WiredTiger cache with irrelevant data, and holds read tickets until the whole instance cascades into cache pressure.&lt;/p></description></item><item><title>MongoDB storage not reclaimed after delete: WiredTiger, compact, and resync</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-storage-not-reclaimed-after-delete/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-storage-not-reclaimed-after-delete/</guid><description>&lt;h1 id="mongodb-storage-not-reclaimed-after-delete-wiredtiger-compact-and-resync">MongoDB storage not reclaimed after delete: WiredTiger, compact, and resync&lt;/h1>
&lt;p>You deleted a large portion of a collection, and &lt;code>db.collection.stats()&lt;/code> confirms the logical size dropped. &lt;code>df -h&lt;/code> on the data volume does not. The filesystem blocks were not returned. WiredTiger maintains internal free-space lists and reclaims pages for future writes rather than shrinking files. Dropping a database or collection removes the underlying files and returns space to the OS immediately; this article addresses the case where documents are deleted but the collection files remain. This guide explains how to verify the state and the two operational paths to return blocks to the OS: the &lt;code>compact&lt;/code> command and replica set member resync.&lt;/p></description></item><item><title>MongoDB swapping: why mongod must never swap and how to tune the OS</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-swapping-and-swappiness/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-swapping-and-swappiness/</guid><description>&lt;h1 id="mongodb-swapping-why-mongod-must-never-swap-and-how-to-tune-the-os">MongoDB swapping: why mongod must never swap and how to tune the OS&lt;/h1>
&lt;p>Application timeouts climb. MongoDB latency jumps from milliseconds to seconds or minutes, yet &lt;code>mongod&lt;/code> is still running and accepting connections. CPU is low, disk I/O is not saturated, and the MongoDB log shows no errors. The process has not crashed. It has entered swap death. When the Linux kernel evicts mongod pages to swap, the database continues to function at roughly 1/1000th of normal speed. MongoDB relies on the WiredTiger cache and OS page cache to remain resident in RAM.&lt;/p></description></item><item><title>MongoDB ticket exhaustion: WiredTiger read/write tickets and queued operations</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-ticket-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-ticket-exhaustion/</guid><description>&lt;h1 id="mongodb-ticket-exhaustion-wiredtiger-readwrite-tickets-and-queued-operations">MongoDB ticket exhaustion: WiredTiger read/write tickets and queued operations&lt;/h1>
&lt;p>Your application times out while the OS shows idle CPU and disk utilisation looks survivable. The MongoDB log shows no obvious errors, yet operations stall. The likely cause is WiredTiger ticket exhaustion: the storage engine has run out of read or write concurrency tokens, and new work queues behind slow operations. Confirm ticket starvation, find the root cause, and fix it without raising the ticket limit.&lt;/p></description></item><item><title>MongoDB TLS certificate expiry: rotating certs without dropping connections</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-tls-certificate-expiry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-tls-certificate-expiry/</guid><description>&lt;h1 id="mongodb-tls-certificate-expiry-rotating-certs-without-dropping-connections">MongoDB TLS certificate expiry: rotating certs without dropping connections&lt;/h1>
&lt;p>An expired TLS certificate does not degrade gracefully. In MongoDB, client drivers, replica set members, and mongos routers validate certificates during every TLS handshake. When the server certificate&amp;rsquo;s notAfter date passes, OpenSSL rejects the handshake immediately. Existing connections persist until they close, but nothing reconnects once they do. The result is a sudden cascade: secondaries miss heartbeats and transition to DOWN, application connection pools exhaust, and monitoring probes fail. During an active incident, check certificate dates first. For preventive work, your rotation strategy depends on whether you run MongoDB 5.0 or later.&lt;/p></description></item><item><title>MongoDB Too many open files: file descriptor exhaustion and ulimit tuning</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-too-many-open-files/</guid><description>&lt;h1 id="mongodb-too-many-open-files-file-descriptor-exhaustion-and-ulimit-tuning">MongoDB Too many open files: file descriptor exhaustion and ulimit tuning&lt;/h1>
&lt;p>Connection timeouts appear in application logs. MongoDB logs show &lt;code>Too many open files&lt;/code> or &lt;code>error accepting new connection&lt;/code>. New secondaries fail to sync, or a stable node rejects connections after restart. The &lt;code>mongod&lt;/code> process hit the OS file descriptor limit. The failure is often silent until a dependent system breaks.&lt;/p>
&lt;p>FD exhaustion is not always about connection count. &lt;!-- TODO: verify whether WiredTiger holds every collection/index file open indefinitely or only active ones --> WiredTiger maintains open file descriptors for data files, indexes, and journals. A dense deployment with thousands of collections can hold tens of thousands of FDs in steady state. A connection surge, index build, or restarted node rebuilding caches can push the process over the limit. MongoDB then cannot accept new connections, open new data files, or continue replication.&lt;/p></description></item><item><title>MongoDB too stale to catch up: secondary stuck in RECOVERING and how to resync</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-too-stale-to-catch-up/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-too-stale-to-catch-up/</guid><description>&lt;h1 id="mongodb-too-stale-to-catch-up-secondary-stuck-in-recovering-and-how-to-resync">MongoDB too stale to catch up: secondary stuck in RECOVERING and how to resync&lt;/h1>
&lt;p>You check &lt;code>rs.status()&lt;/code> during an incident and see a member stuck in &lt;code>RECOVERING&lt;/code> with an &lt;code>errmsg&lt;/code> reading &lt;code>error RS102 too stale to catch up&lt;/code>. The node is alive but will never transition back to &lt;code>SECONDARY&lt;/code> on its own. Its last replicated oplog entry is older than the oldest entry still available on the primary, so the history it needs has already been overwritten. Incremental replication is impossible from this state. The only path forward is a full initial sync, which on large datasets can take hours to days, adds significant read load to the sync source, and leaves the cluster with reduced redundancy until it completes. If the stale member is a voting node, you are now one failure away from losing majority. This guide covers how to confirm the condition, identify why the secondary fell off, and recover without pushing the remaining cluster members into the same trap.&lt;/p></description></item><item><title>MongoDB unused indexes: $indexStats, write amplification, and safe removal</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-unused-indexes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-unused-indexes/</guid><description>&lt;h1 id="mongodb-unused-indexes-indexstats-write-amplification-and-safe-removal">MongoDB unused indexes: $indexStats, write amplification, and safe removal&lt;/h1>
&lt;p>Every insert and delete updates all indexes on a MongoDB collection; every update that modifies an indexed field updates every index containing that field. An unused index consumes WiredTiger cache space, increases write latency, and amplifies disk I/O. Over time, redundant indexes accumulate from schema migrations, abandoned query patterns, and exploratory tuning.&lt;/p>
&lt;p>The &lt;code>$indexStats&lt;/code> aggregation stage exposes per-index operation counts, but counters reset on mongod restart and internal operations can inflate values. This guide covers detecting truly unused indexes, validating redundancy, and removing them without causing query regressions.&lt;/p></description></item><item><title>MongoDB w:majority write concern timeout (wtimeout): replication lag and at-risk writes</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-write-concern-timeout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-write-concern-timeout/</guid><description>&lt;h1 id="mongodb-wmajority-write-concern-timeout-wtimeout-replication-lag-and-at-risk-writes">MongoDB w:majority write concern timeout (wtimeout): replication lag and at-risk writes&lt;/h1>
&lt;p>Your application logs show write concern timeouts, or worse, the driver returned an error that the application swallowed and moved on. The write succeeded on the primary, but the cluster could not confirm it across a majority of data-bearing members before the deadline expired. If the primary crashes right now, that data is gone. This article explains how to diagnose the root cause, whether it is replication lag, a down secondary, or network saturation, and what to do before the next primary failure.&lt;/p></description></item><item><title>MongoDB WiredTiger cache dirty ratio high: the leading indicator nobody watches</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-cache-dirty-ratio-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-cache-dirty-ratio-high/</guid><description>&lt;h1 id="mongodb-wiredtiger-cache-dirty-ratio-high-the-leading-indicator-nobody-watches">MongoDB WiredTiger cache dirty ratio high: the leading indicator nobody watches&lt;/h1>
&lt;p>Cache fill at 70% looks safe, but if dirty ratio is climbing past 15%, a latency spike is already forming. Dirty ratio measures modified pages not yet flushed to disk. While fill ratio tells you how much cache is in use, dirty ratio tells you how fast the storage engine is falling behind. It often leads checkpoint stalls and eviction-driven latency spikes by minutes.&lt;/p></description></item><item><title>MongoDB WiredTiger cache pressure cascade: eviction stalls and latency spikes</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-cache-pressure-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-cache-pressure-cascade/</guid><description>&lt;h1 id="mongodb-wiredtiger-cache-pressure-cascade-eviction-stalls-and-latency-spikes">MongoDB WiredTiger cache pressure cascade: eviction stalls and latency spikes&lt;/h1>
&lt;p>Latency jumps from milliseconds to seconds for both reads and writes. The slow query log shows no single offender, but connection count climbs as clients retry and timeout. This is the cache pressure cascade. It starts in the storage engine and becomes a self-reinforcing spiral through replication, admission control, and connection handling. This guide covers the mechanism, confirmation under pressure, and how to stop it.&lt;/p></description></item><item><title>MongoDB WriteConflict errors: optimistic concurrency retries under contention</title><link>https://www.netdata.cloud/guides/mongodb/mongodb-writeconflict-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mongodb/mongodb-writeconflict-errors/</guid><description>&lt;h1 id="mongodb-writeconflict-errors-optimistic-concurrency-retries-under-contention">MongoDB WriteConflict errors: optimistic concurrency retries under contention&lt;/h1>
&lt;p>WriteConflict exceptions (error 112) in application logs, or unexplained write latency spikes, point to document-level contention under WiredTiger optimistic concurrency control. Outside of transactions, MongoDB retries single-document writes internally; the client sees slower responses rather than errors. Inside multi-document transactions, MongoDB aborts immediately and returns error 112. Either way, the root cause is concurrent writers targeting the same document.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>WiredTiger uses optimistic concurrency control at the document level. Two concurrent writes to the same document do not block indefinitely; one proceeds and the other encounters a conflict.&lt;/p></description></item><item><title>Monit</title><link>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/monit/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/monit/</guid><description/></item><item><title>Monit Monitoring</title><link>https://www.netdata.cloud/monitoring-101/monit-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/monit-monitoring/</guid><description>&lt;h2 id="monit-monitoring">Monit Monitoring&lt;/h2>
&lt;h3 id="what-is-monit">What Is Monit?&lt;/h3>
&lt;p>Monit is a small Open Source utility for managing and monitoring Unix systems. It conducts automatic maintenance and repair and can engage in stage-based escalation if there are problems.&lt;/p>
&lt;h3 id="monitoring-monit-with-netdata">Monitoring Monit With Netdata&lt;/h3>
&lt;p>Integrating the Monit monitoring tool with Netdata provides real-time, comprehensive monitoring of system services. With Netdata, you can visualize all the data that Monit collects, helping you monitor Monit and optimize your infrastructure more effectively. To see it in action, &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">check out our live demo&lt;/a>.&lt;/p></description></item><item><title>Monitoring 101</title><link>https://www.netdata.cloud/monitoring-101/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/</guid><description/></item><item><title>Monitoring Kubernetes across vanilla, EKS, GKE, AKS, k3s, and rke2</title><link>https://www.netdata.cloud/guides/kubernetes/kubernetes-deployment-variants-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/kubernetes/kubernetes-deployment-variants-monitoring/</guid><description>&lt;h1 id="monitoring-kubernetes-across-vanilla-eks-gke-aks-k3s-and-rke2">Monitoring Kubernetes across vanilla, EKS, GKE, AKS, k3s, and rke2&lt;/h1>
&lt;p>Every Kubernetes cluster exposes the same core control plane components, but what you can actually scrape, query, and alert on depends entirely on how the distribution packages those components. An operator migrating between vanilla kubeadm, Amazon EKS, Google GKE, Azure AKS, k3s, and rke2 needs to know which metrics endpoints exist, which are reachable without extra tooling, and which are gated behind vendor lock-in or architectural quirks. This article maps the native monitoring surface of each variant so you can adapt your collection strategy without discovering gaps during an incident.&lt;/p></description></item><item><title>Monitoring MIG-partitioned NVIDIA GPUs: per-instance metrics</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-mig-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-mig-monitoring/</guid><description>&lt;h1 id="monitoring-mig-partitioned-nvidia-gpus-per-instance-metrics">Monitoring MIG-partitioned NVIDIA GPUs: per-instance metrics&lt;/h1>
&lt;p>On a MIG-enabled GPU (A100, A30, H100 and later), the physical card is partitioned into isolated GPU instances, each with its own SMs, memory partition, and failure domain. Every monitoring habit built around whole-GPU metrics breaks here: aggregate utilization, aggregate memory, and even some health counters describe the card, not the tenant. An instance can be OOM while &lt;code>nvidia-smi&lt;/code> shows healthy free memory on the card.&lt;/p></description></item><item><title>Monitoring overlay tunnels (IPsec/GRE/VXLAN): the signals that matter</title><link>https://www.netdata.cloud/guides/network/network-overlay-tunnel-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-overlay-tunnel-monitoring/</guid><description>&lt;h1 id="monitoring-overlay-tunnels-ipsecgrevxlan-the-signals-that-matter">Monitoring overlay tunnels (IPsec/GRE/VXLAN): the signals that matter&lt;/h1>
&lt;p>Overlay tunnels share the underlay&amp;rsquo;s physical path but add their own failure modes: encapsulation overhead, separate control and data planes, and type-specific state machines. Monitoring only the tunnel interface&amp;rsquo;s link state is the fundamental trap. An interface reporting &amp;ldquo;UP&amp;rdquo; can still be forwarding into a black hole because the peer is unreachable, the SA has expired, or the underlay is dropping fragments.&lt;/p></description></item><item><title>Monitoring route origin and AS-path changes for hijack detection</title><link>https://www.netdata.cloud/guides/network/network-route-origin-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-route-origin-monitoring/</guid><description>&lt;h1 id="monitoring-route-origin-and-as-path-changes-for-hijack-detection">Monitoring route origin and AS-path changes for hijack detection&lt;/h1>
&lt;p>BGP route hijacks do not announce themselves. A prefix legitimately originated by AS2906 starts appearing in the global table with origin AS65001. Traffic that should reach your infrastructure follows a different path, gets blackholed, or lands on an interception point. The control-plane signals are there, but they are distributed across route collectors, RPKI validators, and per-prefix state that most monitoring stacks do not track at the granularity needed.&lt;/p></description></item><item><title>Monnit Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/monnit-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/monnit-corporation-snmp-traps/</guid><description/></item><item><title>Moser Baer AG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/moser-baer-ag-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/moser-baer-ag-snmp-traps/</guid><description/></item><item><title>mosquitto</title><link>https://www.netdata.cloud/integrations/data-collection/databases/mosquitto/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/mosquitto/</guid><description/></item><item><title>Mosquitto Monitoring</title><link>https://www.netdata.cloud/monitoring-101/mosquitto-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/mosquitto-monitoring/</guid><description>&lt;h2 id="mosquitto-monitoring">Mosquitto Monitoring&lt;/h2>
&lt;h3 id="what-is-mosquitto">What Is Mosquitto?&lt;/h3>
&lt;p>Mosquitto is a lightweight MQTT broker, designed to facilitate IoT communications efficiently. As a central hub for sending and receiving messages, Mosquitto ensures that devices can publish and subscribe to channels seamlessly, making it a critical component for IoT infrastructure and applications.&lt;/p>
&lt;h3 id="monitoring-mosquitto-with-netdata">Monitoring Mosquitto With Netdata&lt;/h3>
&lt;p>Monitoring Mosquitto can enhance your IoT message transport and system performance. By using Netdata, you can effectively monitor Mosquitto with ease. Netdata utilizes an openmetrics (Prometheus) exporter to gather and ingest data from the &lt;a href="https://github.com/sapcc/mosquitto-exporter">Mosquitto exporter&lt;/a>. This allows users to enjoy automated dashboards and alerts, without the need for a standalone Prometheus server or Grafana setup.&lt;/p></description></item><item><title>Motorola SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/motorola-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/motorola-snmp-traps/</guid><description/></item><item><title>Moxa Technologies Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/moxa-technologies-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/moxa-technologies-co-ltd-snmp-traps/</guid><description/></item><item><title>Mpb Communications Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mpb-communications-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mpb-communications-inc-snmp-traps/</guid><description/></item><item><title>MQTT Blackbox</title><link>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/mqtt-blackbox/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/mqtt-blackbox/</guid><description/></item><item><title>MQTT Blackbox Monitoring</title><link>https://www.netdata.cloud/monitoring-101/mqtt_blackbox-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/mqtt_blackbox-monitoring/</guid><description>&lt;h2 id="mqtt-blackbox-monitoring">MQTT Blackbox Monitoring&lt;/h2>
&lt;h3 id="what-is-mqtt-blackbox">What Is MQTT Blackbox?&lt;/h3>
&lt;p>MQTT Blackbox is a specialized monitoring technique designed to test and track the performance of MQTT message transport using blackbox testing methods. It leverages the MQTT Blackbox Exporter to simulate client interactions and analyze the reliability and efficiency of message brokers in real-time.&lt;/p>
&lt;h3 id="monitoring-mqtt-blackbox-with-netdata">Monitoring MQTT Blackbox With Netdata&lt;/h3>
&lt;p>Netdata excels at comprehensive MQTT Blackbox monitoring by utilizing an openmetrics (prometheus) exporter. With Netdata, you can effortlessly ingest data from any Prometheus exporter, eliminating the need for a Prometheus server or Grafana. This integration provides users with automated dashboards, real-time alerts, and more, making it an incredibly efficient tool for monitoring MQTT Blackbox.&lt;/p></description></item><item><title>Mrv Communications In Reach Product Division SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mrv-communications-in-reach-product-division-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mrv-communications-in-reach-product-division-snmp-traps/</guid><description/></item><item><title>MS Exchange</title><link>https://www.netdata.cloud/integrations/data-collection/applications/ms-exchange/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/ms-exchange/</guid><description/></item><item><title>mtail</title><link>https://www.netdata.cloud/integrations/data-collection/applications/mtail/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/mtail/</guid><description/></item><item><title>Mtail Monitoring</title><link>https://www.netdata.cloud/monitoring-101/mtail-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/mtail-monitoring/</guid><description>&lt;h2 id="mtail-monitoring">Mtail Monitoring&lt;/h2>
&lt;h3 id="what-is-mtail">What Is Mtail?&lt;/h3>
&lt;p>Mtail is a vital tool used for extracting and parsing log data. Developed by Google, Mtail assists in monitoring and showcasing real-time logs. This is particularly beneficial for DevOps, SREs, and IT professionals who need to ensure the smooth operation of systems through log data metrics. Mtail acts as a log data extractor tailored to plug gaps in monitoring solutions, perfect for complex environments.&lt;/p>
&lt;h3 id="monitoring-mtail-with-netdata">Monitoring Mtail With Netdata&lt;/h3>
&lt;p>Netdata offers a seamlessly integrated environment to monitor mtail using the OpenMetrics (Prometheus) exporter. Netdata&amp;rsquo;s advanced platform can ingest data from any Prometheus exporter, providing automated dashboards, alerts, and visualizations without the need for a Prometheus server or Grafana. This attribute makes Netdata an exceptional mtail monitoring tool.&lt;/p></description></item><item><title>Mts Allstream Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mts-allstream-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mts-allstream-inc-snmp-traps/</guid><description/></item><item><title>Multiple drives failing at once: it's the power supply, not the drives</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-power-supply-multi-drive/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-power-supply-multi-drive/</guid><description>&lt;h1 id="multiple-drives-failing-at-once-its-the-power-supply-not-the-drives">Multiple drives failing at once: it&amp;rsquo;s the power supply, not the drives&lt;/h1>
&lt;p>Three, five, or twelve drives light up at once. Spin_Retry_Count has jumped on every HDD in a shelf. Unsafe Shutdown counts are climbing across every NVMe device. A few drives may have dropped off the bus and come back. The instinct is to start filing RMAs.&lt;/p>
&lt;p>Stop. When the same symptom appears on many drives at the same time, the probability that every drive independently decided to fail in the same hour approaches zero. The root cause is almost certainly external: a degrading PSU, an overloaded PDU, a failing UPS battery, or aggregate spin-up inrush sagging the 12V rail because staggered spin-up is disabled.&lt;/p></description></item><item><title>Mylex Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mylex-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mylex-corporation-snmp-traps/</guid><description/></item><item><title>MySQL</title><link>https://www.netdata.cloud/integrations/data-collection/databases/mysql/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/mysql/</guid><description/></item><item><title>MySQL Aborted_connects and Aborted_clients climbing: diagnosis</title><link>https://www.netdata.cloud/guides/mysql/mysql-aborted-connections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-aborted-connections/</guid><description>&lt;h1 id="mysql-aborted_connects-and-aborted_clients-climbing-diagnosis">MySQL Aborted_connects and Aborted_clients climbing: diagnosis&lt;/h1>
&lt;p>You notice &lt;code>Aborted_connects&lt;/code> or &lt;code>Aborted_clients&lt;/code> climbing on a production MySQL instance. Because these are cumulative counters, a steady upward slope means something is actively failing or dropping connections. In a busy system, a rising &lt;code>Aborted_connects&lt;/code> rate can trigger host blocking via &lt;code>max_connect_errors&lt;/code>, suddenly preventing legitimate clients from connecting. A rising &lt;code>Aborted_clients&lt;/code> rate usually shows up as application-side exceptions about closed connections, forcing retries that can cascade into connection exhaustion. The two counters track completely different failure modes: &lt;code>Aborted_connects&lt;/code> counts connection attempts that never finished authentication; &lt;code>Aborted_clients&lt;/code> counts connections that authenticated but then died unexpectedly. Treating them as the same metric leads to wrong fixes. Restarting the network stack will not fix a credential rotation bug, and increasing &lt;code>max_connections&lt;/code> will not fix a client killed by &lt;code>wait_timeout&lt;/code>.&lt;/p></description></item><item><title>MySQL adaptive hash index latch contention: high CPU, low throughput</title><link>https://www.netdata.cloud/guides/mysql/mysql-adaptive-hash-index-latch-contention/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-adaptive-hash-index-latch-contention/</guid><description>&lt;h1 id="mysql-adaptive-hash-index-latch-contention-high-cpu-low-throughput">MySQL adaptive hash index latch contention: high CPU, low throughput&lt;/h1>
&lt;p>Burning CPU while barely answering queries: OS user time near 100%, application latency spiking, and &lt;code>Questions&lt;/code> flat or falling. &lt;code>Threads_connected&lt;/code> is high but &lt;code>Threads_running&lt;/code> stays low, and disk I/O is quiet. This pattern matches an Adaptive Hash Index (AHI) latch storm.&lt;/p>
&lt;p>InnoDB builds an in-memory hash index over frequently accessed B-tree pages to speed lookups. Since MySQL 5.7, the latch is partitioned into &lt;code>innodb_adaptive_hash_index_parts&lt;/code> (default 8). Under high concurrency, especially point lookups on a hot index, threads can collide on the same partition latch and spin instead of executing. CPU saturates while throughput collapses.&lt;/p></description></item><item><title>MySQL authentication failure spike: brute force vs broken credential rotation</title><link>https://www.netdata.cloud/guides/mysql/mysql-auth-failure-rate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-auth-failure-rate/</guid><description>&lt;h1 id="mysql-authentication-failure-spike-brute-force-vs-broken-credential-rotation">MySQL authentication failure spike: brute force vs broken credential rotation&lt;/h1>
&lt;p>You get paged because &lt;code>Aborted_connects&lt;/code> is climbing. The error log shows a wall of authentication failures, application dashboards are yellow, and someone in security is asking if this is an attack. Before you block IP addresses or rotate passwords, you need to know which failure mode you are dealing with. External brute force and internal broken credential rotation produce the same MySQL status counters, but the fixes are opposite. One requires closing the network window and flushing blocked hosts. The other requires finding the application instance that missed the new secret.&lt;/p></description></item><item><title>MySQL binary logs filling the disk: expiry, lagging replicas, and purge</title><link>https://www.netdata.cloud/guides/mysql/mysql-binary-log-disk-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-binary-log-disk-full/</guid><description>&lt;h1 id="mysql-binary-logs-filling-the-disk-expiry-lagging-replicas-and-purge">MySQL binary logs filling the disk: expiry, lagging replicas, and purge&lt;/h1>
&lt;p>You get a disk-full alert on the MySQL primary. &lt;code>df -h&lt;/code> shows the data partition at 95%, and &lt;code>du&lt;/code> points to &lt;code>/var/lib/mysql/binlog.*&lt;/code> consuming hundreds of gigabytes. Writes are about to fail.&lt;/p>
&lt;p>Binary logs are append-only. MySQL rotates to a new file at &lt;code>max_binlog_size&lt;/code> (default 1 GB), but rotation does not delete old files. Deletion only happens via automatic expiry or &lt;code>PURGE BINARY LOGS&lt;/code>. MySQL refuses to delete any binlog file that a connected replica has not yet consumed. In MySQL 5.7, the default expiry is never (&lt;code>expire_logs_days = 0&lt;/code>). In 8.0, the default is 30 days, but automatic purge cannot remove files that a connected replica still needs. A replica lagging past the expiry window therefore causes unbounded growth.&lt;/p></description></item><item><title>MySQL connection exhaustion: detection, diagnosis, and prevention</title><link>https://www.netdata.cloud/guides/mysql/mysql-connection-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-connection-exhaustion/</guid><description>&lt;h1 id="mysql-connection-exhaustion-detection-diagnosis-and-prevention">MySQL connection exhaustion: detection, diagnosis, and prevention&lt;/h1>
&lt;p>ERROR 1040 (HY000): Too many connections. Health checks fail. Users cannot sign in. If you configured an admin account correctly, a break-glass session may still get through on the reserved slot, and &lt;code>SELECT 1&lt;/code> returns. MySQL is running, but the connection pool is a wall.&lt;/p>
&lt;p>This is a hard cliff. One moment queries flow; the next, every new TCP handshake to port 3306 is rejected. The server does not queue connections. Whether the root cause is a connection leak, a retry storm, or genuine overload from slow queries, the symptom is the same: &lt;code>Threads_connected&lt;/code> has reached &lt;code>max_connections&lt;/code>.&lt;/p></description></item><item><title>MySQL Created_tmp_disk_tables: temp tables spilling to disk</title><link>https://www.netdata.cloud/guides/mysql/mysql-temp-tables-on-disk/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-temp-tables-on-disk/</guid><description>&lt;h1 id="mysql-created_tmp_disk_tables-temp-tables-spilling-to-disk">MySQL Created_tmp_disk_tables: temp tables spilling to disk&lt;/h1>
&lt;p>Queries that &lt;code>GROUP BY&lt;/code>, &lt;code>ORDER BY&lt;/code>, &lt;code>DISTINCT&lt;/code>, or materialize derived results build internal temporary tables. When those tables exceed memory limits, MySQL spills them to disk. The status counter &lt;code>Created_tmp_disk_tables&lt;/code> tracks this, and its ratio to &lt;code>Created_tmp_tables&lt;/code> is the signal to watch during a latency incident.&lt;/p>
&lt;p>A low ratio is normal. A high ratio means queries pay a random-disk penalty for intermediate results. You see that as rising query latency, increased disk I/O on the data directory, and connection pile-up. This guide covers how to read the ratio, find the offending patterns, and fix them without just adding RAM.&lt;/p></description></item><item><title>MySQL ERROR 1040 (HY000): Too many connections - causes and fixes</title><link>https://www.netdata.cloud/guides/mysql/mysql-too-many-connections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-too-many-connections/</guid><description>&lt;h1 id="mysql-error-1040-hy000-too-many-connections---causes-and-fixes">MySQL ERROR 1040 (HY000): Too many connections - causes and fixes&lt;/h1>
&lt;p>Once &lt;code>Threads_connected&lt;/code> reaches &lt;code>max_connections&lt;/code>, MySQL returns &lt;code>ERROR 1040 (HY000)&lt;/code> to every new connection attempt before authentication. The server is not down; it is full.&lt;/p>
&lt;p>This error is usually a symptom. Connections may be leaking from the application, a cache stampede may have flooded the pool, or slow queries may be holding slots open longer than expected. MySQL reserves one extra connection for users with &lt;code>CONNECTION_ADMIN&lt;/code> (or the deprecated &lt;code>SUPER&lt;/code> privilege). If that slot is free, an operator can still connect. If an app user has taken it, you may need to restart.&lt;/p></description></item><item><title>MySQL ERROR 1045 (28000): Access denied for user - diagnosis</title><link>https://www.netdata.cloud/guides/mysql/mysql-access-denied-for-user/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-access-denied-for-user/</guid><description>&lt;h1 id="mysql-error-1045-28000-access-denied-for-user---diagnosis">MySQL ERROR 1045 (28000): Access denied for user - diagnosis&lt;/h1>
&lt;p>When you see &lt;code>ERROR 1045 (28000): Access denied for user 'app'@'10.0.0.5' (using password: YES)&lt;/code>, the connection never reached the query parser. MySQL rejected it during authentication, incremented &lt;code>Aborted_connects&lt;/code>, and returned SQLSTATE &lt;code>28000&lt;/code>. The error always includes the effective user and host as seen by the server. This is your first and most important clue.&lt;/p>
&lt;p>This error has four root causes in production: wrong credentials, user@host mismatch, authentication plugin incompatibility, and host-level blocking from accumulated connection errors. A brute-force probe and a misconfigured deploy look identical from the server side, so diagnosis hinges on correlating the error pattern with connection metrics and the host cache.&lt;/p></description></item><item><title>MySQL ERROR 1205: Lock wait timeout exceeded; try restarting transaction</title><link>https://www.netdata.cloud/guides/mysql/mysql-lock-wait-timeout-exceeded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-lock-wait-timeout-exceeded/</guid><description>&lt;h1 id="mysql-error-1205-lock-wait-timeout-exceeded-try-restarting-transaction">MySQL ERROR 1205: Lock wait timeout exceeded; try restarting transaction&lt;/h1>
&lt;p>&lt;code>ERROR 1205 (HY000): Lock wait timeout exceeded; try restarting transaction&lt;/code> appears when one transaction holds an InnoDB row lock so long that another transaction exhausts &lt;code>innodb_lock_wait_timeout&lt;/code> (default 50 seconds). Unlike a deadlock, which InnoDB resolves automatically, a timeout means a blocker is still active. The root cause is almost always a long-running transaction or hot row contention. Find it before it cascades into a wider outage.&lt;/p></description></item><item><title>MySQL ERROR 1213: Deadlock found when trying to get lock; try restarting transaction</title><link>https://www.netdata.cloud/guides/mysql/mysql-deadlock-found/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-deadlock-found/</guid><description>&lt;h1 id="mysql-error-1213-deadlock-found-when-trying-to-get-lock-try-restarting-transaction">MySQL ERROR 1213: Deadlock found when trying to get lock; try restarting transaction&lt;/h1>
&lt;p>Your application logs show &lt;code>ERROR 1213 (40001): Deadlock found when trying to get lock; try restarting transaction&lt;/code>. One transaction was rolled back; the other completed normally. InnoDB broke a circular lock wait by selecting the cheaper transaction as the victim. A few deadlocks per hour in high-concurrency OLTP is normal. A sustained storm is not: it means transactions are colliding under load.&lt;/p></description></item><item><title>MySQL ERROR 1290: --read-only option so it cannot execute this statement</title><link>https://www.netdata.cloud/guides/mysql/mysql-replica-read-only-write-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-replica-read-only-write-error/</guid><description>&lt;h1 id="mysql-error-1290---read-only-option-so-it-cannot-execute-this-statement">MySQL ERROR 1290: &amp;ndash;read-only option so it cannot execute this statement&lt;/h1>
&lt;p>A write that fails with &lt;code>ERROR 1290 (HY000): The MySQL server is running with the --read-only option so it cannot execute this statement&lt;/code> is a routing or topology problem, not a query problem. The server refuses the statement because &lt;code>read_only&lt;/code> or &lt;code>super_read_only&lt;/code> is enabled. You are likely sending writes to a replica, a recently promoted primary that never cleared its read-only flag, or a node correctly enforcing &lt;code>super_read_only&lt;/code>. Confirm the instance role and fix the write path.&lt;/p></description></item><item><title>MySQL ERROR 2006/2013: MySQL server has gone away -- causes and fixes</title><link>https://www.netdata.cloud/guides/mysql/mysql-server-has-gone-away/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-server-has-gone-away/</guid><description>&lt;h1 id="mysql-error-20062013-mysql-server-has-gone-away----causes-and-fixes">MySQL ERROR 2006/2013: MySQL server has gone away &amp;ndash; causes and fixes&lt;/h1>
&lt;p>ERROR 2006 (CR_SERVER_GONE_ERROR) and ERROR 2013 (CR_SERVER_LOST) produce the messages &amp;ldquo;MySQL server has gone away&amp;rdquo; and &amp;ldquo;Lost connection during query.&amp;rdquo; They often appear together, but they indicate different failures. ERROR 2006 means the server closed a connection the client believed was valid. ERROR 2013 means the connection dropped while a query was executing or results were transferring.&lt;/p>
&lt;p>Treating them identically leads to incorrect fixes. Raising &lt;code>max_connections&lt;/code> does not stop idle timeouts. Tuning timeouts does not fix an OOM kill. Before changing configuration, determine whether the MySQL process restarted. A stable process points to timeout or packet mismatches. A restarted process points to a crash or memory exhaustion.&lt;/p></description></item><item><title>MySQL ERROR 24 (HY000): Too many open files -- raising open_files_limit</title><link>https://www.netdata.cloud/guides/mysql/mysql-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-too-many-open-files/</guid><description>&lt;h1 id="mysql-error-24-hy000-too-many-open-files----raising-open_files_limit">MySQL ERROR 24 (HY000): Too many open files &amp;ndash; raising open_files_limit&lt;/h1>
&lt;p>&lt;code>ERROR 24 (HY000): Too many open files&lt;/code> means the OS refused a file descriptor request because the &lt;code>mysqld&lt;/code> process hit its hard ceiling. Clients may see &lt;code>Can't create/write to file&lt;/code> with &lt;code>errno: 24&lt;/code>. Tables that were accessible moments ago fail to open, and new connections may be rejected even though &lt;code>Threads_connected&lt;/code> is well below &lt;code>max_connections&lt;/code>.&lt;/p>
&lt;p>MySQL holds file descriptors for cached tables, client connections, binary logs, relay logs, and on-disk temporary tables. When combined demand exceeds the OS-enforced limit, the server returns errno 24. Unlike connection exhaustion, which is gated by &lt;code>max_connections&lt;/code>, file descriptor exhaustion is gated by &lt;code>open_files_limit&lt;/code> and the OS limits that enforce it. Raising &lt;code>open_files_limit&lt;/code> in &lt;code>my.cnf&lt;/code> is not enough if systemd, PAM, or the shell ulimit blocks the request.&lt;/p></description></item><item><title>MySQL FLUSH TABLES WITH READ LOCK stall: backups that freeze the server</title><link>https://www.netdata.cloud/guides/mysql/mysql-flush-tables-with-read-lock-stall/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-flush-tables-with-read-lock-stall/</guid><description>&lt;h1 id="mysql-flush-tables-with-read-lock-stall-backups-that-freeze-the-server">MySQL FLUSH TABLES WITH READ LOCK stall: backups that freeze the server&lt;/h1>
&lt;p>Your application suddenly cannot write. &lt;code>Threads_connected&lt;/code> climbs toward &lt;code>max_connections&lt;/code>, but &lt;code>Questions&lt;/code> flatlines. The last change was a backup job that started ten minutes ago. The culprit is almost always &lt;code>FLUSH TABLES WITH READ LOCK&lt;/code> (FTWRL), and the damage is caused not by the lock itself but by what happens while the server waits to acquire it.&lt;/p>
&lt;p>FTWRL is triggered by &lt;code>mysqldump --master-data&lt;/code>, &lt;code>mysqldump --lock-tables&lt;/code>, and similar tools that need a consistent logical backup. It attempts to close all open tables and acquire a global read lock. While it waits for a long-running query to finish, new writes are already blocked. The result is a whole-server freeze: the query rate drops to zero, connections pile up, and the application sees cascading timeouts.&lt;/p></description></item><item><title>MySQL full table scans: Handler_read_rnd_next and the missing index</title><link>https://www.netdata.cloud/guides/mysql/mysql-full-table-scans/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-full-table-scans/</guid><description>&lt;h1 id="mysql-full-table-scans-handler_read_rnd_next-and-the-missing-index">MySQL full table scans: Handler_read_rnd_next and the missing index&lt;/h1>
&lt;p>You get paged because query latency is spiking. &lt;code>Slow_queries&lt;/code> is climbing. &lt;code>Handler_read_rnd_next&lt;/code> has jumped five-fold over its baseline and keeps rising. Connections, CPU, and buffer pool hit ratio look fine. The culprit is usually a new full table scan after a deploy, schema change, or query pattern shift.&lt;/p>
&lt;p>&lt;code>Handler_read_rnd_next&lt;/code> increments when the storage engine reads the next row during a table scan or sorted retrieval. In a healthy OLTP system, &lt;code>Handler_read_key&lt;/code> dominates and &lt;code>Handler_read_rnd_next&lt;/code> stays flat relative to query volume. When the ratio inverts, you are looking at a query plan regression, a missing index, or an optimizer decision gone wrong.&lt;/p></description></item><item><title>MySQL gap locks and next-key locks: surprising deadlocks under REPEATABLE READ</title><link>https://www.netdata.cloud/guides/mysql/mysql-gap-locks-next-key-locks/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-gap-locks-next-key-locks/</guid><description>&lt;h1 id="mysql-gap-locks-and-next-key-locks-surprising-deadlocks-under-repeatable-read">MySQL gap locks and next-key locks: surprising deadlocks under REPEATABLE READ&lt;/h1>
&lt;p>MySQL deadlocks in production frequently involve transactions that modify apparently unrelated rows. Two concurrent UPDATEs with narrow WHERE clauses collide, or an UPDATE blocks every INSERT into the table. The cause is usually REPEATABLE READ combined with InnoDB next-key locking and a missing or non-unique index.&lt;/p>
&lt;p>Under REPEATABLE READ, the default in MySQL, InnoDB does not lock only matching rows. To prevent phantom reads, it locks the gaps between index entries. When no suitable index exists, or when the optimizer scans a non-unique index, a targeted UPDATE escalates into a range lock covering far more of the table than the predicate suggests.&lt;/p></description></item><item><title>MySQL Got error 28 from storage engine / No space left on device — recovery</title><link>https://www.netdata.cloud/guides/mysql/mysql-disk-full-no-space-left/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-disk-full-no-space-left/</guid><description>&lt;h1 id="mysql-got-error-28-from-storage-engine--no-space-left-on-device--recovery">MySQL Got error 28 from storage engine / No space left on device — recovery&lt;/h1>
&lt;p>When &lt;code>Got error 28 from storage engine&lt;/code> appears, the underlying filesystem is at 100%. MySQL needs writable space for InnoDB data files, redo logs, binary logs, temporary tables, relay logs, and slow query logs. Writes fail immediately. Affected threads enter a retry loop, logging warnings until an operator frees space or kills the operation.&lt;/p>
&lt;p>The most dangerous scenario is the redo log filling up. If checkpoint age approaches capacity, InnoDB forces aggressive synchronous flushing. Under sustained pressure this can trigger an unclean shutdown and crash recovery on restart. Replicas are similarly vulnerable: a full relay log partition stops the I/O thread and lets lag grow unboundedly, risking binlog expiry on the source before catch-up.&lt;/p></description></item><item><title>MySQL GTID errant transactions: detecting replication divergence</title><link>https://www.netdata.cloud/guides/mysql/mysql-gtid-errant-transactions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-gtid-errant-transactions/</guid><description>&lt;h1 id="mysql-gtid-errant-transactions-detecting-replication-divergence">MySQL GTID errant transactions: detecting replication divergence&lt;/h1>
&lt;p>Before a planned failover, or after an orchestrator aborts with a GTID consistency error, replication can appear healthy: &lt;code>SHOW REPLICA STATUS&lt;/code> reports both threads running, &lt;code>Seconds_Behind_Source&lt;/code> is near zero, and the error log is quiet. Yet comparing GTID sets between source and replica reveals mismatching numbers. Errant transactions are GTIDs present in a replica&amp;rsquo;s &lt;code>gtid_executed&lt;/code> set that the source never generated.&lt;/p>
&lt;p>Ordinary lag resolves as the replica catches up. Errant transactions represent true divergence. Promoting a replica with extra GTIDs propagates those transactions into the new source and downstream replicas, creating split-brain that is expensive to reverse. This guide covers distinguishing benign lag from dangerous divergence, pinpointing offending GTIDs, and deciding between empty-transaction injection and a full rebuild.&lt;/p></description></item><item><title>MySQL InnoDB buffer pool hit ratio collapse: the cliff edge</title><link>https://www.netdata.cloud/guides/mysql/mysql-buffer-pool-hit-ratio-collapse/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-buffer-pool-hit-ratio-collapse/</guid><description>&lt;h1 id="mysql-innodb-buffer-pool-hit-ratio-collapse-the-cliff-edge">MySQL InnoDB buffer pool hit ratio collapse: the cliff edge&lt;/h1>
&lt;p>Your OLTP queries were running in single-digit milliseconds. Now every query is taking seconds, the disk subsystem is saturated, &lt;code>Threads_running&lt;/code> is climbing toward &lt;code>max_connections&lt;/code>, and the buffer pool hit ratio, which sat at 99.9% for months, just fell through 95% and keeps dropping.&lt;/p>
&lt;p>This is the InnoDB buffer pool cliff edge. When the working set exceeds the buffer pool, pages are evicted before they can be reused. Every miss becomes a physical disk read. The degradation is non-linear: 99.9% to 99% is a slow bleed, 99% to 95% is rapid, and below 95% disk saturation, uniform latency inflation, and connection exhaustion turn a capacity problem into an availability incident.&lt;/p></description></item><item><title>MySQL InnoDB checkpoint age: the redo log capacity signal nobody watches</title><link>https://www.netdata.cloud/guides/mysql/mysql-checkpoint-age-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-checkpoint-age-monitoring/</guid><description>&lt;h1 id="mysql-innodb-checkpoint-age-the-redo-log-capacity-signal-nobody-watches">MySQL InnoDB checkpoint age: the redo log capacity signal nobody watches&lt;/h1>
&lt;p>Checkpoint age is the distance between the current InnoDB log sequence number (LSN) and the last checkpoint LSN. It measures how much of the circular redo log is occupied by changes not yet flushed to data files. When this age approaches redo log capacity, InnoDB escalates flushing aggressiveness until it runs out of options and stalls all writes.&lt;/p>
&lt;p>Standard MySQL does not expose &lt;code>Innodb_checkpoint_age&lt;/code> as a status variable. Depending on version, you must parse &lt;code>SHOW ENGINE INNODB STATUS&lt;/code>, enable a disabled-by-default &lt;code>INNODB_METRICS&lt;/code> counter, or compute the delta from MySQL 8.0.30+ redo status variables. MySQL 8.0.30 replaced &lt;code>innodb_log_file_size&lt;/code> and &lt;code>innodb_log_files_in_group&lt;/code> with &lt;code>innodb_redo_log_capacity&lt;/code>. &lt;!-- TODO: verify whether leaving the new variable unset on upgrade always reduces capacity, or if MySQL computes it from legacy settings when present. -->&lt;/p></description></item><item><title>MySQL InnoDB history list length growing: purge lag explained</title><link>https://www.netdata.cloud/guides/mysql/mysql-history-list-length-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-history-list-length-growing/</guid><description>&lt;h1 id="mysql-innodb-history-list-length-growing-purge-lag-explained">MySQL InnoDB history list length growing: purge lag explained&lt;/h1>
&lt;p>If query latency climbs across every table and no single query stands out, check the InnoDB history list length. It tracks unpurged undo records: MVCC debt. When purge falls behind, every consistent read walks longer version chains, and the slowdown is global. Unlike a slow query, this debt is invisible to conventional query analysis.&lt;/p>
&lt;p>In MySQL 8.0, the programmatic source for this metric is &lt;code>trx_rseg_history_len&lt;/code> in &lt;code>information_schema.INNODB_METRICS&lt;/code>. &lt;!-- TODO: verify whether Innodb_history_list_length was removed from SHOW GLOBAL STATUS in MySQL 8.0 or only in specific versions --> Despite being labeled &lt;code>status_counter&lt;/code> in the metrics table, &lt;code>trx_rseg_history_len&lt;/code> behaves as a gauge: it rises when purge lags and falls when purge catches up.&lt;/p></description></item><item><title>MySQL InnoDB redo log checkpoint stall: when all writes freeze</title><link>https://www.netdata.cloud/guides/mysql/mysql-redo-log-checkpoint-stall/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-redo-log-checkpoint-stall/</guid><description>&lt;h1 id="mysql-innodb-redo-log-checkpoint-stall-when-all-writes-freeze">MySQL InnoDB redo log checkpoint stall: when all writes freeze&lt;/h1>
&lt;p>InnoDB redo log checkpoint stalls freeze all writes synchronously when checkpoint age reaches the end of the circular redo log. INSERTs, UPDATEs, and DELETEs that complete in milliseconds hang for seconds. &lt;code>Threads_running&lt;/code> climbs while the &lt;code>Questions&lt;/code> rate collapses. Read queries may still return initially, but soon every connection waits for a write lock or commit acknowledgment. Then throughput recovers just as abruptly. This is not a disk failure, deadlock, or runaway query. It is MySQL&amp;rsquo;s write path running out of runway.&lt;/p></description></item><item><title>MySQL InnoDB row lock contention: finding who blocks whom</title><link>https://www.netdata.cloud/guides/mysql/mysql-row-lock-contention/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-row-lock-contention/</guid><description>&lt;h1 id="mysql-innodb-row-lock-contention-finding-who-blocks-whom">MySQL InnoDB row lock contention: finding who blocks whom&lt;/h1>
&lt;p>Queries that normally finish in milliseconds take seconds. &lt;code>Threads_running&lt;/code> climbs while &lt;code>Questions&lt;/code> stalls. &lt;code>SHOW PROCESSLIST&lt;/code> shows active threads, yet the database is frozen. That is the shape of InnoDB row lock contention: transactions are waiting to release row-level locks, and the queue is growing.&lt;/p>
&lt;p>Row lock contention differs from a metadata lock cascade. Metadata locks block DDL and DML at the table level and live in &lt;code>performance_schema.metadata_locks&lt;/code>. Row locks are held by open InnoDB transactions and block at the row, gap, or next-key level. Rising &lt;code>Innodb_row_lock_current_waits&lt;/code> alongside rising &lt;code>Threads_running&lt;/code> signals a live contention crisis. The goal is to identify the blocker, the waiter, and the lock footprint so you can break the chain without guessing.&lt;/p></description></item><item><title>MySQL innodb_buffer_pool_size tuning: 60-80% of RAM and when that breaks</title><link>https://www.netdata.cloud/guides/mysql/mysql-buffer-pool-sizing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-buffer-pool-sizing/</guid><description>&lt;h1 id="mysql-innodb_buffer_pool_size-tuning-60-80-of-ram-and-when-that-breaks">MySQL innodb_buffer_pool_size tuning: 60-80% of RAM and when that breaks&lt;/h1>
&lt;p>The 60-80% rule works for a bare-metal host running only MySQL. In containers, shared hardware, and high-connection-count environments, it is a hazard. Size the pool against aligned allocation, cgroup limits, per-connection memory, and OS headroom. Do not size it as a percentage of total RAM.&lt;/p>
&lt;h2 id="what-the-buffer-pool-costs">What the buffer pool costs&lt;/h2>
&lt;p>InnoDB caches data and index pages in a fixed-size pool. Every read and write touches it. When the working set fits entirely in memory, queries avoid disk. When it does not, InnoDB evicts pages and disk reads dominate latency.&lt;/p></description></item><item><title>MySQL Innodb_buffer_pool_wait_free > 0: buffer pool memory pressure</title><link>https://www.netdata.cloud/guides/mysql/mysql-buffer-pool-wait-free/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-buffer-pool-wait-free/</guid><description>&lt;h1 id="mysql-innodb_buffer_pool_wait_free--0-buffer-pool-memory-pressure">MySQL Innodb_buffer_pool_wait_free &amp;gt; 0: buffer pool memory pressure&lt;/h1>
&lt;p>&lt;code>Innodb_buffer_pool_wait_free&lt;/code> increments when InnoDB must synchronously flush dirty pages to make room for new reads. A sustained nonzero rate means queries are waiting on disk writes before they can proceed. Unlike the buffer pool hit ratio, which can stay above 99% while the system stalls, &lt;code>wait_free&lt;/code> confirms the buffer pool is operating at its limit.&lt;/p>
&lt;p>This is the Buffer Pool Cliff pattern: once the working set exceeds available clean pages, performance degrades non-linearly, disk I/O saturates, and threads pile up.&lt;/p></description></item><item><title>MySQL innodb_deadlock_detect=OFF: when deadlock detection becomes the bottleneck</title><link>https://www.netdata.cloud/guides/mysql/mysql-deadlock-detect-off-high-concurrency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-deadlock-detect-off-high-concurrency/</guid><description>&lt;h1 id="mysql-innodb_deadlock_detectoff-when-deadlock-detection-becomes-the-bottleneck">MySQL innodb_deadlock_detect=OFF: when deadlock detection becomes the bottleneck&lt;/h1>
&lt;p>At extreme concurrency, InnoDB&amp;rsquo;s deadlock detector can become the bottleneck. It traverses the wait-for graph on every lock enqueue. That cost is usually negligible until thousands of concurrent writing transactions hit a small set of hot rows. When traversal itself saturates CPU, throughput collapses even though disks are idle and the buffer pool is warm.&lt;/p>
&lt;p>&lt;code>innodb_deadlock_detect=OFF&lt;/code> removes that traversal. Instead of instant deadlock detection, InnoDB relies exclusively on &lt;code>innodb_lock_wait_timeout&lt;/code> to break cycles. The blocked transaction waits, holding its locks, until the timeout expires. This is not a generic performance tuning knob. It is a narrow, high-risk tradeoff that converts instant rollback into delayed timeout, and it is appropriate only when you have confirmed through testing that deadlock detection is the bottleneck.&lt;/p></description></item><item><title>MySQL Innodb_log_waits > 0: the log buffer is too small (not a checkpoint stall)</title><link>https://www.netdata.cloud/guides/mysql/mysql-log-buffer-waits/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-log-buffer-waits/</guid><description>&lt;h1 id="mysql-innodb_log_waits--0-the-log-buffer-is-too-small-not-a-checkpoint-stall">MySQL Innodb_log_waits &amp;gt; 0: the log buffer is too small (not a checkpoint stall)&lt;/h1>
&lt;p>InnoDB increments &lt;code>Innodb_log_waits&lt;/code> every time a thread waits because the log buffer is full. A sustained nonzero rate means the buffer cannot absorb your write burst, or the storage layer cannot drain it fast enough.&lt;/p>
&lt;p>This is not a checkpoint stall. Checkpoint stalls happen when the on-disk redo log fills and InnoDB forces synchronous dirty-page flushing. &lt;code>Innodb_log_waits&lt;/code> measures in-memory log buffer pressure only. Confusing the two leads to tuning redo log file size while the real problem persists.&lt;/p></description></item><item><title>MySQL innodb_redo_log_capacity sizing: how big should the redo log be</title><link>https://www.netdata.cloud/guides/mysql/mysql-redo-log-capacity-sizing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-redo-log-capacity-sizing/</guid><description>&lt;h1 id="mysql-innodb_redo_log_capacity-sizing-how-big-should-the-redo-log-be">MySQL innodb_redo_log_capacity sizing: how big should the redo log be&lt;/h1>
&lt;p>An undersized InnoDB redo log is a common cause of MySQL write stalls. When redo generation outpaces dirty-page flushing, checkpoint age advances until InnoDB enters synchronous flushing and user writes block.&lt;/p>
&lt;p>Since MySQL 8.0.30, &lt;code>innodb_redo_log_capacity&lt;/code> replaces &lt;code>innodb_log_file_size&lt;/code> and &lt;code>innodb_log_files_in_group&lt;/code>. Resizing no longer requires a restart, but the sizing logic is unchanged: capacity must absorb peak write rates without pushing checkpoint age into the synchronous-flush zone.&lt;/p></description></item><item><title>MySQL InnoDB: Database page corruption — detection and recovery</title><link>https://www.netdata.cloud/guides/mysql/mysql-innodb-page-corruption/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-innodb-page-corruption/</guid><description>&lt;h1 id="mysql-innodb-database-page-corruption--detection-and-recovery">MySQL InnoDB: Database page corruption — detection and recovery&lt;/h1>
&lt;p>Your error log shows &lt;code>InnoDB: Database page corruption on disk or a failed file read&lt;/code>. MySQL may abort the connection, mark the tablespace read-only, or crash. Page corruption is a storage-layer failure, not a query bug: a 16KB InnoDB page failed checksum validation.&lt;/p>
&lt;p>The critical distinction is between secondary index corruption, which is rebuildable without data loss, and clustered index corruption, which affects the table data itself. Detect the failure, classify the damage, extract what you can, and rebuild safely.&lt;/p></description></item><item><title>MySQL long-running transactions: detecting and killing the silent blocker</title><link>https://www.netdata.cloud/guides/mysql/mysql-long-running-transactions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-long-running-transactions/</guid><description>&lt;h1 id="mysql-long-running-transactions-detecting-and-killing-the-silent-blocker">MySQL long-running transactions: detecting and killing the silent blocker&lt;/h1>
&lt;p>&lt;code>Threads_running&lt;/code> climbs. A DDL operation that should take seconds is stuck for minutes. Storage grows steadily with no matching data increase. The cause is often a single idle transaction in &lt;code>INNODB_TRX&lt;/code>, holding row locks and an MVCC read view long after its last query finished. It does not appear in the slow query log and may have no active query string. Left alone, it blocks InnoDB purge, inflates the history list, and can trigger a metadata lock cascade that fills the connection pool.&lt;/p></description></item><item><title>MySQL metadata lock cascade: how one ALTER TABLE freezes a whole table</title><link>https://www.netdata.cloud/guides/mysql/mysql-metadata-lock-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-metadata-lock-cascade/</guid><description>&lt;h1 id="mysql-metadata-lock-cascade-how-one-alter-table-freezes-a-whole-table">MySQL metadata lock cascade: how one ALTER TABLE freezes a whole table&lt;/h1>
&lt;p>You run an &lt;code>ALTER TABLE&lt;/code> to add an index. Seconds later, health checks fail for queries that touch only that table. Other tables work fine. CPU and disk are idle. The connection pool fills. &lt;code>SHOW ENGINE INNODB STATUS&lt;/code> shows no row lock waits. The culprit is a metadata lock cascade, not InnoDB contention.&lt;/p>
&lt;p>This happens when DDL requests an exclusive metadata lock (MDL) on a table already protected by a shared MDL held by a long-running or idle transaction. The DDL waits. Subsequent DML on that table queues behind it. The queue grows until the connection pool exhausts and the application times out. Because the outage is isolated to one table and leaves no trace in InnoDB lock metrics, it is easy to misdiagnose as a network blip or a runaway query.&lt;/p></description></item><item><title>MySQL Monitoring</title><link>https://www.netdata.cloud/monitoring-101/mysql-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/mysql-monitoring/</guid><description>&lt;h2 id="mysql-monitoring">MySQL Monitoring&lt;/h2>
&lt;h3 id="what-is-mysql">What Is MySQL?&lt;/h3>
&lt;p>&lt;a href="https://www.mysql.com/">MySQL&lt;/a> is a widely-used open-source relational database management system. It is a crucial component in web application architecture and is known for its reliability, scalability, and performance. MySQL supports a variety of database-driven applications by enabling efficient storage, retrieval, and management of data.&lt;/p>
&lt;h3 id="monitoring-mysql-with-netdata">Monitoring MySQL With Netdata&lt;/h3>
&lt;p>Netdata offers real-time, insightful metrics and monitoring capabilities required for effective MySQL monitoring. With Netdata, you gain access to detailed visualizations and alerts to ensure that your MySQL server operates optimally. This is achieved via &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/mysql/">Netdata&amp;rsquo;s MySQL monitoring tool&lt;/a> which seamlessly integrates and enables deep analytics of your MySQL instances.&lt;/p></description></item><item><title>MySQL monitoring checklist: the signals every production instance needs</title><link>https://www.netdata.cloud/guides/mysql/mysql-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-monitoring-checklist/</guid><description>&lt;h1 id="mysql-monitoring-checklist-the-signals-every-production-instance-needs">MySQL monitoring checklist: the signals every production instance needs&lt;/h1>
&lt;p>Most production incidents involving MySQL stem from missing observability into specific internal subsystems. Connection saturation, buffer pool thrashing, and replication lag can degrade a healthy cluster into an outage if the right signals are not monitored.&lt;/p>
&lt;p>This checklist serves as a tiered operational reference for senior engineers and SREs. It defines the exact metrics needed to keep MySQL instances available, fast, and reliable, scaling from basic survival metrics to expert-level internals. Use it to validate existing dashboards or baseline monitoring for a new deployment.&lt;/p></description></item><item><title>MySQL monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/mysql/mysql-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-monitoring-maturity-model/</guid><description>&lt;h1 id="mysql-monitoring-maturity-model-from-survival-to-expert">MySQL monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most MySQL monitoring stacks grow reactively. A team starts with a liveness check and disk alerting. An incident exposes a gap, a new signal gets added, and the cycle repeats. The result is uneven coverage: strong in areas that caused past outages, blind in areas that will cause the next one.&lt;/p>
&lt;p>This article maps MySQL monitoring into four capability levels. Use it as a self-assessment, not a checklist to max out. Every level includes signals the playbook identifies as operationally important, along with the failure modes they catch. If you are missing an entire category at your current level, that is your highest-priority gap.&lt;/p></description></item><item><title>MySQL online DDL still blocking: ALGORITHM, LOCK, and the copy phase</title><link>https://www.netdata.cloud/guides/mysql/mysql-online-ddl-blocking/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-online-ddl-blocking/</guid><description>&lt;h1 id="mysql-online-ddl-still-blocking-algorithm-lock-and-the-copy-phase">MySQL online DDL still blocking: ALGORITHM, LOCK, and the copy phase&lt;/h1>
&lt;p>Online DDL in MySQL is not lock-free. Even with &lt;code>ALGORITHM=INPLACE&lt;/code> or &lt;code>ALGORITHM=INSTANT&lt;/code>, and even with &lt;code>LOCK=NONE&lt;/code>, every online DDL operation passes through a brief window where it upgrades its metadata lock (MDL) to exclusive. That window is short, but it is real, and it is the single most common reason an &amp;ldquo;online&amp;rdquo; schema change still stalls a busy production table.&lt;/p></description></item><item><title>MySQL OOM-killed: buffer pool, per-connection buffers, and the kernel killer</title><link>https://www.netdata.cloud/guides/mysql/mysql-out-of-memory-oom-killed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-out-of-memory-oom-killed/</guid><description>&lt;h1 id="mysql-oom-killed-buffer-pool-per-connection-buffers-and-the-kernel-killer">MySQL OOM-killed: buffer pool, per-connection buffers, and the kernel killer&lt;/h1>
&lt;p>MySQL disappears. Application logs fill with connection timeouts. &lt;code>Uptime&lt;/code> resets to near zero. In the kernel log: &lt;code>oom-kill: task mysqld&lt;/code> and its anonymous RSS. Orchestrators mark the pod &lt;code>OOMKilled&lt;/code> (exit code 137). To the application, this is a crash. To the kernel, mysqld was the largest memory consumer and the system ran out.&lt;/p>
&lt;p>OOM kills are gradual, then catastrophic. Memory pressure builds as connections open, temp tables materialize, and dirty pages accumulate. The buffer pool is the obvious consumer, but the killer often enters through the back door: a connection burst multiplies per-thread buffers, a container limit sits too close to the buffer pool size, or Transparent Huge Pages block reclamation. Map all allocators to prevent recurrence.&lt;/p></description></item><item><title>MySQL Opened_tables climbing: table_open_cache and open-files pressure</title><link>https://www.netdata.cloud/guides/mysql/mysql-table-open-cache-pressure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-table-open-cache-pressure/</guid><description>&lt;h1 id="mysql-opened_tables-climbing-table_open_cache-and-open-files-pressure">MySQL Opened_tables climbing: table_open_cache and open-files pressure&lt;/h1>
&lt;p>Steadily climbing &lt;code>Opened_tables&lt;/code> with &lt;code>Open_tables&lt;/code> pinned near &lt;code>table_open_cache&lt;/code> means the instance is churning table file descriptors instead of reusing them. Each cache miss opens a table, consumes a file descriptor, and adds metadata lock overhead. The first symptom is usually sporadic query latency during peak traffic. Left unchecked, the instance exhausts its file descriptor limit and returns errors like &lt;code>Too many open files&lt;/code> or refuses connections.&lt;/p></description></item><item><title>MySQL privilege change auditing: GRANT, REVOKE, and unexpected SUPER</title><link>https://www.netdata.cloud/guides/mysql/mysql-privilege-changes-auditing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-privilege-changes-auditing/</guid><description>&lt;h1 id="mysql-privilege-change-auditing-grant-revoke-and-unexpected-super">MySQL privilege change auditing: GRANT, REVOKE, and unexpected SUPER&lt;/h1>
&lt;p>Unexpected privilege changes are high-impact, low-frequency events. A spike in &lt;code>Com_grant&lt;/code>, an unknown account with &lt;code>SUPER&lt;/code>, or a &lt;code>GRANT ALL&lt;/code> outside a change window can signal privilege escalation, misconfigured automation, or migration debt. This guide covers how to detect, diagnose, and remove unauthorized grants and deprecated &lt;code>SUPER&lt;/code> privileges in production.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>MySQL increments status counters for &lt;code>GRANT&lt;/code>, &lt;code>REVOKE&lt;/code>, &lt;code>CREATE USER&lt;/code>, &lt;code>ALTER USER&lt;/code>, and &lt;code>DROP USER&lt;/code>. The counters are cumulative; compute deltas over fixed intervals instead of reading absolute values.&lt;/p></description></item><item><title>MySQL purge lag from an idle transaction: the slow bleed</title><link>https://www.netdata.cloud/guides/mysql/mysql-purge-lag-idle-transaction/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-purge-lag-idle-transaction/</guid><description>&lt;h1 id="mysql-purge-lag-from-an-idle-transaction-the-slow-bleed">MySQL purge lag from an idle transaction: the slow bleed&lt;/h1>
&lt;p>Your OLTP workload slows over hours. No single query dominates the slow log. CPU and disk I/O are not saturated. Yet read latency climbs, throughput drifts downward, and the undo tablespace grows. The culprit is often a single idle transaction that opened a read view and never closed it. Under InnoDB&amp;rsquo;s default REPEATABLE READ isolation, that transaction pins the purge boundary. Every subsequent write appends to the undo log. The history list length grows without bound, and all MVCC reads traverse longer version chains.&lt;/p></description></item><item><title>MySQL query plan regression: when a query gets slow after a deploy or ANALYZE</title><link>https://www.netdata.cloud/guides/mysql/mysql-query-plan-regression/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-query-plan-regression/</guid><description>&lt;h1 id="mysql-query-plan-regression-when-a-query-gets-slow-after-a-deploy-or-analyze">MySQL query plan regression: when a query gets slow after a deploy or ANALYZE&lt;/h1>
&lt;p>A query that completed in milliseconds yesterday now times out. The deploy was clean, or maybe someone ran ANALYZE TABLE during maintenance. Slow_queries climbs, Handler_read_rnd_next jumps, and the same application code suddenly kills the database. The optimizer has changed how it executes one or more queries, and the new plan is far more expensive.&lt;/p>
&lt;p>Plan regressions are sharp step-changes, not gradual growth. They correlate tightly with specific events: a schema migration, an index drop, a bulk load that triggered auto-recalc, a manual ANALYZE TABLE, or a MySQL version upgrade. The cause is usually stale or shifted statistics, an optimizer behavior change, or a missing index the optimizer can no longer use.&lt;/p></description></item><item><title>MySQL relay log filling the replica disk: Relay_Log_Space and recovery</title><link>https://www.netdata.cloud/guides/mysql/mysql-relay-log-disk-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-relay-log-disk-full/</guid><description>&lt;h1 id="mysql-relay-log-filling-the-replica-disk-relay_log_space-and-recovery">MySQL relay log filling the replica disk: Relay_Log_Space and recovery&lt;/h1>
&lt;p>When &lt;code>SHOW REPLICA STATUS&lt;/code> shows &lt;code>Seconds_Behind_Source&lt;/code> climbing and &lt;code>Relay_Log_Space&lt;/code> growing, the SQL thread is not applying events as fast as the IO thread fetches them. Normally MySQL purges each relay log after the SQL thread finishes it, so total disk usage stays bounded. When apply cannot keep up, files accumulate. If the partition fills, the IO thread stops. No new events arrive, lag becomes unbounded, and if the source purges binary logs before recovery, the replica requires a full resync.&lt;/p></description></item><item><title>MySQL Replica_IO_Running / Replica_SQL_Running not Yes: replication stopped</title><link>https://www.netdata.cloud/guides/mysql/mysql-replica-io-sql-thread-stopped/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-replica-io-sql-thread-stopped/</guid><description>&lt;h1 id="mysql-replica_io_running--replica_sql_running-not-yes-replication-stopped">MySQL Replica_IO_Running / Replica_SQL_Running not Yes: replication stopped&lt;/h1>
&lt;p>&lt;code>SHOW REPLICA STATUS\G&lt;/code> (MySQL 8.0+) or &lt;code>SHOW SLAVE STATUS\G&lt;/code> (MySQL 5.7) shows &lt;code>Replica_IO_Running&lt;/code> or &lt;code>Replica_SQL_Running&lt;/code> is not &lt;code>Yes&lt;/code>. The IO thread may hang in &lt;code>Connecting&lt;/code>, or the SQL thread may stop with an error. Replication no longer makes progress. Lag grows without bound, and if the source purges binary logs before recovery, you face a full rebuild.&lt;/p>
&lt;p>Unlike many MySQL problems, a stopped replication thread does not self-heal. The SQL thread halts on the first apply error and waits for operator intervention. The IO thread retries automatically, but only up to the configured retry limit. &lt;code>Seconds_Behind_Source&lt;/code> can show &lt;code>0&lt;/code> even when the IO thread is disconnected and the SQL thread has consumed the relay log, so lag alone is not a reliable health signal.&lt;/p></description></item><item><title>MySQL replication lag spiral: Seconds_Behind_Source growing without bound</title><link>https://www.netdata.cloud/guides/mysql/mysql-replication-lag-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-replication-lag-spiral/</guid><description>&lt;h1 id="mysql-replication-lag-spiral-seconds_behind_source-growing-without-bound">MySQL replication lag spiral: Seconds_Behind_Source growing without bound&lt;/h1>
&lt;p>You check &lt;code>SHOW REPLICA STATUS&lt;/code> and see &lt;code>Seconds_Behind_Source&lt;/code> is not just elevated; it is climbing. Five minutes ago it was 60. Now it is 120. The replica is not catching up; it is falling further behind.&lt;/p>
&lt;p>This is a replication lag death spiral. The SQL apply thread (or threads) cannot replay events as fast as the source generates them. Relay logs accumulate. Lag compounds because longer apply queues mean larger effective transactions, longer lock holds, and greater exposure to blocking queries on the replica.&lt;/p></description></item><item><title>MySQL Seconds_Behind_Source is unreliable: measuring real replication lag</title><link>https://www.netdata.cloud/guides/mysql/mysql-seconds-behind-source-unreliable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-seconds-behind-source-unreliable/</guid><description>&lt;h1 id="mysql-seconds_behind_source-is-unreliable-measuring-real-replication-lag">MySQL Seconds_Behind_Source is unreliable: measuring real replication lag&lt;/h1>
&lt;p>You run &lt;code>SHOW REPLICA STATUS&lt;/code> and see &lt;code>Seconds_Behind_Source = 0&lt;/code>. Both replication threads say &lt;code>Yes&lt;/code>. You assume the replica is current. During a source write burst, you fail over and discover hours of missing transactions. This happens because &lt;code>Seconds_Behind_Source&lt;/code> is not a real-time lag measurement. It is a timestamp diff between the replica&amp;rsquo;s clock and the timestamp of the event the SQL thread is currently applying. When the SQL thread reaches the end of the relay log, the metric snaps to zero even if the I/O thread has not fetched newer events. When the SQL thread stops, it returns NULL, which most lag alerts ignore. When the source was idle and then bursts, it jumps to the idle duration even though real lag is minimal. With parallel replication, it oscillates based on worker timing. For failover-critical decisions, you need a different measurement strategy.&lt;/p></description></item><item><title>MySQL Select_full_join > 0: joins running without an index</title><link>https://www.netdata.cloud/guides/mysql/mysql-select-full-join/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-select-full-join/</guid><description>&lt;h1 id="mysql-select_full_join--0-joins-running-without-an-index">MySQL Select_full_join &amp;gt; 0: joins running without an index&lt;/h1>
&lt;p>Select_full_join is climbing. In an OLTP system, this counter should stay at zero. Any sustained nonzero rate means at least one query is executing a join without a usable index, effectively performing a full cross-product scan on the joined table. The damage scales with table size. A single unindexed join against a million-row table can stall an otherwise healthy instance.&lt;/p>
&lt;p>This is not a buffer pool problem or a connection storm. It is a query correctness issue. The optimizer cannot find an index to resolve the join predicate, so it reads every row of the referenced table for each row of the driving table. In MySQL 8.0.20 and later, the execution engine uses hash join for these cases (block nested loop was removed), but the underlying pathology remains identical: missing index, catastrophic read amplification.&lt;/p></description></item><item><title>MySQL semi-synchronous replication stall: commits hanging on ACK</title><link>https://www.netdata.cloud/guides/mysql/mysql-semi-sync-replication-stall/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-semi-sync-replication-stall/</guid><description>&lt;h1 id="mysql-semi-synchronous-replication-stall-commits-hanging-on-ack">MySQL semi-synchronous replication stall: commits hanging on ACK&lt;/h1>
&lt;p>Commits suddenly take 10 seconds or more and then time out. On the MySQL primary, &lt;code>Threads_running&lt;/code> climbs while the commit rate flatlines. A moment later, commits resume, but &lt;code>Rpl_semi_sync_source_no_tx&lt;/code> is ticking upward. The primary is waiting for a semi-synchronous replica to acknowledge receipt of binlog events, and that ACK is not arriving fast enough or at all.&lt;/p>
&lt;p>Confirm the stall, find the missing ACK, and decide whether to fix the replica path or temporarily fall back to asynchronous replication.&lt;/p></description></item><item><title>MySQL slow after restart: buffer pool warm-up and the cold cache</title><link>https://www.netdata.cloud/guides/mysql/mysql-buffer-pool-not-warming-up/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-buffer-pool-not-warming-up/</guid><description>&lt;h1 id="mysql-slow-after-restart-buffer-pool-warm-up-and-the-cold-cache">MySQL slow after restart: buffer pool warm-up and the cold cache&lt;/h1>
&lt;p>You restart MySQL for maintenance, crash recovery, or a failover. Seconds later, query latency jumps from sub-millisecond to tens or hundreds of milliseconds. Disk read I/O saturates. The buffer pool hit ratio that normally sits above 99% has collapsed. There are no runaway queries, no lock waits, and no checkpoint stalls. The server is simply cold.&lt;/p>
&lt;p>After any restart, the InnoDB buffer pool starts empty. Every data page read misses memory and goes to disk, so &lt;code>Innodb_buffer_pool_reads&lt;/code> spikes. If the adaptive hash index is enabled, it is empty too. The server may accept client connections long before the pool is ready to serve production traffic. The issue looks like a capacity crisis, but it is usually transient and expected. Distinguishing a cold cache from a true capacity problem keeps you from chasing the wrong fixes during an incident.&lt;/p></description></item><item><title>MySQL slow commits with idle CPU: redo and binlog fsync pressure</title><link>https://www.netdata.cloud/guides/mysql/mysql-fsync-pressure-slow-commits/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-fsync-pressure-slow-commits/</guid><description>&lt;h1 id="mysql-slow-commits-with-idle-cpu-redo-and-binlog-fsync-pressure">MySQL slow commits with idle CPU: redo and binlog fsync pressure&lt;/h1>
&lt;p>Slow commits or transaction timeouts with idle CPU usually mean the bottleneck is in the durability path, not the query engine.&lt;/p>
&lt;p>With &lt;code>innodb_flush_log_at_trx_commit=1&lt;/code> and &lt;code>sync_binlog=1&lt;/code>, every &lt;code>COMMIT&lt;/code> triggers an InnoDB redo log &lt;code>fsync()&lt;/code> and a binary log &lt;code>fsync()&lt;/code> before returning to the client. These are sequential, latency-bound operations. When storage cannot drain fsync requests as fast as the server generates them, commits queue up while the CPU waits.&lt;/p></description></item><item><title>MySQL slow queries: from Slow_queries to the slow log to the fix</title><link>https://www.netdata.cloud/guides/mysql/mysql-slow-queries-diagnosis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-slow-queries-diagnosis/</guid><description>&lt;h1 id="mysql-slow-queries-from-slow_queries-to-the-slow-log-to-the-fix">MySQL slow queries: from Slow_queries to the slow log to the fix&lt;/h1>
&lt;p>When &lt;code>Slow_queries&lt;/code> climbs but the slow log is empty, or &lt;code>long_query_time&lt;/code> is still at the default 10 seconds, you have an instrumentation gap. &lt;code>Slow_queries&lt;/code> increments regardless of whether &lt;code>slow_query_log&lt;/code> is &lt;code>ON&lt;/code>, so a rising counter with no log detail is a dead end.&lt;/p>
&lt;p>This guide moves from the status counter to slow log configuration, then to &lt;code>performance_schema&lt;/code> digests, and finally to resource signals that explain why queries slowed. The goal is to help you decide in minutes whether you face a query plan regression, lock contention, buffer pool pressure, or a configuration gap.&lt;/p></description></item><item><title>MySQL Sort_merge_passes climbing: filesort and sort_buffer_size</title><link>https://www.netdata.cloud/guides/mysql/mysql-filesort-sort-merge-passes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-filesort-sort-merge-passes/</guid><description>&lt;h1 id="mysql-sort_merge_passes-climbing-filesort-and-sort_buffer_size">MySQL Sort_merge_passes climbing: filesort and sort_buffer_size&lt;/h1>
&lt;p>&lt;code>Sort_merge_passes&lt;/code> climbing fast is the signature of filesort spilling to disk. You notice it in &lt;code>SHOW GLOBAL STATUS&lt;/code>, along with rising disk I/O on the tmpdir partition and a growing &lt;code>Handler_read_rnd&lt;/code>. Queries that returned in milliseconds now take seconds.&lt;/p>
&lt;p>The absolute value of this cumulative counter is meaningless. Rate of change is what matters. A rapid climb within hours means sorts that used to stay in memory are now writing chunks to disk and merging them in multiple passes. Each pass costs disk I/O and adds latency. Active threads pile up behind slow sorts, driving &lt;code>Threads_running&lt;/code> higher and the &lt;code>Questions&lt;/code> rate lower.&lt;/p></description></item><item><title>MySQL stuck in InnoDB crash recovery: why startup hangs after an unclean shutdown</title><link>https://www.netdata.cloud/guides/mysql/mysql-crash-recovery-slow-startup/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-crash-recovery-slow-startup/</guid><description>&lt;h1 id="mysql-stuck-in-innodb-crash-recovery-why-startup-hangs-after-an-unclean-shutdown">MySQL stuck in InnoDB crash recovery: why startup hangs after an unclean shutdown&lt;/h1>
&lt;p>After an unclean shutdown, mysqld may look healthy: the process is running, port 3306 accepts connections, and &lt;code>SHOW GLOBAL STATUS LIKE 'Uptime'&lt;/code> counts up. Yet applications time out, &lt;code>SELECT 1&lt;/code> hangs, and the error log prints InnoDB recovery messages. InnoDB is replaying the redo log to make data files consistent, rolling back incomplete transactions, merging the change buffer, and running purge. Until that finishes, queries are rejected.&lt;/p></description></item><item><title>MySQL tablespace fragmentation: reclaiming space with OPTIMIZE TABLE</title><link>https://www.netdata.cloud/guides/mysql/mysql-tablespace-fragmentation-data-free/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-tablespace-fragmentation-data-free/</guid><description>&lt;h1 id="mysql-tablespace-fragmentation-reclaiming-space-with-optimize-table">MySQL tablespace fragmentation: reclaiming space with OPTIMIZE TABLE&lt;/h1>
&lt;p>InnoDB tablespaces fragment after bulk deletes, updates, and purge cycles. Logical data volume shrinks, but the on-disk file retains allocated extents because InnoDB returns freed pages to its internal free lists rather than the OS. For file-per-table &lt;code>.ibd&lt;/code> tables, that unused space is visible as &lt;code>DATA_FREE&lt;/code> in &lt;code>information_schema.TABLES&lt;/code> and can be reclaimed with &lt;code>OPTIMIZE TABLE&lt;/code>.&lt;/p>
&lt;p>Tables inside the shared system tablespace &lt;code>ibdata1&lt;/code> do not shrink at the OS level. Freed space returns to InnoDB&amp;rsquo;s internal free lists, but &lt;code>ibdata1&lt;/code> never releases extents back to the filesystem. If the goal is reclaiming disk space on a full volume, &lt;code>OPTIMIZE TABLE&lt;/code> against tables in &lt;code>ibdata1&lt;/code> is ineffective. The rest of this guide covers identifying reclaimable tables, running the rebuild safely, and verifying the result.&lt;/p></description></item><item><title>MySQL Threads_connected vs Threads_running: which one to actually alert on</title><link>https://www.netdata.cloud/guides/mysql/mysql-threads-connected-vs-threads-running/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-threads-connected-vs-threads-running/</guid><description>&lt;h1 id="mysql-threads_connected-vs-threads_running-which-one-to-actually-alert-on">MySQL Threads_connected vs Threads_running: which one to actually alert on&lt;/h1>
&lt;p>&lt;code>Threads_connected&lt;/code> counts every client holding an open socket, including idle pooled connections. &lt;code>Threads_running&lt;/code> counts only threads actively executing a statement. Paging on the first while ignoring the second is a common postmortem mistake. &lt;code>Threads_running&lt;/code> relative to CPU cores is the real load signal. The gap between the two metrics diagnoses failure modes. Connection count still deserves a ticket, but rarely a page.&lt;/p></description></item><item><title>MySQL Threads_created climbing: thread cache churn and missing pooling</title><link>https://www.netdata.cloud/guides/mysql/mysql-thread-cache-churn/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-thread-cache-churn/</guid><description>&lt;h1 id="mysql-threads_created-climbing-thread-cache-churn-and-missing-pooling">MySQL Threads_created climbing: thread cache churn and missing pooling&lt;/h1>
&lt;p>When the rate of &lt;code>Threads_created&lt;/code> climbs to tens or hundreds of new threads per minute, the thread cache is not absorbing your connection churn. In a pooled deployment this rate should stay near zero. Every cache miss pays the full cost of OS thread creation and initialization, plus a TLS handshake if &lt;code>require_secure_transport&lt;/code> is enabled. The first visible symptom is usually connection latency spikes or CPU time diverted to thread management.&lt;/p></description></item><item><title>MySQL TLS and encryption monitoring: unencrypted connections and weak auth plugins</title><link>https://www.netdata.cloud/guides/mysql/mysql-tls-encryption-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-tls-encryption-monitoring/</guid><description>&lt;h1 id="mysql-tls-and-encryption-monitoring-unencrypted-connections-and-weak-auth-plugins">MySQL TLS and encryption monitoring: unencrypted connections and weak auth plugins&lt;/h1>
&lt;p>MySQL 8.0 generates self-signed certificates on startup and enables TLS by default, but this does not guarantee encrypted traffic. Clients must explicitly request TLS over TCP, and misconfigurations often leave sessions in plaintext without errors. Long-lived accounts frequently remain on &lt;code>mysql_native_password&lt;/code> despite its deprecation in MySQL 8.0.34, while &lt;code>caching_sha2_password&lt;/code> requires secure transport or RSA key exchange on first connection, causing sporadic authentication failures when client drivers lack encryption configuration. These risks are invisible to standard availability monitoring. This guide covers the status variables, Performance Schema tables, and account audits needed to detect unencrypted connections and weak authentication plugins in production.&lt;/p></description></item><item><title>MySQL undo tablespace growing: ibdata1 bloat and undo truncation</title><link>https://www.netdata.cloud/guides/mysql/mysql-undo-tablespace-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-undo-tablespace-growing/</guid><description>&lt;h1 id="mysql-undo-tablespace-growing-ibdata1-bloat-and-undo-truncation">MySQL undo tablespace growing: ibdata1 bloat and undo truncation&lt;/h1>
&lt;p>Disk space alert: &lt;code>/var/lib/mysql/ibdata1&lt;/code> is 40 GB larger than last week, or the &lt;code>undo_*.ibu&lt;/code> files in MySQL 8.0 are consuming terabytes. Table sizes from &lt;code>information_schema&lt;/code> do not justify the growth. This is almost always undo log accumulation from a blocked purge thread.&lt;/p>
&lt;p>Undo logs let InnoDB reconstruct older row versions for MVCC and transaction rollback. The purge thread deletes them when no active transaction needs them. If one long-running transaction holds a read view, purge stops. History list length then grows linearly with write rate, and the undo tablespace expands to hold the backlog. In MySQL 5.7, that space lands in the system tablespace &lt;code>ibdata1&lt;/code>, which never shrinks. In MySQL 8.0, separate undo tablespaces can truncate automatically, but only after the history list drains and the purge thread frees rollback segments. Until then, the disk fills.&lt;/p></description></item><item><title>MySQL Waiting for table metadata lock: diagnosing the DDL stall</title><link>https://www.netdata.cloud/guides/mysql/mysql-waiting-for-table-metadata-lock/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-waiting-for-table-metadata-lock/</guid><description>&lt;h1 id="mysql-waiting-for-table-metadata-lock-diagnosing-the-ddl-stall">MySQL Waiting for table metadata lock: diagnosing the DDL stall&lt;/h1>
&lt;p>You run &lt;code>SHOW PROCESSLIST&lt;/code> and see a wall of threads in &lt;code>Waiting for table metadata lock&lt;/code>. An &lt;code>ALTER TABLE&lt;/code> that should finish in seconds hangs for minutes. Queries against one table stop returning, the application connection pool drains, and CPU and disk look calm. This is a metadata lock (MDL) stall. It is a queueing failure, not resource exhaustion. Left alone, it cascades into connection exhaustion and a partial outage where every other table continues to work normally.&lt;/p></description></item><item><title>MySQL: got a packet bigger than 'max_allowed_packet' bytes - causes and fixes</title><link>https://www.netdata.cloud/guides/mysql/mysql-max-allowed-packet-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/mysql/mysql-max-allowed-packet-errors/</guid><description>&lt;h1 id="mysql-got-a-packet-bigger-than-max_allowed_packet-bytes---causes-and-fixes">MySQL: got a packet bigger than &amp;lsquo;max_allowed_packet&amp;rsquo; bytes - causes and fixes&lt;/h1>
&lt;p>ER_NET_PACKET_TOO_LARGE (MySQL error 1153) drops the connection immediately. The payload crossed a limit on the client, the server, or a replica. A common mistake is tuning only the server: a source set to 256 MB still aborts a 20 MB payload if the &lt;code>mysql&lt;/code> client is at its default 16 MB. Replicas add their own hard and soft ceilings. This guide shows how to identify which boundary failed, raise it without a restart where possible, and prevent recurrence.&lt;/p></description></item><item><title>Mystrotv SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mystrotv-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/mystrotv-snmp-traps/</guid><description/></item><item><title>Nag LLC SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nag-llc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nag-llc-snmp-traps/</guid><description/></item><item><title>Nagios Monitoring</title><link>https://www.netdata.cloud/monitoring-101/nagios-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/nagios-monitoring/</guid><description>&lt;h2 id="nagios-monitoring">Nagios Monitoring&lt;/h2>
&lt;h3 id="what-is-nagios">What Is Nagios?&lt;/h3>
&lt;p>Nagios is a powerful, enterprise-grade IT infrastructure monitoring solution. It enables organizations to identify and rectify problems related to networks, systems, and applications by alerting users of the fault instances to ensure your IT systems are running efficiently. With a strong reputation in the field of &lt;a href="https://www.nagios.org/">network monitoring&lt;/a> and &lt;a href="https://www.nagios.org/what-is-nagios/">performance management&lt;/a>, Nagios has been a go-to solution for IT admins and engineers wanting to monitor critical IT infrastructure components.&lt;/p></description></item><item><title>Nagios Plugins and Custom Scripts</title><link>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/nagios-plugins-and-custom-scripts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/nagios-plugins-and-custom-scripts/</guid><description/></item><item><title>Nagios SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nagios-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nagios-snmp-traps/</guid><description/></item><item><title>named not responding on port 53: total outage versus UDP-works-TCP-fails</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-not-responding/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-not-responding/</guid><description>&lt;h1 id="named-not-responding-on-port-53-total-outage-versus-udp-works-tcp-fails">named not responding on port 53: total outage versus UDP-works-TCP-fails&lt;/h1>
&lt;p>&lt;code>named&lt;/code> can be alive (&lt;code>pgrep -x named&lt;/code> returns a PID) yet not answering on the serving interface. Process liveness checks pass, but clients experience timeouts and retries that propagate to every service depending on DNS.&lt;/p>
&lt;p>Two diagnostic forks determine the fix path. The first: total outage (both UDP and TCP fail on the serving interface) versus UDP-works-TCP-fails. A total outage means the main DNS path is broken. A UDP-works-TCP-fails condition leaves large responses, DNSSEC-heavy answers, and zone transfers at risk while UDP-only health checks stay green. The second fork: loopback success versus serving-interface success. A query that succeeds against &lt;code>127.0.0.1&lt;/code> confirms the daemon processes queries, not that real clients can reach it.&lt;/p></description></item><item><title>Nasuni Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nasuni-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nasuni-corporation-snmp-traps/</guid><description/></item><item><title>Nasuni Filer</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/nasuni-filer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/nasuni-filer/</guid><description/></item><item><title>NAT and session-table exhaustion: catching it before connections fail</title><link>https://www.netdata.cloud/guides/network/network-nat-session-table-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-nat-session-table-exhaustion/</guid><description>&lt;h1 id="nat-and-session-table-exhaustion-catching-it-before-connections-fail">NAT and session-table exhaustion: catching it before connections fail&lt;/h1>
&lt;p>New connections fail while existing ones keep working. Applications report &amp;ldquo;connection refused&amp;rdquo; or timeouts. Open SSH sessions stay alive, but new SSH attempts hang. Your monitoring shows the firewall or NAT gateway is up, interfaces are healthy, and CPU is normal. The session or NAT translation table is full.&lt;/p>
&lt;p>Session-table exhaustion is a cliff-edge failure. The table degrades gracefully until it hits its limit, then every new connection is denied. Existing flows continue because their entries are already in the table. The symptom pattern is distinctive but easy to misdiagnose as application failure, DNS issues, or upstream provider problems, because the applications are the ones reporting errors.&lt;/p></description></item><item><title>Nateks Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nateks-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nateks-ltd-snmp-traps/</guid><description/></item><item><title>National Standardization Committee Of Radio Television SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/national-standardization-committee-of-radio-television-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/national-standardization-committee-of-radio-television-snmp-traps/</guid><description/></item><item><title>NATS</title><link>https://www.netdata.cloud/integrations/data-collection/databases/nats/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/nats/</guid><description/></item><item><title>NATS /healthz explained: js-server-only vs js-enabled-only vs the bare check</title><link>https://www.netdata.cloud/guides/nats/nats-healthz-endpoint-explained/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-healthz-endpoint-explained/</guid><description>&lt;h1 id="nats-healthz-explained-js-server-only-vs-js-enabled-only-vs-the-bare-check">NATS /healthz explained: js-server-only vs js-enabled-only vs the bare check&lt;/h1>
&lt;p>A NATS server restarts with a large JetStream store. The process is fine, clients will be served shortly, but your pager fires and Kubernetes kills the pod before recovery finishes. The server starts recovering again, the probe fails again, and you are in a restart loop that looks like an outage but is a health check misconfiguration.&lt;/p>
&lt;p>The cause is almost always the same: a liveness probe or a PAGE alert pointed at the bare &lt;code>/healthz&lt;/code> endpoint on a JetStream-enabled server. On such a server, bare &lt;code>/healthz&lt;/code> does not answer &amp;ldquo;is the process alive?&amp;rdquo; It runs the full JetStream health suite, and that suite legitimately fails for minutes while the server replays and recovers its assets after a restart.&lt;/p></description></item><item><title>NATS Authentication Timeout: clients that connect but never finish the handshake</title><link>https://www.netdata.cloud/guides/nats/nats-authentication-timeout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-authentication-timeout/</guid><description>&lt;h1 id="nats-authentication-timeout-clients-that-connect-but-never-finish-the-handshake">NATS Authentication Timeout: clients that connect but never finish the handshake&lt;/h1>
&lt;p>Your NATS server log shows lines like this, and clients are complaining they cannot connect:&lt;/p>
&lt;pre tabindex="0">&lt;code>[ERR] 10.2.3.14:51824 - cid:1042 - Authentication Timeout
&lt;/code>&lt;/pre>&lt;p>The confusing part: the TCP connection succeeded. The client reached the server. But the authentication handshake never finished, so the server sent &lt;code>-ERR 'Authentication Timeout'&lt;/code> and closed the connection. This is not a bad-credentials problem. It is a timing problem: credentials (or the TLS handshake that must precede them) never arrived within the server&amp;rsquo;s auth window. For the related case where credentials arrive but are rejected, the log string is &lt;code>Authorization Violation&lt;/code> and the diagnosis is different.&lt;/p></description></item><item><title>NATS Authorization Violation: authentication failures and credential rotation</title><link>https://www.netdata.cloud/guides/nats/nats-authorization-violation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-authorization-violation/</guid><description>&lt;h1 id="nats-authorization-violation-authentication-failures-and-credential-rotation">NATS Authorization Violation: authentication failures and credential rotation&lt;/h1>
&lt;p>Your NATS server log is filling with &lt;code>Authorization Violation&lt;/code> entries, clients are failing to connect, and the error message tells you almost nothing. That is deliberate: NATS keeps auth error messages vague so they do not leak information to attackers. The side effect is that the same log line covers a typo&amp;rsquo;d password, an expired user JWT, a bad credential file deployed to a fleet, and someone port-scanning your cluster from the internet.&lt;/p></description></item><item><title>NATS connection churn: a stable connection count hiding constant reconnects</title><link>https://www.netdata.cloud/guides/nats/nats-connection-churn/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-connection-churn/</guid><description>&lt;h1 id="nats-connection-churn-a-stable-connection-count-hiding-constant-reconnects">NATS connection churn: a stable connection count hiding constant reconnects&lt;/h1>
&lt;p>Your NATS dashboard shows 4,000 client connections. It showed 4,000 an hour ago, and 4,000 yesterday. Everything looks stable. Meanwhile, clients are connecting and disconnecting hundreds of times per minute. Every reconnect burns CPU on protocol handshakes (and TLS handshakes, if enabled), the server logs fill with connect and disconnect events, and your auth system processes a constant stream of authentication attempts.&lt;/p></description></item><item><title>NATS connection storm: reconnect thundering herd after a network event</title><link>https://www.netdata.cloud/guides/nats/nats-connection-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-connection-storm/</guid><description>&lt;h1 id="nats-connection-storm-reconnect-thundering-herd-after-a-network-event">NATS connection storm: reconnect thundering herd after a network event&lt;/h1>
&lt;p>A network device fails and recovers. A load balancer health check flaps. A rolling restart drops a node. For a few seconds, every client attached to a NATS server loses its connection. Then the network heals, and every one of those clients tries to reconnect in the same second.&lt;/p>
&lt;p>Each reconnect is not cheap: TCP handshake, optional TLS handshake, protocol negotiation, authentication, and a full resubscribe of every subscription the client held. Multiply that by hundreds or thousands of clients arriving simultaneously and you have a connection storm: a sharp spike in &lt;code>connections&lt;/code> and &lt;code>total_connections&lt;/code>, a CPU spike dominated by TLS handshakes if encryption is enabled, memory climbing as per-connection state is allocated, and in the worst case the server hitting &lt;code>max_connections&lt;/code> or the OS file descriptor limit and rejecting clients. Rejected clients retry. Now you have a reject-and-retry loop on top of the storm.&lt;/p></description></item><item><title>NATS consumer stalled at MaxAckPending: delivery stops until messages are acked</title><link>https://www.netdata.cloud/guides/nats/nats-consumer-max-ack-pending/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-consumer-max-ack-pending/</guid><description>&lt;h1 id="nats-consumer-stalled-at-maxackpending-delivery-stops-until-messages-are-acked">NATS consumer stalled at MaxAckPending: delivery stops until messages are acked&lt;/h1>
&lt;p>Your JetStream consumer was processing messages fine, and then deliveries just stopped. No error on the client. No error in the server log. The stream keeps growing, CPU and memory are normal, and the consumer application is still connected. Everything is green except the one thing that matters: no messages are moving.&lt;/p>
&lt;p>This is the most common cause of &amp;ldquo;JetStream consumer stopped receiving messages,&amp;rdquo; and it is silent by design. When a consumer&amp;rsquo;s count of delivered-but-unacknowledged messages (&lt;code>num_ack_pending&lt;/code>) reaches its configured &lt;code>MaxAckPending&lt;/code> limit, the server stops delivering new messages to that consumer. No error is raised, no advisory is emitted. Delivery pauses until pending messages are acknowledged, negatively acknowledged, or expire past &lt;code>AckWait&lt;/code>.&lt;/p></description></item><item><title>NATS context deadline exceeded: JetStream publish and request timeouts</title><link>https://www.netdata.cloud/guides/nats/nats-context-deadline-exceeded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-context-deadline-exceeded/</guid><description>&lt;h1 id="nats-context-deadline-exceeded-jetstream-publish-and-request-timeouts">NATS context deadline exceeded: JetStream publish and request timeouts&lt;/h1>
&lt;p>Your application logs fill up with &lt;code>nats: context deadline exceeded&lt;/code>. It shows up on &lt;code>js.Publish()&lt;/code>, on &lt;code>js.StreamInfo()&lt;/code>, on consumer fetches, sometimes on &lt;code>nats&lt;/code> CLI commands like &lt;code>nats stream view&lt;/code>. The NATS server is running, &lt;code>/healthz&lt;/code> returns ok, and core NATS pub/sub traffic is flowing fine. Yet every JetStream operation hangs until the client&amp;rsquo;s context fires.&lt;/p>
&lt;p>This error is a client-side deadline expiring. The client sent a JetStream API request and no response came back before the deadline. The request is not being rejected; it is not being answered at all. Somewhere between the client and the JetStream subsystem, the request is stuck, and the server is usually still healthy enough to look innocent.&lt;/p></description></item><item><title>NATS crash loop: unexpected uptime resets and repeated restarts</title><link>https://www.netdata.cloud/guides/nats/nats-crash-loop-restarts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-crash-loop-restarts/</guid><description>&lt;h1 id="nats-crash-loop-unexpected-uptime-resets-and-repeated-restarts">NATS crash loop: unexpected uptime resets and repeated restarts&lt;/h1>
&lt;p>Your NATS server&amp;rsquo;s &lt;code>/varz&lt;/code> uptime keeps resetting. The server was up for 4 minutes, then 2 minutes, then 6. Clients reconnect repeatedly, JetStream streams flap between unavailable and recovering, and every restart replays the WAL from scratch. This is a crash loop, and the fix depends entirely on why the process is dying.&lt;/p>
&lt;p>The uptime field on &lt;code>/varz&lt;/code> is the fastest confirmation. It reports time since process start as a NATS-specific duration string with y/d/h/m/s suffixes (for example &amp;ldquo;1d2h3m4s&amp;rdquo;), not a standard Go duration. When that value drops between scrapes, the process restarted. A single restart may be maintenance. More than 3 restarts in 30 minutes is a crash loop and needs an owner.&lt;/p></description></item><item><title>NATS file descriptor exhaustion: too many open files and the ulimit cliff</title><link>https://www.netdata.cloud/guides/nats/nats-file-descriptor-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-file-descriptor-exhaustion/</guid><description>&lt;h1 id="nats-file-descriptor-exhaustion-too-many-open-files-and-the-ulimit-cliff">NATS file descriptor exhaustion: too many open files and the ulimit cliff&lt;/h1>
&lt;p>Clients suddenly cannot connect. The server process is alive, CPU is fine, memory is fine, but every new connection fails and the logs are full of &lt;code>too many open files&lt;/code>. Existing clients keep working, which makes it look like a network problem until you count sockets.&lt;/p>
&lt;p>This is file descriptor exhaustion, and it is a cliff-edge failure. There is no graceful degradation: the moment the process hits its OS &lt;code>ulimit -n&lt;/code>, &lt;code>accept()&lt;/code> starts failing, JetStream cannot open new storage files, and cluster routes cannot establish. It is one of the most common NATS incidents, because the default &lt;code>ulimit -n&lt;/code> of 1024 on Linux is far too low for a production message broker, while the default &lt;code>max_connections&lt;/code> in NATS is 65536. The OS limit is the real ceiling, and it is usually the one nobody set.&lt;/p></description></item><item><title>NATS gateway disconnected: cross-cluster traffic cut in a supercluster</title><link>https://www.netdata.cloud/guides/nats/nats-gateway-disconnected/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-gateway-disconnected/</guid><description>&lt;h1 id="nats-gateway-disconnected-cross-cluster-traffic-cut-in-a-supercluster">NATS gateway disconnected: cross-cluster traffic cut in a supercluster&lt;/h1>
&lt;p>Subscribers in cluster A have stopped receiving messages published in cluster B. Publishers see no errors. Every server&amp;rsquo;s &lt;code>/healthz&lt;/code> returns ok, CPU and memory look normal, and yet a whole class of traffic has silently stopped flowing between two sites. In a NATS supercluster, this is the signature of a gateway disconnection: the outbound gateway connection to the remote cluster is missing, so no cross-cluster forwarding happens at all.&lt;/p></description></item><item><title>NATS goroutine leak: connection cleanup that never completes</title><link>https://www.netdata.cloud/guides/nats/nats-goroutine-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-goroutine-leak/</guid><description>&lt;h1 id="nats-goroutine-leak-connection-cleanup-that-never-completes">NATS goroutine leak: connection cleanup that never completes&lt;/h1>
&lt;p>The symptom is a steady, one-directional climb. Goroutine count goes up day after day, RSS follows it, and nothing on the traffic side explains it: client connections are flat, routes are flat, throughput is flat. Then, weeks later, the server gets OOM-killed or starts showing GC and scheduling overhead that has nothing to do with message load.&lt;/p>
&lt;p>A NATS goroutine leak is connection (or subsystem) cleanup that never completes. A goroutine is spawned to handle a connection, a timer, or a Raft loop; the work ends, but the goroutine never exits. It sits blocked on a channel receive that will never fire, a lock that will never be released, or a retry loop with nothing to retry. Each one is cheap individually, so the leak is invisible until it is not: at roughly 4-8KB of stack per goroutine, 100k leaked goroutines is about 400-800MB of RSS.&lt;/p></description></item><item><title>NATS high CPU: subject matching, TLS, GC, and the container GOMAXPROCS trap</title><link>https://www.netdata.cloud/guides/nats/nats-high-cpu-usage/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-high-cpu-usage/</guid><description>&lt;h1 id="nats-high-cpu-subject-matching-tls-gc-and-the-container-gomaxprocs-trap">NATS high CPU: subject matching, TLS, GC, and the container GOMAXPROCS trap&lt;/h1>
&lt;p>A NATS server that is pegged on CPU usually looks worse than it is, or better than it is, depending on what you are measuring. The &lt;code>cpu&lt;/code> field in &lt;code>/varz&lt;/code> is process CPU where 100 means one full core, so on a 16-core host a value of 800 is only 50% busy. Operators regularly page themselves on a number that is not actually saturation, or dismiss real saturation because the number &amp;ldquo;only&amp;rdquo; reads 300 on a 2-core container.&lt;/p></description></item><item><title>NATS insufficient storage / maximum bytes exceeded: JetStream publishes rejected</title><link>https://www.netdata.cloud/guides/nats/nats-insufficient-storage/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-insufficient-storage/</guid><description>&lt;h1 id="nats-insufficient-storage--maximum-bytes-exceeded-jetstream-publishes-rejected">NATS insufficient storage / maximum bytes exceeded: JetStream publishes rejected&lt;/h1>
&lt;p>Your producers start throwing JetStream publish errors: &amp;ldquo;insufficient storage&amp;rdquo;, &amp;ldquo;maximum bytes exceeded&amp;rdquo;, or &amp;ldquo;maximum messages exceeded&amp;rdquo;. The server itself looks healthy: &lt;code>/healthz&lt;/code> returns ok, connections are stable, CPU and memory are normal. Writes to one or more streams are failing anyway.&lt;/p>
&lt;p>This is a storage-limit failure, not a server failure. JetStream enforces limits at three levels: per-stream (&lt;code>max_bytes&lt;/code>, &lt;code>max_msgs&lt;/code>, &lt;code>max_age&lt;/code>), per-account storage quotas, and the server-wide JetStream storage reservation. When any of them is hit, the discard policy decides the failure mode. With &lt;code>DiscardNew&lt;/code>, new publishes are rejected and the publisher sees an error. With &lt;code>DiscardOld&lt;/code>, the server silently evicts the oldest messages to make room, which is data loss for any consumer that has not caught up.&lt;/p></description></item><item><title>NATS JetStream AckWait tuning: matching the ack timeout to processing time</title><link>https://www.netdata.cloud/guides/nats/nats-ack-wait-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-ack-wait-tuning/</guid><description>&lt;h1 id="nats-jetstream-ackwait-tuning-matching-the-ack-timeout-to-processing-time">NATS JetStream AckWait tuning: matching the ack timeout to processing time&lt;/h1>
&lt;p>AckWait is the timer the JetStream server starts when it delivers a message to a consumer. If the client does not ack before the timer expires, the server redelivers the message. That is the entire mechanism, and most &amp;ldquo;JetStream consumer is stuck&amp;rdquo; incidents that are not crashes come down to this timer being mismatched to the actual processing time of the work behind it.&lt;/p></description></item><item><title>NATS JetStream API errors: reading the /jsz api.errors counter without false alarms</title><link>https://www.netdata.cloud/guides/nats/nats-jetstream-api-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-jetstream-api-errors/</guid><description>&lt;h1 id="nats-jetstream-api-errors-reading-the-jsz-apierrors-counter-without-false-alarms">NATS JetStream API errors: reading the /jsz api.errors counter without false alarms&lt;/h1>
&lt;p>You opened &lt;code>/jsz&lt;/code> and saw &lt;code>api.errors&lt;/code> in the thousands, or an alert fired because the counter moved. The first question is not &amp;ldquo;what broke&amp;rdquo; but &amp;ldquo;is this counter telling me something real.&amp;rdquo; The JetStream API error counter is cumulative, coarse-grained, and incremented by perfectly healthy client behavior. Alerting on its absolute value or on any movement at all is a guaranteed false-alarm generator.&lt;/p></description></item><item><title>NATS JetStream consumer lag growing: falling behind the stream</title><link>https://www.netdata.cloud/guides/nats/nats-consumer-lag-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-consumer-lag-growing/</guid><description>&lt;h1 id="nats-jetstream-consumer-lag-growing-falling-behind-the-stream">NATS JetStream consumer lag growing: falling behind the stream&lt;/h1>
&lt;p>The NATS server is healthy. &lt;code>/healthz&lt;/code> returns ok, throughput looks normal, CPU and memory are fine. But one stream keeps growing, and the consumer attached to it is not keeping up. &lt;code>num_pending&lt;/code> climbs minute after minute, and every dashboard that only watches server-level metrics shows green.&lt;/p>
&lt;p>Consumer lag requires per-consumer polling, not a single server endpoint, so most monitoring setups never see it. The result: a stream with millions of pending messages, effectively down while the server looks healthy.&lt;/p></description></item><item><title>NATS JetStream consumer stopped receiving messages: the diagnostic tree</title><link>https://www.netdata.cloud/guides/nats/nats-consumer-stopped-receiving/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-consumer-stopped-receiving/</guid><description>&lt;h1 id="nats-jetstream-consumer-stopped-receiving-messages-the-diagnostic-tree">NATS JetStream consumer stopped receiving messages: the diagnostic tree&lt;/h1>
&lt;p>A JetStream consumer that stops receiving messages almost never raises an error. The stream keeps accepting publishes, the health endpoint stays green, and your application simply goes quiet. The signals that explain this live per-consumer, not per-server, which is why server-level dashboards show nothing.&lt;/p>
&lt;p>There are only a handful of mechanisms that produce this symptom, and each has a distinct fingerprint in consumer state. This article is the diagnostic tree. Start at the top with &lt;code>nats consumer info&lt;/code> (or &lt;code>/jsz?consumers=true&lt;/code>), read five counters, and the tree tells you which failure you are in and where to go next.&lt;/p></description></item><item><title>NATS JetStream disabled unexpectedly: the persistence subsystem failed to come up</title><link>https://www.netdata.cloud/guides/nats/nats-jetstream-disabled-unexpectedly/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-jetstream-disabled-unexpectedly/</guid><description>&lt;h1 id="nats-jetstream-disabled-unexpectedly-the-persistence-subsystem-failed-to-come-up">NATS JetStream disabled unexpectedly: the persistence subsystem failed to come up&lt;/h1>
&lt;p>The server is up. Clients connect, core NATS routes messages, &lt;code>/healthz?js-server-only=true&lt;/code> returns ok. But &lt;code>/jsz&lt;/code> reports &lt;code>disabled: true&lt;/code> on a node where JetStream should be running, and every stream, consumer, and KV bucket backed by this node is gone or degraded. Publishers that rely on persistence fail; request-reply still works, which is why this gets noticed late.&lt;/p>
&lt;p>This state means one of two things: JetStream failed to initialize at startup, or it initialized and was later shut down by the server itself. In both cases the root cause is almost always below the NATS layer: the storage directory, the disk underneath it, or the configuration. The server rarely recovers on its own, and a blind restart can make a corrupt store worse.&lt;/p></description></item><item><title>NATS JetStream disk I/O stall: the disk has space but is too slow</title><link>https://www.netdata.cloud/guides/nats/nats-jetstream-disk-io-stall/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-jetstream-disk-io-stall/</guid><description>&lt;h1 id="nats-jetstream-disk-io-stall-the-disk-has-space-but-is-too-slow">NATS JetStream disk I/O stall: the disk has space but is too slow&lt;/h1>
&lt;p>JetStream publishes are timing out or creeping upward in latency. &lt;code>df -h&lt;/code> shows plenty of free space. The &lt;code>/jsz&lt;/code> endpoint shows storage nowhere near its reserved limits, yet &lt;code>api.inflight&lt;/code> sits high and &lt;code>api.errors&lt;/code> keeps climbing. On the host, iowait is elevated and disk latency looks bad.&lt;/p>
&lt;p>This is the JetStream disk I/O stall pattern: the disk has space but is too slow. It is a different failure from storage exhaustion, and it is frequently misdiagnosed because the obvious capacity checks all pass.&lt;/p></description></item><item><title>NATS JetStream meta cluster leader flapping: admin operations that keep failing</title><link>https://www.netdata.cloud/guides/nats/nats-meta-cluster-leader-flapping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-meta-cluster-leader-flapping/</guid><description>&lt;h1 id="nats-jetstream-meta-cluster-leader-flapping-admin-operations-that-keep-failing">NATS JetStream meta cluster leader flapping: admin operations that keep failing&lt;/h1>
&lt;p>Stream and consumer administration is failing, but existing JetStream traffic mostly continues. Creating a stream times out, consumer updates return errors, and retries sometimes succeed only to fail again a minute later.&lt;/p>
&lt;p>This pattern points at the JetStream meta Raft group. The meta group manages cluster-wide JetStream metadata: stream and consumer create, update, and delete operations. Existing streams use separate Raft groups for replicated data, so the data plane can remain available while the administrative plane is unstable.&lt;/p></description></item><item><title>NATS JetStream mirror and source lag: stale replicas and DR recovery point</title><link>https://www.netdata.cloud/guides/nats/nats-mirror-source-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-mirror-source-lag/</guid><description>&lt;h1 id="nats-jetstream-mirror-and-source-lag-stale-replicas-and-dr-recovery-point">NATS JetStream mirror and source lag: stale replicas and DR recovery point&lt;/h1>
&lt;p>A JetStream stream configured as a mirror of another stream, or sourcing from one or more upstream streams, keeps a local copy that is only as fresh as its sync connection. Stream info reports two fields per mirror or source: &lt;code>lag&lt;/code>, the number of messages the local copy is behind, and &lt;code>active&lt;/code>, the time since the last sync activity. When those numbers move the wrong way, the copy is going stale.&lt;/p></description></item><item><title>NATS JetStream no leader / cluster not currently available: writes blocked</title><link>https://www.netdata.cloud/guides/nats/nats-no-leader/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-no-leader/</guid><description>&lt;h1 id="nats-jetstream-no-leader--cluster-not-currently-available-writes-blocked">NATS JetStream no leader / cluster not currently available: writes blocked&lt;/h1>
&lt;p>Your publishers are failing with &lt;code>nats: no leader for stream&lt;/code> or &lt;code>JetStream cluster not currently available&lt;/code>, or every JetStream API call times out with &lt;code>no responders&lt;/code> on &lt;code>$JS.API&lt;/code>. Stream info shows an empty &lt;code>cluster.leader&lt;/code>. Writes stay paused until the affected Raft group elects a leader again.&lt;/p>
&lt;p>This is a quorum problem, not a load problem. JetStream uses Raft for consensus: one meta group manages all JetStream metadata and admin operations, and each replicated stream (and consumer) has its own Raft group. A group without a leader blocks everything that depends on it. If the meta group is leaderless, all JetStream admin operations fail cluster-wide. If only a stream&amp;rsquo;s group is leaderless, that stream stops accepting messages while the rest of JetStream keeps working.&lt;/p></description></item><item><title>NATS JetStream not enabled for account: persistence calls failing on a core server</title><link>https://www.netdata.cloud/guides/nats/nats-jetstream-not-enabled/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-jetstream-not-enabled/</guid><description>&lt;h1 id="nats-jetstream-not-enabled-for-account-persistence-calls-failing-on-a-core-server">NATS JetStream not enabled for account: persistence calls failing on a core server&lt;/h1>
&lt;p>Your application connects to NATS successfully, then fails on the first persistence call: stream creation, consumer creation, or a publish that waits for an ack. The client error is one of two strings: &lt;code>nats: jetstream not enabled for account&lt;/code> (err_code 10039) or &lt;code>nats: jetstream not enabled&lt;/code> (err_code 10076). Core pub/sub keeps working the whole time, which is what makes this confusing during an incident.&lt;/p></description></item><item><title>NATS JetStream Raft election storm: leaders flapping and writes pausing repeatedly</title><link>https://www.netdata.cloud/guides/nats/nats-raft-election-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-raft-election-storm/</guid><description>&lt;h1 id="nats-jetstream-raft-election-storm-leaders-flapping-and-writes-pausing-repeatedly">NATS JetStream Raft election storm: leaders flapping and writes pausing repeatedly&lt;/h1>
&lt;p>Your JetStream cluster is electing leaders over and over. Logs alternate rapidly between &lt;code>Stepping down&lt;/code> and &lt;code>JetStream cluster new leader&lt;/code>. Publishes intermittently time out, stream and consumer management calls error out, and some streams briefly report no leader at all. Then it settles for a few minutes and starts again.&lt;/p>
&lt;p>This is a Raft election storm. It is not a clean failover and it is not usually a software bug. It is a symptom of resource starvation or latency somewhere in the cluster, and it feeds itself: elections consume CPU and I/O, which delays heartbeats further, which triggers more elections.&lt;/p></description></item><item><title>NATS JetStream redelivery loop: num_redelivered climbing and messages reprocessed</title><link>https://www.netdata.cloud/guides/nats/nats-consumer-redelivery-loop/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-consumer-redelivery-loop/</guid><description>&lt;h1 id="nats-jetstream-redelivery-loop-num_redelivered-climbing-and-messages-reprocessed">NATS JetStream redelivery loop: num_redelivered climbing and messages reprocessed&lt;/h1>
&lt;p>A JetStream consumer&amp;rsquo;s &lt;code>num_redelivered&lt;/code> counter is climbing and your application is processing the same messages over and over. Nothing is lost, but nothing completes either: deliveries happen, acks do not, and the server keeps trying again.&lt;/p>
&lt;p>This is a redelivery loop. The server delivers a message, the consumer fails to acknowledge it within &lt;code>AckWait&lt;/code>, the server redelivers it, and each pass increments &lt;code>num_redelivered&lt;/code>. Every redelivery is wasted work: duplicate side effects if your handler is not idempotent, duplicate load on downstream systems, and a &lt;code>num_ack_pending&lt;/code> count that never drains.&lt;/p></description></item><item><title>NATS JetStream replica lag: a non-current replica that would lose data on failover</title><link>https://www.netdata.cloud/guides/nats/nats-stream-replica-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-stream-replica-lag/</guid><description>&lt;h1 id="nats-jetstream-replica-lag-a-non-current-replica-that-would-lose-data-on-failover">NATS JetStream replica lag: a non-current replica that would lose data on failover&lt;/h1>
&lt;p>A replicated JetStream stream shows &lt;code>current: false&lt;/code> with a non-zero &lt;code>lag&lt;/code> for one of its replicas in &lt;code>nats stream info&lt;/code>. Everything still works: publishes succeed, consumers receive messages, no alerts fire on server health. But the replication factor you configured is not the replication factor you have. If the leader fails right now, a failover to the lagging replica loses every message the leader acknowledged but the replica has not yet written.&lt;/p></description></item><item><title>NATS JetStream retention: limits vs interest vs workqueue and the /dev/null stream</title><link>https://www.netdata.cloud/guides/nats/nats-retention-policy-confusion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-retention-policy-confusion/</guid><description>&lt;h1 id="nats-jetstream-retention-limits-vs-interest-vs-workqueue-and-the-devnull-stream">NATS JetStream retention: limits vs interest vs workqueue and the /dev/null stream&lt;/h1>
&lt;p>JetStream&amp;rsquo;s retention policy decides when a stored message is deleted, and the failure modes are asymmetric: get it wrong in one direction and storage grows until publishes are rejected; get it wrong in the other and every message you publish is deleted on arrival while the server reports itself perfectly healthy.&lt;/p>
&lt;p>This article covers the three policies (&lt;code>limits&lt;/code>, &lt;code>interest&lt;/code>, &lt;code>workqueue&lt;/code>), the exact conditions under which each one deletes a message, the silent &lt;code>/dev/null&lt;/code> failure mode, and the checks that confirm your streams are retaining what you think they are. It assumes a working mental model of streams, consumers, and acks. If not, start with &lt;a href="https://www.netdata.cloud/guides/nats/nats-how-it-works-in-production/">how NATS actually works in production&lt;/a>.&lt;/p></description></item><item><title>NATS JetStream storage exhaustion spiral: stalled consumers that starve retention</title><link>https://www.netdata.cloud/guides/nats/nats-storage-exhaustion-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-storage-exhaustion-spiral/</guid><description>&lt;h1 id="nats-jetstream-storage-exhaustion-spiral-stalled-consumers-that-starve-retention">NATS JetStream storage exhaustion spiral: stalled consumers that starve retention&lt;/h1>
&lt;p>JetStream storage is climbing steadily. Not a spike, a slope. Then publishes start failing with API errors, producers back up, and the pipeline degrades. The server looks healthy: connections stable, CPU and memory fine. The disk is filling anyway.&lt;/p>
&lt;p>This is the JetStream storage exhaustion spiral: a deadlock where one or more consumers have stalled (usually pinned at MaxAckPending) on a stream using &lt;code>interest&lt;/code> or &lt;code>workqueue&lt;/code> retention. The server cannot delete a message until every interested consumer has acknowledged it, so a single dead consumer makes its entire backlog undeletable. Messages accumulate, storage approaches the configured limit, and new publishes are rejected. The system is stuck: consumers must process messages to free space, but the consumers are the broken part.&lt;/p></description></item><item><title>NATS leaf node disconnected: an edge server isolated from the hub</title><link>https://www.netdata.cloud/guides/nats/nats-leaf-node-disconnected/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-leaf-node-disconnected/</guid><description>&lt;h1 id="nats-leaf-node-disconnected-an-edge-server-isolated-from-the-hub">NATS leaf node disconnected: an edge server isolated from the hub&lt;/h1>
&lt;p>A leaf node connection is the single TCP session that ties an edge NATS server to your hub cluster. When it drops, the edge site keeps running locally, but it is cut off from the rest of the messaging fabric. Subscribers on the edge stop receiving messages published at the hub, and subscribers at the hub stop receiving anything published at the edge.&lt;/p></description></item><item><title>NATS Maximum Connections Exceeded: new clients rejected at the max_connections wall</title><link>https://www.netdata.cloud/guides/nats/nats-maximum-connections-exceeded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-maximum-connections-exceeded/</guid><description>&lt;h1 id="nats-maximum-connections-exceeded-new-clients-rejected-at-the-max_connections-wall">NATS Maximum Connections Exceeded: new clients rejected at the max_connections wall&lt;/h1>
&lt;p>New clients cannot connect to your NATS server. Existing connections keep working. Client logs show the protocol error &lt;code>-ERR 'Maximum Connections Exceeded'&lt;/code> and then the connection closes. The server itself looks healthy: &lt;code>/healthz&lt;/code> returns ok, messages still flow for connected clients, and nothing crashed.&lt;/p>
&lt;p>This is the &lt;code>max_connections&lt;/code> wall. The server has reached its configured connection limit and is rejecting every new connection during the handshake. There is no queueing and no graceful degradation. One slot short of the limit, everything works. At the limit, every new client is turned away.&lt;/p></description></item><item><title>NATS Maximum Payload Violation: messages rejected for exceeding max_payload</title><link>https://www.netdata.cloud/guides/nats/nats-maximum-payload-violation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-maximum-payload-violation/</guid><description>&lt;h1 id="nats-maximum-payload-violation-messages-rejected-for-exceeding-max_payload">NATS Maximum Payload Violation: messages rejected for exceeding max_payload&lt;/h1>
&lt;p>A publisher calls publish, the call fails, and the NATS server log shows a Maximum Payload Violation. Seconds later the same client reconnects, publishes again, and gets disconnected again. The server is not broken: it is enforcing the configured &lt;code>max_payload&lt;/code> limit, which defaults to 1 MB. The application, meanwhile, is in a publish-disconnect-reconnect loop and its messages are not flowing.&lt;/p>
&lt;p>The protocol behavior is strict. When a client sends a message whose payload exceeds &lt;code>max_payload&lt;/code>, the server responds with &lt;code>-ERR 'Maximum Payload Violation'&lt;/code> and closes the connection. Client libraries that auto-reconnect come straight back and repeat the offense, which is why a single oversized publish looks like connection churn rather than a clean, one-time error.&lt;/p></description></item><item><title>NATS memory growth and OOM: reading RSS past the Go GC sawtooth</title><link>https://www.netdata.cloud/guides/nats/nats-memory-growth-oom/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-memory-growth-oom/</guid><description>&lt;h1 id="nats-memory-growth-and-oom-reading-rss-past-the-go-gc-sawtooth">NATS memory growth and OOM: reading RSS past the Go GC sawtooth&lt;/h1>
&lt;p>The nats-server process is climbing toward its container memory limit. Grafana shows a jagged line that spikes to nearly double the baseline and drops back, except lately the drops are getting shallower and the floor keeps rising. Then the pod restarts, &lt;code>uptime&lt;/code> resets to a few minutes, and the cycle repeats. If JetStream is enabled, the numbers look even worse, and half of what you see is not memory the server is actually holding.&lt;/p></description></item><item><title>NATS messages published but not received: subject mismatches and cross-server gaps</title><link>https://www.netdata.cloud/guides/nats/nats-messages-published-not-received/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-messages-published-not-received/</guid><description>&lt;h1 id="nats-messages-published-but-not-received-subject-mismatches-and-cross-server-gaps">NATS messages published but not received: subject mismatches and cross-server gaps&lt;/h1>
&lt;p>The publisher&amp;rsquo;s &lt;code>Publish()&lt;/code> call returns no error. The server is healthy. The subscriber receives nothing. This is almost never a server bug. It is how core NATS works: fire-and-forget routing against an in-memory subject tree, with no persistence, no delivery guarantee, and no signal when a message matches zero subscriptions.&lt;/p>
&lt;p>In core NATS, a message published to a subject with zero matching subscribers is silently discarded. No error, no log line, no metric. The only server-level evidence is an asymmetry between &lt;code>in_msgs&lt;/code> and &lt;code>out_msgs&lt;/code>, which most teams never graph.&lt;/p></description></item><item><title>NATS Monitoring</title><link>https://www.netdata.cloud/monitoring-101/nats-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/nats-monitoring/</guid><description>&lt;h2 id="nats-monitoring">NATS Monitoring&lt;/h2>
&lt;h3 id="what-is-nats">What Is NATS?&lt;/h3>
&lt;p>&lt;a href="https://nats.io/">NATS&lt;/a> is a high-performance messaging system that enables distributed applications to generate and consume microservices in a seamless manner. It is widely used in various industries for cloud, IoT, and edge applications due to its lightweight and scalable architecture.&lt;/p>
&lt;h3 id="monitoring-nats-with-netdata">Monitoring NATS With Netdata&lt;/h3>
&lt;p>Netdata offers a cutting-edge NATS monitoring tool that provides real-time insights into the performance and health of your NATS servers. By leveraging Netdata&amp;rsquo;s capabilities, organizations can keep their messaging infrastructure efficient and reliable. For an interactive experience, you can try out our &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">Live Demo&lt;/a>.&lt;/p></description></item><item><title>NATS monitoring checklist: the signals every production server needs</title><link>https://www.netdata.cloud/guides/nats/nats-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-monitoring-checklist/</guid><description>&lt;h1 id="nats-monitoring-checklist-the-signals-every-production-server-needs">NATS monitoring checklist: the signals every production server needs&lt;/h1>
&lt;p>A NATS server can pass every health check and still be losing messages. Core NATS drops messages to slow or absent subscribers by design, JetStream consumers can stall while every server-level metric looks green, and a node can be partitioned from its cluster while its local health endpoint returns 200. &amp;ldquo;Is the process running?&amp;rdquo; is the start of a monitoring strategy, not the end of one.&lt;/p></description></item><item><title>NATS monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/nats/nats-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-monitoring-maturity-model/</guid><description>&lt;h1 id="nats-monitoring-maturity-model-from-survival-to-expert">NATS monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most NATS outages are not caused by a lack of metrics. The server exposes a rich monitoring API on port 8222. The failures happen because teams collect the wrong tier of signals for the failures they actually experience. A &lt;code>/healthz&lt;/code> check tells you the process is alive. It tells you nothing about a consumer stalled at &lt;code>MaxAckPending&lt;/code>, a route connection backing up with pending bytes, or a Raft meta cluster electing a new leader every ninety seconds.&lt;/p></description></item><item><title>NATS no responders available for request: request-reply into the void</title><link>https://www.netdata.cloud/guides/nats/nats-no-responders-available/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-no-responders-available/</guid><description>&lt;h1 id="nats-no-responders-available-for-request-request-reply-into-the-void">NATS no responders available for request: request-reply into the void&lt;/h1>
&lt;p>Your service just started throwing &lt;code>nats: no responders available for request&lt;/code>. The client got an answer back from the server almost instantly, and the answer was: nobody is listening on that subject. This is the request-reply counterpart to NATS&amp;rsquo;s silent message loss: instead of the request vanishing and the client waiting out its timeout, the server short-circuits the call and fails fast with a 503 status in the reply headers.&lt;/p></description></item><item><title>NATS pending bytes growing: catching a slow consumer before it is disconnected</title><link>https://www.netdata.cloud/guides/nats/nats-pending-bytes-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-pending-bytes-growing/</guid><description>&lt;h1 id="nats-pending-bytes-growing-catching-a-slow-consumer-before-it-is-disconnected">NATS pending bytes growing: catching a slow consumer before it is disconnected&lt;/h1>
&lt;p>By the time &lt;code>slow_consumers&lt;/code> in &lt;code>/varz&lt;/code> increments, the damage is done. The server has declared a connection slow, and in core NATS the default response is to disconnect it. Messages buffered for that client are dropped, not queued, not retried. The client library auto-reconnects, resubscribes, immediately falls behind on the same backlog, and you are in a churn loop.&lt;/p></description></item><item><title>NATS Permissions Violation: authorized clients publishing or subscribing out of scope</title><link>https://www.netdata.cloud/guides/nats/nats-permissions-violation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-permissions-violation/</guid><description>&lt;h1 id="nats-permissions-violation-authorized-clients-publishing-or-subscribing-out-of-scope">NATS Permissions Violation: authorized clients publishing or subscribing out of scope&lt;/h1>
&lt;p>A &lt;code>Permissions Violation&lt;/code> entry in the NATS server log means an authenticated client tried to publish or subscribe to a subject outside its permitted scope. The server rejected the operation, logged the event with the subject, client IP, and account, and sent an &lt;code>-ERR&lt;/code> back to the client. The connection stays open. This is not an authentication failure: the credentials were accepted, but the permissions attached to them do not cover what the client tried to do.&lt;/p></description></item><item><title>NATS route disconnected: a missing cluster route means a partition</title><link>https://www.netdata.cloud/guides/nats/nats-route-disconnected/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-route-disconnected/</guid><description>&lt;h1 id="nats-route-disconnected-a-missing-cluster-route-means-a-partition">NATS route disconnected: a missing cluster route means a partition&lt;/h1>
&lt;p>One of your NATS servers shows fewer routes than it should. In a full-mesh cluster of N servers, every server should hold N-1 route connections, one to each peer. When that count drops, you do not have a degraded link. You have a partition.&lt;/p>
&lt;p>The impact is asymmetric and easy to underestimate. The isolated server still accepts client connections, still reports healthy on &lt;code>/healthz&lt;/code>, and still routes messages locally. But subscribers connected to it stop receiving messages published on the other side of the partition, and publishers on it vanish from the rest of the cluster&amp;rsquo;s view. If you run JetStream with replicated streams, Raft groups whose members span the partition can lose quorum, which turns a messaging gap into write failures.&lt;/p></description></item><item><title>NATS route RTT high: inter-server latency that triggers Raft elections</title><link>https://www.netdata.cloud/guides/nats/nats-cluster-route-rtt-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-cluster-route-rtt-high/</guid><description>&lt;h1 id="nats-route-rtt-high-inter-server-latency-that-triggers-raft-elections">NATS route RTT high: inter-server latency that triggers Raft elections&lt;/h1>
&lt;p>You are looking at a JetStream cluster that keeps electing new leaders. The &lt;code>/jsz&lt;/code> meta cluster leader changes every few minutes, stream writes intermittently fail with API errors, and &lt;code>nats server list&lt;/code> or &lt;code>/routez&lt;/code> shows round-trip times between servers far above what a same-datacenter cluster should produce. Server CPU and memory look fine. Nothing has crashed. Yet the cluster cannot hold a stable leader.&lt;/p></description></item><item><title>NATS route slow consumer: inter-server forwarding backing up cluster-wide</title><link>https://www.netdata.cloud/guides/nats/nats-route-slow-consumer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-route-slow-consumer/</guid><description>&lt;h1 id="nats-route-slow-consumer-inter-server-forwarding-backing-up-cluster-wide">NATS route slow consumer: inter-server forwarding backing up cluster-wide&lt;/h1>
&lt;p>Your NATS server&amp;rsquo;s &lt;code>slow_consumer_stats.routes&lt;/code> counter just went non-zero, or &lt;code>/routez&lt;/code> shows a &lt;code>pending_size&lt;/code> that keeps climbing on one route. This is not the same problem as a slow client. A route is the TCP connection that carries inter-server traffic between two NATS servers in a cluster. When it backs up, every subscriber reachable through that peer falls behind or stops receiving messages entirely, and in a JetStream cluster the degradation can extend to Raft heartbeat timing and leader stability.&lt;/p></description></item><item><title>NATS server not responding: healthz failing and the process down or hung</title><link>https://www.netdata.cloud/guides/nats/nats-server-not-responding/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-server-not-responding/</guid><description>&lt;h1 id="nats-server-not-responding-healthz-failing-and-the-process-down-or-hung">NATS server not responding: healthz failing and the process down or hung&lt;/h1>
&lt;p>Your probe against the NATS monitoring port is failing: &lt;code>/healthz&lt;/code> returns non-200, or the port does not answer at all. Clients may still be connected, reconnecting in a storm, or already failing over to other cluster nodes.&lt;/p>
&lt;p>Do not start with a restart. A crashed process, a hung event loop, lame-duck shutdown, resource exhaustion, and JetStream recovery can all look like &amp;ldquo;NATS is down,&amp;rdquo; and the first safe action is different for each. Restarting too early destroys evidence and can turn a recoverable hang into data loss.&lt;/p></description></item><item><title>NATS silent message loss: zero-subscriber drops and the in/out asymmetry</title><link>https://www.netdata.cloud/guides/nats/nats-silent-message-loss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-silent-message-loss/</guid><description>&lt;h1 id="nats-silent-message-loss-zero-subscriber-drops-and-the-inout-asymmetry">NATS silent message loss: zero-subscriber drops and the in/out asymmetry&lt;/h1>
&lt;p>Your NATS server is healthy. Health checks pass, connection counts are stable, there are no slow consumers, no errors in the logs. And yet a downstream service insists it has not received a single message in the last hour, while the producer&amp;rsquo;s metrics show it published thousands.&lt;/p>
&lt;p>Both are telling the truth. In core NATS, a message published to a subject with zero subscribers is discarded at the moment of publish. The publisher gets no error. The server writes no log line. No counter records the drop. The message simply does not exist anymore.&lt;/p></description></item><item><title>NATS slow consumer breakdown: clients vs routes vs gateways and blast radius</title><link>https://www.netdata.cloud/guides/nats/nats-slow-consumer-clients-routes-gateways/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-slow-consumer-clients-routes-gateways/</guid><description>&lt;h1 id="nats-slow-consumer-breakdown-clients-vs-routes-vs-gateways-and-blast-radius">NATS slow consumer breakdown: clients vs routes vs gateways and blast radius&lt;/h1>
&lt;p>Your NATS alert fired: &lt;code>slow_consumers&lt;/code> is incrementing. You open &lt;code>/varz&lt;/code>, see a non-zero counter, and start hunting for a misbehaving client. That is the right move only some of the time. The aggregate &lt;code>slow_consumers&lt;/code> counter mixes four very different kinds of events, and two of them mean your cluster fabric itself is backing up, not a single subscriber.&lt;/p></description></item><item><title>NATS slow consumer detected: the write buffer overflowed and messages were dropped</title><link>https://www.netdata.cloud/guides/nats/nats-slow-consumer-detected/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-slow-consumer-detected/</guid><description>&lt;h1 id="nats-slow-consumer-detected-the-write-buffer-overflowed-and-messages-were-dropped">NATS slow consumer detected: the write buffer overflowed and messages were dropped&lt;/h1>
&lt;p>Your NATS server log shows &lt;code>Slow Consumer Detected&lt;/code>, or your client logged &lt;code>nats: slow consumer, messages dropped&lt;/code>. In core NATS (no JetStream), that line means messages were dropped. Not queued, not retried, not parked somewhere for later.&lt;/p>
&lt;p>The server buffers outbound messages per connection, and when a subscriber cannot drain its write buffer fast enough, the server sheds load by dropping messages for that subscriber or disconnecting it outright. Delivery in core NATS is at-most-once, and the slow subscriber is not told what was lost.&lt;/p></description></item><item><title>NATS stalled clients and stale connections: half-dead sockets and write-path distress</title><link>https://www.netdata.cloud/guides/nats/nats-stalled-clients-stale-connections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-stalled-clients-stale-connections/</guid><description>&lt;h1 id="nats-stalled-clients-and-stale-connections-half-dead-sockets-and-write-path-distress">NATS stalled clients and stale connections: half-dead sockets and write-path distress&lt;/h1>
&lt;p>Two counters in &lt;code>/varz&lt;/code> tell you about connections that are unhealthy but not yet dead: &lt;code>stalled_clients&lt;/code> and &lt;code>stale_connections&lt;/code>. Neither one means a connection has been dropped. That is the point: both describe states that precede visible failure. Stalled clients are on the way to becoming slow consumers; stale connections are sockets that look open but whose peer has stopped responding.&lt;/p></description></item><item><title>NATS stream max_bytes: one stream hitting its limit while the server has room</title><link>https://www.netdata.cloud/guides/nats/nats-stream-max-bytes-limit/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-stream-max-bytes-limit/</guid><description>&lt;h1 id="nats-stream-max_bytes-one-stream-hitting-its-limit-while-the-server-has-room">NATS stream max_bytes: one stream hitting its limit while the server has room&lt;/h1>
&lt;p>Publishers to one JetStream stream start getting rejected with a &amp;ldquo;maximum bytes exceeded&amp;rdquo; error. You check the server: JetStream storage is at 40% of &lt;code>reserved_storage&lt;/code>, the disk has hundreds of gigabytes free, and every server-level dashboard is green. Nothing about the server looks full.&lt;/p>
&lt;p>What happened: &lt;code>max_bytes&lt;/code> is a per-stream limit, independent of server-level and account-level JetStream quotas. One stream reached its own configured ceiling and started enforcing its discard policy while every other stream kept working. If your monitoring only watches aggregate JetStream storage (&lt;code>/jsz&lt;/code> without parameters), this failure is invisible until publishers start erroring.&lt;/p></description></item><item><title>NATS subscription leak: unbounded subscription growth and subject-trie bloat</title><link>https://www.netdata.cloud/guides/nats/nats-subscription-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-subscription-leak/</guid><description>&lt;h1 id="nats-subscription-leak-unbounded-subscription-growth-and-subject-trie-bloat">NATS subscription leak: unbounded subscription growth and subject-trie bloat&lt;/h1>
&lt;p>The symptom is a number that only goes up. You open &lt;code>/varz&lt;/code> on a NATS server and &lt;code>subscriptions&lt;/code> is higher than it was yesterday, higher than it was last week, climbing for days while &lt;code>connections&lt;/code> stays flat. Memory follows the same slope. Nothing is erroring. No slow consumers, no restarts, no complaints from applications. The server just keeps getting heavier.&lt;/p>
&lt;p>That is a subscription leak: an application (or a bug) is registering subscriptions faster than it removes them. Every subscription lives in the server&amp;rsquo;s subject trie and in per-connection tracking state. Left alone, the leak consumes memory, slows subject matching, and in a cluster propagates interest across routes so that one leaky client inflates the subject trie on every server.&lt;/p></description></item><item><title>NATS TLS certificate expiry: an expired cert that locks out every new connection</title><link>https://www.netdata.cloud/guides/nats/nats-tls-certificate-expiry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-tls-certificate-expiry/</guid><description>&lt;h1 id="nats-tls-certificate-expiry-an-expired-cert-that-locks-out-every-new-connection">NATS TLS certificate expiry: an expired cert that locks out every new connection&lt;/h1>
&lt;p>Every new client connection to your NATS server is failing at the TLS handshake with a verification error, typically &lt;code>x509: certificate has expired or is not yet valid&lt;/code>. Clients that were already connected still work. That is what makes this failure mode deceptive: the server passes every process-level check while refusing all new work.&lt;/p>
&lt;p>If you run mutual TLS on cluster routes and gateways, the blast radius is bigger. Route and gateway connections also fail to establish, so servers that restart or reconnect after a network event cannot rejoin the cluster. A certificate that expires during an unrelated incident can turn a recoverable blip into a full partition.&lt;/p></description></item><item><title>NATS write_deadline and buffer sizing: tuning how long the server waits on a slow writer</title><link>https://www.netdata.cloud/guides/nats/nats-write-deadline-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nats/nats-write-deadline-tuning/</guid><description>&lt;h1 id="nats-write_deadline-and-buffer-sizing-tuning-how-long-the-server-waits-on-a-slow-writer">NATS write_deadline and buffer sizing: tuning how long the server waits on a slow writer&lt;/h1>
&lt;p>A client gets disconnected with &lt;code>Slow Consumer Detected&lt;/code> in the server log. Someone raises &lt;code>write_deadline&lt;/code>, the disconnects stop, and the ticket gets closed. Two weeks later the same client is back in the log, and now the server is also showing memory growth because the larger buffer window lets more messages pile up per connection. This is the standard lifecycle of a write_deadline tuning mistake.&lt;/p></description></item><item><title>Nature Remo E lite devices</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/nature-remo-e-lite-devices/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/nature-remo-e-lite-devices/</guid><description/></item><item><title>Nature Remo E lite devices Monitoring</title><link>https://www.netdata.cloud/monitoring-101/nature_remo-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/nature_remo-monitoring/</guid><description>&lt;h2 id="nature-remo-monitoring">Nature Remo Monitoring&lt;/h2>
&lt;h3 id="what-is-nature-remo">What Is Nature Remo?&lt;/h3>
&lt;p>Nature Remo E lite devices are innovative smart home solutions designed to enhance home automation and energy management. These devices enable you to control your home environment, focusing on convenience and efficiency, by integrating seamlessly into your smart home ecosystem.&lt;/p>
&lt;h3 id="monitoring-nature-remo-with-netdata">Monitoring Nature Remo With Netdata&lt;/h3>
&lt;p>To monitor Nature Remo, Netdata employs an openmetrics (Prometheus) exporter, allowing the ingestion of data from any Prometheus exporter. This setup ensures you get automated dashboards, real-time alerts, and a comprehensive monitoring experience without needing a Prometheus server or Grafana. &lt;a href="https://github.com/kenfdev/remo-exporter">Get the community exporter here&lt;/a> to easily set up Nature Remo monitoring.&lt;/p></description></item><item><title>Nbase Switch Communication SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nbase-switch-communication-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nbase-switch-communication-snmp-traps/</guid><description/></item><item><title>NEC BGP</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/nec-bgp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/nec-bgp/</guid><description/></item><item><title>Nec Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nec-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nec-corporation-snmp-traps/</guid><description/></item><item><title>NEC Univerge</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/nec-univerge/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/nec-univerge/</guid><description/></item><item><title>Neoteris Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/neoteris-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/neoteris-inc-snmp-traps/</guid><description/></item><item><title>NET Framework</title><link>https://www.netdata.cloud/integrations/data-collection/applications/net-framework/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/net-framework/</guid><description/></item><item><title>Net Insight AB SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/net-insight-ab-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/net-insight-ab-snmp-traps/</guid><description/></item><item><title>Net Snmp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/net-snmp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/net-snmp-snmp-traps/</guid><description/></item><item><title>Net To Net Technologies SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/net-to-net-technologies-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/net-to-net-technologies-snmp-traps/</guid><description/></item><item><title>Net Track GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/net-track-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/net-track-gmbh-snmp-traps/</guid><description/></item><item><title>Net-SNMP Host</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/net-snmp-host/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/net-snmp-host/</guid><description/></item><item><title>net.inet.icmp.stats</title><link>https://www.netdata.cloud/integrations/data-collection/networking/net.inet.icmp.stats/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/net.inet.icmp.stats/</guid><description/></item><item><title>net.inet.ip.stats</title><link>https://www.netdata.cloud/integrations/data-collection/networking/net.inet.ip.stats/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/net.inet.ip.stats/</guid><description/></item><item><title>net.inet.tcp.states</title><link>https://www.netdata.cloud/integrations/data-collection/networking/net.inet.tcp.states/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/net.inet.tcp.states/</guid><description/></item><item><title>net.inet.tcp.stats</title><link>https://www.netdata.cloud/integrations/data-collection/networking/net.inet.tcp.stats/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/net.inet.tcp.stats/</guid><description/></item><item><title>net.inet.udp.stats</title><link>https://www.netdata.cloud/integrations/data-collection/networking/net.inet.udp.stats/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/net.inet.udp.stats/</guid><description/></item><item><title>net.inet6.icmp6.stats</title><link>https://www.netdata.cloud/integrations/data-collection/networking/net.inet6.icmp6.stats/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/net.inet6.icmp6.stats/</guid><description/></item><item><title>net.inet6.ip6.stats</title><link>https://www.netdata.cloud/integrations/data-collection/networking/net.inet6.ip6.stats/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/net.inet6.ip6.stats/</guid><description/></item><item><title>net.isr</title><link>https://www.netdata.cloud/integrations/data-collection/networking/net.isr/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/net.isr/</guid><description/></item><item><title>Netapp</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/netapp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/netapp/</guid><description/></item><item><title>Netapp ONTAP API</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/netapp-ontap-api/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/netapp-ontap-api/</guid><description/></item><item><title>NetApp ONTAP API Monitoring</title><link>https://www.netdata.cloud/monitoring-101/netapp_ontap-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/netapp_ontap-monitoring/</guid><description>&lt;h2 id="netapp-ontap-api-monitoring">NetApp ONTAP API Monitoring&lt;/h2>
&lt;h3 id="what-is-netapp-ontap-api">What Is NetApp ONTAP API?&lt;/h3>
&lt;p>NetApp ONTAP API, renowned for its robust data management capabilities, is a storage operating system designed to efficiently manage and protect petabytes of data. Its performance directly impacts storage efficiency, making it essential for technical professionals, such as DevOps, IT admins, and SREs, to keep an attentive eye on its metrics. Monitoring the performance and health of the NetApp ONTAP API is critical in ensuring data integrity and optimizing storage functionalities.&lt;/p></description></item><item><title>NetApp Solidfire</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/netapp-solidfire/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/netapp-solidfire/</guid><description/></item><item><title>NetApp Solidfire Monitoring</title><link>https://www.netdata.cloud/monitoring-101/netapp_solidfire-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/netapp_solidfire-monitoring/</guid><description>&lt;h2 id="netapp-solidfire-monitoring">NetApp Solidfire Monitoring&lt;/h2>
&lt;h3 id="what-is-netapp-solidfire">What Is NetApp Solidfire?&lt;/h3>
&lt;p>NetApp Solidfire is a high-performing storage solution designed to manage data efficiently across your storage infrastructure. It provides seamless data handling capabilities that ensure optimal resource allocation and management. Solidfire is instrumental in handling workloads with demanding storage and performance requirements.&lt;/p>
&lt;h3 id="monitoring-netapp-solidfire-with-netdata">Monitoring NetApp Solidfire With Netdata&lt;/h3>
&lt;p>When it comes to monitoring NetApp Solidfire, Netdata takes advantage of an openmetrics (Prometheus) exporter. The &lt;a href="https://github.com/mjavier2k/solidfire-exporter">NetApp Solidfire Exporter&lt;/a> is utilized to gather pertinent metrics, which are essential for monitoring and enhancing overall system performance.&lt;/p></description></item><item><title>Netatmo sensors</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/netatmo-sensors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/netatmo-sensors/</guid><description/></item><item><title>Netatmo Sensors Monitoring</title><link>https://www.netdata.cloud/monitoring-101/netatmo-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/netatmo-monitoring/</guid><description>&lt;h2 id="netatmo-sensors-monitoring">Netatmo Sensors Monitoring&lt;/h2>
&lt;h3 id="what-are-netatmo-sensors">What Are Netatmo Sensors?&lt;/h3>
&lt;p>Netatmo sensors are a range of smart home devices designed to monitor environmental conditions such as temperature, humidity, and air quality. These IoT devices play a crucial role in home automation and energy management, enabling users to adapt their environments for comfort and efficiency.&lt;/p>
&lt;h3 id="monitoring-netatmo-sensors-with-netdata">Monitoring Netatmo Sensors With Netdata&lt;/h3>
&lt;p>Effective tools for monitoring Netatmo Sensors are essential for ensuring optimal performance and timely insights. Netdata provides a comprehensive solution for monitoring Netatmo Sensors without the complexity of setting up separate Prometheus servers or Grafana dashboards. By using an openmetrics (Prometheus) exporter, Netdata can ingest data from the &lt;a href="https://github.com/xperimental/netatmo-exporter">Netatmo exporter&lt;/a>, offering automated dashboards and alerts to keep track of your smart home devices&amp;rsquo; performance in real time.&lt;/p></description></item><item><title>Netbotz SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netbotz-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netbotz-snmp-traps/</guid><description/></item><item><title>NetBox</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/netbox/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/netbox/</guid><description/></item><item><title>Netcomm Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netcomm-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netcomm-ltd-snmp-traps/</guid><description/></item><item><title>Netdata at AWS Cloud Days 2023</title><link>https://www.netdata.cloud/events/aws-cloud-days-2023/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/aws-cloud-days-2023/</guid><description>&lt;p>Ralph Meijer, VP of Technology at Netdata, joined a panel at AWS Cloud Days Athens 2023 on October 10. The session, &amp;ldquo;Running Managed Containers in AI-Powered Startups,&amp;rdquo; ran from 16:45 to 17:15 and focused on the operational side of containerized workloads in fast-moving companies.&lt;/p>
&lt;p>The discussion covered what it actually takes to run containers in production when your team is small and your infrastructure is changing weekly. Observability came up repeatedly &amp;ndash; specifically the gap between what managed container services give you out of the box and what you actually need to debug a problem at 2 AM. Ralph brought Netdata&amp;rsquo;s perspective: auto-detection of containers, per-second metrics without manual configuration, and the ability to monitor ephemeral workloads that spin up and disappear before traditional tools finish their first collection cycle.&lt;/p></description></item><item><title>Netdata at Civo Navigate 2023</title><link>https://www.netdata.cloud/events/civo-navigate-2023/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/civo-navigate-2023/</guid><description>&lt;p>Netdata sent two speakers to Civo Navigate Europe 2023 in London on September 5&amp;ndash;6. Ralph Meijer, VP of Technology, opened with &amp;ldquo;Opinionated Observability&amp;rdquo; at noon on day one &amp;ndash; a talk about the design trade-offs behind monitoring defaults and why those choices matter more than most vendors admit. The argument: when your tool ships with good opinions about what to collect, how to visualize it, and when to alert, you remove an entire class of configuration toil.&lt;/p></description></item><item><title>Netdata at Conf42 Cloud Native 2024</title><link>https://www.netdata.cloud/events/conf42-cloud-native-2024/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/conf42-cloud-native-2024/</guid><description>&lt;p>Costa Tsaousis spoke at Conf42 Cloud Native 2024 (virtual, March 2024) with &amp;ldquo;Practical AI with Machine Learning for Observability in Netdata.&amp;rdquo; The talk was a technical walkthrough of how Netdata applies unsupervised machine learning to metrics &amp;ndash; not as a feature checkbox, but as a way to surface problems that static thresholds miss.&lt;/p>
&lt;p>The key insight Costa presented: individual anomalies on individual metrics are often noise. A CPU spike on one node, a latency bump on one service &amp;ndash; these happen constantly and mean nothing on their own. But when anomalies converge across multiple metrics and services simultaneously, that convergence is a strong signal that something unusual is actually happening. As he put it: &amp;ldquo;The power of ML becomes evident when seemingly noisy anomalies converge across various services, serving as indicators of something exceedingly unusual.&amp;rdquo;&lt;/p></description></item><item><title>Netdata at Conf42 DevOps 2024</title><link>https://www.netdata.cloud/events/conf42-devops-2024/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/conf42-devops-2024/</guid><description>&lt;p>Costa Tsaousis spoke at Conf42 DevOps 2024 on January 25 with a talk titled &amp;ldquo;Observability Standardization: the elephant is still in the room!&amp;rdquo; The event was virtual, so the audience was global.&lt;/p>
&lt;p>The premise of the talk was blunt: the industry talks endlessly about open standards and protocol compatibility for observability, but the real question is whether we actually need custom monitoring solutions for every infrastructure scenario. OpenTelemetry, Prometheus exposition format, StatsD, SNMP &amp;ndash; there is no shortage of standards. The problem is not the number of protocols. It is that the underlying data models, collection patterns, and retention strategies are so different across tools that &amp;ldquo;standardization&amp;rdquo; often just means another translation layer rather than true interoperability.&lt;/p></description></item><item><title>Netdata at Conf42 DevOps 2025</title><link>https://www.netdata.cloud/events/conf42-devops-2025/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/conf42-devops-2025/</guid><description>&lt;p>Shyam Sreevalsan spoke at Conf42 DevOps 2025 on January 23 (virtual) about the future of observability. The talk covered three threads: AI-driven anomaly detection, real-time processing at the edge, and predictive analytics as a replacement for reactive alerting.&lt;/p>
&lt;p>The central argument was that monitoring is shifting from &amp;ldquo;tell me when something breaks&amp;rdquo; to &amp;ldquo;tell me before something breaks.&amp;rdquo; Proactive monitoring requires two things that most traditional tools lack: per-second data collection (because you cannot predict what you cannot see) and ML models that run continuously rather than on a query schedule. Shyam walked through how Netdata handles both &amp;ndash; distributed agents collecting every second, unsupervised ML training on each metric individually, and anomaly correlation that surfaces coordinated deviations before they become incidents.&lt;/p></description></item><item><title>Netdata at Data Centre World London 2025</title><link>https://www.netdata.cloud/events/data-centre-world-london-2025/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/data-centre-world-london-2025/</guid><description>&lt;p>Netdata exhibited at Data Centre World London 2025, March 12&amp;ndash;13 at ExCeL London, at Booth DC476 in the co-located Tech Show London event. We ran live infrastructure monitoring demos throughout both days and held a LEGO draw that produced two winners.&lt;/p>
&lt;p>The audience here was different from a developer conference. Data centre operators and infrastructure managers care about things like power monitoring, cooling efficiency, network throughput, and server fleet health &amp;ndash; all at scale, all in real time. The conversations at the booth centered on three recurring pain points: the cost of existing observability tools (especially per-GB ingestion pricing), the complexity of maintaining multiple monitoring systems for different parts of the stack, and the time wasted on manual troubleshooting when dashboards show 60-second averages that hide the actual problem.&lt;/p></description></item><item><title>Netdata at DevOops Athens Meetup 2024</title><link>https://www.netdata.cloud/events/devops-athens-meetup-2024/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/devops-athens-meetup-2024/</guid><description>&lt;p>Costa Tsaousis spoke at the DevOops Athens Meetup in December 2024, telling the origin story of Netdata. The project was born out of a production incident &amp;ndash; a real problem that existing monitoring tools failed to diagnose quickly enough. That experience shaped everything that followed: the insistence on per-second granularity, the zero-configuration philosophy, the decision to run monitoring at the edge rather than depending on a remote backend.&lt;/p>
&lt;p>For a local meetup in Athens, the tone was more personal than a typical conference talk. Costa could talk about the early days of the project, the decisions that felt risky at the time, and how those choices played out over the years. The audience was Athens-based DevOps practitioners &amp;ndash; people who work in the Greek tech ecosystem and, in many cases, had been following the project since its early days.&lt;/p></description></item><item><title>Netdata at DevOps Meetup Finland 2024</title><link>https://www.netdata.cloud/events/devops-meetup-finland-2024/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/devops-meetup-finland-2024/</guid><description>&lt;p>Shyam Sreevalsan presented &amp;ldquo;Opinionated Observability&amp;rdquo; at the DevOps Finland Meetup in May 2024. The talk picked up a thread from Ralph Meijer&amp;rsquo;s earlier presentation at Civo Navigate and adapted it for a meetup audience &amp;ndash; shorter, more conversational, and grounded in specific examples.&lt;/p>
&lt;p>The core question: what happens when you move from collecting metrics to actually visualizing and alerting on them? That transition is where most monitoring setups break down. You install an agent, metrics flow into a database, and then you spend weeks building dashboards, tuning thresholds, and configuring notification rules. Shyam walked through the trade-offs involved in making those decisions for the user &amp;ndash; good defaults for visualization, sensible alert thresholds out of the box, and the risks of being too opinionated versus not opinionated enough.&lt;/p></description></item><item><title>Netdata at DevOps Summer Retreat Latvia 2024</title><link>https://www.netdata.cloud/events/devops-summer-retreat-2024/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/devops-summer-retreat-2024/</guid><description>&lt;p>Shyam Sreevalsan spoke at the DevOps &amp;amp; Agile Summer Retreat in Latvia in July 2024. His talk, &amp;ldquo;The Future of DevOps: The Next 10 Years,&amp;rdquo; was an attempt to lay out a technical roadmap for where the discipline is heading &amp;ndash; not vague predictions, but specific shifts already underway.&lt;/p>
&lt;p>Three themes anchored the talk. First, AI-driven automation: not replacing engineers, but handling the repetitive diagnostic and remediation tasks that consume on-call time. Second, DevSecOps becoming the default rather than the exception &amp;ndash; security as a continuous, embedded practice rather than a gate at the end of a pipeline. Third, decentralized autonomous teams with decentralized tools and workflows, where each team owns its full stack including observability, rather than relying on a central platform team to configure monitoring for them.&lt;/p></description></item><item><title>Netdata at DevOpsDays Geneva 2024</title><link>https://www.netdata.cloud/events/devops-days-geneva-2024/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/devops-days-geneva-2024/</guid><description>&lt;p>Costa Tsaousis spoke at DevOpsDays Geneva 2024 with a talk that had one of the more specific titles we have brought to a conference: &amp;ldquo;A powerful logs management solution we all have and use, but we underestimate: systemd-journal.&amp;rdquo; We also had a booth.&lt;/p>
&lt;p>The argument was simple. Every Linux system running systemd already has a structured, indexed, binary log store built in. Most teams ignore it and ship logs to an external system immediately, paying for ingestion, storage, and query infrastructure that duplicates something the OS already provides. Costa walked through how systemd-journal works, what it stores, how Netdata integrates with it for log exploration and correlation with metrics, and where it falls short (multi-node aggregation being the obvious gap).&lt;/p></description></item><item><title>Netdata at DEVworld Amsterdam 2025</title><link>https://www.netdata.cloud/events/devworld-amsterdam-2025/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/devworld-amsterdam-2025/</guid><description>&lt;p>Netdata exhibited at DEVworld Conference in Amsterdam in February 2025, at Booth 18E. No talk at this one &amp;ndash; just the team at the booth, running live demos and talking to developers.&lt;/p>
&lt;p>DEVworld attracts a broad developer audience, not just infrastructure or DevOps specialists. That meant explaining Netdata to people who might not have thought much about monitoring beyond &amp;ldquo;is my app up or down.&amp;rdquo; The live demos worked well for this: showing real-time, per-second dashboards to someone who has only ever seen delayed, aggregated metrics makes the difference immediately obvious.&lt;/p></description></item><item><title>Netdata at FOCS 2026 — The Future of Cybersecurity Summit</title><link>https://www.netdata.cloud/events/focs-malaysia-2026/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/focs-malaysia-2026/</guid><description>&lt;p>Netdata is at the 7th Future of Cybersecurity Summit (FOCS) 2026, organized by the PIKOM Cybersecurity Chapter in Petaling Jaya, Malaysia, alongside our regional partner &lt;a href="https://www.cloudengined.com/">Cloud Engine Digital (CED)&lt;/a>. CED is an exhibition partner at FOCS and brings expertise in cloud infrastructure, security, and data engineering across Southeast Asia. This year&amp;rsquo;s theme is &amp;ldquo;Cybersecurity 2030: Building Resilience in an AI-Driven, Borderless World,&amp;rdquo; and the agenda covers AI-driven security operations, predictive risk management, and the regulatory landscape across the region.&lt;/p></description></item><item><title>Netdata at FOSDEM 2024</title><link>https://www.netdata.cloud/events/fosdem-2024/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/fosdem-2024/</guid><description>&lt;p>Costa Tsaousis presented in the Observability Devroom at FOSDEM 2024 in Brussels, starting at 16:30. The talk covered the journey of Netdata &amp;ndash; how it started, the architectural decisions that shaped it, and the challenges of building a distributed monitoring system in the open.&lt;/p>
&lt;p>FOSDEM is not a vendor event. The audience is developers and maintainers who care about how things work, not how they are marketed. Costa&amp;rsquo;s talk leaned into that: the early design goal of per-second granularity, the choice to run everything at the edge rather than relying on a centralized backend, and the ongoing tension between keeping the project simple for individual users while scaling it for organizations with thousands of nodes.&lt;/p></description></item><item><title>Netdata at FOSSCOMM 2024</title><link>https://www.netdata.cloud/events/fosscomm-2024/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/fosscomm-2024/</guid><description>&lt;p>Costa Tsaousis spoke at FOSSCOMM 2024, held November 9&amp;ndash;10 at the University of Macedonia in Thessaloniki. FOSSCOMM is Greece&amp;rsquo;s largest open-source conference, and for Netdata &amp;ndash; a project with Greek roots &amp;ndash; it was a homecoming of sorts.&lt;/p>
&lt;p>The talk, &amp;ldquo;Netdata: Open Source, Distributed Observability Pipeline &amp;ndash; Journey and Challenges,&amp;rdquo; covered the full arc: how the project started, the architectural decisions that defined it, the community that grew around it, and the ongoing challenges of maintaining a large open-source codebase that millions of nodes depend on. Costa did not shy away from the hard parts &amp;ndash; the difficulty of balancing open-source community expectations with commercial product development, and the engineering cost of supporting an agent that runs on everything from a Raspberry Pi to a 256-core production server.&lt;/p></description></item><item><title>Netdata at Gartner IOCS Las Vegas 2025</title><link>https://www.netdata.cloud/events/gartner-iocs-2025/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/gartner-iocs-2025/</guid><description>&lt;p>Netdata exhibited at the Gartner IT Infrastructure, Operations &amp;amp; Cloud Strategies Conference (IOCS) 2025, December 9&amp;ndash;11 at The Venetian in Las Vegas. We were at Booth #636 in the Solution Village: Operations section, running live demos every 30 minutes and meeting with IT leaders from enterprises and mid-market organizations.&lt;/p>
&lt;p>The highlight was Costa Tsaousis&amp;rsquo;s talk on December 11 at 12:30 PM PST in Theater 3: &amp;ldquo;Why 80% of Organizations Will Overspend for Observability in 2026.&amp;rdquo; The thesis was direct. Legacy observability tools have a broken economic model &amp;ndash; they charge by data volume, which means the more you monitor, the more you pay. That creates a perverse incentive to sample, average, and reduce data, which in turn creates blind spots. Organizations end up paying more for less visibility. Costa walked through how sampling at 15&amp;ndash;60 second intervals masks the very spikes and anomalies that cause incidents, and how that hidden cost &amp;ndash; longer MTTR, more outages, more manual investigation &amp;ndash; often exceeds the monitoring bill itself.&lt;/p></description></item><item><title>Netdata at GITEX AI ASIA 2026</title><link>https://www.netdata.cloud/events/gitex-asia-2026/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/gitex-asia-2026/</guid><description>&lt;p>Netdata is at GITEX AI ASIA 2026, April 9–10 at Marina Bay Sands in Singapore, together with our regional partner &lt;a href="https://www.cloudengined.com/">Cloud Engine Digital (CED)&lt;/a>. GITEX AI ASIA is the region&amp;rsquo;s largest tech and AI event, drawing enterprise leaders, startups, and investors from over 110 countries.&lt;/p>
&lt;p>We&amp;rsquo;re showing live demos of Netdata&amp;rsquo;s real-time monitoring—per-second granularity across infrastructure, from Kubernetes clusters to bare metal servers. For teams running AI and ML workloads, this means capturing GPU utilization spikes, training pipeline anomalies, and inference latency fluctuations that 15–60 second sampling intervals simply cannot see.&lt;/p></description></item><item><title>Netdata at Howard Conference and Expo 2026</title><link>https://www.netdata.cloud/events/howard-expo-2026/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/howard-expo-2026/</guid><description>&lt;p>Netdata exhibited at the Howard Conference and Expo &amp;ldquo;Game On&amp;rdquo; event, February 24&amp;ndash;26 at the Grand Hotel Marriott Resort in Fairhope, Alabama. Constantine Nikitiadis and Stuart McMurran were on the ground, running demos and talking with IT leaders attending the conference.&lt;/p>
&lt;p>The demos focused on two things: per-second granularity versus the 15&amp;ndash;60 second sampling that most monitoring tools default to, and AI-assisted troubleshooting that turns hours of dashboard investigation into minutes of directed analysis. For an audience of IT practitioners managing diverse infrastructure &amp;ndash; bare metal, VMs, cloud, some Kubernetes &amp;ndash; the &amp;ldquo;works everywhere&amp;rdquo; message landed well. One agent, one install, coverage across the full stack.&lt;/p></description></item><item><title>Netdata at India DevOps Show 2025</title><link>https://www.netdata.cloud/events/india-devops-show-2025/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/india-devops-show-2025/</guid><description>&lt;p>Netdata was a Co-Presenting Partner at the 7th Edition India DevOps Show on May 23, 2025 at the Holiday Inn in Mumbai. Satyadeep Ashwathnarayana, Constantine Nikitiadis, and Shyam Sreevalsan represented the team.&lt;/p>
&lt;p>The talk &amp;ndash; &amp;ldquo;Stop Building Dashboards. Start Solving Problems.&amp;rdquo; &amp;ndash; challenged the default approach to observability that most DevOps teams fall into: install a monitoring tool, spend weeks building dashboards, and then stare at those dashboards during incidents trying to find the problem. The argument is that dashboards are a means, not an end. If your monitoring tool can surface the problem directly &amp;ndash; through anomaly detection, automated correlation, and AI-driven investigation &amp;ndash; then the dashboard becomes a verification step, not a diagnostic one.&lt;/p></description></item><item><title>Netdata at India DevOps Show 2026</title><link>https://www.netdata.cloud/events/india-devops-show-2026/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/india-devops-show-2026/</guid><description>&lt;p>Netdata participated as a Silver Partner at the 10th Edition India DevOps Show on February 13, 2026 at the Aloft ORR Hotel in Bengaluru. Shyam Sreevalsan and Constantine Nikitiadis represented the team, meeting with DevOps practitioners and tech leaders from across India.&lt;/p>
&lt;p>The focus was on showing how per-second metrics change the way teams respond to deployments and incidents. When you can see the impact of a deployment within seconds rather than waiting for the next polling interval, rollback decisions get faster and more confident. We demonstrated Netdata&amp;rsquo;s AI-assisted troubleshooting &amp;ndash; how it analyzes anomalies across the stack and produces actionable investigation reports in minutes rather than the hours of manual dashboard-digging that most teams are used to.&lt;/p></description></item><item><title>Netdata at ObservabilityCon 2023</title><link>https://www.netdata.cloud/events/observabilitycon-2023/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/observabilitycon-2023/</guid><description>&lt;p>Netdata had a booth at ObservabilityCon 2023 in London, hosted by Grafana Labs. No talk this time &amp;ndash; just the team at a booth, ready for deep-dive conversations about how different observability tools fit together.&lt;/p>
&lt;p>The event drew people who are already invested in open-source observability, many of them running Grafana alongside other tools. That made for pointed, technical conversations. A common thread: people were interested in how Netdata handles high-resolution metrics collection at the edge without requiring a heavy backend. The idea of collecting per-second data, running ML-based anomaly detection locally, and then feeding results into Grafana dashboards resonated with teams already comfortable in that ecosystem.&lt;/p></description></item><item><title>Netdata at Open Source Monitoring Conference 2024</title><link>https://www.netdata.cloud/events/open-source-monitoring-conference-2024/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/open-source-monitoring-conference-2024/</guid><description>&lt;p>Costa Tsaousis opened the final day of the Open Source Monitoring Conference (OSMC) 2024 in Nuremberg on November 21, speaking from 9:30 to 10:00 at the Jacobi venue. His talk, &amp;ldquo;Netdata: Open Source, Distributed Observability Pipeline &amp;ndash; Journey and Challenge,&amp;rdquo; covered the project&amp;rsquo;s evolution and the architectural decisions that set it apart from other open-source monitoring tools.&lt;/p>
&lt;p>OSMC runs November 19&amp;ndash;21 and is one of the more established events in the European IT monitoring space. The attendees are people who run Icinga, Checkmk, Prometheus, Zabbix, and similar tools in production. They know monitoring. What they wanted to hear was how Netdata&amp;rsquo;s distributed approach &amp;ndash; agents collecting and processing data at the edge, with no mandatory centralized storage &amp;ndash; compares to the centralized architectures they are used to.&lt;/p></description></item><item><title>Netdata at Open Source Observability Day 2024</title><link>https://www.netdata.cloud/events/open-source-observability-day-2024/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/open-source-observability-day-2024/</guid><description>&lt;p>Costa Tsaousis spoke at Open Source Observability Day 2024 on October 24. The talk focused on the evolution of Netdata and the broader challenges facing open-source observability tooling.&lt;/p>
&lt;p>The tension Costa addressed is one that every open-source monitoring project faces: how do you keep things simple for a developer who just wants to monitor a few servers, while also scaling to organizations with thousands of nodes and complex compliance requirements? Netdata&amp;rsquo;s answer has been a distributed architecture &amp;ndash; agents at the edge doing the heavy lifting, with optional cloud coordination &amp;ndash; but that comes with its own set of challenges around consistency, aggregation, and user experience.&lt;/p></description></item><item><title>Netdata at OpenConf 2025</title><link>https://www.netdata.cloud/events/openconf-2025/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/openconf-2025/</guid><description>&lt;p>Costa Tsaousis spoke at OpenConf 2025, held November 21&amp;ndash;22 at Dais Events in Athens. The talk, &amp;ldquo;Practical AI and Machine Learning for Observability in Netdata,&amp;rdquo; presented a specific framing of anomaly detection that Costa has been developing across multiple conferences: ML as an advisor, not just an alert trigger.&lt;/p>
&lt;p>The distinction matters. Most ML-in-monitoring implementations boil down to &amp;ldquo;replace static thresholds with dynamic ones.&amp;rdquo; Netdata&amp;rsquo;s approach is different. Multiple independent ML models run on each node, each trained on a single metric&amp;rsquo;s behavior. When one model flags an anomaly, that is information but not necessarily action. When dozens of models across multiple services flag anomalies simultaneously, that convergence is a strong signal. The system acts as an advisor &amp;ndash; surfacing unusual patterns, predicting potential failures, and detecting early signs of security breaches &amp;ndash; rather than firing off yet another alert.&lt;/p></description></item><item><title>Netdata at QBITS 2025</title><link>https://www.netdata.cloud/events/qbits-2025/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/qbits-2025/</guid><description>&lt;p>Netdata sponsored QBITS 2025 in Montreal, April 8&amp;ndash;10. Shyam Sreevalsan (VP - Product &amp;amp; Strategy) and Stuart McMurran (Enterprise Sales) represented the team at this Quadbridge event, which brings together IT leaders from across Canada and the US.&lt;/p>
&lt;p>The audience was heavily tilted toward infrastructure decision-makers &amp;ndash; CTOs, VPs of IT, directors of operations. These are people who sign off on monitoring tool purchases and live with the consequences. Three pain points came up in nearly every conversation at the booth: cost (observability bills that scale unpredictably with data volume), complexity (too many tools, too many dashboards, too much configuration), and manual troubleshooting (spending hours correlating data across systems during an incident).&lt;/p></description></item><item><title>Netdata at SREcon 2023</title><link>https://www.netdata.cloud/events/srecon-2023/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/srecon-2023/</guid><description>&lt;p>Costa Tsaousis, Netdata&amp;rsquo;s founder and CEO, joined a panel at SREcon23 Americas on October 11, 2023 in San Francisco. The session, &amp;ldquo;Open-source Development as a Full-time Pursuit,&amp;rdquo; ran from 16:50 to 17:30 and brought together maintainers who build open-source infrastructure tooling as their day job &amp;ndash; not as a side project.&lt;/p>
&lt;p>The panel dug into the realities of sustaining an open-source project when it is also the foundation of a company. Topics included funding models, the tension between community contributions and product roadmap, and the practical challenge of keeping a project healthy when you have both volunteer contributors and a paid engineering team pulling in potentially different directions.&lt;/p></description></item><item><title>Netdata at SREday London 2024</title><link>https://www.netdata.cloud/events/sreday-london-2024/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/sreday-london-2024/</guid><description>&lt;p>Costa Tsaousis spoke at SREday London 2024 on September 19&amp;ndash;20 with &amp;ldquo;Practical AI with Machine Learning for Observability in Netdata.&amp;rdquo; The event ran a promo code &amp;ndash; COSTA10 &amp;ndash; for 10% off tickets, which brought some extra traffic our way.&lt;/p>
&lt;p>The talk was tailored for an SRE audience. These are people who live in dashboards, write alert rules, and get paged at night. They are skeptical of ML claims because they have seen too many tools that promise &amp;ldquo;intelligent alerting&amp;rdquo; and deliver more noise. Costa focused on the mechanics: how Netdata trains unsupervised models per metric at the edge, why anomaly convergence across metrics matters more than any single anomaly score, and how this translates to fewer false positives in practice.&lt;/p></description></item><item><title>Netdata at stackconf 2024</title><link>https://www.netdata.cloud/events/stackconf-2024/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/stackconf-2024/</guid><description>&lt;p>Costa Tsaousis spoke at stackconf 2024 in Berlin on June 19, in a 14:30&amp;ndash;15:00 slot. The talk traced the history of Netdata from its inception: the original goals, the architectural bets, and what held up over time.&lt;/p>
&lt;p>The starting point was straightforward. When Costa began building Netdata, the goal was high-resolution metrics &amp;ndash; per-second granularity, not the 10- or 60-second averages that were standard at the time. That required a fundamentally different collection architecture: lightweight agents that process data locally, real-time visualization that does not depend on a query round-trip to a central database, and auto-detection of services so that adding a new node does not require writing configuration files.&lt;/p></description></item><item><title>Netdata at Tech Show London 2026</title><link>https://www.netdata.cloud/events/techshow-london-2026/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/techshow-london-2026/</guid><description>&lt;p>Netdata is exhibiting at Tech Show London 2026, March 4&amp;ndash;5 at ExCeL London. We are at Booth F223 in the Cloud &amp;amp; AI Infrastructure zone. Constantine Nikitiadis and Shyam Sreevalsan are running the booth, with 1:1 meetings available for anyone who wants dedicated time.&lt;/p>
&lt;p>The booth features live monitoring demos comparing per-second granularity against the 10&amp;ndash;60 second averages that most tools deliver. For teams running AI and cloud infrastructure, that difference is not academic &amp;ndash; GPU utilization spikes, model training anomalies, and container scheduling events happen on sub-second timescales that traditional monitoring simply misses. We are also showing Netdata&amp;rsquo;s AI-powered infrastructure observability: how unsupervised ML at the edge catches anomalies across your stack without requiring manual threshold configuration.&lt;/p></description></item><item><title>Netdata at WeAreDevelopers World Congress 2024</title><link>https://www.netdata.cloud/events/wearedevelopers-2024/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/wearedevelopers-2024/</guid><description>&lt;p>Netdata was at WeAreDevelopers World Congress 2024 in Berlin, July 17&amp;ndash;19. We had a booth in Hall A at stand A_S10, and Costa Tsaousis gave a talk on &amp;ldquo;Practical AI with Machine Learning in Observability.&amp;rdquo;&lt;/p>
&lt;p>The booth was packed for most of the event. We ran live demos of Netdata&amp;rsquo;s real-time dashboards, had a spin-the-wheel game, and did a prize draw. The wheel brought people in; the demos kept them. Developers who stopped by expecting a quick spin ended up watching per-second metrics streaming across a live infrastructure and asking how the anomaly detection works under the hood.&lt;/p></description></item><item><title>Netdata at WeAreDevelopers World Congress 2025</title><link>https://www.netdata.cloud/events/wearedevelopers-2025/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/events/wearedevelopers-2025/</guid><description>&lt;p>Netdata returned to WeAreDevelopers World Congress in Berlin as a partner for the 2025 edition, July 9&amp;ndash;11. Costa Tsaousis gave &amp;ldquo;Practical AI with Machine Learning in Observability&amp;rdquo; &amp;ndash; a talk he has been refining throughout 2024 and into 2025, each time sharpening the examples and responding to questions from previous audiences.&lt;/p>
&lt;p>The booth was busy all three days. WeAreDevelopers draws tens of thousands of developers, and the crowd skews younger and more curious than a typical infrastructure conference. Many visitors had not thought deeply about monitoring before &amp;ndash; they knew they needed it, but had not compared tools or architectures. Live demos of Netdata running on real infrastructure, showing per-second metrics with zero configuration, gave them a baseline to compare against whatever they try next.&lt;/p></description></item><item><title>Netdata Enterprise Agent</title><link>https://www.netdata.cloud/secure-foss-agent/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/secure-foss-agent/</guid><description/></item><item><title>Netdata Mobile App</title><link>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/netdata-mobile-app/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/netdata-mobile-app/</guid><description/></item><item><title>Netdata Pricing: Free Up to 5 Nodes | From $4.50/node</title><link>https://www.netdata.cloud/pricing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/pricing/</guid><description/></item><item><title>Netdata Streaming Topology</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/netdata-streaming-topology/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/netdata-streaming-topology/</guid><description/></item><item><title>Netdata usage survey</title><link>https://www.netdata.cloud/netdata-usage-survey/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/netdata-usage-survey/</guid><description/></item><item><title>Netdata vs Chronosphere | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/chronosphere/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/chronosphere/</guid><description>Netdata provides edge-native observability with per-second granularity, automatic ML anomaly detection, and transparent pricing—eliminating PromQL learning curves and SaaS-only limitations that challenge Chronosphere users.</description></item><item><title>Netdata vs KloudFuse | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/kloudfuse/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/kloudfuse/</guid><description>Discover why Netdata delivers instant infrastructure visibility with per-second granularity and transparent pricing, while KloudFuse requires days of setup and quote-based costs.</description></item><item><title>Netdata vs Microsoft SCOM | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/msscom/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/msscom/</guid><description/></item><item><title>Netdata vs Site24x7 | Monitoring Tools Comparison</title><link>https://www.netdata.cloud/comparisons/site24x7/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/comparisons/site24x7/</guid><description/></item><item><title>Netdata: Monitoring and troubleshooting transformed</title><link>https://www.netdata.cloud/partnerships/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/partnerships/</guid><description/></item><item><title>Netfilter</title><link>https://www.netdata.cloud/integrations/data-collection/networking/netfilter/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/netfilter/</guid><description/></item><item><title>NetFlow</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/flow-protocols/netflow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/flow-protocols/netflow/</guid><description/></item><item><title>NetFlow storage sizing: how much disk your flow collector really needs</title><link>https://www.netdata.cloud/guides/network/network-flow-collector-disk-sizing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-flow-collector-disk-sizing/</guid><description>&lt;h1 id="netflow-storage-sizing-how-much-disk-your-flow-collector-really-needs">NetFlow storage sizing: how much disk your flow collector really needs&lt;/h1>
&lt;p>Flow records arrive at thousands to tens of thousands per second, and every record hits disk. The bottleneck is almost always disk throughput or capacity, not CPU.&lt;/p>
&lt;p>This article covers the math: raw record sizes, effective storage after columnar compression, the capacity formula with worked examples, IOPS considerations, and the operational pitfalls that make disks fill faster than the formula predicts. The guidance applies to NetFlow v5/v9, IPFIX, and sFlow collectors using ClickHouse or similar columnar backends.&lt;/p></description></item><item><title>NetFlow v9/IPFIX template desync: flows decoded wrong or dropped after a reboot</title><link>https://www.netdata.cloud/guides/network/network-netflow-template-desync/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-netflow-template-desync/</guid><description>&lt;h1 id="netflow-v9ipfix-template-desync-flows-decoded-wrong-or-dropped-after-a-reboot">NetFlow v9/IPFIX template desync: flows decoded wrong or dropped after a reboot&lt;/h1>
&lt;p>You rebooted a router or upgraded its firmware. Minutes later, your flow collector shows a gap or anomaly. The exporter is still sending data: UDP packet counters are nonzero and climbing. But decoded flow records are zero, suspiciously low, or the field values are shifted and garbled.&lt;/p>
&lt;p>This is NetFlow v9 or IPFIX template desync. The collector holds cached template definitions that no longer match what the exporter is sending. Until it receives and caches the correct templates, it either drops records silently or misinterprets the byte layout, producing garbage fields.&lt;/p></description></item><item><title>NetFlow vs sFlow vs IPFIX: what they measure and how each one fails</title><link>https://www.netdata.cloud/guides/network/network-netflow-vs-sflow-vs-ipfix/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-netflow-vs-sflow-vs-ipfix/</guid><description>&lt;h1 id="netflow-vs-sflow-vs-ipfix-what-they-measure-and-how-each-one-fails">NetFlow vs sFlow vs IPFIX: what they measure and how each one fails&lt;/h1>
&lt;p>Flow telemetry protocols are often lumped together as &amp;ldquo;flow data,&amp;rdquo; but they measure fundamentally different things. NetFlow and IPFIX build stateful flow records by tracking conversations in device memory. sFlow captures random packet samples without maintaining any flow state. This architectural split determines not only what you can see but how the data breaks when something goes wrong.&lt;/p></description></item><item><title>Netgear</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/netgear/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/netgear/</guid><description/></item><item><title>Netgear Access Point</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/netgear-access-point/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/netgear-access-point/</guid><description/></item><item><title>Netgear Readynas</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/netgear-readynas/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/netgear-readynas/</guid><description/></item><item><title>Netgear SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netgear-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netgear-snmp-traps/</guid><description/></item><item><title>Netgear Switch</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/netgear-switch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/netgear-switch/</guid><description/></item><item><title>Netline SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netline-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netline-snmp-traps/</guid><description/></item><item><title>Netpartner S R O SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netpartner-s-r-o-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netpartner-s-r-o-snmp-traps/</guid><description/></item><item><title>Netquest Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netquest-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netquest-corp-snmp-traps/</guid><description/></item><item><title>Netrake Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netrake-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netrake-corporation-snmp-traps/</guid><description/></item><item><title>Netreality Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netreality-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netreality-inc-snmp-traps/</guid><description/></item><item><title>Netscaler SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netscaler-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netscaler-snmp-traps/</guid><description/></item><item><title>Netscreen Technologies Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netscreen-technologies-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netscreen-technologies-inc-snmp-traps/</guid><description/></item><item><title>Netstar Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netstar-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/netstar-inc-snmp-traps/</guid><description/></item><item><title>Network Alchemy Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/network-alchemy-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/network-alchemy-inc-snmp-traps/</guid><description/></item><item><title>Network Appliance Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/network-appliance-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/network-appliance-corporation-snmp-traps/</guid><description/></item><item><title>Network Connections</title><link>https://www.netdata.cloud/integrations/data-collection/networking/network-connections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/network-connections/</guid><description/></item><item><title>Network interfaces</title><link>https://www.netdata.cloud/integrations/data-collection/networking/network-interfaces/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/network-interfaces/</guid><description/></item><item><title>Network monitoring checklist: the signals every production network needs</title><link>https://www.netdata.cloud/guides/network/network-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-monitoring-checklist/</guid><description>&lt;h1 id="network-monitoring-checklist-the-signals-every-production-network-needs">Network monitoring checklist: the signals every production network needs&lt;/h1>
&lt;p>This checklist covers the signals production networks need, organized by detection priority and mapped to maturity levels from survival to expert.&lt;/p>
&lt;p>An NPM stack is a federation of collectors, parsers, enrichment services, storage tiers, and an analytics core. Most production incidents are not &amp;ldquo;the network broke&amp;rdquo; but &amp;ldquo;a collector&amp;rsquo;s UDP buffer dropped packets,&amp;rdquo; &amp;ldquo;the NetFlow v9 template cache went stale after a device reboot,&amp;rdquo; or &amp;ldquo;the polling worker pool fell behind and now a healthy device looks down.&amp;rdquo; The checklist is organized to surface those failure modes, not just the top-level symptoms.&lt;/p></description></item><item><title>Network statistics</title><link>https://www.netdata.cloud/integrations/data-collection/networking/network-statistics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/network-statistics/</guid><description/></item><item><title>Network Subsystem</title><link>https://www.netdata.cloud/integrations/data-collection/networking/network-subsystem/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/network-subsystem/</guid><description/></item><item><title>Network Technologies Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/network-technologies-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/network-technologies-inc-snmp-traps/</guid><description/></item><item><title>Networth Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/networth-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/networth-inc-snmp-traps/</guid><description/></item><item><title>New Oak Communications Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/new-oak-communications-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/new-oak-communications-inc-snmp-traps/</guid><description/></item><item><title>New Relic</title><link>https://www.netdata.cloud/integrations/exporters/new-relic/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/new-relic/</guid><description/></item><item><title>Newbridge Networks Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/newbridge-networks-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/newbridge-networks-corporation-snmp-traps/</guid><description/></item><item><title>Newtec Cy SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/newtec-cy-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/newtec-cy-snmp-traps/</guid><description/></item><item><title>Nextcloud Monitoring</title><link>https://www.netdata.cloud/monitoring-101/nextcloud-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/nextcloud-monitoring/</guid><description>&lt;h2 id="nextcloud-monitoring">Nextcloud Monitoring&lt;/h2>
&lt;h3 id="what-is-nextcloud">What Is Nextcloud?&lt;/h3>
&lt;p>Nextcloud is a widely-used, self-hosted productivity platform that offers a suite of client-server software for creating and using file hosting services. It is designed to allow users to share and collaborate on documents, manage files, and streamline communication all within the cloud. As a cornerstone of cloud services and computing, ensuring the optimal performance of your Nextcloud servers is critical for scalability and efficient operations.&lt;/p></description></item><item><title>Nextcloud servers</title><link>https://www.netdata.cloud/integrations/data-collection/applications/nextcloud-servers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/nextcloud-servers/</guid><description/></item><item><title>NextDNS</title><link>https://www.netdata.cloud/integrations/data-collection/networking/nextdns/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/nextdns/</guid><description/></item><item><title>NextDNS Monitoring</title><link>https://www.netdata.cloud/monitoring-101/nextdns-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/nextdns-monitoring/</guid><description>&lt;h2 id="nextdns-monitoring">NextDNS Monitoring&lt;/h2>
&lt;h3 id="what-is-nextdns">What Is NextDNS?&lt;/h3>
&lt;p>NextDNS is a powerful DNS resolver and security platform, tailored to provide efficient DNS management and enhance security. It&amp;rsquo;s an essential tool for IT teams looking to optimize their network infrastructure by offering better privacy, security, and performance.&lt;/p>
&lt;h3 id="monitoring-nextdns-with-netdata">Monitoring NextDNS With Netdata&lt;/h3>
&lt;p>Monitor NextDNS effortlessly with Netdata, which employs an OpenMetrics (Prometheus) exporter for seamless integration. This enables you to gather crucial DNS resolver and security metrics, helping you manage your domain&amp;rsquo;s name system more efficiently. One of the standout features of Netdata is its ability to ingest data from any Prometheus exporter, providing automated dashboards and alerts without requiring a dedicated Prometheus server or Grafana setup.&lt;/p></description></item><item><title>Nextnet SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nextnet-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nextnet-snmp-traps/</guid><description/></item><item><title>NFS Client</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/nfs-client/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/nfs-client/</guid><description/></item><item><title>NFS Server</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/nfs-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/nfs-server/</guid><description/></item><item><title>NGINX</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/nginx/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/nginx/</guid><description/></item><item><title>NGINX $request_time vs $upstream_response_time: isolating where latency lives</title><link>https://www.netdata.cloud/guides/nginx/nginx-request-time-vs-upstream-response-time/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-request-time-vs-upstream-response-time/</guid><description>&lt;h1 id="nginx-request_time-vs-upstream_response_time-isolating-where-latency-lives">NGINX $request_time vs $upstream_response_time: isolating where latency lives&lt;/h1>
&lt;p>P95 $request_time doubles. The default assumption is upstream degradation, so you scale the backend, tune the database, and add instances. The latency barely moves. The bottleneck was a slow mobile client, proxy temp-file disk I/O, or a large request body on a lossy network. This is the most common nginx misdiagnosis.&lt;/p>
&lt;p>$request_time measures the full cycle from the first byte read from the client to the last byte sent to the client. It includes reading the request, waiting for the upstream, and writing the response. $upstream_response_time measures only the backend portion, from establishing the upstream connection to receiving the last byte of the response body. The gap between them is where client-side, network, and nginx-internal delays live. To avoid chasing phantom backend problems, log both variables and compare them.&lt;/p></description></item><item><title>nginx 413 Request Entity Too Large: client_max_body_size explained</title><link>https://www.netdata.cloud/guides/nginx/nginx-413-request-entity-too-large/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-413-request-entity-too-large/</guid><description>&lt;h1 id="nginx-413-request-entity-too-large-client_max_body_size-explained">nginx 413 Request Entity Too Large: client_max_body_size explained&lt;/h1>
&lt;p>A &lt;code>413 Request Entity Too Large&lt;/code> after adding &lt;code>client_max_body_size 50m&lt;/code> to &lt;code>nginx.conf&lt;/code> usually means a more specific context still overrides it, or the upstream application rejects the body after nginx accepts it. The directive applies to &lt;code>http&lt;/code>, &lt;code>server&lt;/code>, and &lt;code>location&lt;/code> blocks. Raising it at the edge only moves the failure deeper if the rest of the stack is not adjusted. This guide covers how the directive inherits, when nginx fires the 413, how &lt;code>proxy_request_buffering&lt;/code> changes the failure mode, and why you must verify every hop including the upstream application.&lt;/p></description></item><item><title>nginx 499 status code: why clients close connections before the response</title><link>https://www.netdata.cloud/guides/nginx/nginx-499-client-closed-connection/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-499-client-closed-connection/</guid><description>&lt;h1 id="nginx-499-status-code-why-clients-close-connections-before-the-response">nginx 499 status code: why clients close connections before the response&lt;/h1>
&lt;p>Status 499 in nginx access logs means the client closed the TCP connection before nginx finished responding. It is an nginx-specific code that never reaches the client, so it is easy to dismiss. In practice, a 499 surge is an early warning: users or intermediaries abandon requests before upstreams officially time out and before 5xx errors spike. Ignore 499s and you usually see 502s or 504s minutes later.&lt;/p></description></item><item><title>nginx 500 Internal Server Error: how to diagnose it</title><link>https://www.netdata.cloud/guides/nginx/nginx-500-internal-server-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-500-internal-server-error/</guid><description>&lt;h1 id="nginx-500-internal-server-error-how-to-diagnose-it">nginx 500 Internal Server Error: how to diagnose it&lt;/h1>
&lt;p>A 500 from nginx tells you something failed in the request path, but not whether the failure originated in your application, FastCGI/uWSGI backend, or nginx itself. When 500s spike during an incident, first determine which side of the nginx boundary is breaking.&lt;/p>
&lt;p>Unlike 502 Bad Gateway or 504 Gateway Time-out, which point upstream, a 500 can be an application bug passed through by nginx, a configuration error, a permission failure, or resource exhaustion inside an nginx worker.&lt;/p></description></item><item><title>NGINX 502 Bad Gateway: Causes And How To Fix It</title><link>https://www.netdata.cloud/guides/nginx/nginx-502-bad-gateway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-502-bad-gateway/</guid><description>&lt;p>A 502 Bad Gateway means the upstream server returned an invalid response, refused the connection, or terminated before completing the response. Unlike 504, which signals upstream slowness, 502 means the upstream never produced a valid response or nginx could not reach it.&lt;/p>
&lt;p>Start with the error log. A single line like &lt;code>connect() failed (111: Connection refused)&lt;/code> tells you the upstream is not listening. A line like &lt;code>upstream prematurely closed connection&lt;/code> tells you the backend died mid-request. Match the exact message to the root cause.&lt;/p></description></item><item><title>nginx 503 Service Temporarily Unavailable: causes and fixes</title><link>https://www.netdata.cloud/guides/nginx/nginx-503-service-unavailable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-503-service-unavailable/</guid><description>&lt;h1 id="nginx-503-service-temporarily-unavailable-causes-and-fixes">nginx 503 service temporarily unavailable: causes and fixes&lt;/h1>
&lt;p>A &lt;code>503 Service Temporarily Unavailable&lt;/code> from nginx does not always mean the upstream application is broken. A healthy nginx process returns 503 by design when rate limits reject traffic, or when every backend in an upstream block is unavailable. The same status code covers three different failure paths: intentional throttling, upstream exhaustion, or resource saturation on the nginx host. The fix depends on which path the request took.&lt;/p></description></item><item><title>NGINX 504 Gateway Time-Out: Causes &amp; Fixes</title><link>https://www.netdata.cloud/guides/nginx/nginx-504-gateway-timeout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-504-gateway-timeout/</guid><description>&lt;p>A 504 Gateway Time-out means nginx reached the upstream but the upstream did not finish its response before &lt;code>proxy_read_timeout&lt;/code> expired. The default is 60 seconds. Unlike a 502 Bad Gateway, which means nginx never established a valid upstream connection, a 504 means the connection succeeded but the response did not complete in time.&lt;/p>
&lt;p>This guide covers isolating slow upstreams via access log variables, tuning timeouts and retries, and distinguishing 504 from 502.&lt;/p></description></item><item><title>NGINX access log performance: buffering, sampling, and the event loop</title><link>https://www.netdata.cloud/guides/nginx/nginx-access-log-performance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-access-log-performance/</guid><description>&lt;h1 id="nginx-access-log-performance-buffering-sampling-and-the-event-loop">NGINX access log performance: buffering, sampling, and the event loop&lt;/h1>
&lt;p>Every NGINX worker is a single-threaded event loop. By default, finishing a request triggers an immediate, synchronous write of the access log line. Under normal load the cost is negligible. Under incident conditions, when error rates and log volume spike, that synchronous write becomes a compounding failure: disk I/O stalls the event loop, responses slow down, clients time out, and the resulting errors and abandonments generate even more log lines.&lt;/p></description></item><item><title>NGINX active connections climbing: reading, writing, waiting explained</title><link>https://www.netdata.cloud/guides/nginx/nginx-active-connections-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-active-connections-high/</guid><description>&lt;h1 id="nginx-active-connections-climbing-reading-writing-waiting-explained">NGINX active connections climbing: reading, writing, waiting explained&lt;/h1>
&lt;p>When operators see Active connections climbing in &lt;code>stub_status&lt;/code>, the first instinct is often to add capacity. That instinct is usually wrong. The &lt;code>stub_status&lt;/code> module exposes exactly seven metrics, and the most useful of them is the breakdown of active connections into Reading, Writing, and Waiting. The absolute number of active connections is almost meaningless without the ratio between these three states. A server with 10,000 active connections where 9,000 are Waiting is healthy. A server with 500 active connections where 400 are Reading may be under a slowloris-style attack.&lt;/p></description></item><item><title>NGINX backend cascade failure: when slow upstreams take down everything</title><link>https://www.netdata.cloud/guides/nginx/nginx-backend-cascade-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-backend-cascade-failure/</guid><description>&lt;h1 id="nginx-backend-cascade-failure-when-slow-upstreams-take-down-everything">NGINX backend cascade failure: when slow upstreams take down everything&lt;/h1>
&lt;p>Users report timeouts. 502 Bad Gateway and 504 Gateway Time-out responses are climbing, and nginx error logs show upstream timeouts. On the nginx host, CPU and memory are normal, and the master process is alive. The proxy is healthy but out of connections.&lt;/p>
&lt;p>This is a backend cascade failure. One slow upstream causes nginx workers to hold connections open while waiting for responses, consuming finite &lt;code>worker_connections&lt;/code> slots. As slots fill, new requests cannot be forwarded. Traffic concentrates on the remaining healthy backends, which overload and slow down. Eventually every backend times out or fails health checks, and nginx returns 502/504 to all clients while the proxy process remains up.&lt;/p></description></item><item><title>nginx connect() failed (111: Connection refused) while connecting to upstream</title><link>https://www.netdata.cloud/guides/nginx/nginx-connect-failed-connection-refused-upstream/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-connect-failed-connection-refused-upstream/</guid><description>&lt;h1 id="nginx-connect-failed-111-connection-refused-while-connecting-to-upstream">nginx connect() failed (111: Connection refused) while connecting to upstream&lt;/h1>
&lt;p>HTTP 502 Bad Gateway and the error &lt;code>connect() failed (111: Connection refused) while connecting to upstream&lt;/code> mean nginx reached the upstream IP, but the target port actively refused the TCP connection. The backend is either not running, not listening on the interface nginx expects, or a firewall is blocking the port.&lt;/p>
&lt;p>This is distinct from &lt;code>upstream timed out (110: Connection timed out)&lt;/code>. A timeout means the TCP SYN never received a response, usually because a firewall silently dropped the packet or the host is unreachable. Errno 111 means the network path is open but no process is accepting connections. The error log includes the upstream address, such as &lt;code>upstream: &amp;quot;fastcgi://127.0.0.1:9000&amp;quot;&lt;/code>. Read that line first to isolate the exact backend.&lt;/p></description></item><item><title>NGINX connection exhaustion: detection, diagnosis, and prevention</title><link>https://www.netdata.cloud/guides/nginx/nginx-connection-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-connection-exhaustion/</guid><description>&lt;h1 id="nginx-connection-exhaustion-detection-diagnosis-and-prevention">NGINX connection exhaustion: detection, diagnosis, and prevention&lt;/h1>
&lt;p>Users see connection timeouts while load balancer health checks and the NGINX &lt;code>stub_status&lt;/code> endpoint still return HTTP 200. New connections are silently dropped. Connection exhaustion is a cliff-edge failure: once the limit is hit, there is no graceful degradation. Connections are refused at the kernel level, or accepted into the TCP backlog but discarded by NGINX because no worker has a free connection slot.&lt;/p></description></item><item><title>NGINX DNS resolution failures on dynamic upstreams: 502s and resolver_timeout</title><link>https://www.netdata.cloud/guides/nginx/nginx-dns-resolution-failures-upstream/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-dns-resolution-failures-upstream/</guid><description>&lt;h1 id="nginx-dns-resolution-failures-on-dynamic-upstreams-502s-and-resolver_timeout">NGINX DNS resolution failures on dynamic upstreams: 502s and resolver_timeout&lt;/h1>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>Intermittent 502 Bad Gateway responses that only hit locations using a variable in &lt;code>proxy_pass&lt;/code>, such as &lt;code>proxy_pass http://$backend;&lt;/code>, point to dynamic DNS resolution failure. Static upstream locations are unaffected. &lt;code>dig&lt;/code> from the host may succeed instantly while nginx logs show 502s with latency spikes clustering at exactly 30 seconds, the default &lt;code>resolver_timeout&lt;/code>.&lt;/p>
&lt;p>nginx resolves upstream hostnames through two paths.&lt;/p></description></item><item><title>NGINX dropped connections: the accepts vs handled gap</title><link>https://www.netdata.cloud/guides/nginx/nginx-dropped-connections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-dropped-connections/</guid><description>&lt;h1 id="nginx-dropped-connections-the-accepts-vs-handled-gap">NGINX dropped connections: the accepts vs handled gap&lt;/h1>
&lt;p>Users report intermittent connection timeouts. Your HTTP 5xx rate is flat. The error log is quiet. Something is dropping traffic before it ever becomes a request.&lt;/p>
&lt;p>On every NGINX instance, the &lt;code>stub_status&lt;/code> page exposes two cumulative counters: &lt;code>accepts&lt;/code> and &lt;code>handled&lt;/code>. When &lt;code>accepts&lt;/code> grows faster than &lt;code>handled&lt;/code>, NGINX is taking connections from the kernel and then discarding them. This gap is a leading indicator of connection-slot or file-descriptor exhaustion. It often starts increasing minutes before the system hits the hard wall.&lt;/p></description></item><item><title>NGINX limit_req burst and nodelay tuning: rate limiting without blocking real users</title><link>https://www.netdata.cloud/guides/nginx/nginx-limit-req-burst-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-limit-req-burst-tuning/</guid><description>&lt;h1 id="nginx-limit_req-burst-and-nodelay-tuning-rate-limiting-without-blocking-real-users">NGINX limit_req burst and nodelay tuning: rate limiting without blocking real users&lt;/h1>
&lt;p>Most production nginx rate limiting configs fall into two camps: no burst at all, which rejects legitimate traffic during harmless spikes, or burst without nodelay, which queues real users into artificial delays that mimic upstream slowness. Neither is what you want. The &lt;code>limit_req&lt;/code> module implements a leaky bucket at millisecond granularity, and the interaction between &lt;code>burst&lt;/code> and &lt;code>nodelay&lt;/code> determines whether a request is delayed, rejected, or forwarded immediately. Understanding that interaction, and sizing the shared memory zone to match your traffic profile, is the difference between rate limiting that protects upstreams and rate limiting that creates incidents during normal user behavior.&lt;/p></description></item><item><title>nginx limiting requests, excess -- understanding limit_req rejections</title><link>https://www.netdata.cloud/guides/nginx/nginx-limiting-requests-excess/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-limiting-requests-excess/</guid><description>&lt;h1 id="nginx-limiting-requests-excess----understanding-limit_req-rejections">nginx limiting requests, excess &amp;ndash; understanding limit_req rejections&lt;/h1>
&lt;p>When &lt;code>[error] ... limiting requests, excess&lt;/code> appears in nginx error logs alongside 503 responses in access logs, determine whether you are under attack, misconfigured, or out of shared memory. The &lt;code>ngx_http_limit_req_module&lt;/code> implements a leaky bucket rate limiter. Its interaction with &lt;code>burst&lt;/code>, &lt;code>nodelay&lt;/code>, and shared memory sizing determines whether you reject malicious traffic, delay legitimate users, or silently stop enforcing limits.&lt;/p>
&lt;h2 id="what-it-is-and-why-it-matters">What it is and why it matters&lt;/h2>
&lt;p>&lt;code>limit_req&lt;/code> is nginx&amp;rsquo;s request-level rate limiter. It uses a shared memory zone, configured via &lt;code>limit_req_zone&lt;/code>, to track request rates per key, typically &lt;code>$binary_remote_addr&lt;/code>. The zone is mapped into every worker process. When a request arrives, nginx checks the key&amp;rsquo;s current rate against the configured limit. Depending on &lt;code>burst&lt;/code> and &lt;code>nodelay&lt;/code>, it delays the request, rejects it, or processes it immediately.&lt;/p></description></item><item><title>NGINX listen queue overflow: somaxconn, backlog, and silent connection drops</title><link>https://www.netdata.cloud/guides/nginx/nginx-listen-queue-overflow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-listen-queue-overflow/</guid><description>&lt;h1 id="nginx-listen-queue-overflow-somaxconn-backlog-and-silent-connection-drops">NGINX listen queue overflow: somaxconn, backlog, and silent connection drops&lt;/h1>
&lt;p>Clients report intermittent connection timeouts. Your load balancer health checks pass. NGINX error logs are clean and access logs show no 5xx spikes. The issue is not in NGINX workers or upstream applications. It is in the kernel accept queue.&lt;/p>
&lt;p>When the accept queue fills, the kernel drops new connections silently. NGINX never sees them, so it logs nothing. Evidence is client-side timeouts and the kernel counter &lt;code>TcpExtListenOverflows&lt;/code>.&lt;/p></description></item><item><title>NGINX log disk full: when logging silently stops and how to recover</title><link>https://www.netdata.cloud/guides/nginx/nginx-log-disk-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-log-disk-full/</guid><description>&lt;h1 id="nginx-log-disk-full-when-logging-silently-stops-and-how-to-recover">NGINX log disk full: when logging silently stops and how to recover&lt;/h1>
&lt;p>During an incident, upstream latency climbs and 5xx errors appear. You run &lt;code>tail -f /var/log/nginx/error.log&lt;/code> and the cursor sits there. The last entry is hours old. Requests still return 200s. The server is serving, but it has stopped logging. Disk utilization shows &lt;code>/var/log&lt;/code> at 100 percent.&lt;/p>
&lt;p>This is the NGINX log disk-full failure mode. The process never signals the client or the operator that it can no longer write diagnostics. Visibility evaporates when it is most needed.&lt;/p></description></item><item><title>NGINX Monitoring</title><link>https://www.netdata.cloud/monitoring-101/nginx-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/nginx-monitoring/</guid><description>&lt;h2 id="nginx-monitoring">NGINX Monitoring&lt;/h2>
&lt;h3 id="what-is-nginx">What Is NGINX?&lt;/h3>
&lt;p>&lt;a href="https://www.nginx.com/">NGINX&lt;/a> is a high-performance web server, reverse proxy server, and load balancer designed to handle a large number of concurrent connections efficiently. Its modular architecture allows it to be extended with additional features, making it a versatile component in modern application infrastructures.&lt;/p>
&lt;h3 id="monitoring-nginx-with-netdata">Monitoring NGINX With Netdata&lt;/h3>
&lt;p>When it comes to monitoring NGINX, Netdata offers an intuitive and real-time monitoring solution. With &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/nginx/">Netdata&amp;rsquo;s NGINX monitoring tool&lt;/a>, you can observe your server&amp;rsquo;s metrics in real time, use interactive charts, and detect any anomalies swiftly.&lt;/p></description></item><item><title>NGINX monitoring checklist: the signals every production server needs</title><link>https://www.netdata.cloud/guides/nginx/nginx-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-monitoring-checklist/</guid><description>&lt;h1 id="nginx-monitoring-checklist-the-signals-every-production-server-needs">NGINX monitoring checklist: the signals every production server needs&lt;/h1>
&lt;p>NGINX is an event-driven, single-threaded-per-worker process. Most production failures follow predictable patterns: connection exhaustion, backend cascades, file descriptor limits, or silent kernel-level drops. This article maps the signals that expose those failures into four cumulative maturity levels: Survival, Operational, Mature, and Expert. Use it to audit your current coverage or to justify instrumentation work before the next incident.&lt;/p>
&lt;p>Each level adds depth. Survival answers &amp;ldquo;Is it up?&amp;rdquo; Operational answers &amp;ldquo;Is it healthy?&amp;rdquo; Mature adds leading indicators. Expert adds the signals you instrument after your third postmortem. The tables below list each signal, why it matters, and the threshold that should trigger a response.&lt;/p></description></item><item><title>NGINX monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/nginx/nginx-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-monitoring-maturity-model/</guid><description>&lt;h1 id="nginx-monitoring-maturity-model-from-survival-to-expert">NGINX monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>nginx exposes exactly seven scalars through &lt;code>stub_status&lt;/code>. Latency distributions, upstream health, cache efficiency, and kernel-level drops live in access logs, error logs, or OS counters. Teams that collect only the stub_status numbers assume they have visibility. They do not.&lt;/p>
&lt;p>This article defines four monitoring maturity levels. Level 1 tells you if nginx is alive. Level 2 tells you if it is healthy. Level 3 gives you leading indicators of saturation. Level 4 exposes the blind spots that only appear after repeated incidents. Use these levels to audit your current coverage and decide which signals to add next.&lt;/p></description></item><item><title>nginx no live upstreams while connecting to upstream: what it means</title><link>https://www.netdata.cloud/guides/nginx/nginx-no-live-upstreams/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-no-live-upstreams/</guid><description>&lt;h1 id="nginx-no-live-upstreams-while-connecting-to-upstream-what-it-means">nginx no live upstreams while connecting to upstream: what it means&lt;/h1>
&lt;p>When nginx logs &lt;code>no live upstreams while connecting to upstream&lt;/code>, every server in the affected upstream block is marked unavailable. The proxied request has no eligible backend, so nginx returns 502 Bad Gateway. This is not an nginx defect; it signals that all backends have failed open-source nginx&amp;rsquo;s passive health checks, or a network partition has made them unreachable from the nginx host.&lt;/p></description></item><item><title>NGINX old worker processes accumulating after reload: worker_shutdown_timeout</title><link>https://www.netdata.cloud/guides/nginx/nginx-old-worker-processes-accumulating/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-old-worker-processes-accumulating/</guid><description>&lt;h1 id="nginx-old-worker-processes-accumulating-after-reload-worker_shutdown_timeout">NGINX old worker processes accumulating after reload: worker_shutdown_timeout&lt;/h1>
&lt;p>After &lt;code>nginx -s reload&lt;/code>, worker process count climbs above &lt;code>worker_processes&lt;/code>. Memory and file descriptor usage grow. The workers are not crashing; they are old generations waiting for long-lived connections to close.&lt;/p>
&lt;p>This pattern is normal in small doses. During a reload, the master keeps old workers alive until active connections drain. Without a shutdown deadline, a single WebSocket, gRPC stream, or long-polling connection can pin an old worker indefinitely. In environments that reload frequently, such as Kubernetes ingress controllers reacting to endpoint changes, accumulation becomes a resource leak that can exhaust memory or file descriptors.&lt;/p></description></item><item><title>NGINX Plus</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/nginx-plus/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/nginx-plus/</guid><description/></item><item><title>NGINX Plus Monitoring</title><link>https://www.netdata.cloud/monitoring-101/nginxplus-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/nginxplus-monitoring/</guid><description>&lt;h2 id="nginx-plus-monitoring">NGINX Plus Monitoring&lt;/h2>
&lt;h3 id="what-is-nginx-plus">What Is NGINX Plus?&lt;/h3>
&lt;p>&lt;a href="https://www.nginx.com/products/nginx/">NGINX Plus&lt;/a> is a premium version of NGINX, offering additional features such as advanced load balancing, reliability, security, and flexibility to deploy applications. It is widely used for web serving, reverse proxying, caching, load balancing, media streaming, and more.&lt;/p>
&lt;h3 id="monitoring-nginx-plus-with-netdata">Monitoring NGINX Plus With Netdata&lt;/h3>
&lt;p>Netdata provides a robust NGINX Plus monitoring tool that gives real-time insight into the performance and health of your NGINX Plus servers. With the ability to visualize key metrics and diagnose performance issues swiftly, Netdata becomes an invaluable tool for any DevOps, SRE, or IT professional managing NGINX Plus instances. To learn more, see the &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/nginxplus/?utm_source=website&amp;amp;utm_content=monitoring101">NGINX Plus collector documentation&lt;/a>.&lt;/p></description></item><item><title>NGINX proxy buffer spill to disk: proxy_buffers and temp file latency</title><link>https://www.netdata.cloud/guides/nginx/nginx-proxy-buffer-spill-to-disk/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-proxy-buffer-spill-to-disk/</guid><description>&lt;h1 id="nginx-proxy-buffer-spill-to-disk-proxy_buffers-and-temp-file-latency">NGINX proxy buffer spill to disk: proxy_buffers and temp file latency&lt;/h1>
&lt;p>You notice some proxied requests are crawling. Upstream logs show sub-50 ms response times. The network path is clean. The nginx error log is quiet. In the access log, however, &lt;code>$request_time&lt;/code> is ten times larger than &lt;code>$upstream_response_time&lt;/code>. For large responses, this gap is the signature of proxy buffer spill.&lt;/p>
&lt;p>When an upstream response exceeds the memory buffers allocated by &lt;code>proxy_buffers&lt;/code>, nginx writes the overflow to a temporary file under &lt;code>proxy_temp_path&lt;/code> and reads it back later. Because the log message for this event is emitted at debug level only, the delay is silent. Standard upstream monitoring gives no hint; the delay hides entirely inside nginx.&lt;/p></description></item><item><title>NGINX proxy buffer tuning: proxy_buffers, proxy_buffer_size, and busy buffers</title><link>https://www.netdata.cloud/guides/nginx/nginx-proxy-buffer-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-proxy-buffer-tuning/</guid><description>&lt;h1 id="nginx-proxy-buffer-tuning-proxy_buffers-proxy_buffer_size-and-busy-buffers">NGINX proxy buffer tuning: proxy_buffers, proxy_buffer_size, and busy buffers&lt;/h1>
&lt;p>When &lt;code>proxy_buffering&lt;/code> is &lt;code>on&lt;/code> (the default), NGINX absorbs the upstream response in memory before sending it to the client. This shields backends from slow clients and enables compression, but three directives control it: &lt;code>proxy_buffer_size&lt;/code> for headers, &lt;code>proxy_buffers&lt;/code> for the body, and &lt;code>proxy_busy_buffers_size&lt;/code> for the in-flight flush window. Misconfiguration causes 502s, silent disk spills, and reload failures.&lt;/p>
&lt;p>The defaults are modest: eight body buffers of one memory page each, and one header page (typically 4K or 8K). That works for static sites and small JSON, but it fails for modern workloads: APIs with large JWT tokens in headers, bulk exports returning multi-megabyte JSON, and Server-Sent Events streams. Undersized body buffers spill to disk. Undersized header buffers return 502. Invalid &lt;code>proxy_busy_buffers_size&lt;/code> values prevent NGINX from starting or reloading.&lt;/p></description></item><item><title>NGINX proxy cache hit rate is low: measuring and improving it</title><link>https://www.netdata.cloud/guides/nginx/nginx-cache-hit-rate-low/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-cache-hit-rate-low/</guid><description>&lt;h1 id="nginx-proxy-cache-hit-rate-is-low-measuring-and-improving-it">NGINX proxy cache hit rate is low: measuring and improving it&lt;/h1>
&lt;p>Low NGINX proxy cache hit rate shows up as upstream CPU climbing or origin traffic higher than expected. Access logs show MISS and BYPASS where you expect HIT. When caching fails, every request reaches the backend, adding latency and load. The symptom is usually a gradual slide from 85% to 40% over a day, or a collapse to zero after a deployment or restart. Root cause is often configuration drift: a new header, a changed query parameter, or a keys_zone sized for last year&amp;rsquo;s traffic. Diagnose by measuring which cache status dominates, then trace that status back to the directive or upstream behavior that produces it.&lt;/p></description></item><item><title>NGINX proxy_cache not caching: why responses bypass the cache</title><link>https://www.netdata.cloud/guides/nginx/nginx-proxy-cache-not-caching/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-proxy-cache-not-caching/</guid><description>&lt;h1 id="nginx-proxy_cache-not-caching-why-responses-bypass-the-cache">NGINX proxy_cache not caching: why responses bypass the cache&lt;/h1>
&lt;p>After enabling &lt;code>proxy_cache&lt;/code> and defining the cache path, upstreams still take every hit. Access logs show &lt;code>$upstream_cache_status&lt;/code> as &lt;code>BYPASS&lt;/code> or &lt;code>MISS&lt;/code>, hit rate stays near zero, and nothing appears in the error log. NGINX applies a strict request-phase and response-phase decision tree before anything enters the cache. If any condition fails, the response is never stored. The symptom looks like upstream overload, but the root cause is usually a directive, a header, or a missing validity window.&lt;/p></description></item><item><title>NGINX rate limiting returns 503 not 429: limit_req_status explained</title><link>https://www.netdata.cloud/guides/nginx/nginx-rate-limiting-503-not-429/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-rate-limiting-503-not-429/</guid><description>&lt;h1 id="nginx-rate-limiting-returns-503-not-429-limit_req_status-explained">NGINX rate limiting returns 503 not 429: limit_req_status explained&lt;/h1>
&lt;p>When &lt;code>limit_req&lt;/code> rejects a request, nginx returns 503 by default. This is the same code used for genuine capacity exhaustion, so a 503 spike pages the on-call rotation even when the upstream is healthy and the infrastructure is fine. Changing &lt;code>limit_req_status&lt;/code> to 429 is one directive, but the implications for alerting, monitoring, and client behavior are not trivial.&lt;/p>
&lt;h2 id="what-it-is-and-why-it-matters">What it is and why it matters&lt;/h2>
&lt;p>The &lt;code>limit_req_status&lt;/code> directive sets the HTTP status code returned when &lt;code>limit_req&lt;/code> rejects a request. The companion directive &lt;code>limit_conn_status&lt;/code> does the same for connection limits enforced by &lt;code>limit_conn&lt;/code>. Both default to 503.&lt;/p></description></item><item><title>nginx recv() failed (104: Connection reset by peer) while reading from upstream</title><link>https://www.netdata.cloud/guides/nginx/nginx-connection-reset-by-peer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-connection-reset-by-peer/</guid><description>&lt;h1 id="nginx-recv-failed-104-connection-reset-by-peer-while-reading-from-upstream">nginx recv() failed (104: Connection reset by peer) while reading from upstream&lt;/h1>
&lt;p>You tail the nginx error log and see this:&lt;/p>
&lt;pre tabindex="0">&lt;code>[error] ... recv() failed (104: Connection reset by peer) while reading response header from upstream
&lt;/code>&lt;/pre>&lt;p>The client gets a 502 Bad Gateway. The error is not a configuration syntax problem, and the service was working five minutes ago. This message means the upstream server sent a TCP RST while nginx was mid-read on an upstream connection. The reset originates from the backend or from a network middlebox between nginx and the backend. It does not come from nginx itself, and it does not come from the client.&lt;/p></description></item><item><title>NGINX reload not applying config: why old workers keep serving</title><link>https://www.netdata.cloud/guides/nginx/nginx-reload-not-applying-config/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-reload-not-applying-config/</guid><description>&lt;h1 id="nginx-reload-not-applying-config-why-old-workers-keep-serving">NGINX reload not applying config: why old workers keep serving&lt;/h1>
&lt;p>You pushed a config change, ran &lt;code>nginx -s reload&lt;/code>, and moved on. Hours later, the new certificate is not being served, the updated upstream is not receiving traffic, or the tightened rate limit never took effect. NGINX did not stop running, but the reload never applied. This is the silent rollback: when a reload fails validation, the master keeps the previous configuration active and old workers continue serving. Even when validation passes, old workers can remain alive for hours if long-lived connections prevent them from draining and &lt;code>worker_shutdown_timeout&lt;/code> is not set. This guide shows how to confirm the failure, find the root cause, and prevent config drift from going undetected.&lt;/p></description></item><item><title>NGINX slow requests: from access log to root cause</title><link>https://www.netdata.cloud/guides/nginx/nginx-slow-requests/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-slow-requests/</guid><description>&lt;h1 id="nginx-slow-requests-from-access-log-to-root-cause">NGINX slow requests: from access log to root cause&lt;/h1>
&lt;p>Elevated &lt;code>$request_time&lt;/code> in access logs does not mean the upstream is slow. The variable measures the full lifecycle: from reading the first client byte through sending the last response byte. That includes client upload, upstream wait, nginx processing, and client download. Blaming the backend by reflex is the most common nginx latency mistake.&lt;/p>
&lt;p>To split the time accurately, confirm your &lt;code>log_format&lt;/code> includes &lt;code>$request_time&lt;/code>, &lt;code>$upstream_response_time&lt;/code>, &lt;code>$upstream_connect_time&lt;/code>, and &lt;code>$upstream_header_time&lt;/code>.&lt;/p></description></item><item><title>NGINX slowloris and slow-client attacks: detection and mitigation</title><link>https://www.netdata.cloud/guides/nginx/nginx-slowloris-attack/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-slowloris-attack/</guid><description>&lt;h1 id="nginx-slowloris-and-slow-client-attacks-detection-and-mitigation">NGINX slowloris and slow-client attacks: detection and mitigation&lt;/h1>
&lt;p>You check &lt;code>stub_status&lt;/code> and the Reading count is triple its normal value and not dropping. Requests per second has collapsed to near zero. Active connections are climbing toward &lt;code>worker_connections * worker_processes&lt;/code> while worker CPU stays idle. This is not a backend slowdown. It is a slowloris or slow-client attack: connections open faster than they complete, and NGINX waits for data that arrives one byte at a time.&lt;/p></description></item><item><title>NGINX SSL certificate expired: detection and emergency renewal</title><link>https://www.netdata.cloud/guides/nginx/nginx-ssl-certificate-expired/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-ssl-certificate-expired/</guid><description>&lt;h1 id="nginx-ssl-certificate-expired-detection-and-emergency-renewal">NGINX SSL certificate expired: detection and emergency renewal&lt;/h1>
&lt;p>An expired SSL certificate on NGINX is an immediate outage, not gradual degradation. Browsers and API clients reject the connection at the TLS handshake, often before NGINX logs anything useful. The fix is rarely as simple as running a renewal script again. Verify what is on disk, confirm the running configuration is using it, and force a reload so workers load the new certificate.&lt;/p></description></item><item><title>NGINX SSL session cache: improving TLS resumption and cutting CPU</title><link>https://www.netdata.cloud/guides/nginx/nginx-ssl-session-cache/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-ssl-session-cache/</guid><description>&lt;h1 id="nginx-ssl-session-cache-improving-tls-resumption-and-cutting-cpu">NGINX SSL session cache: improving TLS resumption and cutting CPU&lt;/h1>
&lt;p>TLS handshakes are the most CPU-intensive routine operation in nginx. A worker terminating RSA-2048 TLS can handle only a few hundred full handshakes per second, compared to tens of thousands of plain HTTP requests. In production, every reconnecting client that repeats a full handshake wastes CPU and adds latency. The &lt;code>ssl_session_cache&lt;/code> directive exists to eliminate that waste by allowing session resumption across connections. Yet many configurations either omit it, under-size it, or misunderstand how TLS 1.3 changes resumption behavior. This article explains the mechanism, sizing, and the signals that tell you whether your cache is working.&lt;/p></description></item><item><title>nginx SSL_do_handshake() failed — diagnosing TLS handshake errors</title><link>https://www.netdata.cloud/guides/nginx/nginx-ssl-handshake-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-ssl-handshake-failed/</guid><description>&lt;h1 id="nginx-ssl_do_handshake-failed--diagnosing-tls-handshake-errors">nginx SSL_do_handshake() failed — diagnosing TLS handshake errors&lt;/h1>
&lt;p>You see &lt;code>[crit]&lt;/code> or &lt;code>[error]&lt;/code> entries for &lt;code>SSL_do_handshake() failed&lt;/code> in &lt;code>/var/log/nginx/error.log&lt;/code>. The message might involve a client connecting to nginx, or nginx connecting to an upstream server. These errors mean the TLS negotiation never completed, so no application data was exchanged. The impact ranges from a few browsers showing security warnings to all proxied traffic returning 502 Bad Gateway.&lt;/p>
&lt;p>Before chasing certificates, determine which side of a connection is failing. nginx logs two variants: &lt;code>while SSL handshaking to client&lt;/code> and &lt;code>while SSL handshaking to upstream&lt;/code>. The first affects visitors hitting your server directly. The second affects nginx as a reverse proxy calling a backend over HTTPS. The symptoms, root causes, and fixes are different.&lt;/p></description></item><item><title>NGINX SSL/TLS handshake CPU saturation: detection and tuning</title><link>https://www.netdata.cloud/guides/nginx/nginx-ssl-handshake-cpu-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-ssl-handshake-cpu-saturation/</guid><description>&lt;h1 id="nginx-ssltls-handshake-cpu-saturation-detection-and-tuning">NGINX SSL/TLS handshake CPU saturation: detection and tuning&lt;/h1>
&lt;p>NGINX latency climbs while requests per second flatline. Worker processes are pinned near 100% CPU, yet active connections are nowhere near the &lt;code>worker_connections&lt;/code> limit. Access logs show fast upstream response times, but &lt;code>$request_time&lt;/code> is an order of magnitude larger. The bottleneck is not the network, disk, or backends: it is the TLS handshake.&lt;/p>
&lt;p>When workers burn CPU on cryptography, the single-threaded event loop has no time left for request processing. Every new TCP connection that requires a full SSL handshake adds asymmetric crypto workload. If clients do not resume sessions and the connection rate is high, throughput collapses even though the machine has plenty of idle connection slots. This guide shows how to detect, diagnose, and tune for SSL handshake CPU saturation.&lt;/p></description></item><item><title>NGINX Unit</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/nginx-unit/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/nginx-unit/</guid><description/></item><item><title>NGINX Unit Monitoring</title><link>https://www.netdata.cloud/monitoring-101/nginxunit-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/nginxunit-monitoring/</guid><description>&lt;h2 id="nginx-unit-monitoring">NGINX Unit Monitoring&lt;/h2>
&lt;p>NGINX Unit monitoring is crucial for ensuring the optimal performance and reliability of web applications. With the capabilities offered by the &lt;a href="https://unit.nginx.org/">NGINX Unit&lt;/a>, a dynamic application server, it&amp;rsquo;s possible to handle various configurations for multiple languages seamlessly. Using robust monitoring tools like the Netdata platform can provide real-time insights into your NGINX Unit, allowing IT engineers, DevOps, and SRE teams to swiftly detect issues and optimize their infrastructure.&lt;/p></description></item><item><title>NGINX upstream keepalive: eliminating per-request TCP and TLS overhead</title><link>https://www.netdata.cloud/guides/nginx/nginx-upstream-keepalive/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-upstream-keepalive/</guid><description>&lt;h1 id="nginx-upstream-keepalive-eliminating-per-request-tcp-and-tls-overhead">NGINX upstream keepalive: eliminating per-request TCP and TLS overhead&lt;/h1>
&lt;p>Every proxied request that opens a fresh TCP connection burns latency on the handshake and, if the upstream uses HTTPS, on TLS negotiation. At low volume this cost is invisible. At production throughput it becomes a measurable tax on every request, and at extreme scale it can exhaust the kernel&amp;rsquo;s ephemeral port range and bury the host in TIME_WAIT sockets. For HTTPS upstreams, the CPU cost of TLS handshakes across thousands of requests per second consumes worker cycles that could be spent proxying traffic.&lt;/p></description></item><item><title>nginx upstream prematurely closed connection while reading response header</title><link>https://www.netdata.cloud/guides/nginx/nginx-upstream-prematurely-closed-connection/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-upstream-prematurely-closed-connection/</guid><description>&lt;h1 id="nginx-upstream-prematurely-closed-connection-while-reading-response-header">nginx upstream prematurely closed connection while reading response header&lt;/h1>
&lt;p>&lt;code>upstream prematurely closed connection while reading response header from upstream&lt;/code> means the upstream server closed the TCP socket while nginx was still reading response headers. This produces a 502 Bad Gateway. Unlike a timeout, the upstream actively terminated the connection.&lt;/p>
&lt;p>The root cause is typically on the backend: a crash, worker recycle, request size limit, or stale keepalive connection the backend closed while nginx tried to reuse it. nginx retries the request on another backend only if &lt;code>proxy_next_upstream&lt;/code> includes &lt;code>error&lt;/code> (the default for idempotent methods). Retries improve availability but do not fix the underlying issue.&lt;/p></description></item><item><title>nginx upstream sent too big header while reading response header from upstream</title><link>https://www.netdata.cloud/guides/nginx/nginx-upstream-sent-too-big-header/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-upstream-sent-too-big-header/</guid><description>&lt;h1 id="nginx-upstream-sent-too-big-header-while-reading-response-header-from-upstream">nginx upstream sent too big header while reading response header from upstream&lt;/h1>
&lt;p>The error log contains &lt;code>upstream sent too big header while reading response header from upstream&lt;/code>. Clients receive 502 Bad Gateway. This is not an upstream crash or network timeout. It is a hard size limit: an upstream server is sending response headers larger than the fixed buffer nginx allocates for reading them, so nginx aborts the request and returns 502.&lt;/p></description></item><item><title>nginx upstream timed out (110: Connection timed out) while connecting/reading</title><link>https://www.netdata.cloud/guides/nginx/nginx-upstream-timed-out/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-upstream-timed-out/</guid><description>&lt;h1 id="nginx-upstream-timed-out-110-connection-timed-out-while-connectingreading">nginx upstream timed out (110: Connection timed out) while connecting/reading&lt;/h1>
&lt;p>&lt;code>upstream timed out (110: Connection timed out)&lt;/code> in the nginx error log usually surfaces to clients as a 504 Gateway Timeout. The suffix after the error string tells you which phase failed: connecting, sending, or reading. That phase determines whether you are looking at a dead backend, a network partition, or a retry storm hiding the real problem.&lt;/p>
&lt;p>The defaults are unforgiving. &lt;code>proxy_connect_timeout&lt;/code>, &lt;code>proxy_send_timeout&lt;/code>, and &lt;code>proxy_read_timeout&lt;/code> all default to 60 seconds, and &lt;code>proxy_next_upstream&lt;/code> implicitly retries on &lt;code>error&lt;/code> and &lt;code>timeout&lt;/code>. Retries can mask the root cause while exhausting upstream capacity.&lt;/p></description></item><item><title>NGINX VTS</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/nginx-vts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/nginx-vts/</guid><description/></item><item><title>NGINX VTS Monitoring</title><link>https://www.netdata.cloud/monitoring-101/nginxvts-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/nginxvts-monitoring/</guid><description>&lt;h2 id="nginx-vts-monitoring">NGINX VTS Monitoring&lt;/h2>
&lt;h3 id="what-is-nginx-vts">What Is NGINX VTS?&lt;/h3>
&lt;p>NGINX VTS (Virtual Traffic Status) is a third-party module for NGINX, a high-performance web server and reverse proxy. The VTS module provides detailed traffic status and metrics crucial for web server monitoring and management. It enables users to keep track of the overall health and performance of their NGINX server instances.&lt;/p>
&lt;h3 id="monitoring-nginx-vts-with-netdata">Monitoring NGINX VTS With Netdata&lt;/h3>
&lt;p>Using Netdata as an NGINX VTS monitoring tool allows you to collect and visualize vital metrics in real time, offering unprecedented visibility into your web server&amp;rsquo;s performance. With &lt;a href="https://www.netdata.cloud/">Netdata&lt;/a>, you can monitor NGINX VTS effortlessly by leveraging its extensive integration capabilities and live data streaming features.&lt;/p></description></item><item><title>NGINX worker_connections and worker_processes: sizing for real traffic</title><link>https://www.netdata.cloud/guides/nginx/nginx-worker-connections-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-worker-connections-tuning/</guid><description>&lt;h1 id="nginx-worker_connections-and-worker_processes-sizing-for-real-traffic">NGINX worker_connections and worker_processes: sizing for real traffic&lt;/h1>
&lt;p>NGINX defaults leave most CPU cores idle and exhaust quickly under load. The real limit is often the OS file descriptor ceiling, which silently overrides the directive.&lt;/p>
&lt;p>Sizing these parameters means understanding the capacity chain: kernel queue, connection slot, file descriptor limit, event loop. This guide provides concrete rules for static and proxy workloads and the signals that reveal when headroom has disappeared.&lt;/p></description></item><item><title>NGINX worker_rlimit_nofile: setting file descriptor limits correctly</title><link>https://www.netdata.cloud/guides/nginx/nginx-worker-rlimit-nofile/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-worker-rlimit-nofile/</guid><description>&lt;h1 id="nginx-worker_rlimit_nofile-setting-file-descriptor-limits-correctly">NGINX worker_rlimit_nofile: setting file descriptor limits correctly&lt;/h1>
&lt;p>nginx file descriptor exhaustion is a silent failure mode. Existing connections continue to be served, but new connections queue in the kernel backlog and are eventually dropped. The client sees a timeout, while nginx error logs may show nothing until the limit is hit.&lt;/p>
&lt;p>&lt;code>worker_rlimit_nofile&lt;/code> raises the per-worker file descriptor limit above conservative operating system defaults. It does not operate in isolation. It sits inside a layered stack of kernel parameters, systemd service limits, container runtime defaults, and PAM session policies. A value written into &lt;code>nginx.conf&lt;/code> is only effective if the master process inherits a hard limit at least as high, and the kernel ceiling permits it. If any layer caps the limit below your intended value, workers inherit that lower ceiling and your tuning is silently ignored.&lt;/p></description></item><item><title>nginx: a client request body is buffered to a temporary file — what it means</title><link>https://www.netdata.cloud/guides/nginx/nginx-buffered-to-temporary-file/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-buffered-to-temporary-file/</guid><description>&lt;h1 id="nginx-a-client-request-body-is-buffered-to-a-temporary-file--what-it-means">nginx: a client request body is buffered to a temporary file — what it means&lt;/h1>
&lt;p>You are tailing the nginx error log during a latency investigation and see the line: a client request body is buffered to a temporary file. It is logged at [warn], not [error], so it is easy to ignore. But the message means a request body has exceeded the in-memory buffer and nginx is now writing that data to disk. On a busy reverse proxy or file-upload endpoint, this behavior can add hundreds of milliseconds or seconds of latency before your upstream application receives the payload. It also consumes file descriptors and disk I/O capacity without ever surfacing as a 5xx status code.&lt;/p></description></item><item><title>nginx: bind() to 0.0.0.0:80 failed (98: Address already in use)</title><link>https://www.netdata.cloud/guides/nginx/nginx-address-already-in-use/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-address-already-in-use/</guid><description>&lt;h1 id="nginx-bind-to-000080-failed-98-address-already-in-use">nginx: bind() to 0.0.0.0:80 failed (98: Address already in use)&lt;/h1>
&lt;p>The error log shows &lt;code>[emerg] bind() to 0.0.0.0:80 failed (98: Address already in use)&lt;/code> and the master process exits. During a reload, old workers keep running with the previous configuration, so users may not notice immediately. During system boot or a manual start, the service is down.&lt;/p>
&lt;p>Error 98 is &lt;code>EADDRINUSE&lt;/code>. The nginx master binds listening sockets before forking workers. If the kernel reports port 80 is occupied, nginx cannot start or apply the new configuration. The holder might be a different service, a stale nginx master after a crash, or another nginx instance. Inside the same configuration, conflicting socket options for the same address:port can also trigger the error.&lt;/p></description></item><item><title>nginx: configuration file test failed - finding the syntax error</title><link>https://www.netdata.cloud/guides/nginx/nginx-configuration-file-test-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-configuration-file-test-failed/</guid><description>&lt;h1 id="nginx-configuration-file-test-failed---finding-the-syntax-error">nginx: configuration file test failed - finding the syntax error&lt;/h1>
&lt;p>An &lt;code>nginx -t&lt;/code> or &lt;code>nginx -s reload&lt;/code> ending with &lt;code>nginx: configuration file /path/to/nginx.conf test failed&lt;/code> means the configuration tree is syntactically invalid or references a missing file. The master process rejects the change, so the running server continues with the previous working configuration. That prevents an outage, but your intended change is silently inactive. Read the exact error message, map it to the real source, fix it, and validate with &lt;code>nginx -t&lt;/code> before reloading.&lt;/p></description></item><item><title>nginx: could not build server_names_hash -- server_names_hash_bucket_size</title><link>https://www.netdata.cloud/guides/nginx/nginx-server-names-hash-bucket-size/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-server-names-hash-bucket-size/</guid><description>&lt;h1 id="nginx-could-not-build-server_names_hash----server_names_hash_bucket_size">nginx: could not build server_names_hash &amp;ndash; server_names_hash_bucket_size&lt;/h1>
&lt;p>Running &lt;code>nginx -t&lt;/code> or &lt;code>nginx -s reload&lt;/code> stops with &lt;code>[emerg] could not build the server_names_hash, you should increase server_names_hash_bucket_size: 32&lt;/code>. Alternatively, nginx starts but logs &lt;code>[warn] could not build optimal server_names_hash, you should increase either server_names_hash_max_size: 512 or server_names_hash_bucket_size: 64&lt;/code>.&lt;/p>
&lt;p>This error means your configuration has exceeded a limit in nginx&amp;rsquo;s server name hashing logic. nginx builds this hash at configuration load time, not at request time. If the build fails, the configuration test fails. If the build succeeds but is suboptimal, nginx starts with degraded lookup performance. Both directives, &lt;code>server_names_hash_bucket_size&lt;/code> and &lt;code>server_names_hash_max_size&lt;/code>, are valid only inside the &lt;code>http&lt;/code> context.&lt;/p></description></item><item><title>nginx: no resolver defined to resolve - dynamic upstream DNS</title><link>https://www.netdata.cloud/guides/nginx/nginx-no-resolver-defined/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-no-resolver-defined/</guid><description>&lt;h1 id="nginx-no-resolver-defined-to-resolve---dynamic-upstream-dns">nginx: no resolver defined to resolve - dynamic upstream DNS&lt;/h1>
&lt;p>502 Bad Gateway responses paired with &lt;code>no resolver defined to resolve example.com&lt;/code> in the error log mean &lt;code>proxy_pass&lt;/code> uses a variable - for example, &lt;code>proxy_pass http://$backend;&lt;/code> - and the enclosing context has no &lt;code>resolver&lt;/code> directive.&lt;/p>
&lt;p>With a literal &lt;code>proxy_pass&lt;/code>, nginx resolves the upstream hostname once at startup or reload and caches the result indefinitely. It never queries DNS again until restart or reload. With a variable-based &lt;code>proxy_pass&lt;/code>, nginx resolves the hostname at request time through its internal async resolver. Without a &lt;code>resolver&lt;/code> directive, the lookup fails immediately and returns 502.&lt;/p></description></item><item><title>nginx: too many open files - diagnosing file descriptor exhaustion</title><link>https://www.netdata.cloud/guides/nginx/nginx-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-too-many-open-files/</guid><description>&lt;h1 id="nginx-too-many-open-files---diagnosing-file-descriptor-exhaustion">nginx: too many open files - diagnosing file descriptor exhaustion&lt;/h1>
&lt;p>After a traffic spike, the error log shows &lt;code>accept4() failed (24: Too many open files)&lt;/code>, then goes silent. Existing connections still serve, but new ones cannot land.&lt;/p>
&lt;p>File descriptor exhaustion is a hard failure. Once the limit is hit, nginx cannot accept new connections, open upstream sockets, or write to the error log. Default OS limits of 1024 are too low for production reverse proxies. Each proxied request consumes at least two FDs, and idle keepalive connections hold them indefinitely. The effective limit is the lower of &lt;code>worker_rlimit_nofile&lt;/code> and the OS hard limit enforced by systemd or the container runtime.&lt;/p></description></item><item><title>nginx: worker_connections are not enough — causes and fixes</title><link>https://www.netdata.cloud/guides/nginx/nginx-worker-connections-are-not-enough/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nginx/nginx-worker-connections-are-not-enough/</guid><description>&lt;h1 id="nginx-worker_connections-are-not-enough--causes-and-fixes">nginx: worker_connections are not enough — causes and fixes&lt;/h1>
&lt;p>Your error log shows &lt;code>worker_connections are not enough while connecting to upstream&lt;/code>. New clients time out while existing connections may still work. This is a hard capacity cliff: once a worker exhausts its connection slots, it cannot accept new connections until a slot frees. The default limit is 512 per worker, not 1024, and in reverse-proxy mode each request consumes at least two slots. Raising the number is often the first reaction, but if a slow backend is holding connections open, the slots will fill again no matter how high you set the limit.&lt;/p></description></item><item><title>NIC RSS misconfiguration: one CPU core silently dropping your telemetry</title><link>https://www.netdata.cloud/guides/network/network-collector-nic-rss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-collector-nic-rss/</guid><description>&lt;h1 id="nic-rss-misconfiguration-one-cpu-core-silently-dropping-your-telemetry">NIC RSS misconfiguration: one CPU core silently dropping your telemetry&lt;/h1>
&lt;p>Your flow collector has 16 CPU cores, but one is pinned at 100% while the other 15 sit idle. NIC receive drop counters are climbing. UDP socket buffer errors (&lt;code>Udp_RcvbufErrors&lt;/code>) are incrementing. Your bandwidth charts show traffic declining during what is actually a traffic spike. The box looks underpowered, so you start sizing a bigger one. The real problem: Receive Side Scaling (RSS) is funneling every inbound packet to a single receive queue serviced by a single CPU core. No amount of additional cores or RAM fixes this until RSS distributes interrupts across them.&lt;/p></description></item><item><title>Nice Systems Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nice-systems-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nice-systems-ltd-snmp-traps/</guid><description/></item><item><title>Nimble Storage SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nimble-storage-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nimble-storage-snmp-traps/</guid><description/></item><item><title>Nokia BGP</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/nokia-bgp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/nokia-bgp/</guid><description/></item><item><title>Nokia Formerly Alcatel Lucent SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nokia-formerly-alcatel-lucent-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nokia-formerly-alcatel-lucent-snmp-traps/</guid><description/></item><item><title>Nokia Formerly Coriant R D GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nokia-formerly-coriant-r-d-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nokia-formerly-coriant-r-d-gmbh-snmp-traps/</guid><description/></item><item><title>Nokia Formerly Infinera Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nokia-formerly-infinera-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nokia-formerly-infinera-corp-snmp-traps/</guid><description/></item><item><title>Nokia Formerly Lumentis AB SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nokia-formerly-lumentis-ab-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nokia-formerly-lumentis-ab-snmp-traps/</guid><description/></item><item><title>Nokia Formerly Transmode Systems AB SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nokia-formerly-transmode-systems-ab-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nokia-formerly-transmode-systems-ab-snmp-traps/</guid><description/></item><item><title>Nokia Networks Formerly Nokia Siemens Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nokia-networks-formerly-nokia-siemens-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nokia-networks-formerly-nokia-siemens-networks-snmp-traps/</guid><description/></item><item><title>Nokia Service Router OS</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/nokia-service-router-os/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/nokia-service-router-os/</guid><description/></item><item><title>Nokia SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nokia-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nokia-snmp-traps/</guid><description/></item><item><title>Nomad Containers</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/nomad-containers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/nomad-containers/</guid><description/></item><item><title>Nomadix SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nomadix-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nomadix-snmp-traps/</guid><description/></item><item><title>Non-Uniform Memory Access</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/non-uniform-memory-access/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/non-uniform-memory-access/</guid><description/></item><item><title>Normalizing syslog severity across vendors: why 'critical' isn't critical</title><link>https://www.netdata.cloud/guides/network/network-syslog-severity-normalization/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-syslog-severity-normalization/</guid><description>&lt;h1 id="normalizing-syslog-severity-across-vendors-why-critical-isnt-critical">Normalizing syslog severity across vendors: why &amp;lsquo;critical&amp;rsquo; isn&amp;rsquo;t critical&lt;/h1>
&lt;p>RFC 5424 defines eight syslog severity levels, numbered 0 through 7: Emergency, Alert, Critical, Error, Warning, Notice, Informational, and Debug. Every major network vendor implements the same numeric scale. The integer that means &amp;ldquo;Critical&amp;rdquo; on a Cisco router means &amp;ldquo;Critical&amp;rdquo; on a Juniper switch.&lt;/p>
&lt;p>But the severity a device assigns to a given event is not standardized. A BGP session reset might arrive as severity 5 (Notice) from one vendor and severity 3 (Error) from another. A hardware alarm that one platform logs as Critical (2), another logs as Alert (1) or Warning (4). Same operational condition, different severity label, same RFC.&lt;/p></description></item><item><title>Northern Telecom Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/northern-telecom-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/northern-telecom-ltd-snmp-traps/</guid><description/></item><item><title>Novell SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/novell-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/novell-snmp-traps/</guid><description/></item><item><title>Novelsat SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/novelsat-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/novelsat-snmp-traps/</guid><description/></item><item><title>NRPE daemon</title><link>https://www.netdata.cloud/integrations/data-collection/applications/nrpe-daemon/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/nrpe-daemon/</guid><description/></item><item><title>NRPE daemon Monitoring</title><link>https://www.netdata.cloud/monitoring-101/nrpe-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/nrpe-monitoring/</guid><description>&lt;h2 id="nrpe-daemon-monitoring">NRPE daemon Monitoring&lt;/h2>
&lt;h3 id="what-is-nrpe-daemon">What Is NRPE daemon?&lt;/h3>
&lt;p>The Nagios Remote Plugin Executor (NRPE) daemon is an integral part of system and network monitoring, widely utilized to execute remote commands and scripts. This daemon facilitates the collection of key performance metrics from remote systems, providing insights into system health and performance. It is especially crucial for environments relying on Nagios for extensive monitoring capabilities, allowing for seamless integration of remote checks into the central monitoring framework.&lt;/p></description></item><item><title>Nsc Communications Siberia Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nsc-communications-siberia-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nsc-communications-siberia-ltd-snmp-traps/</guid><description/></item><item><title>Nsc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nsc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nsc-snmp-traps/</guid><description/></item><item><title>NSD</title><link>https://www.netdata.cloud/integrations/data-collection/networking/nsd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/nsd/</guid><description/></item><item><title>NSD Monitoring</title><link>https://www.netdata.cloud/monitoring-101/nsd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/nsd-monitoring/</guid><description>&lt;h2 id="nsd-monitoring">NSD Monitoring&lt;/h2>
&lt;h3 id="what-is-nsd">What Is NSD?&lt;/h3>
&lt;p>NSD is an authoritative DNS name server developed by NLnet Labs. It is renowned for its high performance, robust security measures, and support for various DNS standards. NSD is a cornerstone for businesses looking for a reliable DNS solution to ensure smooth and secure network communication. You can find more in-depth information on the &lt;a href="https://nsd.docs.nlnetlabs.nl/en/latest">NSD official documentation&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-nsd-with-netdata">Monitoring NSD With Netdata&lt;/h3>
&lt;p>Monitoring NSD effectively is crucial to maintaining optimal server performance and securing uptime. Netdata offers a comprehensive NSD monitoring tool that easily integrates with your system to provide in-depth insights into your DNS server&amp;rsquo;s performance. By leveraging Netdata, you can continuously monitor NSD metrics in real time, enabling swift diagnosis and resolution of potential issues.&lt;/p></description></item><item><title>Nsi Software SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nsi-software-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nsi-software-snmp-traps/</guid><description/></item><item><title>ntfy</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/ntfy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/ntfy/</guid><description/></item><item><title>NTP drift on network devices: the silent killer of event correlation</title><link>https://www.netdata.cloud/guides/network/network-ntp-drift/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-ntp-drift/</guid><description>&lt;h1 id="ntp-drift-on-network-devices-the-silent-killer-of-event-correlation">NTP drift on network devices: the silent killer of event correlation&lt;/h1>
&lt;p>Clock drift on network devices produces no visible symptom. The device stays up, interfaces carry traffic, BGP sessions remain Established, SNMP keeps responding. The damage surfaces hours or days later, in a postmortem where two devices&amp;rsquo; timestamps disagree by hundreds of milliseconds and the analyst cannot reconstruct the event sequence. Every cross-device correlation in the monitoring stack depends on accurate, monotonic time across every collector and every polled device.&lt;/p></description></item><item><title>NTPd</title><link>https://www.netdata.cloud/integrations/data-collection/networking/ntpd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/ntpd/</guid><description/></item><item><title>NTPd Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ntpd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ntpd-monitoring/</guid><description>&lt;h2 id="ntpd-monitoring">NTPd Monitoring&lt;/h2>
&lt;h3 id="what-is-ntpd">What Is NTPd?&lt;/h3>
&lt;p>NTPd stands for Network Time Protocol daemon, an essential component for time synchronization across computer networks. By utilizing the NTP protocol, NTPd ensures that timekeeping is accurate and consistent across systems, which is crucial for numerous applications and services.&lt;/p>
&lt;h3 id="monitoring-ntpd-with-netdata">Monitoring NTPd With Netdata&lt;/h3>
&lt;p>When you monitor NTPd with Netdata, you gain real-time insights into your time synchronization processes. The &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/ntpd/?utm_source=website&amp;amp;utm_content=monitoring101">NTPd monitoring tool&lt;/a> by Netdata offers comprehensive visibility into NTPd&amp;rsquo;s operational metrics, helping you to ensure your network&amp;rsquo;s timing accuracy.&lt;/p></description></item><item><title>NUMA Architecture</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/numa-architecture/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/numa-architecture/</guid><description/></item><item><title>Nutanix Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nutanix-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nutanix-inc-snmp-traps/</guid><description/></item><item><title>Nvidia</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/nvidia/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/nvidia/</guid><description/></item><item><title>NVIDIA BAR1 memory exhaustion: mapping failures with free framebuffer</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-bar1-memory-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-bar1-memory-exhaustion/</guid><description>&lt;h1 id="nvidia-bar1-memory-exhaustion-mapping-failures-with-free-framebuffer">NVIDIA BAR1 memory exhaustion: mapping failures with free framebuffer&lt;/h1>
&lt;p>Your job dies with a CUDA allocation or mapping error, but &lt;code>nvidia-smi&lt;/code> shows gigabytes of framebuffer free. You check for fragmentation, restart the job, and it fails again at the same place. The resource that ran out is not VRAM. It is the BAR1 aperture, the PCIe-mapped window the CPU uses to reach GPU memory directly.&lt;/p>
&lt;p>BAR1 exhaustion is a distinct failure mode from framebuffer OOM. A GPU can have most of its framebuffer free and still refuse new mappings because every process, IPC handle, and GPUDirect RDMA registration consumes space in a small, shared aperture. It is most common in multi-process, MPS, containerized, and multi-tenant environments where many processes map GPU memory at once.&lt;/p></description></item><item><title>NVIDIA BGP</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/nvidia-bgp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/bgp-monitoring/nvidia-bgp/</guid><description/></item><item><title>Nvidia Cumulus Linux Switch</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/nvidia-cumulus-linux-switch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/nvidia-cumulus-linux-switch/</guid><description/></item><item><title>Nvidia Data Center GPU Manager (DCGM)</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/nvidia-data-center-gpu-manager-dcgm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/nvidia-data-center-gpu-manager-dcgm/</guid><description/></item><item><title>NVIDIA DCGM exporter on Kubernetes: mapping GPU metrics to pods</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-dcgm-exporter-kubernetes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-dcgm-exporter-kubernetes/</guid><description>&lt;h1 id="nvidia-dcgm-exporter-on-kubernetes-mapping-gpu-metrics-to-pods">NVIDIA DCGM exporter on Kubernetes: mapping GPU metrics to pods&lt;/h1>
&lt;p>On a bare-metal GPU node, attribution is simple: &lt;code>nvidia-smi&lt;/code> shows you a PID, you look up the process, done. On Kubernetes there are three layers between you and that answer. The device plugin assigns GPUs to pods, the kubelet tracks those assignments, and dcgm-exporter reads GPU state through NVML/DCGM, which has no idea what a pod is. Unless the exporter explicitly joins these two views, your GPU metrics are half-blind: you can see GPU 3 on node 17 at 98% utilization, but not which workload is responsible.&lt;/p></description></item><item><title>NVIDIA driver and CUDA version mismatch: subtle failures and crashes</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-driver-cuda-version-mismatch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-driver-cuda-version-mismatch/</guid><description>&lt;h1 id="nvidia-driver-and-cuda-version-mismatch-subtle-failures-and-crashes">NVIDIA driver and CUDA version mismatch: subtle failures and crashes&lt;/h1>
&lt;p>The loud version of this problem is easy: a job dies at startup with &lt;code>CUDA driver version is insufficient for CUDA runtime version&lt;/code>, and the fix is a driver upgrade. The version that ruins your week is quieter. A kernel update goes out fleet-wide, the NVIDIA module rebuilds on most nodes but not all, and now a subset of training jobs crash with kernel-launch failures or, worse, run to completion with subtly wrong results. Nothing in your dashboards looks broken: GPUs up, memory fine, temperatures normal.&lt;/p></description></item><item><title>NVIDIA Fabric Manager not running: NVSwitch GPUs lose NVLink</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-fabric-manager-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-fabric-manager-down/</guid><description>&lt;h1 id="nvidia-fabric-manager-not-running-nvswitch-gpus-lose-nvlink">NVIDIA Fabric Manager not running: NVSwitch GPUs lose NVLink&lt;/h1>
&lt;p>Your multi-GPU training job hangs during NCCL initialization, or fails with &lt;code>cudaErrorSystemNotReady&lt;/code> (error 802) the moment a process touches CUDA. Each GPU shows up in &lt;code>nvidia-smi&lt;/code> with normal temperature, memory, and utilization. You burn hours on NCCL debug logs, InfiniBand checks, and application-level bisection before someone runs &lt;code>systemctl status nvidia-fabricmanager&lt;/code> and finds the service failed two days ago after a reboot.&lt;/p></description></item><item><title>Nvidia GPU</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/nvidia-gpu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/nvidia-gpu/</guid><description/></item><item><title>NVIDIA GPU clock throttle reasons: why the GPU isn't running at full speed</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-throttle-reasons/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-throttle-reasons/</guid><description>&lt;h1 id="nvidia-gpu-clock-throttle-reasons-why-the-gpu-isnt-running-at-full-speed">NVIDIA GPU clock throttle reasons: why the GPU isn&amp;rsquo;t running at full speed&lt;/h1>
&lt;p>A training job that used to finish an epoch in 40 minutes now takes 90. &lt;code>nvidia-smi&lt;/code> shows 100% GPU utilization, memory usage looks normal, and the workload is running, just slowly. This is the classic throttling symptom: the GPU is executing kernels at reduced clock speeds, so everything takes longer while utilization stays high.&lt;/p>
&lt;p>The mistake most teams make here is guessing. They check temperature, see 82C, and debate whether that is &amp;ldquo;too hot.&amp;rdquo; They see power draw pinned at the limit. None of that says why the clocks are down. The GPU already knows, and it reports the answer directly through the clock event reasons bitmask.&lt;/p></description></item><item><title>NVIDIA GPU configuration drift: ECC, persistence, power limits, compute mode</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-configuration-drift/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-configuration-drift/</guid><description>&lt;h1 id="nvidia-gpu-configuration-drift-ecc-persistence-power-limits-compute-mode">NVIDIA GPU configuration drift: ECC, persistence, power limits, compute mode&lt;/h1>
&lt;p>A GPU that passes every health check can still be misconfigured. ECC turned off, persistence mode lost after a reboot, a power limit set below default, compute mode flipped to Exclusive. None of these produce an XID error or a crashed job on their own. They produce silent data corruption, unexplained throttling, scheduling failures, and monitoring gaps that surface days later as a different incident.&lt;/p></description></item><item><title>NVIDIA GPU ECC disabled: the silent data-corruption risk</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-ecc-disabled/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-ecc-disabled/</guid><description>&lt;h1 id="nvidia-gpu-ecc-disabled-the-silent-data-corruption-risk">NVIDIA GPU ECC disabled: the silent data-corruption risk&lt;/h1>
&lt;p>A datacenter GPU with ECC disabled does not crash when memory goes bad. It keeps running, keeps returning answers, and keeps writing checkpoints, while single-bit errors corrupt whatever it is computing. There is no error counter incrementing, no XID in dmesg, no page in the middle of the night. The usual discovery path is a training run that diverges to NaN on one specific node, or an inference service returning subtly wrong results, days or weeks after someone flipped the setting.&lt;/p></description></item><item><title>NVIDIA GPU ECC errors: corrected, uncorrected, volatile, and aggregate</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-ecc-errors-corrected-uncorrected/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-ecc-errors-corrected-uncorrected/</guid><description>&lt;h1 id="nvidia-gpu-ecc-errors-corrected-uncorrected-volatile-and-aggregate">NVIDIA GPU ECC errors: corrected, uncorrected, volatile, and aggregate&lt;/h1>
&lt;p>You opened &lt;code>nvidia-smi -q -d ECC&lt;/code> because something looked off, and now you are staring at four sets of counters with non-zero values and no idea which ones matter. Maybe a monitoring alert fired on &amp;ldquo;ECC errors present&amp;rdquo;. Maybe a training run produced NaN loss on one node and someone told you to check ECC. Either way, the raw output does not tell you the two things you need to know: has data been corrupted, and is this GPU dying.&lt;/p></description></item><item><title>NVIDIA GPU fan at 0%: fan failure on air-cooled cards</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-fan-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-fan-failure/</guid><description>&lt;h1 id="nvidia-gpu-fan-at-0-fan-failure-on-air-cooled-cards">NVIDIA GPU fan at 0%: fan failure on air-cooled cards&lt;/h1>
&lt;p>&lt;code>nvidia-smi&lt;/code> shows &lt;code>Fan: 0%&lt;/code> and the GPU is sitting at 70C under load. Either the fan is dead, or the card is doing exactly what its firmware told it to do. Telling those two apart quickly is the whole job: the wrong guess in one direction means a cooked GPU, the wrong guess in the other means a pointless RMA.&lt;/p></description></item><item><title>NVIDIA GPU HBM (memory) temperature: the thermal limit most teams miss</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-hbm-memory-temperature/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-hbm-memory-temperature/</guid><description>&lt;h1 id="nvidia-gpu-hbm-memory-temperature-the-thermal-limit-most-teams-miss">NVIDIA GPU HBM (memory) temperature: the thermal limit most teams miss&lt;/h1>
&lt;p>Most GPU fleets monitor one temperature: &lt;code>temperature.gpu&lt;/code>, the die temperature. Alerts, dashboards, and throttle runbooks are built around it. On datacenter GPUs with HBM (A100, H100, H200), that is only half the picture, and for memory-bound work it is often the wrong half.&lt;/p>
&lt;p>The HBM stacks sit on the same package as the die but have a worse heat path. Under LLM inference, embedding-heavy workloads, and memory-bound training, HBM temperature commonly runs 10 to 20 C hotter than the die. HBM also has its own throttle point. When it trips, the GPU reduces memory bandwidth with no Xid, no error, and no log entry most teams watch. The job keeps running. It just gets slower while die temperature still looks healthy.&lt;/p></description></item><item><title>NVIDIA GPU HBM progressive failure: from single-bit errors to a dead GPU</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-hbm-progressive-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-hbm-progressive-failure/</guid><description>&lt;h1 id="nvidia-gpu-hbm-progressive-failure-from-single-bit-errors-to-a-dead-gpu">NVIDIA GPU HBM progressive failure: from single-bit errors to a dead GPU&lt;/h1>
&lt;p>HBM failure on a datacenter GPU is almost never a surprise if you are watching the right signals. It is a staged degradation: single-bit ECC errors rise over days or weeks, the GPU burns through its self-repair capacity (row remaps on Ampere and later, page retirement on pre-Ampere), the first uncorrectable double-bit error lands, and eventually the GPU can no longer repair itself or falls off the PCIe bus entirely.&lt;/p></description></item><item><title>NVIDIA GPU HW Power Brake Slowdown: the chassis is cutting GPU power</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-hw-power-brake-slowdown/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-hw-power-brake-slowdown/</guid><description>&lt;h1 id="nvidia-gpu-hw-power-brake-slowdown-the-chassis-is-cutting-gpu-power">NVIDIA GPU HW Power Brake Slowdown: the chassis is cutting GPU power&lt;/h1>
&lt;p>You are looking at a GPU that is running, reachable, not overheating, and still delivering a fraction of its normal throughput. &lt;code>nvidia-smi -q -d PERFORMANCE&lt;/code> shows &lt;code>HW Power Brake Slowdown : Active&lt;/code> (and usually &lt;code>HW Slowdown : Active&lt;/code> alongside it), SM clocks are at half their rated speed or lower, and the temperature looks fine. The workload did not change. The GPU did not change. Something outside the GPU decided it gets less power.&lt;/p></description></item><item><title>NVIDIA GPU memory leak: framebuffer usage climbing without a plateau</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-memory-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-memory-leak/</guid><description>&lt;h1 id="nvidia-gpu-memory-leak-framebuffer-usage-climbing-without-a-plateau">NVIDIA GPU memory leak: framebuffer usage climbing without a plateau&lt;/h1>
&lt;p>Your training job has been running for six hours. &lt;code>nvidia-smi&lt;/code> showed 40 GiB used after warmup. Then 45. Then 52. The curve has not flattened, and at this rate the job dies with &lt;code>CUDA out of memory&lt;/code> sometime tomorrow, taking a day of checkpoint progress with it.&lt;/p>
&lt;p>GPU memory is a cliff resource: there is no swap and no graceful degradation. When an allocation fails, it fails immediately. The saving grace is that a true leak is visible hours before it kills you, as sustained linear growth in framebuffer usage. The hard part is telling that apart from the many things that look like a leak but are not, because ML frameworks deliberately fill memory and hold it.&lt;/p></description></item><item><title>NVIDIA GPU memory-bandwidth bound: high DRAM activity, idle SMs</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-memory-bandwidth-bound/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-memory-bandwidth-bound/</guid><description>&lt;h1 id="nvidia-gpu-memory-bandwidth-bound-high-dram-activity-idle-sms">NVIDIA GPU memory-bandwidth bound: high DRAM activity, idle SMs&lt;/h1>
&lt;p>Your training job or inference service is running, &lt;code>nvidia-smi&lt;/code> shows high GPU utilization, power draw looks healthy, and throughput is far below what the hardware should deliver. The profiling metrics show the pattern: the device memory interface is active nearly 100% of cycles while the streaming multiprocessors sit mostly idle. The GPU is memory-bandwidth bound. It is not short on compute. It is waiting on HBM.&lt;/p></description></item><item><title>Nvidia GPU Monitoring</title><link>https://www.netdata.cloud/monitoring-101/nvidia-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/nvidia-monitoring/</guid><description>&lt;h2 id="what-is-nvidia-gpu">What is Nvidia GPU?&lt;/h2>
&lt;p>Nvidia GPU (Graphic Processing Unit) is a specialized electronic circuit designed to rapidly process and manipulate graphics data. Nvidia GPUs are typically found in high-end gaming computers, workstations, and servers. They are used to power complex video games and other graphics-intensive tasks.&lt;/p>
&lt;h2 id="monitoring-nvidia-gpu-with-netdata">Monitoring Nvidia GPU with Netdata&lt;/h2>
&lt;p>The prerequisites for monitoring Nvidia GPU with Netdata are to have a system with an Nvidia GPU and &lt;a href="https://learn.netdata.cloud/docs/cloud/get-started/">Netdata installed&lt;/a> on your system.&lt;/p></description></item><item><title>Nvidia GPU Monitoring</title><link>https://www.netdata.cloud/monitoring-101/nvidia_smi-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/nvidia_smi-monitoring/</guid><description>&lt;h2 id="nvidia-gpu-monitoring">Nvidia GPU Monitoring&lt;/h2>
&lt;h3 id="what-is-nvidia-gpu">What Is Nvidia GPU?&lt;/h3>
&lt;p>Nvidia GPUs are specialized processing units designed by &lt;a href="https://www.nvidia.com/en-us/">Nvidia&lt;/a> primarily for graphics rendering, though they are widely used in computational tasks such as deep learning, scientific simulations, and cryptocurrency mining. Nvidia&amp;rsquo;s advanced GPU technology empowers applications to perform complex tasks efficiently.&lt;/p>
&lt;h3 id="monitoring-nvidia-gpu-with-netdata">Monitoring Nvidia GPU With Netdata&lt;/h3>
&lt;p>Netdata provides real-time monitoring for Nvidia GPUs by leveraging the &lt;code>nvidia-smi&lt;/code> CLI tool. This setup allows you to keep an eye on various performance metrics, ensuring optimal operation and helping diagnose potential issues as they occur.&lt;/p></description></item><item><title>NVIDIA GPU monitoring checklist: the signals every production GPU fleet needs</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-monitoring-checklist/</guid><description>&lt;h1 id="nvidia-gpu-monitoring-checklist-the-signals-every-production-gpu-fleet-needs">NVIDIA GPU monitoring checklist: the signals every production GPU fleet needs&lt;/h1>
&lt;p>Most GPU monitoring setups fail in one of two ways. Either they collect a handful of nvidia-smi counters and page on the wrong things (raw temperature, memory percentage, idle PCIe downgrade), or they collect everything and alert on nothing meaningful. Both failure modes come from the same root cause: no shared vocabulary for which signals matter, at what fidelity, and with what alert semantics.&lt;/p></description></item><item><title>NVIDIA GPU monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-monitoring-maturity-model/</guid><description>&lt;h1 id="nvidia-gpu-monitoring-maturity-model-from-survival-to-expert">NVIDIA GPU monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most GPU fleets sit at one of two extremes: nothing beyond &amp;ldquo;does &lt;code>nvidia-smi&lt;/code> still respond,&amp;rdquo; or a wall of dashboards nobody wired to an alert. Neither survives a real incident. A GPU falling off the bus at 3 a.m., a training job running 5x slower because one card is thermally throttled, an accelerating ECC error rate that becomes silent data corruption next week: all detectable, but only if you collect the right signals at the right fidelity.&lt;/p></description></item><item><title>NVIDIA GPU NVLink down or degraded: a link that dropped in multi-GPU training</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-nvlink-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-nvlink-down/</guid><description>&lt;h1 id="nvidia-gpu-nvlink-down-or-degraded-a-link-that-dropped-in-multi-gpu-training">NVIDIA GPU NVLink down or degraded: a link that dropped in multi-GPU training&lt;/h1>
&lt;p>A link that drops mid-training is one of the nastiest GPU failures to catch, because the job usually does not crash. It hangs at a synchronization barrier, or it keeps running at a fraction of its normal step rate. Bandwidth between GPUs collapses from hundreds of GB/s (600 GB/s on A100, 900 GB/s on H100) to PCIe speeds, and every AllReduce now waits on the slowest path in the collective.&lt;/p></description></item><item><title>NVIDIA GPU NVLink errors: CRC and replay errors on GPU-to-GPU links</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-nvlink-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-nvlink-errors/</guid><description>&lt;h1 id="nvidia-gpu-nvlink-errors-crc-and-replay-errors-on-gpu-to-gpu-links">NVIDIA GPU NVLink errors: CRC and replay errors on GPU-to-GPU links&lt;/h1>
&lt;p>Your training job is running. No crashes, no CUDA errors, no OOM. But step time is up roughly 20%, and the slowdown shows up exactly during the collective phases: AllReduce, gradient sync. The team has already blamed the data loader, shuffled hyperparameters, and re-run the job twice. The actual problem is one degraded NVLink cable retransmitting corrupted flits, and nothing in the default monitoring stack is looking at it.&lt;/p></description></item><item><title>NVIDIA GPU PCIe link running below maximum: x16 silently training at x8</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-pcie-link-degraded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-pcie-link-degraded/</guid><description>&lt;h1 id="nvidia-gpu-pcie-link-running-below-maximum-x16-silently-training-at-x8">NVIDIA GPU PCIe link running below maximum: x16 silently training at x8&lt;/h1>
&lt;p>The job runs. Health checks pass. nvidia-smi shows the GPU, temperature is fine, ECC is clean, and utilization looks normal. Yet training throughput is half of what it should be, or one GPU in an eight-GPU node drags every collective operation. A common root cause is a PCIe link that negotiated below its capability: a Gen4 x16 card silently running at Gen3, or worse, at x8 or x4.&lt;/p></description></item><item><title>NVIDIA GPU PCIe replay counter rising: physical-layer signal integrity</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-pcie-replay-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-pcie-replay-errors/</guid><description>&lt;h1 id="nvidia-gpu-pcie-replay-counter-rising-physical-layer-signal-integrity">NVIDIA GPU PCIe replay counter rising: physical-layer signal integrity&lt;/h1>
&lt;p>The PCIe replay counter on one of your GPUs is climbing. Not a burst at boot, but a sustained rate of hundreds or thousands per second. Nothing has crashed yet: training still runs, nvidia-smi still answers, and there are no XID errors in dmesg. That is exactly why this signal gets ignored, and exactly why it should not be.&lt;/p>
&lt;p>A PCIe replay is a link-layer retransmission. The GPU (or the root complex, switch, or retimer on the path) received a packet that failed its CRC check and asked for it again. Occasional replays are normal PCIe behavior. A sustained high rate means bits are being corrupted on the wire often enough that retransmission is routine, not exceptional. That is a physical-layer signal integrity problem: a cable, connector, slot, riser, or an overheating retimer or PCIe switch somewhere between the CPU and the GPU.&lt;/p></description></item><item><title>NVIDIA GPU PCIe straggler: one slow link throttling multi-GPU training</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-pcie-straggler-multi-gpu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-pcie-straggler-multi-gpu/</guid><description>&lt;h1 id="nvidia-gpu-pcie-straggler-one-slow-link-throttling-multi-gpu-training">NVIDIA GPU PCIe straggler: one slow link throttling multi-GPU training&lt;/h1>
&lt;p>Your distributed training job is running. No errors, no crashes, no XIDs. But step time jumped 30-60% and never came back, and scaling from 4 GPUs to 8 GPUs bought almost nothing. Per-GPU utilization shows one GPU a few points below the rest, and its PCIe throughput is consistently below its peers.&lt;/p>
&lt;p>That is the PCIe straggler pattern. One GPU&amp;rsquo;s link negotiated at a lower generation or narrower width than it should have: Gen3 instead of Gen4, x8 instead of x16, or worse. That GPU now feeds data and exchanges gradients at half (or less) of its peers&amp;rsquo; bandwidth. In synchronous data-parallel training, every collective waits for the slowest participant, so the whole job runs at the speed of the worst link.&lt;/p></description></item><item><title>NVIDIA GPU Remapping Failure Occurred: HBM spare rows exhausted (Ampere+)</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-row-remapping-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-row-remapping-failure/</guid><description>&lt;h1 id="nvidia-gpu-remapping-failure-occurred-hbm-spare-rows-exhausted-ampere">NVIDIA GPU Remapping Failure Occurred: HBM spare rows exhausted (Ampere+)&lt;/h1>
&lt;p>You ran &lt;code>nvidia-smi -q -d ROW_REMAPPER&lt;/code> (or your monitoring did) and saw &lt;code>Remapping Failure Occurred : Yes&lt;/code>. This is not a driver problem, not a configuration problem, and not something a reboot clears. It is the GPU telling you that its hardware self-repair budget is spent.&lt;/p>
&lt;p>On Ampere and later GPUs (A100, H100, H200, Blackwell), the memory subsystem does not retire whole pages the way Volta and Turing did. Instead, when HBM rows start failing, the GPU remaps them to a finite pool of spare rows built into each DRAM bank. While spares remain, the GPU heals itself transparently: correctable ECC events trigger remaps, XID 63 appears in the kernel log as an informational note, and workloads never notice.&lt;/p></description></item><item><title>NVIDIA GPU retired pages: page retirement, pending retirements, and end of life</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-retired-pages/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-retired-pages/</guid><description>&lt;h1 id="nvidia-gpu-retired-pages-page-retirement-pending-retirements-and-end-of-life">NVIDIA GPU retired pages: page retirement, pending retirements, and end of life&lt;/h1>
&lt;p>Retired pages are the permanent record of a GPU&amp;rsquo;s memory hardware degrading. Every retirement means the driver found a framebuffer page it could no longer trust and removed it from the allocatable pool for good. The counts persist in the GPU&amp;rsquo;s InfoROM across reboots and driver reloads, which makes them one of the few genuinely cumulative hardware health signals nvidia-smi exposes.&lt;/p></description></item><item><title>NVIDIA GPU single-bit ECC error rate rising: the leading indicator of HBM failure</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-single-bit-ecc-rate-rising/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-single-bit-ecc-rate-rising/</guid><description>&lt;h1 id="nvidia-gpu-single-bit-ecc-error-rate-rising-the-leading-indicator-of-hbm-failure">NVIDIA GPU single-bit ECC error rate rising: the leading indicator of HBM failure&lt;/h1>
&lt;p>A GPU in your fleet has started logging corrected ECC errors faster than it used to. Maybe you noticed a bump in &lt;code>ecc.errors.corrected.volatile.total&lt;/code> during a routine check, or Xid 92 events started appearing in dmesg. Nothing is broken yet: jobs run, loss curves look normal, nvidia-smi reports a healthy card. This is exactly the moment most teams get wrong. They look at the absolute count, decide it is small, and move on.&lt;/p></description></item><item><title>NVIDIA GPU SW Power Cap: performance capped by the power limit</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-sw-power-cap-throttling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-sw-power-cap-throttling/</guid><description>&lt;h1 id="nvidia-gpu-sw-power-cap-performance-capped-by-the-power-limit">NVIDIA GPU SW Power Cap: performance capped by the power limit&lt;/h1>
&lt;p>&lt;code>nvidia-smi&lt;/code> reports &lt;code>SW Power Cap: Active&lt;/code> under Clocks Event Reasons, the workload is running slower than it should, and the GPU is not hot. This is the power-limit throttle: the GPU&amp;rsquo;s power scaling algorithm is holding clocks below what the workload requested because the board is consuming as much power as it is allowed to.&lt;/p>
&lt;p>Two things make this symptom confusing. First, SW Power Cap is not always a fault. A GPU at full load sitting exactly at its power limit is doing the maximum work its TDP allows, and datacenter operators often cap GPUs below the factory limit on purpose. Second, the number that actually matters is &lt;code>enforced.power.limit&lt;/code>, not &lt;code>power.limit&lt;/code>, and the two can disagree. Reading the wrong field sends you chasing a misconfiguration that does not exist, or missing one that does.&lt;/p></description></item><item><title>NVIDIA GPU temperature too high: die temperature and thermal limits</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-temperature-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-temperature-high/</guid><description>&lt;h1 id="nvidia-gpu-temperature-too-high-die-temperature-and-thermal-limits">NVIDIA GPU temperature too high: die temperature and thermal limits&lt;/h1>
&lt;p>A GPU reporting 85 °C is not necessarily a GPU with a temperature problem. Under sustained load, datacenter GPUs like the A100 and H100 are designed to run hot, and 80 °C is routine under full load. The number alone tells you almost nothing.&lt;/p>
&lt;p>What matters is where the temperature sits relative to the SKU&amp;rsquo;s own thermal limits, whether the GPU has started throttling, and whether the pattern is one hot GPU or a whole chassis running hot. A GPU at 85 °C with no throttle reasons active is fine. A GPU at 75 °C with &lt;code>hw_thermal_slowdown&lt;/code> active has a cooling system that is already failing.&lt;/p></description></item><item><title>NVIDIA GPU Tensor Cores idle: mixed precision not being used</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-tensor-cores-idle/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-tensor-cores-idle/</guid><description>&lt;h1 id="nvidia-gpu-tensor-cores-idle-mixed-precision-not-being-used">NVIDIA GPU Tensor Cores idle: mixed precision not being used&lt;/h1>
&lt;p>Your training job is running, &lt;code>nvidia-smi&lt;/code> shows high SM utilization, power draw looks healthy, and step time is 5 to 10 times worse than the hardware spec says it should be. DCGM profiling shows Tensor Core Active pinned at or near zero. The GPU is busy, but it is doing the work on CUDA cores instead of Tensor Cores.&lt;/p>
&lt;p>Nothing errors. Nothing crashes. The job just runs at a fraction of the throughput you paid for, and because SM utilization still reads 90%+, it looks healthy in every dashboard that only tracks &lt;code>utilization.gpu&lt;/code>.&lt;/p></description></item><item><title>NVIDIA GPU thermal cascade in dense servers: one hot GPU heats its neighbours</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-thermal-cascade-dense-servers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-thermal-cascade-dense-servers/</guid><description>&lt;h1 id="nvidia-gpu-thermal-cascade-in-dense-servers-one-hot-gpu-heats-its-neighbours">NVIDIA GPU thermal cascade in dense servers: one hot GPU heats its neighbours&lt;/h1>
&lt;p>Three of the eight GPUs in a node are thermal throttling. Training step time has doubled. The instinct is to suspect the cooling system as a whole, or the workload. But line up per-GPU temperatures and one unit stands out: GPU 5 is running 12 degrees hotter than everything else, and the GPUs downstream of it in the airflow path are the ones throttling.&lt;/p></description></item><item><title>NVIDIA GPU thermal throttling: clocks dropping under heat</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-thermal-throttling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-thermal-throttling/</guid><description>&lt;h1 id="nvidia-gpu-thermal-throttling-clocks-dropping-under-heat">NVIDIA GPU thermal throttling: clocks dropping under heat&lt;/h1>
&lt;p>Training throughput fell 25% overnight. No errors in the application logs. &lt;code>nvidia-smi&lt;/code> shows every GPU at 100% utilization. Nothing looks broken, and yet the job is measurably slower.&lt;/p>
&lt;p>This is what NVIDIA GPU thermal throttling looks like in practice: the GPU keeps working, keeps reporting full utilization, and silently delivers 10-40% less real work because its SM clocks have dropped. Utilization measures time busy, not work done. A kernel running at half clock speed still occupies the SMs for the whole sampling window, so &lt;code>utilization.gpu&lt;/code> stays pinned at 100% while tokens/sec or samples/sec collapse.&lt;/p></description></item><item><title>NVIDIA GPU utilization low during training: a data-pipeline bottleneck</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-low-utilization-data-starvation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-low-utilization-data-starvation/</guid><description>&lt;h1 id="nvidia-gpu-utilization-low-during-training-a-data-pipeline-bottleneck">NVIDIA GPU utilization low during training: a data-pipeline bottleneck&lt;/h1>
&lt;p>Your training job is running, the loss is decreasing, and &lt;code>nvidia-smi&lt;/code> shows the GPU at 20-40% utilization, or oscillating between 0% and 100% in a sawtooth. The job finishes, so it looks healthy. It is not: you are paying for a GPU that is idle most of the time because the host cannot feed it.&lt;/p>
&lt;p>Low utilization during active training is almost never a GPU fault. It means the GPU is starved: waiting on CPU preprocessing, disk or network data loading, pageable-memory copies across PCIe, or a synchronization barrier. The hardware is fine; the pipeline upstream of it is the bottleneck.&lt;/p></description></item><item><title>Nvidia Mellanox Switchx</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/nvidia-mellanox-switchx/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/nvidia-mellanox-switchx/</guid><description/></item><item><title>NVIDIA persistence mode: why the GPU keeps re-initializing and P-state flaps</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-persistence-mode/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-persistence-mode/</guid><description>&lt;h1 id="nvidia-persistence-mode-why-the-gpu-keeps-re-initializing-and-p-state-flaps">NVIDIA persistence mode: why the GPU keeps re-initializing and P-state flaps&lt;/h1>
&lt;p>A training node accepts a job, and the first CUDA call takes two seconds before any kernel runs. The next job lands, and it happens again. Your dashboard shows the GPU bouncing between P0 and P8 all day, clocks collapsing to idle between launches, power draw dropping toward zero. It looks like throttling. It looks like the GPU is restarting constantly. It is neither: the driver is unloading GPU state every time the last CUDA context exits, and reinitializing it on the next one.&lt;/p></description></item><item><title>NVIDIA Xid 13: Graphics Engine Exception</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-13-graphics-engine-exception/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-13-graphics-engine-exception/</guid><description>&lt;h1 id="nvidia-xid-13-graphics-engine-exception">NVIDIA Xid 13: Graphics Engine Exception&lt;/h1>
&lt;p>You found a line like this in the kernel log:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-text" data-lang="text">&lt;span style="display:flex;">&lt;span>NVRM: Xid (PCI:0000:65:00.0): 13, pid=&amp;#39;&amp;lt;unknown&amp;gt;&amp;#39;, name=&amp;lt;unknown&amp;gt;, Graphics Engine Exception.
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Or, more likely, you found the application-side symptom first: a training job died with a generic &lt;code>CUDA error&lt;/code>, a run produced NaN loss, or an inference service started returning garbage, and only after digging into &lt;code>dmesg&lt;/code> did Xid 13 show up.&lt;/p>
&lt;p>Xid 13 is the most ambiguous Xid code NVIDIA emits. It means the graphics engine raised an exception while executing work. The three candidate causes are a bad CUDA kernel (out-of-bounds access, illegal instruction), a driver bug, or genuine hardware degradation. Most isolated occurrences are software. The operational problem is telling which case you are in, because the correct response ranges from &amp;ldquo;file a bug against the application&amp;rdquo; to &amp;ldquo;RMA the GPU&amp;rdquo;.&lt;/p></description></item><item><title>NVIDIA Xid 31: GPU memory page fault (invalid address)</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-31-memory-page-fault/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-31-memory-page-fault/</guid><description>&lt;h1 id="nvidia-xid-31-gpu-memory-page-fault-invalid-address">NVIDIA Xid 31: GPU memory page fault (invalid address)&lt;/h1>
&lt;p>You found &lt;code>NVRM: Xid (PCI:0000:xx:00.x): 31, ...&lt;/code> in the kernel log, usually right after a training job or inference process died with &lt;code>CUDA error: an illegal memory access was encountered&lt;/code>. The GPU raised a memory page fault: some unit on the chip accessed a virtual address that was not mapped to valid GPU memory.&lt;/p>
&lt;p>The important thing to know up front: Xid 31 almost always means the application is wrong, not the GPU. Accessing freed GPU memory, bad pointer arithmetic, an out-of-bounds index in a kernel. These produce exactly this fault. The exception is when the same GPU produces Xid 31 across multiple unrelated applications. Then the suspicion flips to hardware.&lt;/p></description></item><item><title>NVIDIA Xid 43: GPU stopped processing (the GPU hang)</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-43-gpu-stopped-processing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-43-gpu-stopped-processing/</guid><description>&lt;h1 id="nvidia-xid-43-gpu-stopped-processing-the-gpu-hang">NVIDIA Xid 43: GPU stopped processing (the GPU hang)&lt;/h1>
&lt;p>You see &lt;code>NVRM: Xid (PCI:0000:XX:00.0): 43, ... GPU stopped processing&lt;/code> in dmesg, and the job on that GPU has gone silent. CUDA calls do not return. &lt;code>nvidia-smi&lt;/code> either hangs outright or returns a stale view of the card: utilization frozen at the last sampled value, memory allocation unchanged, power draw flat. The GPU has not disappeared from the bus the way it does with Xid 79. It is still enumerated and still answering the driver at some level, but the compute engine is wedged and nothing it was running will complete.&lt;/p></description></item><item><title>NVIDIA Xid 48: Double Bit ECC Error</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-48-double-bit-ecc/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-48-double-bit-ecc/</guid><description>&lt;h1 id="nvidia-xid-48-double-bit-ecc-error">NVIDIA Xid 48: Double Bit ECC Error&lt;/h1>
&lt;p>You found &lt;code>NVRM: Xid (PCI:....): 48, Double Bit ECC Error&lt;/code> in dmesg, a training job just died, or a monitoring alert fired on a new uncorrected ECC error. This error is unambiguous: data corruption has occurred in GPU memory, and whatever was running on that GPU when it happened cannot be trusted.&lt;/p>
&lt;p>The CUDA context that triggered the error is typically killed, recent outputs (weights, checkpoints, inference results) may be corrupt, and the GPU needs a reset or node reboot before it returns to clean service. There is no &amp;ldquo;watch and see&amp;rdquo; path. The correct response is to drain the GPU, verify whether the faulting memory was retired or remapped, validate or discard recent work, and schedule replacement.&lt;/p></description></item><item><title>NVIDIA Xid 63 and 64: ECC page retirement and row-remap events</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-63-64-page-retirement/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-63-64-page-retirement/</guid><description>&lt;h1 id="nvidia-xid-63-and-64-ecc-page-retirement-and-row-remap-events">NVIDIA Xid 63 and 64: ECC page retirement and row-remap events&lt;/h1>
&lt;p>You found a line like this in the kernel log:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-text" data-lang="text">&lt;span style="display:flex;">&lt;span>NVRM: Xid (PCI:0000:10:1c): 63, pid=1896, Row Remapper: New row marked for remapping, reset gpu to activate.
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Or the less welcome variant, Xid 64. Both come from the GPU&amp;rsquo;s memory self-repair machinery: ECC detected a bad region of HBM or GDDR memory, and the driver tried to permanently remove it from service. Xid 63 means the repair was recorded successfully. Xid 64 means the recording failed, which is a different and more serious situation.&lt;/p></description></item><item><title>NVIDIA Xid 74: NVLink error</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-74-nvlink-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-74-nvlink-error/</guid><description>&lt;h1 id="nvidia-xid-74-nvlink-error">NVIDIA Xid 74: NVLink error&lt;/h1>
&lt;p>You found &lt;code>NVRM: Xid 74&lt;/code> in &lt;code>dmesg&lt;/code>. The driver detected an error on an NVLink connection between this GPU and another GPU or an NVSwitch. The confusing part: nothing crashed. Training is still running. What actually happened is that the link&amp;rsquo;s error recovery machinery fired - CRC failures, packet replays, link recovery events - and the driver logged it.&lt;/p>
&lt;p>Xid 74 is the classic silent straggler. Retransmissions eat NVLink bandwidth, and collective operations like AllReduce synchronize at the speed of the slowest participant. Step time creeps up, throughput drops, and nobody gets paged because nothing &amp;ldquo;failed.&amp;rdquo; The job is just slower, and the GPU-hour bill is higher.&lt;/p></description></item><item><title>NVIDIA Xid 79: GPU has fallen off the bus</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-79-fallen-off-bus/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-79-fallen-off-bus/</guid><description>&lt;h1 id="nvidia-xid-79-gpu-has-fallen-off-the-bus">NVIDIA Xid 79: GPU has fallen off the bus&lt;/h1>
&lt;p>You will usually find this one by accident: a training job crashes with an opaque CUDA error, &lt;code>nvidia-smi&lt;/code> hangs or shows &lt;code>ERR!&lt;/code> for one device, and when you check &lt;code>dmesg&lt;/code> you find &lt;code>NVRM: Xid (PCI:0000:XX:00): 79, GPU has fallen off the bus&lt;/code>. The GPU is gone. Not throttled, not wedged, not holding stale contexts. The PCIe link between the host and the GPU has failed, and the device is no longer reachable at all.&lt;/p></description></item><item><title>NVIDIA Xid 92: high single-bit ECC error rate</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-92-high-sbe-rate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-92-high-sbe-rate/</guid><description>&lt;h1 id="nvidia-xid-92-high-single-bit-ecc-error-rate">NVIDIA Xid 92: high single-bit ECC error rate&lt;/h1>
&lt;p>You found a line like this in the kernel log:&lt;/p>
&lt;pre tabindex="0">&lt;code>NVRM: Xid (PCI:0000:3b:00.0): 92, pid=&amp;#39;&amp;lt;unknown&amp;gt;&amp;#39;, name=&amp;lt;unknown&amp;gt;, High single-bit ECC error rate
&lt;/code>&lt;/pre>&lt;p>Nothing crashed. No job failed. nvidia-smi still works, the GPU is still training or serving, and there is no Xid 48 or 79 demanding an immediate reboot. That is exactly the point of Xid 92: the driver has seen the correctable single-bit ECC error rate cross its internal threshold. Nothing is broken yet. Something will be.&lt;/p></description></item><item><title>NVIDIA Xid 94 and 95: contained vs uncontained ECC errors (Ampere+)</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-94-95-contained-uncontained-ecc/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-94-95-contained-uncontained-ecc/</guid><description>&lt;h1 id="nvidia-xid-94-and-95-contained-vs-uncontained-ecc-errors-ampere">NVIDIA Xid 94 and 95: contained vs uncontained ECC errors (Ampere+)&lt;/h1>
&lt;p>You are looking at a kernel log line like &lt;code>NVRM: Xid (PCI:0000:01:00): 94, pid=7194, Contained: ...&lt;/code> or &lt;code>...: 95, pid=7062, Uncontained: ...&lt;/code> on an A100, H100, or newer GPU, and you need to know two things fast: how bad is it, and what do you drain.&lt;/p>
&lt;p>Both Xid 94 and Xid 95 are uncorrectable (double-bit class) ECC error events. The distinction NVIDIA draws, starting with the A100, is containment: whether the corruption was confined to the faulting application context or may have spread across everything running on the GPU. That distinction drives the entire response.&lt;/p></description></item><item><title>NVIDIA Xid errors: reading NVRM Xid messages in the kernel log</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-xid-errors/</guid><description>&lt;h1 id="nvidia-xid-errors-reading-nvrm-xid-messages-in-the-kernel-log">NVIDIA Xid errors: reading NVRM Xid messages in the kernel log&lt;/h1>
&lt;p>You found a line like this in the kernel log, or a user reported a dead training job and you went looking:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-text" data-lang="text">&lt;span style="display:flex;">&lt;span>NVRM: Xid (PCI:0000:41:00): 79, pid=1432, name=python, GPU has fallen off the bus.
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Xid messages are the NVIDIA driver&amp;rsquo;s primary error-reporting channel. Every significant GPU fault, from a correctable memory event to a catastrophic PCIe link failure, is printed by the NVRM kernel module as an Xid code in the kernel log. Not in nvidia-smi output. Not in a sysfs file. In the log stream, alongside everything else the kernel has to say.&lt;/p></description></item><item><title>nvidia-smi hangs or is unresponsive: a wedged GPU or a stuck driver</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-nvidia-smi-hangs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-nvidia-smi-hangs/</guid><description>&lt;h1 id="nvidia-smi-hangs-or-is-unresponsive-a-wedged-gpu-or-a-stuck-driver">nvidia-smi hangs or is unresponsive: a wedged GPU or a stuck driver&lt;/h1>
&lt;p>You run &lt;code>nvidia-smi&lt;/code> and nothing comes back. No output, no error, no exit code. Ctrl-C does nothing, and &lt;code>kill -9&lt;/code> does nothing either. This is a different class of failure from &amp;ldquo;NVIDIA-SMI has failed because it couldn&amp;rsquo;t communicate with the NVIDIA driver&amp;rdquo;, where the driver at least answers and tells you it is broken. A hang means the driver is stuck inside a kernel call, usually waiting on GPU hardware that will never respond.&lt;/p></description></item><item><title>NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-couldnt-communicate-with-driver/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-couldnt-communicate-with-driver/</guid><description>&lt;h1 id="nvidia-smi-has-failed-because-it-couldnt-communicate-with-the-nvidia-driver">NVIDIA-SMI has failed because it couldn&amp;rsquo;t communicate with the NVIDIA driver&lt;/h1>
&lt;p>You run &lt;code>nvidia-smi&lt;/code> and get:&lt;/p>
&lt;pre tabindex="0">&lt;code>NVIDIA-SMI has failed because it couldn&amp;#39;t communicate with the NVIDIA driver.
Make sure that the latest NVIDIA driver is installed and running.
&lt;/code>&lt;/pre>&lt;p>This error means the userspace side of the driver (NVML, the library nvidia-smi talks to) cannot reach the kernel side. Either the &lt;code>nvidia&lt;/code> kernel module is not loaded, it is loaded but built for a different kernel than the one running, the kernel refused to load it, or the GPU itself is gone from the PCIe bus. nvidia-smi exits non-zero and prints the message to stderr; the error string is the reliable signal, not the exact exit code.&lt;/p></description></item><item><title>NVMe ASPM latency spikes: PCIe power states adding first-request latency</title><link>https://www.netdata.cloud/guides/nvme/nvme-aspm-latency-spikes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-aspm-latency-spikes/</guid><description>&lt;h1 id="nvme-aspm-latency-spikes-pcie-power-states-adding-first-request-latency">NVMe ASPM latency spikes: PCIe power states adding first-request latency&lt;/h1>
&lt;p>You are chasing tail latency on an NVMe-backed service. The p99 graph shows sporadic spikes: tens of microseconds added to reads that should complete in well under 100us. The spikes do not correlate with load, temperature, media errors, queue depth, or garbage collection. SMART is clean. The block layer looks clean. The one pattern you can find: the slow I/Os tend to be the first request after the device sat idle.&lt;/p></description></item><item><title>NVMe available spare below threshold: critical warning bit 0 and end-of-life wear</title><link>https://www.netdata.cloud/guides/nvme/nvme-available-spare-below-threshold/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-available-spare-below-threshold/</guid><description>&lt;h1 id="nvme-available-spare-below-threshold-critical-warning-bit-0-and-end-of-life-wear">NVMe available spare below threshold: critical warning bit 0 and end-of-life wear&lt;/h1>
&lt;p>Your monitoring fired on an NVMe drive: critical warning bit 0 is set, or &lt;code>available_spare&lt;/code> has dropped to (or below) the vendor threshold. The drive still works. Reads and writes complete, latency looks fine, nothing in the application layer is complaining. That is exactly what makes this signal easy to ignore and expensive to ignore.&lt;/p>
&lt;p>Bit 0 means the controller&amp;rsquo;s pool of spare NAND blocks, the reserve it uses to transparently replace failed cells, has fallen below the safety margin the vendor baked into the firmware. The drive is telling you it is approaching end-of-life. It is not telling you it has failed.&lt;/p></description></item><item><title>NVMe Available Spare below threshold: the spare block pool is running out</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-nvme-available-spare/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-nvme-available-spare/</guid><description>&lt;h1 id="nvme-available-spare-below-threshold-the-spare-block-pool-is-running-out">NVMe Available Spare below threshold: the spare block pool is running out&lt;/h1>
&lt;p>When &lt;code>smartctl -A /dev/nvme0n1&lt;/code> reports Available Spare below the Available Spare Threshold, the drive&amp;rsquo;s internal spare block pool is running low. The controller is reporting less reserved NAND capacity than it considers safe for continued reliable operation.&lt;/p>
&lt;p>Available Spare is the percentage of reserved NAND blocks remaining for replacing worn or failed blocks. It starts at 100% and decreases monotonically as the drive consumes spares. When it reaches 0%, the next bad block causes permanent data loss for that block&amp;rsquo;s data: the drive has no spare to remap to.&lt;/p></description></item><item><title>NVMe available spare declining: watching the wear trajectory before the threshold</title><link>https://www.netdata.cloud/guides/nvme/nvme-available-spare-declining/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-available-spare-declining/</guid><description>&lt;h1 id="nvme-available-spare-declining-watching-the-wear-trajectory-before-the-threshold">NVMe available spare declining: watching the wear trajectory before the threshold&lt;/h1>
&lt;p>The drive reported available_spare at 100% for two years. This quarter it reads 94%, and last month it was 96%. Nothing has alerted: critical_warning is zero, I/O is clean, latency is normal. The question is whether you are watching normal aging or the early edge of a failure curve. The answer is almost never in the current value. It is in the rate of change.&lt;/p></description></item><item><title>NVMe controller reset loop: repeated resets from a firmware hang</title><link>https://www.netdata.cloud/guides/nvme/nvme-controller-reset-loop/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-controller-reset-loop/</guid><description>&lt;h1 id="nvme-controller-reset-loop-repeated-resets-from-a-firmware-hang">NVMe controller reset loop: repeated resets from a firmware hang&lt;/h1>
&lt;p>Your kernel log shows the same cycle over and over: an I/O command stalls until the 30-second &lt;code>io_timeout&lt;/code> fires, the NVMe driver resets the controller, I/O briefly recovers, then the whole thing hangs again. Each cycle costs your applications 5 to 30 seconds of stalled I/O. Databases time out, requests fail, clusters rebalance, and minutes later it happens again.&lt;/p>
&lt;p>This is the controller firmware hang loop. The firmware hits a bug triggered by a specific command sequence, queue depth, I/O size, or power state transition. The kernel detects the timeout, resets the controller, and replays the in-flight commands. The same trigger recurs, so the controller hangs again. The loop can continue indefinitely.&lt;/p></description></item><item><title>NVMe controller state not live: reading resetting, deleting, and dead from sysfs</title><link>https://www.netdata.cloud/guides/nvme/nvme-controller-state-not-live/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-controller-state-not-live/</guid><description>&lt;h1 id="nvme-controller-state-not-live-reading-resetting-deleting-and-dead-from-sysfs">NVMe controller state not live: reading resetting, deleting, and dead from sysfs&lt;/h1>
&lt;p>The fastest way to answer &amp;ldquo;is this NVMe device actually usable right now&amp;rdquo; is a single file: &lt;code>/sys/class/nvme/nvmeX/state&lt;/code>. It is the kernel NVMe driver&amp;rsquo;s own view of the controller, exposed as one word: &lt;code>live&lt;/code>, &lt;code>resetting&lt;/code>, &lt;code>connecting&lt;/code>, &lt;code>deleting&lt;/code>, &lt;code>dead&lt;/code>, or &lt;code>new&lt;/code>. No SMART parsing, no log pages, no vendor tooling. If the state is not &lt;code>live&lt;/code>, I/O to that device is stalled, failing, or about to.&lt;/p></description></item><item><title>NVMe Critical Warning bits: decoding the health-log bitmask</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-nvme-critical-warning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-nvme-critical-warning/</guid><description>&lt;h1 id="nvme-critical-warning-bits-decoding-the-health-log-bitmask">NVMe Critical Warning bits: decoding the health-log bitmask&lt;/h1>
&lt;p>The NVMe Critical Warning byte is the single most important health signal on an NVMe drive. Unlike ATA SMART&amp;rsquo;s scattered attribute IDs, this one byte consolidates five distinct critical conditions into a compact bitmask. Most operators who find this article have just seen a non-zero value from smartctl and need to know which bit is set and what to do about it.&lt;/p></description></item><item><title>NVMe critical_warning is nonzero: decoding the SMART critical warning bitmask</title><link>https://www.netdata.cloud/guides/nvme/nvme-critical-warning-nonzero/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-critical-warning-nonzero/</guid><description>&lt;h1 id="nvme-critical_warning-is-nonzero-decoding-the-smart-critical-warning-bitmask">NVMe critical_warning is nonzero: decoding the SMART critical warning bitmask&lt;/h1>
&lt;p>Your monitoring fired because &lt;code>critical_warning&lt;/code> in the NVMe SMART log is nonzero. Or you ran &lt;code>nvme smart-log /dev/nvme0&lt;/code> during an investigation and saw something like &lt;code>critical_warning : 0x04&lt;/code> where you expected &lt;code>0x00&lt;/code>. You need two things fast: which bit is set, and how bad it is.&lt;/p>
&lt;p>The most common mistake here is treating &lt;code>critical_warning != 0&lt;/code> as one alert with one severity. It is a bitmask of six independent conditions with wildly different severity. Bit 3 means the drive has gone read-only and is refusing writes: a page-right-now outage. Bit 0 means spare capacity is below the vendor threshold: a procurement ticket, not a 3 a.m. incident. Bit 1 may be a transient thermal event during a backup run that clears itself.&lt;/p></description></item><item><title>NVMe device disappeared: nvme0: Removing and a drive that fell off the PCIe bus</title><link>https://www.netdata.cloud/guides/nvme/nvme-device-removed-disappeared/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-device-removed-disappeared/</guid><description>&lt;h1 id="nvme-device-disappeared-nvme0-removing-and-a-drive-that-fell-off-the-pcie-bus">NVMe device disappeared: nvme0: Removing and a drive that fell off the PCIe bus&lt;/h1>
&lt;p>You are looking at a host where an NVMe drive was present and is now gone. The kernel log shows some version of this sequence: a &lt;code>pcieport: AER: Uncorrectable (Fatal)&lt;/code> line, possibly followed by &lt;code>nvme nvme0: controller is down; will reset: CSTS=0xffffffff&lt;/code>, then &lt;code>nvme nvme0: Removing&lt;/code> (or &lt;code>nvme nvme0: Removing after probe failure status: -19&lt;/code>), and often &lt;code>nvme0n1: detected capacity change from X to 0&lt;/code>. After that, &lt;code>/dev/nvme0n1&lt;/code> no longer exists and &lt;code>/sys/class/nvme/nvme0&lt;/code> is empty.&lt;/p></description></item><item><title>NVMe devices</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/nvme-devices/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/nvme-devices/</guid><description/></item><item><title>NVMe drive in read-only mode: critical warning bit 3 and rejected writes</title><link>https://www.netdata.cloud/guides/nvme/nvme-read-only-mode/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-read-only-mode/</guid><description>&lt;h1 id="nvme-drive-in-read-only-mode-critical-warning-bit-3-and-rejected-writes">NVMe drive in read-only mode: critical warning bit 3 and rejected writes&lt;/h1>
&lt;p>Your monitoring fires a PAGE: the NVMe device has set &lt;code>critical_warning&lt;/code> bit 3 (0x08), meaning the drive has autonomously placed its media in read-only mode. Every write command the host sends is now rejected by the drive&amp;rsquo;s firmware. Reads still work. Writes do not.&lt;/p>
&lt;p>This is not transient, not load-dependent, and not cleared by a reboot or controller reset. The drive has decided, based on its own internal assessment, that it can no longer safely accept writes, typically because it has run out of spare NAND blocks to remap failing cells into. It has switched itself into data-preservation mode so you can get your data off before it dies completely.&lt;/p></description></item><item><title>NVMe endurance runway: projecting time-to-replacement from wear signals</title><link>https://www.netdata.cloud/guides/nvme/nvme-endurance-runway-planning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-endurance-runway-planning/</guid><description>&lt;h1 id="nvme-endurance-runway-projecting-time-to-replacement-from-wear-signals">NVMe endurance runway: projecting time-to-replacement from wear signals&lt;/h1>
&lt;p>NVMe wear-out is one of the few storage failures that usually announces itself months in advance. The announcement is quiet: a monotonically rising &lt;code>percentage_used&lt;/code>, a slow drift in &lt;code>available_spare&lt;/code>, maybe the first few &lt;code>media_errors&lt;/code>. If you only look at current values, you find out when the drive crosses a threshold. If you trend the rates, you can procure before the drive becomes an incident.&lt;/p></description></item><item><title>NVMe error log entries growing: num_err_log_entries beyond media errors</title><link>https://www.netdata.cloud/guides/nvme/nvme-error-log-entries-increasing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-error-log-entries-increasing/</guid><description>&lt;h1 id="nvme-error-log-entries-growing-num_err_log_entries-beyond-media-errors">NVMe error log entries growing: num_err_log_entries beyond media errors&lt;/h1>
&lt;p>Your NVMe alert fired because &lt;code>num_err_log_entries&lt;/code> is climbing. You pull the SMART log, and &lt;code>media_errors&lt;/code> is zero. The drive looks healthy, but something is writing error entries at a steady, sometimes alarming, rate.&lt;/p>
&lt;p>&lt;code>num_err_log_entries&lt;/code> is a superset counter: it counts every error the controller records, including admin command errors, I/O command errors, internal controller errors, and thermal events. &lt;code>media_errors&lt;/code> counts only data integrity failures against the NAND. A rising &lt;code>num_err_log_entries&lt;/code> with flat &lt;code>media_errors&lt;/code> points away from failing flash and toward firmware, driver, or command-level issues. In a large share of cases, it points at your own monitoring stack.&lt;/p></description></item><item><title>NVMe firmware version tracking: catching the bugs vendors do not advertise</title><link>https://www.netdata.cloud/guides/nvme/nvme-firmware-version-tracking/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-firmware-version-tracking/</guid><description>&lt;h1 id="nvme-firmware-version-tracking-catching-the-bugs-vendors-do-not-advertise">NVMe firmware version tracking: catching the bugs vendors do not advertise&lt;/h1>
&lt;p>The NVMe controller is an embedded computer running proprietary firmware that owns your data path: the flash translation layer, wear leveling, garbage collection, error correction, and the queue interface to the host. When that firmware has a bug, the symptoms show up as reset loops, erratic latency, premature wear, or in the worst cases silent data loss. Controller firmware bugs are more common than vendors admit, and they are usually triggered by specific command sequences or power-state transitions, so two identical drives on different firmware revisions can behave like different hardware.&lt;/p></description></item><item><title>NVMe high I/O latency: reading block-layer latency and the outliers that matter</title><link>https://www.netdata.cloud/guides/nvme/nvme-high-io-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-high-io-latency/</guid><description>&lt;h1 id="nvme-high-io-latency-reading-block-layer-latency-and-the-outliers-that-matter">NVMe high I/O latency: reading block-layer latency and the outliers that matter&lt;/h1>
&lt;p>Latency complaints about NVMe arrive as &amp;ldquo;the database is slow&amp;rdquo; or &amp;ldquo;queries time out randomly.&amp;rdquo; By the time the report reaches you, the mean latency from &lt;code>iostat&lt;/code> often looks fine, because NVMe mean latency is almost always fine. The damage comes from the tail: the occasional 10 ms or 1 s outlier that breaks a request deadline and cascades into application timeouts.&lt;/p></description></item><item><title>NVMe Media and Data Integrity Errors incrementing: confirmed NAND corruption</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-nvme-media-data-integrity-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-nvme-media-data-integrity-errors/</guid><description>&lt;h1 id="nvme-media-and-data-integrity-errors-incrementing-confirmed-nand-corruption">NVMe Media and Data Integrity Errors incrementing: confirmed NAND corruption&lt;/h1>
&lt;p>The NVMe SMART/Health Information log exposes a field called &amp;ldquo;Media and Data Integrity Errors.&amp;rdquo; When this value is zero, the controller has never returned data that failed integrity verification. When it increments, the controller detected an unrecovered data integrity error on data retrieved from NAND flash. These are errors such as uncorrectable ECC failures, CRC checksum failures, or LBA tag mismatches that exceeded the controller&amp;rsquo;s internal correction layer.&lt;/p></description></item><item><title>NVMe media placed in read-only mode: Critical Warning bit 3</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-nvme-read-only-mode/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-nvme-read-only-mode/</guid><description>&lt;h1 id="nvme-media-placed-in-read-only-mode-critical-warning-bit-3">NVMe media placed in read-only mode: Critical Warning bit 3&lt;/h1>
&lt;p>When smartctl reports &lt;code>Critical Warning: 0x08&lt;/code> on an NVMe drive, the controller has locked the media into read-only mode. The drive is refusing all writes to protect existing data. This is a hardware-enforced decision made by the drive firmware. Reboots, firmware updates, and format commands will not clear it.&lt;/p>
&lt;p>The most common cause is endurance exhaustion: the spare block pool is depleted and the controller locks writes rather than risk corruption. Read-only transitions can also result from firmware bugs, thermal events, or electrical issues. An unexpected transition on a drive well below its rated endurance warrants forensic review.&lt;/p></description></item><item><title>NVMe media_errors increasing: uncorrectable data-integrity errors on NAND</title><link>https://www.netdata.cloud/guides/nvme/nvme-media-errors-increasing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-media-errors-increasing/</guid><description>&lt;h1 id="nvme-media_errors-increasing-uncorrectable-data-integrity-errors-on-nand">NVMe media_errors increasing: uncorrectable data-integrity errors on NAND&lt;/h1>
&lt;p>Your monitoring shows the &lt;code>media_errors&lt;/code> counter on an NVMe drive going up. In &lt;code>nvme smart-log&lt;/code> output the field is labelled &lt;code>media_errors&lt;/code> (or &amp;ldquo;Media and Data Integrity Errors&amp;rdquo; in some nvme-cli versions), and it is the one SMART field you should never explain away: each increment is a read or write where the controller could not maintain data integrity even after its internal ECC and retry mechanisms were exhausted.&lt;/p></description></item><item><title>NVMe Monitoring</title><link>https://www.netdata.cloud/monitoring-101/nvme-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/nvme-monitoring/</guid><description>&lt;h2 id="nvme-monitoring">NVMe Monitoring&lt;/h2>
&lt;h3 id="what-is-nvme">What Is NVMe?&lt;/h3>
&lt;p>NVMe, or Non-Volatile Memory Express, is a storage protocol designed to capitalize on the low latency and internal parallelism of solid-state drives (SSDs). It is widely embraced for boosting storage performance and is crucial for applications demanding high data throughput and speed. NVMe allows for direct CPU communication, drastically reducing data transfer overhead compared to legacy protocols like SATA.&lt;/p>
&lt;h3 id="monitoring-nvme-with-netdata">Monitoring NVMe With Netdata&lt;/h3>
&lt;p>Monitoring NVMe is essential for maintaining optimal performance, ensuring data integrity, and identifying potential hardware issues. Netdata provides a robust set of monitoring tools tailored for NVMe devices that allow you to track their health and performance metrics in real time.&lt;/p></description></item><item><title>NVMe monitoring checklist: the signals every production SSD needs</title><link>https://www.netdata.cloud/guides/nvme/nvme-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-monitoring-checklist/</guid><description>&lt;h1 id="nvme-monitoring-checklist-the-signals-every-production-ssd-needs">NVMe monitoring checklist: the signals every production SSD needs&lt;/h1>
&lt;p>Most NVMe monitoring collects a handful of SMART counters and stops there. That works until a controller hangs without touching SMART, a PCIe link silently retrains to half speed, or a drive sets a critical warning bit that your single blanket alert treats as noise. NVMe failures rarely arrive as a clean &amp;ldquo;disk error.&amp;rdquo; They arrive as latency, throttling, stalled queues, or a device that vanishes from the bus.&lt;/p></description></item><item><title>NVMe monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/nvme/nvme-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-monitoring-maturity-model/</guid><description>&lt;h1 id="nvme-monitoring-maturity-model-from-survival-to-expert">NVMe monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most teams discover their NVMe monitoring gaps during an incident: a drive throttling silently at 80 C while every dashboard shows green, or a Gen4 x4 device running at Gen3 x2 for weeks with zero errors. NVMe fails differently from SATA and SAS. It has a PCIe transport layer with its own error reporting, a controller running a flash translation layer that causes latency variance invisible to block-layer averages, and wear signals that only matter as rates of change. Generic disk monitoring misses all of this.&lt;/p></description></item><item><title>NVMe namespaces and the 512e vs 4K sector-size trap</title><link>https://www.netdata.cloud/guides/nvme/nvme-namespace-sector-size/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-namespace-sector-size/</guid><description>&lt;h1 id="nvme-namespaces-and-the-512e-vs-4k-sector-size-trap">NVMe namespaces and the 512e vs 4K sector-size trap&lt;/h1>
&lt;p>A namespace formatted for 512-byte logical blocks (512e) while the workload, filesystem, or encryption layer assumes 4K blocks produces no errors, no SMART warnings, and no kernel log entries. It just makes the controller do more work per host write than it should. This is one of the genuinely silent NVMe failure modes: everything looks healthy while write amplification quietly burns endurance and adds latency.&lt;/p></description></item><item><title>NVMe NVM subsystem reliability degraded: critical warning bit 2</title><link>https://www.netdata.cloud/guides/nvme/nvme-nvm-subsystem-reliability-degraded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-nvm-subsystem-reliability-degraded/</guid><description>&lt;h1 id="nvme-nvm-subsystem-reliability-degraded-critical-warning-bit-2">NVMe NVM subsystem reliability degraded: critical warning bit 2&lt;/h1>
&lt;p>Your monitoring just reported &lt;code>critical_warning : 0x4&lt;/code> on an NVMe device, or smartctl printed &amp;ldquo;SMART overall-health self-assessment test result: FAILED! - NVM subsystem reliability has been degraded.&amp;rdquo; The drive is still serving I/O. Nothing in the kernel log looks catastrophic. Is the drive dying, or is this noise?&lt;/p>
&lt;p>It depends on what else the SMART log says. Critical warning bit 2 is the vaguest bit in the critical warning byte. The specification language is broad: the controller has detected significant media-related errors, or some internal error, that degrades NVM subsystem reliability. What counts as &amp;ldquo;significant&amp;rdquo; and &amp;ldquo;degraded&amp;rdquo; is left to the vendor, and vendors interpret it very differently. Some drives set bit 2 only when the flash is genuinely failing. Others, notably several Samsung consumer models, set it preemptively the moment Percentage Used crosses 100 percent, which is a warranty-consumed signal, not a failure signal.&lt;/p></description></item><item><title>NVMe NVM subsystem reliability degraded: Critical Warning bit 2</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-nvme-reliability-degraded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-nvme-reliability-degraded/</guid><description>&lt;h1 id="nvme-nvm-subsystem-reliability-degraded-critical-warning-bit-2">NVMe NVM subsystem reliability degraded: Critical Warning bit 2&lt;/h1>
&lt;p>&lt;code>smartctl -H&lt;/code> reports FAILED on an NVMe drive that otherwise seems fine. I/O is normal, Available Spare is healthy, and there are no media errors. The Critical Warning byte is &lt;code>0x04&lt;/code>.&lt;/p>
&lt;p>Bit 2 (0x04) in the NVMe Critical Warning byte means &amp;ldquo;NVM subsystem reliability has been degraded.&amp;rdquo; Unlike bit 0 (spare below threshold), bit 1 (temperature), or bit 3 (read-only mode), bit 2 has no single metric that directly explains why firmware set it. The NVMe specification defines it as triggered by &amp;ldquo;significant media related errors or any internal error that degrades NVM subsystem reliability,&amp;rdquo; but leaves the exact conditions vendor-defined.&lt;/p></description></item><item><title>nvme nvme0: I/O timeout, Resetting controller: what an NVMe controller reset means</title><link>https://www.netdata.cloud/guides/nvme/nvme-controller-reset-timeout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-controller-reset-timeout/</guid><description>&lt;h1 id="nvme-nvme0-io-timeout-resetting-controller-what-an-nvme-controller-reset-means">nvme nvme0: I/O timeout, Resetting controller: what an NVMe controller reset means&lt;/h1>
&lt;p>You are here because &lt;code>dmesg&lt;/code> or &lt;code>journalctl -k&lt;/code> shows a sequence like this:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-text" data-lang="text">&lt;span style="display:flex;">&lt;span>nvme nvme0: I/O 24 QID 3 timeout
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>nvme nvme0: Abort status: 0x0
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>nvme nvme0: Resetting controller
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The NVMe driver submitted a command, the controller did not complete it within the I/O timeout (30 seconds by default), the driver tried to abort the command, the abort did not resolve things, and the driver forced a full controller reset. During the reset-recovery cycle, which typically takes 5 to 30 seconds, all I/O to that device stalls. In-flight commands are replayed after the reset completes.&lt;/p></description></item><item><title>NVMe PCIe AER errors: correctable and uncorrectable transport-layer faults</title><link>https://www.netdata.cloud/guides/nvme/nvme-pcie-aer-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-pcie-aer-errors/</guid><description>&lt;h1 id="nvme-pcie-aer-errors-correctable-and-uncorrectable-transport-layer-faults">NVMe PCIe AER errors: correctable and uncorrectable transport-layer faults&lt;/h1>
&lt;p>You opened the kernel log because something felt off, and found lines like this repeating:&lt;/p>
&lt;pre tabindex="0">&lt;code>pcieport 0000:00:1b.0: AER: Corrected error received: 0000:01:00.0
nvme 0000:01:00.0: PCIe Bus Error: severity=Corrected, type=Physical Layer, (Receiver ID)
&lt;/code>&lt;/pre>&lt;p>That is PCI Express Advanced Error Reporting (AER), the transport layer underneath NVMe telling you the link between the root port and the drive is producing errors. These errors happen below the NVMe protocol: the drive&amp;rsquo;s SMART data can look perfectly healthy while the PCIe link is retransmitting constantly, quietly adding latency to every I/O.&lt;/p></description></item><item><title>NVMe PCIe link degraded: current link speed and width below maximum</title><link>https://www.netdata.cloud/guides/nvme/nvme-pcie-link-speed-width-degraded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-pcie-link-speed-width-degraded/</guid><description>&lt;h1 id="nvme-pcie-link-degraded-current-link-speed-and-width-below-maximum">NVMe PCIe link degraded: current link speed and width below maximum&lt;/h1>
&lt;p>An NVMe drive negotiates its PCIe link at boot, at hot-plug, and after some error events. When negotiation lands below the device&amp;rsquo;s capability, nothing breaks. The drive keeps serving I/O, SMART stays clean, and the filesystem sees no errors. The only symptom is a lower bandwidth ceiling: a Gen4 x4 drive running at Gen3 x2 delivers roughly one quarter of its rated throughput, and every application on top of it just runs slower.&lt;/p></description></item><item><title>NVMe percentage used at 100%: reading the endurance-consumed estimate</title><link>https://www.netdata.cloud/guides/nvme/nvme-percentage-used-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-percentage-used-high/</guid><description>&lt;h1 id="nvme-percentage-used-at-100-reading-the-endurance-consumed-estimate">NVMe percentage used at 100%: reading the endurance-consumed estimate&lt;/h1>
&lt;p>Your monitoring just told you an NVMe drive is at 100% &amp;ldquo;percentage used&amp;rdquo;, or 123%, or 167%. The number looks like a fuel gauge hitting empty, and the instinct is to treat it as an emergency. It is not one. &lt;code>percentage_used&lt;/code> is a vendor estimate of rated endurance consumed, the NVMe specification explicitly allows it to exceed 100, and drives routinely operate past 100% for months or years. It is also monotonic: it never goes back down.&lt;/p></description></item><item><title>NVMe Percentage Used at or above 100%: rated endurance consumed</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-nvme-percentage-used/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-nvme-percentage-used/</guid><description>&lt;h1 id="nvme-percentage-used-at-or-above-100-rated-endurance-consumed">NVMe Percentage Used at or above 100%: rated endurance consumed&lt;/h1>
&lt;p>The NVMe SMART/Health Information Log (Log Page 02h) includes a field called Percentage Used that estimates how much of a drive&amp;rsquo;s rated write endurance has been consumed. At 100%, the drive has used up its warranted Total Bytes Written (TBW). Above 100%, it is operating beyond its manufacturer endurance rating.&lt;/p>
&lt;p>This is not a failure signal. The NVMe specification explicitly allows values up to 255%, and drives routinely continue working past their rated endurance. A drive reporting 150% Percentage Used with 90% Available Spare and zero Media Errors is likely fine. Percentage Used is a planning signal: it tells you where the drive is in its endurance lifecycle so you can sequence replacement before the actual failure indicators (Available Spare exhaustion and Media Errors) force an emergency.&lt;/p></description></item><item><title>NVMe power cycles and power-on hours: fleet age and lifecycle tracking</title><link>https://www.netdata.cloud/guides/nvme/nvme-power-cycles-power-on-hours/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-power-cycles-power-on-hours/</guid><description>&lt;h1 id="nvme-power-cycles-and-power-on-hours-fleet-age-and-lifecycle-tracking">NVMe power cycles and power-on hours: fleet age and lifecycle tracking&lt;/h1>
&lt;p>&lt;code>power_cycles&lt;/code> and &lt;code>power_on_hours&lt;/code> become operational when you need to know how old a drive really is, whether it has been losing power independently of the OS, and which fleet cohort is approaching replacement together.&lt;/p>
&lt;p>They are not failure alerts by themselves. Treat them as the lifecycle ledger for local PCIe-attached NVMe devices: &lt;code>power_cycles&lt;/code> counts full device power-off and power-on events, not NVMe controller resets, while &lt;code>power_on_hours&lt;/code> counts electrical on-time, not active I/O time. For controller-reported I/O occupancy, use &lt;code>controller_busy_time&lt;/code>.&lt;/p></description></item><item><title>NVMe power-loss protection: knowing whether your drive has PLP at all</title><link>https://www.netdata.cloud/guides/nvme/nvme-power-loss-protection-plp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-power-loss-protection-plp/</guid><description>&lt;h1 id="nvme-power-loss-protection-knowing-whether-your-drive-has-plp-at-all">NVMe power-loss protection: knowing whether your drive has PLP at all&lt;/h1>
&lt;p>Power-loss protection (PLP) is the difference between an unsafe shutdown being a non-event and an unsafe shutdown being silent data corruption. Enterprise NVMe drives carry capacitors that hold the controller up long enough to flush the volatile write cache to NAND when power drops. Consumer drives do not. On a drive without PLP, every power loss, PSU failure, or hard reset risks losing writes that the kernel and the application already believe are durable.&lt;/p></description></item><item><title>NVMe queue depth saturation: command slots, io_timeout, and deep queues</title><link>https://www.netdata.cloud/guides/nvme/nvme-io-queue-depth-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-io-queue-depth-saturation/</guid><description>&lt;h1 id="nvme-queue-depth-saturation-command-slots-io_timeout-and-deep-queues">NVMe queue depth saturation: command slots, io_timeout, and deep queues&lt;/h1>
&lt;p>You are looking at a host where NVMe latency has climbed, iostat shows the device busy, and nothing is erroring. No media errors, no kernel I/O error lines, SMART looks clean. The question is whether the drive is simply working as hard as it can, or whether something is wrong inside it. Queue depth is the signal that separates those two cases.&lt;/p></description></item><item><title>NVMe sanitize, format, and secure erase: expected events versus red flags</title><link>https://www.netdata.cloud/guides/nvme/nvme-sanitize-format-secure-erase/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-sanitize-format-secure-erase/</guid><description>&lt;h1 id="nvme-sanitize-format-and-secure-erase-expected-events-versus-red-flags">NVMe sanitize, format, and secure erase: expected events versus red flags&lt;/h1>
&lt;p>Sanitize, Format NVM, and secure erase are the three ways an NVMe controller destroys data on its own media. They are admin commands executed inside controller firmware, and once started they are largely outside the host&amp;rsquo;s control. In a healthy fleet they appear during commissioning (occasionally) and during decommission or repurpose. Any other occurrence is a security or integrity event.&lt;/p></description></item><item><title>NVMe silent data degradation: media errors, reliability bit, and failing cold reads</title><link>https://www.netdata.cloud/guides/nvme/nvme-silent-data-degradation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-silent-data-degradation/</guid><description>&lt;h1 id="nvme-silent-data-degradation-media-errors-reliability-bit-and-failing-cold-reads">NVMe silent data degradation: media errors, reliability bit, and failing cold reads&lt;/h1>
&lt;p>The drive is still there. It answers I/O, the filesystem is mounted, latency looks mostly normal, and nothing in &lt;code>dmesg&lt;/code> is screaming. But &lt;code>media_errors&lt;/code> has been climbing for days, &lt;code>critical_warning&lt;/code> now shows bit 2 set, and a few reads per hour are inexplicably slow. A backup verification job just failed a checksum on a file nobody has written to in months.&lt;/p></description></item><item><title>NVMe temperature threshold exceeded: critical warning bit 1, WCTEMP, and CCTEMP</title><link>https://www.netdata.cloud/guides/nvme/nvme-temperature-threshold-exceeded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-temperature-threshold-exceeded/</guid><description>&lt;h1 id="nvme-temperature-threshold-exceeded-critical-warning-bit-1-wctemp-and-cctemp">NVMe temperature threshold exceeded: critical warning bit 1, WCTEMP, and CCTEMP&lt;/h1>
&lt;p>Your monitoring fired on an NVMe drive: &lt;code>critical_warning&lt;/code> is non-zero, and after decoding the bitmask you find bit 1 set (value &lt;code>0x02&lt;/code>). The drive is reporting that its composite temperature has crossed a vendor-defined over-temperature threshold. The question that matters is whether this is a transient self-protecting throttle under heavy load, or a sustained thermal condition that is getting worse.&lt;/p></description></item><item><title>NVMe thermal management transitions: TMT1 and TMT2 as an early throttle signal</title><link>https://www.netdata.cloud/guides/nvme/nvme-thermal-management-transitions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-thermal-management-transitions/</guid><description>&lt;h1 id="nvme-thermal-management-transitions-tmt1-and-tmt2-as-an-early-throttle-signal">NVMe thermal management transitions: TMT1 and TMT2 as an early throttle signal&lt;/h1>
&lt;p>An NVMe drive that is thermally throttling rarely looks broken. There are no I/O errors, no kernel messages, no failed commands. Throughput drifts down, latency drifts up, and the application team files a ticket about &amp;ldquo;the database being slow.&amp;rdquo; By the time &lt;code>critical_warning&lt;/code> bit 1 (temperature threshold exceeded) is set, the drive has already been protecting itself for a while.&lt;/p></description></item><item><title>NVMe thermal throttling: the drive runs hot and performance quietly drops</title><link>https://www.netdata.cloud/guides/nvme/nvme-thermal-throttling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-thermal-throttling/</guid><description>&lt;h1 id="nvme-thermal-throttling-the-drive-runs-hot-and-performance-quietly-drops">NVMe thermal throttling: the drive runs hot and performance quietly drops&lt;/h1>
&lt;p>The database is slow. CPU utilization is low, the query plan has not changed, and there are no I/O errors in &lt;code>dmesg&lt;/code>, no media errors in SMART, no failed disk anywhere. Yet p99 query latency has doubled and IOPS are half of what they were last week.&lt;/p>
&lt;p>This is the classic NVMe thermal throttling incident. The controller detected that it was running too hot and quietly reduced its own performance to protect itself. There is no error, no log entry, no kernel message. The drive simply gets slower until it reaches thermal equilibrium, and it gets faster again when the load drops. If you are only watching error counters and application logs, you will never see it.&lt;/p></description></item><item><title>NVMe thermal throttling: throughput dropping with no errors</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-nvme-thermal-throttling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-nvme-thermal-throttling/</guid><description>&lt;h1 id="nvme-thermal-throttling-throughput-dropping-with-no-errors">NVMe thermal throttling: throughput dropping with no errors&lt;/h1>
&lt;p>NVMe throughput drops during sustained I/O. No I/O errors in dmesg, no command timeouts, SMART health says PASSED, Media and Data Integrity Errors at zero. The workload has not changed, but latency is up and throughput is down by 30, 50, or more percent. Hours later, performance recovers on its own. The pattern repeats: degradation during peak load, recovery during quiet periods or overnight.&lt;/p></description></item><item><title>NVMe throughput collapse: busy controller, low IOPS, and internal contention</title><link>https://www.netdata.cloud/guides/nvme/nvme-throughput-collapse/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-throughput-collapse/</guid><description>&lt;h1 id="nvme-throughput-collapse-busy-controller-low-iops-and-internal-contention">NVMe throughput collapse: busy controller, low IOPS, and internal contention&lt;/h1>
&lt;p>The drive is flat out. Whatever metric you look at says the controller is working constantly. And yet the application is starving: IOPS are a fraction of what the drive is rated for, write throughput has fallen off a cliff, and latency is climbing. Nothing in the kernel log looks broken. No media errors. No resets. The drive is busy, and it is delivering almost nothing.&lt;/p></description></item><item><title>NVMe unsafe shutdowns increasing: power-loss events and silent corruption risk</title><link>https://www.netdata.cloud/guides/nvme/nvme-unsafe-shutdowns-increasing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-unsafe-shutdowns-increasing/</guid><description>&lt;h1 id="nvme-unsafe-shutdowns-increasing-power-loss-events-and-silent-corruption-risk">NVMe unsafe shutdowns increasing: power-loss events and silent corruption risk&lt;/h1>
&lt;p>The &lt;code>unsafe_shutdowns&lt;/code> field in the NVMe SMART log is a lifetime counter. It increments every time the drive loses power without first receiving a shutdown notification (CC.SHN) from the host. A clean reboot increments &lt;code>power_cycles&lt;/code> but not &lt;code>unsafe_shutdowns&lt;/code>. A power cut, kernel panic, or someone holding the power button increments both.&lt;/p>
&lt;p>The counter never goes down, so the absolute number is history. What matters is the rate of change: every new increment means something cut power to the drive unexpectedly. On enterprise drives with power-loss protection (PLP) capacitors, in-flight writes in DRAM get flushed to NAND before the power dies, and the event is a footnote. On consumer drives without PLP, every increment is a data-loss roll of the dice: writes the host believes were completed may never have reached NAND.&lt;/p></description></item><item><title>NVMe volatile memory backup device failed: CriticalWarning bit 4 and lost power-loss protection</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-nvme-volatile-memory-backup-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-nvme-volatile-memory-backup-failed/</guid><description>&lt;h1 id="nvme-volatile-memory-backup-device-failed-criticalwarning-bit-4-and-lost-power-loss-protection">NVMe volatile memory backup device failed: CriticalWarning bit 4 and lost power-loss protection&lt;/h1>
&lt;p>You run &lt;code>smartctl -H /dev/nvme0n1&lt;/code> and get FAILED. The detail line reads &lt;code>- volatile memory backup device has failed&lt;/code>. The Critical Warning byte shows &lt;code>0x10&lt;/code>. The drive is still serving reads and writes at full speed, latency is normal, Available Spare is fine, and Percentage Used is well within spec.&lt;/p>
&lt;p>This is NVMe Critical Warning bit 4. It does not mean the NAND is failing or the controller is dying. It means the drive&amp;rsquo;s power-loss protection (PLP) hardware, typically supercapacitors on enterprise NVMe SSDs, has failed. Under stable power, the drive operates normally. But if power drops unexpectedly, data sitting in the drive&amp;rsquo;s volatile write buffer will be lost because the capacitor bank can no longer flush it to NAND.&lt;/p></description></item><item><title>NVMe volatile memory backup failed: critical warning bit 4 and a dead PLP capacitor</title><link>https://www.netdata.cloud/guides/nvme/nvme-volatile-memory-backup-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-volatile-memory-backup-failed/</guid><description>&lt;h1 id="nvme-volatile-memory-backup-failed-critical-warning-bit-4-and-a-dead-plp-capacitor">NVMe volatile memory backup failed: critical warning bit 4 and a dead PLP capacitor&lt;/h1>
&lt;p>Your monitoring shows &lt;code>critical_warning&lt;/code> nonzero on an enterprise NVMe drive. Decoding the bitmask, it is bit 4: volatile memory backup failed. The drive itself is behaving normally. IOPS are fine, latency is fine, media errors are zero. If you only looked at performance dashboards, nothing would look wrong.&lt;/p>
&lt;p>That is exactly the problem. Bit 4 means the power-loss-protection (PLP) capacitor bank on the drive has failed or is degraded. The drive still accepts and acknowledges writes at full speed, but the moment this host loses power unexpectedly, any writes sitting in the controller&amp;rsquo;s volatile DRAM or FTL state are gone. You have lost the one feature that made every previous unsafe shutdown survivable.&lt;/p></description></item><item><title>NVMe warning and critical temperature time: reading the thermal-stress counters</title><link>https://www.netdata.cloud/guides/nvme/nvme-warning-critical-temp-time/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-warning-critical-temp-time/</guid><description>&lt;h1 id="nvme-warning-and-critical-temperature-time-reading-the-thermal-stress-counters">NVMe warning and critical temperature time: reading the thermal-stress counters&lt;/h1>
&lt;p>Most NVMe thermal monitoring is live-state monitoring: you watch composite temperature and alert when it crosses a threshold. That works while you are looking. It tells you nothing about the 40 minutes last night when a backup job pushed an M.2 drive past its warning threshold and the controller quietly throttled your database.&lt;/p>
&lt;p>Two fields in the NVMe SMART/Health Information log close that gap: &lt;code>warning_temp_time&lt;/code> (Warning Composite Temperature Time) and &lt;code>critical_comp_time&lt;/code> (Critical Composite Temperature Time). They are cumulative counters, in minutes, recording how long the drive has spent above its Warning Composite Temperature Threshold (WCTEMP) and Critical Composite Temperature Threshold (CCTEMP). They are the drive&amp;rsquo;s own thermal-stress history, kept whether or not anyone was watching.&lt;/p></description></item><item><title>NVMe write amplification: why data_units_written understates real NAND wear</title><link>https://www.netdata.cloud/guides/nvme/nvme-write-amplification/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-write-amplification/</guid><description>&lt;h1 id="nvme-write-amplification-why-data_units_written-understates-real-nand-wear">NVMe write amplification: why data_units_written understates real NAND wear&lt;/h1>
&lt;p>The usual incident goes like this: &lt;code>percentage_used&lt;/code> on a drive is climbing faster than planned, someone pulls &lt;code>data_units_written&lt;/code> from the SMART log, divides by power-on hours, and concludes the workload is well within the drive&amp;rsquo;s DWPD rating. Six months later the drive sets critical warning bit 0 and procurement is scrambling. The math was not wrong. The input was.&lt;/p>
&lt;p>&lt;code>data_units_written&lt;/code> counts what the host sent to the controller. It does not count what the controller wrote to the NAND. Between those two numbers sits the flash translation layer: garbage collection relocating valid pages, wear leveling shuffling cold blocks, metadata updates, SLC cache destaging. Every one of those operations programs NAND without appearing in a host-visible counter.&lt;/p></description></item><item><title>NVMe write cliff: SLC cache exhaustion and garbage-collection stalls</title><link>https://www.netdata.cloud/guides/nvme/nvme-gc-write-cliff/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvme/nvme-gc-write-cliff/</guid><description>&lt;h1 id="nvme-write-cliff-slc-cache-exhaustion-and-garbage-collection-stalls">NVMe write cliff: SLC cache exhaustion and garbage-collection stalls&lt;/h1>
&lt;p>Write throughput on an NVMe drive falls off a cliff: 50-90% below baseline, arriving as a step function rather than a gradual decline. Write latency jumps 3-10x at the same moment. The application layer starts reporting slow queries, stalled flush operations, or request timeouts, and the drive looks guilty.&lt;/p>
&lt;p>The confusing part is what is absent. The drive is not hot. There are no media errors. The kernel log is quiet. SMART looks clean. Standard disk monitoring shows a busy device, which tells you nothing you did not already know.&lt;/p></description></item><item><title>Nxnetworks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nxnetworks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/nxnetworks-snmp-traps/</guid><description/></item><item><title>OBS Studio</title><link>https://www.netdata.cloud/integrations/data-collection/applications/obs-studio/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/obs-studio/</guid><description/></item><item><title>OBS Studio Monitoring</title><link>https://www.netdata.cloud/monitoring-101/obs_studio-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/obs_studio-monitoring/</guid><description>&lt;h2 id="obs-studio-monitoring">OBS Studio Monitoring&lt;/h2>
&lt;h3 id="what-is-obs-studio">What Is OBS Studio?&lt;/h3>
&lt;p>OBS Studio is a powerful open-source software often used for live streaming and recording. Its capabilities cater to content creators who produce video content for platforms like YouTube, Twitch, and Facebook Live. With its extensive suite of features, OBS Studio enables seamless video mixing, filtering, and transitions, making it a staple in media streaming and recording environments.&lt;/p>
&lt;h3 id="monitoring-obs-studio-with-netdata">Monitoring OBS Studio With Netdata&lt;/h3>
&lt;p>To effectively monitor OBS Studio, utilizing the &lt;a href="https://github.com/lukegb/obs_studio_exporter">OBS Studio Exporter&lt;/a> is crucial. Netdata, a cutting-edge monitoring solution, leverages openmetrics (prometheus) to integrate with the OBS Studio Exporter. This integration empowers technical users to ingest data from any Prometheus exporter without the need for a standalone Prometheus server or Grafana dashboards. Instead, Netdata provides automated dashboards, real-time alerts, and other interactive functionalities to ensure comprehensive monitoring of your OBS Studio setup.&lt;/p></description></item><item><title>Occam Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/occam-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/occam-networks-inc-snmp-traps/</guid><description/></item><item><title>Octel Communications Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/octel-communications-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/octel-communications-corp-snmp-traps/</guid><description/></item><item><title>Offline_Uncorrectable climbing: permanent data loss at the media level</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-offline-uncorrectable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-offline-uncorrectable/</guid><description>&lt;h1 id="offline_uncorrectable-climbing-permanent-data-loss-at-the-media-level">Offline_Uncorrectable climbing: permanent data loss at the media level&lt;/h1>
&lt;p>Offline_Uncorrectable (SMART attribute 198) is the most unambiguous media-failure signal in the ATA SMART attribute set. When this counter increments, a sector was read during an offline scan or self-test and the drive firmware could not recover the data even after ECC correction and multiple retries. The data at that LBA is gone. If the filesystem or application layer did not provide redundancy (RAID, checksums, backups), the loss is permanent.&lt;/p></description></item><item><title>Oid 0 SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/oid-0-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/oid-0-snmp-traps/</guid><description/></item><item><title>Oid 1 SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/oid-1-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/oid-1-snmp-traps/</guid><description/></item><item><title>OIDC</title><link>https://www.netdata.cloud/integrations/authentication/oidc/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/authentication/oidc/</guid><description/></item><item><title>Oki Data Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/oki-data-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/oki-data-corporation-snmp-traps/</guid><description/></item><item><title>Okta SSO</title><link>https://www.netdata.cloud/integrations/authentication/okta-sso/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/authentication/okta-sso/</guid><description/></item><item><title>Olicom A S SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/olicom-a-s-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/olicom-a-s-snmp-traps/</guid><description/></item><item><title>Omnitron Systems Technology SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/omnitron-systems-technology-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/omnitron-systems-technology-snmp-traps/</guid><description/></item><item><title>Omron CJ Ethernet IP</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/omron-cj-ethernet-ip/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/omron-cj-ethernet-ip/</guid><description/></item><item><title>One4Net GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/one4net-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/one4net-gmbh-snmp-traps/</guid><description/></item><item><title>Oneaccess SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/oneaccess-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/oneaccess-snmp-traps/</guid><description/></item><item><title>Onstream Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/onstream-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/onstream-networks-snmp-traps/</guid><description/></item><item><title>OpeanSearch</title><link>https://www.netdata.cloud/integrations/exporters/opeansearch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/opeansearch/</guid><description/></item><item><title>Open vSwitch</title><link>https://www.netdata.cloud/integrations/data-collection/networking/open-vswitch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/open-vswitch/</guid><description/></item><item><title>Open vSwitch Monitoring</title><link>https://www.netdata.cloud/monitoring-101/openvswitch-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/openvswitch-monitoring/</guid><description>&lt;h2 id="open-vswitch-monitoring">Open vSwitch Monitoring&lt;/h2>
&lt;h3 id="what-is-open-vswitch">What Is Open vSwitch?&lt;/h3>
&lt;p>Open vSwitch (OVS) is a multilayer software switch used in virtualized environments to manage network traffic. It is designed to enable network automation through programmatic extensions, while supporting standard management interfaces. OVS plays a significant role in creating highly scalable distributed networking environments.&lt;/p>
&lt;h3 id="monitoring-open-vswitch-with-netdata">Monitoring Open vSwitch With Netdata&lt;/h3>
&lt;p>Monitoring Open vSwitch with Netdata is a straightforward process which can greatly enhance your understanding of network performance and reliability. Netdata utilizes an &lt;a href="https://github.com/digitalocean/openvswitch_exporter">OpenMetrics (Prometheus) exporter&lt;/a> specifically designed for Open vSwitch, allowing users to collect and visualize key network metrics in real-time. Notably, Netdata does not require a separate Prometheus server or Grafana, simplifying the deployment process. Once data is collected, users receive automated dashboards and alerts, enabling proactive troubleshooting and network optimization.&lt;/p></description></item><item><title>Opencode Systems Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/opencode-systems-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/opencode-systems-ltd-snmp-traps/</guid><description/></item><item><title>Opengear Console Manager</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/opengear-console-manager/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/opengear-console-manager/</guid><description/></item><item><title>Opengear Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/opengear-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/opengear-inc-snmp-traps/</guid><description/></item><item><title>Opengear Infrastructure Manager</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/opengear-infrastructure-manager/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/opengear-infrastructure-manager/</guid><description/></item><item><title>OpenLDAP</title><link>https://www.netdata.cloud/integrations/data-collection/applications/openldap/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/openldap/</guid><description/></item><item><title>OpenRC</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/openrc/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/openrc/</guid><description/></item><item><title>OpenRC Monitoring</title><link>https://www.netdata.cloud/monitoring-101/openrc-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/openrc-monitoring/</guid><description>&lt;h2 id="openrc-monitoring">OpenRC Monitoring&lt;/h2>
&lt;h3 id="what-is-openrc">What Is OpenRC?&lt;/h3>
&lt;p>OpenRC is a dependency-based init system that provides services management for Unix-like operating systems. Known for its versatility and compatibility with various environments, OpenRC is an appealing choice for managing system startup processes.&lt;/p>
&lt;h3 id="monitoring-openrc-with-netdata">Monitoring OpenRC With Netdata&lt;/h3>
&lt;p>To effectively monitor OpenRC, using a comprehensive and robust tool like Netdata is crucial. Netdata employs an openmetrics (Prometheus) exporter to collect data from OpenRC systems. This allows for seamless ingestion of metrics and the creation of automated dashboards and alerts without the need for an extensive setup involving a Prometheus server or Grafana. With Netdata, you can gather insights about system performance, detect anomalies, and ensure the health of your OpenRC-managed environments with ease.&lt;/p></description></item><item><title>OpenROADM devices</title><link>https://www.netdata.cloud/integrations/data-collection/networking/openroadm-devices/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/openroadm-devices/</guid><description/></item><item><title>OpenROADM devices Monitoring</title><link>https://www.netdata.cloud/monitoring-101/openroadm-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/openroadm-monitoring/</guid><description>&lt;h2 id="openroadm-devices-monitoring">OpenROADM devices Monitoring&lt;/h2>
&lt;h3 id="what-is-openroadm-devices">What Is OpenROADM devices?&lt;/h3>
&lt;p>OpenROADM (Reconfigurable Optical Add-Drop Multiplexer) devices are integral components of optical transport networks, enabling data to traverse long distances with high fidelity. These devices ensure efficient utilization of optical bandwidth and support the dynamic reconfiguration of network paths, making them critical for modern, adaptable network infrastructure.&lt;/p>
&lt;h3 id="monitoring-openroadm-devices-with-netdata">Monitoring OpenROADM devices With Netdata&lt;/h3>
&lt;p>To monitor OpenROADM devices, Netdata employs an openmetrics (Prometheus) exporter, providing a seamless integration process. Netdata supports data ingestion from any Prometheus exporter, which provides automated dashboards, alerts, and more - all without needing a dedicated Prometheus server or Grafana instance. This makes Netdata an excellent $name monitoring tool as it streamlines the network performance monitoring process, giving you insights in real-time.&lt;/p></description></item><item><title>OpenSearch</title><link>https://www.netdata.cloud/integrations/data-collection/databases/opensearch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/opensearch/</guid><description/></item><item><title>OpenShift Containers</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/openshift-containers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/openshift-containers/</guid><description/></item><item><title>OpenSIPS</title><link>https://www.netdata.cloud/integrations/data-collection/applications/opensips/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/opensips/</guid><description/></item><item><title>OpenStack VMs</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/openstack-vms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/openstack-vms/</guid><description/></item><item><title>OpenTelemetry</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/opentelemetry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/opentelemetry/</guid><description/></item><item><title>OpenTelemetry Logs</title><link>https://www.netdata.cloud/integrations/logs/opentelemetry-logs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/logs/opentelemetry-logs/</guid><description/></item><item><title>OpenTSDB</title><link>https://www.netdata.cloud/integrations/exporters/opentsdb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/opentsdb/</guid><description/></item><item><title>Openvision Technologies Limited SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/openvision-technologies-limited-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/openvision-technologies-limited-snmp-traps/</guid><description/></item><item><title>OpenVPN</title><link>https://www.netdata.cloud/integrations/data-collection/networking/openvpn/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/openvpn/</guid><description/></item><item><title>OpenVPN Monitoring</title><link>https://www.netdata.cloud/monitoring-101/openvpn-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/openvpn-monitoring/</guid><description>&lt;h2 id="openvpn-monitoring">OpenVPN Monitoring&lt;/h2>
&lt;h3 id="what-is-openvpn">What Is OpenVPN?&lt;/h3>
&lt;p>OpenVPN is a renowned open-source VPN protocol that offers secure point-to-point and site-to-site connections. Often utilized to bypass restrictions, enhance online privacy, or securely connect remote workers to enterprise networks, OpenVPN provides a robust and flexible solution for various VPN needs. It is a key tool in the arsenal of IT admins and DevOps engineers who aim to ensure secure network access across distributed environments.&lt;/p></description></item><item><title>OpenVPN status log</title><link>https://www.netdata.cloud/integrations/data-collection/networking/openvpn-status-log/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/openvpn-status-log/</guid><description/></item><item><title>OpenVPN Status Log Monitoring</title><link>https://www.netdata.cloud/monitoring-101/openvpn-status-log-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/openvpn-status-log-monitoring/</guid><description>&lt;h2 id="what-is-openvpn-status-log">What is OpenVPN Status Log?&lt;/h2>
&lt;p>Netdata parses server log files and provides summary (client, traffic) metrics. Unlike the OpenVPN collector which requires management interface to be enabled.&lt;/p>
&lt;h2 id="monitoring-openvpn-status-log-with-netdata">Monitoring OpenVPN Status Log with Netdata&lt;/h2>
&lt;p>Netdata auto discovers hundreds of services, and for those it doesn&amp;rsquo;t turning on manual discovery is a one line configuration. For more information on configuring Netdata for OpenVPN Status Log monitoring please read the collector &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/openvpn_status_log/">documentation&lt;/a>.&lt;/p>
&lt;p>Netdata has a public &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/">demo space&lt;/a> (no login required) where you can explore different monitoring use-cases and get a feel for Netdata.&lt;/p></description></item><item><title>OpenVPN Status Log Monitoring</title><link>https://www.netdata.cloud/monitoring-101/openvpn_status_log-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/openvpn_status_log-monitoring/</guid><description>&lt;h2 id="openvpn-monitoring">OpenVPN Monitoring&lt;/h2>
&lt;h3 id="what-is-openvpn">What Is OpenVPN?&lt;/h3>
&lt;p>OpenVPN is a widely-used open-source VPN protocol designed to secure point-to-point or site-to-site connections in routed or bridged configurations. With a versatile range of applications, it provides secure communication by encrypting data and passing it through secure virtual tunnels. Learn more about OpenVPN on their &lt;a href="https://openvpn.net/">official website&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-openvpn-with-netdata">Monitoring OpenVPN With Netdata&lt;/h3>
&lt;p>Monitoring an OpenVPN setup ensures that your VPN service maintains high performance and reliability. With Netdata, you can monitor OpenVPN by collecting and visualizing metrics in real-time. This allows you to assess your VPN&amp;rsquo;s health proactively and troubleshoot efficiently, safeguarding your network&amp;rsquo;s privacy and integrity.&lt;/p></description></item><item><title>OpenWeatherMap</title><link>https://www.netdata.cloud/integrations/data-collection/applications/openweathermap/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/openweathermap/</guid><description/></item><item><title>OpenWeatherMap Monitoring</title><link>https://www.netdata.cloud/monitoring-101/openweathermap-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/openweathermap-monitoring/</guid><description>&lt;h2 id="openweathermap-monitoring">OpenWeatherMap Monitoring&lt;/h2>
&lt;h3 id="what-is-openweathermap">What Is OpenWeatherMap?&lt;/h3>
&lt;p>OpenWeatherMap is a widely-used platform that delivers a comprehensive range of weather data and air pollution metrics. It serves as an invaluable tool for developers, IT admins, and engineers looking to incorporate real-time weather data into their applications. With OpenWeatherMap, you can access detailed weather information that is crucial for environmental monitoring and analysis, enabling smarter business decisions.&lt;/p>
&lt;h3 id="monitoring-openweathermap-with-netdata">Monitoring OpenWeatherMap With Netdata&lt;/h3>
&lt;p>Monitoring OpenWeatherMap with Netdata involves using an openmetrics (Prometheus) exporter. &lt;a href="https://www.netdata.cloud/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata&lt;/a> is capable of ingesting data from any Prometheus-compatible exporter, including the &lt;a href="https://github.com/Tenzer/openweathermap-exporter">OpenWeatherMap Exporter&lt;/a>. This integration eliminates the need for a separate Prometheus server or Grafana dashboards, as Netdata provides automatically generated dashboards, real-time alerts, and more. By leveraging Netdata&amp;rsquo;s powerful monitoring capabilities, you gain seamless access to critical weather and environmental metrics, ensuring efficient monitoring and analysis.&lt;/p></description></item><item><title>Oplink Communications Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/oplink-communications-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/oplink-communications-inc-snmp-traps/</guid><description/></item><item><title>Opsgenie</title><link>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/opsgenie/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/opsgenie/</guid><description/></item><item><title>OpsGenie</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/opsgenie/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/opsgenie/</guid><description/></item><item><title>Optical Access Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/optical-access-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/optical-access-inc-snmp-traps/</guid><description/></item><item><title>Optical Data Systems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/optical-data-systems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/optical-data-systems-snmp-traps/</guid><description/></item><item><title>Optical modules</title><link>https://www.netdata.cloud/integrations/data-collection/networking/optical-modules/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/optical-modules/</guid><description/></item><item><title>Optical Modules Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ethtool-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ethtool-monitoring/</guid><description>&lt;h2 id="optical-modules-monitoring">Optical Modules Monitoring&lt;/h2>
&lt;h3 id="what-is-optical-modules">What Is Optical Modules?&lt;/h3>
&lt;p>Optical modules are integral components in network environments, tasked with converting electrical signals to optical signals and vice versa for data transmission over fiber optic networks. They are pivotal in ensuring seamless connectivity across high-speed networks. These modules, such as SFP and DDM, support a range of diagnostic features that can be leveraged for efficient monitoring and maintenance.&lt;/p>
&lt;h3 id="monitoring-optical-modules-with-netdata">Monitoring Optical Modules With Netdata&lt;/h3>
&lt;p>Using Netdata, you can effectively monitor optical modules with the &lt;a href="https://man7.org/linux/man-pages/man8/ethtool.8.html">ethtool&lt;/a> collector. Built into Netdata&amp;rsquo;s go.d.plugin, this tool provides real-time insights into essential diagnostic parameters like temperature, voltage, laser bias current, and power levels. Access the &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/ethtool/">Optical Modules collector documentation&lt;/a> to get started quickly.&lt;/p></description></item><item><title>Optical Transmission Labs Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/optical-transmission-labs-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/optical-transmission-labs-inc-snmp-traps/</guid><description/></item><item><title>ORA-00020: maximum number of processes exceeded</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-00020-maximum-processes-exceeded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-00020-maximum-processes-exceeded/</guid><description>&lt;h1 id="ora-00020-maximum-number-of-processes-exceeded">ORA-00020: maximum number of processes exceeded&lt;/h1>
&lt;p>ORA-00020 is a connection-admission cliff-edge. At 99% of the configured PROCESSES limit, everything works. At 100%, new connections are refused and the error cascades into every monitoring tool, DBA session, and application pool. Existing sessions keep running, which is why this is often discovered late.&lt;/p>
&lt;p>The symptom: applications report connection failures, the listener answers TCP but the database rejects the actual connect with ORA-00020, and your usual SYSDBA session may also be refused. The alert log fills with the error. If your connection pool retries in a tight loop, the situation worsens before it improves, because each retry is another failed admission attempt.&lt;/p></description></item><item><title>ORA-00060: deadlock detected while waiting for resource</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-00060-deadlock-detected/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-00060-deadlock-detected/</guid><description>&lt;h1 id="ora-00060-deadlock-detected-while-waiting-for-resource">ORA-00060: deadlock detected while waiting for resource&lt;/h1>
&lt;p>ORA-00060 fires when Oracle&amp;rsquo;s background deadlock detector finds a wait cycle among sessions competing for enqueues. By the time you see it, Oracle has already resolved it: one session was picked as the victim, its current statement was rolled back, and the other sessions in the cycle continued. The victim session is still connected and its transaction is still open. The application must commit or roll it back.&lt;/p></description></item><item><title>ORA-00257: archiver error, connect internal only until freed</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-00257-archiver-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-00257-archiver-error/</guid><description>&lt;h1 id="ora-00257-archiver-error-connect-internal-only-until-freed">ORA-00257: archiver error, connect internal only until freed&lt;/h1>
&lt;p>Mid-incident: users report the application is hung. New connections from application servers fail with ORA-00257. Existing sessions are stuck and return no error. The instance reports OPEN and ACTIVE. The listener responds to TCP probes. A read-only health check says the database is up. The database is not up.&lt;/p>
&lt;p>ORA-00257 is the connect-time symptom of an Archive Hang. The archiver process (ARCn) cannot copy filled online redo logs to the archive destination. Online redo logs fill, the log writer (LGWR) cannot switch to a new log group, and every session that needs to generate redo freezes. Only non-SYSDBA sessions attempting to connect see an error. Everyone else just waits.&lt;/p></description></item><item><title>ORA-00600: internal error code, arguments - triage and what to capture</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-00600-internal-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-00600-internal-error/</guid><description>&lt;h1 id="ora-00600-internal-error-code-arguments---triage-and-what-to-capture">ORA-00600: internal error code, arguments - triage and what to capture&lt;/h1>
&lt;p>ORA-00600 is Oracle&amp;rsquo;s catch-all internal error code. A server process hit an unexpected condition inside the kernel and failed an internal assertion. The database may stay up, the failing call rolls back, and the system may look normal for minutes or hours, but the engine has reported a condition you cannot fix from SQL.&lt;/p>
&lt;p>The message format is &lt;code>ORA-00600: internal error code, arguments: [kdsgrp1], [], [], [], [], [], [], []&lt;/code>. The first bracketed token, here &lt;code>[kdsgrp1]&lt;/code>, is the internal message number and the single most important input for Oracle Support. Multiple distinct bugs can assert in the same internal function, so the argument narrows the search but does not uniquely identify the bug. The full argument list plus the exact database version, down to patch level, is what Support needs to find the matching known bug.&lt;/p></description></item><item><title>ORA-01555: snapshot too old, rollback segment too small</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-01555-snapshot-too-old/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-01555-snapshot-too-old/</guid><description>&lt;h1 id="ora-01555-snapshot-too-old-rollback-segment-too-small">ORA-01555: snapshot too old, rollback segment too small&lt;/h1>
&lt;p>ORA-01555 is an Oracle read-path failure. A long-running query needs an old undo version of a block that has already been overwritten, so the query errors out:&lt;/p>
&lt;pre tabindex="0">&lt;code>ORA-01555: snapshot too old (rollback segment too small)
&lt;/code>&lt;/pre>&lt;p>The &amp;ldquo;rollback segment&amp;rdquo; wording is a legacy artifact. Manual rollback segments were deprecated when Automatic Undo Management (AUM) was introduced in 9i. &lt;!-- TODO: verify exact deprecation timeline across 9i/10g/11g --> The actual cause is undo pressure inside an AUM undo tablespace. Treating this as a &amp;ldquo;rollback segment&amp;rdquo; problem leads down the wrong path.&lt;/p></description></item><item><title>ORA-01578: ORACLE data block corrupted (file #, block #)</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-01578-block-corruption/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-01578-block-corruption/</guid><description>&lt;h1 id="ora-01578-oracle-data-block-corrupted-file--block-">ORA-01578: ORACLE data block corrupted (file #, block #)&lt;/h1>
&lt;p>ORA-01578 is raised when Oracle reads a data block whose contents fail internal validation. The block is marked corrupt and the read fails. The instance keeps running because corruption rarely crashes it, but every query touching that block errors until the block is repaired or bypassed.&lt;/p>
&lt;p>Two risks follow. First, the data in that block is inaccessible or lost. Second, backups may be compromised: RMAN detects corrupt blocks during backup by default (MAXCORRUPT=0), but without proactive VALIDATE you may not know when corruption started or which backup pieces are clean.&lt;/p></description></item><item><title>ORA-01652: unable to extend temp segment in tablespace</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-01652-unable-to-extend-temp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-01652-unable-to-extend-temp/</guid><description>&lt;h1 id="ora-01652-unable-to-extend-temp-segment-in-tablespace">ORA-01652: unable to extend temp segment in tablespace&lt;/h1>
&lt;p>ORA-01652 fires when Oracle cannot allocate another extent for a temporary segment in a tablespace. In production the tablespace is almost always the default &lt;code>TEMP&lt;/code>, and the consumers are operations that have spilled out of the PGA: large sorts, hash joins, global temporary tables, and temporary LOBs. The error is a hard stop for the failing statement, while every other session that needs temp behind it queues on &lt;code>direct path write temp&lt;/code>.&lt;/p></description></item><item><title>ORA-01653: unable to extend table in tablespace</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-01653-unable-to-extend-table/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-01653-unable-to-extend-table/</guid><description>&lt;h1 id="ora-01653-unable-to-extend-table-in-tablespace">ORA-01653: unable to extend table in tablespace&lt;/h1>
&lt;p>ORA-01653 is the hard stop. A DML statement needs to extend a table segment, Oracle cannot allocate the next extent in the tablespace, and the statement fails immediately. There is no graceful degradation: the tablespace was performing normally a moment ago, and now any write that needs new space is rejected. Existing committed data is intact, but inserts, updates that grow row size, index maintenance behind constraints, and any operation that requires new extents all error out.&lt;/p></description></item><item><title>ORA-01654: unable to extend index in tablespace</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-01654-unable-to-extend-index/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-01654-unable-to-extend-index/</guid><description>&lt;h1 id="ora-01654-unable-to-extend-index-in-tablespace">ORA-01654: unable to extend index in tablespace&lt;/h1>
&lt;p>When ORA-01654 fires, an index segment tried to allocate a new extent and could not. The statement fails. Depending on the index, this can stop a bulk load, an &lt;code>ALTER INDEX ... REBUILD&lt;/code>, or any DML that modifies the index. It is the index-specific sibling of ORA-01653: same mechanism (segment extent allocation failure), same fix surface (tablespace capacity), different failing object.&lt;/p>
&lt;p>Two operator surprises are common. First, the failing index often lives in a different tablespace than its base table. Default schemas create indexes alongside tables, but production layouts commonly split indexes into a dedicated INDX tablespace. When the error fires, investigate the index&amp;rsquo;s tablespace, not the table&amp;rsquo;s.&lt;/p></description></item><item><title>ORA-04031: unable to allocate bytes of shared memory (shared pool)</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-04031-shared-pool/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-04031-shared-pool/</guid><description>&lt;h1 id="ora-04031-unable-to-allocate-bytes-of-shared-memory-shared-pool">ORA-04031: unable to allocate bytes of shared memory (shared pool)&lt;/h1>
&lt;p>ORA-04031 fires when a session needs a chunk of shared pool memory and Oracle cannot find a single contiguous free region large enough. The full message lists four values: bytes requested, pool name, allocation type, and heap name. A typical first sighting: &lt;code>ORA-04031: unable to allocate 4032 bytes of shared memory (&amp;quot;shared pool&amp;quot;,&amp;quot;unknown object&amp;quot;,&amp;quot;sga heap&amp;quot;,&amp;quot;row cache buffers&amp;quot;)&lt;/code>. It can also fire against the large pool, java pool, or streams pool.&lt;/p></description></item><item><title>ORA-04036: PGA memory used by the instance exceeds PGA_AGGREGATE_LIMIT</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-04036-pga-limit-exceeded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-04036-pga-limit-exceeded/</guid><description>&lt;h1 id="ora-04036-pga-memory-used-by-the-instance-exceeds-pga_aggregate_limit">ORA-04036: PGA memory used by the instance exceeds PGA_AGGREGATE_LIMIT&lt;/h1>
&lt;p>ORA-04036 fires when instance-wide PGA consumption crosses &lt;code>PGA_AGGREGATE_LIMIT&lt;/code>, the hard cap introduced in 12c. This is not a tuning warning. The database is defending itself by killing or interrupting work.&lt;/p>
&lt;p>&lt;code>PGA_AGGREGATE_TARGET&lt;/code> is a soft target Oracle tries to honor. &lt;code>PGA_AGGREGATE_LIMIT&lt;/code> is an enforced ceiling. Sessions can exceed the target legitimately. Crossing the limit triggers Oracle to abort the call of, and then terminate, the sessions holding the most untunable PGA. SYS and most background processes are not eligible for termination, which produces a distinct and more dangerous failure mode described below.&lt;/p></description></item><item><title>ORA-07445: exception encountered, core dump — a process crash in Oracle code</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-07445-exception/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-07445-exception/</guid><description>&lt;h1 id="ora-07445-exception-encountered-core-dump--a-process-crash-in-oracle-code">ORA-07445: exception encountered, core dump — a process crash in Oracle code&lt;/h1>
&lt;p>ORA-07445 in the alert log means an Oracle server process received a fatal operating system signal and dumped core. A foreground or background process crashed inside Oracle kernel code, PMON cleaned up the session, and Oracle wrote an incident to the Automatic Diagnostic Repository (ADR). Unlike a normal ORA- error returned to a client, ORA-07445 is the kernel telling you a process died underneath the database.&lt;/p></description></item><item><title>ORA-30036: unable to extend segment in undo tablespace</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-30036-unable-to-extend-undo/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-ora-30036-unable-to-extend-undo/</guid><description>&lt;h1 id="ora-30036-unable-to-extend-segment-in-undo-tablespace">ORA-30036: unable to extend segment in undo tablespace&lt;/h1>
&lt;p>A session executing DML fails with &lt;code>ORA-30036: unable to extend segment by N in undo tablespace&lt;/code>. The transaction rolls back and any other session needing undo for the same tablespace also fails. Unlike ORA-01555 (&amp;ldquo;snapshot too old&amp;rdquo;), which is a read-path failure, ORA-30036 is a write-path failure: Oracle cannot find space to record the before-image of a data change.&lt;/p>
&lt;p>The error is cliff-edge. The undo tablespace is the shared resource every DML writes to. When it is exhausted by ACTIVE undo (undo records belonging to uncommitted transactions that Oracle cannot overwrite), no new DML can record its undo and the operation fails immediately.&lt;/p></description></item><item><title>Oracle 'buffer busy waits': hot blocks, sequence headers, and index leaf splits</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-buffer-busy-waits/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-buffer-busy-waits/</guid><description>&lt;h1 id="oracle-buffer-busy-waits-hot-blocks-sequence-headers-and-index-leaf-splits">Oracle &amp;lsquo;buffer busy waits&amp;rsquo;: hot blocks, sequence headers, and index leaf splits&lt;/h1>
&lt;p>&lt;code>buffer busy waits&lt;/code> fires when a session needs a buffer that another session currently has pinned in the buffer cache. It is a contention signal, not an I/O signal. The waiting session is blocked by a session holding the buffer, not by storage latency.&lt;/p>
&lt;p>Since Oracle 10.1 this event has been distinct from &lt;code>read by other session&lt;/code>, which fires when a session waits for another session to finish reading a block from disk into cache. Before 10.1 both conditions collapsed into &lt;code>buffer busy waits&lt;/code>. Modern Oracle (19c, 23ai) keeps the four-event split: &lt;code>buffer busy waits&lt;/code>, &lt;code>read by other session&lt;/code>, &lt;code>gc buffer busy acquire&lt;/code>, and &lt;code>gc buffer busy release&lt;/code>. The last two are RAC-only.&lt;/p></description></item><item><title>Oracle 'Checkpoint not complete': redo log sizing, DBWn, and log-switch stalls</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-checkpoint-not-complete/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-checkpoint-not-complete/</guid><description>&lt;h1 id="oracle-checkpoint-not-complete-redo-log-sizing-dbwn-and-log-switch-stalls">Oracle &amp;lsquo;Checkpoint not complete&amp;rsquo;: redo log sizing, DBWn, and log-switch stalls&lt;/h1>
&lt;p>The alert log says &lt;code>Checkpoint not complete&lt;/code> followed by &lt;code>Current log# N seq# N mem# N: &amp;lt;path&amp;gt;&lt;/code>. Foreground sessions stall on &lt;code>log file switch (checkpoint incomplete)&lt;/code> at every log switch. Commits hesitate for tens of milliseconds to seconds, and the stall repeats each time LGWR wraps to the next redo log group. The wrong fix (enlarging redo logs when the real bottleneck is DBWn I/O) only buys minutes.&lt;/p></description></item><item><title>Oracle 'cursor: pin S wait on X': mutex contention on hot cursors</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-cursor-pin-s-wait-on-x/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-cursor-pin-s-wait-on-x/</guid><description>&lt;h1 id="oracle-cursor-pin-s-wait-on-x-mutex-contention-on-hot-cursors">Oracle &amp;lsquo;cursor: pin S wait on X&amp;rsquo;: mutex contention on hot cursors&lt;/h1>
&lt;p>Sessions accumulating on &lt;code>cursor: pin S wait on X&lt;/code> want a shared (S) mutex pin on a cached cursor while another session holds an exclusive (X) pin on the same cursor object. The X holder is usually hard parsing, invalidating the cursor, or doing library cache maintenance. Waiters queue behind a single mutex.&lt;/p>
&lt;p>This event sits in the library cache mutex family alongside &lt;code>library cache: mutex X&lt;/code>, &lt;code>cursor: mutex S&lt;/code>, and &lt;code>cursor: mutex X&lt;/code>. All of them reflect contention on the shared SQL area. &lt;code>cursor: pin S wait on X&lt;/code> specifically points at cursor-level pin contention, almost always driven by a small number of hot SQL IDs that are parsed, invalidated, or executed frequently, or that have accumulated an unreasonable number of child cursors.&lt;/p></description></item><item><title>Oracle 'db file scattered read': multiblock reads, full scans, and plan regressions</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-db-file-scattered-read/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-db-file-scattered-read/</guid><description>&lt;h1 id="oracle-db-file-scattered-read-multiblock-reads-full-scans-and-plan-regressions">Oracle &amp;lsquo;db file scattered read&amp;rsquo;: multiblock reads, full scans, and plan regressions&lt;/h1>
&lt;p>A sudden spike in &lt;code>db file scattered read&lt;/code> wait time on an OLTP database is one of the most reliable signals of an execution plan regression. The wait event itself is benign on analytics and warehouse workloads, where multiblock full scans are the expected access path. On a transactional system it usually means a query that used to do a handful of index reads is now scanning whole tables.&lt;/p></description></item><item><title>Oracle 'db file sequential read': single-block index reads and buffer cache misses</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-db-file-sequential-read/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-db-file-sequential-read/</guid><description>&lt;h1 id="oracle-db-file-sequential-read-single-block-index-reads-and-buffer-cache-misses">Oracle &amp;lsquo;db file sequential read&amp;rsquo;: single-block index reads and buffer cache misses&lt;/h1>
&lt;p>&lt;code>db file sequential read&lt;/code> is the dominant single-block I/O wait event in Oracle OLTP. Each time a foreground process needs one block, finds it missing from the buffer cache, and waits for the read from a datafile into the SGA, the session accrues time on this event. The name is historical: &amp;ldquo;sequential&amp;rdquo; refers to the single-block read into a specific buffer cache slot, not to a sequential scan pattern. Multiblock scans are tracked separately as &lt;code>db file scattered read&lt;/code>.&lt;/p></description></item><item><title>Oracle 'enq: TM - contention': unindexed foreign keys and table-level locks</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-enq-tm-contention-unindexed-fk/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-enq-tm-contention-unindexed-fk/</guid><description>&lt;h1 id="oracle-enq-tm---contention-unindexed-foreign-keys-and-table-level-locks">Oracle &amp;rsquo;enq: TM - contention&amp;rsquo;: unindexed foreign keys and table-level locks&lt;/h1>
&lt;p>&lt;code>enq: TM - contention&lt;/code> is the wait event a session emits when it wants a DML (table) enqueue on an object and another session already holds a conflicting mode. Unlike &lt;code>enq: TX - row lock contention&lt;/code>, which is row-level and usually about uncommitted transactions, TM contention is almost always structural: an unindexed foreign key that forces Oracle to take a full table lock where it would otherwise take a row lock.&lt;/p></description></item><item><title>Oracle 'enq: TX - row lock contention': blocking sessions and uncommitted DML</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-enq-tx-row-lock-contention/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-enq-tx-row-lock-contention/</guid><description>&lt;h1 id="oracle-enq-tx---row-lock-contention-blocking-sessions-and-uncommitted-dml">Oracle &amp;rsquo;enq: TX - row lock contention&amp;rsquo;: blocking sessions and uncommitted DML&lt;/h1>
&lt;p>The classic symptom is partial slowness. Some transactions succeed. Some hang indefinitely. Application logs show threads stuck inside database calls. There are no ORA- errors in the alert log, CPU is low, I/O latency is normal, the instance is OPEN, and the listener responds.&lt;/p>
&lt;p>When you query V$SESSION for the waiters, they are all parked on the same event: &lt;code>enq: TX - row lock contention&lt;/code>. When you walk &lt;code>BLOCKING_SESSION&lt;/code> up the chain, you usually land on a session whose &lt;code>STATUS&lt;/code> is &lt;code>INACTIVE&lt;/code> and whose wait is &lt;code>SQL*Net message from client&lt;/code>. That session is sitting idle with an open, uncommitted transaction. One stuck holder, many waiters behind it.&lt;/p></description></item><item><title>Oracle 'free buffer waits': when DBWn can't clean buffers fast enough</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-free-buffer-waits/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-free-buffer-waits/</guid><description>&lt;h1 id="oracle-free-buffer-waits-when-dbwn-cant-clean-buffers-fast-enough">Oracle &amp;lsquo;free buffer waits&amp;rsquo;: when DBWn can&amp;rsquo;t clean buffers fast enough&lt;/h1>
&lt;p>Sessions waiting on &lt;code>free buffer waits&lt;/code> cannot find a clean buffer in the cache to read a new block into. The buffer cache is full of dirty buffers that DBWn has not yet flushed, so foreground processes that need to read a new block have nowhere to put it.&lt;/p>
&lt;p>This is a write-path bottleneck, not a read-path problem. The root cause is DBWn falling behind on writes: either the storage cannot absorb the write rate, or DBWn itself is CPU-starved or process-limited. The fix is on the write side. Throwing a bigger buffer cache at the problem does not help and frequently makes it worse.&lt;/p></description></item><item><title>Oracle 'library cache: mutex X' waits: parsing pressure and cursor contention</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-library-cache-mutex-x/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-library-cache-mutex-x/</guid><description>&lt;h1 id="oracle-library-cache-mutex-x-waits-parsing-pressure-and-cursor-contention">Oracle &amp;rsquo;library cache: mutex X&amp;rsquo; waits: parsing pressure and cursor contention&lt;/h1>
&lt;p>When &lt;code>library cache: mutex X&lt;/code> shows up as a dominant wait event, the database is burning CPU on parsing rather than on query execution. Sessions serialize behind exclusive mutexes that protect the shared SQL area, throughput erodes, and the instance still reports OPEN/ACTIVE on every basic availability check. The shared SQL cache cannot keep up with the rate of new SQL text the application is sending.&lt;/p></description></item><item><title>Oracle 'log file sync' waits: slow commits, LGWR, and the redo path</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-log-file-sync-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-log-file-sync-high/</guid><description>&lt;h1 id="oracle-log-file-sync-waits-slow-commits-lgwr-and-the-redo-path">Oracle &amp;rsquo;log file sync&amp;rsquo; waits: slow commits, LGWR, and the redo path&lt;/h1>
&lt;p>Every COMMIT in Oracle is a synchronous handshake with the Log Writer (LGWR). The foreground session hands off its redo, waits for LGWR to flush that redo to the online redo logs, and only then returns control to the application. The wait event that covers this round trip is &lt;code>log file sync&lt;/code>. When LGWR is slow, every committing transaction is slow. This is the most common &amp;ldquo;everything is uniformly slow&amp;rdquo; pattern in Oracle and the single most important latency signal to monitor.&lt;/p></description></item><item><title>Oracle 'Thread N cannot allocate new log': the archive hang that masquerades as up</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-cannot-allocate-new-log/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-cannot-allocate-new-log/</guid><description>&lt;h1 id="oracle-thread-n-cannot-allocate-new-log-the-archive-hang-that-masquerades-as-up">Oracle &amp;lsquo;Thread N cannot allocate new log&amp;rsquo;: the archive hang that masquerades as up&lt;/h1>
&lt;p>The alert log shows &lt;code>Thread 1 cannot allocate new log, sequence 12345&lt;/code>. The instance is OPEN. &lt;code>lsnrctl status&lt;/code> returns services. Your basic availability check, the one that runs &lt;code>SELECT 1 FROM DUAL&lt;/code>, passes. The dashboard is green.&lt;/p>
&lt;p>Meanwhile, every session that needs to commit is frozen on &lt;code>log file switch (archiving needed)&lt;/code>. TPS is at zero. Application connection pools are hung on their next write. New non-SYSDBA logins get ORA-00257; new SYSDBA logins may succeed but immediately block on the first redo-generating statement. This is the archive hang: instance up, commits dead.&lt;/p></description></item><item><title>Oracle archive log destination full: V$ARCHIVE_DEST_STATUS, the ERROR state, and space</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-archive-destination-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-archive-destination-full/</guid><description>&lt;h1 id="oracle-archive-log-destination-full-varchive_dest_status-the-error-state-and-space">Oracle archive log destination full: V$ARCHIVE_DEST_STATUS, the ERROR state, and space&lt;/h1>
&lt;p>The most dangerous Oracle outage masquerades as &amp;ldquo;database up.&amp;rdquo; The instance shows OPEN and ACTIVE in &lt;code>V$INSTANCE&lt;/code>, the listener answers TCP probes, existing sessions stay connected, and basic availability checks pass. But the database is frozen because ARCn cannot write archived redo logs and LGWR cannot switch online redo log groups. Every session that needs to generate redo hangs on &lt;code>log file switch (archiving needed)&lt;/code>, and new non-SYSDBA connections receive ORA-00257.&lt;/p></description></item><item><title>Oracle autoextend hit MAXSIZE: the space gotcha with a half-empty filesystem</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-autoextend-maxsize-reached/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-autoextend-maxsize-reached/</guid><description>&lt;h1 id="oracle-autoextend-hit-maxsize-the-space-gotcha-with-a-half-empty-filesystem">Oracle autoextend hit MAXSIZE: the space gotcha with a half-empty filesystem&lt;/h1>
&lt;p>ORA-01653 (table) or ORA-01654 (index) fires on a production instance. Application writes fail. You log in, check &lt;code>df&lt;/code> on the datafile filesystem, and see hundreds of gigabytes free. ASM disk group shows plenty of headroom. There is no obvious space problem at the storage layer. Yet Oracle insists the tablespace cannot extend.&lt;/p>
&lt;p>This is the autoextend ceiling gotcha. The datafile has AUTOEXTEND ON, but it has reached its MAXSIZE. Oracle refuses to grow the file further even though the disk beneath has abundant space. Disk-only monitoring never sees this coming because the disk is not the constraint. The per-datafile MAXSIZE is.&lt;/p></description></item><item><title>Oracle blocking sessions: finding the blocker at the head of the chain</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-blocking-sessions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-blocking-sessions/</guid><description>&lt;h1 id="oracle-blocking-sessions-finding-the-blocker-at-the-head-of-the-chain">Oracle blocking sessions: finding the blocker at the head of the chain&lt;/h1>
&lt;p>Blocking sessions are the most common cause of &amp;ldquo;the database is slow but everything looks healthy.&amp;rdquo; One session holds an uncommitted transaction on a row; every other session that wants to touch that row queues behind it on &lt;code>enq: TX - row lock contention&lt;/code>. The instance is OPEN, the listener responds, CPU is often low, and TPS quietly decays. From the outside it looks like a performance regression rather than a lock incident.&lt;/p></description></item><item><title>Oracle buffer cache hit ratio: the most misused metric in Oracle monitoring</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-buffer-cache-hit-ratio/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-buffer-cache-hit-ratio/</guid><description>&lt;h1 id="oracle-buffer-cache-hit-ratio-the-most-misused-metric-in-oracle-monitoring">Oracle buffer cache hit ratio: the most misused metric in Oracle monitoring&lt;/h1>
&lt;p>The buffer cache hit ratio appears in nearly every legacy monitoring template, executive dashboard, and database health report. It is also one of the least useful signals for diagnosing real production problems.&lt;/p>
&lt;p>The formula is simple: &lt;code>1 - (physical reads / (db block gets + consistent gets))&lt;/code>, computed from cumulative counters in &lt;code>V$SYSSTAT&lt;/code>. A high percentage looks reassuring. A low percentage looks alarming. Neither reaction is reliably correct. A system doing nothing but &lt;code>SELECT * FROM dual&lt;/code> in a loop has a 99.99% hit ratio and zero useful work. A system running large parallel analytics might sit at 70% and be performing exactly as designed.&lt;/p></description></item><item><title>Oracle connection and session exhaustion: PROCESSES, SESSIONS, and pool sizing</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-connection-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-connection-exhaustion/</guid><description>&lt;h1 id="oracle-connection-and-session-exhaustion-processes-sessions-and-pool-sizing">Oracle connection and session exhaustion: PROCESSES, SESSIONS, and pool sizing&lt;/h1>
&lt;p>New Oracle connections start failing with ORA-00020 (maximum number of processes exceeded) or ORA-00018 (maximum number of sessions exceeded). The listener still answers TCP on 1521 but refuses new connections with TNS-12516 or TNS-12519 because the instance has no free handler to hand over. Existing sessions often keep working, masking the problem until an application tier&amp;rsquo;s pool needs to grow or refresh and fails.&lt;/p></description></item><item><title>Oracle Data Guard transport and apply lag: RPO, RTO, and gap detection</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-data-guard-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-data-guard-lag/</guid><description>&lt;h1 id="oracle-data-guard-transport-and-apply-lag-rpo-rto-and-gap-detection">Oracle Data Guard transport and apply lag: RPO, RTO, and gap detection&lt;/h1>
&lt;p>Data Guard lag is two measurements operators often treat as one. Transport lag is how far behind the standby is in receiving redo: your real-time Recovery Point Objective (RPO) exposure, the data you would lose if the primary failed right now. Apply lag is how far behind the standby is in applying that redo: your real-time Recovery Time Objective (RTO) exposure, the extra time the standby needs to finish catching up before it can open after a failover.&lt;/p></description></item><item><title>Oracle Database monitoring checklist: the signals every production instance needs</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-monitoring-checklist/</guid><description>&lt;h1 id="oracle-database-monitoring-checklist-the-signals-every-production-instance-needs">Oracle Database monitoring checklist: the signals every production instance needs&lt;/h1>
&lt;p>Oracle production instances have a wide surface area: CPU, memory (SGA plus PGA), storage I/O, network, process slots, locks, file descriptors. They also have multiple single points of failure (LGWR, the archiver, the listener) and several failure modes that masquerade as &amp;ldquo;database up&amp;rdquo; while the application is frozen.&lt;/p>
&lt;p>The four-level framework below is drawn from the Signal Catalog and Maturity Levels in the Oracle Database playbook. Each level is additive: Level 2 includes everything in Level 1, and so on. Signals are grouped by the question they answer. Use this checklist to audit your own monitoring against what actually pages you at 3 a.m.&lt;/p></description></item><item><title>Oracle Database monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-monitoring-maturity-model/</guid><description>&lt;h1 id="oracle-database-monitoring-maturity-model-from-survival-to-expert">Oracle Database monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Oracle monitoring is a stack of progressively deeper signals, each layer catching failure modes the layer below cannot see. Teams that jump from &amp;ldquo;is the instance up&amp;rdquo; directly to &amp;ldquo;ASH analysis&amp;rdquo; usually miss the middle, where the highest-frequency production incidents live.&lt;/p>
&lt;p>This model maps four levels of monitoring maturity. It is a prioritization framework, not a tooling roadmap: which signals earn their place at each level, which failure modes they expose, and what blind spots remain if you stop there. Most production Oracle estates sit between Level 1 and Level 2, with a few critical Level 3 signals missing.&lt;/p></description></item><item><title>Oracle DB</title><link>https://www.netdata.cloud/integrations/data-collection/databases/oracle-db/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/oracle-db/</guid><description/></item><item><title>Oracle Fast Recovery Area full: db_recovery_file_dest_size, reclaimable space, and DELETE OBSOLETE</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-fra-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-fra-full/</guid><description>&lt;h1 id="oracle-fast-recovery-area-full-db_recovery_file_dest_size-reclaimable-space-and-delete-obsolete">Oracle Fast Recovery Area full: db_recovery_file_dest_size, reclaimable space, and DELETE OBSOLETE&lt;/h1>
&lt;p>The Fast Recovery Area (FRA) is Oracle&amp;rsquo;s self-managing location for recovery-related files: archived redo logs, RMAN backups, flashback logs, and control file autobackups. Its size is bounded by the &lt;code>DB_RECOVERY_FILE_DEST_SIZE&lt;/code> parameter, a hard quota Oracle enforces internally on the total bytes these files can occupy.&lt;/p>
&lt;p>When the FRA fills, every consumer that writes to it stalls at once. ARCn cannot write archived redo logs. RMAN backups fail. Flashback log creation fails. The most dangerous consequence is the archive stall: online redo logs cannot be reused until they are archived, so LGWR eventually cannot switch to a new log group, and every session that needs to generate redo freezes.&lt;/p></description></item><item><title>Oracle hard parse storm: literal SQL, bind variables, and shared pool churn</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-hard-parse-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-hard-parse-storm/</guid><description>&lt;h1 id="oracle-hard-parse-storm-literal-sql-bind-variables-and-shared-pool-churn">Oracle hard parse storm: literal SQL, bind variables, and shared pool churn&lt;/h1>
&lt;p>CPU is pegged. The top non-idle wait event is &lt;code>library cache: mutex X&lt;/code> or &lt;code>cursor: pin S wait on X&lt;/code>. Active sessions are climbing, but logical reads are flat or falling: sessions are burning CPU on parsing, not execution. Transactions per second is unstable or declining. If the storm runs long enough, ORA-04031 starts appearing in the alert log.&lt;/p></description></item><item><title>Oracle HugePages not configured: page-table overhead and wasted SGA memory</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-hugepages-not-configured/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-hugepages-not-configured/</guid><description>&lt;h1 id="oracle-hugepages-not-configured-page-table-overhead-and-wasted-sga-memory">Oracle HugePages not configured: page-table overhead and wasted SGA memory&lt;/h1>
&lt;p>When Oracle&amp;rsquo;s SGA runs on regular 4KB pages, the Linux kernel maintains a page table entry for every page the SGA touches. A 100GB SGA spans roughly 26 million 4KB pages. Every Oracle process that maps the SGA (dedicated servers, background processes, parallel slaves) carries page table structures tracking those mappings. This per-process page table overhead is invisible to Oracle&amp;rsquo;s own memory views: &lt;code>V$SGA&lt;/code>, &lt;code>V$PGASTAT&lt;/code>, and &lt;code>V$SGASTAT&lt;/code> report the SGA as configured, not as the OS actually consumes it. You only see the overhead at the OS level in &lt;code>/proc/&amp;lt;oracle_pid&amp;gt;/status&lt;/code> (the &lt;code>VmPTE&lt;/code> field) or system-wide in &lt;code>/proc/meminfo&lt;/code> (&lt;code>PageTables&lt;/code>).&lt;/p></description></item><item><title>Oracle instance status: OPEN, MOUNTED, RESTRICTED, and detecting a real outage</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-instance-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-instance-down/</guid><description>&lt;h1 id="oracle-instance-status-open-mounted-restricted-and-detecting-a-real-outage">Oracle instance status: OPEN, MOUNTED, RESTRICTED, and detecting a real outage&lt;/h1>
&lt;p>An Oracle instance that reports &lt;code>STATUS = OPEN&lt;/code> and &lt;code>DATABASE_STATUS = ACTIVE&lt;/code> is not necessarily serving production traffic. Several intermediate and degraded states look healthy to a naive health check while blocking user work: restricted mode left on after maintenance, a quiesce in progress, a resumable operation suspended on a space error, or a shutdown pending behind stuck sessions. Equally common is the opposite failure: a physical standby correctly reports &lt;code>READ ONLY WITH APPLY&lt;/code> and gets paged as &amp;ldquo;down&amp;rdquo; because the check assumed &lt;code>READ WRITE&lt;/code>.&lt;/p></description></item><item><title>Oracle Linux</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/oracle-linux/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/oracle-linux/</guid><description/></item><item><title>Oracle listener errors TNS-12516 / TNS-12519: no available handler and connection refusals</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-listener-down-tns-12516/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-listener-down-tns-12516/</guid><description>&lt;h1 id="oracle-listener-errors-tns-12516--tns-12519-no-available-handler-and-connection-refusals">Oracle listener errors TNS-12516 / TNS-12519: no available handler and connection refusals&lt;/h1>
&lt;p>Applications start failing with TNS-12516 or TNS-12519. Your TCP health probe to port 1521 still returns green. Existing database sessions keep running and serving queries, but every new connection attempt from the application pool is refused. The listener is up, the instance is up, and basic health checks pass, while new connections cannot be established.&lt;/p>
&lt;p>TNS-12516 (&amp;ldquo;listener could not find available handler with matching protocol stack&amp;rdquo;) and TNS-12519 (&amp;ldquo;no appropriate service handler found&amp;rdquo;) mean the listener accepted the TCP socket but could not route the client to a database service handler. The listener accepts the TCP connection, checks its registered services and handlers, and only then spawns or hands off to a dedicated server process. When no handler is available, the client gets a TNS error after the TCP layer already succeeded.&lt;/p></description></item><item><title>Oracle lock contention cascade: one idle session that stalls the whole application</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-lock-contention-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-lock-contention-cascade/</guid><description>&lt;h1 id="oracle-lock-contention-cascade-one-idle-session-that-stalls-the-whole-application">Oracle lock contention cascade: one idle session that stalls the whole application&lt;/h1>
&lt;p>An application opens a transaction, updates a few rows on a hot table, and never commits. A developer&amp;rsquo;s SQL tool is waiting for input, a connection pool returned a dirty connection, or a batch job is mid-update. From the database&amp;rsquo;s perspective the session is INACTIVE and waiting on &lt;code>SQL*Net message from client&lt;/code>. From the application&amp;rsquo;s perspective, every other transaction touching those rows is stuck.&lt;/p></description></item><item><title>Oracle logical reads spiking: the system-wide symptom of a bad plan</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-high-logical-reads/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-high-logical-reads/</guid><description>&lt;h1 id="oracle-logical-reads-spiking-the-system-wide-symptom-of-a-bad-plan">Oracle logical reads spiking: the system-wide symptom of a bad plan&lt;/h1>
&lt;p>A sudden spike in &lt;code>session logical reads&lt;/code> from &lt;code>V$SYSSTAT&lt;/code> is usually the system-wide fingerprint of a plan regression. One query switches from an index scan doing 10 buffer gets per execution to a full table scan doing a million, and at 100 executions per second the database is suddenly doing 100 million additional buffer gets per second. CPU saturates, response times climb, and the whole system feels slow even though no individual component has failed.&lt;/p></description></item><item><title>Oracle out of memory: the Linux OOM killer, SGA, PGA, and random session deaths</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-out-of-memory-oom/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-out-of-memory-oom/</guid><description>&lt;h1 id="oracle-out-of-memory-the-linux-oom-killer-sga-pga-and-random-session-deaths">Oracle out of memory: the Linux OOM killer, SGA, PGA, and random session deaths&lt;/h1>
&lt;p>Oracle sessions are dying at random. Users report sudden disconnects with no application-side explanation. There is no ORA- error returned to the client, no blocking session, no lock chain. The sessions simply vanish. In &lt;code>dmesg&lt;/code> or &lt;code>journalctl -k&lt;/code> you find the evidence: &lt;code>Out of memory: Kill process &amp;lt;pid&amp;gt; (oracle...)&lt;/code>.&lt;/p>
&lt;!-- TODO: verify whether ORA-27300/ORA-27301 appear in the alert log as a direct symptom of OOM kills, or only when fork fails under memory pressure. Client-side errors for OOM-killed sessions are more typically ORA-03113/ORA-03135. -->
&lt;p>The root cause: Oracle&amp;rsquo;s combined memory footprint exceeds physical RAM. The SGA (ideally pinned in hugepages), aggregate PGA across all dedicated server processes, and OS overhead push the system past available memory. The Linux OOM killer does not understand Oracle&amp;rsquo;s internal memory model. It ranks processes by memory consumption and kills the highest-scoring victim. Oracle server processes score high because they map the SGA and carry their own PGA, so they die first.&lt;/p></description></item><item><title>Oracle PGA memory pressure: over-allocation, temp spills, and PGA_AGGREGATE_TARGET</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-pga-memory-pressure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-pga-memory-pressure/</guid><description>&lt;h1 id="oracle-pga-memory-pressure-over-allocation-temp-spills-and-pga_aggregate_target">Oracle PGA memory pressure: over-allocation, temp spills, and PGA_AGGREGATE_TARGET&lt;/h1>
&lt;p>Program Global Area (PGA) is the private memory each dedicated server process uses for sorting, hashing, bitmap operations, and session state. It is allocated outside the SGA, one chunk per process, and managed as an aggregate pool with a soft target (&lt;code>PGA_AGGREGATE_TARGET&lt;/code>) and, from Oracle 12c onward, a hard ceiling (&lt;code>PGA_AGGREGATE_LIMIT&lt;/code>). PGA pressure has two opposite failure signatures, and the operator&amp;rsquo;s job is to tell them apart quickly.&lt;/p></description></item><item><title>Oracle physical I/O latency: per-datafile hotspots and AVGIOTIM</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-physical-io-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-physical-io-latency/</guid><description>&lt;h1 id="oracle-physical-io-latency-per-datafile-hotspots-and-avgiotim">Oracle physical I/O latency: per-datafile hotspots and AVGIOTIM&lt;/h1>
&lt;p>When a datafile&amp;rsquo;s read or write latency rises, every session touching that file slows. Wait events like &lt;code>db file sequential read&lt;/code> and &lt;code>db file scattered read&lt;/code> dominate the top timed events, and throughput drops. Oracle reports I/O statistics across several views with different units, coverage, and granularity. Querying &lt;code>V$FILESTAT&lt;/code> without understanding that &lt;code>AVGIOTIM&lt;/code> is in centiseconds, not milliseconds, leads to misdiagnosis by 10x.&lt;/p></description></item><item><title>Oracle RAC global cache waits: gc buffer busy, interconnect health, and workload affinity</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-rac-gc-buffer-busy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-rac-gc-buffer-busy/</guid><description>&lt;h1 id="oracle-rac-global-cache-waits-gc-buffer-busy-interconnect-health-and-workload-affinity">Oracle RAC global cache waits: gc buffer busy, interconnect health, and workload affinity&lt;/h1>
&lt;p>In Oracle RAC, the global cache (gc) wait events describe time spent moving or coordinating access to data blocks across instances. When these waits dominate the top wait list, the cluster is paying for cross-instance coordination instead of serving work. The signal is clear in &lt;code>V$SYSTEM_EVENT&lt;/code>, but the cause is not.&lt;/p>
&lt;p>&lt;code>gc buffer busy&lt;/code> waits are the most misunderstood of these events. They are not transfer waits. They are contention waits. The same hot blocks are being requested from multiple instances, and sessions queue behind an in-flight transfer instead of getting their own block right away. The fix is almost always workload affinity, not interconnect tuning.&lt;/p></description></item><item><title>Oracle redo generation rate: capacity planning for archiving and Data Guard</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-redo-generation-rate-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-redo-generation-rate-high/</guid><description>&lt;h1 id="oracle-redo-generation-rate-capacity-planning-for-archiving-and-data-guard">Oracle redo generation rate: capacity planning for archiving and Data Guard&lt;/h1>
&lt;p>Redo generation rate is the single most important capacity signal on the Oracle write path. It measures, in bytes per second, how fast change vectors are produced, and therefore how fast every downstream consumer must drain them: online redo logs, the archiver, Data Guard redo transport, and archive log storage.&lt;/p>
&lt;p>Unlike &lt;code>log file sync&lt;/code> wait time, which tells you commits are already slow, redo generation rate is a planning signal. A database generating 50 MB/s of redo needs redo log groups sized for that rate, an archiver that can sustain it, a Data Guard network link that can carry it, and archive storage that can absorb it. If any consumer falls behind the rate, redo backs up until the database hangs.&lt;/p></description></item><item><title>Oracle redo log switch frequency: undersized logs and checkpoint pressure</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-redo-log-switch-frequency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-redo-log-switch-frequency/</guid><description>&lt;h1 id="oracle-redo-log-switch-frequency-undersized-logs-and-checkpoint-pressure">Oracle redo log switch frequency: undersized logs and checkpoint pressure&lt;/h1>
&lt;p>Redo log switch frequency is a direct proxy for redo throughput pressure. When the database cycles through online redo log groups faster than DBWn can flush dirty buffers or ARCn can archive filled logs, every committing session starts waiting. The database does not crash, but transactions stall.&lt;/p>
&lt;p>This article covers what switch frequency measures, the thresholds that separate normal operation from pressure, the cascade from undersized logs into &lt;code>checkpoint not complete&lt;/code> waits, and the operational levers for fixing it. See the &lt;a href="https://www.netdata.cloud/guides/oracle-database/how-oracle-database-works-in-production/">Oracle Database mental model&lt;/a> if you need background on Oracle&amp;rsquo;s redo mechanism and background processes.&lt;/p></description></item><item><title>Oracle RMAN backup failures: V$RMAN_BACKUP_JOB_DETAILS and silent RPO loss</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-rman-backup-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-rman-backup-failures/</guid><description>&lt;h1 id="oracle-rman-backup-failures-vrman_backup_job_details-and-silent-rpo-loss">Oracle RMAN backup failures: V$RMAN_BACKUP_JOB_DETAILS and silent RPO loss&lt;/h1>
&lt;p>RMAN backups can fail for days without visible application impact. The database stays OPEN, queries succeed, transactions commit, and the alert log may show nothing actionable. The only authoritative record is in V$RMAN_BACKUP_JOB_DETAILS, a view that many shops either do not query or query incorrectly by filtering on STATUS = &amp;lsquo;FAILED&amp;rsquo; and missing the more common partial-failure states.&lt;/p>
&lt;p>The operational consequence is silent RPO loss. If the last successful backup is 9 days old and you discover this only when you need to restore, your effective recovery point objective is 9 days, regardless of what your runbook says. RMAN does not raise a pager when backups stop working. The scheduler runs, the script exits 0, and the database continues serving traffic with no safety net.&lt;/p></description></item><item><title>Oracle slow commit cascade: when redo storage degrades and every transaction waits</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-slow-commit-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-slow-commit-cascade/</guid><description>&lt;h1 id="oracle-slow-commit-cascade-when-redo-storage-degrades-and-every-transaction-waits">Oracle slow commit cascade: when redo storage degrades and every transaction waits&lt;/h1>
&lt;p>Every write transaction is slow. TPS is down, p99 commit latency is up, and the slowdown hits writes uniformly across transaction types. No single SQL_ID is the culprit. CPU is underutilized while sessions pile up on a wait event. This is the Oracle slow commit cascade: the storage behind your online redo logs can no longer service LGWR&amp;rsquo;s commit flushes fast enough, and every COMMIT in the database pays the price.&lt;/p></description></item><item><title>Oracle slow query diagnosis: top SQL by buffer gets and elapsed time</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-slow-query-diagnosis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-slow-query-diagnosis/</guid><description>&lt;h1 id="oracle-slow-query-diagnosis-top-sql-by-buffer-gets-and-elapsed-time">Oracle slow query diagnosis: top SQL by buffer gets and elapsed time&lt;/h1>
&lt;p>The application is slow and you suspect one or a few SQL statements, but you do not know which. Oracle answers this precisely: rank statements in V$SQL by BUFFER_GETS (logical I/O) and ELAPSED_TIME (database time), convert cumulative counters into per-execution rates, and compare against the wait-event profile.&lt;/p>
&lt;p>This is a triage article for the case where the instance is OPEN and ACTIVE, the listener responds, and the symptom is degraded query or transaction latency rather than a total hang. If every session is frozen on a single wait event, start with the blocking sessions guide instead.&lt;/p></description></item><item><title>Oracle SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/oracle-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/oracle-snmp-traps/</guid><description/></item><item><title>Oracle SQL plan regression: when a good query suddenly starts doing full scans</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-sql-plan-regression/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-sql-plan-regression/</guid><description>&lt;h1 id="oracle-sql-plan-regression-when-a-good-query-suddenly-starts-doing-full-scans">Oracle SQL plan regression: when a good query suddenly starts doing full scans&lt;/h1>
&lt;p>A query that returned in 10 milliseconds for months now takes 10 seconds. The application is timing out. CPU on the database host is pinned. The instance is OPEN, the listener responds, existing connections work, but every page load crawls. There is no error in the alert log. Nothing crashed.&lt;/p>
&lt;p>This is SQL plan regression. Oracle&amp;rsquo;s cost-based optimizer chose a new execution plan for a statement that used to perform well. The old plan walked an index and touched a handful of blocks. The new plan does a full table scan and touches millions. The SQL_ID is the same, but the PLAN_HASH_VALUE changed, and BUFFER_GETS per execution jumped by one to three orders of magnitude.&lt;/p></description></item><item><title>Oracle tablespace full: monitoring used percent against max capacity</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-tablespace-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-tablespace-full/</guid><description>&lt;h1 id="oracle-tablespace-full-monitoring-used-percent-against-max-capacity">Oracle tablespace full: monitoring used percent against max capacity&lt;/h1>
&lt;p>A tablespace full event in Oracle is a cliff-edge failure. Operations that worked seconds ago return &lt;code>ORA-01653&lt;/code> (table) or &lt;code>ORA-01654&lt;/code> (index). The instance is OPEN, the listener responds, existing SELECTs may still succeed, but any INSERT, UPDATE, or index maintenance that needs to allocate a new extent fails. There is no graceful degradation.&lt;/p>
&lt;p>Oracle allocates space on demand, extent by extent. The moment a segment cannot extend, the operation errors out. If the failing tablespace is UNDO, the failure cascades to every transaction in the database. If it is SYSTEM or SYSAUX, internal operations can stall.&lt;/p></description></item><item><title>Oracle temp tablespace full: sorts, hash joins, and PGA spill</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-temp-tablespace-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-temp-tablespace-full/</guid><description>&lt;h1 id="oracle-temp-tablespace-full-sorts-hash-joins-and-pga-spill">Oracle temp tablespace full: sorts, hash joins, and PGA spill&lt;/h1>
&lt;p>ORA-01652 hits your alert log and a batch job, an ETL run, or an analytical query dies mid-execution. Sessions report &lt;code>ORA-01652: unable to extend temp segment by N in tablespace TEMP&lt;/code>. Throughput to the rest of the database may be fine, but anything that needs work space on disk is now blocked.&lt;/p>
&lt;p>The temp tablespace is the overflow for PGA. Sorts, hash joins, bitmap operations, global temporary tables (GTTs), and temporary LOBs all allocate here when work does not fit in per-process memory. Unlike a permanent tablespace full event, temp full does not usually mean the database is down. It means a specific class of query cannot run.&lt;/p></description></item><item><title>Oracle undo pressure spiral: long transactions, long queries, and ORA-01555/30036</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-undo-pressure-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-undo-pressure-spiral/</guid><description>&lt;h1 id="oracle-undo-pressure-spiral-long-transactions-long-queries-and-ora-0155530036">Oracle undo pressure spiral: long transactions, long queries, and ORA-01555/30036&lt;/h1>
&lt;p>ORA-01555 (&amp;ldquo;snapshot too old&amp;rdquo;) and ORA-30036 (&amp;ldquo;unable to extend undo segment&amp;rdquo;) are two faces of the same pressure: undo extents are being consumed faster than they can be reclaimed or retained. In production they usually arrive together, in a spiral where one large uncommitted transaction starves a fleet of long-running queries, then starves new writes.&lt;/p>
&lt;p>The classic trigger is a batch job that updates or deletes millions of rows in a single transaction with no intermediate commits. It generates enormous undo. At the same time, reports and ETL reads need older undo blocks for read consistency. The undo tablespace fills with ACTIVE extents that cannot be reclaimed. Under the default &lt;code>RETENTION NOGUARANTEE&lt;/code>, Oracle starts stealing UNEXPIRED undo to make room, and the long reads start failing with ORA-01555. If ACTIVE undo fills everything, new DML fails outright with ORA-30036.&lt;/p></description></item><item><title>Oracle undo tablespace usage: ACTIVE vs UNEXPIRED vs EXPIRED and UNDO_RETENTION</title><link>https://www.netdata.cloud/guides/oracle-database/oracle-database-undo-tablespace-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/oracle-database/oracle-database-undo-tablespace-full/</guid><description>&lt;h1 id="oracle-undo-tablespace-usage-active-vs-unexpired-vs-expired-and-undo_retention">Oracle undo tablespace usage: ACTIVE vs UNEXPIRED vs EXPIRED and UNDO_RETENTION&lt;/h1>
&lt;p>Oracle&amp;rsquo;s undo tablespace holds the before-images of changed blocks for transaction rollback and read consistency. A session can undo its own work, PMON can roll back a dead session, and a long-running query can see the database as it existed at the query&amp;rsquo;s start SCN. Those uses compete for the same finite space, and the three extent states in &lt;code>DBA_UNDO_EXTENTS&lt;/code> (ACTIVE, UNEXPIRED, EXPIRED) track which extents can be reused and which cannot.&lt;/p></description></item><item><title>Os Nexus Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/os-nexus-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/os-nexus-inc-snmp-traps/</guid><description/></item><item><title>OSPF Adjacency Topology</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/ospf-adjacency-topology/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/ospf-adjacency-topology/</guid><description/></item><item><title>Overland Data Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/overland-data-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/overland-data-inc-snmp-traps/</guid><description/></item><item><title>oVirt VMs</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/ovirt-vms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/ovirt-vms/</guid><description/></item><item><title>P Cube Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/p-cube-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/p-cube-ltd-snmp-traps/</guid><description/></item><item><title>Pacific Broadband Communications SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pacific-broadband-communications-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pacific-broadband-communications-snmp-traps/</guid><description/></item><item><title>Pacific Broadbank Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pacific-broadbank-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pacific-broadbank-networks-snmp-traps/</guid><description/></item><item><title>Pacific Softworks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pacific-softworks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pacific-softworks-inc-snmp-traps/</guid><description/></item><item><title>Packeteer Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/packeteer-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/packeteer-inc-snmp-traps/</guid><description/></item><item><title>Packetlight Networks Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/packetlight-networks-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/packetlight-networks-ltd-snmp-traps/</guid><description/></item><item><title>Padtec Optical Components And Systems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/padtec-optical-components-and-systems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/padtec-optical-components-and-systems-snmp-traps/</guid><description/></item><item><title>Page types</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/page-types/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/page-types/</guid><description/></item><item><title>PagerDuty</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/pagerduty/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/pagerduty/</guid><description/></item><item><title>PagerDuty</title><link>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/pagerduty/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/pagerduty/</guid><description/></item><item><title>Pairgain Technologies Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pairgain-technologies-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pairgain-technologies-inc-snmp-traps/</guid><description/></item><item><title>Palo Alto</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/palo-alto/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/palo-alto/</guid><description/></item><item><title>Palo Alto Cloudgenix</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/palo-alto-cloudgenix/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/palo-alto-cloudgenix/</guid><description/></item><item><title>Palo Alto Networks PAN-OS</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/palo-alto-networks-pan-os/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/palo-alto-networks-pan-os/</guid><description/></item><item><title>Palo Alto Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/palo-alto-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/palo-alto-networks-snmp-traps/</guid><description/></item><item><title>Pan Dacom Direkt GmbH Formerly Pan Dacom Networking AG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pan-dacom-direkt-gmbh-formerly-pan-dacom-networking-ag-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pan-dacom-direkt-gmbh-formerly-pan-dacom-networking-ag-snmp-traps/</guid><description/></item><item><title>Pan Dacom Telekommunikations SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pan-dacom-telekommunikations-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pan-dacom-telekommunikations-snmp-traps/</guid><description/></item><item><title>Panasas Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/panasas-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/panasas-inc-snmp-traps/</guid><description/></item><item><title>Pandas</title><link>https://www.netdata.cloud/integrations/data-collection/databases/pandas/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/pandas/</guid><description/></item><item><title>Panduit Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/panduit-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/panduit-corp-snmp-traps/</guid><description/></item><item><title>Papouch Elektronika SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/papouch-elektronika-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/papouch-elektronika-snmp-traps/</guid><description/></item><item><title>Paradyne SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/paradyne-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/paradyne-snmp-traps/</guid><description/></item><item><title>Parameter LLC SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/parameter-llc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/parameter-llc-snmp-traps/</guid><description/></item><item><title>Patroni</title><link>https://www.netdata.cloud/integrations/data-collection/databases/patroni/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/patroni/</guid><description/></item><item><title>Patroni Monitoring</title><link>https://www.netdata.cloud/monitoring-101/patroni-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/patroni-monitoring/</guid><description>&lt;h2 id="patroni-monitoring">Patroni Monitoring&lt;/h2>
&lt;h3 id="what-is-patroni">What Is Patroni?&lt;/h3>
&lt;p>Patroni is an open-source high-availability solution for PostgreSQL databases, providing automatic failover and reliable replication solutions. It is designed to be easy to configure and highly reliable, making it a popular choice for database administrators who require robust failover management.&lt;/p>
&lt;h3 id="monitoring-patroni-with-netdata">Monitoring Patroni With Netdata&lt;/h3>
&lt;p>Monitoring Patroni with Netdata gives users real-time insights into their Patroni clusters. By utilizing an openmetrics (Prometheus) exporter like &lt;a href="https://github.com/gopaytech/patroni_exporter">Patroni Exporter&lt;/a>, Netdata can seamlessly ingest metrics from any Prometheus exporter, allowing for automated dashboards, alerts, and more—all without the need for a Prometheus server or Grafana setup. This approach ensures that you have everything you need to monitor Patroni efficiently with minimal setup.&lt;/p></description></item><item><title>Pdu5 SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pdu5-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pdu5-snmp-traps/</guid><description/></item><item><title>Pentair Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pentair-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pentair-inc-snmp-traps/</guid><description/></item><item><title>Peplink</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/peplink/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/peplink/</guid><description/></item><item><title>Percona MySQL</title><link>https://www.netdata.cloud/integrations/data-collection/databases/percona-mysql/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/percona-mysql/</guid><description/></item><item><title>Peribit Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/peribit-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/peribit-networks-snmp-traps/</guid><description/></item><item><title>Periphonics Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/periphonics-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/periphonics-corporation-snmp-traps/</guid><description/></item><item><title>Perle Systems Limited SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/perle-systems-limited-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/perle-systems-limited-snmp-traps/</guid><description/></item><item><title>Personal Weather Station</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/personal-weather-station/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/personal-weather-station/</guid><description/></item><item><title>Personal Weather Station Monitoring</title><link>https://www.netdata.cloud/monitoring-101/pws-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/pws-monitoring/</guid><description>&lt;h2 id="personal-weather-station-monitoring">Personal Weather Station Monitoring&lt;/h2>
&lt;h3 id="what-is-personal-weather-station">What Is Personal Weather Station?&lt;/h3>
&lt;p>A Personal Weather Station (PWS) allows individuals and enthusiasts to set up their own weather monitoring system, offering localized weather metrics that aren&amp;rsquo;t typically available from larger public weather services. These stations can provide real-time data on temperature, humidity, wind speed, rainfall, and other atmospheric conditions.&lt;/p>
&lt;h3 id="monitoring-personal-weather-station-with-netdata">Monitoring Personal Weather Station With Netdata&lt;/h3>
&lt;p>To monitor a Personal Weather Station effectively, Netdata employs an openmetrics (Prometheus) exporter. This seamless integration allows Netdata to ingest data from any Prometheus exporter, providing automated dashboards, alerts, and more—without the need for setting up a Prometheus server or Grafana. The &lt;a href="https://github.com/JohnOrthoefer/pws-exporter">Personal Weather Station Exporter&lt;/a> is an essential tool in this process, enabling efficient real-time monitoring and tracking of weather data.&lt;/p></description></item><item><title>PF Sense</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/pf-sense/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/pf-sense/</guid><description/></item><item><title>pgBackRest</title><link>https://www.netdata.cloud/integrations/data-collection/databases/pgbackrest/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/pgbackrest/</guid><description/></item><item><title>pgBackRest Monitoring</title><link>https://www.netdata.cloud/monitoring-101/pgbackrest-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/pgbackrest-monitoring/</guid><description>&lt;h2 id="pgbackrest-monitoring">pgBackRest Monitoring&lt;/h2>
&lt;h3 id="what-is-pgbackrest">What Is pgBackRest?&lt;/h3>
&lt;p>pgBackRest is a reliable, efficient, and secure backup solution for PostgreSQL databases. It offers advanced functionality for backup and restore, including full, differential, and incremental backups, along with support for parallelism, compression, and encryption. This ensures that your databases are efficiently managed and protected against data loss.&lt;/p>
&lt;h3 id="monitoring-pgbackrest-with-netdata">Monitoring pgBackRest With Netdata&lt;/h3>
&lt;p>Monitoring pgBackRest can be seamlessly integrated into your infrastructure using Netdata. Netdata uses an openmetrics (Prometheus) exporter for pgBackRest, allowing you to efficiently track metrics. Netdata&amp;rsquo;s pgBackRest monitoring tool can ingest data from any Prometheus exporter, automatically generating insightful dashboards and setting up alerts — all without needing a separate Prometheus server or Grafana setup. This integration simplifies the monitoring process and provides real-time visualization of pgBackRest&amp;rsquo;s performance metrics.&lt;/p></description></item><item><title>PgBouncer</title><link>https://www.netdata.cloud/integrations/data-collection/databases/pgbouncer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/pgbouncer/</guid><description/></item><item><title>PgBouncer advisory locks in transaction mode: orphaned locks and mysterious contention</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-advisory-locks-transaction-mode/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-advisory-locks-transaction-mode/</guid><description>&lt;h1 id="pgbouncer-advisory-locks-in-transaction-mode-orphaned-locks-and-mysterious-contention">PgBouncer advisory locks in transaction mode: orphaned locks and mysterious contention&lt;/h1>
&lt;p>Your application takes an advisory lock with &lt;code>pg_advisory_lock()&lt;/code>, does its work, releases it, and moves on. Except under load, other parts of the application start blocking on that same lock, or timing out waiting for it. The application is sure it released the lock. PostgreSQL&amp;rsquo;s lock views, looked at from the wrong place, show nothing. PgBouncer&amp;rsquo;s metrics are completely healthy: no waiting clients, no queue, normal wait times.&lt;/p></description></item><item><title>PgBouncer and PostgreSQL max_connections: when pool_size outruns the backend limit</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-postgresql-max-connections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-postgresql-max-connections/</guid><description>&lt;h1 id="pgbouncer-and-postgresql-max_connections-when-pool_size-outruns-the-backend-limit">PgBouncer and PostgreSQL max_connections: when pool_size outruns the backend limit&lt;/h1>
&lt;p>PgBouncer exists to reduce the number of connections PostgreSQL has to hold, so it is easy to assume that adding PgBouncer makes the connection limit problem go away. It does not. PgBouncer is itself a consumer of PostgreSQL connection slots, and its worst-case demand is arithmetic you control in config files: one pool per (database, user) pair, each pool allowed to open pool_size server connections, plus reserve_pool_size overflow, multiplied by every PgBouncer instance pointing at the same backend.&lt;/p></description></item><item><title>PgBouncer auth failed / password authentication failed: client login rejected</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-auth-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-auth-failed/</guid><description>&lt;p>PgBouncer logs &lt;code>closing because: auth failed&lt;/code> when a client fails to authenticate against the pooler itself. This is client-side authentication: the connection between your application and PgBouncer, not between PgBouncer and PostgreSQL. The connection is rejected before it enters any pool.&lt;/p>
&lt;p>These events are log-only. PgBouncer exposes no SHOW command counter for authentication failures. There is no &lt;code>auth_failed_count&lt;/code> in SHOW STATS, no per-user rejection tally, no rate metric. If you are not parsing the log, you are blind to auth failures.&lt;/p></description></item><item><title>PgBouncer auth_query failures: the authentication dependency loop on PostgreSQL</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-auth-query-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-auth-query-failures/</guid><description>&lt;p>Clients cannot connect. The PgBouncer log fills with &amp;ldquo;password authentication failed&amp;rdquo; and &amp;ldquo;S: login failed&amp;rdquo; messages. &lt;code>SHOW POOLS&lt;/code> shows &lt;code>sv_login&lt;/code> connections stuck in the login state, &lt;code>cl_waiting&lt;/code> climbing. If you use &lt;code>auth_query&lt;/code> to authenticate clients against PostgreSQL, the root cause may be neither a wrong password nor a down database: it may be the circular dependency that &lt;code>auth_query&lt;/code> creates between the pooler and the backend.&lt;/p>
&lt;p>&lt;code>auth_query&lt;/code> tells PgBouncer to authenticate each connecting client by running a SQL query against PostgreSQL to look up that user&amp;rsquo;s password hash. You still need &lt;code>auth_file&lt;/code> (userlist.txt) for &lt;code>auth_user&lt;/code> credentials, but you do not need to sync every application user&amp;rsquo;s password into it. The cost: every new client connection requires a working PostgreSQL connection before the client can be pooled. When PostgreSQL is healthy, this is invisible. When PostgreSQL is under load, or when PgBouncer restarts and hundreds of clients re-authenticate simultaneously, &lt;code>auth_query&lt;/code> becomes the bottleneck that amplifies the outage.&lt;/p></description></item><item><title>PgBouncer avg_query_time high: reading backend slowdown through the pooler</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-avg-query-time-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-avg-query-time-high/</guid><description>&lt;h1 id="pgbouncer-avg_query_time-high-reading-backend-slowdown-through-the-pooler">PgBouncer avg_query_time high: reading backend slowdown through the pooler&lt;/h1>
&lt;p>&lt;code>avg_query_time&lt;/code> in PgBouncer&amp;rsquo;s &lt;code>SHOW STATS&lt;/code> output just doubled, and now &lt;code>cl_waiting&lt;/code> is starting to flicker above zero. This is the classic early-warning sequence for a pool exhaustion cascade: queries take longer, server connections are held longer, pool utilization climbs, and clients begin to queue. Catching the rise at the &lt;code>avg_query_time&lt;/code> stage, before &lt;code>cl_waiting&lt;/code> climbs, is the difference between a quiet Tuesday fix and a paged incident.&lt;/p></description></item><item><title>PgBouncer avg_wait_time high: the latency the pool itself is injecting</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-avg-wait-time-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-avg-wait-time-high/</guid><description>&lt;h1 id="pgbouncer-avg_wait_time-high-the-latency-the-pool-itself-is-injecting">PgBouncer avg_wait_time high: the latency the pool itself is injecting&lt;/h1>
&lt;p>Your application latency is up, PostgreSQL looks fine, and then PgBouncer&amp;rsquo;s stats show &lt;code>avg_wait_time&lt;/code> at tens or hundreds of milliseconds. That number answers &amp;ldquo;where did the latency come from&amp;rdquo;: it is time clients spent queued inside PgBouncer waiting for a server connection, before their query even reached PostgreSQL.&lt;/p>
&lt;p>&lt;code>avg_wait_time&lt;/code> is purely PgBouncer-induced latency. A direct connection to PostgreSQL would not have it. When it is high, the pool is not keeping up with demand, and every queued client pays that delay on top of normal query execution time.&lt;/p></description></item><item><title>PgBouncer backend unreachable: PostgreSQL down and the pool draining</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-backend-unreachable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-backend-unreachable/</guid><description>&lt;h1 id="pgbouncer-backend-unreachable-postgresql-down-and-the-pool-draining">PgBouncer backend unreachable: PostgreSQL down and the pool draining&lt;/h1>
&lt;p>Clients are queueing in PgBouncer, wait times are climbing, and &lt;code>SHOW POOLS&lt;/code> shows fewer server connections than there were an hour ago. Nobody changed the config. The likely cause: PostgreSQL is down, network-partitioned, or rejecting connections, and PgBouncer cannot establish new server connections to replace the ones it is losing.&lt;/p>
&lt;p>This failure mode is deceptive because it degrades slowly. Existing server connections keep serving queries until they expire (&lt;code>server_lifetime&lt;/code>, default 3600s), go idle past &lt;code>server_idle_timeout&lt;/code> (default 600s), or error out. The pool drains gradually rather than failing all at once. Meanwhile clients pile into the wait queue and eventually get disconnected at &lt;code>query_wait_timeout&lt;/code> (default 120s).&lt;/p></description></item><item><title>PgBouncer capacity planning: runway for pools, clients, and PostgreSQL slots</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-capacity-planning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-capacity-planning/</guid><description>&lt;h1 id="pgbouncer-capacity-planning-runway-for-pools-clients-and-postgresql-slots">PgBouncer capacity planning: runway for pools, clients, and PostgreSQL slots&lt;/h1>
&lt;p>PgBouncer capacity planning fails in a specific way: operators size one resource, usually &lt;code>default_pool_size&lt;/code>, and forget the other three. Then a deployment doubles the app fleet, or a config change raises &lt;code>max_client_conn&lt;/code> without touching the OS file descriptor limit, and the first sign of trouble is an incident rather than a dashboard trend.&lt;/p>
&lt;p>There are four independent ceilings in any PgBouncer deployment, each with its own runway. The server pool determines how many queries can execute concurrently. Client slots determine how many application connections PgBouncer will accept. File descriptors determine what the operating system will let PgBouncer open. PostgreSQL backend slots determine whether the database has room for every pool PgBouncer might fill. Exhausting any one of them takes traffic down, and they degrade differently: some are cliffs, some are walls.&lt;/p></description></item><item><title>PgBouncer client connection leak: idle clients that never disconnect</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-client-connection-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-client-connection-leak/</guid><description>&lt;h1 id="pgbouncer-client-connection-leak-idle-clients-that-never-disconnect">PgBouncer client connection leak: idle clients that never disconnect&lt;/h1>
&lt;p>&lt;code>used_clients&lt;/code> keeps climbing. Traffic is flat, &lt;code>cl_waiting&lt;/code> is zero, queries are fast, and yet the client count creeps toward &lt;code>max_client_conn&lt;/code> day after day. When it gets there, PgBouncer starts rejecting new connections with &lt;code>no more connections allowed (max_client_conn)&lt;/code>, even though the database itself is completely healthy.&lt;/p>
&lt;p>This is a client-side connection leak: application instances open connections to PgBouncer and never close them. PgBouncer is doing exactly what it is told to do, holding those sockets open. The pooler just makes the leak visible earlier and more painfully, because it has a hard front-door limit.&lt;/p></description></item><item><title>PgBouncer database paused or disabled: maintenance state that looks like an outage</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-database-paused/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-database-paused/</guid><description>&lt;h1 id="pgbouncer-database-paused-or-disabled-maintenance-state-that-looks-like-an-outage">PgBouncer database paused or disabled: maintenance state that looks like an outage&lt;/h1>
&lt;p>&lt;code>cl_waiting&lt;/code> is climbing, &lt;code>sv_active&lt;/code> just dropped to zero, &lt;code>maxwait&lt;/code> is ticking upward. The pattern looks identical to pool exhaustion or a backend failure. But if someone is running planned PostgreSQL maintenance with &lt;code>PAUSE&lt;/code>, the metrics are behaving as designed: &lt;code>PAUSE&lt;/code> stops new query routing, existing transactions finish, server connections close, and every pending client queues. From the client&amp;rsquo;s perspective, it looks like an outage. The difference is that it is planned and reversible with &lt;code>RESUME&lt;/code>.&lt;/p></description></item><item><title>PgBouncer event loop stall: the single thread that freezes every pool at once</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-event-loop-stall/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-event-loop-stall/</guid><description>&lt;h1 id="pgbouncer-event-loop-stall-the-single-thread-that-freezes-every-pool-at-once">PgBouncer event loop stall: the single thread that freezes every pool at once&lt;/h1>
&lt;p>Every client connection, server connection, DNS lookup, and admin command runs on a single libevent thread inside PgBouncer. When something blocks that thread, every pool on every database freezes at the same time. Clients queue everywhere. The admin console itself becomes slow or unresponsive. This is not pool exhaustion; it is the entire process stalled.&lt;/p>
&lt;p>The signature is simultaneity. In normal pool exhaustion, one (database, user) pool saturates while others stay healthy. In an event loop stall, all pools degrade together, and the &lt;code>SHOW LISTS&lt;/code> command you run to diagnose the problem takes seconds to return. That admin console latency is the tell.&lt;/p></description></item><item><title>PgBouncer high CPU: single-core saturation, TLS, and connection churn</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-high-cpu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-high-cpu/</guid><description>&lt;h1 id="pgbouncer-high-cpu-single-core-saturation-tls-and-connection-churn">PgBouncer high CPU: single-core saturation, TLS, and connection churn&lt;/h1>
&lt;p>PgBouncer is consuming a disproportionate share of one CPU core. The process sits at 70%, 90%, or 100% of a single core while system-wide CPU looks normal. Application queries slow down and the admin console feels sluggish.&lt;/p>
&lt;p>PgBouncer is a single-threaded, event-driven process built on libevent. It runs on exactly one CPU core regardless of how many cores the machine has. A busy PgBouncer looks nearly idle on a multi-core box because system-wide CPU averages dilute the one saturated core. Measure per-process or per-core CPU, not system-wide averages.&lt;/p></description></item><item><title>PgBouncer idle in transaction: the silent pool killer in transaction mode</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-idle-in-transaction/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-idle-in-transaction/</guid><description>&lt;h1 id="pgbouncer-idle-in-transaction-the-silent-pool-killer-in-transaction-mode">PgBouncer idle in transaction: the silent pool killer in transaction mode&lt;/h1>
&lt;p>Your PgBouncer pool looks busy: &lt;code>sv_active&lt;/code> is pinned at &lt;code>pool_size&lt;/code>, &lt;code>cl_waiting&lt;/code> is climbing, and applications are timing out. But PostgreSQL is barely doing anything. CPU is low, &lt;code>avg_query_time&lt;/code> is a few milliseconds, and there are no slow queries to kill. The pool is full of connections that are &amp;ldquo;active&amp;rdquo; yet running nothing at all.&lt;/p>
&lt;p>This is the idle-in-transaction pattern, and it is the most common silent killer in transaction-mode PgBouncer deployments. An application runs &lt;code>BEGIN&lt;/code>, gets assigned a server connection, runs a query, then goes off to do non-database work (HTTP calls, serialization, computation, waiting on another service) before coming back to &lt;code>COMMIT&lt;/code>. In transaction pooling mode, that server connection stays assigned to the client for the entire transaction, including every second the client spends not talking to the database. Enough clients doing this and every server connection is checked out but idle. Everyone else queues.&lt;/p></description></item><item><title>PgBouncer LISTEN/NOTIFY not working: why pub/sub needs session pooling</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-listen-notify-transaction-mode/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-listen-notify-transaction-mode/</guid><description>&lt;h1 id="pgbouncer-listennotify-not-working-why-pubsub-needs-session-pooling">PgBouncer LISTEN/NOTIFY not working: why pub/sub needs session pooling&lt;/h1>
&lt;p>Your application issues &lt;code>LISTEN job_events&lt;/code>, the command succeeds, and PostgreSQL&amp;rsquo;s logs show &lt;code>NOTIFY&lt;/code> firing on schedule. But the listener never receives anything. No error in the application. No error in PgBouncer. No metric anywhere that moves. The feature worked in staging, worked before you put PgBouncer in front of the database, and now it silently does nothing.&lt;/p>
&lt;p>This is the pool mode mismatch failure pattern, and LISTEN/NOTIFY is its most confusing variant because the failure is completely silent. Unlike prepared statements (which at least produce &amp;ldquo;prepared statement does not exist&amp;rdquo; errors), a lost LISTEN registration produces no error at all. The notification is delivered to a backend connection your client no longer holds, or to whichever client happens to hold that connection next.&lt;/p></description></item><item><title>PgBouncer max_client_conn tuning: setting the client limit against real FD headroom</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-max-client-conn-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-max-client-conn-tuning/</guid><description>&lt;h1 id="pgbouncer-max_client_conn-tuning-setting-the-client-limit-against-real-fd-headroom">PgBouncer max_client_conn tuning: setting the client limit against real FD headroom&lt;/h1>
&lt;p>Applications fail with connection errors, but PgBouncer&amp;rsquo;s pools look healthy: &lt;code>sv_active&lt;/code> is well below &lt;code>pool_size&lt;/code>, &lt;code>cl_waiting&lt;/code> is zero, and PostgreSQL is idle. The log tells the real story: &lt;code>no more connections allowed (max_client_conn)&lt;/code> or &lt;code>accept failed: Too many open files&lt;/code>. Clients are being refused before they ever reach a pool.&lt;/p>
&lt;p>The usual root cause is a mismatch between two limits operators treat as one. &lt;code>max_client_conn&lt;/code> is a configuration value. The OS file-descriptor limit (&lt;code>ulimit -n&lt;/code>) is a hard kernel ceiling. PgBouncer cannot accept more client connections than it has file descriptors for, no matter what the config says. If you set &lt;code>max_client_conn = 10000&lt;/code> while the process runs with the default 1024 FD limit, PgBouncer lowers the value at startup or hits the FD ceiling under load and refuses connections long before you expect it to. &lt;!-- TODO: verify whether the startup lowering is logged as a warning or fully silent in current PgBouncer versions -->&lt;/p></description></item><item><title>PgBouncer maxwait high: the oldest client waiter and how close it is to timing out</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-maxwait-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-maxwait-high/</guid><description>&lt;h1 id="pgbouncer-maxwait-high-the-oldest-client-waiter-and-how-close-it-is-to-timing-out">PgBouncer maxwait high: the oldest client waiter and how close it is to timing out&lt;/h1>
&lt;p>You opened &lt;code>SHOW POOLS&lt;/code> because an application is slow, and one column stands out: &lt;code>maxwait&lt;/code> is 8, 20, maybe 90 seconds. That number is the age of the oldest client sitting in PgBouncer&amp;rsquo;s FIFO wait queue, computed from the &lt;code>query_start&lt;/code> of the first waiter. It is the worst-case queuing latency any client is experiencing right now, before its query even reaches PostgreSQL.&lt;/p></description></item><item><title>PgBouncer memory growth: RSS, pkt_buf, and the slab allocator</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-memory-growth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-memory-growth/</guid><description>&lt;h1 id="pgbouncer-memory-growth-rss-pkt_buf-and-the-slab-allocator">PgBouncer memory growth: RSS, pkt_buf, and the slab allocator&lt;/h1>
&lt;p>PgBouncer&amp;rsquo;s RSS is bounded by configuration: &lt;code>max_client_conn&lt;/code> plus the total server connection budget. The process pre-allocates connection structures at startup, so RSS stabilizes after warmup and should not grow monotonically under stable load. When it does, the cause is usually one of: &lt;code>pkt_buf&lt;/code> set too high, TLS session overhead, or a version-specific leak.&lt;/p>
&lt;h2 id="why-rss-matters-for-pgbouncer">Why RSS matters for PgBouncer&lt;/h2>
&lt;p>PgBouncer allocates memory proportional to connection count. The base cost is roughly 2KB per idle connection for socket buffer bookkeeping and connection metadata. For 10,000 clients, that is approximately 30-50MB of RSS at default settings. Active connections cost more because packet buffers are allocated to handle I/O.&lt;/p></description></item><item><title>PgBouncer Monitoring</title><link>https://www.netdata.cloud/monitoring-101/pgbouncer-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/pgbouncer-monitoring/</guid><description>&lt;h2 id="pgbouncer-monitoring">PgBouncer Monitoring&lt;/h2>
&lt;h3 id="what-is-pgbouncer">What Is PgBouncer?&lt;/h3>
&lt;p>&lt;a href="https://www.pgbouncer.org/">PgBouncer&lt;/a> is a lightweight connection pooler for PostgreSQL that aims to reduce the overhead of establishing connections to a PostgreSQL database server. It excels in handling large numbers of connection requests to ensure efficient resource usage and smoother database operations.&lt;/p>
&lt;h3 id="monitoring-pgbouncer-with-netdata">Monitoring PgBouncer With Netdata&lt;/h3>
&lt;p>Netdata, a comprehensive real-time monitoring solution, provides an invaluable toolset to monitor PgBouncer. With Netdata, you can visualize and understand your PgBouncer instances deeply, using real-time visualizations of key metrics that help to diagnose issues and improve performance.&lt;/p></description></item><item><title>PgBouncer monitoring checklist: the signals every connection pooler needs</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-monitoring-checklist/</guid><description>&lt;h1 id="pgbouncer-monitoring-checklist-the-signals-every-connection-pooler-needs">PgBouncer monitoring checklist: the signals every connection pooler needs&lt;/h1>
&lt;p>PgBouncer is not a database. It is a single-threaded, event-driven proxy that multiplexes many client connections onto a smaller set of PostgreSQL connections, and it should be monitored the way you monitor HAProxy or nginx: queueing, connection exhaustion, and process health. The teams that get burned monitor PgBouncer with their PostgreSQL playbook (replication lag, WAL, bloat) and never check whether clients are actually waiting for connections.&lt;/p></description></item><item><title>PgBouncer monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-monitoring-maturity-model/</guid><description>&lt;h1 id="pgbouncer-monitoring-maturity-model-from-survival-to-expert">PgBouncer monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>PgBouncer usually fails as a proxy, not as a process: it is alive, the port is open, the dashboards are green, and clients are still waiting two minutes for a server connection because nobody watched the wait queue. Its important failure modes are queueing, connection exhaustion, event loop stalls, and stale DNS. Most default database checks do not see them.&lt;/p>
&lt;p>This model has four levels, from &amp;ldquo;is it alive&amp;rdquo; to &amp;ldquo;correlate pool behavior with PostgreSQL and the application.&amp;rdquo; Each level answers a specific operational question. The goal is not to reach Level 4. The goal is to know which level you are actually at, and what you are blind to because of it.&lt;/p></description></item><item><title>PgBouncer no more connections allowed (max_client_conn): the front door is full</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-no-more-connections-allowed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-no-more-connections-allowed/</guid><description>&lt;h1 id="pgbouncer-no-more-connections-allowed-max_client_conn-the-front-door-is-full">PgBouncer no more connections allowed (max_client_conn): the front door is full&lt;/h1>
&lt;p>Your application logs fill with connection errors, and every new connection attempt to PgBouncer fails immediately with:&lt;/p>
&lt;pre tabindex="0">&lt;code>ERROR: no more connections allowed (max_client_conn)
&lt;/code>&lt;/pre>&lt;p>This is not pool exhaustion. The client never gets in the door. There is no queue, no wait, no &lt;code>query_wait_timeout&lt;/code>. PgBouncer counts the client connection, sees it would exceed &lt;code>max_client_conn&lt;/code>, and refuses it on the spot. Existing clients keep working; only new ones are turned away.&lt;/p></description></item><item><title>PgBouncer pool exhausted: how to diagnose and fix client waits</title><link>https://www.netdata.cloud/guides/postgres/postgres-pgbouncer-pool-exhausted/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-pgbouncer-pool-exhausted/</guid><description>&lt;h1 id="pgbouncer-pool-exhausted-how-to-diagnose-and-fix-client-waits">PgBouncer pool exhausted: how to diagnose and fix client waits&lt;/h1>
&lt;p>Connection timeouts appear in application logs, but &lt;code>pg_stat_activity&lt;/code> shows plenty of &lt;code>idle&lt;/code> PostgreSQL backends. The bottleneck is usually the connection pooler. When PgBouncer exhausts server connections, clients queue at the pooler instead of reaching PostgreSQL. Sustained queuing raises latency and can cause timeouts before the database sees the query.&lt;/p>
&lt;p>Confirm pool exhaustion, distinguish it from PostgreSQL-side connection limits, and fix the root cause without restarting services.&lt;/p></description></item><item><title>PgBouncer pool exhaustion: clients queue, wait times climb, and the retry cascade</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-pool-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-pool-exhaustion/</guid><description>&lt;h1 id="pgbouncer-pool-exhaustion-clients-queue-wait-times-climb-and-the-retry-cascade">PgBouncer pool exhaustion: clients queue, wait times climb, and the retry cascade&lt;/h1>
&lt;p>Every server connection in a pool is busy. &lt;code>sv_active&lt;/code> equals &lt;code>pool_size&lt;/code>, &lt;code>sv_idle&lt;/code> is zero, and &lt;code>cl_waiting&lt;/code> is climbing. Clients that were getting sub-millisecond connection assignment are now sitting in a FIFO queue, and the oldest waiter (&lt;code>maxwait&lt;/code>) is old enough that application timeouts are firing. This is PgBouncer pool exhaustion, the most common PgBouncer incident, and it has a nasty property: it feeds itself.&lt;/p></description></item><item><title>PgBouncer pool utilization high: sv_active approaching pool_size before clients queue</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-pool-utilization-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-pool-utilization-high/</guid><description>&lt;h1 id="pgbouncer-pool-utilization-high-sv_active-approaching-pool_size-before-clients-queue">PgBouncer pool utilization high: sv_active approaching pool_size before clients queue&lt;/h1>
&lt;p>Your alert fired on &lt;code>sv_active / pool_size &amp;gt; 85%&lt;/code>, or you spotted the ratio creeping up on a dashboard. No clients are waiting yet. &lt;code>cl_waiting&lt;/code> is zero. Latency looks normal. This is exactly the moment this signal exists for: it is the last cheap warning you get before the pool goes over the cliff.&lt;/p>
&lt;p>PgBouncer pool saturation is not a gradual degradation. Below 100% utilization, client wait time is approximately zero because server connection assignment is instant. At 100%, the next client request has nowhere to go and enters a FIFO queue with unbounded wait. There is no &amp;ldquo;slow but working&amp;rdquo; middle state. The ratio of &lt;code>sv_active&lt;/code> to &lt;code>pool_size&lt;/code> tells you how close you are to that edge, and it moves before &lt;code>cl_waiting&lt;/code>, &lt;code>maxwait&lt;/code>, and &lt;code>avg_wait_time&lt;/code> show anything.&lt;/p></description></item><item><title>PgBouncer pool_size sizing: matching pool capacity to transaction time and throughput</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-sizing-pool-size/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-sizing-pool-size/</guid><description>&lt;h1 id="pgbouncer-pool_size-sizing-matching-pool-capacity-to-transaction-time-and-throughput">PgBouncer pool_size sizing: matching pool capacity to transaction time and throughput&lt;/h1>
&lt;p>&lt;code>pool_size&lt;/code> decides how many queries can execute at once through PgBouncer. It is also the setting most teams guess at: leave the default, raise it when something breaks, lower it when PostgreSQL complains about connections. The correct value falls out of two numbers you can measure in a minute: peak transactions per second, and how long the average transaction holds a server connection.&lt;/p></description></item><item><title>PgBouncer prepared statement does not exist: transaction pooling and lost session state</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-prepared-statement-does-not-exist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-prepared-statement-does-not-exist/</guid><description>&lt;h1 id="pgbouncer-prepared-statement-does-not-exist-transaction-pooling-and-lost-session-state">PgBouncer prepared statement does not exist: transaction pooling and lost session state&lt;/h1>
&lt;p>Your application starts throwing &lt;code>prepared statement &amp;quot;...&amp;quot; does not exist&lt;/code> errors in production. PgBouncer is healthy: no queuing, no wait time, pools well under capacity, PostgreSQL is fast. Single-user testing never reproduces it. Restarting the app makes it go away for a while, then it comes back under load.&lt;/p>
&lt;p>This is the pool mode mismatch failure pattern, and it is one of the nastier PgBouncer failure modes because every infrastructure signal looks green. The error is a SQL-level error from PostgreSQL, not a connectivity error from PgBouncer. Nothing in &lt;code>SHOW POOLS&lt;/code> or &lt;code>SHOW STATS&lt;/code> points at it. The only place it shows up is application error logs.&lt;/p></description></item><item><title>PgBouncer process down: the single point of failure in front of PostgreSQL</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-process-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-process-down/</guid><description>&lt;h1 id="pgbouncer-process-down-the-single-point-of-failure-in-front-of-postgresql">PgBouncer process down: the single point of failure in front of PostgreSQL&lt;/h1>
&lt;p>PgBouncer is one process on one host. Its single-threaded libevent loop handles every client socket, server socket, DNS lookup, and admin command on one CPU core. This gives it low overhead (roughly 2KB per idle connection) but makes it a hard single point of failure. When the PID disappears, the event loop stalls, or the process crash-loops, all database traffic through that instance is severed at once.&lt;/p></description></item><item><title>PgBouncer query rate drop: throughput falling without a traffic change</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-query-rate-drop/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-query-rate-drop/</guid><description>&lt;h1 id="pgbouncer-query-rate-drop-throughput-falling-without-a-traffic-change">PgBouncer query rate drop: throughput falling without a traffic change&lt;/h1>
&lt;p>Your PgBouncer dashboard shows queries per second falling off a cliff, but the application team insists nothing changed: no deploy, no traffic shift, no feature flag. The &lt;code>avg_query_count&lt;/code> from &lt;code>SHOW STATS&lt;/code> (or the delta of &lt;code>total_query_count&lt;/code>) is down 40, 60, maybe 90 percent from baseline, and it is not coming back.&lt;/p>
&lt;p>A query rate drop with stable inbound traffic is almost never a PgBouncer bug. It is a symptom of something downstream or upstream: queries are taking longer so fewer complete per second, clients are queued instead of executing, the backend is unreachable, or the application itself has stopped sending work. The query rate is the smoke; your job is to find which fire is producing it.&lt;/p></description></item><item><title>PgBouncer query_wait_timeout: clients disconnected after waiting too long for a connection</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-query-wait-timeout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-query-wait-timeout/</guid><description>&lt;h1 id="pgbouncer-query_wait_timeout-clients-disconnected-after-waiting-too-long-for-a-connection">PgBouncer query_wait_timeout: clients disconnected after waiting too long for a connection&lt;/h1>
&lt;p>Your PgBouncer log is filling with &lt;code>pooler error: query_wait_timeout&lt;/code> lines and application teams are reporting intermittent database errors. The error string sounds like a query problem. It is not. &lt;code>query_wait_timeout&lt;/code> fires when a client has been sitting in PgBouncer&amp;rsquo;s wait queue, blocked on getting a server connection, for longer than the configured timeout. The client never reached PostgreSQL.&lt;/p>
&lt;p>The default is 120 seconds. When this error appears, the pool has been exhausted for at least that long. &lt;code>query_wait_timeout&lt;/code> is a lagging indicator: it tells you an incident happened, not that one is starting.&lt;/p></description></item><item><title>PgBouncer reserve pool activation: overflow capacity that hides an undersized pool</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-reserve-pool-activation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-reserve-pool-activation/</guid><description>&lt;h1 id="pgbouncer-reserve-pool-activation-overflow-capacity-that-hides-an-undersized-pool">PgBouncer reserve pool activation: overflow capacity that hides an undersized pool&lt;/h1>
&lt;p>PgBouncer&amp;rsquo;s reserve pool is overflow capacity: extra server connections beyond &lt;code>pool_size&lt;/code> that PgBouncer may open when clients have waited too long. Used as designed, it absorbs a short traffic spike and goes quiet. Used as a crutch, it hides a chronically undersized pool for months, until the day both base pool and reserve are exhausted and the queuing cliff is steeper than it would have been otherwise.&lt;/p></description></item><item><title>PgBouncer SCRAM / auth_type mismatch: md5 vs scram-sha-256 login failures</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-scram-auth-mismatch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-scram-auth-mismatch/</guid><description>&lt;h1 id="pgbouncer-scram--auth_type-mismatch-md5-vs-scram-sha-256-login-failures">PgBouncer SCRAM / auth_type mismatch: md5 vs scram-sha-256 login failures&lt;/h1>
&lt;p>Clients cannot log in through PgBouncer. Authentication fails with errors that look like wrong passwords, but the passwords are correct and PostgreSQL accepts them on direct connection.&lt;/p>
&lt;p>The root cause is a mismatch between PgBouncer&amp;rsquo;s &lt;code>auth_type&lt;/code> and PostgreSQL&amp;rsquo;s &lt;code>password_encryption&lt;/code> or &lt;code>pg_hba.conf&lt;/code> method. PostgreSQL 12+ defaults to &lt;code>scram-sha-256&lt;/code> for &lt;code>password_encryption&lt;/code>. Deployments that upgrade PostgreSQL without updating PgBouncer&amp;rsquo;s auth configuration break silently because the authentication mechanisms are incompatible, not the credentials.&lt;/p></description></item><item><title>PgBouncer server DNS lookup failed: stale cache and failed failover</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-server-dns-lookup-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-server-dns-lookup-failed/</guid><description>&lt;h1 id="pgbouncer-server-dns-lookup-failed-stale-cache-and-failed-failover">PgBouncer server DNS lookup failed: stale cache and failed failover&lt;/h1>
&lt;p>You see &lt;code>server DNS lookup failed&lt;/code> in the PgBouncer log, usually right after a PostgreSQL failover, a DNS change, or a network event. New server connections cannot be established. Depending on timing, you may instead see the nastier variant: no error at all, just a pool quietly draining because PgBouncer&amp;rsquo;s DNS cache still points at the old primary IP.&lt;/p>
&lt;p>PgBouncer maintains its own DNS cache, independent of the TTL your DNS records publish. The cache lifetime is controlled by &lt;code>dns_max_ttl&lt;/code> (default: 15 seconds). After a failover, up to &lt;code>dns_max_ttl&lt;/code> can pass before PgBouncer even becomes eligible to re-resolve the backend hostname. Cached results are only re-queried when a new server connection is needed, so existing connections to the old IP keep running (against a dead or read-only host) while nothing forces a fresh lookup.&lt;/p></description></item><item><title>PgBouncer server login failed: PgBouncer cannot authenticate to PostgreSQL</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-server-login-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-server-login-failed/</guid><description>&lt;h1 id="pgbouncer-server-login-failed-pgbouncer-cannot-authenticate-to-postgresql">PgBouncer server login failed: PgBouncer cannot authenticate to PostgreSQL&lt;/h1>
&lt;p>Your PgBouncer log is filling with &lt;code>closing because: server login failed&lt;/code> (or &lt;code>server login timed out&lt;/code>), and application latency is climbing. PgBouncer accepts client connections fine. It can even reach PostgreSQL over the network. But every attempt to complete the backend login handshake fails, so no new server connections enter the pool.&lt;/p>
&lt;p>The symptom pattern is distinctive: &lt;code>sv_login&lt;/code> stays elevated while &lt;code>sv_idle&lt;/code> drains toward zero and &lt;code>cl_waiting&lt;/code> grows. Existing server connections keep working until they expire or are recycled, but nothing replaces them. The pool shrinks from the inside while clients pile up in the wait queue. If nothing is fixed, the pool empties completely and every client waits until &lt;code>query_wait_timeout&lt;/code> (default 120s) fires.&lt;/p></description></item><item><title>PgBouncer server_lifetime recycling waves: synchronized reconnects and capacity dips</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-server-lifetime-recycling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-server-lifetime-recycling/</guid><description>&lt;h1 id="pgbouncer-server_lifetime-recycling-waves-synchronized-reconnects-and-capacity-dips">PgBouncer server_lifetime recycling waves: synchronized reconnects and capacity dips&lt;/h1>
&lt;p>Your dashboards show a repeating pattern: every hour (or whatever &lt;code>server_lifetime&lt;/code> is set to), &lt;code>sv_login&lt;/code> spikes, &lt;code>sv_idle&lt;/code> drops, and there is a brief bump in &lt;code>cl_waiting&lt;/code> or &lt;code>avg_wait_time&lt;/code>. It lasts seconds to a couple of minutes, then everything is green again. Application latency ticks up at the same moment. It looks like a flaky backend, but PostgreSQL is fine and the timing is suspiciously regular.&lt;/p></description></item><item><title>PgBouncer server_reset_query and DISCARD ALL: the hidden per-return overhead</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-server-reset-query-discard-all/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-server-reset-query-discard-all/</guid><description>&lt;h1 id="pgbouncer-server_reset_query-and-discard-all-the-hidden-per-return-overhead">PgBouncer server_reset_query and DISCARD ALL: the hidden per-return overhead&lt;/h1>
&lt;p>Every time PgBouncer returns a server connection to the pool in session pooling mode, it runs a cleanup query on that connection before anyone else can use it. By default that query is &lt;code>DISCARD ALL&lt;/code>, which wipes every piece of session state PostgreSQL is holding: temp tables, prepared statements, &lt;code>SET&lt;/code> variables, advisory locks, cursors, &lt;code>LISTEN&lt;/code> subscriptions. The reset is what makes connection sharing safe. It is also a backend round trip you pay for on every connection return, and it is rarely accounted for.&lt;/p></description></item><item><title>PgBouncer SET search_path lost between queries: session variables in transaction mode</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-set-search-path-lost/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-set-search-path-lost/</guid><description>&lt;h1 id="pgbouncer-set-search_path-lost-between-queries-session-variables-in-transaction-mode">PgBouncer SET search_path lost between queries: session variables in transaction mode&lt;/h1>
&lt;p>Your application sets &lt;code>search_path&lt;/code> (or &lt;code>timezone&lt;/code>, or &lt;code>role&lt;/code>, or &lt;code>statement_timeout&lt;/code>) after connecting, and everything works in staging. In production, under concurrent load, queries intermittently hit the wrong schema, run with the wrong role context, or fail with &amp;ldquo;relation does not exist&amp;rdquo; for tables that clearly exist. Restarting the app &amp;ldquo;fixes&amp;rdquo; it briefly. Nothing in PgBouncer&amp;rsquo;s metrics looks wrong.&lt;/p>
&lt;p>This is pool mode mismatch: the application depends on session-level state, but PgBouncer is running in transaction pooling mode, where a client is assigned a different server connection for every transaction. Session state set on one server connection is not present on the next one. The failure is silent, load-dependent, and looks exactly like an application logic bug. PgBouncer itself reports nothing.&lt;/p></description></item><item><title>PgBouncer sv_idle at zero: no headroom and one slow query from a cascade</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-sv-idle-zero/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-sv-idle-zero/</guid><description>&lt;h1 id="pgbouncer-sv_idle-at-zero-no-headroom-and-one-slow-query-from-a-cascade">PgBouncer sv_idle at zero: no headroom and one slow query from a cascade&lt;/h1>
&lt;p>Your PgBouncer dashboards look green. &lt;code>cl_waiting&lt;/code> is zero, &lt;code>maxwait&lt;/code> is zero, no clients are queuing, no errors in the log. But &lt;code>SHOW POOLS&lt;/code> tells a different story: &lt;code>sv_idle&lt;/code> is 0 and &lt;code>sv_active&lt;/code> equals &lt;code>pool_size&lt;/code>. Every server connection in the pool is checked out. Nothing is waiting yet, but nothing is available either.&lt;/p>
&lt;p>This is the &amp;ldquo;looks green, is actually yellow&amp;rdquo; state, and it is one of the most dangerous steady states a connection pooler can sit in. The next request that arrives while all connections are busy queues immediately. There is no buffer, no graceful degradation. PgBouncer&amp;rsquo;s saturation curve is cliff-edge: below 100% utilization, assignment latency is effectively zero; at 100%, latency jumps to unbounded FIFO queuing.&lt;/p></description></item><item><title>PgBouncer thundering herd after restart: a login storm against PostgreSQL</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-thundering-herd-restart/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-thundering-herd-restart/</guid><description>&lt;h1 id="pgbouncer-thundering-herd-after-restart-a-login-storm-against-postgresql">PgBouncer thundering herd after restart: a login storm against PostgreSQL&lt;/h1>
&lt;p>PgBouncer just restarted. Maybe you pushed a config change that required it, maybe the process crashed, maybe the OOM killer took it out and systemd&amp;rsquo;s &lt;code>Restart=always&lt;/code> brought it right back. Within seconds, your dashboards light up: clients are queueing, wait times spike, and PostgreSQL is suddenly absorbing hundreds of simultaneous connection attempts. Then, usually, it calms down on its own within 10 to 60 seconds.&lt;/p></description></item><item><title>PgBouncer TLS: encrypted client and server connections, and certificate expiry</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-tls-configuration/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-tls-configuration/</guid><description>&lt;h1 id="pgbouncer-tls-encrypted-client-and-server-connections-and-certificate-expiry">PgBouncer TLS: encrypted client and server connections, and certificate expiry&lt;/h1>
&lt;p>PgBouncer terminates or initiates TLS on two independent legs: the client leg (application to PgBouncer) and the server leg (PgBouncer to PostgreSQL). Each leg has its own configuration, certificate chain, and failure modes. Enabling TLS on one tells you nothing about the other.&lt;/p>
&lt;p>This independence is the source of most TLS operational surprises. A deployment can have fully encrypted client connections while the backend leg runs in plaintext, and nothing in PgBouncer&amp;rsquo;s metrics flags the discrepancy. The only reliable audit is the &lt;code>tls&lt;/code> column in &lt;code>SHOW CLIENTS&lt;/code> and &lt;code>SHOW SERVERS&lt;/code>.&lt;/p></description></item><item><title>PgBouncer Too many open files: file descriptor exhaustion and refused connections</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-too-many-open-files/</guid><description>&lt;h1 id="pgbouncer-too-many-open-files-file-descriptor-exhaustion-and-refused-connections">PgBouncer Too many open files: file descriptor exhaustion and refused connections&lt;/h1>
&lt;p>PgBouncer is logging &lt;code>Too many open files&lt;/code> and clients are being turned away, or connections to PostgreSQL are failing with the same OS error. The confusing part: &lt;code>used_clients&lt;/code> is nowhere near &lt;code>max_client_conn&lt;/code>, the pools look healthy, and yet new connections are refused. The front door is not full. The kernel is out of file descriptors.&lt;/p>
&lt;p>PgBouncer is FD-hungry by design. Every proxied connection consumes roughly two file descriptors, one for the client socket and one for the server socket, plus a baseline of listening sockets, admin console sockets, DNS resolver sockets, and the log file. When the process hits its &lt;code>Max open files&lt;/code> limit, three things break at once: &lt;code>accept()&lt;/code> on the listen socket fails so new clients cannot connect, new server connections to PostgreSQL cannot be opened so existing clients start queuing, and log writes can fail so diagnostics disappear exactly when you need them.&lt;/p></description></item><item><title>PgBouncer transaction vs session pooling: what each mode breaks and when to use it</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-transaction-vs-session-pooling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-transaction-vs-session-pooling/</guid><description>&lt;h1 id="pgbouncer-transaction-vs-session-pooling-what-each-mode-breaks-and-when-to-use-it">PgBouncer transaction vs session pooling: what each mode breaks and when to use it&lt;/h1>
&lt;p>&lt;code>pool_mode&lt;/code> is the most consequential setting in &lt;code>pgbouncer.ini&lt;/code>. It decides when a server connection is returned to the pool, which in turn decides how many PostgreSQL backends you need and which PostgreSQL features your application is no longer allowed to use.&lt;/p>
&lt;p>The failure pattern that brings people to this page is consistent: someone switches from &lt;code>session&lt;/code> to &lt;code>transaction&lt;/code> for efficiency, deploys, and hours or days later the application starts throwing intermittent &amp;ldquo;prepared statement does not exist&amp;rdquo; errors, losing temp tables mid-request, or leaking advisory locks. Every PgBouncer metric looks healthy. The breakage only shows up under concurrency, because single-user testing keeps landing on the same backend.&lt;/p></description></item><item><title>PgBouncer vs Pgpool-II vs Odyssey: choosing a PostgreSQL connection pooler</title><link>https://www.netdata.cloud/guides/postgres/postgres-pgbouncer-vs-pgpool/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-pgbouncer-vs-pgpool/</guid><description>&lt;h1 id="pgbouncer-vs-pgpool-ii-vs-odyssey-choosing-a-postgresql-connection-pooler">PgBouncer vs Pgpool-II vs Odyssey: choosing a PostgreSQL connection pooler&lt;/h1>
&lt;p>PostgreSQL uses one backend process per connection. Past a few hundred concurrent connections, memory overhead and context switching dominate and throughput collapses. A connection pooler becomes architecturally mandatory at scale. PgBouncer, pgpool-II, and Odyssey make different trade-offs around session state, prepared statements, throughput, and operational complexity. The wrong choice leads to session-leak bugs, throughput ceilings, or split-brain failures unrelated to PostgreSQL itself.&lt;/p></description></item><item><title>PgBouncer wait time vs query time: is it the pool or the database?</title><link>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-wait-time-vs-query-time/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/pgbouncer/pgbouncer-wait-time-vs-query-time/</guid><description>&lt;h1 id="pgbouncer-wait-time-vs-query-time-is-it-the-pool-or-the-database">PgBouncer wait time vs query time: is it the pool or the database?&lt;/h1>
&lt;p>Your application is slow. PgBouncer sits between the application and PostgreSQL, so the first question is always the same: is the latency coming from the pool itself, or from the database behind it? Operators routinely get this split wrong. They see high end-to-end latency, blame PostgreSQL, spend an hour in pg_stat_statements, and then discover avg_query_time was 5ms the whole time while avg_wait_time was 2000ms. The database was fast. The pool was too small.&lt;/p></description></item><item><title>Pgpool-II</title><link>https://www.netdata.cloud/integrations/data-collection/databases/pgpool-ii/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/pgpool-ii/</guid><description/></item><item><title>Pgpool-II Monitoring</title><link>https://www.netdata.cloud/monitoring-101/pgpool2-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/pgpool2-monitoring/</guid><description>&lt;h2 id="pgpool-ii-monitoring">Pgpool-II Monitoring&lt;/h2>
&lt;h3 id="what-is-pgpool-ii">What Is Pgpool-II?&lt;/h3>
&lt;p>Pgpool-II is an essential middleware tool for PostgreSQL databases, allowing for significant improvements in connection pooling, load balancing, and data replication. Designed to handle database connections more efficiently, Pgpool-II helps to maximize your database&amp;rsquo;s performance and scalability. This makes it an ideal choice for developers, DevOps, and SRE professionals who focus on optimizing database operations.&lt;/p>
&lt;h3 id="monitoring-pgpool-ii-with-netdata">Monitoring Pgpool-II With Netdata&lt;/h3>
&lt;p>Monitoring Pgpool-II is crucial for ensuring the performance and reliability of your database systems. Netdata offers a powerful solution for Pgpool-II monitoring by utilizing an openmetrics (Prometheus) exporter. This means that Netdata can seamlessly integrate data from any Prometheus exporter, providing automated dashboards, intelligent alerts, and more, all without the need for a separate Prometheus server or Grafana setup. With Netdata, monitoring Pgpool-II becomes effortless, allowing you to focus on what truly matters—analyzing performance data and making informed decisions.&lt;/p></description></item><item><title>Phihong USA SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/phihong-usa-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/phihong-usa-snmp-traps/</guid><description/></item><item><title>Philips Communication D Entreprise Claude Lubin SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/philips-communication-d-entreprise-claude-lubin-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/philips-communication-d-entreprise-claude-lubin-snmp-traps/</guid><description/></item><item><title>Philips Hue</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/philips-hue/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/philips-hue/</guid><description/></item><item><title>Philips Hue Monitoring</title><link>https://www.netdata.cloud/monitoring-101/philips_hue-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/philips_hue-monitoring/</guid><description>&lt;h2 id="philips-hue-monitoring">Philips Hue Monitoring&lt;/h2>
&lt;h3 id="what-is-philips-hue">What Is Philips Hue?&lt;/h3>
&lt;p>Philips Hue is a smart lighting system that allows users to control light fixtures remotely and automate lighting setups to improve home automation and energy efficiency. With a range of bulbs and accessories, Philips Hue integrates seamlessly into a connected home environment.&lt;/p>
&lt;h3 id="monitoring-philips-hue-with-netdata">Monitoring Philips Hue With Netdata&lt;/h3>
&lt;p>To monitor Philips Hue, Netdata employs an openmetrics (Prometheus) exporter. The integration allows Netdata to collect metrics from Philips Hue using the &lt;a href="https://github.com/aexel90/hue_exporter">Philips Hue Exporter&lt;/a>. This setup enables users to have access to automated dashboards, alerts, and more, all without requiring a separate Prometheus server or Grafana installation. With Netdata&amp;rsquo;s real-time monitoring capabilities, you can ensure that your Philips Hue system is running efficiently and effectively.&lt;/p></description></item><item><title>Phobos Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/phobos-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/phobos-corporation-snmp-traps/</guid><description/></item><item><title>Phoenix Contact GmbH Co KG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/phoenix-contact-gmbh-co-kg-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/phoenix-contact-gmbh-co-kg-snmp-traps/</guid><description/></item><item><title>Phoenix Technologies Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/phoenix-technologies-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/phoenix-technologies-inc-snmp-traps/</guid><description/></item><item><title>Phoenixtec Power Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/phoenixtec-power-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/phoenixtec-power-co-ltd-snmp-traps/</guid><description/></item><item><title>PHP-FPM</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/php-fpm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/php-fpm/</guid><description/></item><item><title>PHP-FPM "child N exited on signal 11 (SIGSEGV)": worker segfaults</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-child-exited-on-signal-sigsegv/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-child-exited-on-signal-sigsegv/</guid><description>&lt;h1 id="php-fpm-child-n-exited-on-signal-11-sigsegv-worker-segfaults">PHP-FPM &amp;ldquo;child N exited on signal 11 (SIGSEGV)&amp;rdquo;: worker segfaults&lt;/h1>
&lt;p>The log line is unambiguous:&lt;/p>
&lt;pre tabindex="0">&lt;code>WARNING: [pool www] child 12345 exited on signal 11 (SIGSEGV) after 3.4 seconds from start
&lt;/code>&lt;/pre>&lt;p>A worker hit a memory protection fault and the kernel killed it. FPM forks a replacement, but the in-flight request is gone: the client sees a truncated response or a 502, and in a tight loop a whole pool can die to a crash storm. A SIGSEGV is almost never a bug in your PHP code. PHP-level fatal errors exit cleanly with a 500 and a stack trace. Signal 11 means native code dereferenced a bad pointer, so the fault lives in a C extension, OPcache, the JIT, or the PHP runtime itself.&lt;/p></description></item><item><title>PHP-FPM "connect() failed (111: Connection refused) while connecting to upstream"</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-connect-failed-connection-refused/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-connect-failed-connection-refused/</guid><description>&lt;h1 id="php-fpm-connect-failed-111-connection-refused-while-connecting-to-upstream">PHP-FPM &amp;ldquo;connect() failed (111: Connection refused) while connecting to upstream&amp;rdquo;&lt;/h1>
&lt;p>The error string in nginx &lt;code>error.log&lt;/code>:&lt;/p>
&lt;pre tabindex="0">&lt;code>*1234 connect() failed (111: Connection refused) while connecting to upstream, client: 10.0.0.5, server: example.com, request: &amp;#34;GET / HTTP/1.1&amp;#34;, upstream: &amp;#34;fastcgi://unix:/run/php/php8.2-fpm.sock:&amp;#34;
&lt;/code>&lt;/pre>&lt;p>&lt;code>errno 111&lt;/code> is &lt;code>ECONNREFUSED&lt;/code>. nginx tried to open a connection to the FastCGI backend and the kernel on the PHP-FPM side rejected it before any FastCGI bytes were exchanged. Users see HTTP 502 Bad Gateway in batches.&lt;/p></description></item><item><title>PHP-FPM "connect() to unix:/... failed (11: Resource temporarily unavailable)"</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-resource-temporarily-unavailable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-resource-temporarily-unavailable/</guid><description>&lt;h1 id="php-fpm-connect-to-unix-failed-11-resource-temporarily-unavailable">PHP-FPM &amp;ldquo;connect() to unix:/&amp;hellip; failed (11: Resource temporarily unavailable)&amp;rdquo;&lt;/h1>
&lt;p>The nginx error &lt;code>connect() to unix:/run/php/php-fpm.sock failed (11: Resource temporarily unavailable) while connecting to upstream&lt;/code> looks like a connection problem. It is not. Errno 11 is &lt;code>EAGAIN&lt;/code>. nginx uses non-blocking sockets, and in this context the kernel&amp;rsquo;s accept queue for the PHP-FPM Unix socket is full. The kernel had nowhere to buffer the new FastCGI connection, so it refused the &lt;code>connect(2)&lt;/code> immediately.&lt;/p></description></item><item><title>PHP-FPM "connect() to unix:/... failed (13: Permission denied)": socket ownership</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-socket-permission-denied/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-socket-permission-denied/</guid><description>&lt;h1 id="php-fpm-connect-to-unix-failed-13-permission-denied-socket-ownership">PHP-FPM &amp;ldquo;connect() to unix:/&amp;hellip; failed (13: Permission denied)&amp;rdquo;: socket ownership&lt;/h1>
&lt;p>The nginx error &lt;code>connect() to unix:/var/run/php/php-fpm.sock failed (13: Permission denied)&lt;/code> is a Unix socket ownership problem between the web server and PHP-FPM. Every PHP request returns 502 while static assets serve fine. The PHP-FPM master is up, the socket file exists, workers are idle, and yet no traffic flows.&lt;/p>
&lt;p>The &lt;code>(13: Permission denied)&lt;/code> suffix is errno 13, &lt;code>EACCES&lt;/code>. nginx calls &lt;code>connect(2)&lt;/code> on the socket path and the kernel rejects it before any FastCGI bytes are exchanged. The cause is on the socket inode or on a parent directory, not in the FastCGI protocol, the application, or pool sizing.&lt;/p></description></item><item><title>PHP-FPM "server reached pm.max_children setting (N), consider raising it"</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-max-children-reached/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-max-children-reached/</guid><description>&lt;h1 id="php-fpm-server-reached-pmmax_children-setting-n-consider-raising-it">PHP-FPM &amp;ldquo;server reached pm.max_children setting (N), consider raising it&amp;rdquo;&lt;/h1>
&lt;p>The warning appears in your PHP-FPM error log verbatim:&lt;/p>
&lt;pre tabindex="0">&lt;code>WARNING: [pool www] server reached pm.max_children setting (5), consider raising it
&lt;/code>&lt;/pre>&lt;p>The pool tried to fork another worker and hit the configured ceiling. Each occurrence is a request that had to wait for a worker to free up instead of being served immediately. The line is emitted at warning level. &lt;!-- TODO: verify "cannot be suppressed through configuration as of PHP 8.x" - log_level=error would mask it at the cost of all warnings -->&lt;/p></description></item><item><title>PHP-FPM "Too many open files": file descriptor exhaustion</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-too-many-open-files/</guid><description>&lt;h1 id="php-fpm-too-many-open-files-file-descriptor-exhaustion">PHP-FPM &amp;ldquo;Too many open files&amp;rdquo;: file descriptor exhaustion&lt;/h1>
&lt;p>The signature line in the PHP-FPM error log is:&lt;/p>
&lt;pre tabindex="0">&lt;code>ERROR: failed to prepare the stderr pipe: Too many open files (24)
&lt;/code>&lt;/pre>&lt;p>The &lt;code>(24)&lt;/code> is the C errno &lt;code>EMFILE&lt;/code>: the calling process has hit its per-process open file descriptor limit. By the time this line appears, the master has already failed to fork a usable child, and application code has probably been failing intermittently for minutes beforehand with random connection refusals, unreadable files, and sessions that refuse to start.&lt;/p></description></item><item><title>PHP-FPM "upstream prematurely closed connection while reading response header"</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-upstream-prematurely-closed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-upstream-prematurely-closed/</guid><description>&lt;h1 id="php-fpm-upstream-prematurely-closed-connection-while-reading-response-header">PHP-FPM &amp;ldquo;upstream prematurely closed connection while reading response header&amp;rdquo;&lt;/h1>
&lt;p>nginx logs this error at &lt;code>[error]&lt;/code> level when the FastCGI connection to PHP-FPM closes before response headers arrive. The client gets a 502 Bad Gateway. Unlike &lt;code>connect() failed (111: Connection refused)&lt;/code> (no worker available) or &lt;code>upstream timed out (110: Connection timed out)&lt;/code> (worker too slow), this error means the connection was established, the worker began executing PHP, and the worker vanished before delivering a response.&lt;/p></description></item><item><title>PHP-FPM 502 Bad Gateway: the web server cannot reach the pool</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-502-bad-gateway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-502-bad-gateway/</guid><description>&lt;h1 id="php-fpm-502-bad-gateway-the-web-server-cannot-reach-the-pool">PHP-FPM 502 Bad Gateway: the web server cannot reach the pool&lt;/h1>
&lt;p>A 502 Bad Gateway means the upstream PHP-FPM pool never returned a valid FastCGI response. nginx, Apache, or Caddy tried to open a connection to the FPM socket and either the connection never established, was refused by the kernel, or was torn down mid-request.&lt;/p>
&lt;p>The fastest triage move is to test both ends of the FastCGI connection: probe the FPM ping endpoint, then read the web server error log for the exact upstream message. Those two signals narrow the cause in seconds.&lt;/p></description></item><item><title>PHP-FPM 504 Gateway Timeout: requests accepted but never finishing in time</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-504-gateway-timeout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-504-gateway-timeout/</guid><description>&lt;h1 id="php-fpm-504-gateway-timeout-requests-accepted-but-never-finishing-in-time">PHP-FPM 504 Gateway Timeout: requests accepted but never finishing in time&lt;/h1>
&lt;p>A 504 from nginx means the FastCGI connection was established, the request was handed to a worker, and the worker never produced a response before nginx gave up. This is a different failure from a 502, where nginx could not connect at all (socket missing, backlog full, or master dead). The FastCGI channel was healthy. The work was not.&lt;/p></description></item><item><title>PHP-FPM active processes near max_children: reading pool utilization</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-active-processes-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-active-processes-high/</guid><description>&lt;h1 id="php-fpm-active-processes-near-max_children-reading-pool-utilization">PHP-FPM active processes near max_children: reading pool utilization&lt;/h1>
&lt;p>The ratio of &lt;code>active processes&lt;/code> to &lt;code>pm.max_children&lt;/code> is the most important saturation signal PHP-FPM exposes. Each worker handles one request at a time, so the active worker count is your current concurrent request load. Divided by the configured ceiling, it behaves like a capacity gauge: near 1.0 means no headroom; sustained near 1.0 means one slow query away from queuing.&lt;/p>
&lt;p>This is a reading guide for that ratio: what the numerator and denominator actually mean, how to distinguish a normal burst from a sustained problem, how to use the high-water mark (&lt;code>max active processes&lt;/code>) for capacity planning, and which correlated signals disambiguate &amp;ldquo;busy&amp;rdquo;, &amp;ldquo;saturated&amp;rdquo;, and &amp;ldquo;broken&amp;rdquo;. It is not a tuning guide; see the related guides at the end for that.&lt;/p></description></item><item><title>PHP-FPM and nginx timeout mismatch: phantom workers and confusing 504s</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-nginx-timeout-mismatch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-nginx-timeout-mismatch/</guid><description>&lt;h1 id="php-fpm-and-nginx-timeout-mismatch-phantom-workers-and-confusing-504s">PHP-FPM and nginx timeout mismatch: phantom workers and confusing 504s&lt;/h1>
&lt;p>Users see 504 Gateway Timeout. You check PHP-FPM and the pool is not saturated, or it is saturated but the active workers are running requests that should have finished minutes ago. The nginx error log shows upstream timeouts. The FPM error log shows nothing unusual.&lt;/p>
&lt;p>nginx and PHP-FPM each have their own notion of how long a request may run. When those notions disagree, you get two failure modes: phantom workers (nginx gave up, FPM kept going) and confusing 502/504 patterns (FPM killed a worker nginx was still waiting on). The configuration is incoherent across the request path.&lt;/p></description></item><item><title>PHP-FPM crash loop and fork storm: workers dying faster than they serve</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-crash-loop/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-crash-loop/</guid><description>&lt;h1 id="php-fpm-crash-loop-and-fork-storm-workers-dying-faster-than-they-serve">PHP-FPM crash loop and fork storm: workers dying faster than they serve&lt;/h1>
&lt;p>Workers start, accept a request, hit a crash-inducing code path, die with SIGSEGV or SIGBUS, and get respawned by the master. The replacement picks up another request from the same traffic pattern, hits the same code, and dies again. The pool spends more time forking and dying than serving. Each fork costs the kernel time to copy page tables and costs PHP time to initialize the runtime. Users see intermittent 502s while &lt;code>total processes&lt;/code> oscillates and the &lt;code>accepted conn&lt;/code> rate collapses.&lt;/p></description></item><item><title>PHP-FPM dynamic mode scaling lag: why the pool cannot keep up with bursts</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-dynamic-mode-scaling-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-dynamic-mode-scaling-lag/</guid><description>&lt;h1 id="php-fpm-dynamic-mode-scaling-lag-why-the-pool-cannot-keep-up-with-bursts">PHP-FPM dynamic mode scaling lag: why the pool cannot keep up with bursts&lt;/h1>
&lt;p>During a traffic burst, your PHP-FPM pool reports total processes well below &lt;code>pm.max_children&lt;/code>, yet the listen queue fills and latency spikes. The status page shows idle workers hitting zero, then climbing back a few seconds later as new workers come online. By then, requests have already queued and some users have already seen 502s or 504s.&lt;/p>
&lt;p>This is dynamic mode scaling lag. The dynamic process manager is reactive, not predictive. It checks idle-worker counts on a timer and forks replacement workers after the deficit is already visible. When a burst arrives faster than the check-and-fork cycle can respond, the idle pool drains to zero and the listen backlog absorbs the overflow until new workers are ready.&lt;/p></description></item><item><title>PHP-FPM emergency restart: "failed processes threshold reached, initiating reload"</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-emergency-restart/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-emergency-restart/</guid><description>&lt;h1 id="php-fpm-emergency-restart-failed-processes-threshold-reached-initiating-reload">PHP-FPM emergency restart: &amp;ldquo;failed processes threshold reached, initiating reload&amp;rdquo;&lt;/h1>
&lt;p>You see a NOTICE line in the PHP-FPM error log:&lt;/p>
&lt;pre tabindex="0">&lt;code>NOTICE: failed processes threshold (N in M sec) is reached, initiating reload
&lt;/code>&lt;/pre>&lt;p>The master has watched enough children die from SIGSEGV or SIGBUS inside a configured interval and is giving up on soft recovery. It is about to &lt;code>execvp()&lt;/code> itself: every pool is recycled, the OPcache shared memory segment is destroyed, and for a few seconds no PHP request can be served. When traffic returns, every worker pays a compilation penalty while the cache warms.&lt;/p></description></item><item><title>PHP-FPM executing PHP in upload directories: webshell RCE detection</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-php-execution-restricted-paths/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-php-execution-restricted-paths/</guid><description>&lt;h1 id="php-fpm-executing-php-in-upload-directories-webshell-rce-detection">PHP-FPM executing PHP in upload directories: webshell RCE detection&lt;/h1>
&lt;p>A correctly configured web server and PHP-FPM stack never returns HTTP 200 for a &lt;code>.php&lt;/code> file inside an upload, temporary, or media directory. If it does, that file executed as PHP code. In production, that file is a webshell, and the host is compromised.&lt;/p>
&lt;p>This is a binary condition. The PHP-FPM playbook assigns it PAGE severity because a healthy server cannot produce this signal. If your access logs show a 200 response for a PHP file in a restricted path, start incident response immediately: isolate the host, preserve forensic data, and assume the attacker has achieved code execution within the PHP-FPM worker&amp;rsquo;s user context.&lt;/p></description></item><item><title>PHP-FPM graceful reload: the brief no-worker window on SIGUSR2</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-graceful-reload-window/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-graceful-reload-window/</guid><description>&lt;h1 id="php-fpm-graceful-reload-the-brief-no-worker-window-on-sigusr2">PHP-FPM graceful reload: the brief no-worker window on SIGUSR2&lt;/h1>
&lt;p>A PHP-FPM graceful reload (&lt;code>kill -USR2 &amp;lt;master_pid&amp;gt;&lt;/code>, &lt;code>systemctl reload php8.3-fpm&lt;/code>) is not graceful in the nginx sense. There is no overlap window where new workers handle fresh traffic while old workers drain. The old pool is torn down first, the master re-execs itself, and only then does it fork replacement workers. For a measurable interval zero workers are serving requests.&lt;/p>
&lt;p>If you page on a 502 spike, a listen-queue jump, or a failed ping probe during a deploy, you have probably already met this window. The pattern is narrow, predictable, and bounded by &lt;code>process_control_timeout&lt;/code>. Treating it as an incident is one of the most common PHP-FPM false alarms.&lt;/p></description></item><item><title>PHP-FPM idle processes at zero: no burst headroom left</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-idle-processes-zero/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-idle-processes-zero/</guid><description>&lt;h1 id="php-fpm-idle-processes-at-zero-no-burst-headroom-left">PHP-FPM idle processes at zero: no burst headroom left&lt;/h1>
&lt;p>The PHP-FPM status page reports &lt;code>idle processes&lt;/code> at or near zero. In &lt;code>static&lt;/code> and &lt;code>dynamic&lt;/code> modes this is the leading indicator that the pool is about to start queuing. The next request that arrives has no worker ready to accept it, so it lands in the kernel-managed socket backlog. If demand keeps rising, the backlog fills and the web server starts returning 502s.&lt;/p></description></item><item><title>PHP-FPM in containers: cgroup limits and the silent OOM kill</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-cgroup-oom-container/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-cgroup-oom-container/</guid><description>&lt;h1 id="php-fpm-in-containers-cgroup-limits-and-the-silent-oom-kill">PHP-FPM in containers: cgroup limits and the silent OOM kill&lt;/h1>
&lt;p>The container restarts, or the PHP-FPM error log fills with &lt;code>child N exited on signal 9 (SIGKILL)&lt;/code> entries, but &lt;code>dmesg&lt;/code> shows no global OOM event. The host has plenty of free RAM. The PHP &lt;code>memory_limit&lt;/code> is set well below the kill threshold. None of the obvious explanations fit.&lt;/p>
&lt;p>PHP-FPM has no awareness of cgroup memory limits. Workers allocate against the PHP heap (capped by &lt;code>memory_limit&lt;/code>), against extension memory (Imagick, libxml, glibc arenas), and against page cache for mmap&amp;rsquo;d files. None of those allocations consult &lt;code>memory.max&lt;/code> in the cgroup. When the cgroup as a whole crosses its limit, the kernel cgroup OOM killer fires a SIGKILL at a process inside the cgroup. Whether the victim is a single worker or the master process determines whether you see a silent kill or a container restart.&lt;/p></description></item><item><title>PHP-FPM listen backlog overflow: the kernel silently dropping connections</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-listen-backlog-overflow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-listen-backlog-overflow/</guid><description>&lt;h1 id="php-fpm-listen-backlog-overflow-the-kernel-silently-dropping-connections">PHP-FPM listen backlog overflow: the kernel silently dropping connections&lt;/h1>
&lt;p>The signature: a PHP-FPM pool where every status page field looks healthy yet the web server sporadically returns 502s. Active processes are not pinned at &lt;code>pm.max_children&lt;/code>, the &lt;code>listen queue&lt;/code> field reads zero, the slow log is quiet, opcache hit rate is normal. The 502s cluster into short bursts that may or may not line up with known load events.&lt;/p>
&lt;p>The kernel is dropping incoming FastCGI connections because the listen backlog has filled. PHP-FPM has no visibility into these drops. Once the backlog is at capacity, the kernel refuses to enqueue new connections, and the status page cannot report a queue depth above the configured &lt;code>listen.backlog&lt;/code> value because those connections never enter a queue FPM could observe. The failure happens entirely below PHP-FPM&amp;rsquo;s instrumentation layer.&lt;/p></description></item><item><title>PHP-FPM listen queue growing: the earliest signal of saturation</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-listen-queue-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-listen-queue-growing/</guid><description>&lt;h1 id="php-fpm-listen-queue-growing-the-earliest-signal-of-saturation">PHP-FPM listen queue growing: the earliest signal of saturation&lt;/h1>
&lt;p>The listen queue is the kernel-managed socket backlog between the web server and PHP-FPM workers. When it carries a sustained non-zero depth, requests are arriving faster than workers can drain them. By the time the web server logs a 502, the queue has already overflowed and the kernel has started dropping connections.&lt;/p>
&lt;p>The signal is easy to miss for two reasons. First, the FPM status page field is a point-in-time snapshot: a 10-second poller can report 0 even while bursts spill into the queue and drain between polls. Second, on Unix domain sockets (the most common production transport), the status page always reports 0 for all three queue fields due to a long-standing PHP bug. You have to read the kernel&amp;rsquo;s Recv-Q via &lt;code>ss&lt;/code>, or poll at 1-second intervals, to see the real depth.&lt;/p></description></item><item><title>PHP-FPM listen.backlog tuning: somaxconn and version defaults</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-listen-backlog-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-listen-backlog-tuning/</guid><description>&lt;h1 id="php-fpm-listenbacklog-tuning-somaxconn-and-version-defaults">PHP-FPM listen.backlog tuning: somaxconn and version defaults&lt;/h1>
&lt;p>When the PHP-FPM pool saturates, the socket backlog is the buffer between &amp;ldquo;requests are waiting&amp;rdquo; and &amp;ldquo;requests are dropped.&amp;rdquo; Operators who hit 502 storms often reach for &lt;code>listen.backlog&lt;/code> first, bump it to a large number, reload PHP-FPM, and see no change. The reason is almost always one of two things: the kernel silently clamped the request to &lt;code>net.core.somaxconn&lt;/code>, or the backlog was never the bottleneck.&lt;/p></description></item><item><title>PHP-FPM memory leak: per-worker RSS climbing until the box runs out</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-memory-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-memory-leak/</guid><description>&lt;h1 id="php-fpm-memory-leak-per-worker-rss-climbing-until-the-box-runs-out">PHP-FPM memory leak: per-worker RSS climbing until the box runs out&lt;/h1>
&lt;p>PHP-FPM workers are slowly bloating. RSS climbs hour by hour, the box runs out of RAM, the OOM killer shoots workers (or the master), and a restart makes everything look fine again. Hours or days later, the cycle repeats.&lt;/p>
&lt;p>This is the slow-burn memory leak pattern, and it is almost always enabled by a single configuration value: &lt;code>pm.max_requests = 0&lt;/code>. With worker recycling disabled, every byte a worker fails to release accumulates indefinitely. The leak itself may live in your application code, in a C extension, or in the PHP runtime. The diagnosis is not the same as the fix.&lt;/p></description></item><item><title>PHP-FPM memory_limit vs worker RSS: why workers exceed the limit you set</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-memory-limit-vs-rss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-memory-limit-vs-rss/</guid><description>&lt;h1 id="php-fpm-memory_limit-vs-worker-rss-why-workers-exceed-the-limit-you-set">PHP-FPM memory_limit vs worker RSS: why workers exceed the limit you set&lt;/h1>
&lt;p>You set &lt;code>memory_limit = 256M&lt;/code> in php.ini. You check &lt;code>ps&lt;/code> and see PHP-FPM workers sitting at 350 MB, 400 MB, sometimes more. No PHP error fired. No &amp;ldquo;Allowed memory size exhausted&amp;rdquo; in the logs. The workers did not crash. They are just bigger than the limit you set.&lt;/p>
&lt;p>This is not a bug, and usually not a leak. It is the expected result of how PHP&amp;rsquo;s memory manager is scoped, which most operators learn only after being paged at 3 a.m. for an OOM kill they cannot explain.&lt;/p></description></item><item><title>PHP-FPM Monitoring</title><link>https://www.netdata.cloud/monitoring-101/phpfpm-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/phpfpm-monitoring/</guid><description>&lt;h2 id="php-fpm-monitoring">PHP-FPM Monitoring&lt;/h2>
&lt;h3 id="what-is-php-fpm">What Is PHP-FPM?&lt;/h3>
&lt;p>PHP-FPM (FastCGI Process Manager) is a PHP FastCGI implementation, primarily focused on managing heavy-loaded web applications that demand high performance and fast execution. It is an alternative PHP FastCGI implementation with some additional features useful for sites of any size, especially busier sites. &lt;a href="https://php-fpm.org/">Learn more about PHP-FPM&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-php-fpm-with-netdata">Monitoring PHP-FPM With Netdata&lt;/h3>
&lt;p>To effectively monitor PHP-FPM, you need real-time insights into its performance and workload. Netdata provides a comprehensive PHP-FPM monitoring tool capable of capturing in-depth metrics with minimal configuration. Netdata’s intuitive UI and detailed charts allow you to visualize PHP-FPM’s operation and pinpoint performance bottlenecks swiftly.&lt;/p></description></item><item><title>PHP-FPM monitoring checklist: the signals every production pool needs</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-monitoring-checklist/</guid><description>&lt;h1 id="php-fpm-monitoring-checklist-the-signals-every-production-pool-needs">PHP-FPM monitoring checklist: the signals every production pool needs&lt;/h1>
&lt;p>PHP-FPM is a process-based concurrency model. Each worker handles exactly one request at a time, so the worker pool is the binding constraint. When all workers are occupied, new requests queue in the socket backlog. Once that fills, the kernel drops connections silently. Every PHP-FPM incident is a story about worker capacity, worker health, or what workers are blocked on.&lt;/p>
&lt;p>This checklist organizes production signals into four maturity levels: survival, operational, mature, and expert. Use it as a gap audit, or as a triage guide during incidents when workers are exhausted or the site returns 502s.&lt;/p></description></item><item><title>PHP-FPM monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-monitoring-maturity-model/</guid><description>&lt;h1 id="php-fpm-monitoring-maturity-model-from-survival-to-expert">PHP-FPM monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>PHP-FPM uses a master-worker process model where each worker handles exactly one request at a time. Concurrent request capacity equals the active worker count, and the path from healthy to users-seeing-502 runs through worker slots, the socket backlog, and kernel connection drops. This is a four-level reference, from minimum liveness checks to the expert signals that explain phantom workers and cgroup OOM kills. Use it as an audit: read down the levels, mark which signals you already collect, and fill the first gap. Levels are cumulative: Level 2 assumes Level 1, Level 3 assumes Level 2.&lt;/p></description></item><item><title>PHP-FPM ondemand cold start: fork and warmup latency on the first request</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-ondemand-cold-start/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-ondemand-cold-start/</guid><description>&lt;h1 id="php-fpm-ondemand-cold-start-fork-and-warmup-latency-on-the-first-request">PHP-FPM ondemand cold start: fork and warmup latency on the first request&lt;/h1>
&lt;p>PHP-FPM&amp;rsquo;s &lt;code>ondemand&lt;/code> process manager exists to reclaim memory when a pool is doing nothing. The tradeoff is paid in latency on the first request after an idle period: the master must fork a worker, that worker must handle its first request, and if OPcache has no compiled bytecode for the requested scripts, PHP must parse and compile them from disk. On a warm &lt;code>dynamic&lt;/code> or &lt;code>static&lt;/code> pool, none of that happens. On an &lt;code>ondemand&lt;/code> pool that has been idle, all of it happens on the critical path of a single request.&lt;/p></description></item><item><title>PHP-FPM OPcache cold start: the CPU stampede after a restart or deploy</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-opcache-cold-start/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-opcache-cold-start/</guid><description>&lt;h1 id="php-fpm-opcache-cold-start-the-cpu-stampede-after-a-restart-or-deploy">PHP-FPM OPcache cold start: the CPU stampede after a restart or deploy&lt;/h1>
&lt;p>After a PHP-FPM restart, deploy, or graceful reload, CPU climbs, latency rises, and throughput drops for the first few minutes. The OPcache shared memory segment is empty, so every worker compiles PHP source from disk before it can execute. Under enough traffic, those slow compiles tie up workers long enough to cascade into pool exhaustion, listen queue growth, and 502 or 504 errors at the edge.&lt;/p></description></item><item><title>PHP-FPM OPcache hit rate below 99%: silent CPU and latency tax</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-opcache-hit-rate-low/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-opcache-hit-rate-low/</guid><description>&lt;h1 id="php-fpm-opcache-hit-rate-below-99-silent-cpu-and-latency-tax">PHP-FPM OPcache hit rate below 99%: silent CPU and latency tax&lt;/h1>
&lt;p>A PHP-FPM pool with a healthy OPcache sits above 99% hit rate after warmup. Every request that misses recompiles PHP source into bytecode: parse, compile, optimize, and store. That work costs CPU on the worker handling the request and adds latency. When the cache is full and cannot admit new scripts, every worker handling those uncached scripts pays the compile tax on every request.&lt;/p></description></item><item><title>PHP-FPM OPcache out of memory: oom_restarts and the recompile cliff</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-opcache-memory-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-opcache-memory-full/</guid><description>&lt;h1 id="php-fpm-opcache-out-of-memory-oom_restarts-and-the-recompile-cliff">PHP-FPM OPcache out of memory: oom_restarts and the recompile cliff&lt;/h1>
&lt;p>Hours or days of normal operation. Then CPU spikes across every PHP-FPM worker at once, request latency jumps uniformly, and throughput drops. When you check OPcache status, &lt;code>oom_restarts&lt;/code> has incremented and &lt;code>free_memory&lt;/code> is near zero. The shared memory segment filled, OPcache force-cleared the entire cache, and every worker is now recompiling PHP from disk simultaneously.&lt;/p>
&lt;p>This is the recompile cliff. OPcache has no LRU eviction. When it runs out of space it either restarts (clearing everything) or silently stops caching new scripts.&lt;/p></description></item><item><title>PHP-FPM OPcache thrashing: a full cache recompiling PHP on every request</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-opcache-thrashing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-opcache-thrashing/</guid><description>&lt;h1 id="php-fpm-opcache-thrashing-a-full-cache-recompiling-php-on-every-request">PHP-FPM OPcache thrashing: a full cache recompiling PHP on every request&lt;/h1>
&lt;p>Every PHP request is suddenly slow. Latency is up across every endpoint, not just one. CPU is pinned across every worker, not a subset. The PHP-FPM status page shows workers running, the listen queue may be empty, and &lt;code>pm.max_children&lt;/code> is not the binding constraint. A graceful reload helps for a minute, then the slowness returns. You are probably inside an OPcache thrash.&lt;/p></description></item><item><title>PHP-FPM OPcache wasted memory: deploy fragmentation and when to reset</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-opcache-wasted-memory/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-opcache-wasted-memory/</guid><description>&lt;h1 id="php-fpm-opcache-wasted-memory-deploy-fragmentation-and-when-to-reset">PHP-FPM OPcache wasted memory: deploy fragmentation and when to reset&lt;/h1>
&lt;p>OPcache&amp;rsquo;s &lt;code>wasted_memory&lt;/code> counter grows on every deploy and never goes down on its own. When you ship new PHP files without clearing the cache, the old compiled bytecode stays in the shared memory segment alongside the new bytecode, occupying space that cannot be reclaimed until a full reset. Over a day of frequent deploys, this fragmentation starves the cache of room for hot scripts.&lt;/p></description></item><item><title>PHP-FPM opcache.validate_timestamps: the development setting that costs production</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-opcache-validate-timestamps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-opcache-validate-timestamps/</guid><description>&lt;h1 id="php-fpm-opcachevalidate_timestamps-the-development-setting-that-costs-production">PHP-FPM opcache.validate_timestamps: the development setting that costs production&lt;/h1>
&lt;p>The default &lt;code>opcache.validate_timestamps = 1&lt;/code> is convenient for development: edit a file, refresh the browser, see the change immediately. In production, that same convenience becomes a per-request filesystem tax and, with frequent deploys, a silent source of OPcache fragmentation.&lt;/p>
&lt;p>The trade-off is not subtle. Leave &lt;code>validate_timestamps&lt;/code> enabled and PHP stats source files on every request (or at the interval set by &lt;code>revalidate_freq&lt;/code>), burning syscalls and CPU. Disable it and PHP serves cached bytecode forever, which means a deploy that does not explicitly clear OPcache will serve stale code to every user until someone intervenes. Neither option is free. The question is which cost you can control.&lt;/p></description></item><item><title>PHP-FPM open_basedir and disable_functions: sandboxing worker pools</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-open-basedir-disable-functions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-open-basedir-disable-functions/</guid><description>&lt;h1 id="php-fpm-open_basedir-and-disable_functions-sandboxing-worker-pools">PHP-FPM open_basedir and disable_functions: sandboxing worker pools&lt;/h1>
&lt;p>PHP-FPM runs each worker as an OS user. Whatever that user can read or execute, the PHP code inside the worker can read or execute. On a host running more than one application, or on any host where a compromise is plausible, that is a wide blast radius. A single uploaded webshell in one pool can read &lt;code>/etc/passwd&lt;/code>, walk into adjacent application directories for credentials, or shell out to enumerate the network.&lt;/p></description></item><item><title>PHP-FPM phpinfo() exposed in production: mapping the attack surface</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-phpinfo-exposed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-phpinfo-exposed/</guid><description>&lt;p>A reachable &lt;code>phpinfo()&lt;/code> page in production is not a direct code execution vulnerability, which is why teams often deprioritize it. But the output is a reconnaissance tool handed to an attacker for free: PHP version, every loaded extension, absolute file paths, environment variables (which frequently contain database passwords, API keys, and S3 credentials in containerized deployments), database connection settings, and internal network topology.&lt;/p>
&lt;p>Automated scanners probe for &lt;code>info.php&lt;/code>, &lt;code>phpinfo.php&lt;/code>, &lt;code>test.php&lt;/code>, &lt;code>debug.php&lt;/code>, and similar filenames continuously. If any return a 200 with the full PHP configuration page, the attacker has a detailed map of your runtime. From there, version-specific CVEs, extension-specific exploits, and credential reuse attacks become targeted rather than speculative.&lt;/p></description></item><item><title>PHP-FPM pm.max_requests: worker recycling as the memory-leak safety net</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-max-requests-recycling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-max-requests-recycling/</guid><description>&lt;h1 id="php-fpm-pmmax_requests-worker-recycling-as-the-memory-leak-safety-net">PHP-FPM pm.max_requests: worker recycling as the memory-leak safety net&lt;/h1>
&lt;p>PHP-FPM runs fine for hours or days, then per-worker RSS slowly climbs, the box runs low on RAM, and the OOM killer shoots workers or the master. Restarting PHP-FPM fixes it immediately, and the cycle repeats. The root cause is usually a memory leak in application code, an extension, or allocator fragmentation. The reason it becomes an incident instead of a slow nuisance is almost always the same missing safety net: &lt;code>pm.max_requests = 0&lt;/code>.&lt;/p></description></item><item><title>PHP-FPM pool running as root or a shared user: privilege and isolation risk</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-pool-running-as-root/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-pool-running-as-root/</guid><description>&lt;p>A single line of PHP-FPM configuration decides how far an attacker travels after a remote code execution bug in your application. The pool&amp;rsquo;s &lt;code>user&lt;/code> and &lt;code>group&lt;/code> directives set the identity of every worker that runs your code. That identity is the blast radius of a successful exploit.&lt;/p>
&lt;p>The detection commands below are read-only. The lockdown changes require a graceful reload (&lt;code>SIGUSR2&lt;/code> or &lt;code>systemctl reload php-fpm&lt;/code>), which re-reads the pool configuration and cycles workers while in-flight requests finish. Plan each change as a deploy.&lt;/p></description></item><item><title>PHP-FPM request duration climbing: spotting stuck and outlier workers</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-request-duration-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-request-duration-high/</guid><description>&lt;h1 id="php-fpm-request-duration-climbing-spotting-stuck-and-outlier-workers">PHP-FPM request duration climbing: spotting stuck and outlier workers&lt;/h1>
&lt;p>When you pull &lt;code>?full&lt;/code> from the PHP-FPM status page during an incident and the &lt;code>request duration&lt;/code> column looks alarming, verify your units first. The field is microseconds, not seconds or milliseconds, and it is the single most misread value in the FPM status surface. A worker showing &lt;code>4500000&lt;/code> has been running for 4.5 seconds, not 4.5 million seconds. Convert before you escalate.&lt;/p></description></item><item><title>PHP-FPM request_terminate_timeout: stopping stuck requests from eroding the pool</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-request-terminate-timeout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-request-terminate-timeout/</guid><description>&lt;h1 id="php-fpm-request_terminate_timeout-stopping-stuck-requests-from-eroding-the-pool">PHP-FPM request_terminate_timeout: stopping stuck requests from eroding the pool&lt;/h1>
&lt;p>PHP-FPM pools silently lose capacity when workers get stuck and never return to idle. The default &lt;code>request_terminate_timeout = 0&lt;/code> means a single deadlocked request, infinite loop, or hung database call occupies a worker forever. Over time, these stuck workers accumulate. The pool keeps running, requests keep being accepted, but effective concurrency shrinks with no error and no alert.&lt;/p>
&lt;p>The symptom is subtle. Stuck workers are still counted as active in the FPM status page, so the count looks normal. In static mode, &lt;code>active processes&lt;/code> sits at &lt;code>pm.max_children&lt;/code> with &lt;code>idle processes&lt;/code> at 0, but throughput is lower than the traffic should produce. The web server does not complain. The full status page reveals workers with request durations far exceeding your baseline. Eventually a traffic spike hits the reduced effective pool, the listen queue fills, and users see 502s. By the time the outage is visible, capacity has been eroding for hours or days.&lt;/p></description></item><item><title>PHP-FPM session lock contention: file sessions serializing a user's requests</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-session-lock-contention/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-session-lock-contention/</guid><description>&lt;p>A single user reports AJAX-heavy pages loading slowly, but the server has spare worker capacity and CPU is barely loaded. The PHP-FPM status page shows active workers climbing toward &lt;code>max_children&lt;/code>. You raise &lt;code>max_children&lt;/code> and nothing improves. The slow log, when configured, shows stack traces parked at &lt;code>session_start()&lt;/code>.&lt;/p>
&lt;p>This is session lock contention. PHP&amp;rsquo;s default file-based session handler acquires an exclusive &lt;code>flock(LOCK_EX)&lt;/code> on the session file at &lt;code>session_start()&lt;/code> and holds it until the script ends or &lt;code>session_write_close()&lt;/code> is called. When one browser session makes concurrent requests (parallel AJAX calls, SPA data fetching, long-polling, upload progress checks), those requests serialize completely behind that single lock.&lt;/p></description></item><item><title>PHP-FPM session_write_close and Redis sessions: fixing session serialization</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-session-write-close/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-session-write-close/</guid><description>&lt;h1 id="php-fpm-session_write_close-and-redis-sessions-fixing-session-serialization">PHP-FPM session_write_close and Redis sessions: fixing session serialization&lt;/h1>
&lt;p>Same-user requests that should run in parallel are completing one after another. AJAX panels load sequentially, dashboards stall on parallel fetches, and the slow log shows workers blocked at &lt;code>session_start()&lt;/code>. The PHP-FPM pool has spare capacity and CPU is low. This is session lock serialization, and the fix is almost always one of two moves: release the file lock earlier with &lt;code>session_write_close()&lt;/code>, or move sessions to a backend whose locking semantics do not serialize the same user.&lt;/p></description></item><item><title>PHP-FPM sizing pm.max_children: by memory, not by CPU cores</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-sizing-max-children/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-sizing-max-children/</guid><description>&lt;h1 id="php-fpm-sizing-pmmax_children-by-memory-not-by-cpu-cores">PHP-FPM sizing pm.max_children: by memory, not by CPU cores&lt;/h1>
&lt;p>&lt;code>pm.max_children&lt;/code> is the hard ceiling on concurrent request processing in a PHP-FPM pool. Each worker handles exactly one request at a time. Once all workers are busy, new requests queue in the socket backlog. Once the backlog fills, the kernel drops connections and users see 502 errors.&lt;/p>
&lt;p>The common sizing mistake is anchoring this number to CPU core count. Teams pick 4 workers for a 4-core box, or 8 for an 8-core box, and assume they have sized correctly. They have not. PHP-FPM workers spend most of their lifetime blocked on I/O: database queries, external API calls, filesystem reads, session locks. While a worker waits on a slow database response, it holds a process slot but uses essentially zero CPU. A 4-core machine can run 50 to 200 workers if memory allows.&lt;/p></description></item><item><title>PHP-FPM slow log: turning on request_slowlog_timeout to see what is slow</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-slow-log-setup/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-slow-log-setup/</guid><description>&lt;h1 id="php-fpm-slow-log-turning-on-request_slowlog_timeout-to-see-what-is-slow">PHP-FPM slow log: turning on request_slowlog_timeout to see what is slow&lt;/h1>
&lt;p>When PHP-FPM workers pile up, the status page tells you they are busy. It does not tell you why. Active processes climb toward &lt;code>pm.max_children&lt;/code>, the listen queue fills, and you end up correlating timestamps against database slow query logs, APM traces, and application error logs to reconstruct what happened. The slow log closes that gap. When a request exceeds &lt;code>request_slowlog_timeout&lt;/code>, the FPM master ptraces the worker and writes a backtrace naming the exact script and call that blocked.&lt;/p></description></item><item><title>PHP-FPM slow request cascade: one slow dependency drains the whole pool</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-slow-request-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-slow-request-cascade/</guid><description>&lt;h1 id="php-fpm-slow-request-cascade-one-slow-dependency-drains-the-whole-pool">PHP-FPM slow request cascade: one slow dependency drains the whole pool&lt;/h1>
&lt;p>Your PHP-FPM pool is pinned at &lt;code>max_children&lt;/code>. The listen queue is climbing. Users see 502 Bad Gateway or 504 Gateway Timeout. CPU on the box is flat and the application error log is quiet. The PHP code is fine; it is waiting on a slow dependency.&lt;/p>
&lt;p>This is the slow request cascade. A database query regresses, an external API starts timing out, an NFS mount stalls, or DNS resolution hangs. A request that normally completes in 50ms now takes 5 to 30 seconds. Each slow request holds a worker for the full duration. The pool drains in seconds, the socket backlog fills, and the kernel starts refusing connections. This is the most common PHP-FPM failure mode in production, and the fix is almost never inside PHP-FPM itself.&lt;/p></description></item><item><title>PHP-FPM static vs dynamic vs ondemand: choosing a process manager mode</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-static-vs-dynamic-vs-ondemand/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-static-vs-dynamic-vs-ondemand/</guid><description>&lt;h1 id="php-fpm-static-vs-dynamic-vs-ondemand-choosing-a-process-manager-mode">PHP-FPM static vs dynamic vs ondemand: choosing a process manager mode&lt;/h1>
&lt;p>PHP-FPM&amp;rsquo;s &lt;code>pm&lt;/code> directive controls how the master process decides how many worker processes to keep alive. There are three modes: &lt;code>static&lt;/code>, &lt;code>dynamic&lt;/code>, and &lt;code>ondemand&lt;/code>. The choice changes how your pool reacts to a traffic burst, how much idle memory you pay for, and which status page counters are meaningful.&lt;/p>
&lt;p>Each worker handles exactly one request at a time, so worker count is your concurrency ceiling. The pm mode determines whether that ceiling is pre-allocated, scaled up reactively, or created on demand. Pick the wrong mode for your traffic pattern and you get one of three failure shapes: RAM burned on idle workers, latency spikes when the pool cannot fork fast enough, or a pool that scales down to zero and has to cold-start back up at the worst moment.&lt;/p></description></item><item><title>PHP-FPM status and ping page exposed: operational information disclosure</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-status-page-exposed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-status-page-exposed/</guid><description>&lt;h1 id="php-fpm-status-and-ping-page-exposed-operational-information-disclosure">PHP-FPM status and ping page exposed: operational information disclosure&lt;/h1>
&lt;p>The PHP-FPM status page and ping endpoint are operational tools meant for monitoring and health checks from inside the network. When reachable from the public internet, they leak data that aids reconnaissance: pool names, process manager configuration, worker counts, PIDs, request durations, and in full mode the exact script paths and query strings of every active request. This is a binary condition. If the status page returns HTTP 200 from an external IP, it is exposed and needs immediate restriction.&lt;/p></description></item><item><title>PHP-FPM total processes below max_children: workers not spawning</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-total-processes-mismatch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-total-processes-mismatch/</guid><description>&lt;h1 id="php-fpm-total-processes-below-max_children-workers-not-spawning">PHP-FPM total processes below max_children: workers not spawning&lt;/h1>
&lt;p>The PHP-FPM status page reports &lt;code>total processes&lt;/code> below &lt;code>pm.max_children&lt;/code>, and the count is not climbing. Whether this is a problem depends on which process manager mode the pool runs and whether traffic is present.&lt;/p>
&lt;p>In &lt;code>static&lt;/code> mode, &lt;code>total processes&lt;/code> must equal &lt;code>pm.max_children&lt;/code> at all times. Any shortfall means workers have died and the master has not replaced them, or the master cannot fork new ones. This is always abnormal.&lt;/p></description></item><item><title>PHP-FPM will not start: bind failures, config errors, and PID file problems</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-wont-start/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-wont-start/</guid><description>&lt;h1 id="php-fpm-will-not-start-bind-failures-config-errors-and-pid-file-problems">PHP-FPM will not start: bind failures, config errors, and PID file problems&lt;/h1>
&lt;p>&lt;code>systemctl start php-fpm&lt;/code> returns failed, the master process never appears in the process table, and the web server returns 502 for every PHP request. Workers are never forked. The failure happens during master initialization, before the &lt;code>ready to handle connections&lt;/code> notice.&lt;/p>
&lt;p>This is a control-plane failure, distinct from worker exhaustion or saturation cascades. Those affect a running master. Here, the usual FPM metrics (active processes, listen queue, max children reached) are unavailable because the status page and ping endpoint do not exist. The diagnostic path is different: read the first fatal log line.&lt;/p></description></item><item><title>PHP-FPM worker exhaustion: all workers busy and requests piling into the backlog</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-worker-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-worker-exhaustion/</guid><description>&lt;h1 id="php-fpm-worker-exhaustion-all-workers-busy-and-requests-piling-into-the-backlog">PHP-FPM worker exhaustion: all workers busy and requests piling into the backlog&lt;/h1>
&lt;p>PHP-FPM worker exhaustion is the most common PHP-FPM failure mode in production. The signature: &lt;code>active processes&lt;/code> equals &lt;code>pm.max_children&lt;/code>, the listen queue is growing, the &lt;code>max children reached&lt;/code> counter is incrementing, and the web server starts returning 502 and 504 responses.&lt;/p>
&lt;p>The counterintuitive part: CPU is often LOW. Workers spend most of their time waiting on I/O, not computing. When all workers block on a slow database query or an unresponsive external API, CPU drops while requests pile up. If you are judging saturation by CPU alone, you will miss the problem.&lt;/p></description></item><item><title>PHP-FPM worker exit rate: normal recycling vs abnormal deaths</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-worker-exit-rate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-worker-exit-rate/</guid><description>&lt;h1 id="php-fpm-worker-exit-rate-normal-recycling-vs-abnormal-deaths">PHP-FPM worker exit rate: normal recycling vs abnormal deaths&lt;/h1>
&lt;p>PHP-FPM workers exit constantly in a healthy pool. &lt;code>pm.max_requests&lt;/code> exists precisely to make workers exit on purpose after serving a fixed number of requests. The master forks a replacement, and the pool continues serving traffic. An exit rate of zero over hours means &lt;code>pm.max_requests&lt;/code> is set to 0 (unlimited), which means workers never recycle and any memory leak accumulates without bound.&lt;/p></description></item><item><title>PHP-FPM worker memory sizing: RSS, PSS, and honest capacity math</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-worker-memory-sizing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-worker-memory-sizing/</guid><description>&lt;h1 id="php-fpm-worker-memory-sizing-rss-pss-and-honest-capacity-math">PHP-FPM worker memory sizing: RSS, PSS, and honest capacity math&lt;/h1>
&lt;p>The standard PHP-FPM capacity formula is &lt;code>max_children = available_memory / average_worker_RSS&lt;/code>. It shows up in every tuning guide, and it overcounts unique memory by a predictable margin. Forked PHP-FPM workers share a large block of read-only memory (OPcache bytecode, shared libraries, interned strings), and tools like &lt;code>ps&lt;/code> count those shared pages in every worker&amp;rsquo;s RSS. Sum RSS across workers and divide, and you overcount by roughly 30 to 50 percent.&lt;/p></description></item><item><title>PHP-FPM workers busy at low CPU: blocked on the database, API, or filesystem</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-workers-blocked-on-io/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-workers-blocked-on-io/</guid><description>&lt;h1 id="php-fpm-workers-busy-at-low-cpu-blocked-on-the-database-api-or-filesystem">PHP-FPM workers busy at low CPU: blocked on the database, API, or filesystem&lt;/h1>
&lt;p>PHP-FPM &lt;code>active processes&lt;/code> is climbing toward &lt;code>pm.max_children&lt;/code>, the listen queue is building, users are seeing slow responses or intermittent 502s, and yet system CPU is flat at 10 to 20 percent. The pool looks saturated in terms of worker slots, but the CPU headroom suggests the box is barely working.&lt;/p>
&lt;p>Each PHP-FPM worker handles exactly one request at a time. When a worker blocks on I/O (a database query, an external API response, a DNS lookup, or a filesystem operation), it holds its worker slot for the entire duration of that wait while consuming essentially zero CPU. It is &amp;ldquo;active&amp;rdquo; from FPM&amp;rsquo;s scoreboard perspective but idle from the kernel scheduler&amp;rsquo;s perspective.&lt;/p></description></item><item><title>PHP-FPM workers OOM-killed: "child N exited on signal 9 (SIGKILL)" and the memory cliff</title><link>https://www.netdata.cloud/guides/php-fpm/php-fpm-oom-killed-workers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/php-fpm/php-fpm-oom-killed-workers/</guid><description>&lt;h1 id="php-fpm-workers-oom-killed-child-n-exited-on-signal-9-sigkill-and-the-memory-cliff">PHP-FPM workers OOM-killed: &amp;ldquo;child N exited on signal 9 (SIGKILL)&amp;rdquo; and the memory cliff&lt;/h1>
&lt;p>The PHP-FPM error log shows the same line: &lt;code>WARNING: [pool www] child 12345 exited on signal 9 (SIGKILL)&lt;/code>. No segfault, no stack trace, no &amp;ldquo;core dumped&amp;rdquo;. Just signal 9. Users see 502s, requests fail in batches, and reloading PHP-FPM makes it go away for a while. Hours or days later, it returns.&lt;/p>
&lt;p>Signal 9 is not a PHP crash. It is the kernel or cgroup OOM killer terminating the worker because the process exceeded available memory. The master sees the worker die, logs the SIGKILL, and forks a replacement. It has no idea why the worker was killed. The log line is a consequence, not a cause.&lt;/p></description></item><item><title>phpDaemon</title><link>https://www.netdata.cloud/integrations/data-collection/applications/phpdaemon/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/phpdaemon/</guid><description/></item><item><title>phpDaemon Monitoring</title><link>https://www.netdata.cloud/monitoring-101/phpdaemon-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/phpdaemon-monitoring/</guid><description>&lt;h2 id="phpdaemon-monitoring">phpDaemon Monitoring&lt;/h2>
&lt;h3 id="what-is-phpdaemon">What Is phpDaemon?&lt;/h3>
&lt;p>phpDaemon is an advanced asynchronous server-side daemon for PHP that can efficiently handle diverse application workloads by relying on an event-driven architecture. It is particularly suited for applications requiring persistent connections and rapid response times.&lt;/p>
&lt;h3 id="monitoring-phpdaemon-with-netdata">Monitoring phpDaemon With Netdata&lt;/h3>
&lt;p>Netdata offers a comprehensive monitoring solution for phpDaemon, showcasing a range of metrics in real time. By integrating the Netdata Agent with your phpDaemon instance, you can enjoy insightful visual data right at your fingertips, utilize historical insights, and diagnose issues seamlessly. &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/phpdaemon/">Read the phpDaemon collector documentation&lt;/a> to get started.&lt;/p></description></item><item><title>Physical and Logical Disk Performance Metrics</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/physical-and-logical-disk-performance-metrics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/physical-and-logical-disk-performance-metrics/</guid><description/></item><item><title>Pi-hole</title><link>https://www.netdata.cloud/integrations/data-collection/networking/pi-hole/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/pi-hole/</guid><description/></item><item><title>Pi-hole Monitoring</title><link>https://www.netdata.cloud/monitoring-101/pihole-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/pihole-monitoring/</guid><description>&lt;h2 id="pi-hole-monitoring">Pi-hole Monitoring&lt;/h2>
&lt;h3 id="what-is-pi-hole">What Is Pi-hole?&lt;/h3>
&lt;p>Pi-hole is a popular network-wide ad blocker that functions as a DNS sinkhole. It blocks unwanted content by intercepting DNS queries and preventing ads and trackers, enabling a cleaner and faster web browsing experience. Its efficiency and open-source nature make it a favorite among privacy-conscious users.&lt;/p>
&lt;h3 id="monitoring-pi-hole-with-netdata">Monitoring Pi-hole With Netdata&lt;/h3>
&lt;p>To ensure Pi-hole runs optimally, monitoring is essential. With &lt;a href="https://www.netdata.cloud">Netdata&lt;/a>, you gain real-time insights into your Pi-hole metrics, identifying trends and spot anomalies quickly. Netdata supports Pi-hole monitoring via its &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/pihole/">pi-hole collector module&lt;/a>, collecting various DNS statistics from the Pi-hole API 6.0, including total queries, blocked domains, and client information.&lt;/p></description></item><item><title>Picturetel Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/picturetel-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/picturetel-corporation-snmp-traps/</guid><description/></item><item><title>Pika</title><link>https://www.netdata.cloud/integrations/data-collection/databases/pika/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/pika/</guid><description/></item><item><title>Pika Monitoring</title><link>https://www.netdata.cloud/monitoring-101/pika-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/pika-monitoring/</guid><description>&lt;h2 id="pika-monitoring">Pika Monitoring&lt;/h2>
&lt;h3 id="what-is-pika">What Is Pika?&lt;/h3>
&lt;p>Pika is a NoSQL database server, designed for high concurrency and scalability, compatible with Redis protocol. Pika adds persistence to the Redis dataset and is engineered to support large datasets while maintaining excellent performance.&lt;/p>
&lt;h3 id="monitoring-pika-with-netdata">Monitoring Pika With Netdata&lt;/h3>
&lt;p>Using &lt;a href="https://app.netdata.cloud/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata&lt;/a> as your Pika monitoring tool provides real-time insights and comprehensive metrics from your Pika servers. By leveraging Netdata&amp;rsquo;s monitoring capabilities, you&amp;rsquo;ll gain visibility into system operations, enabling you to diagnose issues and optimize performance efficiently.&lt;/p></description></item><item><title>Pimoroni Enviro+</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/pimoroni-enviro+/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/pimoroni-enviro+/</guid><description/></item><item><title>Pimoroni Enviro+ Monitoring</title><link>https://www.netdata.cloud/monitoring-101/pimoroni_enviro_plus-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/pimoroni_enviro_plus-monitoring/</guid><description>&lt;h2 id="pimoroni-enviro-monitoring">Pimoroni Enviro+ Monitoring&lt;/h2>
&lt;h3 id="what-is-pimoroni-enviro">What Is Pimoroni Enviro+?&lt;/h3>
&lt;p>The Pimoroni Enviro+ is a sophisticated environmental monitoring device equipped to track air quality and other environmental data. Outfitted with various sensors, it allows developers and environmental enthusiasts to collect crucial data for analysis and decision-making. Its capabilities are particularly essential for IoT projects that require precise environmental readings.&lt;/p>
&lt;h3 id="monitoring-pimoroni-enviro-with-netdata">Monitoring Pimoroni Enviro+ With Netdata&lt;/h3>
&lt;p>When it comes to effectively monitor Pimoroni Enviro+, Netdata stands out as a robust monitoring tool. Netdata utilizes an openmetrics (Prometheus) exporter to seamlessly ingest data from the Pimoroni Enviro+. By using this method, you can avoid the hassle of setting up and maintaining a separate Prometheus server or Grafana dashboards. Netdata provides automated dashboards and alerting mechanisms, making the monitoring process straightforward and efficient.&lt;/p></description></item><item><title>Ping</title><link>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/ping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/ping/</guid><description/></item><item><title>Ping Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ping-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ping-monitoring/</guid><description>&lt;h2 id="ping-monitoring">Ping Monitoring&lt;/h2>
&lt;h3 id="what-is-ping">What Is Ping?&lt;/h3>
&lt;p>Ping is a network administration tool used to test the reachability of a host on an Internet Protocol (IP) network. It works by sending Internet Control Message Protocol (ICMP) Echo Request packets to the target host and waiting for a response. This basic utility is fundamental in diagnosing network connections and performance issues, making it a critical tool for IT professionals.&lt;/p>
&lt;h3 id="monitoring-ping-with-netdata">Monitoring Ping With Netdata&lt;/h3>
&lt;p>When it comes to monitoring ping, Netdata provides a comprehensive solution through its go.d.plugin with the ping module. This $name monitoring tool assesses network performance by measuring round-trip time (RTT) and packet loss, offering valuable insights into the health and reliability of your network connections. With Netdata, you can monitor ping in real-time, visualize trends, identify bottlenecks, and troubleshoot issues efficiently.&lt;/p></description></item><item><title>Plaintree Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/plaintree-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/plaintree-systems-inc-snmp-traps/</guid><description/></item><item><title>Planet Technology Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/planet-technology-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/planet-technology-corp-snmp-traps/</guid><description/></item><item><title>Platform Computing Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/platform-computing-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/platform-computing-corporation-snmp-traps/</guid><description/></item><item><title>Plexcom Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/plexcom-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/plexcom-inc-snmp-traps/</guid><description/></item><item><title>Podman</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/podman/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/podman/</guid><description/></item><item><title>Podman Containers</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/podman-containers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/podman-containers/</guid><description/></item><item><title>Podman Monitoring</title><link>https://www.netdata.cloud/monitoring-101/podman-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/podman-monitoring/</guid><description>&lt;h2 id="podman-monitoring">Podman Monitoring&lt;/h2>
&lt;h3 id="what-is-podman">What Is Podman?&lt;/h3>
&lt;p>Podman is a platform and service designed to manage and run containers from the CLI. It allows developers and IT administrators to create, deploy, and maintain containerized applications. Unlike Docker, Podman operates without a central daemon and can run in a rootless mode, offering more flexibility and security.&lt;/p>
&lt;h3 id="monitoring-podman-with-netdata">Monitoring Podman With Netdata&lt;/h3>
&lt;p>Monitoring Podman can be a game-changer for maintaining optimal performance of your containerized applications. Netdata provides a robust Podman monitoring tool powered by an openmetrics (Prometheus) exporter—&lt;a href="https://github.com/containers/prometheus-podman-exporter">Prometheus Podman Exporter&lt;/a>. Netdata can ingest data from any Prometheus exporter, enabling you to monitor Podman extensively without requiring a Prometheus server or Grafana. This functionality comes with ready-to-use dashboards and alerts for a seamless monitoring experience.&lt;/p></description></item><item><title>Polycom Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/polycom-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/polycom-inc-snmp-traps/</guid><description/></item><item><title>Postfix</title><link>https://www.netdata.cloud/integrations/data-collection/applications/postfix/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/postfix/</guid><description/></item><item><title>Postfix active queue saturation: hitting qmgr_message_active_limit</title><link>https://www.netdata.cloud/guides/postfix/postfix-active-queue-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-active-queue-saturation/</guid><description>&lt;h1 id="postfix-active-queue-saturation-hitting-qmgr_message_active_limit">Postfix active queue saturation: hitting qmgr_message_active_limit&lt;/h1>
&lt;p>Mail stops flowing. The deferred queue grows. CPU is low, network paths are healthy, and DNS resolves fine. You run &lt;code>postqueue -p&lt;/code> and see tens of thousands of messages in the active queue. The queue manager has hit &lt;code>qmgr_message_active_limit&lt;/code> and cannot schedule new deliveries.&lt;/p>
&lt;p>This is head-of-line blocking by design. The active queue has a hard ceiling (default 20,000 messages). Once reached, the queue manager stops scanning both the incoming and deferred queues. No new messages enter active delivery until an existing one is delivered or deferred. The system does not degrade gradually; it works, then it stops.&lt;/p></description></item><item><title>Postfix backscatter storm: bounces to forged senders and blocklisting</title><link>https://www.netdata.cloud/guides/postfix/postfix-backscatter-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-backscatter-storm/</guid><description>&lt;h1 id="postfix-backscatter-storm-bounces-to-forged-senders-and-blocklisting">Postfix backscatter storm: bounces to forged senders and blocklisting&lt;/h1>
&lt;p>Your mail queue is growing with bounces to envelope senders at domains you have never heard of. Within hours, your IP lands on a DNSBL and legitimate mail stops reaching its destination.&lt;/p>
&lt;p>This is a backscatter storm. Postfix accepted mail for recipients it could not deliver to, then generated non-delivery reports (NDRs) to the forged envelope sender. Those NDRs hit innocent third parties whose addresses the spammer forged.&lt;/p></description></item><item><title>Postfix bounce rate spike: 5xx failures, bad address lists, and reputation risk</title><link>https://www.netdata.cloud/guides/postfix/postfix-bounce-rate-spike/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-bounce-rate-spike/</guid><description>&lt;h1 id="postfix-bounce-rate-spike-5xx-failures-bad-address-lists-and-reputation-risk">Postfix bounce rate spike: 5xx failures, bad address lists, and reputation risk&lt;/h1>
&lt;p>A bounce rate spike is easy to miss. Postfix records a bounce as a completed transaction, so naive delivery metrics stay green while mail permanently fails. The &lt;code>status=bounced&lt;/code> entries pile up in the logs, sender reputation degrades with every message, and sustained high bounce rates trigger blocklist listings at major destinations.&lt;/p>
&lt;p>This guide covers how to identify a bounce rate spike, classify the 5xx failure codes driving it, and stop the bleeding before reputation damage compounds.&lt;/p></description></item><item><title>postfix check warnings: configuration drift and permission problems</title><link>https://www.netdata.cloud/guides/postfix/postfix-check-warnings-config-drift/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-check-warnings-config-drift/</guid><description>&lt;h1 id="postfix-check-warnings-configuration-drift-and-permission-problems">postfix check warnings: configuration drift and permission problems&lt;/h1>
&lt;p>&lt;code>postfix check&lt;/code> walks the Postfix file tree, verifies permissions against the &lt;code>postfix-files&lt;/code> manifest, and reports drift. It returns non-zero when it finds problems but does not stop or restart the running master. Safe to run in production at any time, which also makes the warnings easy to ignore.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>The expected filesystem state is defined in &lt;code>postfix-files&lt;/code>, which ships with the package and encodes ownership, group, and permission bits for every file and directory Postfix touches. When reality diverges from the manifest, you get a warning.&lt;/p></description></item><item><title>Postfix Connection refused: blocked port 25 and rejected outbound delivery</title><link>https://www.netdata.cloud/guides/postfix/postfix-connection-refused/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-connection-refused/</guid><description>&lt;h1 id="postfix-connection-refused-blocked-port-25-and-rejected-outbound-delivery">Postfix Connection refused: blocked port 25 and rejected outbound delivery&lt;/h1>
&lt;p>You see &lt;code>status=deferred (connect to mx.example.com[192.0.2.1]:25: Connection refused)&lt;/code> filling your mail logs. The deferred queue is growing. Postfix cannot establish outbound TCP connections to destination mail servers.&lt;/p>
&lt;p>&amp;ldquo;Connection refused&amp;rdquo; is not the same as &amp;ldquo;Connection timed out.&amp;rdquo; Connection refused means the TCP handshake was actively rejected with a RST packet: the destination IP is reachable, but nothing is listening on that port, or something in the path is rejecting the connection. Connection timed out means the SYN packet was silently dropped: a firewall rule, a network partition, or an ISP dropping egress traffic.&lt;/p></description></item><item><title>Postfix connection storm: dictionary attacks, anvil rate limits, and connection floods</title><link>https://www.netdata.cloud/guides/postfix/postfix-connection-storm-dictionary-attack/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-connection-storm-dictionary-attack/</guid><description>&lt;h1 id="postfix-connection-storm-dictionary-attacks-anvil-rate-limits-and-connection-floods">Postfix connection storm: dictionary attacks, anvil rate limits, and connection floods&lt;/h1>
&lt;p>A flood of inbound SMTP connections hits your Postfix server. The smtpd process count climbs, file descriptors tighten, and legitimate mail starts queuing behind connection churn. Throughput drops even though the system is busy. The likely cause is a connection storm: a dictionary attack harvesting recipient addresses, a credential-stuffing run against your submission port, or a misbehaving client with a broken connection pool.&lt;/p></description></item><item><title>Postfix Connection timed out: delivery deferrals to unreachable destinations</title><link>https://www.netdata.cloud/guides/postfix/postfix-connection-timed-out/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-connection-timed-out/</guid><description>&lt;h1 id="postfix-connection-timed-out-delivery-deferrals-to-unreachable-destinations">Postfix Connection timed out: delivery deferrals to unreachable destinations&lt;/h1>
&lt;p>The mail log shows the same deferral line for every message to one or more destinations:&lt;/p>
&lt;pre tabindex="0">&lt;code>status=deferred (connect to mx.example.com[192.0.2.1]:25: Connection timed out)
&lt;/code>&lt;/pre>&lt;p>Postfix sent a TCP SYN to the destination MX on port 25 and received no SYN-ACK within &lt;code>smtp_connect_timeout&lt;/code> (default 30 seconds). Something on the network path silently dropped the SYN. Postfix defers the message and schedules a retry with exponential backoff.&lt;/p></description></item><item><title>Postfix content_filter backpressure: incoming queue growth when Amavis or Rspamd slows</title><link>https://www.netdata.cloud/guides/postfix/postfix-content-filter-backpressure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-content-filter-backpressure/</guid><description>&lt;h1 id="postfix-content_filter-backpressure-incoming-queue-growth-when-amavis-or-rspamd-slows">Postfix content_filter backpressure: incoming queue growth when Amavis or Rspamd slows&lt;/h1>
&lt;p>When your content filter (Amavis, Rspamd, or a commercial scanner) slows down or stalls, mail accumulates in Postfix&amp;rsquo;s incoming queue while the active queue stays small. This is the reverse of the more familiar queue gridlock pattern where a slow destination fills the active queue, and it confuses operators who expect active to be large whenever mail is stuck.&lt;/p></description></item><item><title>Postfix daemon memory growth: leaking RSS and OOM risk</title><link>https://www.netdata.cloud/guides/postfix/postfix-daemon-memory-growth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-daemon-memory-growth/</guid><description>&lt;h1 id="postfix-daemon-memory-growth-leaking-rss-and-oom-risk">Postfix daemon memory growth: leaking RSS and OOM risk&lt;/h1>
&lt;p>Postfix daemons are disposable. Each smtpd, smtp, cleanup, or local process handles one session or message, then exits, keeping per-process RSS bounded at roughly 1 to 5 MB. When RSS climbs steadily on a qmgr process or a delivery agent that should have been recycled hours ago, you are looking at either a memory leak or a queue structure grown beyond normal operating size.&lt;/p></description></item><item><title>Postfix deferred queue growing: why mail piles up and how to drain it</title><link>https://www.netdata.cloud/guides/postfix/postfix-deferred-queue-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-deferred-queue-growing/</guid><description>&lt;h1 id="postfix-deferred-queue-growing-why-mail-piles-up-and-how-to-drain-it">Postfix deferred queue growing: why mail piles up and how to drain it&lt;/h1>
&lt;p>Sustained growth in &lt;code>/var/spool/postfix/deferred/&lt;/code> means delivery failures are outpacing retry successes. The queue is accumulating messages faster than the queue manager can drain them.&lt;/p>
&lt;p>The problem compounds in two ways. First, every message that fails retry stays in the deferred queue while fresh mail continues to enter the system. Second, the queue manager uses exponential backoff, so older deferred messages may not be retried for up to &lt;code>maximal_backoff_time&lt;/code> (default 4000s, roughly 66 minutes). Even after you fix the root cause, the queue takes hours to drain because most messages are in their backoff cool-off period.&lt;/p></description></item><item><title>Postfix destination concurrency limit: tuning per-destination delivery</title><link>https://www.netdata.cloud/guides/postfix/postfix-destination-concurrency-limit/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-destination-concurrency-limit/</guid><description>&lt;h1 id="postfix-destination-concurrency-limit-tuning-per-destination-delivery">Postfix destination concurrency limit: tuning per-destination delivery&lt;/h1>
&lt;p>Postfix does not open unlimited outbound connections to a single destination. The queue manager (&lt;code>qmgr&lt;/code>) enforces a per-destination concurrency cap controlling how many simultaneous delivery attempts may target the same recipient domain at once. This parameter, &lt;code>smtp_destination_concurrency_limit&lt;/code>, is behind two common operational failures: queue gridlock from one slow destination monopolizing active queue slots, and reputation penalties from hammering a fragile destination that then throttles or blocklists your IP.&lt;/p></description></item><item><title>Postfix DNS resolver failure: when a broken resolver defers mail to everyone</title><link>https://www.netdata.cloud/guides/postfix/postfix-dns-resolver-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-dns-resolver-failure/</guid><description>&lt;h1 id="postfix-dns-resolver-failure-when-a-broken-resolver-defers-mail-to-everyone">Postfix DNS resolver failure: when a broken resolver defers mail to everyone&lt;/h1>
&lt;p>Mail is deferring to every destination. Gmail, Outlook, corporate partners, even your own relayhost. The master process is running, the queue manager is responsive, and the SMTP listener accepts connections. But the deferred queue keeps growing, and the deferral reasons point to DNS: &amp;ldquo;Host or domain name not found&amp;rdquo; or &amp;ldquo;Name service error for name=&amp;hellip; type=MX: Host not found, try again.&amp;rdquo;&lt;/p></description></item><item><title>Postfix double-bounce loop: MAILER-DAEMON mail multiplying in the queue</title><link>https://www.netdata.cloud/guides/postfix/postfix-double-bounce-mailer-daemon-loop/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-double-bounce-mailer-daemon-loop/</guid><description>&lt;h1 id="postfix-double-bounce-loop-mailer-daemon-mail-multiplying-in-the-queue">Postfix double-bounce loop: MAILER-DAEMON mail multiplying in the queue&lt;/h1>
&lt;p>The Postfix queue is growing fast and &lt;code>postqueue -p&lt;/code> output is dominated by messages from &lt;code>MAILER-DAEMON&lt;/code> or &lt;code>double-bounce@&amp;lt;your hostname&amp;gt;&lt;/code>. The count climbs steadily despite no corresponding increase in legitimate mail volume. Queue slots, disk I/O, and inodes are being consumed by bounce generation, and legitimate mail is getting delayed or stuck behind the noise.&lt;/p>
&lt;p>Left unchecked, the loop can exhaust inodes on the queue filesystem (each queued message is a separate file), starve the queue manager of active delivery slots, and delay legitimate mail for hours.&lt;/p></description></item><item><title>Postfix file descriptor limits: raising ulimit and systemd LimitNOFILE</title><link>https://www.netdata.cloud/guides/postfix/postfix-file-descriptor-limits/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-file-descriptor-limits/</guid><description>&lt;h1 id="postfix-file-descriptor-limits-raising-ulimit-and-systemd-limitnofile">Postfix file descriptor limits: raising ulimit and systemd LimitNOFILE&lt;/h1>
&lt;p>Postfix logs &lt;code>fatal: socket: Too many open files&lt;/code> or begins silently dropping connections when per-process file descriptor limits are too low. The default soft limit on many distributions is 1024, which is adequate for a development mail server but insufficient for any production MTA handling concurrent SMTP sessions and active queue processing simultaneously.&lt;/p>
&lt;p>The most common operator mistake is editing &lt;code>/etc/security/limits.conf&lt;/code>, restarting Postfix, and finding the limit unchanged. On systemd-managed distributions, services do not read PAM limits. The effective file descriptor limit for a systemd service comes from the unit file&amp;rsquo;s &lt;code>LimitNOFILE&lt;/code> directive, the global &lt;code>DefaultLimitNOFILE&lt;/code> in &lt;code>/etc/systemd/system.conf&lt;/code>, or the kernel&amp;rsquo;s compiled-in default. Editing only one layer leaves the limit silently unchanged.&lt;/p></description></item><item><title>Postfix flushing and clearing the deferred queue: postqueue and postsuper</title><link>https://www.netdata.cloud/guides/postfix/postfix-flush-clear-deferred-queue/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-flush-clear-deferred-queue/</guid><description>&lt;h1 id="postfix-flushing-and-clearing-the-deferred-queue-postqueue-and-postsuper">Postfix flushing and clearing the deferred queue: postqueue and postsuper&lt;/h1>
&lt;p>A deferred queue is growing and you need to act. The difference between using &lt;code>postqueue&lt;/code> and &lt;code>postsuper&lt;/code> well and badly is the difference between controlled recovery and a self-inflicted stampede.&lt;/p>
&lt;p>Postfix holds messages that failed temporary delivery in the deferred queue for retry with exponential backoff. The retry scheduler uses &lt;code>minimal_backoff_time&lt;/code> (default 300s) and &lt;code>maximal_backoff_time&lt;/code> (default 4000s) to space out attempts. When you intervene manually, you are overriding that scheduler. Sometimes that is the right call, like when a relay host has recovered and the mail is now deliverable. Often it is not.&lt;/p></description></item><item><title>Postfix greylisting delays: 450 4.7.1 deferrals and slow first delivery</title><link>https://www.netdata.cloud/guides/postfix/postfix-greylisting-delays/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-greylisting-delays/</guid><description>&lt;h1 id="postfix-greylisting-delays-450-471-deferrals-and-slow-first-delivery">Postfix greylisting delays: 450 4.7.1 deferrals and slow first delivery&lt;/h1>
&lt;p>Mail to Gmail, Yahoo, Hotmail, and other major destinations shows a consistent delay of roughly 300 seconds on first delivery. The mail log shows &lt;code>450 4.7.1&lt;/code> or &lt;code>451 4.7.1&lt;/code> deferrals from the destination MX. The deferred queue climbs slowly but not explosively. On retry, the messages deliver successfully. This is greylisting, not a Postfix fault.&lt;/p>
&lt;p>Greylisting intentionally rejects the first delivery attempt from an unknown sender triplet (client IP, sender address, recipient address) with a 4xx temporary failure. The destination expects legitimate MTAs to retry and spam bots to give up. Postfix handles this correctly by deferring the message and scheduling a retry. The goal is to align retry timing with the destination&amp;rsquo;s greylisting window, and to whitelist trusted senders if you run postgrey for inbound filtering.&lt;/p></description></item><item><title>Postfix HELO command rejected: need fully-qualified hostname</title><link>https://www.netdata.cloud/guides/postfix/postfix-helo-hostname-rejected/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-helo-hostname-rejected/</guid><description>&lt;p>You see lines like this filling your mail log:&lt;/p>
&lt;pre tabindex="0">&lt;code>NOQUEUE: reject: RCPT from unknown[203.0.113.45]: 504 5.5.2 &amp;lt;desktop-abc123&amp;gt;: Helo command rejected: need fully-qualified hostname; from=&amp;lt;user@example.com&amp;gt; to=&amp;lt;recipient@example.org&amp;gt; proto=ESMTP helo=&amp;lt;desktop-abc123&amp;gt;
&lt;/code>&lt;/pre>&lt;p>Postfix is rejecting the connection because the client sent a bare hostname in its HELO or EHLO command instead of a fully-qualified domain name. The restriction responsible is &lt;code>reject_non_fqdn_helo_hostname&lt;/code> in your &lt;code>smtpd_helo_restrictions&lt;/code>.&lt;/p>
&lt;p>Most rejections are legitimate: bots and spam scripts routinely send garbage or bare names in HELO. The problem starts when a legitimate client gets caught &amp;ndash; a desktop mail client sending its machine name, an internal monitoring server using a short hostname, or an application hardcoded with &lt;code>localhost&lt;/code> as its HELO string. All trigger the same 504 rejection.&lt;/p></description></item><item><title>Postfix Host or domain name not found: DNS name service errors deferring mail</title><link>https://www.netdata.cloud/guides/postfix/postfix-host-not-found-name-service-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-host-not-found-name-service-error/</guid><description>&lt;h1 id="postfix-host-or-domain-name-not-found-dns-name-service-errors-deferring-mail">Postfix Host or domain name not found: DNS name service errors deferring mail&lt;/h1>
&lt;p>Mail is piling up in the deferred queue. Postfix logs the same deferral repeatedly:&lt;/p>
&lt;pre tabindex="0">&lt;code>status=deferred (Host or domain name not found. Name service error for name=example.com type=MX: Host not found, try again)
&lt;/code>&lt;/pre>&lt;p>The enhanced status code is 4.4.3 (directory server failure). Postfix will retry, but retries keep producing the same error because the resolver is broken, not the destination.&lt;/p></description></item><item><title>Postfix inode exhaustion: the queue filesystem blind spot behind No space left on device</title><link>https://www.netdata.cloud/guides/postfix/postfix-inode-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-inode-exhaustion/</guid><description>&lt;h1 id="postfix-inode-exhaustion-the-queue-filesystem-blind-spot-behind-no-space-left-on-device">Postfix inode exhaustion: the queue filesystem blind spot behind No space left on device&lt;/h1>
&lt;p>Postfix reports &lt;code>No space left on device&lt;/code> in your logs. You check &lt;code>df -h /var/spool/postfix&lt;/code> and see 50% free space. Mail acceptance is failing and queue operations are erroring. The obvious explanation does not match reality.&lt;/p>
&lt;p>The cause is almost always inode exhaustion on the queue filesystem. Postfix creates one file per queue entry: every incoming message, every deferred retry, every bounce notification. Each file consumes one inode regardless of how few bytes it occupies. A deferred queue with 500,000 tiny messages might use only a couple of GB of block space but exhaust every inode on the partition.&lt;/p></description></item><item><title>Postfix IP blocklisted: deliverability collapse and sender reputation</title><link>https://www.netdata.cloud/guides/postfix/postfix-blocklisted-ip-reputation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-blocklisted-ip-reputation/</guid><description>&lt;h1 id="postfix-ip-blocklisted-deliverability-collapse-and-sender-reputation">Postfix IP blocklisted: deliverability collapse and sender reputation&lt;/h1>
&lt;p>A blocklist entry is a symptom, not the root cause. Your sending IP appeared on Spamhaus, SpamCop, SORBS, or a provider-internal list because something in your mail stream triggered their detection. Delisting requests fail or the IP is relisted within hours until you identify and stop that activity.&lt;/p>
&lt;p>The first sign is usually deferrals spiking against one provider, bounces climbing, and the deferred queue growing while injection holds steady. Two patterns dominate: sudden listing from a compromised account or spam outbreak flooding garbage through your server, and gradual erosion from chronic issues like bad recipient lists, backscatter, or weak authentication accumulating negative reputation signals over weeks until a threshold is crossed. The diagnostic path differs for each.&lt;/p></description></item><item><title>Postfix mail flow: injection rate outpacing delivery rate</title><link>https://www.netdata.cloud/guides/postfix/postfix-mail-flow-injection-vs-delivery/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-mail-flow-injection-vs-delivery/</guid><description>&lt;h1 id="postfix-mail-flow-injection-rate-outpacing-delivery-rate">Postfix mail flow: injection rate outpacing delivery rate&lt;/h1>
&lt;p>The queue is growing, delivery latency is climbing, and the gap between what Postfix accepts and what it delivers is sustained. This divergence surfaces before queue-depth alerts fire.&lt;/p>
&lt;p>Injection rate and delivery rate should track within about 10% over any 5-minute window. When they diverge consistently, the outbound path is constrained. The constraint might be a single slow destination monopolizing active queue slots, a DNS resolver failure silently deferring all deliveries, a content filter that has stopped responding, or destination-side rate limiting. Each cause has a distinct signature in the logs and queue directories.&lt;/p></description></item><item><title>Postfix mail for domain loops back to myself: MX and myhostname misconfiguration</title><link>https://www.netdata.cloud/guides/postfix/postfix-mail-loops-back-to-myself/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-mail-loops-back-to-myself/</guid><description>&lt;h1 id="postfix-mail-for-domain-loops-back-to-myself-mx-and-myhostname-misconfiguration">Postfix mail for domain loops back to myself: MX and myhostname misconfiguration&lt;/h1>
&lt;p>The bounce message &amp;ldquo;mail for example.com loops back to myself&amp;rdquo; is a routing error, not a delivery timeout. It appears in logs as &lt;code>status=bounced (mail for example.com loops back to myself)&lt;/code> with DSN 5.4.6. Postfix resolved the destination domain&amp;rsquo;s MX record, found the best-preference MX points at this host, but found no matching entry in &lt;code>mydestination&lt;/code>, &lt;code>virtual_mailbox_domains&lt;/code>, &lt;code>virtual_alias_domains&lt;/code>, or &lt;code>relay_domains&lt;/code>. Instead of looping, Postfix bounces the message.&lt;/p></description></item><item><title>Postfix maildrop queue growing: pickup daemon and local submission failures</title><link>https://www.netdata.cloud/guides/postfix/postfix-maildrop-queue-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-maildrop-queue-growing/</guid><description>&lt;h1 id="postfix-maildrop-queue-growing-pickup-daemon-and-local-submission-failures">Postfix maildrop queue growing: pickup daemon and local submission failures&lt;/h1>
&lt;p>Files are accumulating in &lt;code>/var/spool/postfix/maildrop/&lt;/code>. Cron output is not arriving in mailboxes. Monitoring alerts that depend on local mail delivery are silently failing. &lt;code>postqueue -p&lt;/code> shows normal SMTP traffic flowing, but locally submitted messages are stuck in a directory that should be transient.&lt;/p>
&lt;p>The maildrop queue is the only Postfix queue that does not receive mail from the network. It holds messages submitted locally through the &lt;code>sendmail&lt;/code> command or the &lt;code>postdrop&lt;/code> setgid helper. Every cron job that produces output, every monitoring script that sends mail, and every local application that invokes &lt;code>/usr/sbin/sendmail&lt;/code> writes here. The single-threaded pickup daemon drains it within seconds, handing each message to cleanup for header rewriting and content normalization before it enters the incoming queue.&lt;/p></description></item><item><title>Postfix master process not running: the whole MTA is down</title><link>https://www.netdata.cloud/guides/postfix/postfix-master-process-not-running/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-master-process-not-running/</guid><description>&lt;h1 id="postfix-master-process-not-running-the-whole-mta-is-down">Postfix master process not running: the whole MTA is down&lt;/h1>
&lt;p>When the Postfix master process is not running, the entire MTA is down. No mail is received on port 25. No delivery agents are spawned. The queue manager cannot run. Local submissions via sendmail queue up but go nowhere. Upstream senders either timeout, defer, or bounce.&lt;/p>
&lt;p>Symptoms are uniform regardless of cause: &lt;code>postqueue -p&lt;/code> returns a fatal error, SMTP connections to port 25 are refused, and no &lt;code>master&lt;/code> process appears in the process table. The underlying cause varies: OOM kill, stale lock file from an unclean shutdown, port conflict, configuration error, or PID namespace confusion in a container.&lt;/p></description></item><item><title>Postfix message size exceeds fixed limit: 552 5.3.4 and message_size_limit</title><link>https://www.netdata.cloud/guides/postfix/postfix-message-size-exceeds-limit/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-message-size-exceeds-limit/</guid><description>&lt;h1 id="postfix-message-size-exceeds-fixed-limit-552-534-and-message_size_limit">Postfix message size exceeds fixed limit: 552 5.3.4 and message_size_limit&lt;/h1>
&lt;p>The error &lt;code>552 5.3.4 Message size exceeds fixed limit&lt;/code> appears in Postfix logs when a message, including headers, body, and MIME encoding overhead, exceeds &lt;code>message_size_limit&lt;/code>. This is a hard rejection, not a deferral. The sender receives a bounce, and the message never enters the queue.&lt;/p>
&lt;p>The default &lt;code>message_size_limit&lt;/code> is 10240000 bytes (approximately 10 MB). Base64 encoding inflates binary attachments by roughly 33 percent, so a 7 MB attachment becomes approximately 9.3 MB of encoded data before headers are added. Many operators discover this only when users report that &amp;ldquo;sometimes attachments work, sometimes they don&amp;rsquo;t.&amp;rdquo;&lt;/p></description></item><item><title>Postfix Milter timeouts: smtpd hangs when a milter stops responding</title><link>https://www.netdata.cloud/guides/postfix/postfix-milter-timeouts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-milter-timeouts/</guid><description>&lt;h1 id="postfix-milter-timeouts-smtpd-hangs-when-a-milter-stops-responding">Postfix Milter timeouts: smtpd hangs when a milter stops responding&lt;/h1>
&lt;p>When a before-queue milter like OpenDKIM, milter-greylist, or a custom policy daemon stops responding, smtpd processes block during the SMTP transaction. Each connection that hits a milter callback waits for the full timeout duration before Postfix gives up. During that wait, the smtpd process handling that connection is occupied and counts toward the service&amp;rsquo;s maxproc limit.&lt;/p>
&lt;p>The first visible symptom is usually not mail bouncing or deferring. It is mail acceptance slowing down. SMTP clients connect, get a 220 greeting, then stall during the DATA phase. As more smtpd processes accumulate waiting for milter responses, fewer remain available for new connections. Eventually new connections queue in the listen backlog or are refused entirely.&lt;/p></description></item><item><title>Postfix Monitoring</title><link>https://www.netdata.cloud/monitoring-101/postfix-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/postfix-monitoring/</guid><description>&lt;h2 id="postfix-monitoring">Postfix Monitoring&lt;/h2>
&lt;h3 id="what-is-postfix">What Is Postfix?&lt;/h3>
&lt;p>Postfix is a well-known open-source mail transfer agent (MTA) used by systems administrators to manage and route email communications across networks. It is reputed for its robustness and efficiency, making it a popular choice among users looking to handle mail server tasks effectively. &lt;a href="https://www.postfix.org/">Learn more about Postfix&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-postfix-with-netdata">Monitoring Postfix With Netdata&lt;/h3>
&lt;p>To effectively monitor Postfix, leveraging a real-time monitoring tool like Netdata is invaluable. By utilizing &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/postfix/">Netdata&amp;rsquo;s collector for Postfix&lt;/a>, you can access critical insights about your mail server&amp;rsquo;s activities. The integration is seamless, allowing you to diagnose issues promptly and maintain the health of your email infrastructure. To see it in action, &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">check out the Live Demo&lt;/a>.&lt;/p></description></item><item><title>Postfix monitoring checklist: the signals every production mail server needs</title><link>https://www.netdata.cloud/guides/postfix/postfix-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-monitoring-checklist/</guid><description>&lt;h1 id="postfix-monitoring-checklist-the-signals-every-production-mail-server-needs">Postfix monitoring checklist: the signals every production mail server needs&lt;/h1>
&lt;p>A production Postfix server fails in ways that look healthy until they suddenly do not. Deferred queues fill with retrying messages while disk space appears fine. Inodes exhaust on a filesystem that &lt;code>df&lt;/code> reports at 50% free. A single slow destination silently monopolizes the active queue, starving all other mail. Outbound TLS certificates expire on connections nobody monitors, and deliveries start deferring with cryptic SSL errors.&lt;/p></description></item><item><title>Postfix monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/postfix/postfix-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-monitoring-maturity-model/</guid><description>&lt;h1 id="postfix-monitoring-maturity-model-from-survival-to-expert">Postfix monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Postfix is a modular, queue-based MTA where mail flow depends on a chain of cooperating daemons, filesystem-backed queues, DNS resolution, and external filters. Monitoring maturity is not about collecting more metrics for their own sake. It is about closing the gap between what you can detect and what actually causes incidents: inode exhaustion hiding behind &amp;ldquo;No space left on device,&amp;rdquo; active queue saturation from one slow destination, silent filter backpressure, or deferred queues that grow for hours before anyone notices.&lt;/p></description></item><item><title>Postfix No space left on device: disk and inode exhaustion halting the queue</title><link>https://www.netdata.cloud/guides/postfix/postfix-no-space-left-on-device/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-no-space-left-on-device/</guid><description>&lt;h1 id="postfix-no-space-left-on-device-disk-and-inode-exhaustion-halting-the-queue">Postfix No space left on device: disk and inode exhaustion halting the queue&lt;/h1>
&lt;p>Postfix logs show &lt;code>fatal: ... No space left on device&lt;/code> or &lt;code>warning: not enough free space in mail queue&lt;/code>. You run &lt;code>df -h&lt;/code> and the queue filesystem has free space. Mail acceptance stalls, delivery stops, and the queue grows.&lt;/p>
&lt;p>The most common cause is inode exhaustion. Postfix creates one file per queued message across its queue subdirectories: maildrop, incoming, active, deferred, bounce, defer, hold, and corrupt. A mail storm, a backscatter loop, or months of accumulated deferred messages can exhaust millions of inodes while disk space barely moves. &lt;code>df -h&lt;/code> looks fine. &lt;code>df -i&lt;/code> shows 100%.&lt;/p></description></item><item><title>Postfix not listening on port 25 or 587: connection refused</title><link>https://www.netdata.cloud/guides/postfix/postfix-smtp-port-not-listening/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-smtp-port-not-listening/</guid><description>&lt;h1 id="postfix-not-listening-on-port-25-or-587-connection-refused">Postfix not listening on port 25 or 587: connection refused&lt;/h1>
&lt;p>&amp;ldquo;Connection refused&amp;rdquo; on port 25 or 587 means no process on the host is completing the TCP handshake on that port. This is distinct from a timeout, where something accepts the SYN but never responds. &amp;ldquo;Refused&amp;rdquo; means either no process is listening, or the kernel is actively rejecting the connection.&lt;/p>
&lt;p>The most common real-world causes: the master process is down, Postfix is bound to loopback only, the smtpd process pool is exhausted, or a firewall is intercepting traffic before Postfix sees it. Less common but worth checking: &lt;code>inet_protocols&lt;/code> mismatch on IPv6-disabled hosts, port conflicts with other MTAs, and postscreen handoff failures.&lt;/p></description></item><item><title>Postfix open relay: unauthorized relay through a misconfigured server</title><link>https://www.netdata.cloud/guides/postfix/postfix-open-relay-unauthorized/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-open-relay-unauthorized/</guid><description>&lt;h1 id="postfix-open-relay-unauthorized-relay-through-a-misconfigured-server">Postfix open relay: unauthorized relay through a misconfigured server&lt;/h1>
&lt;p>A Postfix server that accepts mail from unauthenticated, non-local clients and forwards it to arbitrary external domains is an open relay. Treat any confirmed occurrence as a PAGE-level security incident: automated scanners find open relays within hours, and a single relayed spam batch can trigger blocklist listings that take days to clear.&lt;/p>
&lt;p>The most common causes are configuration changes that widen trust: &lt;code>mynetworks&lt;/code> set too broadly, restrictions evaluated in the wrong order, a SASL backend that accepts empty credentials, or a backup-MX setup without recipient validation.&lt;/p></description></item><item><title>Postfix process limit reached: smtpd 'to limit' warnings and connection delays</title><link>https://www.netdata.cloud/guides/postfix/postfix-process-limit-to-limit/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-process-limit-to-limit/</guid><description>&lt;h1 id="postfix-process-limit-reached-smtpd-to-limit-warnings-and-connection-delays">Postfix process limit reached: smtpd &amp;rsquo;to limit&amp;rsquo; warnings and connection delays&lt;/h1>
&lt;p>The &lt;code>warning: service smtpd: ... to limit&lt;/code> message in your mail log means Postfix has run out of smtpd processes. The master daemon caps concurrent instances of each service at a maxproc value defined in master.cf. For smtpd, that cap defaults to 100, inherited from &lt;code>default_process_limit&lt;/code>. When all 100 smtpd processes are busy, new TCP connections on port 25 either queue in the kernel listen backlog or get refused.&lt;/p></description></item><item><title>Postfix queue gridlock: one slow destination stalls all mail</title><link>https://www.netdata.cloud/guides/postfix/postfix-queue-gridlock-slow-destination/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-queue-gridlock-slow-destination/</guid><description>&lt;p>Mail delivery has stopped for everyone, but nothing looks broken. CPU is low. Network throughput is low. The active queue is full, the deferred queue is growing, and none of the usual suspects (DNS failure, content filter backpressure, disk exhaustion) explain it. This is Postfix queue gridlock: a single slow or throttling destination has consumed the active queue, and the queue manager&amp;rsquo;s fair scheduler is letting it starve every other destination.&lt;/p></description></item><item><title>Postfix queue message age: the delivery latency queue depth cannot show</title><link>https://www.netdata.cloud/guides/postfix/postfix-queue-message-age/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-queue-message-age/</guid><description>&lt;h1 id="postfix-queue-message-age-the-delivery-latency-queue-depth-cannot-show">Postfix queue message age: the delivery latency queue depth cannot show&lt;/h1>
&lt;p>A queue of 10 messages whose oldest is 3 hours old is worse than 10,000 messages aged 5 seconds. Depth tells you volume. Age tells you latency. They answer different questions, and conflating them leads to missed incidents.&lt;/p>
&lt;p>Postfix exposes no built-in &amp;ldquo;oldest message age&amp;rdquo; metric. The showq daemon reports per-message arrival times to postqueue, but no aggregate age, no percentile, no histogram. Most monitoring setups track queue depth as a count and stop there. A flat depth line looks healthy. But if messages enter and leave the queue at the same rate, depth stays constant even when every message takes 4 hours to deliver instead of 4 seconds.&lt;/p></description></item><item><title>Postfix queue partition disk full: /var/spool/postfix out of space</title><link>https://www.netdata.cloud/guides/postfix/postfix-disk-space-queue-partition/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-disk-space-queue-partition/</guid><description>&lt;h1 id="postfix-queue-partition-disk-full-varspoolpostfix-out-of-space">Postfix queue partition disk full: /var/spool/postfix out of space&lt;/h1>
&lt;p>When the filesystem holding &lt;code>/var/spool/postfix&lt;/code> runs out of bytes, all mail I/O stops. The queue manager cannot write new queue files, cleanup cannot inject messages, and pickup cannot move mail from maildrop into incoming. Postfix may still listen on port 25 and accept TCP connections, but every accepted message eventually fails with &amp;ldquo;No space left on device&amp;rdquo; or is rejected by the SMTP server&amp;rsquo;s free-space check.&lt;/p></description></item><item><title>Postfix rate limited by destination: 421 and 4.7.1 throttling from major providers</title><link>https://www.netdata.cloud/guides/postfix/postfix-rate-limited-by-destination/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-rate-limited-by-destination/</guid><description>&lt;h1 id="postfix-rate-limited-by-destination-421-and-471-throttling-from-major-providers">Postfix rate limited by destination: 421 and 4.7.1 throttling from major providers&lt;/h1>
&lt;p>When a major email provider starts throttling your Postfix server, the symptom is distinctive: mail to that one provider defers while everything else flows. You see 421 &amp;ldquo;too many connections&amp;rdquo; or 4.7.1 rate-limit responses in your mail logs, the deferred queue fills with messages to that domain, and retries with exponential backoff make the pile worse. Other destinations deliver normally.&lt;/p></description></item><item><title>Postfix Relay access denied: 554 5.7.1 and blocked senders</title><link>https://www.netdata.cloud/guides/postfix/postfix-relay-access-denied/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-relay-access-denied/</guid><description>&lt;h1 id="postfix-relay-access-denied-554-571-and-blocked-senders">Postfix Relay access denied: 554 5.7.1 and blocked senders&lt;/h1>
&lt;p>The mail log shows a pattern like this:&lt;/p>
&lt;pre tabindex="0">&lt;code>NOQUEUE: reject: RCPT from unknown[10.0.2.15]: 554 5.7.1 &amp;lt;user@external.example&amp;gt;: Relay access denied; from=&amp;lt;app@internal.corp&amp;gt; to=&amp;lt;user@external.example&amp;gt; proto=ESMTP helo=&amp;lt;app01.internal.corp&amp;gt;
&lt;/code>&lt;/pre>&lt;p>A client connected to Postfix, sent MAIL FROM and RCPT TO for a recipient on an external domain, and Postfix rejected the message before queuing it. The recipient domain is not one Postfix considers local or a configured relay destination, and the client failed every authorization check. This is Postfix enforcing relay control exactly as designed.&lt;/p></description></item><item><title>Postfix relay recipient map stale or slow: rejected valid recipients and smtpd hangs</title><link>https://www.netdata.cloud/guides/postfix/postfix-relay-recipient-map-stale/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-relay-recipient-map-stale/</guid><description>&lt;h1 id="postfix-relay-recipient-map-stale-or-slow-rejected-valid-recipients-and-smtpd-hangs">Postfix relay recipient map stale or slow: rejected valid recipients and smtpd hangs&lt;/h1>
&lt;p>In a relay or gateway Postfix setup, &lt;code>relay_recipient_maps&lt;/code> validates recipients before mail enters the queue. When this map goes stale or responds slowly, you see one of two failure modes: valid recipients rejected with a 550 response (stale &lt;code>.db&lt;/code> file, no queue entry to investigate), or &lt;code>smtpd&lt;/code> processes hanging during RCPT TO because a network-backed lookup never returns (connection capacity drops, new connections queue or fail).&lt;/p></description></item><item><title>Postfix SASL authentication failed: brute force on submission and credential errors</title><link>https://www.netdata.cloud/guides/postfix/postfix-sasl-authentication-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-sasl-authentication-failed/</guid><description>&lt;h1 id="postfix-sasl-authentication-failed-brute-force-on-submission-and-credential-errors">Postfix SASL authentication failed: brute force on submission and credential errors&lt;/h1>
&lt;p>Postfix logs &amp;ldquo;SASL LOGIN authentication failed&amp;rdquo; or &amp;ldquo;SASL PLAIN authentication failed&amp;rdquo; when an SMTP client on port 587 (submission) or 465 (smtps) sends an AUTH command that the SASL backend rejects. Postfix does not verify credentials itself. It delegates to a backend via &lt;code>smtpd_sasl_type&lt;/code> (&lt;code>dovecot&lt;/code> or &lt;code>cyrus&lt;/code>), which may in turn proxy to LDAP, Active Directory, PAM, or SQL.&lt;/p></description></item><item><title>Postfix TLS certificate expiry: expired certs and handshake failures</title><link>https://www.netdata.cloud/guides/postfix/postfix-tls-certificate-expiry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-tls-certificate-expiry/</guid><description>&lt;h1 id="postfix-tls-certificate-expiry-expired-certs-and-handshake-failures">Postfix TLS certificate expiry: expired certs and handshake failures&lt;/h1>
&lt;p>An expired TLS certificate on a Postfix server breaks mail delivery in two distinct ways, and teams frequently detect only one. The inbound symptom is loud: clients that require STARTTLS fail handshakes, connections drop, and complaints arrive quickly. The outbound symptom is quiet: mail to destinations enforcing mutual TLS or DANE silently defers into the deferred queue, where it sits under Postfix&amp;rsquo;s increasing backoff schedule until someone notices the queue growing.&lt;/p></description></item><item><title>Postfix TLS downgrade: mandatory TLS falling back to plaintext</title><link>https://www.netdata.cloud/guides/postfix/postfix-tls-downgrade-plaintext/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-tls-downgrade-plaintext/</guid><description>&lt;h1 id="postfix-tls-downgrade-mandatory-tls-falling-back-to-plaintext">Postfix TLS downgrade: mandatory TLS falling back to plaintext&lt;/h1>
&lt;p>When mail to a destination configured for mandatory TLS is delivered in plaintext, that is a security incident, not a delivery problem. Postfix TLS security levels exist to prevent this: at &lt;code>encrypt&lt;/code> and above, the message should be deferred if TLS is unavailable, not silently downgraded. If plaintext delivery is happening where policy forbids it, the policy enforcement chain is broken.&lt;/p></description></item><item><title>Postfix TLS handshake failures: SSL_accept errors and cipher mismatches</title><link>https://www.netdata.cloud/guides/postfix/postfix-tls-handshake-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-tls-handshake-failures/</guid><description>&lt;h1 id="postfix-tls-handshake-failures-ssl_accept-errors-and-cipher-mismatches">Postfix TLS handshake failures: SSL_accept errors and cipher mismatches&lt;/h1>
&lt;p>&lt;code>SSL_accept error&lt;/code> in your mail logs means the TLS handshake on an inbound connection failed mid-negotiation. The remote SMTP client connected, STARTTLS was offered, and the OpenSSL handshake aborted before a session was established. Postfix either rejects the message (mandatory TLS) or silently falls back to plaintext (opportunistic TLS).&lt;/p>
&lt;p>Default Postfix does not log failed opportunistic TLS handshakes. Both &lt;code>smtpd_tls_loglevel&lt;/code> and &lt;code>smtp_tls_loglevel&lt;/code> default to 0. You see successful TLS connections but not the ones that failed and fell back to cleartext. You may be delivering mail in plaintext to destinations you believe require encryption, with the only visible symptom being a deferred queue growing with delivery delays.&lt;/p></description></item><item><title>Postfix too many open files: file descriptor exhaustion and refused connections</title><link>https://www.netdata.cloud/guides/postfix/postfix-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-too-many-open-files/</guid><description>&lt;h1 id="postfix-too-many-open-files-file-descriptor-exhaustion-and-refused-connections">Postfix too many open files: file descriptor exhaustion and refused connections&lt;/h1>
&lt;p>Postfix logs &amp;ldquo;too many open files&amp;rdquo; or &amp;ldquo;unable to fork.&amp;rdquo; Inbound SMTP connections fail. Queue files cannot be opened or written. The master process is still running, ports are still bound, and the disk has free space and inodes.&lt;/p>
&lt;p>This is file descriptor exhaustion: a Postfix process has reached its per-process soft limit on open file descriptors. Each open socket, queue file, and pipe handle counts toward the limit. On most Linux distributions the default soft limit is 1024, which is inadequate for a production MTA. Under enough concurrent load, a process hits the ceiling and Postfix fails.&lt;/p></description></item><item><title>Postfix User unknown in recipient table: 550 recipient address rejected</title><link>https://www.netdata.cloud/guides/postfix/postfix-user-unknown-recipient-rejected/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postfix/postfix-user-unknown-recipient-rejected/</guid><description>&lt;h1 id="postfix-user-unknown-in-recipient-table-550-recipient-address-rejected">Postfix User unknown in recipient table: 550 recipient address rejected&lt;/h1>
&lt;p>A sender gets &lt;code>550 5.1.1 &amp;lt;user@domain&amp;gt;: Recipient address rejected: User unknown in local recipient table&lt;/code> (or the virtual mailbox or relay variant). Postfix rejected the recipient at SMTP &lt;code>RCPT TO&lt;/code> time, before queueing, because the address was not found in the map for that address class. This validation is controlled by &lt;code>smtpd_reject_unlisted_recipient&lt;/code> (default: &lt;code>yes&lt;/code>).&lt;/p>
&lt;p>The problem is either that a legitimate recipient is missing from the map, or that the domain is in the wrong address class and Postfix is checking the wrong map. Before changing restrictions or disabling validation, identify which map Postfix consulted, test the lookup with &lt;code>postmap -q&lt;/code>, and reconcile the map against your authoritative directory. The reject string identifies the address class.&lt;/p></description></item><item><title>PostgreSQL</title><link>https://www.netdata.cloud/integrations/data-collection/databases/postgresql/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/postgresql/</guid><description/></item><item><title>PostgreSQL</title><link>https://www.netdata.cloud/integrations/exporters/postgresql/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/postgresql/</guid><description/></item><item><title>PostgreSQL ALTER TABLE blocked: zero-downtime DDL patterns</title><link>https://www.netdata.cloud/guides/postgres/postgres-alter-table-blocked/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-alter-table-blocked/</guid><description>&lt;h1 id="postgresql-alter-table-blocked-zero-downtime-ddl-patterns">PostgreSQL ALTER TABLE blocked: zero-downtime DDL patterns&lt;/h1>
&lt;p>An &lt;code>ALTER TABLE&lt;/code> to add a column or change a type hangs. Application queries time out. Connection pools saturate. What looked like a simple schema change becomes a production incident.&lt;/p>
&lt;p>By default, &lt;code>ALTER TABLE&lt;/code> acquires an &lt;code>ACCESS EXCLUSIVE&lt;/code> lock. It conflicts with every other lock mode, including &lt;code>ACCESS SHARE&lt;/code> held by a plain &lt;code>SELECT&lt;/code>. Once the DDL statement queues behind a blocker, every subsequent read and write on that table queues behind the waiting DDL. The table goes offline before the &lt;code>ALTER TABLE&lt;/code> executes any work.&lt;/p></description></item><item><title>PostgreSQL autovacuum blocked by long-running transaction: detection and fix</title><link>https://www.netdata.cloud/guides/postgres/postgres-autovacuum-blocked-by-long-transaction/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-autovacuum-blocked-by-long-transaction/</guid><description>&lt;h1 id="postgresql-autovacuum-blocked-by-long-running-transaction-detection-and-fix">PostgreSQL autovacuum blocked by long-running transaction: detection and fix&lt;/h1>
&lt;p>Autovacuum workers are active, but &lt;code>n_dead_tup&lt;/code> climbs and queries slow down as they scan dead tuples. Table bloat grows because &lt;code>VACUUM&lt;/code> cannot reclaim dead row versions: a long-running transaction, abandoned replication slot, or hot-standby feedback is pinning the &lt;strong>xmin horizon&lt;/strong> cluster-wide. The worker is not broken; it is blocked by an older transaction ID that must remain visible.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>PostgreSQL&amp;rsquo;s MVCC keeps old tuple versions in the table until &lt;code>VACUUM&lt;/code> removes them. &lt;code>VACUUM&lt;/code> cannot remove any tuple that might still be visible to an active transaction. The boundary is the &lt;strong>xmin horizon&lt;/strong>: the oldest transaction ID still active anywhere in the cluster.&lt;/p></description></item><item><title>PostgreSQL autovacuum not running: detection, causes, and fixes</title><link>https://www.netdata.cloud/guides/postgres/postgres-autovacuum-not-running/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-autovacuum-not-running/</guid><description>&lt;h1 id="postgresql-autovacuum-not-running-detection-causes-and-fixes">PostgreSQL autovacuum not running: detection, causes, and fixes&lt;/h1>
&lt;p>Table sizes grow faster than insert rates. Query latency creeps up on UPDATE-heavy workloads. &lt;code>pg_stat_user_tables.n_dead_tup&lt;/code> climbs while &lt;code>last_autovacuum&lt;/code> stays frozen. Autovacuum should clean this up, but it is not. Dead tuple accumulation degrades performance and can eventually trigger transaction-ID-wraparound shutdown. Detect why autovacuum is stalled, find the blocker, and fix it without making things worse.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>Autovacuum is a background subsystem that spawns workers to run &lt;code>VACUUM&lt;/code> and &lt;code>ANALYZE&lt;/code> based on table-level thresholds. When it works, it reclaims dead tuple space, updates the free space map, maintains the visibility map, and freezes old transaction IDs to prevent wraparound. When it stops, effects cascade: table and index bloat grow, sequential scans read more dead pages, the planner chooses worse plans as statistics stale, and database age advances toward the 2-billion-transaction hard limit.&lt;/p></description></item><item><title>PostgreSQL autovacuum tuning: per-table thresholds for high-churn workloads</title><link>https://www.netdata.cloud/guides/postgres/postgres-autovacuum-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-autovacuum-tuning/</guid><description>&lt;h1 id="postgresql-autovacuum-tuning-per-table-thresholds-for-high-churn-workloads">PostgreSQL autovacuum tuning: per-table thresholds for high-churn workloads&lt;/h1>
&lt;p>Default autovacuum settings target modest OLTP workloads. On a 500-million-row table, the global &lt;code>autovacuum_vacuum_scale_factor&lt;/code> of 0.2 means autovacuum ignores the table until dead tuples exceed twenty percent of the row count. For high-churn tables, that delay lets bloat accumulate, degrades indexes, slows scans, and pushes transaction ID age toward wraparound. PostgreSQL overrides these thresholds per table via storage parameters (reloptions), so a 50 GB hot table can run aggressive settings without saturating workers across a 500 MB reference table. This guide covers calculating overrides, sizing worker memory, and verifying vacuum keeps up without drowning the cluster in background I/O.&lt;/p></description></item><item><title>PostgreSQL backup strategy: pg_dump, pg_basebackup, and pgBackRest compared</title><link>https://www.netdata.cloud/guides/postgres/postgres-backup-strategy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-backup-strategy/</guid><description>&lt;h1 id="postgresql-backup-strategy-pg_dump-pg_basebackup-and-pgbackrest-compared">PostgreSQL backup strategy: pg_dump, pg_basebackup, and pgBackRest compared&lt;/h1>
&lt;p>Your backup tool determines your RPO, RTO, and whether you can restore to a point in time or only to the backup moment. This guide compares logical dumps via &lt;code>pg_dump&lt;/code>, built-in physical copies via &lt;code>pg_basebackup&lt;/code>, and the third-party tool pgBackRest. PostgreSQL 17 introduced native incremental physical backups. Operators must decide whether built-in capabilities are sufficient or whether to adopt alternatives such as Barman or WAL-G. &lt;!-- TODO: verify pgBackRest maintenance status and claimed April 2026 EOL -->&lt;/p></description></item><item><title>PostgreSQL blocking queries: finding the root blocker in a lock cascade</title><link>https://www.netdata.cloud/guides/postgres/postgres-blocking-queries/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-blocking-queries/</guid><description>&lt;h1 id="postgresql-blocking-queries-finding-the-root-blocker-in-a-lock-cascade">PostgreSQL blocking queries: finding the root blocker in a lock cascade&lt;/h1>
&lt;p>A query that normally finishes in milliseconds is now running for minutes. &lt;code>pg_stat_activity&lt;/code> shows a queue of sessions with &lt;code>wait_event_type = 'Lock'&lt;/code>. You identify one session holding the contested lock, but terminating it does not clear the queue. That session was itself blocked by another, which was blocked by another. Until you find the session at the head of the chain, the cascade continues.&lt;/p></description></item><item><title>PostgreSQL checkpoint storms: detection, causes, and tuning</title><link>https://www.netdata.cloud/guides/postgres/postgres-checkpoint-storms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-checkpoint-storms/</guid><description>&lt;h1 id="postgresql-checkpoint-storms-detection-causes-and-tuning">PostgreSQL checkpoint storms: detection, causes, and tuning&lt;/h1>
&lt;p>If query latency spikes and TPS drops coincide with your &lt;code>checkpoint_timeout&lt;/code> schedule or follow a large bulk load, you are likely hitting a checkpoint storm. PostgreSQL checkpoints guarantee that dirty buffers are on disk. In a healthy system, timed checkpoints occur at &lt;code>checkpoint_timeout&lt;/code> intervals and the background writer spreads that I/O. A storm happens when WAL generation hits &lt;code>max_wal_size&lt;/code> before the timeout, forcing a requested checkpoint that flushes most dirty buffers at once, saturating disk bandwidth and stalling queries.&lt;/p></description></item><item><title>PostgreSQL connection exhaustion: detection, diagnosis, and prevention</title><link>https://www.netdata.cloud/guides/postgres/postgres-connection-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-connection-exhaustion/</guid><description>&lt;h1 id="postgresql-connection-exhaustion-detection-diagnosis-and-prevention">PostgreSQL connection exhaustion: detection, diagnosis, and prevention&lt;/h1>
&lt;p>Application logs show &lt;code>FATAL: sorry, too many clients already&lt;/code>. Health checks are failing. A rolling deploy just finished, and now the database is rejecting connections. Because PostgreSQL uses one process per connection, every slot consumes memory and scheduler overhead. Once &lt;code>max_connections&lt;/code> is reached, the server refuses new backends entirely. Raising &lt;code>max_connections&lt;/code> without fixing the root cause increases memory pressure and context-switch thrashing. This guide covers how to distinguish a true capacity shortage from a leak or pool misconfiguration, how to recover safely, and how to prevent recurrence.&lt;/p></description></item><item><title>PostgreSQL connection refused: pg_hba, listen_addresses, and TCP diagnosis</title><link>https://www.netdata.cloud/guides/postgres/postgres-connection-refused/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-connection-refused/</guid><description>&lt;h1 id="postgresql-connection-refused-pg_hba-listen_addresses-and-tcp-diagnosis">PostgreSQL connection refused: pg_hba, listen_addresses, and TCP diagnosis&lt;/h1>
&lt;p>An application cannot reach PostgreSQL. The client reports either &amp;ldquo;Connection refused,&amp;rdquo; a hang until timeout, or &amp;ldquo;no pg_hba.conf entry.&amp;rdquo; These three symptoms point to different layers. Mixing them up leads to wasted restarts, overly broad firewall rules, or &lt;code>pg_hba.conf&lt;/code> edits that never take effect. Work through the transport layer first, then the network path, then the authorization layer.&lt;/p>
&lt;p>PostgreSQL defaults to binding only to the loopback interface. A fresh installation or container image rejects remote TCP attempts before &lt;code>pg_hba.conf&lt;/code> is consulted. Changing &lt;code>listen_addresses&lt;/code> requires a restart; changing &lt;code>pg_hba.conf&lt;/code> only requires a reload. The diagnostic sequence is: confirm the listener, confirm the path, then confirm the rule.&lt;/p></description></item><item><title>PostgreSQL dead tuples piling up: why autovacuum can't keep up</title><link>https://www.netdata.cloud/guides/postgres/postgres-dead-tuples-piling-up/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-dead-tuples-piling-up/</guid><description>&lt;h1 id="postgresql-dead-tuples-piling-up-why-autovacuum-cant-keep-up">PostgreSQL dead tuples piling up: why autovacuum can&amp;rsquo;t keep up&lt;/h1>
&lt;p>Table sizes grow faster than insert rates, and queries that scanned thousands of rows last week now scan millions. &lt;code>pg_stat_user_tables&lt;/code> shows &lt;code>n_dead_tup&lt;/code> climbing into the millions while &lt;code>last_autovacuum&lt;/code> is stale or absent. Autovacuum is running somewhere in the cluster, but it is losing the race.&lt;/p>
&lt;p>PostgreSQL creates dead tuples on every &lt;code>UPDATE&lt;/code> and &lt;code>DELETE&lt;/code>. Autovacuum reclaims them only when the dead tuple count crosses a threshold. On large or high-churn tables, the default threshold is far too conservative, and even when vacuum does fire, long-running transactions, worker starvation, or streaming replica feedback can prevent tuple removal. The result is bloat: wasted space, slower scans, and eventually transaction ID wraparound risk.&lt;/p></description></item><item><title>PostgreSQL deadlock detected: how to diagnose and prevent deadlocks</title><link>https://www.netdata.cloud/guides/postgres/postgres-deadlock-detected/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-deadlock-detected/</guid><description>&lt;h1 id="postgresql-deadlock-detected-how-to-diagnose-and-prevent-deadlocks">PostgreSQL deadlock detected: how to diagnose and prevent deadlocks&lt;/h1>
&lt;p>&lt;code>ERROR: deadlock detected&lt;/code> means PostgreSQL aborted one transaction to break a circular wait-for graph. The victim returns SQLSTATE &lt;code>40P01&lt;/code>; the application must retry it. Deadlocks are a safety mechanism, not a bug: they fire when concurrent transactions acquire locks in incompatible orders. Even a few per minute degrade user experience, burn retry budget, and mask deeper contention. This guide shows how to read the deadlock output, find the root cause, and stop the cycle.&lt;/p></description></item><item><title>PostgreSQL disk full: emergency recovery and root cause analysis</title><link>https://www.netdata.cloud/guides/postgres/postgres-disk-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-disk-full/</guid><description>&lt;h1 id="postgresql-disk-full-emergency-recovery-and-root-cause-analysis">PostgreSQL disk full: emergency recovery and root cause analysis&lt;/h1>
&lt;p>When &lt;code>df -h&lt;/code> shows 100% utilization on the PostgreSQL data volume, queries fail with &amp;ldquo;could not write to file&amp;rdquo;. If &lt;code>pg_wal&lt;/code> fills, the server enters PANIC and refuses to restart until space is freed. The fastest way to make the incident worse is to delete WAL files from &lt;code>pg_wal&lt;/code> manually. PostgreSQL needs those files for crash recovery; removing them causes data inconsistency that forces a restore from backup. Identify which subsystem is consuming space, reclaim it safely, and fix the root cause before the cycle repeats.&lt;/p></description></item><item><title>PostgreSQL ERROR: could not obtain lock — diagnosis and recovery</title><link>https://www.netdata.cloud/guides/postgres/postgres-lock-not-available/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-lock-not-available/</guid><description>&lt;h1 id="postgresql-error-could-not-obtain-lock--diagnosis-and-recovery">PostgreSQL ERROR: could not obtain lock — diagnosis and recovery&lt;/h1>
&lt;p>&lt;code>ERROR: could not obtain lock on row in relation&lt;/code> and &lt;code>ERROR: canceling statement due to lock timeout&lt;/code> mean a query requested a lock but PostgreSQL refused to wait. The database is not down; a session is holding a resource another transaction needs.&lt;/p>
&lt;p>Three variants produce these errors:&lt;/p>
&lt;ul>
&lt;li>&lt;code>SELECT ... FOR UPDATE NOWAIT&lt;/code> fails immediately if the row is locked.&lt;/li>
&lt;li>A statement that exceeds &lt;code>lock_timeout&lt;/code> fails after waiting.&lt;/li>
&lt;li>DDL such as &lt;code>ALTER TABLE&lt;/code> or &lt;code>CREATE INDEX&lt;/code> requires &lt;code>AccessExclusiveLock&lt;/code> and will wait or fail depending on session configuration.&lt;/li>
&lt;/ul>
&lt;p>These often cascade: one long-running query blocks a schema change, the schema change queues behind it, and subsequent queries queue behind the DDL until the connection pool exhausts.&lt;/p></description></item><item><title>PostgreSQL failover with Patroni: detection, promotion, and rollback</title><link>https://www.netdata.cloud/guides/postgres/postgres-failover-with-patroni/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-failover-with-patroni/</guid><description>&lt;h1 id="postgresql-failover-with-patroni-detection-promotion-and-rollback">PostgreSQL failover with Patroni: detection, promotion, and rollback&lt;/h1>
&lt;p>Patroni automates PostgreSQL failover with machine-enforced lease expiration. An agent on each node maintains a leader lock in a distributed consensus store such as etcd or Consul. When the primary stops renewing that lock, a surviving replica acquires it and promotes itself. Detection to promotion typically completes in seconds rather than the minutes a manual procedure requires.&lt;/p>
&lt;p>Automation does not remove operational risk. A network blip between Patroni and its consensus store can look exactly like a primary death. A promoted replica with unbounded lag can wipe out your RPO. An old primary that restarts outside Patroni&amp;rsquo;s control can create a split brain. Understanding detection, promotion, and rollback mechanics is essential in production.&lt;/p></description></item><item><title>PostgreSQL FATAL: Too Many Connections - Causes &amp; Fixes</title><link>https://www.netdata.cloud/guides/postgres/postgres-too-many-connections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-too-many-connections/</guid><description>&lt;p>Applications throw &lt;code>FATAL: sorry, too many clients already&lt;/code>. Health checks fail. Retries amplify the problem. The database is not down, but it rejects new traffic.&lt;/p>
&lt;p>&lt;a href="https://www.netdata.cloud/guides/postgres/">PostgreSQL&lt;/a> uses a process-per-connection model. Every backend holds memory and scheduler time even when idle. Raising &lt;code>max_connections&lt;/code> usually deepens the problem because the root cause is why slots are occupied and what those backends are doing.&lt;/p>
&lt;p>This guide shows how to diagnose connection exhaustion, distinguish PostgreSQL saturation from pooler exhaustion, and fix common root causes safely.&lt;/p></description></item><item><title>PostgreSQL frozen XID monitoring: catching wraparound 6 months early</title><link>https://www.netdata.cloud/guides/postgres/postgres-frozen-xid-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-frozen-xid-monitoring/</guid><description>&lt;h1 id="postgresql-frozen-xid-monitoring-catching-wraparound-6-months-early">PostgreSQL frozen XID monitoring: catching wraparound 6 months early&lt;/h1>
&lt;p>Transaction ID wraparound is a slow burn, not a sudden crisis. The 32-bit XID counter advances with every transaction. If VACUUM freeze does not keep pace, PostgreSQL will eventually refuse new transactions to prevent data corruption. By the time built-in log warnings fire, you may have only hours of runway left. Monitor &lt;code>age(datfrozenxid)&lt;/code> and &lt;code>age(relfrozenxid)&lt;/code> with tiered thresholds so you act with months of lead time.&lt;/p></description></item><item><title>PostgreSQL fsync=off: why disabling it ruins durability</title><link>https://www.netdata.cloud/guides/postgres/postgres-fsync-disabled/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-fsync-disabled/</guid><description>&lt;h1 id="postgresql-fsyncoff-why-disabling-it-ruins-durability">PostgreSQL fsync=off: why disabling it ruins durability&lt;/h1>
&lt;p>You are reviewing a PostgreSQL instance that crashed during a kernel update and now fails to start with checksum or WAL errors. Or you are benchmarking write throughput and a guide suggests turning off fsync to remove the &amp;ldquo;fsync bottleneck.&amp;rdquo; Maybe you are auditing configuration drift and found &lt;code>fsync = off&lt;/code> in a &lt;code>postgresql.conf&lt;/code> that was copied from an old developer environment.&lt;/p>
&lt;p>PostgreSQL&amp;rsquo;s default is &lt;code>fsync = on&lt;/code> in all supported versions. Disabling it can make bulk loads and heavy write workloads appear dramatically faster because the database stops waiting for the operating system to flush write-ahead log (WAL) records to durable storage. The cost is that an operating system crash or power loss can leave the data directory in an unrecoverable state. WAL is the source of truth; &lt;code>fsync = off&lt;/code> means the database trusts the OS page cache with that truth. When that trust is broken by a power event, the resulting corruption may not be detected until you attempt to start the server or query a specific table.&lt;/p></description></item><item><title>PostgreSQL idle in transaction: detecting and killing zombie sessions</title><link>https://www.netdata.cloud/guides/postgres/postgres-idle-in-transaction/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-idle-in-transaction/</guid><description>&lt;h1 id="postgresql-idle-in-transaction-detecting-and-killing-zombie-sessions">PostgreSQL idle in transaction: detecting and killing zombie sessions&lt;/h1>
&lt;p>When &lt;code>pg_stat_activity&lt;/code> fills with &lt;code>idle in transaction&lt;/code> sessions, autovacuum stalls, dead tuples accumulate, and applications throw &amp;ldquo;too many clients&amp;rdquo; errors while backends sit idle. These zombie sessions hold a transaction snapshot. Under MVCC, this prevents VACUUM from reclaiming dead tuples created after the transaction started, causing table bloat, blocked DDL, and transaction-ID wraparound pressure.&lt;/p>
&lt;p>This guide shows how to detect these sessions, determine whether they are blocking work, terminate them safely, and prevent recurrence.&lt;/p></description></item><item><title>PostgreSQL index bloat: detection and REINDEX CONCURRENTLY recovery</title><link>https://www.netdata.cloud/guides/postgres/postgres-index-bloat/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-index-bloat/</guid><description>&lt;h1 id="postgresql-index-bloat-detection-and-reindex-concurrently-recovery">PostgreSQL index bloat: detection and REINDEX CONCURRENTLY recovery&lt;/h1>
&lt;p>Queries that used to return in milliseconds are now spiking to hundreds. Disk usage grows faster than insert volume. &lt;code>EXPLAIN&lt;/code> shows a bitmap index scan pulling thousands of heap pages for a selective filter. The cause is usually index bloat: dead pages and fragmentation in B-tree indexes. Unlike table bloat, index bloat is not visible in &lt;code>pg_stat_user_tables&lt;/code>, and it often degrades query latency or fills a volume before you notice it. Detect bloat with &lt;code>pgstattuple&lt;/code> and recover online with &lt;code>REINDEX CONCURRENTLY&lt;/code>.&lt;/p></description></item><item><title>PostgreSQL logical replication failures: conflicts, schema drift, and recovery</title><link>https://www.netdata.cloud/guides/postgres/postgres-logical-replication-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-logical-replication-failures/</guid><description>&lt;h1 id="postgresql-logical-replication-failures-conflicts-schema-drift-and-recovery">PostgreSQL logical replication failures: conflicts, schema drift, and recovery&lt;/h1>
&lt;p>Your subscriber is behind. &lt;code>pg_stat_subscription&lt;/code> shows a stalled LSN, the apply worker is throwing errors, or worse: replication appears healthy while the subscriber silently diverges from the publisher. Logical replication does not replicate DDL, does not resolve row conflicts automatically, and will halt on the first integrity error it encounters. When it breaks, the failure is often on the subscriber, but the root cause may be schema drift on either side, a forgotten replication slot on the publisher, or a local write that violated the assumption that the subscriber is read-only.&lt;/p></description></item><item><title>PostgreSQL major version upgrade: pg_upgrade, logical replication, and rollback plans</title><link>https://www.netdata.cloud/guides/postgres/postgres-major-version-upgrade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-major-version-upgrade/</guid><description>&lt;h1 id="postgresql-major-version-upgrade-pg_upgrade-logical-replication-and-rollback-plans">PostgreSQL major version upgrade: pg_upgrade, logical replication, and rollback plans&lt;/h1>
&lt;p>A major version upgrade is one of the highest-risk maintenance operations on a PostgreSQL cluster. Storage format, system catalogs, and the query planner all change between major versions, so you cannot simply restart the server with new binaries. Every major upgrade is a migration, even when it happens on the same host.&lt;/p>
&lt;p>Most teams choose between two paths. &lt;code>pg_upgrade&lt;/code> rewrites system catalogs while reusing or copying data files. It is fast but requires a downtime window. Logical replication streams row changes to a new cluster running the target version, enabling near-zero-downtime cutover but adding operational complexity. The wrong choice is usually the one made without understanding rollback boundaries.&lt;/p></description></item><item><title>PostgreSQL missing indexes: detection from pg_stat_statements and logs</title><link>https://www.netdata.cloud/guides/postgres/postgres-missing-indexes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-missing-indexes/</guid><description>&lt;h1 id="postgresql-missing-indexes-detection-from-pg_stat_statements-and-logs">PostgreSQL missing indexes: detection from pg_stat_statements and logs&lt;/h1>
&lt;p>When the planner cannot find a suitable index path, it falls back to a sequential scan. On large tables, this turns millisecond queries into multi second outages.&lt;/p>
&lt;p>PostgreSQL exposes evidence through built-in instrumentation: &lt;code>pg_stat_user_tables&lt;/code> tracks sequential versus index scan ratios, &lt;code>pg_stat_statements&lt;/code> surfaces the most time-consuming query fingerprints, and &lt;code>auto_explain&lt;/code> logs execution plans that reveal &lt;code>Seq Scan&lt;/code> nodes directly.&lt;/p>
&lt;p>This guide gives an operator workflow to detect missing indexes using read-only checks and hypothetical index simulation. It assumes &lt;code>pg_stat_statements&lt;/code> is enabled; if not, adding it to &lt;code>shared_preload_libraries&lt;/code> requires a server restart.&lt;/p></description></item><item><title>PostgreSQL Monitoring</title><link>https://www.netdata.cloud/monitoring-101/postgres-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/postgres-monitoring/</guid><description>&lt;h2 id="postgresql-monitoring">PostgreSQL Monitoring&lt;/h2>
&lt;h3 id="what-is-postgresql">What Is PostgreSQL?&lt;/h3>
&lt;p>PostgreSQL is a powerful, open-source object-relational database management system known for its robustness, extensibility, and standards compliance. &lt;a href="https://www.postgresql.org/">Learn more about PostgreSQL&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-postgresql-with-netdata">Monitoring PostgreSQL With Netdata&lt;/h3>
&lt;p>Netdata provides a comprehensive PostgreSQL monitoring tool that allows you to keep an eye on crucial metrics in real-time and diagnose potential performance issues effectively. Using the &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/postgres/">Netdata Agent&amp;rsquo;s PostgreSQL module&lt;/a>, you can monitor various aspects of your database to ensure optimal performance.&lt;/p></description></item><item><title>PostgreSQL Monitoring</title><link>https://www.netdata.cloud/monitoring-101/postgresql-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/postgresql-monitoring/</guid><description>&lt;h2 id="whats-postgresql-and-why-monitor-it">What&amp;rsquo;s PostgreSQL and why monitor it?&lt;/h2>
&lt;p>PostgreSQL is a popular open source object-relational database system designed to work for a wide range of workloads from single machines to data warehouses to web services with many concurrent users. PostgreSQL runs on all major operating systems and is used by teams and organizations across the world, including Netdata.&lt;/p>
&lt;p>If you are using PostgreSQL in production, it is crucial that you monitor it for potential issues. And the more comprehensive the monitoring the better!&lt;/p></description></item><item><title>PostgreSQL Monitoring Checklist: The Signals Every Production Database Needs</title><link>https://www.netdata.cloud/guides/postgres/postgres-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-monitoring-checklist/</guid><description>&lt;p>&lt;a href="https://www.netdata.cloud/guides/postgres/">PostgreSQL&lt;/a> exposes hundreds of counters across the &lt;code>pg_stat_*&lt;/code> views, yet most production outages trace back to a small set of undetected conditions. Bloat accumulates silently until vacuum cannot catch up. Replication lag grows until failover becomes a data-loss event. Transaction ID age crosses a threshold and the database stops accepting writes. The problem is rarely a lack of metrics. It is knowing which signals to instrument at each stage of operational maturity.&lt;/p></description></item><item><title>PostgreSQL monitoring maturity model: from reactive to self-healing</title><link>https://www.netdata.cloud/guides/postgres/postgres-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-monitoring-maturity-model/</guid><description>&lt;h1 id="postgresql-monitoring-maturity-model-from-reactive-to-self-healing">PostgreSQL monitoring maturity model: from reactive to self-healing&lt;/h1>
&lt;p>Production PostgreSQL does not usually fail catastrophically; it drifts. An unwatched dashboard, an untested backup, a regressing query plan, or a filling replication slot slowly creates an incident. This model gives you eight observable stages to benchmark your operations, with measurable indicators and common stuck points from production runbooks. Use it to find your current stage, the next transition enabler, and the organizational traps that cause regression.&lt;/p></description></item><item><title>PostgreSQL out of memory: OOM killer, shared_buffers, and work_mem</title><link>https://www.netdata.cloud/guides/postgres/postgres-out-of-memory/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-out-of-memory/</guid><description>&lt;h1 id="postgresql-out-of-memory-oom-killer-shared_buffers-and-work_mem">PostgreSQL out of memory: OOM killer, shared_buffers, and work_mem&lt;/h1>
&lt;p>Your PostgreSQL primary restarts without warning, or individual backends vanish from the process list. The kernel log shows &lt;code>Out of Memory: Killed process 12345 (postgres)&lt;/code>. Existing connections may survive, but new connections fail until the postmaster recovers. This is a memory accounting mismatch between Linux overcommit, PostgreSQL shared and private memory allocation, and how you size &lt;code>shared_buffers&lt;/code> and &lt;code>work_mem&lt;/code>.&lt;/p>
&lt;p>Inside Kubernetes or containers, the symptom is identical but the mechanism differs: cgroup v2 &lt;code>memory.max&lt;/code> triggers an immediate SIGKILL with no ENOMEM grace period. In both cases, stop guessing at memory limits and start budgeting.&lt;/p></description></item><item><title>PostgreSQL pg_upgrade failures: extension, collation, and locale gotchas</title><link>https://www.netdata.cloud/guides/postgres/postgres-pg-upgrade-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-pg-upgrade-failures/</guid><description>&lt;h1 id="postgresql-pg_upgrade-failures-extension-collation-and-locale-gotchas">PostgreSQL pg_upgrade failures: extension, collation, and locale gotchas&lt;/h1>
&lt;p>You ran &lt;code>pg_upgrade --check&lt;/code> and it reported a collation version mismatch. Or the upgrade finished, your application reconnected, and PostGIS functions failed with &lt;code>could not access file '$libdir/postgis-3'&lt;/code>. Maybe &lt;code>SELECT&lt;/code> queries against text columns returned rows in the wrong order, or unique indexes threw violations against previously clean data. These are structural gaps between what pg_upgrade copies (catalog metadata) and what it does not verify (shared libraries, OS collation rules, and runtime configuration).&lt;/p></description></item><item><title>PostgreSQL pg_wal directory full: causes and emergency recovery</title><link>https://www.netdata.cloud/guides/postgres/postgres-wal-disk-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-wal-disk-full/</guid><description>&lt;h1 id="postgresql-pg_wal-directory-full-causes-and-emergency-recovery">PostgreSQL pg_wal directory full: causes and emergency recovery&lt;/h1>
&lt;p>Your paging system fires because the PostgreSQL primary has stopped accepting writes. The error log reports a disk-full condition, and the WAL volume is at 100%. You cannot simply delete files from pg_wal to free space: doing so corrupts the database and breaks replication.&lt;/p>
&lt;p>PostgreSQL recycles WAL segments only after a checkpoint, and only when they are no longer needed for crash recovery, archiving, or replication slots. Archiving failures, stalled replication slots, or bulk loads that exceed max_wal_size cause unbounded accumulation. This guide covers identification, safe recovery, and prevention.&lt;/p></description></item><item><title>PostgreSQL prepared statement plan cache: generic vs custom plan pitfalls</title><link>https://www.netdata.cloud/guides/postgres/postgres-prepared-statement-plan-cache/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-prepared-statement-plan-cache/</guid><description>&lt;h1 id="postgresql-prepared-statement-plan-cache-generic-vs-custom-plan-pitfalls">PostgreSQL prepared statement plan cache: generic vs custom plan pitfalls&lt;/h1>
&lt;p>When an application repeats the same query with different parameters, prepared statements avoid repeated parsing. But PostgreSQL&amp;rsquo;s plan cache does not work the way many operators expect. The server does not automatically cache the first plan it generates. It executes the first five invocations with custom plans that see the actual bound values. At the sixth execution, the planner compares the average cost of those custom plans against a generic plan built with placeholders. If the generic plan looks cheaper, the session switches to it permanently. There is no automatic reversion.&lt;/p></description></item><item><title>PostgreSQL replica disconnected: detecting and recovering streaming replication</title><link>https://www.netdata.cloud/guides/postgres/postgres-replica-disconnected/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-replica-disconnected/</guid><description>&lt;h1 id="postgresql-replica-disconnected-detecting-and-recovering-streaming-replication">PostgreSQL replica disconnected: detecting and recovering streaming replication&lt;/h1>
&lt;p>When &lt;code>replay_lag&lt;/code> grows and the primary&amp;rsquo;s &lt;code>pg_stat_replication&lt;/code> no longer lists the standby, queries return stale data and failover is unsafe. PostgreSQL&amp;rsquo;s WAL receiver automatically retries when streaming breaks, cycling through WAL archive, local &lt;code>pg_wal&lt;/code>, and streaming connections. That loop can mask the root cause while the replica drifts toward an unrecoverable gap. Determine whether the replica is temporarily stalled, beyond recovery, or missing its replication slot on the primary.&lt;/p></description></item><item><title>PostgreSQL replica out of sync: timeline mismatches and recovery</title><link>https://www.netdata.cloud/guides/postgres/postgres-replica-out-of-sync/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-replica-out-of-sync/</guid><description>&lt;h1 id="postgresql-replica-out-of-sync-timeline-mismatches-and-recovery">PostgreSQL replica out of sync: timeline mismatches and recovery&lt;/h1>
&lt;p>A PostgreSQL streaming replica that was healthy yesterday now refuses to start with a timeline mismatch error, or a former primary that you brought back online cannot rejoin the cluster as a standby. The log shows &lt;code>requested timeline N is not a child of this server's history&lt;/code> and the replica loops in crash recovery while the current primary continues to diverge. This guide covers identifying the divergence point, choosing between &lt;code>pg_rewind&lt;/code> and a full re-clone, and recovering the replica without introducing split-brain or data loss.&lt;/p></description></item><item><title>PostgreSQL replication lag: detection, diagnosis, and fixes</title><link>https://www.netdata.cloud/guides/postgres/postgres-replication-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-replication-lag/</guid><description>&lt;h1 id="postgresql-replication-lag-detection-diagnosis-and-fixes">PostgreSQL replication lag: detection, diagnosis, and fixes&lt;/h1>
&lt;p>Replication lag is the distance between the last WAL record generated on the primary and the last record applied on a replica. In asynchronous streaming replication, a few seconds of lag is normal. When lag grows without bound, your recovery point objective becomes fiction.&lt;/p>
&lt;p>Lag often grows silently. Replication processes stay connected, WAL streams flow, and uptime checks stay green while the byte gap creeps from megabytes to gigabytes. Promoting a replica that is hours behind destroys the consistency your application assumes.&lt;/p></description></item><item><title>PostgreSQL replication slot bloat: when a stale slot fills the disk</title><link>https://www.netdata.cloud/guides/postgres/postgres-replication-slot-bloat/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-replication-slot-bloat/</guid><description>&lt;h1 id="postgresql-replication-slot-bloat-when-a-stale-slot-fills-the-disk">PostgreSQL replication slot bloat: when a stale slot fills the disk&lt;/h1>
&lt;p>Disk is filling on the primary. Table sizes are stable and active replicas show healthy replication lag, but WAL in pg_wal keeps growing and is not being recycled. The most likely cause is a stale replication slot. When a logical subscriber, CDC connector, or physical replica disconnects without cleaning up its slot, PostgreSQL retains every WAL segment from the slot&amp;rsquo;s restart_lsn onward. This retention ignores max_wal_size unless max_slot_wal_keep_size is set to a finite value. The default is -1 (unlimited). The primary will retain WAL until the disk fills and writes halt. This guide covers confirmation, recovery, and prevention.&lt;/p></description></item><item><title>PostgreSQL row-level lock contention: SELECT FOR UPDATE patterns and fixes</title><link>https://www.netdata.cloud/guides/postgres/postgres-row-level-lock-contention/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-row-level-lock-contention/</guid><description>&lt;h1 id="postgresql-row-level-lock-contention-select-for-update-patterns-and-fixes">PostgreSQL row-level lock contention: SELECT FOR UPDATE patterns and fixes&lt;/h1>
&lt;p>P99 latency doubles. &lt;code>pg_stat_activity&lt;/code> shows &lt;code>wait_event_type = 'Lock'&lt;/code>. Queries are simple, indexes are present, and CPU is idle. The culprit is usually a row-level lock queue: one transaction holds a tuple lock longer than expected, and every subsequent transaction touching the same row waits in FIFO order.&lt;/p>
&lt;p>This is not a deadlock. PostgreSQL detects deadlocks automatically and kills one participant. Lock contention is different: a slow or abandoned transaction acts as a head-of-line blocker. Unless you are polling &lt;code>pg_locks&lt;/code> and transaction age, the queue is invisible in query logs. The pattern is most common with &lt;code>SELECT ... FOR UPDATE&lt;/code> in database-backed job queues, &lt;code>UPDATE&lt;/code> statements that silently take stronger locks than intended, and session-level advisory locks that survive rollback.&lt;/p></description></item><item><title>PostgreSQL sequential scan on a large table: when to add an index and when not to</title><link>https://www.netdata.cloud/guides/postgres/postgres-sequential-scan-on-large-table/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-sequential-scan-on-large-table/</guid><description>&lt;h1 id="postgresql-sequential-scan-on-a-large-table-when-to-add-an-index-and-when-not-to">PostgreSQL sequential scan on a large table: when to add an index and when not to&lt;/h1>
&lt;p>Seeing &lt;code>Seq Scan&lt;/code> on a large table in &lt;code>EXPLAIN&lt;/code> does not mean the planner is wrong. Adding a B-tree index often makes the query slower, burdens every write path with index maintenance, and wastes storage. PostgreSQL&amp;rsquo;s planner compares the estimated cost of reading the table sequentially against the estimated cost of traversing an index and fetching rows from the heap. There is no hardcoded row-percentage threshold. On SSD-backed servers, the default cost parameters are frequently stale, so the planner may be wrong in either direction.&lt;/p></description></item><item><title>PostgreSQL shared_buffers tuning: 25% of RAM and why that rule of thumb breaks</title><link>https://www.netdata.cloud/guides/postgres/postgres-shared-buffers-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-shared-buffers-tuning/</guid><description>&lt;h1 id="postgresql-shared_buffers-tuning-25-of-ram-and-why-that-rule-of-thumb-breaks">PostgreSQL shared_buffers tuning: 25% of RAM and why that rule of thumb breaks&lt;/h1>
&lt;p>The PostgreSQL documentation suggests 25% of RAM as a starting value for shared_buffers on dedicated servers and warns that values above 40% rarely improve performance. In production, operators routinely use 25% as a default and then encounter checkpoint storms, OOM kills, or degraded analytical query performance. The rule breaks because PostgreSQL does not coordinate with the Linux kernel page cache. The same page can exist in both shared_buffers and the OS page cache, so an oversized PostgreSQL cache starves the kernel of memory needed for sequential scans, WAL, and temporary files.&lt;/p></description></item><item><title>PostgreSQL Slow Queries: Diagnosis From Log To Plan To Fix</title><link>https://www.netdata.cloud/guides/postgres/postgres-slow-queries-diagnosis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-slow-queries-diagnosis/</guid><description>&lt;p>A query that returned in 10 ms yesterday is now taking 8 seconds. No deploys, no schema changes. Before adding an index or restarting the database, determine whether the slowness is in the plan, the data, or the environment. This guide covers a three-layer workflow: log-based discovery with &lt;code>log_min_duration_statement&lt;/code>, aggregate profiling with &lt;code>pg_stat_statements&lt;/code>, and per-query execution plan capture with &lt;code>auto_explain&lt;/code> and manual &lt;code>EXPLAIN (ANALYZE, BUFFERS)&lt;/code>.&lt;/p>

&lt;pre class="mermaid" data-source="markdown-fence">flowchart TD
 A[Slow query reported] --> B[Check pg_stat_statements for mean_time and stddev_time]
 B --> C{High stddev relative to mean?}
 C -->|Yes| D[Plan flapping or parameter skew]
 C -->|No| E[Stable plan or system bottleneck]
 D --> F[Capture plan with auto_explain or EXPLAIN]
 E --> F
 F --> G{Estimated rows far from actual?}
 G -->|Yes| H[Stale statistics or skewed data]
 G -->|No| I[Bloat, locks, or cache pressure]&lt;/pre>
&lt;h2 id="what-this-means">What This Means&lt;/h2>
&lt;p>A slow query is a symptom. &lt;a href="https://www.netdata.cloud/guides/postgres/">PostgreSQL&amp;rsquo;s planner&lt;/a> chooses a path based on statistics from &lt;code>ANALYZE&lt;/code>, bound parameter values, and configuration such as &lt;code>work_mem&lt;/code>. When these inputs change, the same query text can switch from a hash join to a nested loop and slow down by orders of magnitude.&lt;/p></description></item><item><title>PostgreSQL split-brain after failover: detection and reconciliation</title><link>https://www.netdata.cloud/guides/postgres/postgres-split-brain-after-failover/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-split-brain-after-failover/</guid><description>&lt;h1 id="postgresql-split-brain-after-failover-detection-and-reconciliation">PostgreSQL split-brain after failover: detection and reconciliation&lt;/h1>
&lt;p>A failover should leave one primary and a clean topology. Split-brain means two instances accept writes, application connections are split across both nodes, and transaction histories diverge. It typically starts with a network partition, a missed demotion signal, or a health-check false positive that promotes a replica while the old primary keeps running. Once both primaries accept transactions, timelines diverge and you must reconcile.&lt;/p></description></item><item><title>PostgreSQL statistics out of date: ANALYZE, default_statistics_target, and bad plans</title><link>https://www.netdata.cloud/guides/postgres/postgres-statistics-out-of-date/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-statistics-out-of-date/</guid><description>&lt;h1 id="postgresql-statistics-out-of-date-analyze-default_statistics_target-and-bad-plans">PostgreSQL statistics out of date: ANALYZE, default_statistics_target, and bad plans&lt;/h1>
&lt;p>When &lt;code>EXPLAIN&lt;/code> shows a nested loop joining a million-row table, or a sequential scan where an index should win, and the query text has not changed, stale planner statistics are the likely culprit.&lt;/p>
&lt;p>PostgreSQL&amp;rsquo;s cost-based optimizer relies on &lt;code>ANALYZE&lt;/code>-derived catalog data to estimate cardinality. &lt;code>ANALYZE&lt;/code> samples table contents and writes histograms, most-common-value lists, and correlation data into &lt;code>pg_statistic&lt;/code>. When that data no longer matches the on-disk distribution, the planner misestimates row counts. A cardinality underestimate favors a nested loop; an overestimate favors a sequential scan. The result is a sudden latency spike that connection pooling or hardware scaling will not fix.&lt;/p></description></item><item><title>PostgreSQL streaming replication broken: how to rebuild without full basebackup</title><link>https://www.netdata.cloud/guides/postgres/postgres-streaming-replication-broken/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-streaming-replication-broken/</guid><description>&lt;h1 id="postgresql-streaming-replication-broken-how-to-rebuild-without-full-basebackup">PostgreSQL streaming replication broken: how to rebuild without full basebackup&lt;/h1>
&lt;p>When a replica logs &lt;code>requested WAL segment has already been removed&lt;/code>, or after failover when the old primary must rejoin as a standby, a full &lt;code>pg_basebackup&lt;/code> on a multi-terabyte database can take hours and saturate network and disk. If data files have diverged but are mostly identical, &lt;code>pg_rewind&lt;/code> can resync by copying only changed blocks. It is faster than a base backup, but has strict prerequisites. If &lt;code>pg_rewind&lt;/code> crashes mid-operation, the target data directory is likely corrupt and only a fresh base backup is safe.&lt;/p></description></item><item><title>PostgreSQL synchronous_commit: durability vs throughput trade-offs</title><link>https://www.netdata.cloud/guides/postgres/postgres-synchronous-commit-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-synchronous-commit-tuning/</guid><description>&lt;h1 id="postgresql-synchronous_commit-durability-vs-throughput-trade-offs">PostgreSQL synchronous_commit: durability vs throughput trade-offs&lt;/h1>
&lt;p>&lt;code>synchronous_commit&lt;/code> moves the durability boundary between a client receiving COMMIT OK and the data actually surviving a crash. The default, &lt;code>on&lt;/code>, flushes local WAL to disk. With synchronous replication enabled, it also waits for a standby to flush. The five values (&lt;code>off&lt;/code>, &lt;code>local&lt;/code>, &lt;code>on&lt;/code>, &lt;code>remote_write&lt;/code>, and &lt;code>remote_apply&lt;/code>) trade latency against survival guarantees. The wrong choice either accepts unplanned data loss or strangles write throughput with network round-trips.&lt;/p></description></item><item><title>PostgreSQL table bloat: detection, measurement, and remediation</title><link>https://www.netdata.cloud/guides/postgres/postgres-table-bloat/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-table-bloat/</guid><description>&lt;h1 id="postgresql-table-bloat-detection-measurement-and-remediation">PostgreSQL table bloat: detection, measurement, and remediation&lt;/h1>
&lt;p>Table bloat appears as tables growing while row counts stay flat, sequential scans slowing despite indexes, and disk alarms that do not track business growth. PostgreSQL&amp;rsquo;s MVCC writes new tuple versions instead of overwriting old ones; every UPDATE leaves a dead row and every DELETE leaves invisible garbage. VACUUM reclaims that space for reuse within the file, but plain VACUUM does not shrink the relation on disk. When dead tuples outpace cleanup, a table can become several times larger than its logical content.&lt;/p></description></item><item><title>PostgreSQL tablespace management: moving data without downtime</title><link>https://www.netdata.cloud/guides/postgres/postgres-tablespace-management/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-tablespace-management/</guid><description>&lt;h1 id="postgresql-tablespace-management-moving-data-without-downtime">PostgreSQL tablespace management: moving data without downtime&lt;/h1>
&lt;p>PostgreSQL tablespaces map relations to specific filesystem paths. Moving data between tablespaces is not an online operation in core PostgreSQL: &lt;code>ALTER TABLE SET TABLESPACE&lt;/code> copies the relation&amp;rsquo;s data files while holding an &lt;code>ACCESS EXCLUSIVE&lt;/code> lock for the entire duration. For a large relation, reads and writes block for minutes or hours. This guide covers three operational paths: the built-in offline move for small or maintenance-tolerant relations, offline symlink relocation for entire tablespace directories, and near-zero-downtime workarounds for heavy tables.&lt;/p></description></item><item><title>PostgreSQL temp file explosion: work_mem tuning and disk pressure</title><link>https://www.netdata.cloud/guides/postgres/postgres-temp-file-explosion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-temp-file-explosion/</guid><description>&lt;h1 id="postgresql-temp-file-explosion-work_mem-tuning-and-disk-pressure">PostgreSQL temp file explosion: work_mem tuning and disk pressure&lt;/h1>
&lt;p>Disk usage on your PostgreSQL primary is climbing by gigabytes per hour. Query latency has doubled, &lt;code>iowait&lt;/code> is spiking, and you have not deployed new code. The database is writing large temporary files to disk. In-memory operations such as sorts, hash joins, and materialization are spilling because they exceed the per-operation memory budget. The root cause is almost always a mismatch between &lt;code>work_mem&lt;/code> and actual query memory demand, compounded by the fact that PostgreSQL applies this limit per plan node, not per query or session.&lt;/p></description></item><item><title>PostgreSQL TOAST table bloat: detection and recovery for large columns</title><link>https://www.netdata.cloud/guides/postgres/postgres-toast-table-bloat/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-toast-table-bloat/</guid><description>&lt;h1 id="postgresql-toast-table-bloat-detection-and-recovery-for-large-columns">PostgreSQL TOAST table bloat: detection and recovery for large columns&lt;/h1>
&lt;p>Unexplained disk growth on a PostgreSQL instance often traces to TOAST bloat. A table with large JSONB, text, or bytea columns grows far beyond live data size, and queries that materialize those columns slow down. A plain VACUUM does not shrink the on-disk footprint because the dead tuples are in the companion pg_toast table, not the main heap.&lt;/p>
&lt;p>TOAST (The Oversized-Attribute Storage Technique) moves values exceeding the TOAST_TUPLE_TARGET out of the main row and into a companion pg_toast.NNNN table. PostgreSQL compresses the value and splits it into roughly 2 KB chunks, each stored as a separate row with a unique index on (chunk_id, chunk_seq). Updates or deletes to the row leave dead chunks in the TOAST table. Because operators often monitor only the parent table, TOAST bloat grows silently until it dominates disk usage.&lt;/p></description></item><item><title>PostgreSQL transaction ID wraparound: detection and emergency recovery</title><link>https://www.netdata.cloud/guides/postgres/postgres-transaction-id-wraparound/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-transaction-id-wraparound/</guid><description>&lt;h1 id="postgresql-transaction-id-wraparound-detection-and-emergency-recovery">PostgreSQL transaction ID wraparound: detection and emergency recovery&lt;/h1>
&lt;p>PostgreSQL can suddenly stop accepting writes and emit warnings that the database must be vacuumed within a shrinking number of transactions. This is transaction ID wraparound. It is not gradual performance degradation; it is a hard stop that can take a database offline for hours if old tuples are not frozen in time.&lt;/p>
&lt;p>Every write transaction consumes a 32-bit XID. After roughly two billion transactions, the counter nears the point where older tuples could appear to belong to the future, corrupting visibility. PostgreSQL refuses to hand out new XIDs once fewer than roughly three million remain. Warnings appear at roughly forty million remaining. The defense is VACUUM freeze, which marks old rows with FrozenTransactionId so they no longer depend on the live counter.&lt;/p></description></item><item><title>PostgreSQL VACUUM FULL vs pg_repack: choosing the right bloat fix</title><link>https://www.netdata.cloud/guides/postgres/postgres-vacuum-full-vs-pg-repack/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-vacuum-full-vs-pg-repack/</guid><description>&lt;h1 id="postgresql-vacuum-full-vs-pg_repack-choosing-the-right-bloat-fix">PostgreSQL VACUUM FULL vs pg_repack: choosing the right bloat fix&lt;/h1>
&lt;p>Table bloat is the gap between logical data size and on-disk consumption. Heavy UPDATE and DELETE activity leaves dead tuples in the heap. Autovacuum reclaims those tuples for reuse but does not shrink the underlying file. Once sequential scans slow and disk space vanishes, operators must choose between VACUUM FULL and pg_repack. The wrong choice locks a table for hours, exhausts disk mid-operation, or leaves orphaned triggers behind. This guide explains how each tool rewrites the heap, what locks it takes, and how to pick the right one.&lt;/p></description></item><item><title>PostgreSQL WAL archive failures: archive_command exit codes and recovery</title><link>https://www.netdata.cloud/guides/postgres/postgres-wal-archive-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-wal-archive-failures/</guid><description>&lt;h1 id="postgresql-wal-archive-failures-archive_command-exit-codes-and-recovery">PostgreSQL WAL archive failures: archive_command exit codes and recovery&lt;/h1>
&lt;p>WAL archiving is the durability bridge between your PostgreSQL primary and your ability to recover to any point in time. When &lt;code>archive_command&lt;/code> fails, WAL segments stay pinned in &lt;code>pg_wal/&lt;/code>. The directory grows until the filesystem fills, at which point PostgreSQL performs an emergency PANIC shutdown. Even before that happens, every failed segment is a gap in your backup chain, rendering base backups useless for PITR beyond the first missing file.&lt;/p></description></item><item><title>PostgreSQL: checkpoints are occurring too frequently -- what to tune</title><link>https://www.netdata.cloud/guides/postgres/postgres-checkpoints-occurring-too-frequently/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-checkpoints-occurring-too-frequently/</guid><description>&lt;h1 id="postgresql-checkpoints-are-occurring-too-frequently----what-to-tune">PostgreSQL: checkpoints are occurring too frequently &amp;ndash; what to tune&lt;/h1>
&lt;p>Your PostgreSQL logs show &lt;code>LOG: checkpoints are occurring too frequently (9 seconds apart)&lt;/code> with the hint &lt;code>Consider increasing the configuration parameter 'max_wal_size'&lt;/code>. Sustained I/O latency spikes correlate with WAL segment rotation. On write-heavy primaries, this almost always means &lt;code>max_wal_size&lt;/code> is too small for the workload.&lt;/p>
&lt;p>A checkpoint flushes all dirty shared buffers to disk. By default, a checkpoint fires every 5 minutes (&lt;code>checkpoint_timeout&lt;/code>) or every 1 GB of WAL (&lt;code>max_wal_size&lt;/code>), whichever comes first. On a busy OLTP primary, 1 GB of WAL can accumulate in minutes. When &lt;code>max_wal_size&lt;/code> triggers the checkpoint, you get a forced (requested) checkpoint. These are unpredictable, collide with existing write load, and cause visible latency spikes.&lt;/p></description></item><item><title>PostgreSQL: database is not accepting commands to avoid wraparound data loss</title><link>https://www.netdata.cloud/guides/postgres/postgres-database-not-accepting-commands/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-database-not-accepting-commands/</guid><description>&lt;h1 id="postgresql-database-is-not-accepting-commands-to-avoid-wraparound-data-loss">PostgreSQL: database is not accepting commands to avoid wraparound data loss&lt;/h1>
&lt;p>Writes fail abruptly. SELECT still works, but INSERT, UPDATE, DELETE, and DDL return an error like:&lt;/p>
&lt;pre tabindex="0">&lt;code>ERROR: database is not accepting commands that assign new XIDs to avoid wraparound data loss in database &amp;#34;...&amp;#34;
&lt;/code>&lt;/pre>&lt;p>This is PostgreSQL&amp;rsquo;s emergency brake, not a crash. The engine has stopped issuing new transaction IDs to prevent tuple visibility corruption. Modern PostgreSQL lets you recover without restarting or entering single-user mode. Remove whatever is blocking VACUUM progress, then freeze old rows.&lt;/p></description></item><item><title>Power Capping</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/power-capping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/power-capping/</guid><description/></item><item><title>Power Distribution Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/power-distribution-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/power-distribution-inc-snmp-traps/</guid><description/></item><item><title>Power Supply</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/power-supply/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/power-supply/</guid><description/></item><item><title>Power Supply (Win)</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/power-supply-win/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/power-supply-win/</guid><description/></item><item><title>Power_On_Hours and fleet age: the context every other attribute needs</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-power-on-hours-fleet-age/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-power-on-hours-fleet-age/</guid><description>&lt;h1 id="power_on_hours-and-fleet-age-the-context-every-other-attribute-needs">Power_On_Hours and fleet age: the context every other attribute needs&lt;/h1>
&lt;p>Power_On_Hours (ATA attribute ID 9, or the NVMe SMART/Health &amp;ldquo;Power On Hours&amp;rdquo; field) does not tell you anything is wrong. It tells you how long the drive has been powered on. That is the only attribute that gives every other attribute its meaning.&lt;/p>
&lt;p>Ten reallocated sectors on a drive with 100 power-on hours is a manufacturing defect. Ten reallocated sectors on a drive with 50,000 power-on hours is graceful aging. The raw count is identical; the operational response is completely different. Without the age axis, you cannot distinguish &amp;ldquo;the drive is burning out&amp;rdquo; from &amp;ldquo;the drive is wearing out.&amp;rdquo;&lt;/p></description></item><item><title>PowerDNS Authoritative Server</title><link>https://www.netdata.cloud/integrations/data-collection/networking/powerdns-authoritative-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/powerdns-authoritative-server/</guid><description/></item><item><title>PowerDNS Monitoring</title><link>https://www.netdata.cloud/monitoring-101/powerdns-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/powerdns-monitoring/</guid><description>&lt;h2 id="powerdns-monitoring">PowerDNS Monitoring&lt;/h2>
&lt;h3 id="what-is-powerdns">What Is PowerDNS?&lt;/h3>
&lt;p>PowerDNS Authoritative Server is a versatile and highly efficient DNS server software that forms a pivotal part of many IT infrastructures. Renowned for its performance capabilities and extensive management features, PowerDNS provides authoritative DNS services that cater to a wide range of deployment scenarios.&lt;/p>
&lt;h3 id="monitoring-powerdns-with-netdata">Monitoring PowerDNS With Netdata&lt;/h3>
&lt;p>Netdata offers a comprehensive PowerDNS monitoring solution that allows you to track, visualize, and analyze real-time data from your PowerDNS instances. By leveraging the Netdata platform, which is acclaimed for its simplicity and effectiveness, you can gain insights into server performance, diagnose potential issues, and ensure optimal service delivery.&lt;/p></description></item><item><title>PowerDNS Recursor</title><link>https://www.netdata.cloud/integrations/data-collection/networking/powerdns-recursor/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/powerdns-recursor/</guid><description/></item><item><title>PowerDNS Recursor Monitoring</title><link>https://www.netdata.cloud/monitoring-101/powerdns-recursor-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/powerdns-recursor-monitoring/</guid><description>&lt;h2 id="what-is-powerdns-recursor">What is PowerDNS Recursor?&lt;/h2>
&lt;p>&lt;a href="https://doc.powerdns.com/recursor/">&lt;code>PowerDNS Recursor&lt;/code>&lt;/a> is a high-performance DNS recursor with built-in scripting capabilities.&lt;/p>
&lt;h2 id="monitoring-powerdns-recursor-with-netdata">Monitoring PowerDNS Recursor with Netdata&lt;/h2>
&lt;p>The prerequisites for monitoring PowerDNS Recursor with Netdata are to have PowerDNS Recursor and &lt;a href="https://learn.netdata.cloud/docs/cloud/get-started/">Netdata installed&lt;/a> on your system.&lt;/p>
&lt;p>Netdata auto discovers hundreds of services, and for those it doesn&amp;rsquo;t turning on manual discovery is a one line configuration. For more information on configuring Netdata for PowerDNS Recursor monitoring please read the collector &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/powerdns_recursor/">documentation&lt;/a>.&lt;/p></description></item><item><title>PowerDNS Recursor Monitoring</title><link>https://www.netdata.cloud/monitoring-101/powerdns_recursor-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/powerdns_recursor-monitoring/</guid><description>&lt;h2 id="powerdns-recursor-monitoring">PowerDNS Recursor Monitoring&lt;/h2>
&lt;h3 id="what-is-powerdns-recursor">What Is PowerDNS Recursor?&lt;/h3>
&lt;p>PowerDNS Recursor is a high-performance DNS server used by service providers to resolve web traffic efficiently. Designed with a focus on scalability and security, the PowerDNS Recursor acts as the backbone of DNS infrastructure in complex network setups, maintaining trust and performance at scale.&lt;/p>
&lt;h3 id="monitoring-powerdns-recursor-with-netdata">Monitoring PowerDNS Recursor With Netdata&lt;/h3>
&lt;p>When it comes to monitoring PowerDNS Recursor, Netdata offers a comprehensive &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/powerdns_recursor/">PowerDNS Recursor monitoring tool&lt;/a>. Using Netdata, you can monitor key metrics such as incoming and outgoing questions, answer times, timeouts, and cache usage with real-time precision.&lt;/p></description></item><item><title>Powerdsine SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/powerdsine-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/powerdsine-snmp-traps/</guid><description/></item><item><title>Powerpal devices</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/powerpal-devices/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/powerpal-devices/</guid><description/></item><item><title>Powerpal Devices Monitoring</title><link>https://www.netdata.cloud/monitoring-101/powerpal-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/powerpal-monitoring/</guid><description>&lt;h2 id="powerpal-devices-monitoring">Powerpal Devices Monitoring&lt;/h2>
&lt;h3 id="what-is-powerpal-devices">What Is Powerpal Devices?&lt;/h3>
&lt;p>Powerpal devices are innovative IoT-driven smart meters that allow users to efficiently track their energy consumption in real-time. Designed to promote energy efficiency, they provide detailed insights into electricity usage, enabling homeowners and businesses to make informed decisions about their energy consumption strategies.&lt;/p>
&lt;h3 id="monitoring-powerpal-devices-with-netdata">Monitoring Powerpal Devices With Netdata&lt;/h3>
&lt;p>When it comes to monitoring Powerpal devices, Netdata offers a robust solution that leverages the openmetrics (Prometheus) exporter approach. Using a community exporter, such as the &lt;a href="https://github.com/aashley/powerpal_exporter">Powerpal Exporter&lt;/a>, Netdata can effortlessly ingest and visualize metrics without the need for setting up a Prometheus server or a Grafana dashboard. With Netdata&amp;rsquo;s user-friendly platform, users can benefit from automated dashboards and real-time alerts, ensuring they stay informed about their energy consumption patterns.&lt;/p></description></item><item><title>Powershield Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/powershield-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/powershield-ltd-snmp-traps/</guid><description/></item><item><title>Powertek Limited SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/powertek-limited-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/powertek-limited-snmp-traps/</guid><description/></item><item><title>Pre-register to join the waiting list!</title><link>https://www.netdata.cloud/waiting-list/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/waiting-list/</guid><description/></item><item><title>Premier Network Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/premier-network-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/premier-network-co-ltd-snmp-traps/</guid><description/></item><item><title>Pressure Stall Information</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/pressure-stall-information/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/pressure-stall-information/</guid><description/></item><item><title>Printer Working Group SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/printer-working-group-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/printer-working-group-snmp-traps/</guid><description/></item><item><title>Privacy Policy</title><link>https://www.netdata.cloud/privacy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/privacy/</guid><description/></item><item><title>Processor</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/processor/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/processor/</guid><description/></item><item><title>Product Hunt Form</title><link>https://www.netdata.cloud/product-hunt-form/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/product-hunt-form/</guid><description/></item><item><title>ProFTPD</title><link>https://www.netdata.cloud/integrations/data-collection/applications/proftpd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/proftpd/</guid><description/></item><item><title>ProFTPD Monitoring</title><link>https://www.netdata.cloud/monitoring-101/proftpd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/proftpd-monitoring/</guid><description>&lt;h2 id="proftpd-monitoring">ProFTPD Monitoring&lt;/h2>
&lt;h3 id="what-is-proftpd">What Is ProFTPD?&lt;/h3>
&lt;p>ProFTPD is a popular FTP server used for hosting, transferring, and sharing files over networks. Known for its security features, configurability, and flexibility, ProFTPD is a go-to solution for many organizations that require robust file transfer capabilities. It supports numerous authentication methods and access control options, making it versatile for varied environments.&lt;/p>
&lt;h3 id="monitoring-proftpd-with-netdata">Monitoring ProFTPD With Netdata&lt;/h3>
&lt;p>When it comes to monitoring ProFTPD, Netdata offers a seamless way to collect and visualize vital metrics using the &lt;a href="https://github.com/transnano/proftpd_exporter">ProFTPD Exporter&lt;/a>. Netdata taps into openmetrics with a Prometheus-compatible exporter, meaning you can capture detailed insights into your ProFTPD server’s performance. Unlike traditional setups that require a Prometheus server or a Grafana dashboard, Netdata simplifies the process. With its integration capabilities, users gain access to real-time automated dashboards and alerts without additional complex setups.&lt;/p></description></item><item><title>Prometheus endpoint</title><link>https://www.netdata.cloud/integrations/data-collection/applications/prometheus-endpoint/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/prometheus-endpoint/</guid><description/></item><item><title>Prometheus Endpoint Monitoring</title><link>https://www.netdata.cloud/monitoring-101/generic-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/generic-monitoring/</guid><description>&lt;h2 id="prometheus-endpoint-monitoring">Prometheus Endpoint Monitoring&lt;/h2>
&lt;h3 id="what-is-prometheus">What Is Prometheus?&lt;/h3>
&lt;p>Prometheus is an open-source systems monitoring and alerting toolkit, originally built at SoundCloud. Now, it’s a part of the Cloud Native Computing Foundation, providing powerful capabilities for collecting and querying metrics.&lt;/p>
&lt;h3 id="monitoring-prometheus-endpoints-with-netdata">Monitoring Prometheus Endpoints With Netdata&lt;/h3>
&lt;p>To monitor Prometheus endpoints efficiently, Netdata employs an OpenMetrics (Prometheus) exporter. This allows Netdata to ingest data from any Prometheus exporter with ease. Users can leverage Netdata to gain automated dashboards and alerts without the need for deploying a Prometheus server or setting up Grafana. This enables seamless and efficient monitoring of your systems with minimal overhead.&lt;/p></description></item><item><title>Prometheus Endpoint Monitoring</title><link>https://www.netdata.cloud/monitoring-101/prometheus-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/prometheus-monitoring/</guid><description>&lt;p>Prometheus endpoints are HTTP interfaces exposed by various applications and services that provide metrics in a format consumable by Prometheus. These endpoints follow the OpenMetrics exposition format, which ensures compatibility with Prometheus monitoring. By collecting and analyzing metrics from these endpoints, Prometheus enables users to gain deep insights into their infrastructure&amp;rsquo;s performance and health, helping them detect and resolve issues proactively. The ability to monitor Prometheus endpoints is crucial for maintaining the performance and reliability of modern applications and systems.&lt;/p></description></item><item><title>Prometheus Remote Write</title><link>https://www.netdata.cloud/integrations/exporters/prometheus-remote-write/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/prometheus-remote-write/</guid><description/></item><item><title>Prominet Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/prominet-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/prominet-corporation-snmp-traps/</guid><description/></item><item><title>Promise Technology Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/promise-technology-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/promise-technology-inc-snmp-traps/</guid><description/></item><item><title>Protection One Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/protection-one-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/protection-one-inc-snmp-traps/</guid><description/></item><item><title>Prowl</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/prowl/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/prowl/</guid><description/></item><item><title>Proxim Wireless Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/proxim-wireless-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/proxim-wireless-inc-snmp-traps/</guid><description/></item><item><title>Proxmox VE</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/proxmox-ve/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/proxmox-ve/</guid><description/></item><item><title>Proxmox VE Monitoring</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/proxmox-ve-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/proxmox-ve-monitoring/</guid><description/></item><item><title>Proxmox VE Monitoring</title><link>https://www.netdata.cloud/monitoring-101/proxmox-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/proxmox-monitoring/</guid><description>&lt;h2 id="proxmox-ve-monitoring">Proxmox VE Monitoring&lt;/h2>
&lt;h3 id="what-is-proxmox-ve">What Is Proxmox VE?&lt;/h3>
&lt;p>Proxmox VE (Virtual Environment) is a comprehensive open-source platform for enterprise virtualization, combining powerful KVM hypervisor technology with robust container-based solutions. It allows IT teams to manage virtual machines, containers, highly available clusters, and software-defined storage, all in a single convenient solution.&lt;/p>
&lt;h3 id="monitoring-proxmox-ve-with-netdata">Monitoring Proxmox VE With Netdata&lt;/h3>
&lt;p>To efficiently monitor Proxmox VE, Netdata leverages an openmetrics (prometheus) exporter. Netdata seamlessly ingests data from any Prometheus exporter. This enables IT professionals to benefit from automated dashboards, alerts, and detailed statistics, without the need for a dedicated Prometheus server or Grafana setup. With &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata&lt;/a>, you can visualize metrics in real-time, easily detect anomalies, and ensure your virtual environment operates optimally.&lt;/p></description></item><item><title>Proxmox VMs and Containers</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/proxmox-vms-and-containers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/proxmox-vms-and-containers/</guid><description/></item><item><title>ProxySQL</title><link>https://www.netdata.cloud/integrations/data-collection/databases/proxysql/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/proxysql/</guid><description/></item><item><title>ProxySQL Active_Transactions pinning backend connections: multiplexing lost to open transactions</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-transactions-pinning-connections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-transactions-pinning-connections/</guid><description>&lt;p>Active_Transactions is climbing in stats_mysql_global. ConnUsed is growing across backends. But the Questions rate hasn&amp;rsquo;t changed. You check for slow queries, backend errors, new traffic patterns. Nothing explains it.&lt;/p>
&lt;p>An open transaction pins a backend connection for its entire duration. ProxySQL disables multiplexing on that session until the client issues COMMIT or ROLLBACK. If transactions are long-running, or if sessions hold transactions open and never commit, backend connections accumulate in ConnUsed with no corresponding query throughput increase. ConnFree trends toward zero. Eventually queries queue, ConnPool_get_conn_failure starts climbing, and new queries hit the max connect timeout.&lt;/p></description></item><item><title>ProxySQL all replicas lagging: reads falling back onto the writer</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-all-replicas-lagging/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-all-replicas-lagging/</guid><description>&lt;h1 id="proxysql-all-replicas-lagging-reads-falling-back-onto-the-writer">ProxySQL all replicas lagging: reads falling back onto the writer&lt;/h1>
&lt;p>You configured a read/write split with replicas in the reader hostgroup, each with &lt;code>max_replication_lag&lt;/code> set to keep stale reads away from clients. Now all of them are SHUNNED and the reader hostgroup has zero ONLINE backends. Read queries are either failing or falling back onto the writer, which absorbs the full read and write workload.&lt;/p>
&lt;p>The cascade is the danger. As each replica lags and shuns, its traffic redirects to the remaining ONLINE replicas, increasing their load. The last standing replica absorbs all read traffic, its lag climbs, and it shuns too. The reader hostgroup goes empty and every read hits the writer.&lt;/p></description></item><item><title>ProxySQL backend connection pool exhausted: queries queuing for a free connection</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-backend-connection-pool-exhausted/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-backend-connection-pool-exhausted/</guid><description>&lt;h1 id="proxysql-backend-connection-pool-exhausted-queries-queuing-for-a-free-connection">ProxySQL backend connection pool exhausted: queries queuing for a free connection&lt;/h1>
&lt;p>Queries are queuing. Client-facing latency is climbing. Applications are reporting timeouts or &amp;ldquo;too many connections&amp;rdquo; errors. You look at ProxySQL&amp;rsquo;s backend pool and see &lt;code>ConnFree&lt;/code> at zero across one or more backends. The natural assumption is that the pool is exhausted.&lt;/p>
&lt;p>It might not be. &lt;code>ConnFree == 0&lt;/code> alone is not saturation. ProxySQL can still create new backend connections up to the &lt;code>max_connections&lt;/code> limit configured in &lt;code>mysql_servers&lt;/code>. The pool is under pressure, but ProxySQL has headroom to open more connections if the backend can accept them.&lt;/p></description></item><item><title>ProxySQL backend flapping between ONLINE and SHUNNED: monitor-induced oscillation</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-backend-flapping-online-shunned/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-backend-flapping-online-shunned/</guid><description>&lt;h1 id="proxysql-backend-flapping-between-online-and-shunned-monitor-induced-oscillation">ProxySQL backend flapping between ONLINE and SHUNNED: monitor-induced oscillation&lt;/h1>
&lt;p>A ProxySQL backend alternating between ONLINE and SHUNNED more than 3 times in 10 minutes is a distinct failure pattern from a persistent SHUNNED state. The backend MySQL is healthy when checked directly, but ProxySQL&amp;rsquo;s monitor module produces false-negative health decisions on a jittery network path or under transient backend pressure. Each oscillation kills active connections, causing ConnERR spikes, latency spikes on surviving backends, and dips in Questions as traffic is disrupted and redistributed.&lt;/p></description></item><item><title>ProxySQL backend SHUNNED: why a healthy backend gets pulled out of rotation</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-backend-shunned/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-backend-shunned/</guid><description>&lt;h1 id="proxysql-backend-shunned-why-a-healthy-backend-gets-pulled-out-of-rotation">ProxySQL backend SHUNNED: why a healthy backend gets pulled out of rotation&lt;/h1>
&lt;p>You see &lt;code>status=SHUNNED&lt;/code> on a backend in &lt;code>stats_mysql_connection_pool&lt;/code> or &lt;code>runtime_mysql_servers&lt;/code>. The backend itself looks healthy when you connect directly. Queries may be failing or routing to fewer backends than expected, but the MySQL server is up and accepting connections.&lt;/p>
&lt;p>SHUNNED is ProxySQL&amp;rsquo;s protection mechanism, not a bug. When a backend generates more than &lt;code>mysql-shun_on_failures&lt;/code> connection errors within one second, or when its replication lag exceeds the configured &lt;code>max_replication_lag&lt;/code>, ProxySQL temporarily removes it from rotation. The backend typically returns to ONLINE on its own within &lt;code>mysql-shun_recovery_time_sec&lt;/code> (default 10 seconds), provided there is activity in the connection pool for that hostgroup.&lt;/p></description></item><item><title>ProxySQL backend_lagging_during_query and backend_offline_during_query: in-flight query failures</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-backend-lagging-offline-during-query/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-backend-lagging-offline-during-query/</guid><description>&lt;h1 id="proxysql-backend_lagging_during_query-and-backend_offline_during_query-in-flight-query-failures">ProxySQL backend_lagging_during_query and backend_offline_during_query: in-flight query failures&lt;/h1>
&lt;p>Two counters in &lt;code>stats_mysql_global&lt;/code> track queries that were in flight when the backend changed state underneath them. &lt;code>backend_lagging_during_query&lt;/code> counts queries that may have returned stale data. &lt;code>backend_offline_during_query&lt;/code> counts queries that failed. Both are cumulative since the last ProxySQL restart; a nonzero value after steady-state means something disrupted a running query.&lt;/p>
&lt;p>&lt;code>backend_offline_during_query&lt;/code> sustained increase is a TICKET by the playbook severity model. Each increment means a query was routed to a backend that then disappeared mid-execution: crash, network partition, SHUNNED transition, or operator-triggered OFFLINE_HARD. The diagnostic work is correlating counter increments with the specific backend transitions that caused them.&lt;/p></description></item><item><title>ProxySQL client connections at mysql-max_connections: frontend saturation and rejected clients</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-client-connections-max-reached/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-client-connections-max-reached/</guid><description>&lt;p>When &lt;code>Client_Connections_connected&lt;/code> reaches &lt;code>mysql-max_connections&lt;/code> (default 2048), ProxySQL stops accepting new client sessions. Already-connected clients continue working, but new sessions fail with &amp;ldquo;Too many connections&amp;rdquo; until existing ones close.&lt;/p>
&lt;p>This is a cliff-edge failure. There is no queueing at the frontend and no backpressure signal. Because &lt;code>mysql-max_connections&lt;/code> is a global limit shared across all users and hostgroups, a single application with a connection leak can consume every available slot and lock out every other user of the proxy.&lt;/p></description></item><item><title>ProxySQL Client_Connections_aborted rising: clients rejected or crashing on connect</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-client-connections-aborted/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-client-connections-aborted/</guid><description>&lt;h1 id="proxysql-client_connections_aborted-rising-clients-rejected-or-crashing-on-connect">ProxySQL Client_Connections_aborted rising: clients rejected or crashing on connect&lt;/h1>
&lt;p>&lt;code>Client_Connections_aborted&lt;/code> is a cumulative counter in &lt;code>stats_mysql_global&lt;/code>. What matters is the rate of change, not the absolute value. A sustained non-zero rate means clients are failing to establish or maintain connections through ProxySQL.&lt;/p>
&lt;p>The diagnostic question is whether aborts come from ProxySQL actively rejecting connections (authentication failures, connection limits), clients disconnecting improperly (timeouts, crashes, protocol mismatches), or the Client Error Limit feature banning source addresses. The &lt;code>Access_Denied_*&lt;/code> counters in the same stats table identify which. If aborts rise while the &lt;code>Questions&lt;/code> rate drops, clients are experiencing a real outage. If &lt;code>Questions&lt;/code> is stable, the impact may be limited to a subset of clients.&lt;/p></description></item><item><title>ProxySQL Cluster checksum mismatch: split-brain routing across proxy peers</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-cluster-checksum-mismatch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-cluster-checksum-mismatch/</guid><description>&lt;h1 id="proxysql-cluster-checksum-mismatch-split-brain-routing-across-proxy-peers">ProxySQL Cluster checksum mismatch: split-brain routing across proxy peers&lt;/h1>
&lt;p>ProxySQL Cluster peers exchange configuration via a pull-based sync protocol. When sync stalls or conflicts, peers diverge silently. Clients on proxy A get one set of backends; clients on proxy B get a different set. The same query routes to different hostgroups. Credentials differ. Query rules match differently.&lt;/p>
&lt;p>The detection signal is &lt;code>stats_proxysql_servers_checksums&lt;/code>. Each peer computes a checksum per config module. When peers agree, checksums match. When they disagree, &lt;code>diff_check&lt;/code> climbs for the affected module. Persistent mismatch means the cluster is not converging.&lt;/p></description></item><item><title>ProxySQL config changes not applied: the LOAD TO RUNTIME / SAVE TO DISK trap</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-config-changes-not-applied/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-config-changes-not-applied/</guid><description>&lt;h1 id="proxysql-config-changes-not-applied-the-load-to-runtime--save-to-disk-trap">ProxySQL config changes not applied: the LOAD TO RUNTIME / SAVE TO DISK trap&lt;/h1>
&lt;p>You INSERT or UPDATE a row in &lt;code>mysql_servers&lt;/code>, &lt;code>mysql_users&lt;/code>, or &lt;code>mysql_query_rules&lt;/code> through the admin interface. The SQL succeeds with no error. But traffic behavior does not change. Or it works fine for a day, then ProxySQL restarts and everything reverts.&lt;/p>
&lt;p>The root cause is the three-layer configuration model. Every admin SQL change lands in MEMORY, the staging area. That change is not active until you explicitly run &lt;code>LOAD ... TO RUNTIME&lt;/code>. And it does not survive a restart until you explicitly run &lt;code>SAVE ... TO DISK&lt;/code>. No transition is automatic. No error fires if you skip a step.&lt;/p></description></item><item><title>ProxySQL config lost after restart: runtime never saved to disk</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-config-lost-after-restart/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-config-lost-after-restart/</guid><description>&lt;h1 id="proxysql-config-lost-after-restart-runtime-never-saved-to-disk">ProxySQL config lost after restart: runtime never saved to disk&lt;/h1>
&lt;p>ProxySQL separates live configuration from persistent configuration. If you apply a change to RUNTIME but never run &lt;code>SAVE ... TO DISK&lt;/code>, that change survives only until the next restart. When ProxySQL restarts, it loads from &lt;code>proxysql.db&lt;/code> on disk, and everything unsaved is gone.&lt;/p>
&lt;p>The three-layer model means making a change live and making it durable are two separate, explicit steps. There is no single &amp;ldquo;save everything&amp;rdquo; command. If you forget any module&amp;rsquo;s &lt;code>SAVE ... TO DISK&lt;/code>, that module reverts on the next restart.&lt;/p></description></item><item><title>ProxySQL connection storm after restart: an empty pool meeting a mass reconnect</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-connection-storm-after-restart/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-connection-storm-after-restart/</guid><description>&lt;h1 id="proxysql-connection-storm-after-restart-an-empty-pool-meeting-a-mass-reconnect">ProxySQL connection storm after restart: an empty pool meeting a mass reconnect&lt;/h1>
&lt;p>ProxySQL just restarted. Within seconds, client connections pour in. The backend pool is empty, so every query demands a new &lt;code>connect()&lt;/code> to MySQL. &lt;code>Client_Connections_created&lt;/code> and backend &lt;code>ConnOK&lt;/code> spike in lockstep. The backend&amp;rsquo;s CPU jumps as it authenticates connections and spins up threads instead of executing queries. If TLS is enabled, each handshake multiplies the cost.&lt;/p>
&lt;p>The storm is usually self-resolving in 30 to 120 seconds as the pool fills and steady-state multiplexing resumes. During that window, latency spikes, queries may queue, and if backend &lt;code>max_connections&lt;/code> is exhausted, new connections fail outright with MySQL error 1040.&lt;/p></description></item><item><title>ProxySQL ConnERR climbing: backend connection errors and how to localise them</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-connerr-backend-connection-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-connerr-backend-connection-errors/</guid><description>&lt;h1 id="proxysql-connerr-climbing-backend-connection-errors-and-how-to-localise-them">ProxySQL ConnERR climbing: backend connection errors and how to localise them&lt;/h1>
&lt;p>&lt;code>ConnERR&lt;/code> is a per-backend cumulative counter in &lt;code>stats_mysql_connection_pool&lt;/code> that increments every time ProxySQL fails to open a connection to a backend MySQL server. Because it is cumulative, a single nonzero value is historical: it means at least one attempt failed since the last stats reset or ProxySQL restart. Take two readings 10 to 30 seconds apart and compute the delta. A positive sustained delta means ProxySQL is actively failing to establish backend connections right now.&lt;/p></description></item><item><title>ProxySQL ConnPool_get_conn_failure rising: the most direct pool-starvation signal</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-connpool-get-conn-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-connpool-get-conn-failure/</guid><description>&lt;h1 id="proxysql-connpool_get_conn_failure-rising-the-most-direct-pool-starvation-signal">ProxySQL ConnPool_get_conn_failure rising: the most direct pool-starvation signal&lt;/h1>
&lt;p>When &lt;code>ConnPool_get_conn_failure&lt;/code> in &lt;code>stats_mysql_global&lt;/code> starts climbing, ProxySQL&amp;rsquo;s internal connection pool failed to hand a backend connection to a worker thread on request. ProxySQL asked the pool for a connection to a hostgroup and got nothing back. That makes it the most actionable signal for backend pool starvation, more direct than watching &lt;code>ConnUsed&lt;/code> and &lt;code>ConnFree&lt;/code> in isolation.&lt;/p>
&lt;p>This counter does not mean clients are seeing errors. It increments on every internal pool-lookup miss, even if ProxySQL subsequently creates a new backend connection or retries successfully. A spike of hundreds of thousands of failures in a few seconds can correspond to zero client-visible errors. What it tells you: the pool was pressured enough to miss, and sustained rising rates mean that pressure is not resolving on its own.&lt;/p></description></item><item><title>ProxySQL error 1040 Too many connections: the backend MySQL rejecting the pool</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-too-many-connections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-too-many-connections/</guid><description>&lt;h1 id="proxysql-error-1040-too-many-connections-the-backend-mysql-rejecting-the-pool">ProxySQL error 1040 Too many connections: the backend MySQL rejecting the pool&lt;/h1>
&lt;p>Your application queries are failing with &lt;code>ERROR 1040 (HY000): Too many connections&lt;/code>. ProxySQL is up, the admin interface responds, and some backends show ONLINE. But queries against certain hostgroups error out, and the problem is spreading. The error is coming from the backend MySQL server itself, not from ProxySQL&amp;rsquo;s own pool ceiling.&lt;/p>
&lt;p>Four independent limits can produce error 1040 or something visually identical to the client: ProxySQL&amp;rsquo;s global &lt;code>mysql-max_connections&lt;/code> (default 2048), per-user &lt;code>max_connections&lt;/code> in &lt;code>mysql_users&lt;/code>, per-backend &lt;code>max_connections&lt;/code> in &lt;code>mysql_servers&lt;/code>, and the backend MySQL&amp;rsquo;s server-level &lt;code>max_connections&lt;/code>. When the backend&amp;rsquo;s limit is hit, ProxySQL tries to open a connection, the backend refuses, &lt;code>ConnERR&lt;/code> climbs, and the backend may be shunned after enough failures. This article covers that fourth case.&lt;/p></description></item><item><title>ProxySQL error 1045 Access denied for user: credential rotation not propagated</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-access-denied-for-user/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-access-denied-for-user/</guid><description>&lt;h1 id="proxysql-error-1045-access-denied-for-user-credential-rotation-not-propagated">ProxySQL error 1045 Access denied for user: credential rotation not propagated&lt;/h1>
&lt;p>Error 1045 is MySQL&amp;rsquo;s &amp;ldquo;Access denied for user.&amp;rdquo; In a ProxySQL environment it can surface at three independent authentication boundaries, and fixing the wrong one wastes the incident window. The typical trigger is a password rotation applied to one layer (backend MySQL, application config, or ProxySQL itself) without propagating it through ProxySQL&amp;rsquo;s three-layer configuration model.&lt;/p>
&lt;p>The error string carries an immediate clue. &lt;code>(using password: YES)&lt;/code> means a password was sent but did not match. &lt;code>(using password: NO)&lt;/code> means no password was sent at all, which can indicate a client misconfiguration or, rarely, a version-specific bug. Most credential rotation incidents produce &lt;code>(using password: YES)&lt;/code>.&lt;/p></description></item><item><title>ProxySQL error 1290 The MySQL server is running with the --read-only option</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-server-read-only-option/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-server-read-only-option/</guid><description>&lt;h1 id="proxysql-error-1290-the-mysql-server-is-running-with-the---read-only-option">ProxySQL error 1290 The MySQL server is running with the &amp;ndash;read-only option&lt;/h1>
&lt;p>MySQL error 1290 in a ProxySQL environment is almost always a read/write split misconfiguration: a write query that should have gone to the primary was routed to a read-only replica, and the replica refused it. In ProxySQL, this surfaces in &lt;code>stats_mysql_errors&lt;/code> with &lt;code>errno=1290&lt;/code>, attributed to a reader hostgroup backend. The proxy is not malfunctioning. It is routing queries exactly as its &lt;code>mysql_query_rules&lt;/code> dictate. The rules are wrong.&lt;/p></description></item><item><title>ProxySQL error 1317 Query execution was interrupted: timeout-killed queries</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-query-execution-interrupted/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-query-execution-interrupted/</guid><description>&lt;h1 id="proxysql-error-1317-query-execution-was-interrupted-timeout-killed-queries">ProxySQL error 1317 Query execution was interrupted: timeout-killed queries&lt;/h1>
&lt;p>Error 1317 (&amp;ldquo;Query execution was interrupted&amp;rdquo;) means ProxySQL killed a backend query that exceeded a configured timeout. When a query runs past its limit, ProxySQL opens a separate connection to the backend, issues &lt;code>KILL QUERY &amp;lt;thread_id&amp;gt;&lt;/code>, and returns error 1317 to the client.&lt;/p>
&lt;p>The question is whether 1317 is a healthy safety valve catching runaway queries, or a symptom of backend degradation killing queries that should normally succeed. Raising a timeout is cheap, but masking backend slowness with a higher ceiling only delays the problem.&lt;/p></description></item><item><title>ProxySQL error 9001 Max connect timeout reached while reaching hostgroup</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-max-connect-timeout-reached-hostgroup/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-max-connect-timeout-reached-hostgroup/</guid><description>&lt;h1 id="proxysql-error-9001-max-connect-timeout-reached-while-reaching-hostgroup">ProxySQL error 9001 Max connect timeout reached while reaching hostgroup&lt;/h1>
&lt;p>Your application is receiving MySQL error 9001 (HY000): &amp;ldquo;Max connect timeout reached while reaching hostgroup N after Mms.&amp;rdquo; ProxySQL tried to obtain a working backend connection for the client query and failed within the configured timeout window (&lt;code>mysql-connect_timeout_server_max&lt;/code>, default 10000ms).&lt;/p>
&lt;p>Every backend in the target hostgroup is either down, SHUNNED, OFFLINE, or saturated to the point where no connection could be borrowed or created in time. Queries are failing.&lt;/p></description></item><item><title>ProxySQL Galera and Group Replication hostgroups: automatic writer/reader routing</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-galera-group-replication-hostgroups/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-galera-group-replication-hostgroups/</guid><description>&lt;h1 id="proxysql-galera-and-group-replication-hostgroups-automatic-writerreader-routing">ProxySQL Galera and Group Replication hostgroups: automatic writer/reader routing&lt;/h1>
&lt;p>ProxySQL&amp;rsquo;s &lt;code>mysql_galera_hostgroups&lt;/code> and &lt;code>mysql_group_replication_hostgroups&lt;/code> tables automate writer/reader hostgroup assignment for synchronous replication clusters. Instead of manually assigning nodes to writer and reader hostgroups in &lt;code>mysql_servers&lt;/code>, you define the hostgroup IDs and let the monitor module reassign nodes based on live cluster topology. When a Galera cluster&amp;rsquo;s active writer shifts during a split-brain resolution, or a Group Replication member wins a primary election, ProxySQL re-routes traffic without operator intervention.&lt;/p></description></item><item><title>ProxySQL hostgroup_locked connections: reading the multiplexing-health ratio</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-hostgroup-locked-connections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-hostgroup-locked-connections/</guid><description>&lt;h1 id="proxysql-hostgroup_locked-connections-reading-the-multiplexing-health-ratio">ProxySQL hostgroup_locked connections: reading the multiplexing-health ratio&lt;/h1>
&lt;p>ProxySQL exists to multiplex: N client connections served by M backend connections, where M is ideally much less than N. When that ratio degrades, the proxy adds latency without providing pooling benefit. The single metric that captures this degradation is &lt;code>Client_Connections_hostgroup_locked&lt;/code> divided by &lt;code>Client_Connections_connected&lt;/code>.&lt;/p>
&lt;p>ORMs and connection libraries can silently disable multiplexing by emitting &lt;code>SET&lt;/code> commands or opening transactions on every connection. The proxy runs at near 1:1 for months, and the capacity plan that assumed 10:1 multiplexing is fiction.&lt;/p></description></item><item><title>ProxySQL memory growth and OOM: what drives RSS and how to tell a leak from load</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-memory-growth-oom/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-memory-growth-oom/</guid><description>&lt;h1 id="proxysql-memory-growth-and-oom-what-drives-rss-and-how-to-tell-a-leak-from-load">ProxySQL memory growth and OOM: what drives RSS and how to tell a leak from load&lt;/h1>
&lt;p>When ProxySQL RSS climbs toward system limits, the first question is whether growth is load-driven (more connections, cached result sets, or query digests) or a genuine leak. The fixes are different: load-driven growth requires tuning buffer sizes, cache limits, or digest configuration. A leak requires identifying the subsystem that is not releasing memory and possibly upgrading to a patched version.&lt;/p></description></item><item><title>ProxySQL monitor check failures: connect, ping, read-only, and replication-lag probes failing</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-monitor-check-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-monitor-check-failures/</guid><description>&lt;h1 id="proxysql-monitor-check-failures-connect-ping-read-only-and-replication-lag-probes-failing">ProxySQL monitor check failures: connect, ping, read-only, and replication-lag probes failing&lt;/h1>
&lt;p>When &lt;code>MySQL_Monitor_connect_check_ERR&lt;/code>, &lt;code>MySQL_Monitor_ping_check_ERR&lt;/code>, &lt;code>MySQL_Monitor_read_only_check_ERR&lt;/code>, or &lt;code>MySQL_Monitor_replication_lag_check_ERR&lt;/code> start climbing in &lt;code>stats_mysql_global&lt;/code>, the ProxySQL monitor module has lost visibility into one or more backends. These counters are a leading indicator. The actual user-facing impact shows up later as backend status transitions, typically SHUNNED, when ProxySQL can no longer verify a backend is healthy and pulls it from rotation.&lt;/p>
&lt;p>The monitor module runs four independent check types, each on its own schedule, each writing results to its own log table in the &lt;code>monitor&lt;/code> schema. Monitor connections use credentials from &lt;code>mysql-monitor_username&lt;/code> and &lt;code>mysql-monitor_password&lt;/code>, which are completely separate from the data-plane credentials in &lt;code>mysql_users&lt;/code>. This separation is the single most common source of confusion: application traffic can continue flowing normally while the monitor silently fails in the background.&lt;/p></description></item><item><title>ProxySQL monitor password invalid: healthy backends shunned because health checks fail</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-monitor-password-invalid/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-monitor-password-invalid/</guid><description>&lt;h1 id="proxysql-monitor-password-invalid-healthy-backends-shunned-because-health-checks-fail">ProxySQL monitor password invalid: healthy backends shunned because health checks fail&lt;/h1>
&lt;p>Backends are going SHUNNED, but direct connections to the MySQL servers work fine. The servers are up, responsive, and serving queries normally. Yet ProxySQL has pulled them from rotation.&lt;/p>
&lt;p>The backends are not the problem. The Monitor module is.&lt;/p>
&lt;p>When &lt;code>mysql-monitor_password&lt;/code> is wrong, expired, or out of sync, every monitor connect and ping check fails with an authentication error. ProxySQL interprets these failures as backend health failures. After &lt;code>mysql-monitor_ping_max_failures&lt;/code> (default: 3) consecutive ping failures, it SHUNS the backend and closes its connections. The backend is healthy, but ProxySQL cannot verify that, so it stops sending traffic.&lt;/p></description></item><item><title>ProxySQL Monitoring</title><link>https://www.netdata.cloud/monitoring-101/proxysql-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/proxysql-monitoring/</guid><description>&lt;h2 id="proxysql-monitoring">ProxySQL Monitoring&lt;/h2>
&lt;h3 id="what-is-proxysql">What Is ProxySQL?&lt;/h3>
&lt;p>&lt;a href="https://www.proxysql.com/">ProxySQL&lt;/a> is a high-performance SQL proxy designed to manage multiple back-end MySQL servers behind a single proxy interface. It enhances MySQL database scalability and reliability, supporting advanced query routing and connection pooling, which leads to improved performance for your database workloads.&lt;/p>
&lt;h3 id="monitoring-proxysql-with-netdata">Monitoring ProxySQL With Netdata&lt;/h3>
&lt;p>Monitoring ProxySQL is vital for maintaining the health and performance of your database layer. Netdata provides a comprehensive ProxySQL monitoring tool that offers real-time insights into various metrics, enabling you to diagnose and troubleshoot any issues promptly. With &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/proxysql/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata&amp;rsquo;s ProxySQL integration&lt;/a>, you can easily observe key performance metrics and customize your monitoring according to your needs.&lt;/p></description></item><item><title>ProxySQL monitoring checklist: the signals every production proxy needs</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-monitoring-checklist/</guid><description>&lt;h1 id="proxysql-monitoring-checklist-the-signals-every-production-proxy-needs">ProxySQL monitoring checklist: the signals every production proxy needs&lt;/h1>
&lt;p>ProxySQL sits between your applications and MySQL-compatible backends, fully parsing the MySQL wire protocol and making per-query routing, caching, and connection-pooling decisions.&lt;/p>
&lt;p>This checklist defines the monitoring signals every ProxySQL deployment needs, organized into four maturity levels: survival, operational, mature, and expert. Each level builds on the previous one. If you track query digest anomalies at level 4 but cannot tell whether your backends are ONLINE at level 1, you are debugging blind.&lt;/p></description></item><item><title>ProxySQL monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-monitoring-maturity-model/</guid><description>&lt;h1 id="proxysql-monitoring-maturity-model-from-survival-to-expert">ProxySQL monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>ProxySQL deployments tend to land at one of two monitoring extremes: a process check and a port check, or a full dashboard nobody acts on. The gap between &amp;ldquo;is it running&amp;rdquo; and &amp;ldquo;is it healthy&amp;rdquo; is where most incidents live. This article maps four monitoring maturity levels, each adding signals that catch failure modes the previous level cannot.&lt;/p>
&lt;p>Use this as a self-assessment. Find the highest level where you have every signal covered, then look at what the next level adds. All signals come from ProxySQL&amp;rsquo;s own &lt;code>stats_*&lt;/code> tables and host-level metrics, queryable through the admin interface on port 6032. &lt;!-- TODO: verify "the Prometheus exporter on port 6070" -- ProxySQL has no built-in Prometheus exporter; port depends on which external exporter is deployed -->&lt;/p></description></item><item><title>ProxySQL multiplexing collapse: when connection pooling silently drops to 1:1</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-multiplexing-collapse/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-multiplexing-collapse/</guid><description>&lt;h1 id="proxysql-multiplexing-collapse-when-connection-pooling-silently-drops-to-11">ProxySQL multiplexing collapse: when connection pooling silently drops to 1:1&lt;/h1>
&lt;p>You deployed ProxySQL for connection pooling. The capacity plan assumed a 10:1 multiplexing ratio: 1000 client connections served by 100 backend connections. Instead, the backend pool keeps filling up, &lt;code>ConnUsed&lt;/code> tracks the client count almost linearly, and the MySQL backend is hitting &lt;code>max_connections&lt;/code> even though traffic has not changed. ProxySQL is running, backends are ONLINE, queries are succeeding, but the proxy is providing zero pooling benefit. It has silently degraded into a 1:1 connection relay with overhead.&lt;/p></description></item><item><title>ProxySQL MySQL_Monitor_Workers is zero: health checks stopped and status is stale</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-monitor-workers-zero/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-monitor-workers-zero/</guid><description>&lt;h1 id="proxysql-mysql_monitor_workers-is-zero-health-checks-stopped-and-status-is-stale">ProxySQL MySQL_Monitor_Workers is zero: health checks stopped and status is stale&lt;/h1>
&lt;p>When &lt;code>MySQL_Monitor_Workers&lt;/code> in &lt;code>stats_mysql_global&lt;/code> reads zero and stays there, ProxySQL&amp;rsquo;s monitor module has stopped probing backends. No connect checks, no ping checks, no read-only checks, no replication lag checks. The status values in &lt;code>runtime_mysql_servers&lt;/code> and &lt;code>stats_mysql_connection_pool&lt;/code> are frozen at whatever they were when the last check ran. A backend that crashed five minutes ago still shows ONLINE.&lt;/p>
&lt;p>This state is silent. The data plane keeps routing queries. Nothing crashes, nothing logs an error visible to most dashboards. The only outward signal is that ProxySQL stops reacting to real backend health changes: a writer failover goes undetected, a lagging replica stays in rotation, a dead backend keeps receiving queries until clients time out.&lt;/p></description></item><item><title>ProxySQL OFFLINE_SOFT vs OFFLINE_HARD vs SHUNNED: what each backend status means</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-backend-offline-hard-soft/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-backend-offline-hard-soft/</guid><description>&lt;h1 id="proxysql-offline_soft-vs-offline_hard-vs-shunned-what-each-backend-status-means">ProxySQL OFFLINE_SOFT vs OFFLINE_HARD vs SHUNNED: what each backend status means&lt;/h1>
&lt;p>ProxySQL tracks four backend statuses: ONLINE, SHUNNED, OFFLINE_SOFT, and OFFLINE_HARD. OFFLINE_SOFT and OFFLINE_HARD are operator-controlled drain states. SHUNNED is automatic and monitor-driven. The distinction determines who sets the state, how it recovers, and which table shows the truth.&lt;/p>
&lt;p>A backend that is OFFLINE_SOFT is deliberately draining and will not recover until you change it back. A backend that is SHUNNED is temporarily avoided by the monitor and self-corrects. Conflating the two leads to unnecessary intervention (restarting backends that would recover on their own) or dangerous neglect (assuming a draining backend will fix itself during a failover).&lt;/p></description></item><item><title>ProxySQL per-user max_connections: one application starving the shared proxy</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-per-user-max-connections/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-per-user-max-connections/</guid><description>&lt;h1 id="proxysql-per-user-max_connections-one-application-starving-the-shared-proxy">ProxySQL per-user max_connections: one application starving the shared proxy&lt;/h1>
&lt;p>ProxySQL enforces two layers of frontend connection limits: a global ceiling (&lt;code>mysql-max_connections&lt;/code>, default 2048) and an optional per-user cap (&lt;code>mysql_users.max_connections&lt;/code>, default 10000). The per-user cap is the one most teams leave at default, which is effectively no limit. A single application with a connection leak or misconfigured pool can consume connections toward the global ceiling. Once the global pool is exhausted, every other application sharing that ProxySQL instance is denied new connections.&lt;/p></description></item><item><title>ProxySQL query cache hit rate dropped: backend load about to surge</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-query-cache-hit-rate-dropped/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-query-cache-hit-rate-dropped/</guid><description>&lt;h1 id="proxysql-query-cache-hit-rate-dropped-backend-load-about-to-surge">ProxySQL query cache hit rate dropped: backend load about to surge&lt;/h1>
&lt;p>A query cache hit rate drop is a leading indicator. By the time backend CPU spikes or the connection pool saturates, the cache has already stopped absorbing read load.&lt;/p>
&lt;p>ProxySQL&amp;rsquo;s query cache is a TTL-based in-memory result set cache with no invalidation on data change. When it works, it shields backends from repetitive read queries. When it stops, every previously-cached query becomes a real backend query, and backend load increases proportionally to the miss delta.&lt;/p></description></item><item><title>ProxySQL query cache memory growth: high-cardinality caching toward OOM</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-query-cache-memory-growth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-query-cache-memory-growth/</guid><description>&lt;h1 id="proxysql-query-cache-memory-growth-high-cardinality-caching-toward-oom">ProxySQL query cache memory growth: high-cardinality caching toward OOM&lt;/h1>
&lt;p>&lt;code>Query_Cache_Memory_bytes&lt;/code> sits at the &lt;code>mysql-query_cache_size_MB&lt;/code> ceiling, the purge rate is rising, and hit rate is dropping despite a full cache. This is the high-cardinality caching pattern: the cache is churning through unique keys that never repeat, and the soft limit cannot keep RSS bounded because it only triggers eviction of expired entries.&lt;/p>
&lt;p>The worst case is OOM kill. The kernel terminates ProxySQL, clients disconnect, and on restart the cache is empty, stats tables reset, and the cycle starts again. If you are mid-incident, run &lt;code>PROXYSQL FLUSH QUERY CACHE&lt;/code> for immediate relief. Then fix the root cause before memory grows back.&lt;/p></description></item><item><title>ProxySQL query cache stampede: a hot entry expires and the herd hits the backend</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-query-cache-stampede/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-query-cache-stampede/</guid><description>&lt;h1 id="proxysql-query-cache-stampede-a-hot-entry-expires-and-the-herd-hits-the-backend">ProxySQL query cache stampede: a hot entry expires and the herd hits the backend&lt;/h1>
&lt;p>Backend connection pool usage spikes suddenly. Backend latency climbs. Query cache hit rate drops to near zero for a brief window, then recovers minutes later. If this pattern repeats at regular intervals, a cache stampede is the likely cause.&lt;/p>
&lt;p>ProxySQL&amp;rsquo;s query cache is stale-read, keyed by query digest, user, and schema. It uses TTL-based expiration only. There is no request coalescing, no lock-based regeneration, and no mechanism to ensure only one request repopulates the cache after a miss. When a hot entry expires, every concurrent request for that query independently hits the backend MySQL.&lt;/p></description></item><item><title>ProxySQL query digest memory growth: stats_mysql_query_digest growing unbounded</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-query-digest-memory-growth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-query-digest-memory-growth/</guid><description>&lt;h1 id="proxysql-query-digest-memory-growth-stats_mysql_query_digest-growing-unbounded">ProxySQL query digest memory growth: stats_mysql_query_digest growing unbounded&lt;/h1>
&lt;p>ProxySQL&amp;rsquo;s &lt;code>stats_mysql_query_digest&lt;/code> table grows without bound. There is no built-in memory cap. On high-cardinality workloads where queries embed unique identifiers (timestamps, UUIDs, session tokens, savepoint names), the digest hash table can balloon to gigabytes, tracked as &lt;code>query_digest_memory&lt;/code> in &lt;code>stats_memory_metrics&lt;/code>.&lt;/p>
&lt;p>The symptoms: ProxySQL runs normally for days or weeks, then develops periodic latency spikes that correlate with monitoring scrapes. RSS creeps upward. In extreme cases, the process is OOM-killed. The digest table is rarely inspected for size, only for query content, so the root cause stays hidden.&lt;/p></description></item><item><title>ProxySQL query rule CPU overload: expensive regex saturating the worker threads</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-query-rule-cpu-overload/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-query-rule-cpu-overload/</guid><description>&lt;h1 id="proxysql-query-rule-cpu-overload-expensive-regex-saturating-the-worker-threads">ProxySQL query rule CPU overload: expensive regex saturating the worker threads&lt;/h1>
&lt;p>ProxySQL host CPU is pegged. Client query latency is climbing. But the MySQL backends are idle: &lt;code>ConnFree&lt;/code> is greater than zero across the pool, backend ping latency is normal, and there are no slow queries on the database side. The proxy itself is the bottleneck.&lt;/p>
&lt;p>When the proxy&amp;rsquo;s worker threads are saturated, every query slows down uniformly regardless of which backend it targets or how complex the SQL is. The symptom looks like a backend problem from the application&amp;rsquo;s perspective, but the databases are fine.&lt;/p></description></item><item><title>ProxySQL query rule not matching: zero hits, silent regex failures, and broken routing</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-query-rule-not-matching/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-query-rule-not-matching/</guid><description>&lt;h1 id="proxysql-query-rule-not-matching-zero-hits-silent-regex-failures-and-broken-routing">ProxySQL query rule not matching: zero hits, silent regex failures, and broken routing&lt;/h1>
&lt;p>A critical query rule (write routing, read/write splitting, query caching) shows zero hits in &lt;code>stats_mysql_query_rules&lt;/code>. Queries that should match it are falling through to the default hostgroup or a catch-all rule instead. The rule exists in &lt;code>runtime_mysql_query_rules&lt;/code>, it looks correct, and there are no errors in the ProxySQL log.&lt;/p>
&lt;p>The proxy parsed your rule, compiled the regex, and loaded it to runtime. It just never matches any query. Meanwhile, writes may be landing on read-only replicas (MySQL error 1290) or reads may be piling onto the writer hostgroup unnecessarily. If replicas are not configured read_only, writes can silently succeed on a non-primary backend and never reach the authoritative source.&lt;/p></description></item><item><title>ProxySQL query rule order and apply=1: how rule chains silently mis-route</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-query-rules-order-apply/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-query-rules-order-apply/</guid><description>&lt;h1 id="proxysql-query-rule-order-and-apply1-how-rule-chains-silently-mis-route">ProxySQL query rule order and apply=1: how rule chains silently mis-route&lt;/h1>
&lt;p>ProxySQL query rules look deceptively simple: write a regex, point it at a hostgroup, done. But &lt;code>mysql_query_rules&lt;/code> is an ordered chain, and the chain&amp;rsquo;s behavior depends on three interacting variables that most operators never think about together: &lt;code>rule_id&lt;/code> order, the &lt;code>apply&lt;/code> flag, and &lt;code>flagIN&lt;/code>/&lt;code>flagOUT&lt;/code> chaining. Get any of these wrong and queries route to the wrong backend silently, with no error.&lt;/p></description></item><item><title>ProxySQL read/write split misrouting: writes silently reaching replicas</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-read-write-split-misrouting/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-read-write-split-misrouting/</guid><description>&lt;h1 id="proxysql-readwrite-split-misrouting-writes-silently-reaching-replicas">ProxySQL read/write split misrouting: writes silently reaching replicas&lt;/h1>
&lt;p>ProxySQL read/write split routes queries based on &lt;code>mysql_query_rules&lt;/code> evaluated in ascending &lt;code>rule_id&lt;/code> order. When that chain is misconfigured, write queries can match a read-routing rule and land on a replica hostgroup. If the replica enforces &lt;code>super_read_only&lt;/code>, the write fails with MySQL error 1290 and the client sees an error. If it does not, the write succeeds on the replica, never reaches the primary, and the client receives a success response. Data diverges silently with no error at any layer.&lt;/p></description></item><item><title>ProxySQL replica shunned for replication lag: max_replication_lag and stale-read protection</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-replica-shunned-replication-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-replica-shunned-replication-lag/</guid><description>&lt;h1 id="proxysql-replica-shunned-for-replication-lag-max_replication_lag-and-stale-read-protection">ProxySQL replica shunned for replication lag: max_replication_lag and stale-read protection&lt;/h1>
&lt;p>When a MySQL replica falls behind its primary, ProxySQL&amp;rsquo;s monitor module detects the lag and pulls that replica out of read rotation by marking it SHUNNED. This is deliberate stale-read protection: applications should not read from a replica that has not applied recent writes, because they would see stale results or violate read-after-write consistency.&lt;/p>
&lt;p>The mechanism is governed by &lt;code>max_replication_lag&lt;/code>, a per-backend column in &lt;code>mysql_servers&lt;/code> measured in seconds. When the monitor&amp;rsquo;s replication lag check finds &lt;code>Seconds_Behind_Master&lt;/code> exceeding that threshold, the backend is shunned until lag drops back below it. The tradeoff is direct: shunning protects consistency but reduces read capacity. In a topology with two or three replicas, losing one shifts 33 to 50 percent of read traffic onto the remaining backends, which can push them over the same threshold.&lt;/p></description></item><item><title>ProxySQL runtime vs memory drift: the unapplied change that surfaces on restart</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-runtime-vs-memory-drift/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-runtime-vs-memory-drift/</guid><description>&lt;h1 id="proxysql-runtime-vs-memory-drift-the-unapplied-change-that-surfaces-on-restart">ProxySQL runtime vs memory drift: the unapplied change that surfaces on restart&lt;/h1>
&lt;p>ProxySQL&amp;rsquo;s three-layer configuration model makes changes safe and atomic, but every edit requires explicit promotion through two transitions. A missed step creates silent drift between what is staged (MEMORY), what is active (RUNTIME), and what persists on disk (DISK). The drift produces no error, no log entry, and no metric. The first visible symptom is often a restart that reverts to a stale configuration.&lt;/p></description></item><item><title>ProxySQL Server_Connections_delayed above zero: queries waiting on the backend pool</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-server-connections-delayed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-server-connections-delayed/</guid><description>&lt;h1 id="proxysql-server_connections_delayed-above-zero-queries-waiting-on-the-backend-pool">ProxySQL Server_Connections_delayed above zero: queries waiting on the backend pool&lt;/h1>
&lt;p>When &lt;code>Server_Connections_delayed&lt;/code> in ProxySQL&amp;rsquo;s &lt;code>stats_mysql_global&lt;/code> table climbs above zero, queries are waiting for a backend connection that was not immediately available. This is not an error counter. It is a pressure signal: ProxySQL&amp;rsquo;s backend connection pool could not instantly satisfy a connection request, so the requesting session had to wait. A brief blip during a traffic burst is normal. A sustained increase means the pool is undersized, multiplexing has degraded, or backends are disappearing from rotation faster than the pool can adapt.&lt;/p></description></item><item><title>ProxySQL SET statements disabling multiplexing: session state that pins connections</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-set-session-variables-disable-multiplexing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-set-session-variables-disable-multiplexing/</guid><description>&lt;h1 id="proxysql-set-statements-disabling-multiplexing-session-state-that-pins-connections">ProxySQL SET statements disabling multiplexing: session state that pins connections&lt;/h1>
&lt;p>ProxySQL multiplexes N client sessions across M backend connections, where M is smaller than N. When an application issues a SET statement that changes session state, ProxySQL can no longer safely reuse that backend connection for other clients. The connection pins to that client session until the client disconnects.&lt;/p>
&lt;p>This is correct behavior. A backend connection carrying a modified session variable (such as &lt;code>sql_mode&lt;/code> or &lt;code>time_zone&lt;/code>) would produce different query results if handed to another client expecting the default. ProxySQL detects this state and pins the connection to protect correctness.&lt;/p></description></item><item><title>ProxySQL Too many open files: file descriptor exhaustion takes down connections and health checks</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-too-many-open-files/</guid><description>&lt;h1 id="proxysql-too-many-open-files-file-descriptor-exhaustion-takes-down-connections-and-health-checks">ProxySQL Too many open files: file descriptor exhaustion takes down connections and health checks&lt;/h1>
&lt;p>ProxySQL hits a host-level cliff that its own stats tables cannot see. Every client, backend, monitor, and admin session consumes a file descriptor. When the process reaches its &lt;code>RLIMIT_NOFILE&lt;/code> ceiling, &lt;code>accept()&lt;/code> and &lt;code>connect()&lt;/code> fail simultaneously: new client connections are rejected, backend connections cannot be established, and the monitor module cannot open sockets to check backend health. The proxy appears to hang while the process keeps running.&lt;/p></description></item><item><title>ProxySQL worker thread CPU saturation: mysql-threads as a hard parallelism ceiling</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-cpu-thread-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-cpu-thread-saturation/</guid><description>&lt;h1 id="proxysql-worker-thread-cpu-saturation-mysql-threads-as-a-hard-parallelism-ceiling">ProxySQL worker thread CPU saturation: mysql-threads as a hard parallelism ceiling&lt;/h1>
&lt;p>When every query through ProxySQL slows down simultaneously, regardless of backend or query digest, the proxy itself is the bottleneck. The most common cause is worker thread CPU saturation: all &lt;code>mysql-threads&lt;/code> worker threads are pegged at or near 100% CPU, and their epoll event loops can no longer service connections without adding queuing delay.&lt;/p>
&lt;p>&lt;code>mysql-threads&lt;/code> (default 4) cannot be changed at runtime. It is set at startup and caps query-processing parallelism absolutely. Adding backends or client connections does not raise this ceiling.&lt;/p></description></item><item><title>ProxySQL zero ONLINE backends in a hostgroup: total outage for that traffic class</title><link>https://www.netdata.cloud/guides/proxysql/proxysql-no-online-backends-hostgroup/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/proxysql/proxysql-no-online-backends-hostgroup/</guid><description>&lt;h1 id="proxysql-zero-online-backends-in-a-hostgroup-total-outage-for-that-traffic-class">ProxySQL zero ONLINE backends in a hostgroup: total outage for that traffic class&lt;/h1>
&lt;p>When ProxySQL reports zero ONLINE backends in a hostgroup, every query routed to that hostgroup fails immediately. If the affected hostgroup handles writes, all INSERT, UPDATE, and DELETE operations fail. If it is a reader hostgroup, read queries either error out or, depending on query rules, flood a writer hostgroup not sized for that load.&lt;/p>
&lt;p>Clients see the error &lt;!-- TODO: verify exact wording across ProxySQL versions --> &amp;ldquo;Hostgroup X has no servers available!&amp;rdquo; or generic connection timeouts. Application error rates spike. Questions drops toward zero while Client_Connections_aborted rises.&lt;/p></description></item><item><title>Pulizzi Engineering Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pulizzi-engineering-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pulizzi-engineering-inc-snmp-traps/</guid><description/></item><item><title>Pulse Power And Measurement Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pulse-power-and-measurement-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pulse-power-and-measurement-ltd-snmp-traps/</guid><description/></item><item><title>Puppet</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/puppet/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/puppet/</guid><description/></item><item><title>Puppet Monitoring</title><link>https://www.netdata.cloud/monitoring-101/puppet-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/puppet-monitoring/</guid><description>&lt;h2 id="puppet-monitoring">Puppet Monitoring&lt;/h2>
&lt;h3 id="what-is-puppet">What Is Puppet?&lt;/h3>
&lt;p>Puppet is a powerful tool for managing and automating the configuration of servers and applications in IT environments. It allows for defining infrastructure as code, enabling consistent and repeatable system setups. Visit the &lt;a href="https://www.puppet.com/">official Puppet site&lt;/a> to learn more.&lt;/p>
&lt;h3 id="monitoring-puppet-with-netdata">Monitoring Puppet With Netdata&lt;/h3>
&lt;p>Using a $name monitoring tool like Netdata provides real-time, per-second visibility into the performance and health of your Puppet infrastructure. Netdata’s lightweight and intuitive dashboards can help diagnose performance issues and ensure the smooth operation of your systems.&lt;/p></description></item><item><title>Pure Storage SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pure-storage-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/pure-storage-snmp-traps/</guid><description/></item><item><title>Pushbullet</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/pushbullet/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/pushbullet/</guid><description/></item><item><title>PushOver</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/pushover/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/pushover/</guid><description/></item><item><title>Qlogic SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/qlogic-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/qlogic-snmp-traps/</guid><description/></item><item><title>Qnap Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/qnap-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/qnap-systems-inc-snmp-traps/</guid><description/></item><item><title>Qsan Technology Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/qsan-technology-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/qsan-technology-inc-snmp-traps/</guid><description/></item><item><title>Qtech LLC SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/qtech-llc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/qtech-llc-snmp-traps/</guid><description/></item><item><title>Qualix Group Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/qualix-group-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/qualix-group-inc-snmp-traps/</guid><description/></item><item><title>Quanta Computer Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/quanta-computer-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/quanta-computer-inc-snmp-traps/</guid><description/></item><item><title>Quantum Bridge SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/quantum-bridge-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/quantum-bridge-snmp-traps/</guid><description/></item><item><title>Quantum Corp Formerly Pathlight Technology Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/quantum-corp-formerly-pathlight-technology-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/quantum-corp-formerly-pathlight-technology-inc-snmp-traps/</guid><description/></item><item><title>Quantum Corporation Formerly Advanced Digital Information Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/quantum-corporation-formerly-advanced-digital-information-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/quantum-corporation-formerly-advanced-digital-information-corporation-snmp-traps/</guid><description/></item><item><title>QuasarDB</title><link>https://www.netdata.cloud/integrations/exporters/quasardb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/quasardb/</guid><description/></item><item><title>RabbitMQ</title><link>https://www.netdata.cloud/integrations/data-collection/databases/rabbitmq/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/rabbitmq/</guid><description/></item><item><title>RabbitMQ Monitoring</title><link>https://www.netdata.cloud/monitoring-101/rabbitmq-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/rabbitmq-monitoring/</guid><description>&lt;h2 id="rabbitmq-monitoring">RabbitMQ Monitoring&lt;/h2>
&lt;h3 id="what-is-rabbitmq">What Is RabbitMQ?&lt;/h3>
&lt;p>RabbitMQ is an open-source message broker that facilitates communication between distributed systems by sending messages back and forth. It&amp;rsquo;s widely used in enterprise and cloud environments to build robust messaging environments that excel in scalability and flexibility. Learn more about &lt;a href="https://www.rabbitmq.com/">RabbitMQ&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-rabbitmq-with-netdata">Monitoring RabbitMQ With Netdata&lt;/h3>
&lt;p>Netdata provides a comprehensive RabbitMQ monitoring tool that allows you to keep a real-time watch over your RabbitMQ instances. Using Netdata’s rich UI and advanced visualizations, users can easily understand their RabbitMQ environment, troubleshoot issues, and optimize performance. With automated alerts, you stay informed about critical events without manual effort.&lt;/p></description></item><item><title>Racktivity SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/racktivity-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/racktivity-snmp-traps/</guid><description/></item><item><title>Racom S R O SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/racom-s-r-o-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/racom-s-r-o-snmp-traps/</guid><description/></item><item><title>Rad Data Communications Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rad-data-communications-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rad-data-communications-ltd-snmp-traps/</guid><description/></item><item><title>Radio Thermostat</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/radio-thermostat/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/radio-thermostat/</guid><description/></item><item><title>Radio Thermostat Monitoring</title><link>https://www.netdata.cloud/monitoring-101/radio_thermostat-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/radio_thermostat-monitoring/</guid><description>&lt;h2 id="radio-thermostat-monitoring">Radio Thermostat Monitoring&lt;/h2>
&lt;h3 id="what-is-radio-thermostat">What Is Radio Thermostat?&lt;/h3>
&lt;p>Radio Thermostat is a line of smart thermostats that offer precision control over home temperature settings, contributing to efficient home automation and energy management. These devices can be integrated with home networks, allowing users to manage their heating, ventilation, and air conditioning systems remotely through smart devices.&lt;/p>
&lt;h3 id="monitoring-radio-thermostat-with-netdata">Monitoring Radio Thermostat With Netdata&lt;/h3>
&lt;p>To monitor Radio Thermostat, Netdata leverages an openmetrics (prometheus) exporter, specifically the &lt;a href="https://github.com/andrewlow/radio-thermostat-exporter">Radio Thermostat Exporter&lt;/a>. This empowers users with detailed insights into heating and cooling trends, enabling proactive adjustments and energy savings.&lt;/p></description></item><item><title>RADIUS</title><link>https://www.netdata.cloud/integrations/data-collection/applications/radius/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/radius/</guid><description/></item><item><title>RADIUS Monitoring</title><link>https://www.netdata.cloud/monitoring-101/radius-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/radius-monitoring/</guid><description>&lt;h2 id="radius-monitoring">RADIUS Monitoring&lt;/h2>
&lt;h3 id="what-is-radius">What Is RADIUS?&lt;/h3>
&lt;p>RADIUS, or Remote Authentication Dial-In User Service, is a networking protocol that provides centralized Authentication, Authorization, and Accounting (AAA) management for users who connect and use a network service. It is a vital component for many network security and management setups, allowing for streamlined and secure user access management. RADIUS is often used by ISPs and large organizations to manage user credentials.&lt;/p>
&lt;h3 id="monitoring-radius-with-netdata">Monitoring RADIUS With Netdata&lt;/h3>
&lt;p>Monitoring RADIUS with Netdata is a seamless experience thanks to the openmetrics (Prometheus) exporter support. Netdata&amp;rsquo;s $name monitoring tool leverages the RADIUS exporter to collect essential protocol metrics without the need for a dedicated Prometheus server or Grafana dashboards. This capability allows DevOps, SREs, developers, and IT admins to monitor RADIUS performance in real-time with automated dashboards and alerts, ensuring efficient authentication and access management.&lt;/p></description></item><item><title>Radwin Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/radwin-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/radwin-ltd-snmp-traps/</guid><description/></item><item><title>Rapid City Communication SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rapid-city-communication-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rapid-city-communication-snmp-traps/</guid><description/></item><item><title>Rapidstream Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rapidstream-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rapidstream-inc-snmp-traps/</guid><description/></item><item><title>Raptor Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/raptor-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/raptor-systems-inc-snmp-traps/</guid><description/></item><item><title>Raritan Computer Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/raritan-computer-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/raritan-computer-inc-snmp-traps/</guid><description/></item><item><title>Raritan Dominion</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/raritan-dominion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/raritan-dominion/</guid><description/></item><item><title>Raritan PDU</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/raritan-pdu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/raritan-pdu/</guid><description/></item><item><title>Raritan PDU Monitoring</title><link>https://www.netdata.cloud/monitoring-101/raritan_pdu-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/raritan_pdu-monitoring/</guid><description>&lt;h2 id="raritan-pdu-monitoring">Raritan PDU Monitoring&lt;/h2>
&lt;h3 id="what-is-raritan-pdu">What Is Raritan PDU?&lt;/h3>
&lt;p>Raritan Power Distribution Units (PDUs) are critical devices used for power management in data centers and IT environments. They provide reliable power distribution to network devices and allow remote power control, ensuring efficient and uninterrupted operations. Monitoring Raritan PDU involves tracking metrics such as power usage, voltage, and current, which are essential for maintaining optimal performance and preventing downtime.&lt;/p>
&lt;h3 id="monitoring-raritan-pdu-with-netdata">Monitoring Raritan PDU With Netdata&lt;/h3>
&lt;p>To monitor Raritan PDU effectively, Netdata utilizes an openmetrics (Prometheus) exporter. This approach allows Netdata to ingest data seamlessly from any Prometheus exporter. This simplifies the monitoring architecture by removing the need for a standalone Prometheus server or Grafana for visualization. With Netdata, you gain immediate access to automated dashboards and alerts, enabling real-time monitoring of your Raritan PDU.&lt;/p></description></item><item><title>Raw_Read_Error_Rate looks enormous: the Seagate false alarm explained</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-raw-read-error-rate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-raw-read-error-rate/</guid><description>&lt;h1 id="raw_read_error_rate-looks-enormous-the-seagate-false-alarm-explained">Raw_Read_Error_Rate looks enormous: the Seagate false alarm explained&lt;/h1>
&lt;p>You see SMART attribute ID 1 with a raw value of 200,450,784. Or 60,000,000,000. Your monitoring system pages you at 3 a.m. because &amp;ldquo;raw read error rate is enormous.&amp;rdquo; You check the drive: &lt;code>smartctl -H&lt;/code> says PASSED. Reallocated sectors: zero. Pending sectors: zero. Offline uncorrectable: zero. The raw value is enormous by design.&lt;/p>
&lt;p>Seagate packs error counts and total operation counts together into the 48-bit raw value field for attribute ID 1 (Raw_Read_Error_Rate). The raw value will always be large on a healthy Seagate drive because it includes every read operation the drive has ever performed, not just errors. Alerting on &lt;code>raw &amp;gt; 0&lt;/code> for this attribute generates constant noise for every Seagate drive in your fleet.&lt;/p></description></item><item><title>Rbb Rundfunk Berlin Brandenburg SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rbb-rundfunk-berlin-brandenburg-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rbb-rundfunk-berlin-brandenburg-snmp-traps/</guid><description/></item><item><title>Reading docker system df: where Docker disk usage actually lives</title><link>https://www.netdata.cloud/guides/docker/docker-system-df/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/docker/docker-system-df/</guid><description>&lt;h1 id="reading-docker-system-df-where-docker-disk-usage-actually-lives">Reading docker system df: Where Docker Disk Usage Actually Lives&lt;/h1>
&lt;p>You run &lt;code>docker system df&lt;/code> but the numbers do not add up to what &lt;code>df -h&lt;/code> reports. Maybe the Build Cache row is empty while &lt;code>/var/lib/docker/buildkit/&lt;/code> consumes tens of gigabytes. Maybe RECLAIMABLE is high but &lt;code>docker system prune&lt;/code> barely frees space because overlay2 layer sharing masks the real unique cost. Or the daemon returns &lt;code>Error response from daemon&lt;/code> when the disk is already full. This guide shows how to read &lt;code>docker system df&lt;/code> precisely, what it hides, and how to triage the real consumers on the host.&lt;/p></description></item><item><title>Reading EXPLAIN ANALYZE: the operator's guide to PostgreSQL query plans</title><link>https://www.netdata.cloud/guides/postgres/postgres-explain-analyze-reading/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/postgres/postgres-explain-analyze-reading/</guid><description>&lt;h1 id="reading-explain-analyze-the-operators-guide-to-postgresql-query-plans">Reading EXPLAIN ANALYZE: the operator&amp;rsquo;s guide to PostgreSQL query plans&lt;/h1>
&lt;p>When &lt;code>pg_stat_statements&lt;/code> flags a query as a top consumer, &lt;code>EXPLAIN ANALYZE&lt;/code> is the operator&amp;rsquo;s ground truth. It shows what the executor did, node by node, buffer by buffer. Misreading the output leads to useless indexes and production changes that make performance worse.&lt;/p>
&lt;p>This guide covers the mechanics that matter in production: how actual time accumulates through the node tree, why estimated rows diverge from reality, when buffer counts reveal cache misses versus disk reads, and how artifacts like the loops multiplier hide expensive nodes. It is a field manual for deciding, in the next five minutes, whether the problem is a missing index, stale statistics, a bad plan choice, or something deeper.&lt;/p></description></item><item><title>Reading NVIDIA GPU memory correctly: nvidia-smi vs the PyTorch caching allocator</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-memory-usage-pytorch-caching-allocator/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-memory-usage-pytorch-caching-allocator/</guid><description>&lt;h1 id="reading-nvidia-gpu-memory-correctly-nvidia-smi-vs-the-pytorch-caching-allocator">Reading NVIDIA GPU memory correctly: nvidia-smi vs the PyTorch caching allocator&lt;/h1>
&lt;p>An on-call classic: nvidia-smi shows 38 of 40 GiB used on a training node, the line is flat, and someone declares a memory leak. The team running the job insists the model only needs 12 GiB of tensors and nothing is growing. Both sides are reading real numbers. They are reading different layers of the stack, and neither tool tells you that on its own.&lt;/p></description></item><item><title>Reading the ATA error log: UNC, ICRC, ABRT, CCTO, IDNF, AMNF</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-ata-error-log-unc/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-ata-error-log-unc/</guid><description>&lt;h1 id="reading-the-ata-error-log-unc-icrc-abrt-ccto-idnf-amnf">Reading the ATA error log: UNC, ICRC, ABRT, CCTO, IDNF, AMNF&lt;/h1>
&lt;p>When &lt;code>smartctl -l error&lt;/code> returns entries instead of &amp;ldquo;No Errors Logged&amp;rdquo;, the error type field in each entry tells you what failed. The ATA Summary Error Log records individual I/O failure events with the error type, the LBA where the error occurred, the command that triggered it, and a timestamp relative to the current power cycle.&lt;/p>
&lt;p>The six error type codes are UNC, ICRC, ABRT, CCTO, IDNF, and AMNF. They are not equally serious. UNC means confirmed data loss. ICRC means a cable problem. ABRT may mean nothing at all. Knowing which one you are looking at determines whether you are evacuating data or reseating a cable.&lt;/p></description></item><item><title>Real-Time Observability ROI Calculator</title><link>https://www.netdata.cloud/value/roi/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/value/roi/</guid><description/></item><item><title>Reallocated_Event_Count vs Reallocated_Sector_Ct: reading both together</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-reallocated-event-count/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-reallocated-event-count/</guid><description>&lt;h1 id="reallocated_event_count-vs-reallocated_sector_ct-reading-both-together">Reallocated_Event_Count vs Reallocated_Sector_Ct: reading both together&lt;/h1>
&lt;p>Two SMART attributes track the same underlying process, sector reallocation, but count different things. Attribute ID 5 (Reallocated_Sector_Ct) counts sectors the drive has permanently retired and replaced with spares. Attribute ID 196 (Reallocated_Event_Count) counts remap operations the firmware initiated. When each bad sector fails individually and is remapped one at a time, the two counters move in lockstep. When a single physical event damages multiple sectors at once (head slap, thermal hotspot, manufacturing defect), the counters diverge: ID 196 increments once for the event while ID 5 jumps by the number of sectors involved.&lt;/p></description></item><item><title>Reallocated_Sector_Ct rising: the drive is burning through its spare pool</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-reallocated-sectors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-reallocated-sectors/</guid><description>&lt;h1 id="reallocated_sector_ct-rising-the-drive-is-burning-through-its-spare-pool">Reallocated_Sector_Ct rising: the drive is burning through its spare pool&lt;/h1>
&lt;p>When &lt;code>smartctl -A&lt;/code> shows Reallocated_Sector_Ct (SMART attribute ID 5) climbing, the drive firmware is silently remapping sectors that have become unreliable. Each reallocation consumes a sector or block from a finite spare pool reserved at the factory. Once that pool is exhausted, the next bad sector becomes an uncorrectable read error. Data loss follows.&lt;/p>
&lt;p>The absolute count tells you less than the rate of change. A drive with 10 reallocated sectors that has been stable for 3 years is operating within its tolerances. A drive that gained 10 reallocated sectors in the last week is actively failing.&lt;/p></description></item><item><title>Red Creek Communications Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/red-creek-communications-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/red-creek-communications-inc-snmp-traps/</guid><description/></item><item><title>Red Hat Enterprise Linux</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/red-hat-enterprise-linux/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/red-hat-enterprise-linux/</guid><description/></item><item><title>Red Lion Controls N Tron SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/red-lion-controls-n-tron-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/red-lion-controls-n-tron-snmp-traps/</guid><description/></item><item><title>Red Lion Controls Sixnet SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/red-lion-controls-sixnet-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/red-lion-controls-sixnet-snmp-traps/</guid><description/></item><item><title>Redis</title><link>https://www.netdata.cloud/integrations/data-collection/databases/redis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/redis/</guid><description/></item><item><title>Redis and Transparent Huge Pages: why THP must be disabled</title><link>https://www.netdata.cloud/guides/redis/redis-disable-transparent-huge-pages/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-disable-transparent-huge-pages/</guid><description>&lt;h1 id="redis-and-transparent-huge-pages-why-thp-must-be-disabled">Redis and Transparent Huge Pages: why THP must be disabled&lt;/h1>
&lt;p>Redis latency spikes during background saves, AOF rewrites, or replica full resyncs often trace to a frozen main thread. Clients time out. Replicas disconnect. If the write rate is high, the next reconnection triggers another fork, and the cascade repeats. One common root cause is Transparent Huge Pages (THP), enabled by default on most Linux distributions. Redis detects THP at startup and logs a warning, but provisioning automation often buries it. The impact is severe: THP can increase fork latency by 10 to 100 times by amplifying copy-on-write memory traffic.&lt;/p></description></item><item><title>Redis aof_last_write_status:err: AOF write failures and recovery</title><link>https://www.netdata.cloud/guides/redis/redis-aof-last-write-status-err/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-aof-last-write-status-err/</guid><description>&lt;h1 id="redis-aof_last_write_statuserr-aof-write-failures-and-recovery">Redis aof_last_write_status:err: AOF write failures and recovery&lt;/h1>
&lt;p>&lt;code>INFO persistence&lt;/code> showing &lt;code>aof_last_write_status:err&lt;/code> means Redis failed to flush its Append-Only File buffer to disk on the last attempt. If you depend on AOF for durability, the instance is no longer persisting writes. With the default &lt;code>appendfsync everysec&lt;/code>, Redis logs the failed fsync and retries, but once &lt;code>aof_last_write_status&lt;/code> is &lt;code>err&lt;/code> and &lt;code>stop-writes-on-bgsave-error&lt;/code> is enabled, the server rejects mutations. The error returned to clients references RDB snapshots even when AOF is the actual failure, which often misleads first-line diagnosis.&lt;/p></description></item><item><title>Redis appendfsync always latency: durability vs throughput trade-offs</title><link>https://www.netdata.cloud/guides/redis/redis-appendfsync-always-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-appendfsync-always-latency/</guid><description>&lt;h1 id="redis-appendfsync-always-latency-durability-vs-throughput-trade-offs">Redis appendfsync always latency: durability vs throughput trade-offs&lt;/h1>
&lt;p>Redis AOF persistence logs every write to disk, but the &lt;code>appendfsync&lt;/code> policy controls how aggressively Redis forces that log to physical storage. That choice defines whether a crash loses zero, one, or sixty seconds of data, and whether a single slow disk seek can freeze the entire event loop.&lt;/p>
&lt;p>&lt;code>appendfsync always&lt;/code> on slow storage degrades throughput more severely than dataset growth. &lt;code>appendfsync no&lt;/code> on a critical ledger silently exposes you to massive data loss.&lt;/p></description></item><item><title>Redis big keys: finding the giant key that blocks the event loop</title><link>https://www.netdata.cloud/guides/redis/redis-big-keys-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-big-keys-latency/</guid><description>&lt;h1 id="redis-big-keys-finding-the-giant-key-that-blocks-the-event-loop">Redis big keys: finding the giant key that blocks the event loop&lt;/h1>
&lt;p>Application latency spikes while &lt;code>redis-cli PING&lt;/code> still returns &lt;code>PONG&lt;/code>. Simple &lt;code>GET&lt;/code> commands take hundreds of milliseconds. Aggregate &lt;code>used_memory&lt;/code> looks stable, &lt;code>instantaneous_ops_per_sec&lt;/code> drops, and the slowlog grows. The culprit is often a single oversized key: a sorted set with millions of elements, a hash with millions of fields, or a list fetched with an unbounded range. Redis executes commands sequentially on one main thread; an O(N) command on a giant key blocks every other client until it completes. This guide shows how to find that key and fix it without restarting Redis.&lt;/p></description></item><item><title>Redis blocked_clients growing: dead consumers vs healthy queues</title><link>https://www.netdata.cloud/guides/redis/redis-blocked-clients-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-blocked-clients-growing/</guid><description>&lt;h1 id="redis-blocked_clients-growing-dead-consumers-vs-healthy-queues">Redis blocked_clients growing: dead consumers vs healthy queues&lt;/h1>
&lt;p>&lt;code>blocked_clients&lt;/code> in &lt;code>INFO clients&lt;/code> is climbing. In queue-based architectures this is often normal: workers call &lt;code>BLPOP&lt;/code>, &lt;code>BRPOP&lt;/code>, or &lt;code>XREAD BLOCK&lt;/code> and wait for producers to push work. When &lt;code>blocked_clients&lt;/code> grows while queue depth also grows, consumers are no longer consuming. They may have crashed, been OOM-killed, or stalled on replication lag via &lt;code>WAIT&lt;/code>.&lt;/p>
&lt;p>&lt;code>blocked_clients&lt;/code> counts only clients waiting on explicit blocking commands. It does not capture clients stalled by slow commands like &lt;code>KEYS *&lt;/code> or large &lt;code>SMEMBERS&lt;/code>. A high value is either a healthy signal of an active queue pattern or a pathological signal of dead connections holding slots open forever, especially with timeout &lt;code>0&lt;/code>.&lt;/p></description></item><item><title>Redis BUSY Redis is busy running a script: blocking Lua and how to recover</title><link>https://www.netdata.cloud/guides/redis/redis-busy-running-script/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-busy-running-script/</guid><description>&lt;h1 id="redis-busy-redis-is-busy-running-a-script-blocking-lua-and-how-to-recover">Redis BUSY Redis is busy running a script: blocking Lua and how to recover&lt;/h1>
&lt;p>&lt;code>redis-cli&lt;/code> returns &lt;code>(error) BUSY Redis is busy running a script. You can only call SCRIPT KILL or SHUTDOWN NOSAVE.&lt;/code> Normal commands stall. Redis is not down, but it might as well be: a Lua script is holding the single event loop hostage and will not yield until it finishes or you intervene.&lt;/p>
&lt;p>Redis executes &lt;code>EVAL&lt;/code> and &lt;code>EVALSHA&lt;/code> atomically. While a script runs, no other command processes. Once execution exceeds &lt;code>lua-time-limit&lt;/code> (default 5000 ms), Redis replies with &lt;code>BUSY&lt;/code> to other clients. The script itself continues until it finishes, is killed, or the server shuts down. Whether the script has already performed writes determines whether you can kill it safely or must choose between waiting and a hard shutdown.&lt;/p></description></item><item><title>Redis Can't save in background: fork: Cannot allocate memory - diagnosis and fix</title><link>https://www.netdata.cloud/guides/redis/redis-cant-save-in-background-fork/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-cant-save-in-background-fork/</guid><description>&lt;h1 id="redis-cant-save-in-background-fork-cannot-allocate-memory---diagnosis-and-fix">Redis Can&amp;rsquo;t save in background: fork: Cannot allocate memory - diagnosis and fix&lt;/h1>
&lt;p>Redis logs &lt;code>Can't save in background: fork: Cannot allocate memory&lt;/code>. &lt;code>free -h&lt;/code> shows plenty of free RAM, yet &lt;code>BGSAVE&lt;/code> or &lt;code>BGREWRITEAOF&lt;/code> fails. If &lt;code>stop-writes-on-bgsave-error&lt;/code> is &lt;code>yes&lt;/code> (default), writes fail too. The gap between free RAM and fork failure is the key.&lt;/p>
&lt;p>This is not a simple OOM. It is a kernel commit charge failure. Linux &lt;code>fork()&lt;/code> must account for the worst case where every copy-on-write page is modified. With &lt;code>vm.overcommit_memory=0&lt;/code> (the default), the kernel enforces a heuristic commit limit. When Redis RSS is large, that limit blocks &lt;code>fork()&lt;/code> even with free physical memory. The fix is usually one sysctl, but THP, container limits, and actual RAM headroom determine whether it holds.&lt;/p></description></item><item><title>Redis client output buffer overflow: slow consumers and client-output-buffer-limit</title><link>https://www.netdata.cloud/guides/redis/redis-client-output-buffer-limit/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-client-output-buffer-limit/</guid><description>&lt;h1 id="redis-client-output-buffer-overflow-slow-consumers-and-client-output-buffer-limit">Redis client output buffer overflow: slow consumers and client-output-buffer-limit&lt;/h1>
&lt;p>Redis memory climbs faster than the dataset justifies. &lt;code>used_memory&lt;/code> approaches &lt;code>maxmemory&lt;/code>, keys evict, or the process is OOM-killed, yet the keyspace has not grown. Logs show &amp;ldquo;scheduled to be closed ASAP for overcoming of output buffer limits,&amp;rdquo; or clients vanish and reconnect. The culprit is usually a slow consumer that cannot drain its output buffer as fast as Redis fills it. A forgotten &lt;code>MONITOR&lt;/code> session or an application that left a socket open but stopped reading are the textbook cases. Redis allocates client output buffers from the main heap; unread response data counts against &lt;code>maxmemory&lt;/code>. The default &lt;code>client-output-buffer-limit normal 0 0 0&lt;/code> leaves normal clients unbounded, turning one slow reader into a memory leak that can kill the instance.&lt;/p></description></item><item><title>Redis cluster bus port blocked: the port+10000 firewall gotcha</title><link>https://www.netdata.cloud/guides/redis/redis-cluster-gossip-port-blocked/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-cluster-gossip-port-blocked/</guid><description>&lt;h1 id="redis-cluster-bus-port-blocked-the-port10000-firewall-gotcha">Redis cluster bus port blocked: the port+10000 firewall gotcha&lt;/h1>
&lt;p>&lt;code>CLUSTER INFO&lt;/code> reports &lt;code>cluster_state:fail&lt;/code>. Nodes show non-zero &lt;code>cluster_slots_pfail&lt;/code>. Clients receive &lt;code>CLUSTERDOWN&lt;/code>. Yet &lt;code>redis-cli -p 6379 PING&lt;/code> returns &lt;code>PONG&lt;/code> on every node, application connections are still accepted, and the client port shows no obvious network outage. The cluster behaves like it is partitioned, but only the bus is broken. Port 16379, or your configured client port plus 10000, is missing from a firewall rule, security group, or container port mapping. The cluster bus carries gossip, failure detection, and node discovery over this separate TCP port. When the bus is unreachable, nodes cannot synchronize the cluster map, so they mark peers as failed and withdraw slot coverage even though the data port stays healthy. Because firewall rules often cover the client port but omit the bus port, this failure mode is common after infrastructure changes, node replacements, or environment migrations.&lt;/p></description></item><item><title>Redis cluster_slots_pfail > 0: impending node failure in a cluster</title><link>https://www.netdata.cloud/guides/redis/redis-cluster-slots-pfail/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-cluster-slots-pfail/</guid><description>&lt;h1 id="redis-cluster_slots_pfail--0-impending-node-failure-in-a-cluster">Redis cluster_slots_pfail &amp;gt; 0: impending node failure in a cluster&lt;/h1>
&lt;p>&lt;code>cluster_slots_pfail &amp;gt; 0&lt;/code> means at least one hash slot is mapped to a node that a peer suspects is down. In Redis Cluster, PFAIL is unilateral: any node raises it when another stops answering gossip PINGs for longer than &lt;code>cluster-node-timeout&lt;/code>. Slots continue to serve traffic; the cluster has not yet agreed the node is dead.&lt;/p>
&lt;p>Brief spikes are expected during background saves, AOF rewrites, or any main-thread freeze. Sustained non-zero values indicate a real problem: network partition, node crash, or overload. If the majority of masters confirm the suspicion within twice &lt;code>cluster-node-timeout&lt;/code>, PFAIL escalates to FAIL. The affected slots become unavailable until a replica wins election. In a three-master cluster, losing two primaries leaves the survivor without quorum. The cluster enters a zombie state where no failover can proceed. Investigate PFAIL while you still have quorum and before automatic escalation.&lt;/p></description></item><item><title>Redis CLUSTERDOWN / cluster_state:fail: slot coverage and recovery</title><link>https://www.netdata.cloud/guides/redis/redis-cluster-state-fail/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-cluster-state-fail/</guid><description>&lt;h1 id="redis-clusterdown--cluster_statefail-slot-coverage-and-recovery">Redis CLUSTERDOWN / cluster_state:fail: slot coverage and recovery&lt;/h1>
&lt;p>&lt;code>CLUSTERDOWN The cluster is down&lt;/code> means at least one of the 16384 hash slots lacks a healthy master. With &lt;code>cluster-require-full-coverage yes&lt;/code> (the default), a single missing slot blocks all writes. This guide covers diagnosing the root cause, recovering safely, and preventing recurrence.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>Redis Cluster shards the keyspace across 16384 hash slots. Each slot must be assigned to a master node that is reachable and healthy to count toward &lt;code>cluster_slots_ok&lt;/code>. When &lt;code>cluster_slots_assigned&lt;/code> drops below 16384, or &lt;code>cluster_slots_fail&lt;/code> becomes non-zero because a node has been marked FAIL by quorum, the cluster transitions to &lt;code>cluster_state:fail&lt;/code>. Clients receive &lt;code>CLUSTERDOWN&lt;/code> for operations hashing to affected slots.&lt;/p></description></item><item><title>Redis connected_clients climbing: connection leak detection</title><link>https://www.netdata.cloud/guides/redis/redis-connected-clients-climbing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-connected-clients-climbing/</guid><description>&lt;h1 id="redis-connected_clients-climbing-connection-leak-detection">Redis connected_clients climbing: connection leak detection&lt;/h1>
&lt;p>A sustained climb in &lt;code>connected_clients&lt;/code> over hours or days while application traffic is flat is a classic Redis connection leak. Each connection costs roughly 10 KB of server-side memory. A thousand leaked connections consume ~100 MB independent of your dataset. If the instance is near &lt;code>maxmemory&lt;/code>, that overhead can push Redis into eviction or OOM territory.&lt;/p>
&lt;p>The default &lt;code>timeout&lt;/code> is 0, so idle connections are never closed. Missing &lt;code>close()&lt;/code> calls, connection pool misconfiguration, or unsubscribed pub/sub listeners all accumulate forever.&lt;/p></description></item><item><title>Redis connected_slaves dropped: detecting replica disconnects on the primary</title><link>https://www.netdata.cloud/guides/redis/redis-connected-slaves-dropped/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-connected-slaves-dropped/</guid><description>&lt;h1 id="redis-connected_slaves-dropped-detecting-replica-disconnects-on-the-primary">Redis connected_slaves dropped: detecting replica disconnects on the primary&lt;/h1>
&lt;p>If &lt;code>INFO replication&lt;/code> on a Redis primary shows &lt;code>connected_slaves&lt;/code> lower than expected, the missing replica shrinks read capacity, makes a full resync likely, and widens the data-loss window during failover. The primary does not keep a tombstone: it decrements the counter and drops the corresponding &lt;code>slaveN&lt;/code> line from the next &lt;code>INFO&lt;/code> sample. You need to determine whether the replica crashed, the network partitioned, or Sentinel promoted the replica and the old primary has not caught up.&lt;/p></description></item><item><title>Redis connection exhaustion: leaks, pools, and the retry storm</title><link>https://www.netdata.cloud/guides/redis/redis-connection-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-connection-exhaustion/</guid><description>&lt;h1 id="redis-connection-exhaustion-leaks-pools-and-the-retry-storm">Redis connection exhaustion: leaks, pools, and the retry storm&lt;/h1>
&lt;p>Application logs show connection timeouts and Redis returns &lt;code>ERR max number of clients reached&lt;/code>. Downstream services fail because they cannot reach the cache. &lt;code>INFO clients&lt;/code> shows &lt;code>connected_clients&lt;/code> at the hard limit even though traffic has not increased. This is connection exhaustion. The most dangerous response is a retry storm that turns a small leak into a site-wide cascade.&lt;/p>
&lt;p>Redis enforces a hard upper bound on connections via &lt;code>maxclients&lt;/code>. When the sum of &lt;code>connected_clients&lt;/code>, &lt;code>connected_slaves&lt;/code>, and &lt;code>cluster_connections&lt;/code> reaches that limit, Redis rejects every new TCP connection. Applications that retry immediately without backoff create a feedback loop: existing connections age out slowly while new attempts pile up, keeping the server pinned at the limit even after the original leak stops growing.&lt;/p></description></item><item><title>Redis CPU saturation: hitting the single-core throughput ceiling</title><link>https://www.netdata.cloud/guides/redis/redis-cpu-saturation-single-thread/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-cpu-saturation-single-thread/</guid><description>&lt;h1 id="redis-cpu-saturation-hitting-the-single-core-throughput-ceiling">Redis CPU saturation: hitting the single-core throughput ceiling&lt;/h1>
&lt;p>Redis latency climbs. PING returns PONG, but simple GETs take milliseconds instead of microseconds. Host CPU looks moderate - perhaps 25% across eight cores - yet commands queue. The likely cause is main-thread CPU saturation. Redis executes all commands on a single event-loop thread. Once that thread saturates one core, latency rises linearly with queue depth. There is no performance cliff - only a steady ramp that eventually drives client timeouts. On multi-core hosts, aggregate process CPU hides this bottleneck because background children, I/O threads, and system accounting spread usage across cores.&lt;/p></description></item><item><title>Redis event loop blocked: when one slow command freezes everything</title><link>https://www.netdata.cloud/guides/redis/redis-event-loop-blocked/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-event-loop-blocked/</guid><description>&lt;h1 id="redis-event-loop-blocked-when-one-slow-command-freezes-everything">Redis event loop blocked: when one slow command freezes everything&lt;/h1>
&lt;p>Redis processes every command on a single main thread. Even with Redis 6.0 and later offloading network I/O to threads, command execution itself remains strictly sequential. When one command takes too long, everything behind it waits. You will see clients still connected, but &lt;code>PING&lt;/code> stalls and throughput collapses to zero. This is the Slow Command Snowball pattern: a single expensive operation blocks the event loop, the client queue backs up, and latency compounds across every connected application.&lt;/p></description></item><item><title>Redis eviction policy tuning: allkeys-lru vs volatile-ttl vs noeviction</title><link>https://www.netdata.cloud/guides/redis/redis-eviction-policy-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-eviction-policy-tuning/</guid><description>&lt;h1 id="redis-eviction-policy-tuning-allkeys-lru-vs-volatile-ttl-vs-noeviction">Redis eviction policy tuning: allkeys-lru vs volatile-ttl vs noeviction&lt;/h1>
&lt;p>When Redis reaches &lt;code>maxmemory&lt;/code>, it must either reject new writes or delete existing keys. The &lt;code>maxmemory-policy&lt;/code> directive decides which path it takes, yet many production instances run with a policy that mismatches the workload. A cache running &lt;code>noeviction&lt;/code> returns OOM errors to clients. A database running &lt;code>allkeys-lru&lt;/code> silently deletes committed data. A session store running &lt;code>volatile-ttl&lt;/code> suddenly rejects writes the moment an application bug omits a TTL.&lt;/p></description></item><item><title>Redis exposed without authentication: the CONFIG SET dir crontab attack</title><link>https://www.netdata.cloud/guides/redis/redis-exposed-without-auth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-exposed-without-auth/</guid><description>&lt;h1 id="redis-exposed-without-authentication-the-config-set-dir-crontab-attack">Redis exposed without authentication: the CONFIG SET dir crontab attack&lt;/h1>
&lt;p>An unauthenticated Redis instance on a public interface is remote code execution. The classic attack chains four commands: &lt;code>CONFIG SET dir&lt;/code> to a cron folder, &lt;code>CONFIG SET dbfilename&lt;/code> to a valid cron file, &lt;code>SET&lt;/code> a malicious payload, and &lt;code>SAVE&lt;/code>. If Redis runs as root, the host is compromised immediately. If unprivileged, attackers pivot via SSH keys or systemd timers.&lt;/p>
&lt;p>This guide is an operational audit and lockdown. It covers how the file-write attack works, how to detect exposure, how to check for active compromise, and how to harden the instance. It also covers the follow-on risk: once authenticated access is gained, recently disclosed authenticated RCE bugs can escalate to full system control.&lt;/p></description></item><item><title>Redis FLUSHALL ran in production: detection, prevention, and recovery</title><link>https://www.netdata.cloud/guides/redis/redis-flushall-data-loss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-flushall-data-loss/</guid><description>&lt;h1 id="redis-flushall-ran-in-production-detection-prevention-and-recovery">Redis FLUSHALL ran in production: detection, prevention, and recovery&lt;/h1>
&lt;p>FLUSHALL deletes every key in every logical database. Under default &lt;code>save&lt;/code> policies, the server typically triggers a background save of the empty keyspace within seconds, overwriting the RDB snapshot. There is no undo. If you are responding to an active incident, your priorities are: confirm the scope, isolate unaffected replicas before they process the flush, recover from the freshest intact persistence source, and harden the instance.&lt;/p></description></item><item><title>Redis fork/COW memory storm: why persistence doubles RSS and OOM-kills the box</title><link>https://www.netdata.cloud/guides/redis/redis-fork-cow-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-fork-cow-storm/</guid><description>&lt;h1 id="redis-forkcow-memory-storm-why-persistence-doubles-rss-and-oom-kills-the-box">Redis fork/COW memory storm: why persistence doubles RSS and OOM-kills the box&lt;/h1>
&lt;p>Redis disappeared from your container with only an &lt;code>OOMKilled&lt;/code> status and a metrics gap that aligns with an RDB snapshot or AOF rewrite. The dataset was under its memory limit moments ago, but during persistence the reported RSS doubled and the kernel killed the process.&lt;/p>
&lt;p>This is the Redis fork/copy-on-write memory storm. Redis calls &lt;code>fork()&lt;/code> to spawn a child process for background RDB snapshots, AOF rewrites, and full replication syncs. After the fork, parent and child share pages through copy-on-write. Pages stay read-only until one process writes. If the parent continues serving writes, every modified page is copied. On a write-heavy instance, this can duplicate the entire dataset, pushing RSS to roughly twice the logical data size. Containers with tight memory limits do not see &lt;code>used_memory&lt;/code>; they see RSS. When RSS hits the cgroup ceiling, the OOM killer fires, both processes die, and the instance restarts cold.&lt;/p></description></item><item><title>Redis KEYS command blocking production: why to replace it with SCAN</title><link>https://www.netdata.cloud/guides/redis/redis-keys-command-blocking-production/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-keys-command-blocking-production/</guid><description>&lt;h1 id="redis-keys-command-blocking-production-why-to-replace-it-with-scan">Redis KEYS command blocking production: why to replace it with SCAN&lt;/h1>
&lt;p>Redis P99 latency jumps from sub-millisecond to seconds. Clients time out. &lt;code>instantaneous_ops_per_sec&lt;/code> drops to near zero while &lt;code>connected_clients&lt;/code> stays high. The likely culprit is a single slow command monopolizing the event loop, and &lt;code>KEYS&lt;/code> is the classic offender.&lt;/p>
&lt;p>Redis runs all client commands on one main thread. &lt;code>KEYS&lt;/code> scans the entire keyspace synchronously to match a pattern. Time complexity is O(N) where N is the total number of keys. During the scan, nothing else executes. Every client, including replication streams, health checks, and monitoring probes, waits.&lt;/p></description></item><item><title>Redis keyspace growing unbounded: keys without TTL and memory leaks</title><link>https://www.netdata.cloud/guides/redis/redis-keyspace-growing-unbounded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-keyspace-growing-unbounded/</guid><description>&lt;h1 id="redis-keyspace-growing-unbounded-keys-without-ttl-and-memory-leaks">Redis keyspace growing unbounded: keys without TTL and memory leaks&lt;/h1>
&lt;p>Redis memory climbs. &lt;code>used_memory&lt;/code> approaches &lt;code>maxmemory&lt;/code>, or the OOM killer intervenes. In &lt;code>INFO keyspace&lt;/code>, &lt;code>keys&lt;/code> rises while &lt;code>expires&lt;/code> stays flat. Or both rise, but &lt;code>used_memory_rss&lt;/code> outpaces the dataset. Unbounded keyspace growth is a symptom, not a root cause: missing TTLs, application bugs, expiration backlog, or leaks in client buffers and allocators.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>&lt;code>INFO keyspace&lt;/code> reports &lt;code>db{N}:keys=X,expires=Y,avg_ttl=Z&lt;/code>. The &lt;code>keys&lt;/code> field counts every key; &lt;code>expires&lt;/code> counts only those with a TTL. When &lt;code>keys&lt;/code> grows and &lt;code>expires&lt;/code> does not, new keys are being created without expiration. Even when &lt;code>expires&lt;/code> tracks &lt;code>keys&lt;/code>, memory can still grow if the active expiration cycle cannot delete keys as fast as they are created.&lt;/p></description></item><item><title>Redis latency spikes: diagnosis with the LATENCY subsystem</title><link>https://www.netdata.cloud/guides/redis/redis-latency-spikes-diagnosis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-latency-spikes-diagnosis/</guid><description>&lt;h1 id="redis-latency-spikes-diagnosis-with-the-latency-subsystem">Redis latency spikes: diagnosis with the LATENCY subsystem&lt;/h1>
&lt;p>Your application P99 latency just jumped and clients are timing out. The infrastructure dashboard shows the Redis host is up, memory is not exhausted, and ops per second look normal, yet something is blocking the main thread. It could be a single &lt;code>KEYS *&lt;/code> freezing the event loop, a &lt;code>fork()&lt;/code> duplicating page tables for an RDB save, an AOF &lt;code>fsync&lt;/code> stalling on saturated disk, or the active expire cycle burning CPU. Standard &lt;code>INFO&lt;/code> counters will not tell you which one.&lt;/p></description></item><item><title>Redis latest_fork_usec too high: THP, NUMA, and fork latency</title><link>https://www.netdata.cloud/guides/redis/redis-latest-fork-usec-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-latest-fork-usec-high/</guid><description>&lt;h1 id="redis-latest_fork_usec-too-high-thp-numa-and-fork-latency">Redis latest_fork_usec too high: THP, NUMA, and fork latency&lt;/h1>
&lt;p>&lt;code>INFO stats&lt;/code> shows &lt;code>latest_fork_usec&lt;/code> in the hundreds of milliseconds. Every &lt;code>fork()&lt;/code> blocks the single event loop, so during that window no commands are processed. Clients time out, replicas disconnect, and a full resync can trigger another fork, creating a loop of latency and reconnection storms. A normal fork costs roughly 10-20ms per gigabyte of resident memory with Transparent Huge Pages disabled. If you are seeing 10-100x that, the culprit is usually THP, NUMA, or memory overcommit policy.&lt;/p></description></item><item><title>Redis Linux kernel tuning: vm.overcommit_memory, swappiness, and NUMA</title><link>https://www.netdata.cloud/guides/redis/redis-kernel-tuning-overcommit-swappiness/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-kernel-tuning-overcommit-swappiness/</guid><description>&lt;h1 id="redis-linux-kernel-tuning-vmovercommit_memory-swappiness-and-numa">Redis Linux kernel tuning: vm.overcommit_memory, swappiness, and NUMA&lt;/h1>
&lt;p>Redis uses fork-based copy-on-write for background saves, AOF rewrites, and full replication resyncs. Linux defaults for memory overcommit, swap, page size, and socket queuing suit general-purpose workloads, not an in-memory store that clones multi-gigabyte address spaces. Left unchanged, they produce intermittent &lt;code>Cannot allocate memory&lt;/code> errors during BGSAVE, 10-100x latency spikes during AOF rewrite, and silent OOM kills. This guide covers the five host-level tunables, the failure mode each prevents, and the production signals that expose a misconfigured host.&lt;/p></description></item><item><title>Redis LOADING Redis is loading the dataset in memory - why and how long</title><link>https://www.netdata.cloud/guides/redis/redis-loading-dataset-in-memory/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-loading-dataset-in-memory/</guid><description>&lt;h1 id="redis-loading-redis-is-loading-the-dataset-in-memory---why-and-how-long">Redis LOADING Redis is loading the dataset in memory - why and how long&lt;/h1>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>When &lt;code>INFO persistence&lt;/code> returns &lt;code>loading:1&lt;/code>, Redis is reading an RDB dump or replaying an AOF log into memory. This happens after any restart where persistence files are present. &lt;code>redis-cli PING&lt;/code> returns &lt;code>PONG&lt;/code>, but all data commands such as &lt;code>GET&lt;/code>, &lt;code>SET&lt;/code>, and &lt;code>HGETALL&lt;/code> are rejected with a &lt;code>-LOADING&lt;/code> error. The &lt;code>loading_loaded_perc&lt;/code>, &lt;code>loading_loaded_bytes&lt;/code>, and &lt;code>loading_eta_seconds&lt;/code> fields expose real-time progress. While loading is active, latency, hit rate, and eviction metrics are meaningless. Load duration ranges from seconds for small RDB snapshots to tens of minutes, or even hours, for large AOF files on slow disks.&lt;/p></description></item><item><title>Redis low keyspace hit rate: cache effectiveness and cold-start recovery</title><link>https://www.netdata.cloud/guides/redis/redis-low-keyspace-hit-rate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-low-keyspace-hit-rate/</guid><description>&lt;h1 id="redis-low-keyspace-hit-rate-cache-effectiveness-and-cold-start-recovery">Redis low keyspace hit rate: cache effectiveness and cold-start recovery&lt;/h1>
&lt;p>A low keyspace hit rate turns Redis into a pass-through to the backend. The metric is &lt;code>keyspace_hits / (keyspace_hits + keyspace_misses)&lt;/code>, but interpreting it is not simple. A restarted instance shows 0% for minutes. Workloads heavy on &lt;code>EXISTS&lt;/code> naturally miss. A cache that evicts faster than it is hit degrades silently until backend load spikes. This guide shows how to distinguish real cache degradation from false alarms and recover.&lt;/p></description></item><item><title>Redis mass key expiration spike: TTL jitter and the active expiry cycle</title><link>https://www.netdata.cloud/guides/redis/redis-mass-key-expiration-spike/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-mass-key-expiration-spike/</guid><description>&lt;h1 id="redis-mass-key-expiration-spike-ttl-jitter-and-the-active-expiry-cycle">Redis mass key expiration spike: TTL jitter and the active expiry cycle&lt;/h1>
&lt;p>Your application latency graph just spiked. Redis &lt;code>instantaneous_ops_per_sec&lt;/code> dropped. &lt;code>keyspace_misses&lt;/code> jumped. &lt;code>INFO stats&lt;/code> shows &lt;code>expired_keys&lt;/code> climbing by thousands per second. The cause is likely not a traffic surge, but a wave of keys hitting TTL at the same moment. When millions of keys share an identical expiration time, Redis&amp;rsquo;s active expiry cycle cannot sample and delete them fast enough. The main thread spends increasing time on expiration cleanup, blocking client commands and triggering a cache stampede as expired keys return nil.&lt;/p></description></item><item><title>Redis MASTERDOWN / master_link_status:down: replication link broken</title><link>https://www.netdata.cloud/guides/redis/redis-master-link-status-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-master-link-status-down/</guid><description>&lt;h1 id="redis-masterdown--master_link_statusdown-replication-link-broken">Redis MASTERDOWN / master_link_status:down: replication link broken&lt;/h1>
&lt;p>You see &lt;code>MASTERDOWN&lt;/code> errors from client libraries, or monitoring shows &lt;code>master_link_status:down&lt;/code> on a replica. The replica is still accepting connections and serving reads, but every response is increasingly stale. The primary continues to take writes, so the gap widens. Determine whether this is a transient resync or a real partition, and fix it without forcing an expensive full resync that freezes the primary with a fork.&lt;/p></description></item><item><title>Redis max number of clients reached: maxclients and rejected_connections</title><link>https://www.netdata.cloud/guides/redis/redis-max-number-of-clients-reached/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-max-number-of-clients-reached/</guid><description>&lt;h1 id="redis-max-number-of-clients-reached-maxclients-and-rejected_connections">Redis max number of clients reached: maxclients and rejected_connections&lt;/h1>
&lt;p>&lt;code>ERR max number of clients reached&lt;/code> in application logs, or a rising &lt;code>rejected_connections&lt;/code> metric, means new TCP sockets are refused while existing connections work normally. &lt;code>PING&lt;/code> still returns &lt;code>PONG&lt;/code>. The failure is a hard cliff: once the total connection count hits &lt;code>maxclients&lt;/code>, every new connection is rejected immediately. Because &lt;code>rejected_connections&lt;/code> is cumulative, a flat line is healthy and any upward slope is an active incident.&lt;/p></description></item><item><title>Redis maxmemory not set: why every production instance needs a memory limit</title><link>https://www.netdata.cloud/guides/redis/redis-maxmemory-not-set/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-maxmemory-not-set/</guid><description>&lt;h1 id="redis-maxmemory-not-set-why-every-production-instance-needs-a-memory-limit">Redis maxmemory not set: why every production instance needs a memory limit&lt;/h1>
&lt;p>On 64-bit builds, Redis defaults maxmemory to 0. In production, this is an incident waiting to happen. A value of 0 means Redis tracks used_memory but enforces no limit. The process grows until the host OS OOM killer intervenes, or until a container memory limit triggers a SIGKILL. Redis dies without warning, restarts into a cold cache, and faces a thundering herd of client reconnections that immediately reapply the same write pressure. The cycle repeats until the configuration is fixed.&lt;/p></description></item><item><title>Redis mem_fragmentation_ratio below 1.0: detecting swap death</title><link>https://www.netdata.cloud/guides/redis/redis-swapping-fragmentation-below-one/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-swapping-fragmentation-below-one/</guid><description>&lt;h1 id="redis-mem_fragmentation_ratio-below-10-detecting-swap-death">Redis mem_fragmentation_ratio below 1.0: detecting swap death&lt;/h1>
&lt;p>A &lt;code>mem_fragmentation_ratio&lt;/code> of 0.72 on a substantial dataset means the operating system has swapped out Redis memory pages. Redis stays alive and responds to &lt;code>PING&lt;/code>, but every command touching a swapped key blocks the single event loop on disk I/O. Because Redis logs nothing about swap, the resulting latency catastrophe looks like a mystery.&lt;/p>
&lt;p>&lt;code>mem_fragmentation_ratio&lt;/code> equals &lt;code>used_memory_rss / used_memory&lt;/code>. When the ratio drops below 1.0, the resident set size is smaller than the memory Redis requested from its allocator. The missing bytes are in swap. Operators typically watch for ratios above 1.5, so a low ratio is misread as &amp;ldquo;good fragmentation&amp;rdquo; when it is actually the worst memory-related failure mode short of an OOM kill.&lt;/p></description></item><item><title>Redis mem_fragmentation_ratio high: jemalloc fragmentation and active defrag</title><link>https://www.netdata.cloud/guides/redis/redis-mem-fragmentation-ratio-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-mem-fragmentation-ratio-high/</guid><description>&lt;h1 id="redis-mem_fragmentation_ratio-high-jemalloc-fragmentation-and-active-defrag">Redis mem_fragmentation_ratio high: jemalloc fragmentation and active defrag&lt;/h1>
&lt;p>A &lt;code>mem_fragmentation_ratio&lt;/code> sustained above 1.5 on a production instance means Redis holds significantly more physical memory (RSS) than its logical dataset size (&lt;code>used_memory&lt;/code>), wasting RAM that could hold data or absorb spikes. This is not a memory leak. Redis uses jemalloc by default, which does not return freed pages to the OS eagerly. Deleted or resized keys leave holes in allocator arenas, inflating RSS while &lt;code>used_memory&lt;/code> stays flat or drops.&lt;/p></description></item><item><title>Redis memory pressure spiral: eviction thrashing and how to break it</title><link>https://www.netdata.cloud/guides/redis/redis-memory-pressure-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-memory-pressure-spiral/</guid><description>&lt;h1 id="redis-memory-pressure-spiral-eviction-thrashing-and-how-to-break-it">Redis memory pressure spiral: eviction thrashing and how to break it&lt;/h1>
&lt;p>Redis latency climbs, CPU saturates, and cache hit rate falls. &lt;code>evicted_keys&lt;/code> rises while application writes increase. The backend database gets hammered. This is not a simple capacity shortage; it is a memory pressure spiral. Redis has reached &lt;code>maxmemory&lt;/code> and started evicting keys. The application responds to cache misses by re-fetching from the origin and writing back to Redis. Those writes trigger more evictions, which cause more misses, which cause more writes. Redis does maximum work for minimum value.&lt;/p></description></item><item><title>Redis MONITOR left running: the output-buffer OOM footgun</title><link>https://www.netdata.cloud/guides/redis/redis-monitor-command-oom/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-monitor-command-oom/</guid><description>&lt;h1 id="redis-monitor-left-running-the-output-buffer-oom-footgun">Redis MONITOR left running: the output-buffer OOM footgun&lt;/h1>
&lt;p>A forgotten &lt;code>MONITOR&lt;/code> session typically produces this pattern:&lt;/p>
&lt;ul>
&lt;li>&lt;code>used_memory&lt;/code> and &lt;code>used_memory_rss&lt;/code> climb steadily.&lt;/li>
&lt;li>If &lt;code>maxmemory&lt;/code> is set and an eviction policy is active, keys disappear; otherwise the kernel OOM killer may terminate the process.&lt;/li>
&lt;li>Application latency spikes.&lt;/li>
&lt;li>Network output roughly doubles network input.&lt;/li>
&lt;li>No large keys, no persistence fork, and no replication backlog overflow.&lt;/li>
&lt;/ul>
&lt;p>&lt;code>MONITOR&lt;/code> streams a serialized copy of every executed command into the requesting client&amp;rsquo;s output buffer. That buffer is allocated on the main heap and counts in &lt;code>used_memory&lt;/code>, which contributes to &lt;code>maxmemory&lt;/code> pressure and RSS growth &lt;!-- TODO: verify that normal client output buffers are included in the maxmemory eviction accounting; replica output buffers are explicitly excluded -->. Redis classifies &lt;code>MONITOR&lt;/code> clients as normal clients, and the default &lt;code>client-output-buffer-limit normal 0 0 0&lt;/code> places no bound on normal client output buffers. Under production load, the buffer can grow by multiple gigabytes per minute. The monitor client also consumes roughly one extra copy of outbound network traffic, so output bandwidth approximately doubles.&lt;/p></description></item><item><title>Redis Monitoring</title><link>https://www.netdata.cloud/monitoring-101/redis-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/redis-monitoring/</guid><description>&lt;h2 id="redis-monitoring">Redis Monitoring&lt;/h2>
&lt;h3 id="what-is-redis">What Is Redis?&lt;/h3>
&lt;p>Redis, an in-memory data structure store, is widely used as a distributed, in-memory key-value database, cache, and message broker. With speeds that are difficult to match, Redis plays a crucial role in many real-time applications. You can learn more about Redis on the &lt;a href="https://redis.com/">official Redis website&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-redis-with-netdata">Monitoring Redis With Netdata&lt;/h3>
&lt;p>Monitoring Redis effectively ensures that your applications run smoothly and that issues are diagnosed before they impact your users. The &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/redis/">$name monitoring tool&lt;/a> from Netdata provides real-time, thorough insights into Redis server performance. Netdata automatically detects Redis instances and starts collecting metrics instantly via protocols like TCP or UNIX sockets.&lt;/p></description></item><item><title>Redis monitoring checklist: the signals every production instance needs</title><link>https://www.netdata.cloud/guides/redis/redis-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-monitoring-checklist/</guid><description>&lt;h1 id="redis-monitoring-checklist-the-signals-every-production-instance-needs">Redis monitoring checklist: the signals every production instance needs&lt;/h1>
&lt;p>Redis can return PONG while replicating hours behind, during an OOM kill in a background save, or while a KEYS command wedges the event loop. This checklist structures monitoring into four maturity levels. Level 1 is the survival floor. Level 2 adds workload and resource awareness. Level 3 introduces leading indicators that catch degradation before it becomes an incident. Level 4 exposes allocator and encoding internals for granular diagnostics.&lt;/p></description></item><item><title>Redis monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/redis/redis-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-monitoring-maturity-model/</guid><description>&lt;h1 id="redis-monitoring-maturity-model-from-survival-to-expert">Redis monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Redis can return PONG while fragmentation doubles RSS, a replica falls behind, or a slow command wedges the event loop. This guide maps four cumulative operator levels derived from production runbooks. Level 1 is liveness and memory limits. Level 2 is workload anomalies and capacity pressure. Level 3 is leading indicators for composite failures. Level 4 is forensic depth. Do not skip Level 1 because you are collecting allocator statistics.&lt;/p></description></item><item><title>Redis NOAUTH / WRONGPASS authentication failures: ACL LOG and credential drift</title><link>https://www.netdata.cloud/guides/redis/redis-acl-noauth-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-acl-noauth-errors/</guid><description>&lt;h1 id="redis-noauth--wrongpass-authentication-failures-acl-log-and-credential-drift">Redis NOAUTH / WRONGPASS authentication failures: ACL LOG and credential drift&lt;/h1>
&lt;p>Application logs report &lt;code>(error) NOAUTH Authentication required&lt;/code> or &lt;code>(error) WRONGPASS invalid username-password pair or user is disabled&lt;/code>. Connections drop, transactions abort, and previously working requests are rejected. Redis does not expose an authentication failure counter in &lt;code>INFO stats&lt;/code>; on Redis 6.0 and later, &lt;code>ACL LOG&lt;/code> is the only structured source. On earlier versions, scrape the log file. Correlate the failure source, username, and timing to distinguish credential drift, misconfiguration, and brute-force probes.&lt;/p></description></item><item><title>Redis OOM command not allowed when used memory > 'maxmemory' - causes and fixes</title><link>https://www.netdata.cloud/guides/redis/redis-oom-command-not-allowed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-oom-command-not-allowed/</guid><description>&lt;h1 id="redis-oom-command-not-allowed-when-used-memory--maxmemory---causes-and-fixes">Redis OOM command not allowed when used memory &amp;gt; &amp;lsquo;maxmemory&amp;rsquo; - causes and fixes&lt;/h1>
&lt;p>Redis returns &lt;code>(error) OOM command not allowed when used memory &amp;gt; 'maxmemory'&lt;/code>. The server stays online; reads succeed, writes fail. If the client library suppresses errors, the first symptom may be missing data or backend load spikes. This occurs when &lt;code>used_memory&lt;/code> reaches &lt;code>maxmemory&lt;/code> and the eviction policy cannot free space. Under &lt;code>noeviction&lt;/code>, Redis rejects every write and keeps all keys. Under &lt;code>volatile-*&lt;/code>, the same happens when no keys carry a TTL. Monitoring often misses this because &lt;code>evicted_keys&lt;/code> stays at zero while &lt;code>used_memory&lt;/code> sits just below the limit.&lt;/p></description></item><item><title>Redis OOM-killed by the kernel: RSS, overcommit, and recovery</title><link>https://www.netdata.cloud/guides/redis/redis-out-of-memory-oom-killed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-out-of-memory-oom-killed/</guid><description>&lt;h1 id="redis-oom-killed-by-the-kernel-rss-overcommit-and-recovery">Redis OOM-killed by the kernel: RSS, overcommit, and recovery&lt;/h1>
&lt;p>Redis reports &lt;code>used_memory&lt;/code> at 60% of &lt;code>maxmemory&lt;/code>, then disappears. The container status is &lt;code>OOMKilled&lt;/code>, or &lt;code>dmesg&lt;/code> shows the kernel OOM killer selected &lt;code>redis-server&lt;/code>. The kernel enforces resident memory (RSS), while &lt;code>used_memory&lt;/code> and &lt;code>maxmemory&lt;/code> track logical allocator state. Fragmentation, copy-on-write pages during persistence, and client buffers inflate RSS above the logical figure most operators monitor. When RSS hits the host or cgroup memory ceiling, the kernel terminates the process even though Redis believes it is within limits.&lt;/p></description></item><item><title>Redis Pub/Sub pattern overhead: PSUBSCRIBE scaling and slow subscribers</title><link>https://www.netdata.cloud/guides/redis/redis-pubsub-pattern-overhead/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-pubsub-pattern-overhead/</guid><description>&lt;h1 id="redis-pubsub-pattern-overhead-psubscribe-scaling-and-slow-subscribers">Redis Pub/Sub pattern overhead: PSUBSCRIBE scaling and slow subscribers&lt;/h1>
&lt;p>You see Redis latency spikes that align with PUBLISH bursts. Subscribers disconnect with output buffer limit errors. In cluster mode, CPU and network climb linearly with node count even though PUBLISH volume stays flat. Two mechanisms drive this: PSUBSCRIBE pattern matching adds O(N) work to every PUBLISH, and slow subscribers accumulate per-client output memory until Redis cuts them off. This guide shows how to confirm the bottleneck, apply safe fixes, and decide when to move to sharded Pub/Sub.&lt;/p></description></item><item><title>Redis Queue</title><link>https://www.netdata.cloud/integrations/data-collection/databases/redis-queue/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/redis-queue/</guid><description/></item><item><title>Redis Queue Monitoring</title><link>https://www.netdata.cloud/monitoring-101/redis_queue-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/redis_queue-monitoring/</guid><description>&lt;h2 id="redis-queue-monitoring">Redis Queue Monitoring&lt;/h2>
&lt;h3 id="what-is-redis-queue">What Is Redis Queue?&lt;/h3>
&lt;p>Redis Queue (RQ) is a simple Python library for queueing jobs and processing them in the background with workers. It’s designed to have a low barrier to entry and to be easy to use. RQ is built on top of Redis, which is an in-memory data structure store that supports various data structures. It&amp;rsquo;s an ideal option for those who want a straightforward solution for task queues without the complexity of more robust tools.&lt;/p></description></item><item><title>Redis rdb_last_bgsave_status:err: diagnosing failed background saves</title><link>https://www.netdata.cloud/guides/redis/redis-rdb-last-bgsave-status-err/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-rdb-last-bgsave-status-err/</guid><description>&lt;h1 id="redis-rdb_last_bgsave_statuserr-diagnosing-failed-background-saves">Redis rdb_last_bgsave_status:err: diagnosing failed background saves&lt;/h1>
&lt;p>&lt;code>INFO persistence&lt;/code> returning &lt;code>rdb_last_bgsave_status:err&lt;/code> means the last background save failed. The flag is sticky: it remains &lt;code>err&lt;/code> until a subsequent &lt;code>BGSAVE&lt;/code> succeeds, so the failure may be hours old. If &lt;code>stop-writes-on-bgsave-error&lt;/code> is enabled (the default), Redis rejects writes and your application sees &lt;code>MISCONF&lt;/code> errors. If the setting is disabled, writes continue but durability is broken; the exposure window grows with every update.&lt;/p>
&lt;p>The failure modes are a narrow set: fork failure, disk full, filesystem write rejection, or child death before completion. Follow this sequence to separate them without restarting Redis.&lt;/p></description></item><item><title>Redis READONLY You can't write against a read only replica - causes and fixes</title><link>https://www.netdata.cloud/guides/redis/redis-readonly-cant-write-against-replica/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-readonly-cant-write-against-replica/</guid><description>&lt;h1 id="redis-readonly-you-cant-write-against-a-read-only-replica---causes-and-fixes">Redis READONLY You can&amp;rsquo;t write against a read only replica - causes and fixes&lt;/h1>
&lt;p>Your application hits &lt;code>(error) READONLY You can't write against a read only replica&lt;/code>. Writes fail; reads work. The connected Redis instance thinks it is a replica, so it rejects mutating commands. The replica is behaving correctly. The problem is a write-capable client routed to a node that is not the current primary. This typically happens in three situations: a routing bug that sends writes to a replica endpoint, stale client topology after a failover or upgrade, or an instance that was accidentally demoted at runtime.&lt;/p></description></item><item><title>Redis replication backlog overflow: full-resync storms and the 1MB default</title><link>https://www.netdata.cloud/guides/redis/redis-replication-backlog-overflow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-replication-backlog-overflow/</guid><description>&lt;h1 id="redis-replication-backlog-overflow-full-resync-storms-and-the-1mb-default">Redis replication backlog overflow: full-resync storms and the 1MB default&lt;/h1>
&lt;p>Replicas drop and reconnect, but each reconnection triggers a full resync instead of a partial sync. The primary forks for an RDB dump, latency spikes, and other replicas fall behind. Before recovery, another replica exceeds the backlog window and the cycle repeats. The default &lt;code>repl-backlog-size&lt;/code> of 1 MB triggers this cascade in most production workloads.&lt;/p>
&lt;p>The backlog is a fixed-size circular buffer of recent writes that lets a disconnected replica catch up without a full resync. When writes during a blip exceed the 1 MB default, the replica&amp;rsquo;s offset falls outside the window. Recovery requires a full resync, which forks the primary and turns a brief disconnect into a site-wide latency event.&lt;/p></description></item><item><title>Redis replication lag: detection, diagnosis, and fixes</title><link>https://www.netdata.cloud/guides/redis/redis-replication-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-replication-lag/</guid><description>&lt;h1 id="redis-replication-lag-detection-diagnosis-and-fixes">Redis replication lag: detection, diagnosis, and fixes&lt;/h1>
&lt;p>Reads from Redis replicas returning stale data indicate replication lag. Your monitoring shows a growing gap between the primary&amp;rsquo;s replication offset and what the replica has acknowledged. During failover, every byte of that gap is potential data loss.&lt;/p>
&lt;p>Replication lag in Redis is measured in bytes: the difference between &lt;code>master_repl_offset&lt;/code> on the primary and &lt;code>slave_repl_offset&lt;/code> on the replica. Small, stable lag is normal in asynchronous replication, but lag that grows continuously or exceeds &lt;code>repl-backlog-size&lt;/code> signals a bottleneck that can cascade into full resync storms.&lt;/p></description></item><item><title>Redis Sentinel triggering unnecessary failovers: quorum and split-brain</title><link>https://www.netdata.cloud/guides/redis/redis-sentinel-unnecessary-failover/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-sentinel-unnecessary-failover/</guid><description>&lt;h1 id="redis-sentinel-triggering-unnecessary-failovers-quorum-and-split-brain">Redis Sentinel triggering unnecessary failovers: quorum and split-brain&lt;/h1>
&lt;p>Your logs show a failover, but the old master never restarted. Or failovers flap: one node is promoted, then another, and clients see &lt;code>MASTERDOWN&lt;/code> and &lt;code>READONLY&lt;/code> while Sentinels disagree. The root cause is usually not the Redis data node. It is the Sentinel control plane: quorum too low, &lt;code>down-after-milliseconds&lt;/code> too aggressive, or a network partition that leaves Sentinels on the wrong side of the split.&lt;/p></description></item><item><title>Redis slowlog filling up: finding and fixing the slow commands</title><link>https://www.netdata.cloud/guides/redis/redis-slowlog-filling-up/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-slowlog-filling-up/</guid><description>&lt;h1 id="redis-slowlog-filling-up-finding-and-fixing-the-slow-commands">Redis slowlog filling up: finding and fixing the slow commands&lt;/h1>
&lt;p>Clients report timeouts and latency has a new plateau. Check &lt;code>SLOWLOG LEN&lt;/code>: if it is climbing or already at 128, the slowlog is filling. The slowlog is a circular buffer that logs commands whose execution exceeds &lt;code>slowlog-log-slower-than&lt;/code> (default 10 ms). Rapid rotation means entries evict before you inspect them, and every entry marks a blocked single-threaded event loop. Extract the culprits before the evidence disappears, distinguish execution time from queue wait, and stop the bleed.&lt;/p></description></item><item><title>Redis stale replica promotion: silent data loss at failover</title><link>https://www.netdata.cloud/guides/redis/redis-stale-replica-promotion-data-loss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-stale-replica-promotion-data-loss/</guid><description>&lt;h1 id="redis-stale-replica-promotion-silent-data-loss-at-failover">Redis stale replica promotion: silent data loss at failover&lt;/h1>
&lt;p>After a Redis failover, the new primary accepts writes immediately and clients reconnect without errors. Writes that were acknowledged by the old primary but not yet replicated are gone. This is stale replica promotion.&lt;/p>
&lt;p>Redis replication is asynchronous by default. The primary persists a write locally, replies OK to the client, then streams the change to replicas. If the primary fails before a replica receives the outstanding writes, that replica never sees them. Sentinel or Redis Cluster promotes the best available replica. If that replica is lagging, the writes in the gap are permanently lost. The client receives no error and no log warns you. The data is simply missing.&lt;/p></description></item><item><title>Redis Stream consumer group lag: pending entries and dead consumers</title><link>https://www.netdata.cloud/guides/redis/redis-stream-consumer-group-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-stream-consumer-group-lag/</guid><description>&lt;h1 id="redis-stream-consumer-group-lag-pending-entries-and-dead-consumers">Redis Stream consumer group lag: pending entries and dead consumers&lt;/h1>
&lt;p>Stream processing falls behind. An alert fires on &lt;code>used_memory&lt;/code>, or your application dashboard shows lag. You run &lt;code>XINFO GROUPS&lt;/code> and see &lt;code>lag&lt;/code> or &lt;code>pending&lt;/code> climbing. One means consumers cannot keep up. The other means they are not acknowledging messages. Both increase memory pressure, but they have different fixes. This guide shows how to tell them apart, find dead consumers, and clean up the Pending Entry List before it becomes a memory incident.&lt;/p></description></item><item><title>Redis sync_full incrementing: diagnosing full resync events</title><link>https://www.netdata.cloud/guides/redis/redis-full-resync-storms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/redis/redis-full-resync-storms/</guid><description>&lt;h1 id="redis-sync_full-incrementing-diagnosing-full-resync-events">Redis sync_full incrementing: diagnosing full resync events&lt;/h1>
&lt;p>Your Redis primary&amp;rsquo;s &lt;code>sync_full&lt;/code> counter is climbing. That means replicas are performing full resyncs instead of partial ones. Each full resync forces the primary to fork, write an RDB snapshot, and push it to the replica, which then wipes its own dataset and reloads from scratch. One full resync is a heavy operation. Several in succession, or multiple at once, can freeze the primary&amp;rsquo;s event loop, spike memory via copy-on-write, and trigger a cascade where more replicas fall behind and also need full resyncs.&lt;/p></description></item><item><title>Redline Communications Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/redline-communications-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/redline-communications-inc-snmp-traps/</guid><description/></item><item><title>Redstone Communications Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/redstone-communications-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/redstone-communications-inc-snmp-traps/</guid><description/></item><item><title>Referral</title><link>https://www.netdata.cloud/referral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/referral/</guid><description/></item><item><title>Reltec Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/reltec-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/reltec-corporation-snmp-traps/</guid><description/></item><item><title>Research In Motion Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/research-in-motion-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/research-in-motion-ltd-snmp-traps/</guid><description/></item><item><title>Resetting a wedged NVIDIA GPU: nvidia-smi --gpu-reset and when only a reboot works</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-reset-and-recovery/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-reset-and-recovery/</guid><description>&lt;h1 id="resetting-a-wedged-nvidia-gpu-nvidia-smi---gpu-reset-and-when-only-a-reboot-works">Resetting a wedged NVIDIA GPU: nvidia-smi &amp;ndash;gpu-reset and when only a reboot works&lt;/h1>
&lt;p>A GPU is wedged: jobs on it are dead or hung, nvidia-smi is slow or throwing errors on that index, and the scheduler cannot place new work on the node. The question is whether you can recover the card in place with &lt;code>nvidia-smi --gpu-reset&lt;/code> or whether you are burning time on a node that needs a reboot and possibly an RMA.&lt;/p></description></item><item><title>Response Time Monitoring</title><link>https://www.netdata.cloud/monitoring-101/response-time-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/response-time-monitoring/</guid><description>&lt;h2 id="what-is-response-time-monitoring">What is Response Time Monitoring?&lt;/h2>
&lt;p>Response time monitoring is the process of measuring the time it takes for a system or application to respond to a request made by a user. This measurement can help identify performance bottlenecks and potential issues with the system, and can be used to optimize its performance.&lt;/p>
&lt;p>Response time monitoring is a critical aspect of performance monitoring that can benefit a wide range of systems and applications including &lt;strong>web applications&lt;/strong>, such as e-commerce sites, social media platforms, enterprise web applications, &lt;strong>mobile applications&lt;/strong>, such as gaming apps, social media apps, productivity apps, &lt;strong>API-based systems&lt;/strong>, especially microservice architectures, &lt;strong>database systems&lt;/strong>, including relational and NoSQL databases.&lt;/p></description></item><item><title>RethinkDB</title><link>https://www.netdata.cloud/integrations/data-collection/databases/rethinkdb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/rethinkdb/</guid><description/></item><item><title>RethinkDB Monitoring</title><link>https://www.netdata.cloud/monitoring-101/rethinkdb-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/rethinkdb-monitoring/</guid><description>&lt;h2 id="rethinkdb-monitoring">RethinkDB Monitoring&lt;/h2>
&lt;h3 id="what-is-rethinkdb">What Is RethinkDB?&lt;/h3>
&lt;p>&lt;a href="https://rethinkdb.com">RethinkDB&lt;/a> is an open-source, distributed database built to easily store JSON documents and effortlessly scale to multiple machines. It offers a robust query language that allows developers to seamlessly deal with highly dynamic modern applications.&lt;/p>
&lt;h3 id="monitoring-rethinkdb-with-netdata">Monitoring RethinkDB With Netdata&lt;/h3>
&lt;p>Monitoring RethinkDB using Netdata offers a real-time, comprehensive insight into your database’s performance and health. Netdata’s lightweight architecture and easy-to-use interface make it an exceptional RethinkDB monitoring tool, enabling you to gain actionable alerts and historical data with minimal setup. By leveraging Netdata, you allow your operations team to focus on innovation rather than maintenance.&lt;/p></description></item><item><title>RetroShare Monitoring</title><link>https://www.netdata.cloud/monitoring-101/retroshare-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/retroshare-monitoring/</guid><description>&lt;h2 id="what-is-retroshare">What is RetroShare?&lt;/h2>
&lt;p>RetroShare is a free, open source, cross-platform software for secure file sharing, chat, and VoIP. It enables users to securely communicate and share files with friends, family, and colleagues, with an emphasis on privacy and security. RetroShare uses end-to-end encryption for all communication, and its decentralized architecture ensures there is no central server that can be compromised.&lt;/p>
&lt;h2 id="monitoring-retroshare-with-netdata">Monitoring RetroShare with Netdata&lt;/h2>
&lt;p>The prerequisites for monitoring RetroShare with Netdata are to have RetroShare and &lt;a href="https://learn.netdata.cloud/docs/cloud/get-started/">Netdata installed&lt;/a> on your system.&lt;/p></description></item><item><title>Riak KV</title><link>https://www.netdata.cloud/integrations/data-collection/databases/riak-kv/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/riak-kv/</guid><description/></item><item><title>Riak KV Monitoring</title><link>https://www.netdata.cloud/monitoring-101/riakkv-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/riakkv-monitoring/</guid><description>&lt;h2 id="riak-kv-monitoring">Riak KV Monitoring&lt;/h2>
&lt;h3 id="what-is-riak-kv">What Is Riak KV?&lt;/h3>
&lt;p>Riak KV is a distributed NoSQL database designed for high availability, scalability, and fault tolerance. It is built to handle a variety of data types and volumes, making it a popular choice for applications requiring robust data storage solutions. Learn more about Riak KV &lt;a href="https://riak.com/products/riak-kv/index.html">here&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-riak-kv-with-netdata">Monitoring Riak KV With Netdata&lt;/h3>
&lt;p>Monitoring Riak KV is crucial to ensure it performs optimally under varying load conditions. With Netdata, you gain real-time visibility into your Riak KV instances, allowing you to monitor throughput, latency, and other critical metrics. Netdata&amp;rsquo;s real-time monitoring capabilities make it an ideal Riak KV monitoring tool, giving you detailed insights into your database performance for effective troubleshooting and optimization.&lt;/p></description></item><item><title>Richard Hirschmann GmbH Co SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/richard-hirschmann-gmbh-co-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/richard-hirschmann-gmbh-co-snmp-traps/</guid><description/></item><item><title>RIPE Atlas</title><link>https://www.netdata.cloud/integrations/data-collection/networking/ripe-atlas/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/ripe-atlas/</guid><description/></item><item><title>RIPE Atlas Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ripe_atlas-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ripe_atlas-monitoring/</guid><description>&lt;h2 id="ripe-atlas-monitoring">RIPE Atlas Monitoring&lt;/h2>
&lt;h3 id="what-is-ripe-atlas">What Is RIPE Atlas?&lt;/h3>
&lt;p>RIPE Atlas is a global network measurement platform that provides real-time insights into Internet connectivity and performance across the world. With thousands of probes distributed globally, RIPE Atlas allows network administrators and IT professionals to perform measurements to assess network health, troubleshoot connectivity issues, and ensure optimal performance.&lt;/p>
&lt;h3 id="monitoring-ripe-atlas-with-netdata">Monitoring RIPE Atlas With Netdata&lt;/h3>
&lt;p>When it comes to monitoring RIPE Atlas, Netdata simplifies the process significantly. By utilizing an openmetrics (Prometheus) exporter, such as the &lt;a href="https://github.com/czerwonk/atlas_exporter">RIPE Atlas Exporter&lt;/a>, Netdata efficiently collects data from the RIPE Atlas platform. Netdata can ingest data from any Prometheus exporter, allowing you to get automated dashboards, alerts, and more without the need for a standalone Prometheus server or Grafana setup.&lt;/p></description></item><item><title>Rittal GmbH Co KG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rittal-gmbh-co-kg-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rittal-gmbh-co-kg-snmp-traps/</guid><description/></item><item><title>Riverbed</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/riverbed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/riverbed/</guid><description/></item><item><title>Riverbed Interceptor</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/riverbed-interceptor/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/riverbed-interceptor/</guid><description/></item><item><title>Riverbed Steelhead</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/riverbed-steelhead/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/riverbed-steelhead/</guid><description/></item><item><title>Riverbed Technology Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/riverbed-technology-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/riverbed-technology-inc-snmp-traps/</guid><description/></item><item><title>Riverdelta Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/riverdelta-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/riverdelta-networks-snmp-traps/</guid><description/></item><item><title>Riverstone Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/riverstone-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/riverstone-networks-snmp-traps/</guid><description/></item><item><title>Rnd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rnd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rnd-snmp-traps/</guid><description/></item><item><title>rndc not responding: control-plane failure while queries still work</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-rndc-control-channel-unresponsive/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-rndc-control-channel-unresponsive/</guid><description>&lt;h1 id="rndc-not-responding-control-plane-failure-while-queries-still-work">rndc not responding: control-plane failure while queries still work&lt;/h1>
&lt;p>&lt;code>rndc status&lt;/code> hangs. You Ctrl-C it, try again, same result. But &lt;code>dig @127.0.0.1 example.com A +short&lt;/code> returns instantly with the right answer. The data plane is healthy. The control plane is dead.&lt;/p>
&lt;p>You cannot flush caches, force zone transfers, dump state, reload configuration, or stop the daemon gracefully. If a cache-poisoning event or upstream degradation starts while the control plane is down, your primary incident-response tools are unavailable.&lt;/p></description></item><item><title>RocketChat</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/rocketchat/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/rocketchat/</guid><description/></item><item><title>RocketChat</title><link>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/rocketchat/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/rocketchat/</guid><description/></item><item><title>Rocky Linux</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/rocky-linux/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/rocky-linux/</guid><description/></item><item><title>Rogue Engineering Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rogue-engineering-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rogue-engineering-inc-snmp-traps/</guid><description/></item><item><title>Rohde Schwarz GmbH Co KG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rohde-schwarz-gmbh-co-kg-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rohde-schwarz-gmbh-co-kg-snmp-traps/</guid><description/></item><item><title>Ross Video Limited SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ross-video-limited-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ross-video-limited-snmp-traps/</guid><description/></item><item><title>RPKI invalid routes: monitoring route origin validation</title><link>https://www.netdata.cloud/guides/network/network-rpki-route-validation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-rpki-route-validation/</guid><description>&lt;h1 id="rpki-invalid-routes-monitoring-route-origin-validation">RPKI invalid routes: monitoring route origin validation&lt;/h1>
&lt;p>RPKI (Resource Public Key Infrastructure) route origin validation is a strong signal for detecting BGP hijacks and route leaks before they reroute traffic. A route classified Invalid by RPKI is, with high confidence, not originated by the authorized AS. Accepting one is a security event.&lt;/p>
&lt;p>The monitoring problem: RPKI validation state is not exposed through any standard SNMP MIB. BGP4-MIB (RFC 4273) predates RPKI entirely. CISCO-BGP4-MIB has no validation-state column. There is no portable OID to count invalid routes across a multi-vendor estate. Every platform exposes this through its own CLI, proprietary MIB extensions, or BMP (RFC 7854).&lt;/p></description></item><item><title>Rspamd</title><link>https://www.netdata.cloud/integrations/data-collection/applications/rspamd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/rspamd/</guid><description/></item><item><title>Rspamd Monitoring</title><link>https://www.netdata.cloud/monitoring-101/rspamd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/rspamd-monitoring/</guid><description>&lt;h2 id="rspamd-monitoring">Rspamd Monitoring&lt;/h2>
&lt;h3 id="what-is-rspamd">What Is Rspamd?&lt;/h3>
&lt;p>Rspamd is a fast, open-source, anti-spam system designed to protect email gateways and filter spam. It offers excellent performance, flexibility, and scalability, using a variety of advanced algorithms and statistical analysis tools.&lt;/p>
&lt;h3 id="monitoring-rspamd-with-netdata">Monitoring Rspamd With Netdata&lt;/h3>
&lt;p>Netdata&amp;rsquo;s &lt;a href="https://rspamd.com/">Rspamd monitoring tool&lt;/a> provides unparalleled insights into the performance and activity of your Rspamd instance. By using Netdata, you can monitor Rspamd&amp;rsquo;s critical metrics in real-time, gain actionable insights, and troubleshoot issues effectively.&lt;/p></description></item><item><title>Rtbrick Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rtbrick-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/rtbrick-inc-snmp-traps/</guid><description/></item><item><title>Ruby Tech Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ruby-tech-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ruby-tech-corp-snmp-traps/</guid><description/></item><item><title>Ruckus</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ruckus/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ruckus/</guid><description/></item><item><title>Ruckus Unleashed</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ruckus-unleashed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ruckus-unleashed/</guid><description/></item><item><title>Ruckus WAP</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ruckus-wap/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ruckus-wap/</guid><description/></item><item><title>Ruckus Wireless Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ruckus-wireless-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ruckus-wireless-inc-snmp-traps/</guid><description/></item><item><title>Ruggedcom Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ruggedcom-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ruggedcom-inc-snmp-traps/</guid><description/></item><item><title>Ruijie Networks Co Ltd Formerly Start Network Technology Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ruijie-networks-co-ltd-formerly-start-network-technology-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ruijie-networks-co-ltd-formerly-start-network-technology-co-ltd-snmp-traps/</guid><description/></item><item><title>S.M.A.R.T.</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/s.m.a.r.t./</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/s.m.a.r.t./</guid><description/></item><item><title>S.M.A.R.T. attributes Monitoring</title><link>https://www.netdata.cloud/monitoring-101/smartd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/smartd-monitoring/</guid><description>&lt;h2 id="what-makes-a-storage-device-smart">What makes a storage device S.M.A.R.T.?&lt;/h2>
&lt;p>&lt;a href="https://en.wikipedia.org/wiki/Self-Monitoring,_Analysis_and_Reporting_Technology">S.M.A.R.T.&lt;/a> (Self-Monitoring, Analysis, and Reporting Technology) is a supplementary component built into many modern storage devices through which devices monitor, store, and analyze the health of their operation. Statistics are collected (temperature, number of reallocated sectors, seek errors etc.) which software can use to measure the health of a device, predict possible device failure, and provide notifications on unsafe values.&lt;/p>
&lt;p>When S.M.A.R.T. data indicates a possible imminent drive failure, software running on the host system may notify the user so preventive action can be taken to prevent data loss, and the failing drive can be replaced and data integrity maintained.&lt;/p></description></item><item><title>S.M.A.R.T. Monitoring</title><link>https://www.netdata.cloud/monitoring-101/smartctl-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/smartctl-monitoring/</guid><description>&lt;h2 id="smart-monitoring">S.M.A.R.T. Monitoring&lt;/h2>
&lt;h3 id="what-is-smart">What Is S.M.A.R.T.?&lt;/h3>
&lt;p>S.M.A.R.T. (Self-Monitoring, Analysis, and Reporting Technology) is an integral system used within computers and storage devices to monitor the health and reliability of storage units. Specifically, S.M.A.R.T. helps in foreseeing potential hardware failures and enhances the ability to carry out proactive diagnostics, ultimately saving critical data from unexpected storage disasters. For more technical insight, you can check &lt;a href="https://linux.die.net/man/8/smartd">man page of smartd&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-smart-with-netdata">Monitoring S.M.A.R.T. with Netdata&lt;/h3>
&lt;p>Netdata provides a robust solution to monitor S.M.A.R.T. Enabled with the &lt;code>go.d.plugin&lt;/code> and &lt;code>smartctl&lt;/code> module, Netdata seamlessly assesses the health of your storage devices. Without directly executing potentially risky binaries, Netdata utilizes &lt;code>ndsudo&lt;/code>, a secure, privileged command execution utility that enhances operational security and smoothens permission challenges. Dive deeper by reading the &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/smartctl/">S.M.A.R.T. collector documentation&lt;/a>.&lt;/p></description></item><item><title>SABnzbd</title><link>https://www.netdata.cloud/integrations/data-collection/applications/sabnzbd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/sabnzbd/</guid><description/></item><item><title>SABnzbd Monitoring</title><link>https://www.netdata.cloud/monitoring-101/sabnzbd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/sabnzbd-monitoring/</guid><description>&lt;h2 id="sabnzbd-monitoring">SABnzbd Monitoring&lt;/h2>
&lt;h3 id="what-is-sabnzbd">What Is SABnzbd?&lt;/h3>
&lt;p>SABnzbd is a powerful and user-friendly Usenet client that automates the downloading of binary files from Usenet. It’s a preferred choice for many due to its ease of use, speed, and large number of supported devices and software. By handling NZB files seamlessly, SABnzbd helps users manage their downloads efficiently.&lt;/p>
&lt;h3 id="monitoring-sabnzbd-with-netdata">Monitoring SABnzbd With Netdata&lt;/h3>
&lt;p>To effectively monitor SABnzbd, Netdata utilizes an openmetrics (Prometheus) exporter. This approach allows seamless data collection from any Prometheus exporter, providing you with automated dashboards, alerts, and insights without the need for setting up Prometheus servers or Grafana dashboards. With Netdata, monitoring the performance and availability of your SABnzbd instance becomes effortless, ensuring optimal resource management and file downloads.&lt;/p></description></item><item><title>Saf Tehnika SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/saf-tehnika-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/saf-tehnika-snmp-traps/</guid><description/></item><item><title>Safran Trusted 4D Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/safran-trusted-4d-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/safran-trusted-4d-inc-snmp-traps/</guid><description/></item><item><title>Sagemcom Sas SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sagemcom-sas-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sagemcom-sas-snmp-traps/</guid><description/></item><item><title>Salicru EQX inverter</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/salicru-eqx-inverter/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/salicru-eqx-inverter/</guid><description/></item><item><title>Salicru EQX inverter Monitoring</title><link>https://www.netdata.cloud/monitoring-101/salicru_eqx-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/salicru_eqx-monitoring/</guid><description>&lt;h2 id="salicru-eqx-inverter-monitoring">Salicru EQX inverter Monitoring&lt;/h2>
&lt;h3 id="what-is-salicru-eqx-inverter">What Is Salicru EQX Inverter?&lt;/h3>
&lt;p>The Salicru EQX inverter is pivotal in managing solar energy, a cornerstone in modern sustainable energy solutions. These inverters are crucial for converting direct current (DC) generated by solar panels into alternating current (AC) used in homes and industries. Understanding and monitoring their performance is vital for optimizing energy efficiency and maximizing the longevity of your solar infrastructure.&lt;/p>
&lt;h3 id="monitoring-salicru-eqx-inverters-with-netdata">Monitoring Salicru EQX Inverters With Netdata&lt;/h3>
&lt;p>To monitor Salicru EQX, Netdata employs an openmetrics (Prometheus) exporter. This method allows users to gather data without the necessity of deploying a separate Prometheus server or Grafana for visualization and alerting. By utilizing &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata&lt;/a>, you get automated dashboards, real-time alerts, and comprehensive visualizations directly out-of-the-box. This efficiency makes Netdata a flexible and powerful tool for monitoring your Salicru EQX inverter.&lt;/p></description></item><item><title>Salix Technologies Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/salix-technologies-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/salix-technologies-inc-snmp-traps/</guid><description/></item><item><title>Samba</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/samba/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/samba/</guid><description/></item><item><title>Samba Monitoring</title><link>https://www.netdata.cloud/monitoring-101/samba-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/samba-monitoring/</guid><description>&lt;h2 id="samba-monitoring">Samba Monitoring&lt;/h2>
&lt;h3 id="what-is-samba">What Is Samba?&lt;/h3>
&lt;p>Samba is an open-source software suite that provides seamless file and print services to SMB/CIFS clients. It facilitates interoperability between Linux/Unix servers and Windows-based clients, allowing for file sharing and printer services. For detailed information, visit the &lt;a href="https://www.samba.org/samba/">Samba official website&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-samba-with-netdata">Monitoring Samba With Netdata&lt;/h3>
&lt;p>Netdata offers a comprehensive solution for monitoring Samba by utilizing its advanced collector: the go.d.plugin for Samba. This powerful $name monitoring tool provides real-time insights into Samba operations without the need for cumbersome configurations or extensive permissions, ensuring that you can monitor Samba instances effortlessly.&lt;/p></description></item><item><title>Samlex America Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/samlex-america-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/samlex-america-inc-snmp-traps/</guid><description/></item><item><title>Samsung Electronics Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/samsung-electronics-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/samsung-electronics-co-ltd-snmp-traps/</guid><description/></item><item><title>Sandvine Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sandvine-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sandvine-corporation-snmp-traps/</guid><description/></item><item><title>Sangoma Technologies SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sangoma-technologies-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sangoma-technologies-snmp-traps/</guid><description/></item><item><title>Sap AG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sap-ag-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sap-ag-snmp-traps/</guid><description/></item><item><title>SAS PHY error counters: transport health for SAS drives</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-sas-phy-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-sas-phy-errors/</guid><description>&lt;h1 id="sas-phy-error-counters-transport-health-for-sas-drives">SAS PHY error counters: transport health for SAS drives&lt;/h1>
&lt;p>You are running SAS drives and something is off. I/O latency has spiked on a specific drive, or the kernel log is showing SAS link resets and error recovery messages. You reach for smartctl to check SMART health, but there is no UDMA_CRC_Error_Count attribute. SAS drives do not use ATA SMART attributes. They use SCSI log pages, and their transport health lives in a different set of counters: the SAS PHY error counters.&lt;/p></description></item><item><title>SATA link downshifted (6 to 3 to 1.5 Gbps): CRC errors forcing a slower link</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-sata-link-downshift/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-sata-link-downshift/</guid><description>&lt;h1 id="sata-link-downshifted-6-to-3-to-15-gbps-crc-errors-forcing-a-slower-link">SATA link downshifted (6 to 3 to 1.5 Gbps): CRC errors forcing a slower link&lt;/h1>
&lt;p>&lt;code>smartctl -a&lt;/code> shows &lt;code>SATA Version is: SATA 3.2, 6.0 Gb/s (current: 3.0 Gb/s)&lt;/code>. The advertised maximum and the current link speed do not match. Throughput is halved, but SMART health says PASSED and there are no reallocated, pending, or uncorrectable sectors.&lt;/p>
&lt;p>This is a SATA link downshift. The kernel&amp;rsquo;s libata driver detected persistent CRC errors on the interface and automatically negotiated a lower link speed to maintain data integrity. At 3.0 Gbps the error rate drops enough for transfers to complete, but you have lost half your bandwidth. If the physical layer degrades further, the link may drop again to 1.5 Gbps.&lt;/p></description></item><item><title>SATA SSD wear-out: Wear_Leveling_Count and Media_Wearout_Indicator declining</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-ssd-wear-leveling-count/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-ssd-wear-leveling-count/</guid><description>&lt;h1 id="sata-ssd-wear-out-wear_leveling_count-and-media_wearout_indicator-declining">SATA SSD wear-out: Wear_Leveling_Count and Media_Wearout_Indicator declining&lt;/h1>
&lt;p>SATA SSDs report remaining NAND endurance through vendor-specific SMART attributes, not the standardized NVMe Percentage Used field. Samsung exposes Wear_Leveling_Count at attribute ID 177. Intel and Solidigm expose Media_Wearout_Indicator at ID 233. Some vendors use ID 173 (also labeled Wear_Leveling_Count) or ID 231 (SSD_Life_Left). All express the same concept: firmware estimating how much rated program/erase cycle budget remains.&lt;/p>
&lt;p>The problem for operators: these attributes are not standardized across vendors. The same attribute ID can carry different semantics on different drives. The THRESH column is almost always zero, so smartctl will never flag these attributes as failing regardless of how low the normalized value drops. The raw value encoding is vendor-defined and frequently misleading. A fleet with mixed SSD vendors cannot be monitored with a single threshold rule.&lt;/p></description></item><item><title>Scannex Electronics Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/scannex-electronics-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/scannex-electronics-ltd-snmp-traps/</guid><description/></item><item><title>Schechtertech LLC SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/schechtertech-llc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/schechtertech-llc-snmp-traps/</guid><description/></item><item><title>Scheduling SMART self-tests: short weekly, extended monthly</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-scheduling-self-tests/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-scheduling-self-tests/</guid><description>&lt;h1 id="scheduling-smart-self-tests-short-weekly-extended-monthly">Scheduling SMART self-tests: short weekly, extended monthly&lt;/h1>
&lt;p>Drives do not run self-tests automatically. SMART firmware monitors passively, recording what it sees during normal I/O. If a sector is never read by production traffic, its degradation stays invisible until a backup job, a scrub, or a user request hits it. At that point you get an I/O error in production instead of a warning from the drive.&lt;/p>
&lt;p>Self-tests are the active probing mechanism. The short test (1-2 minutes) exercises electrical and mechanical basics plus a small media sample. The extended test (hours, proportional to drive size) scans the entire surface and is the only routine mechanism that discovers latent bad sectors before production I/O reaches them. The conveyance test (~5 minutes, HDD only) checks for shipping damage and is typically run once at deployment, not on a schedule.&lt;/p></description></item><item><title>Schleifenbauer Products B.V. SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/schleifenbauer-products-b.v.-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/schleifenbauer-products-b.v.-snmp-traps/</guid><description/></item><item><title>Schmid Telecom AG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/schmid-telecom-ag-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/schmid-telecom-ag-snmp-traps/</guid><description/></item><item><title>Schneider Electric Apc Netbotz SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/schneider-electric-apc-netbotz-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/schneider-electric-apc-netbotz-snmp-traps/</guid><description/></item><item><title>Schneider Electric SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/schneider-electric-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/schneider-electric-snmp-traps/</guid><description/></item><item><title>Schneider Koch Co Datensysteme GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/schneider-koch-co-datensysteme-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/schneider-koch-co-datensysteme-gmbh-snmp-traps/</guid><description/></item><item><title>SCIM</title><link>https://www.netdata.cloud/integrations/authentication/scim/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/authentication/scim/</guid><description/></item><item><title>Scte SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/scte-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/scte-snmp-traps/</guid><description/></item><item><title>SCTP Statistics</title><link>https://www.netdata.cloud/integrations/data-collection/networking/sctp-statistics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/sctp-statistics/</guid><description/></item><item><title>ScyllaDB</title><link>https://www.netdata.cloud/integrations/data-collection/databases/scylladb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/scylladb/</guid><description/></item><item><title>SD-WAN tunnel up but degraded: when the control plane lies</title><link>https://www.netdata.cloud/guides/network/network-sdwan-data-plane-degraded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-sdwan-data-plane-degraded/</guid><description>&lt;h1 id="sd-wan-tunnel-up-but-degraded-when-the-control-plane-lies">SD-WAN tunnel up but degraded: when the control plane lies&lt;/h1>
&lt;p>The orchestrator shows your SD-WAN tunnel as UP. Control connections to vSmart or vBond are healthy. OMP sessions are Established. But users at the far end report slow applications, dropped voice calls, or timeouts.&lt;/p>
&lt;p>The control plane reports a healthy tunnel while the data plane is degraded with packet loss, latency spikes, or silent traffic drops. Interface counters show UP/UP because the degradation is on the underlay path or inside the encapsulated data plane, not on the local interface.&lt;/p></description></item><item><title>Seagate Technology SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/seagate-technology-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/seagate-technology-snmp-traps/</guid><description/></item><item><title>Securitymatrix Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/securitymatrix-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/securitymatrix-inc-snmp-traps/</guid><description/></item><item><title>Seek_Error_Rate: head-positioning wear and the Seagate raw-value trap</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-seek-error-rate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-seek-error-rate/</guid><description>&lt;h1 id="seek_error_rate-head-positioning-wear-and-the-seagate-raw-value-trap">Seek_Error_Rate: head-positioning wear and the Seagate raw-value trap&lt;/h1>
&lt;p>A monitoring dashboard shows Seek_Error_Rate with a raw value of 200,009,354,607 on a Seagate drive. The on-call engineer pages the storage team. The replacement drive goes into the same bay and shows the same number. This cycle repeats across fleets because most monitoring tools and most operators do not know how Seagate encodes this attribute.&lt;/p>
&lt;p>Seek_Error_Rate (ATA attribute ID 7) tracks the accuracy of the HDD actuator arm as it positions read/write heads over target tracks. On most non-Seagate drives, the raw value is a straightforward error count or rate. On Seagate drives, the raw value is a packed composite encoding both total seek operations and seek errors, producing numbers in the billions on perfectly healthy hardware. Alerting on the raw value guarantees false positives on every Seagate drive in the fleet.&lt;/p></description></item><item><title>Seh Computertechnik GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/seh-computertechnik-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/seh-computertechnik-gmbh-snmp-traps/</guid><description/></item><item><title>Self-test completed: read failure - a bad sector found by proactive scanning</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-self-test-read-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-self-test-read-failure/</guid><description>&lt;h1 id="self-test-completed-read-failure---a-bad-sector-found-by-proactive-scanning">Self-test completed: read failure - a bad sector found by proactive scanning&lt;/h1>
&lt;p>When &lt;code>smartctl -l selftest /dev/sdX&lt;/code> shows &amp;ldquo;Completed: read failure&amp;rdquo;, the drive&amp;rsquo;s firmware found a sector it cannot read during an active surface scan. The LBA_of_first_error column gives you the exact logical block address of the defect. The firmware tried multiple times, exhausted its error correction, and could not recover the data at that location.&lt;/p>
&lt;p>The extended self-test scans the entire media surface. Without it, a bad sector remains hidden until production I/O hits that exact LBA, producing an application-visible I/O error, a hung process, or a kernel timeout instead of a diagnostic warning.&lt;/p></description></item><item><title>Semaphore statistics</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/semaphore-statistics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/semaphore-statistics/</guid><description/></item><item><title>Senao International Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/senao-international-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/senao-international-co-ltd-snmp-traps/</guid><description/></item><item><title>Sense Energy</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/sense-energy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/sense-energy/</guid><description/></item><item><title>Sense Energy Monitoring</title><link>https://www.netdata.cloud/monitoring-101/sense_energy-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/sense_energy-monitoring/</guid><description>&lt;h2 id="sense-energy-monitoring">Sense Energy Monitoring&lt;/h2>
&lt;h3 id="what-is-sense-energy">What Is Sense Energy?&lt;/h3>
&lt;p>Sense Energy is an advanced, smart home energy monitoring system that helps homeowners understand their electricity usage in real-time. With the integration of devices and IoT technologies, Sense Energy provides detailed insights and data that empower users to optimize energy consumption and reduce utility costs.&lt;/p>
&lt;h3 id="monitoring-sense-energy-with-netdata">Monitoring Sense Energy With Netdata&lt;/h3>
&lt;p>When it comes to efficiently monitor Sense Energy, the Netdata monitoring tool stands out as an ideal choice. Netdata utilizes an openmetrics (Prometheus) exporter to collect data. This means that with Netdata, you can ingest data from any Prometheus exporter, allowing you to gain comprehensive insights without the need for a Prometheus server or Grafana setup. Netdata offers automated dashboards, real-time alerts, and health maps that enable proactive energy management. For more technical details, you can check out the &lt;a href="https://learn.netdata.cloud/docs/collecting-metrics/generic-collecting-metrics/prometheus-endpoint/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata documentation&lt;/a> and explore the &lt;a href="https://github.com/ejsuncy/sense_energy_prometheus_exporter">community exporter&lt;/a> available for Sense Energy.&lt;/p></description></item><item><title>Sensoria Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sensoria-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sensoria-corporation-snmp-traps/</guid><description/></item><item><title>Sensors</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/sensors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/sensors/</guid><description/></item><item><title>Sensu Enterprise SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sensu-enterprise-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sensu-enterprise-snmp-traps/</guid><description/></item><item><title>Server Iron Switch</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/server-iron-switch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/server-iron-switch/</guid><description/></item><item><title>Server Technology Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/server-technology-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/server-technology-inc-snmp-traps/</guid><description/></item><item><title>Servertech</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/servertech/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/servertech/</guid><description/></item><item><title>Servertech Pdu3</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/servertech-pdu3/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/servertech-pdu3/</guid><description/></item><item><title>Servertech Pdu4</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/servertech-pdu4/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/servertech-pdu4/</guid><description/></item><item><title>ServiceNow</title><link>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/servicenow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/servicenow/</guid><description/></item><item><title>sFlow</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/flow-protocols/sflow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/flow-protocols/sflow/</guid><description/></item><item><title>sFlow sampling rate: why your traffic totals are off by 1000x</title><link>https://www.netdata.cloud/guides/network/network-sflow-sampling-rate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-sflow-sampling-rate/</guid><description>&lt;h1 id="sflow-sampling-rate-why-your-traffic-totals-are-off-by-1000x">sFlow sampling rate: why your traffic totals are off by 1000x&lt;/h1>
&lt;p>Your sFlow-derived bandwidth charts show a 10G link carrying 12 Mbps. SNMP counters on the same interface show 8.4 Gbps. The switch is not broken and the collector is not dropping packets. The analytics pipeline is summing raw sampled bytes without multiplying by the sampling rate.&lt;/p>
&lt;p>sFlow is not NetFlow. It does not maintain a flow cache on the device, aggregate bytes per conversation, and export summary totals. sFlow exports individual packet samples, one per datagram, each carrying the packet&amp;rsquo;s header data and metadata about the sampling process. The collector is responsible for turning those samples into traffic estimates through multiplication. When that multiplication is missing, every chart, alert, capacity plan, and billing report built on the data is wrong by the sampling factor.&lt;/p></description></item><item><title>Sgte Ies SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sgte-ies-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sgte-ies-snmp-traps/</guid><description/></item><item><title>Shanghai Baud Data Communication Development Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/shanghai-baud-data-communication-development-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/shanghai-baud-data-communication-development-corp-snmp-traps/</guid><description/></item><item><title>Shanghai Meridian Technologies Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/shanghai-meridian-technologies-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/shanghai-meridian-technologies-co-ltd-snmp-traps/</guid><description/></item><item><title>Shasta Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/shasta-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/shasta-networks-snmp-traps/</guid><description/></item><item><title>Shelly humidity sensor</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/shelly-humidity-sensor/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/shelly-humidity-sensor/</guid><description/></item><item><title>Shelly Humidity Sensor Monitoring</title><link>https://www.netdata.cloud/monitoring-101/shelly-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/shelly-monitoring/</guid><description>&lt;h2 id="shelly-humidity-sensor-monitoring">Shelly Humidity Sensor Monitoring&lt;/h2>
&lt;h3 id="what-is-shelly-humidity-sensor">What Is Shelly Humidity Sensor?&lt;/h3>
&lt;p>A Shelly Humidity Sensor is a smart home device that provides precise humidity readings essential for maintaining comfort and optimal conditions in an indoor environment. It is an integral part of modern IoT setups, offering insights into the air quality and automating climate control systems for improved home automation.&lt;/p>
&lt;h3 id="monitoring-shelly-humidity-sensor-with-netdata">Monitoring Shelly Humidity Sensor With Netdata&lt;/h3>
&lt;p>To monitor the Shelly Humidity Sensor effectively, Netdata utilizes an openmetrics (Prometheus) exporter. The &lt;a href="https://github.com/aexel90/shelly_exporter">Shelly Exporter&lt;/a> collects metrics by periodically sending HTTP requests to the humidity sensor. Netdata can seamlessly ingest data from any Prometheus exporter, providing automated dashboards, alerting mechanisms, and more—all without the need for a Prometheus server or Grafana dashboard.&lt;/p></description></item><item><title>Shenzhen C Data Technology Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/shenzhen-c-data-technology-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/shenzhen-c-data-technology-co-ltd-snmp-traps/</guid><description/></item><item><title>Shenzhen First Mile Communications Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/shenzhen-first-mile-communications-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/shenzhen-first-mile-communications-ltd-snmp-traps/</guid><description/></item><item><title>Shenzhen Smartbyte Technology Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/shenzhen-smartbyte-technology-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/shenzhen-smartbyte-technology-co-ltd-snmp-traps/</guid><description/></item><item><title>Shiva Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/shiva-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/shiva-corporation-snmp-traps/</guid><description/></item><item><title>Siae Microelettronica S P A SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/siae-microelettronica-s-p-a-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/siae-microelettronica-s-p-a-snmp-traps/</guid><description/></item><item><title>Siemens AG Automation Drives SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/siemens-ag-automation-drives-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/siemens-ag-automation-drives-snmp-traps/</guid><description/></item><item><title>Siemens AG SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/siemens-ag-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/siemens-ag-snmp-traps/</guid><description/></item><item><title>Siemens S7 PLC</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/siemens-s7-plc/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/siemens-s7-plc/</guid><description/></item><item><title>Siemens S7 PLC Monitoring</title><link>https://www.netdata.cloud/monitoring-101/s7_plc-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/s7_plc-monitoring/</guid><description>&lt;h2 id="siemens-s7-plc-monitoring">Siemens S7 PLC Monitoring&lt;/h2>
&lt;h3 id="what-is-siemens-s7-plc">What Is Siemens S7 PLC?&lt;/h3>
&lt;p>Siemens S7 Programmable Logic Controllers (PLCs) are used extensively in industrial environments for automation and control. These devices manage operations ranging from simple on-off control to complex continual processes, ensuring efficiency and reliability in factory automation and other industrial systems.&lt;/p>
&lt;h3 id="monitoring-siemens-s7-plc-with-netdata">Monitoring Siemens S7 PLC With Netdata&lt;/h3>
&lt;p>To monitor Siemens S7 PLC with Netdata, an openmetrics (Prometheus) exporter is used. Netdata&amp;rsquo;s $name monitoring tool can seamlessly ingest data from any Prometheus exporter, providing you with automated dashboards, customizable alerts, and comprehensive insights—all without the need for a dedicated Prometheus server or Grafana setup. This capability makes Netdata a versatile tool for monitoring Siemens S7 PLC, ensuring you have real-time visibility and control over your industrial operations. To explore how this works, &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">check out the Live Demo&lt;/a>.&lt;/p></description></item><item><title>Sigma Network Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sigma-network-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sigma-network-systems-inc-snmp-traps/</guid><description/></item><item><title>SIGNL4</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/signl4/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/signl4/</guid><description/></item><item><title>Sigur SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sigur-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sigur-snmp-traps/</guid><description/></item><item><title>Siklu Communication Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/siklu-communication-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/siklu-communication-ltd-snmp-traps/</guid><description/></item><item><title>Silent UDP flow data loss: why your NetFlow collector is dropping records</title><link>https://www.netdata.cloud/guides/network/network-netflow-udp-flow-loss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-netflow-udp-flow-loss/</guid><description>&lt;h1 id="silent-udp-flow-data-loss-why-your-netflow-collector-is-dropping-records">Silent UDP flow data loss: why your NetFlow collector is dropping records&lt;/h1>
&lt;p>Your flow analytics show traffic declining on multiple exporters simultaneously. SNMP interface counters say traffic is rising. No device alarms, no exporter config changes, no visible network events. The most likely cause: your collector is silently dropping UDP datagrams at the kernel socket buffer boundary.&lt;/p>
&lt;p>UDP has no delivery guarantee. When the socket receive buffer fills, the kernel discards incoming datagrams silently. No error is logged. The only signal is &lt;code>UdpRcvbufErrors&lt;/code> in &lt;code>/proc/net/snmp&lt;/code>, a counter most teams do not monitor. During a traffic spike or DDoS, your charts may show &amp;ldquo;normal&amp;rdquo; or declining traffic while actual packet rates are significantly higher.&lt;/p></description></item><item><title>Silver Peak Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/silver-peak-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/silver-peak-systems-inc-snmp-traps/</guid><description/></item><item><title>Silverpeak Edgeconnect</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/silverpeak-edgeconnect/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/silverpeak-edgeconnect/</guid><description/></item><item><title>Sinclair Internetworking Services SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sinclair-internetworking-services-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sinclair-internetworking-services-snmp-traps/</guid><description/></item><item><title>Sinetica Eagle I</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/sinetica-eagle-i/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/sinetica-eagle-i/</guid><description/></item><item><title>Sinetica SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sinetica-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sinetica-snmp-traps/</guid><description/></item><item><title>Sita Ads SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sita-ads-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sita-ads-snmp-traps/</guid><description/></item><item><title>Site 24x7</title><link>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/site-24x7/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/site-24x7/</guid><description/></item><item><title>Site 24x7 Monitoring</title><link>https://www.netdata.cloud/monitoring-101/site24x7-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/site24x7-monitoring/</guid><description>&lt;h2 id="site-24x7-monitoring">Site 24x7 Monitoring&lt;/h2>
&lt;h3 id="what-is-site-24x7">What Is Site 24x7?&lt;/h3>
&lt;p>Site 24x7 is a comprehensive solution that focuses on holistic monitoring of websites, servers, applications, and networks. It provides an array of insights into performance, uptime, and user interactions, thus helping technical teams to ensure robust infrastructure consistency and performance.&lt;/p>
&lt;h3 id="monitoring-site-24x7-with-netdata">Monitoring Site 24x7 With Netdata&lt;/h3>
&lt;p>To effectively monitor Site 24x7, Netdata utilizes an openmetrics (Prometheus) exporter. This allows Netdata to ingest data from any Prometheus exporter, providing automated dashboards, alerts, and more without necessitating a Prometheus server or Grafana. By using the &lt;a href="https://github.com/svenstaro/site24x7_exporter">Site 24x7 exporter&lt;/a>, teams can effortlessly capture key monitoring metrics, ensuring seamless performance and diagnostic interventions. With Netdata, monitoring Site 24x7 transforms into a streamlined process with precise and real-time analytics that boost operational efficiency.&lt;/p></description></item><item><title>Skyhigh Security LLC SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/skyhigh-security-llc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/skyhigh-security-llc-snmp-traps/</guid><description/></item><item><title>Slack</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/slack/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/slack/</guid><description/></item><item><title>Slack</title><link>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/slack/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/slack/</guid><description/></item><item><title>Slurm</title><link>https://www.netdata.cloud/integrations/data-collection/applications/slurm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/slurm/</guid><description/></item><item><title>Slurm Monitoring</title><link>https://www.netdata.cloud/monitoring-101/slurm-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/slurm-monitoring/</guid><description>&lt;h2 id="slurm-monitoring">Slurm Monitoring&lt;/h2>
&lt;h3 id="what-is-slurm">What Is Slurm?&lt;/h3>
&lt;p>Slurm, also known as the Simple Linux Utility for Resource Management, is an open-source workload management system that is specifically tailored for high-performance computing (HPC) and cluster environments. It efficiently allocates resources such as CPU and memory to various jobs, ensuring optimal use of available resources across clustered nodes.&lt;/p>
&lt;h3 id="monitoring-slurm-with-netdata">Monitoring Slurm With Netdata&lt;/h3>
&lt;p>To effectively monitor Slurm, Netdata utilizes an openmetrics (Prometheus) exporter called the &lt;a href="https://github.com/vpenso/prometheus-slurm-exporter">Prometheus Slurm Exporter&lt;/a>. With Netdata, you can ingest data from any Prometheus exporter, streamlining the process by providing automated dashboards, real-time alerts, and comprehensive insights without the need for setting up a standalone Prometheus server or configuring Grafana.&lt;/p></description></item><item><title>SMA Inverters</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/sma-inverters/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/sma-inverters/</guid><description/></item><item><title>SMA Inverters Monitoring</title><link>https://www.netdata.cloud/monitoring-101/sma_inverter-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/sma_inverter-monitoring/</guid><description>&lt;h2 id="sma-inverters-monitoring">SMA Inverters Monitoring&lt;/h2>
&lt;h3 id="what-is-sma-inverters">What Is SMA Inverters?&lt;/h3>
&lt;p>SMA Inverters are a crucial component in solar energy systems, converting the variable direct current (DC) output of a solar panel into alternating current (AC) that can be fed into the electrical grid or used by a local, off-grid network. Ensuring their optimal performance is vital for the efficient management of solar energy resources.&lt;/p>
&lt;h3 id="monitoring-sma-inverters-with-netdata">Monitoring SMA Inverters With Netdata&lt;/h3>
&lt;p>Monitor SMA Inverters seamlessly using Netdata&amp;rsquo;s powerful monitoring tool. Netdata leverages an OpenMetrics (Prometheus) exporter to gather solar inverter metrics. It can ingest data from any Prometheus exporter, providing users with automated dashboards, alerts, and more—all without the need for a separate Prometheus server or Grafana instance. For a versatile monitoring solution, &lt;a href="https://github.com/dr0ps/sma_inverter_exporter">check out the community exporter&lt;/a> used to monitor SMA Inverters, which integrates seamlessly with Netdata.&lt;/p></description></item><item><title>SMART blind spots: VMs, USB bridges, and drives you think you're watching</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-blind-spot-vm-usb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-blind-spot-vm-usb/</guid><description>&lt;h1 id="smart-blind-spots-vms-usb-bridges-and-drives-you-think-youre-watching">SMART blind spots: VMs, USB bridges, and drives you think you&amp;rsquo;re watching&lt;/h1>
&lt;p>SMART monitoring is deployed, the dashboard is green, but some drives are invisible to smartctl. No alert fired. The pattern: expected physical drive count exceeds drives returning SMART data. The drives may be healthy or failing. You cannot tell because no telemetry is collected, and the monitoring system treats &amp;ldquo;smartctl returned no data&amp;rdquo; as &amp;ldquo;no problem.&amp;rdquo;&lt;/p>
&lt;p>Four configurations create this gap: virtual machines, USB-attached drives, cloud block devices, and hardware RAID controllers. In each case, smartctl cannot reach the physical drive&amp;rsquo;s firmware. The fix is to move monitoring to where SMART data is accessible and use the correct device-type flags.&lt;/p></description></item><item><title>Smart meters SML</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/smart-meters-sml/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/smart-meters-sml/</guid><description/></item><item><title>Smart meters SML Monitoring</title><link>https://www.netdata.cloud/monitoring-101/sml-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/sml-monitoring/</guid><description>&lt;h2 id="smart-meters-sml-monitoring">Smart meters SML Monitoring&lt;/h2>
&lt;h3 id="what-is-smart-meters-sml">What Is Smart meters SML?&lt;/h3>
&lt;p>Smart meters SML (Smart Message Language) is a protocol designed for efficient smart metering and energy management. It plays a crucial role in the Internet of Things (IoT) ecosystem by transmitting detailed data regarding energy consumption and other utilities.&lt;/p>
&lt;h3 id="monitoring-smart-meters-sml-with-netdata">Monitoring Smart meters SML With Netdata&lt;/h3>
&lt;p>To effectively monitor Smart meters SML, Netdata leverages an openmetrics (prometheus) exporter. This integration allows Netdata to ingest data from any Prometheus exporter, providing fully automated dashboards, real-time alerts, and granular insights without needing a Prometheus server or Grafana. This seamless monitoring capability makes Netdata an ideal tool for the continuous surveillance of smart metering systems. For a hands-on experience, check out our &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">Live Demo&lt;/a>.&lt;/p></description></item><item><title>SMART monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-monitoring-maturity-model/</guid><description>&lt;h1 id="smart-monitoring-maturity-model-from-survival-to-expert">SMART monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>&lt;code>smartctl -H&lt;/code> returns PASSED or FAILED. Teams that stop there learn the hard way: a drive reports PASSED one day and drops off the bus the next, a &amp;ldquo;healthy&amp;rdquo; drive starts corrupting data, or a RAID rebuild fails because nobody noticed the spare sector pool degrading for months.&lt;/p>
&lt;p>SMART monitoring is a spectrum of signal depth. Each level in this model closes specific diagnostic blind spots that the previous level could not answer. Use this to identify what failure modes your current monitoring is blind to and what to add next.&lt;/p></description></item><item><title>SMART not accessible behind a RAID controller: -d megaraid and cciss passthrough</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-not-accessible-raid/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-not-accessible-raid/</guid><description>&lt;h1 id="smart-not-accessible-behind-a-raid-controller--d-megaraid-and-cciss-passthrough">SMART not accessible behind a RAID controller: -d megaraid and cciss passthrough&lt;/h1>
&lt;p>When you run &lt;code>smartctl -a /dev/sda&lt;/code> on a server with a hardware RAID controller, the query returns information about the controller&amp;rsquo;s virtual device, not the physical drive. Or it fails entirely. The RAID controller presents only virtual devices to the OS, so standard SMART queries never reach the physical hardware.&lt;/p>
&lt;p>Without the correct controller-specific passthrough flag, you have no visibility into individual drive health. Monitoring may report success because it queried a device node and got a response, but the response contains no real drive data. Drives can fail silently while monitoring appears healthy.&lt;/p></description></item><item><title>SMART overall-health self-assessment: FAILED is the drive's own death notice</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-overall-health-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-overall-health-failed/</guid><description>&lt;h1 id="smart-overall-health-self-assessment-failed-is-the-drives-own-death-notice">SMART overall-health self-assessment: FAILED is the drive&amp;rsquo;s own death notice&lt;/h1>
&lt;p>When &lt;code>smartctl -H&lt;/code> reports &lt;code>SMART overall-health self-assessment test result: FAILED&lt;/code>, the drive&amp;rsquo;s own firmware has concluded it is failing. At least one pre-fail attribute has crossed its vendor-defined threshold. This is not smartctl&amp;rsquo;s interpretation or a monitoring heuristic. Treat it as an unconditional page. Unlike individual attributes that require trend analysis and corroboration, FAILED is the drive saying it is done. False positives are rare in the normal case because the drive itself is making the call, not your monitoring system.&lt;/p></description></item><item><title>SMART says PASSED but the drive is failing: why the health check lies</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-health-passed-but-failing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-health-passed-but-failing/</guid><description>&lt;h1 id="smart-says-passed-but-the-drive-is-failing-why-the-health-check-lies">SMART says PASSED but the drive is failing: why the health check lies&lt;/h1>
&lt;p>The PASSED verdict from &lt;code>smartctl -H&lt;/code> tells you one narrow thing: no pre-fail SMART attribute has crossed its vendor-defined threshold. It does not mean the drive is healthy, that data is safe, or that the drive will survive the week.&lt;/p>
&lt;p>A drive can report PASSED while hundreds of sectors are pending reallocation, dozens have already been remapped, I/O latency is spiking to seconds, and the spare pool is burning down. Google&amp;rsquo;s 2007 study of over 100,000 drives found that 36% of failed drives had zero prior SMART warnings. Backblaze&amp;rsquo;s fleet analysis showed that 23.3% of failed drives had no non-zero values across the five attributes they consider most predictive (IDs 5, 187, 188, 197, 198). The health check is a last-resort binary, not a health indicator.&lt;/p></description></item><item><title>SMART self-tests keep aborting: heavy I/O interrupting the extended scan</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-self-test-aborted/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-self-test-aborted/</guid><description>&lt;h1 id="smart-self-tests-keep-aborting-heavy-io-interrupting-the-extended-scan">SMART self-tests keep aborting: heavy I/O interrupting the extended scan&lt;/h1>
&lt;p>The self-test log tells the same story every time. You started an extended self-test with &lt;code>smartctl -t long /dev/sdX&lt;/code>, checked back hours later, and found another entry reading &amp;ldquo;Aborted by host&amp;rdquo; or &amp;ldquo;Interrupted (host reset)&amp;rdquo;. The drive reports no read failures, no servo errors, no electrical faults. But the full surface scan never actually completed.&lt;/p>
&lt;p>This is a scheduling problem, not a drive health problem. ATA extended self-tests run in background mode by default, meaning the test has low priority and yields to host I/O. On a busy production drive, the test keeps getting deferred and eventually aborted. Latent bad sectors remain undiscovered until production I/O hits them.&lt;/p></description></item><item><title>smartctl disk monitoring checklist: the SMART signals every server needs</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-monitoring-checklist/</guid><description>&lt;h1 id="smartctl-disk-monitoring-checklist-the-smart-signals-every-server-needs">smartctl disk monitoring checklist: the SMART signals every server needs&lt;/h1>
&lt;p>S.M.A.R.T. is firmware-level instrumentation built into every modern HDD, SSD, and NVMe drive. The drive reports on its own internal state. The &lt;code>smartctl&lt;/code> tool from smartmontools reads what the firmware already knows. SMART monitoring is necessary but not sufficient: it catches gradual media degradation and endurance wear-out, but cannot predict sudden controller failures, firmware bugs, or silent data corruption. You still need redundancy, backups, and checksumming filesystems.&lt;/p></description></item><item><title>Smartoptics AS SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/smartoptics-as-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/smartoptics-as-snmp-traps/</guid><description/></item><item><title>SMB Server Shares</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/smb-server-shares/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/smb-server-shares/</guid><description/></item><item><title>Smc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/smc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/smc-snmp-traps/</guid><description/></item><item><title>SMS</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/sms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/sms/</guid><description/></item><item><title>SMSEagle</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/smseagle/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/smseagle/</guid><description/></item><item><title>Snarlsnmp Dynamic Web Application Monitor Developers Group SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/snarlsnmp-dynamic-web-application-monitor-developers-group-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/snarlsnmp-dynamic-web-application-monitor-developers-group-snmp-traps/</guid><description/></item><item><title>SNMP</title><link>https://www.netdata.cloud/integrations/all/snmp/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/all/snmp/</guid><description/></item><item><title>SNMP authentication-failure spikes: misconfiguration vs reconnaissance</title><link>https://www.netdata.cloud/guides/network/network-snmp-auth-failure-spikes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-snmp-auth-failure-spikes/</guid><description>&lt;h1 id="snmp-authentication-failure-spikes-misconfiguration-vs-reconnaissance">SNMP authentication-failure spikes: misconfiguration vs reconnaissance&lt;/h1>
&lt;p>SNMP authentication-failure traps are one of the few security signals built into the network monitoring stack. When they spike, the question is never &amp;ldquo;is something wrong?&amp;rdquo; - it is &amp;ldquo;is this a broken poller or someone probing my devices?&amp;rdquo; The answer changes the response from a quiet config fix to a security incident.&lt;/p>
&lt;p>The authenticationFailure trap (OID &lt;code>1.3.6.1.6.3.1.1.5.5&lt;/code>) fires whenever an SNMP agent receives a protocol message that is not properly authenticated. On SNMPv2c, that means a wrong community string. On SNMPv3, it means a wrong username, wrong auth protocol, wrong auth password, or wrong privacy password. The trap is defined in SNMPv2-MIB and every compliant agent can generate it, but many vendors ship with it disabled by default. If you have never explicitly enabled it (for example, &lt;code>snmp-server enable traps snmp authentication&lt;/code> on Cisco IOS), you may have no signal at all.&lt;/p></description></item><item><title>SNMP counter discontinuity after reboot: bogus rate spikes explained</title><link>https://www.netdata.cloud/guides/network/network-snmp-counter-discontinuity/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-snmp-counter-discontinuity/</guid><description>&lt;h1 id="snmp-counter-discontinuity-after-reboot-bogus-rate-spikes-explained">SNMP counter discontinuity after reboot: bogus rate spikes explained&lt;/h1>
&lt;p>A 1-gigabit interface shows 40 terabits per second on your dashboard right after a switch reboot. The traffic never happened. The chart is lying because of how SNMP counters and rate calculations interact.&lt;/p>
&lt;p>SNMP interface counters (ifInOctets, ifHCInOctets, ifOutOctets, and friends) are monotonically increasing integers. Your monitoring platform does not read current bandwidth from the device. It subtracts the previous counter value from the current one, divides by elapsed time, and reports the result as a rate. When a counter resets to zero after a reboot or wraps past its maximum, that subtraction produces a physically impossible number. If your alerting or billing pipeline acts on it, you have a problem.&lt;/p></description></item><item><title>SNMP counter rollover: fake traffic spikes from 32-bit counters</title><link>https://www.netdata.cloud/guides/network/network-snmp-counter-rollover/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-snmp-counter-rollover/</guid><description>&lt;h1 id="snmp-counter-rollover-fake-traffic-spikes-from-32-bit-counters">SNMP counter rollover: fake traffic spikes from 32-bit counters&lt;/h1>
&lt;p>A bandwidth chart suddenly shows a multi-terabit spike on a 10G interface. The on-call engineer investigates and finds the link was nearly idle. The spike is a math artifact: a 32-bit SNMP counter wrapped from near its maximum value (4,294,967,295) back to zero between two polls, and the collector&amp;rsquo;s differencing algorithm produced a nonsensical delta.&lt;/p>
&lt;p>Depending on how the collector handles the wrap, the symptom differs: a fake spike (when the negative delta is treated as unsigned) or a fake traffic drop to zero (when treated as signed and clamped). Both hide real traffic patterns and train operators to ignore chart anomalies, including genuine ones.&lt;/p></description></item><item><title>SNMP devices</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/snmp-devices/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/snmp-devices/</guid><description/></item><item><title>SNMP Monitoring</title><link>https://www.netdata.cloud/monitoring-101/snmp-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/snmp-monitoring/</guid><description>&lt;h2 id="snmp-monitoring">SNMP Monitoring&lt;/h2>
&lt;h3 id="what-is-snmp">What Is SNMP?&lt;/h3>
&lt;p>SNMP (Simple Network Management Protocol) is a protocol used for managing devices on IP networks, such as routers, switches, servers, and workstations. By utilizing an SNMP monitoring tool like Netdata, network administrators are able to collect and organize information about devices in real-time to ensure efficient functioning and to identify issues before they escalate.&lt;/p>
&lt;h3 id="monitoring-snmp-with-netdata">Monitoring SNMP With Netdata&lt;/h3>
&lt;p>Netdata allows you to seamlessly monitor SNMP-enabled devices by leveraging its SNMP collector module. Whether you&amp;rsquo;re looking to track network interface traffic, errors, or overall uptime, Netdata provides tools for monitoring SNMP with minute precision. For those interested in configuring the monitoring setup for their SNMP devices, check out the &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/snmp/">SNMP Collector Documentation&lt;/a>.&lt;/p></description></item><item><title>SNMP poll response latency: diagnosing a slow poller</title><link>https://www.netdata.cloud/guides/network/network-snmp-poll-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-snmp-poll-latency/</guid><description>&lt;h1 id="snmp-poll-response-latency-diagnosing-a-slow-poller">SNMP poll response latency: diagnosing a slow poller&lt;/h1>
&lt;p>SNMP poll response latency is the round-trip time from your collector&amp;rsquo;s GET or GETBULK request to the device&amp;rsquo;s response. When it climbs, rate calculations lose accuracy, worker threads hold their slots longer than expected, and the poller falls behind schedule. Healthy devices start appearing stale or unreachable.&lt;/p>
&lt;p>The most common misdiagnosis is &amp;ldquo;the network is slow.&amp;rdquo; On a LAN, an SNMP GET to &lt;code>sysUpTime&lt;/code> should return in single-digit milliseconds. When the same device takes 2 to 5 seconds to respond, ICMP to the same target will usually confirm the path is fine. The bottleneck is almost always the device&amp;rsquo;s SNMP agent, the collector&amp;rsquo;s scheduler design, or a specific OID family that triggers expensive computation on the device CPU.&lt;/p></description></item><item><title>SNMP poller falling behind: the polling-storm cascade and how to catch it</title><link>https://www.netdata.cloud/guides/network/network-snmp-polling-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-snmp-polling-storm/</guid><description>&lt;h1 id="snmp-poller-falling-behind-the-polling-storm-cascade-and-how-to-catch-it">SNMP poller falling behind: the polling-storm cascade and how to catch it&lt;/h1>
&lt;p>A single slow device is all it takes. Your SNMP poller queue drifts past its scheduled interval. Within minutes, 30 devices show as DOWN in your NMS dashboard. Every one of them responds to ping. The network is fine; your poller is the problem.&lt;/p>
&lt;p>Scheduler fall-behind is the most common false &amp;ldquo;device down&amp;rdquo; trigger in network monitoring. When a poller cannot complete its collection cycle within the configured interval, every subsequent cycle inherits the debt. Devices that are reachable and healthy appear DOWN because their next poll slot arrives late relative to the alerting threshold. The cascade is self-reinforcing: missed polls generate retries, retries consume worker threads, fewer workers means slower polls for all other devices, more devices time out, and queue depth grows unboundedly.&lt;/p></description></item><item><title>SNMP timeouts and retries: why devices show as down when they aren't</title><link>https://www.netdata.cloud/guides/network/network-snmp-timeouts-retries/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-snmp-timeouts-retries/</guid><description>&lt;h1 id="snmp-timeouts-and-retries-why-devices-show-as-down-when-they-arent">SNMP timeouts and retries: why devices show as down when they aren&amp;rsquo;t&lt;/h1>
&lt;p>SNMP runs over UDP port 161, a transport with no delivery guarantee. When your monitoring platform reports that devices are down, the first question is not &amp;ldquo;why is the network broken&amp;rdquo; but &amp;ldquo;is this actually a network problem, or is my polling stack the problem.&amp;rdquo; SNMP timeout and retry behavior is one of the most common causes of false-positive &amp;ldquo;device down&amp;rdquo; alerts, and it is also one of the most misdiagnosed.&lt;/p></description></item><item><title>SNMP trap listener</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/snmp-trap-listener/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/snmp-trap-listener/</guid><description/></item><item><title>SNMP Trap Node Attribution</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/snmp-trap-node-attribution/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/snmp-trap-node-attribution/</guid><description/></item><item><title>SNMP trap receiver dropping traps: silent UDP/162 loss</title><link>https://www.netdata.cloud/guides/network/network-snmp-trap-receiver-drops/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-snmp-trap-receiver-drops/</guid><description>&lt;h1 id="snmp-trap-receiver-dropping-traps-silent-udp162-loss">SNMP trap receiver dropping traps: silent UDP/162 loss&lt;/h1>
&lt;p>When SNMP traps silently disappear, the first place to look is rarely the device. SNMP traps are push-based UDP datagrams on port 162. The kernel buffers them, and the receiver application (typically &lt;code>snmptrapd&lt;/code> or a commercial collector) must drain that buffer faster than it fills. If it does not, the kernel silently drops datagrams and increments a counter the application never sees. No error is logged, and no alert fires.&lt;/p></description></item><item><title>SNMP Trap Relay Source Resolution</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/snmp-trap-relay-source-resolution/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/snmp-trap-relay-source-resolution/</guid><description/></item><item><title>SNMP Trap Reverse DNS Enrichment</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/snmp-trap-reverse-dns-enrichment/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/snmp-trap-reverse-dns-enrichment/</guid><description/></item><item><title>SNMP v2c vs v3: monitoring coverage and security trade-offs</title><link>https://www.netdata.cloud/guides/network/network-snmp-v2c-vs-v3/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-snmp-v2c-vs-v3/</guid><description>&lt;h1 id="snmp-v2c-vs-v3-monitoring-coverage-and-security-trade-offs">SNMP v2c vs v3: monitoring coverage and security trade-offs&lt;/h1>
&lt;p>The choice between SNMPv2c and SNMPv3 is rarely about whether v3 is more secure. It is. The real question is what you give up operationally when you move to v3, what breaks during migration, and where v2c remains the pragmatic default because the cost of v3 exceeds the risk it mitigates on a given segment.&lt;/p>
&lt;p>If you need the full network monitoring signal catalogue for context, see the &lt;a href="https://www.netdata.cloud/guides/network/network-monitoring-checklist/">network monitoring checklist&lt;/a>.&lt;/p></description></item><item><title>SNMPv3 authentication failures: authorizationError and usmStats decoded</title><link>https://www.netdata.cloud/guides/network/network-snmpv3-auth-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-snmpv3-auth-failures/</guid><description>&lt;h1 id="snmpv3-authentication-failures-authorizationerror-and-usmstats-decoded">SNMPv3 authentication failures: authorizationError and usmStats decoded&lt;/h1>
&lt;p>SNMPv3 authentication failures surface as &lt;code>authorizationError&lt;/code> (errorStatus 16, errorIndex 0) on the manager side, with no indication of which step of the User-based Security Model (USM) state machine failed. The manager reports &amp;ldquo;auth failed&amp;rdquo; but not whether the username was unknown, the HMAC digest mismatched, the packet arrived outside the time window, or the engine ID was never discovered.&lt;/p>
&lt;p>The agent knows exactly what went wrong. Every SNMPv3 engine maintains six read-only Counter32 statistics under the &lt;code>usmStats&lt;/code> subtree (RFC 3414). Each failed inbound packet increments exactly one counter before the agent responds with a Report PDU. Manager-side libraries translate that Report PDU into &lt;code>authorizationError&lt;/code>, discarding the counter value that pinpoints the root cause.&lt;/p></description></item><item><title>Socket statistics</title><link>https://www.netdata.cloud/integrations/data-collection/networking/socket-statistics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/socket-statistics/</guid><description/></item><item><title>Socomec Sicon Ups SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/socomec-sicon-ups-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/socomec-sicon-ups-snmp-traps/</guid><description/></item><item><title>SoftEther VPN Server</title><link>https://www.netdata.cloud/integrations/data-collection/networking/softether-vpn-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/softether-vpn-server/</guid><description/></item><item><title>SoftEther VPN Server Monitoring</title><link>https://www.netdata.cloud/monitoring-101/softether-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/softether-monitoring/</guid><description>&lt;h2 id="softether-vpn-server-monitoring">SoftEther VPN Server Monitoring&lt;/h2>
&lt;h3 id="what-is-softether-vpn-server">What Is SoftEther VPN Server?&lt;/h3>
&lt;p>SoftEther VPN Server is a robust and flexible Virtual Private Network (VPN) tool renowned for its scalability and variety of protocols support. It&amp;rsquo;s frequently deployed in business environments looking to secure connections across multiple network endpoints, provide remote access to corporate resources, and implement secure communication channels.&lt;/p>
&lt;h3 id="monitoring-softether-vpn-server-with-netdata">Monitoring SoftEther VPN Server With Netdata&lt;/h3>
&lt;p>When it comes to monitoring SoftEther VPN Server, Netdata stands out by utilizing an openmetrics (Prometheus) exporter facilitated by the &lt;a href="https://github.com/dalance/softether_exporter">SoftEther Exporter&lt;/a>. This exporter provides an efficient and precise way to gather critical metrics from your SoftEther VPN deployments.&lt;/p></description></item><item><title>SoftIRQ statistics</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/softirq-statistics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/softirq-statistics/</guid><description/></item><item><title>Softnet Statistics</title><link>https://www.netdata.cloud/integrations/data-collection/networking/softnet-statistics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/softnet-statistics/</guid><description/></item><item><title>Solar logging stick</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/solar-logging-stick/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/solar-logging-stick/</guid><description/></item><item><title>Solar Logging Stick Monitoring</title><link>https://www.netdata.cloud/monitoring-101/lsx-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/lsx-monitoring/</guid><description>&lt;h2 id="solar-logging-stick-monitoring">Solar Logging Stick Monitoring&lt;/h2>
&lt;h3 id="what-is-solar-logging-stick">What Is Solar Logging Stick?&lt;/h3>
&lt;p>The Solar Logging Stick is a device that facilitates the monitoring of solar energy metrics, contributing to efficient solar energy management and monitoring. As the global shift towards sustainable energy grows, devices like the Solar Logging Stick become crucial in tracking how well solar panels perform, how much energy they produce, and how they contribute to reducing carbon footprints.&lt;/p>
&lt;h3 id="monitoring-solar-logging-stick-with-netdata">Monitoring Solar Logging Stick With Netdata&lt;/h3>
&lt;p>To monitor the Solar Logging Stick, Netdata uses an OpenMetrics (Prometheus) exporter. This approach allows Netdata to ingest data from any Prometheus exporter and provide automated dashboards, alerts, and more, without needing a Prometheus server or Grafana setup. With Netdata, you can quickly set up comprehensive monitoring tools for the Solar Logging Stick, giving you real-time insights into your solar energy operations and helping ensure optimal performance.&lt;/p></description></item><item><title>Solis Ginlong 5G inverters</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/solis-ginlong-5g-inverters/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/solis-ginlong-5g-inverters/</guid><description/></item><item><title>Solis Ginlong 5G Inverters Monitoring</title><link>https://www.netdata.cloud/monitoring-101/solis-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/solis-monitoring/</guid><description>&lt;h2 id="solis-ginlong-5g-inverters-monitoring">Solis Ginlong 5G Inverters Monitoring&lt;/h2>
&lt;h3 id="what-is-solis-ginlong-5g-inverters">What Is Solis Ginlong 5G Inverters?&lt;/h3>
&lt;p>Solis Ginlong 5G inverters are a sophisticated component of solar energy systems that help convert the direct current electricity generated by solar panels into alternating current electricity. This conversion is essential for feeding energy into the grid or using it for home or business consumption. Ensuring these inverters function optimally is crucial for the efficiency and efficacy of solar energy systems.&lt;/p></description></item><item><title>Solr Monitoring</title><link>https://www.netdata.cloud/monitoring-101/solr-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/solr-monitoring/</guid><description>&lt;h2 id="what-is-solr">What is Solr?&lt;/h2>
&lt;p>&lt;a href="https://solr.apache.org/">Apache Solr&lt;/a> is an open source search platform built on top of Apache Lucene. It is used to quickly and easily search large volumes of data. It provides powerful features such as faceting, text analysis, and geo-spatial search. Solr is highly scalable and can be used in a wide variety of applications.&lt;/p>
&lt;h2 id="monitoring-solr-with-netdata">Monitoring Solr with Netdata&lt;/h2>
&lt;p>The prerequisites for monitoring Solr with Netdata are to have Solr and &lt;a href="https://learn.netdata.cloud/docs/cloud/get-started/">Netdata installed&lt;/a> on your system.&lt;/p></description></item><item><title>SONiC NOS</title><link>https://www.netdata.cloud/integrations/data-collection/networking/sonic-nos/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/sonic-nos/</guid><description/></item><item><title>SONiC NOS Monitoring</title><link>https://www.netdata.cloud/monitoring-101/sonic-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/sonic-monitoring/</guid><description>&lt;h2 id="sonic-nos-monitoring">SONiC NOS Monitoring&lt;/h2>
&lt;h3 id="what-is-sonic-nos">What Is SONiC NOS?&lt;/h3>
&lt;p>SONiC NOS (Software for Open Networking in the Cloud) is an open-source network operating system that enables the community to innovate across network hardware and software layers. It provides scalable and high-performance networking solutions, making it an essential asset in the modern data centers and cloud environments.&lt;/p>
&lt;h3 id="monitoring-sonic-nos-with-netdata">Monitoring SONiC NOS With Netdata&lt;/h3>
&lt;p>To effectively monitor SONiC NOS, Netdata uses an openmetrics (Prometheus) exporter. This integration allows Netdata to ingest data from any Prometheus exporter, providing automated dashboards, customized alerts, and in-depth insights without necessitating a Prometheus server or Grafana.&lt;/p></description></item><item><title>Sonicwall Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sonicwall-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sonicwall-inc-snmp-traps/</guid><description/></item><item><title>Sonix Communications Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sonix-communications-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sonix-communications-ltd-snmp-traps/</guid><description/></item><item><title>Sonus Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sonus-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sonus-networks-inc-snmp-traps/</guid><description/></item><item><title>Sophos Licensing</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/licensing-monitoring/sophos-licensing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/licensing-monitoring/sophos-licensing/</guid><description/></item><item><title>Sophos PLC SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sophos-plc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sophos-plc-snmp-traps/</guid><description/></item><item><title>Sophos XGS Firewall</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/sophos-xgs-firewall/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/sophos-xgs-firewall/</guid><description/></item><item><title>Spacelift</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/spacelift/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/spacelift/</guid><description/></item><item><title>Spacelift Monitoring</title><link>https://www.netdata.cloud/monitoring-101/spacelift-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/spacelift-monitoring/</guid><description>&lt;h2 id="spacelift-monitoring">Spacelift Monitoring&lt;/h2>
&lt;h3 id="what-is-spacelift">What Is Spacelift?&lt;/h3>
&lt;p>Spacelift is a powerful infrastructure-as-code (IaC) platform designed to manage and automate your infrastructure efficiently. It offers robust solutions for DevOps teams and IT professionals looking to streamline their workflows. By integrating with popular version control systems, Spacelift makes it easy to manage infrastructure modifications and deployments.&lt;/p>
&lt;h3 id="monitoring-spacelift-with-netdata">Monitoring Spacelift With Netdata&lt;/h3>
&lt;p>To monitor Spacelift effectively, Netdata utilizes an openmetrics (Prometheus) exporter, specifically the &lt;a href="https://github.com/spacelift-io/prometheus-exporter">Spacelift Exporter&lt;/a>. This approach enables the aggregation and visualization of crucial metrics without needing a standalone Prometheus server or Grafana dashboard. Netdata supports data ingestion from any Prometheus exporter, providing automated dashboards, alerts, and insightful analytics to keep your Spacelift environment running smoothly.&lt;/p></description></item><item><title>Spectra Logic SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/spectra-logic-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/spectra-logic-snmp-traps/</guid><description/></item><item><title>Sphinx</title><link>https://www.netdata.cloud/integrations/data-collection/databases/sphinx/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/sphinx/</guid><description/></item><item><title>Sphinx Monitoring</title><link>https://www.netdata.cloud/monitoring-101/sphinx-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/sphinx-monitoring/</guid><description>&lt;h2 id="sphinx-monitoring">Sphinx Monitoring&lt;/h2>
&lt;h3 id="what-is-sphinx">What Is Sphinx?&lt;/h3>
&lt;p>Sphinx is an open-source full-text search engine that provides powerful search capabilities with high-performance indexing. It&amp;rsquo;s designed to handle extensive data search requirements efficiently, making it an excellent choice for developers and organizations looking to enhance their data retrieval processes.&lt;/p>
&lt;h3 id="monitoring-sphinx-with-netdata">Monitoring Sphinx With Netdata&lt;/h3>
&lt;p>Monitoring Sphinx effectively is crucial for ensuring optimal performance in data search and indexing. Netdata offers a robust solution for Sphinx monitoring by utilizing an openmetrics (Prometheus) exporter. This approach allows Netdata to ingest data seamlessly from any Prometheus exporter, offering automated dashboards, alerts, and more without the need for a Prometheus server or Grafana. The &lt;a href="https://github.com/foxdalas/sphinx_exporter">Sphinx Exporter&lt;/a> enables you to gather valuable metrics about your Sphinx instance, providing insights into performance bottlenecks and resource usage.&lt;/p></description></item><item><title>Spidcom Technologies S.A. SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/spidcom-technologies-s.a.-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/spidcom-technologies-s.a.-snmp-traps/</guid><description/></item><item><title>SpigotMC</title><link>https://www.netdata.cloud/integrations/data-collection/applications/spigotmc/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/spigotmc/</guid><description/></item><item><title>SpigotMC Monitoring</title><link>https://www.netdata.cloud/monitoring-101/spigotmc-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/spigotmc-monitoring/</guid><description>&lt;h2 id="spigotmc-monitoring">SpigotMC Monitoring&lt;/h2>
&lt;h3 id="what-is-spigotmc">What Is SpigotMC?&lt;/h3>
&lt;p>SpigotMC is a high-performance Minecraft server platform that is synonymous with ease of customization and optimization, making it a popular choice among Minecraft server administrators. Whether you&amp;rsquo;re running a private server for friends or a large public server, maintaining optimal performance with SpigotMC requires careful monitoring of various metrics.&lt;/p>
&lt;h3 id="monitoring-spigotmc-with-netdata">Monitoring SpigotMC With Netdata&lt;/h3>
&lt;p>Netdata offers a robust and user-friendly monitoring solution for SpigotMC servers. As a comprehensive tool for monitoring SpigotMC, Netdata provides real-time insights into server performance and helps detect potential issues before they escalate. By using the &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/spigotmc/">SpigotMC monitoring tool&lt;/a>, administrators can track crucial metrics and maintain peak server performance effortlessly.&lt;/p></description></item><item><title>Spin_Retry_Count non-zero: the spindle motor is failing to spin up</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-spin-retry-count/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-spin-retry-count/</guid><description>&lt;h1 id="spin_retry_count-non-zero-the-spindle-motor-is-failing-to-spin-up">Spin_Retry_Count non-zero: the spindle motor is failing to spin up&lt;/h1>
&lt;p>Spin_Retry_Count (SMART attribute ID 10) counts how many times a hard drive&amp;rsquo;s spindle motor failed to reach operating RPM on the first attempt and retried. A healthy HDD spins up cleanly every time. Any non-zero value means the motor may not spin up on the next power cycle.&lt;/p>
&lt;p>This attribute is HDD-only. SSDs have no spindle motor and do not report ID 10. If you see this attribute on an SSD, it is either a synthetic placeholder or a vendor repurposing of the ID for unrelated data.&lt;/p></description></item><item><title>Spin_Up_Time climbing: bearing wear and lubricant degradation</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-spin-up-time-rising/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-spin-up-time-rising/</guid><description>&lt;h1 id="spin_up_time-climbing-bearing-wear-and-lubricant-degradation">Spin_Up_Time climbing: bearing wear and lubricant degradation&lt;/h1>
&lt;p>Spin_Up_Time (SMART attribute ID 3) measures how long the spindle motor takes to bring the platters from zero to rated RPM. The raw value is vendor-specific and noisy. A rising Spin_Up_Time alone is ambiguous; it becomes actionable only when you correlate it with Spin_Retry_Count (ID 10), drive temperature, and the scope of the problem across the chassis.&lt;/p>
&lt;p>This attribute applies only to mechanical HDDs. SSDs have no spindle motor and report Spin_Up_Time as zero or a synthetic placeholder.&lt;/p></description></item><item><title>Splunk</title><link>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/splunk/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/splunk/</guid><description/></item><item><title>Splunk SignalFx</title><link>https://www.netdata.cloud/integrations/exporters/splunk-signalfx/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/splunk-signalfx/</guid><description/></item><item><title>Splunk VictorOps</title><link>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/splunk-victorops/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/splunk-victorops/</guid><description/></item><item><title>Spring Tide Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/spring-tide-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/spring-tide-networks-inc-snmp-traps/</guid><description/></item><item><title>SQL Database agnostic Monitoring</title><link>https://www.netdata.cloud/monitoring-101/sql-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/sql-monitoring/</guid><description>&lt;h2 id="sql-database-agnostic-monitoring">SQL Database agnostic Monitoring&lt;/h2>
&lt;h3 id="what-is-sql-database-monitoring">What Is SQL Database Monitoring?&lt;/h3>
&lt;p>SQL databases are an integral part of most modern applications, serving as the backbone for data storage and management. Monitoring SQL databases involves tracking performance metrics, uptime, and other critical parameters to ensure database queries run smoothly and efficiently. It helps in preemptively identifying issues such as slow queries, connection bottlenecks, or resource exhaustion.&lt;/p>
&lt;h3 id="monitoring-sql-databases-with-netdata">Monitoring SQL Databases With Netdata&lt;/h3>
&lt;p>To effectively monitor SQL databases, &lt;a href="https://app.netdata.cloud/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata&lt;/a> utilizes an openmetrics (Prometheus) exporter. Netdata can ingest data from any Prometheus exporter, allowing users to set up automated dashboards, alerts, and more. This is achieved without needing a dedicated Prometheus server or Grafana, making it an efficient and streamlined solution for your database monitoring needs. With &lt;a href="https://github.com/free/sql_exporter">Netdata&amp;rsquo;s integration&lt;/a>, you can maintain real-time insights into your database health and performance seamlessly.&lt;/p></description></item><item><title>SQL databases (generic)</title><link>https://www.netdata.cloud/integrations/data-collection/databases/sql-databases-generic/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/sql-databases-generic/</guid><description/></item><item><title>SQL Server AG send and redo queues growing: replication lag and failover RTO</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-ag-send-redo-queue-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-ag-send-redo-queue-growing/</guid><description>&lt;h1 id="sql-server-ag-send-and-redo-queues-growing-replication-lag-and-failover-rto">SQL Server AG send and redo queues growing: replication lag and failover RTO&lt;/h1>
&lt;p>Two queues decide whether your Always On Availability Group can actually fail over: the send queue (log generated on the primary but not yet shipped to the secondary) and the redo queue (log received by the secondary but not yet replayed). When either grows without bound, replication lag is the visible symptom, but the hidden cost is failover RTO. On forced or automatic failover, the new primary must drain the entire redo queue before it accepts writes, so a queue that looks tolerable during steady state can turn a 30-second failover into a 30-minute one.&lt;/p></description></item><item><title>SQL Server AlwaysOn failover readiness: quorum, health checks, and the failover you assume works</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-ag-failover-readiness/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-ag-failover-readiness/</guid><description>&lt;h1 id="sql-server-alwayson-failover-readiness-quorum-health-checks-and-the-failover-you-assume-works">SQL Server AlwaysOn failover readiness: quorum, health checks, and the failover you assume works&lt;/h1>
&lt;p>An AG that reports &amp;ldquo;healthy&amp;rdquo; in the dashboard can still fail to failover when you need it. The synchronization state tells you data is flowing between replicas. It says nothing about whether the cluster can orchestrate a failover, whether the health detection policy will catch the specific failure you are about to have, or whether the cluster has already exhausted its automatic failover budget for the period.&lt;/p></description></item><item><title>SQL Server Availability Group not synchronizing: NOT_HEALTHY replicas and failover risk</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-ag-not-synchronizing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-ag-not-synchronizing/</guid><description>&lt;h1 id="sql-server-availability-group-not-synchronizing-not_healthy-replicas-and-failover-risk">SQL Server Availability Group not synchronizing: NOT_HEALTHY replicas and failover risk&lt;/h1>
&lt;p>The symptom arrives as an alert or a dashboard color change: a synchronous-commit secondary replica is reporting &lt;code>synchronization_health_desc = NOT_HEALTHY&lt;/code> or &lt;code>connected_state_desc = DISCONNECTED&lt;/code> in &lt;code>sys.dm_hadr_availability_replica_states&lt;/code>. The primary is still accepting writes, but the protection you assumed is degraded or gone.&lt;/p>
&lt;p>In synchronous-commit mode, the primary waits for the secondary to harden log records before acknowledging commits. When the secondary drops or stops keeping up, the primary either continues unprotected or stops accepting writes entirely, depending on &lt;code>required_synchronized_secondaries_to_commit&lt;/code>. Either way, your recovery point objective and your recovery time objective are both at risk.&lt;/p></description></item><item><title>SQL Server backup freshness: the recovery point you only discover you lack during an incident</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-backup-freshness/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-backup-freshness/</guid><description>&lt;h1 id="sql-server-backup-freshness-the-recovery-point-you-only-discover-you-lack-during-an-incident">SQL Server backup freshness: the recovery point you only discover you lack during an incident&lt;/h1>
&lt;p>The recovery point you actually have is the recovery point you can restore to, not the one your schedule promises. Backup freshness is the gap between those two, measured as the time since the last successful full, differential, and transaction log backup per database. When that gap is wrong, you find out during restore: either an analyst files a ticket for missing data, or an incident forces point-in-time recovery and the chain breaks.&lt;/p></description></item><item><title>SQL Server blocking chains: finding the head blocker before workers run out</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-blocking-chain/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-blocking-chain/</guid><description>&lt;h1 id="sql-server-blocking-chains-finding-the-head-blocker-before-workers-run-out">SQL Server blocking chains: finding the head blocker before workers run out&lt;/h1>
&lt;p>SQL Server is unresponsive. CPU and I/O counters are low. Connections succeed but queries hang. Timeouts and login failures follow. This is the shape of a blocking chain that has crossed into worker-thread exhaustion.&lt;/p>
&lt;p>One session holds a lock. Conflicting sessions queue behind it, each waiting on an &lt;code>LCK_M_*&lt;/code> wait and each pinning a worker from SQL Server&amp;rsquo;s fixed-size pool. As the chain deepens, the worker pool drains. Once exhausted, new requests get &lt;code>THREADPOOL&lt;/code> waits and the instance appears down to applications, even though the OS shows &lt;code>sqlservr&lt;/code> healthy and storage idle.&lt;/p></description></item><item><title>SQL Server buffer cache hit ratio low: when the working set no longer fits in memory</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-buffer-cache-hit-ratio-low/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-buffer-cache-hit-ratio-low/</guid><description>&lt;h1 id="sql-server-buffer-cache-hit-ratio-low-when-the-working-set-no-longer-fits-in-memory">SQL Server buffer cache hit ratio low: when the working set no longer fits in memory&lt;/h1>
&lt;p>A low buffer cache hit ratio (BCHR) gets two reactions in production: teams page on-call for a single dip during a maintenance window, or they ignore a sustained decline because the counter is &amp;ldquo;unreliable.&amp;rdquo; Both are wrong. BCHR is a weak signal alone, but paired with Page Life Expectancy (PLE), &lt;code>PAGEIOLATCH_*&lt;/code> waits, and workload context, it tells you whether your working set still fits in the buffer pool.&lt;/p></description></item><item><title>SQL Server CPU utilization high: telling query load apart from a bad plan</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-cpu-utilization-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-cpu-utilization-high/</guid><description>&lt;h1 id="sql-server-cpu-utilization-high-telling-query-load-apart-from-a-bad-plan">SQL Server CPU utilization high: telling query load apart from a bad plan&lt;/h1>
&lt;p>Your monitoring says the SQL Server host is at 98% CPU. Before you page anyone or start killing sessions: SQL Server is designed to use available CPU. A cold buffer pool after restart, backup compression, an ETL window, or a well-parallelized reporting query will all legitimately pin CPU at 90%+. High CPU is a symptom with no severity attached until you answer two questions: who is burning the CPU (SQL Server or something else on the host), and is the work useful (throughput) or wasted (a bad plan, a compilation storm, or spinlock contention).&lt;/p></description></item><item><title>SQL Server CXPACKET and CXCONSUMER waits: parallelism, MAXDOP, and what is actually wrong</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-cxpacket-cxconsumer-waits/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-cxpacket-cxconsumer-waits/</guid><description>&lt;h1 id="sql-server-cxpacket-and-cxconsumer-waits-parallelism-maxdop-and-what-is-actually-wrong">SQL Server CXPACKET and CXCONSUMER waits: parallelism, MAXDOP, and what is actually wrong&lt;/h1>
&lt;p>You opened &lt;code>sys.dm_os_wait_stats&lt;/code>, excluded the idle noise, and CXPACKET is sitting at the top consuming 40, 50, maybe 70 percent of total wait time. The first search result tells you parallelism is out of control. The second tells you to set MAXDOP to 1. Both are usually wrong.&lt;/p>
&lt;p>CXPACKET is routinely the number one wait on healthy systems. Its presence alone means parallel queries are running, and parallel threads spend much of their existence waiting for each other. The wait is a side effect of work being done in parallel, not the disease. The real questions are whether that parallel work is skewed, whether the queries going parallel should be parallel at all, and whether the engine is burning worker threads and CPU on plans that would be faster serial.&lt;/p></description></item><item><title>SQL Server database in SUSPECT or RECOVERY_PENDING: an offline database and how to recover it</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-database-suspect-recovery-pending/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-database-suspect-recovery-pending/</guid><description>&lt;h1 id="sql-server-database-in-suspect-or-recovery_pending-an-offline-database-and-how-to-recover-it">SQL Server database in SUSPECT or RECOVERY_PENDING: an offline database and how to recover it&lt;/h1>
&lt;p>A production database is showing &lt;code>state_desc = SUSPECT&lt;/code> or &lt;code>RECOVERY_PENDING&lt;/code> in &lt;code>sys.databases&lt;/code>. Applications cannot open connections to that database. Users are seeing login failures, query timeouts, or generic &amp;ldquo;database cannot be opened&amp;rdquo; errors. The SQL Server instance itself is up, and every other database on it may be fine.&lt;/p>
&lt;p>&lt;code>RECOVERY_PENDING&lt;/code> rarely means corruption. It usually means SQL Server could not get the resources it needed during recovery: a missing file, a full log volume, a permissions change, or a transient I/O failure at startup. &lt;code>SUSPECT&lt;/code> is more serious because recovery actually ran and failed, but it still does not automatically mean data loss. The wrong move is to jump straight to &lt;code>DBCC CHECKDB&lt;/code> with &lt;code>REPAIR_ALLOW_DATA_LOSS&lt;/code>. The right move is to fix the underlying resource, re-run recovery, and only fall back to repair or restore when that fails.&lt;/p></description></item><item><title>SQL Server Error 1205: transaction was deadlocked and chosen as the deadlock victim</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-error-1205-deadlock-victim/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-error-1205-deadlock-victim/</guid><description>&lt;h1 id="sql-server-error-1205-transaction-was-deadlocked-and-chosen-as-the-deadlock-victim">SQL Server Error 1205: transaction was deadlocked and chosen as the deadlock victim&lt;/h1>
&lt;p>The error text returned to the client is explicit:&lt;/p>
&lt;blockquote>
&lt;p>Transaction (Process ID %d) was deadlocked on %.*ls resources with another process and has been chosen as the deadlock victim. Rerun the transaction.&lt;/p>
&lt;/blockquote>
&lt;p>&lt;code>%d&lt;/code> is the SPID. &lt;code>%.*ls&lt;/code> names the resource type, typically &lt;code>lock&lt;/code>. The message tells the application to rerun the transaction but not why the deadlock happened, which resource was contended, or which other session was involved. To answer those questions you need the deadlock graph.&lt;/p></description></item><item><title>SQL Server Error 18456: login failed for user, and what the state code means</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-error-18456-login-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-error-18456-login-failed/</guid><description>&lt;p>SQL Server Error 18456 is the universal &amp;ldquo;Login failed for user X&amp;rdquo; message. It is deliberately vague: every client, from &lt;code>sqlcmd&lt;/code> to the application&amp;rsquo;s connection pool, sees the same string with severity 14 and state 1. The client never learns whether the password was wrong, the login does not exist, the database is offline, or the account is disabled. That information lives only in the SQL Server error log, encoded as a state code.&lt;/p></description></item><item><title>SQL Server Error 701: there is insufficient system memory to run this query</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-error-701-insufficient-memory/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-error-701-insufficient-memory/</guid><description>&lt;p>Error 701 is one of SQL Server&amp;rsquo;s bluntest messages: &amp;ldquo;There is insufficient system memory in resource pool &amp;lsquo;default&amp;rsquo; to run this query.&amp;rdquo; When it fires, the engine could not satisfy an allocation. Queries that were running fine seconds ago start failing, and the failure cascades into application timeouts, retry storms, and a flood of related errors (17890, 8645).&lt;/p>
&lt;p>The error itself tells you very little. It does not say whether the buffer pool is starved, whether a single query is hoarding a memory grant, whether an Extended Events ring buffer has eaten 50 GB, or whether the OS is reclaiming memory from a VM balloon driver. Investigate by source.&lt;/p></description></item><item><title>SQL Server Error 823 and 824: I/O and logical consistency errors</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-error-823-824-io-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-error-823-824-io-errors/</guid><description>&lt;p>Errors 823 and 824 are SQL Server&amp;rsquo;s severity-24 storage integrity alarms. 823 means the operating system reported a hard failure on a file API call. 824 means the call succeeded but the page failed an internal integrity check. Both are PAGE-worthy the moment they appear; they do not self-resolve, and continued use of the affected files risks losing data that was fine minutes earlier.&lt;/p>
&lt;p>Error 825 is the soft warning that usually precedes both: SQL Server retried a read that initially failed and eventually succeeded. The query did not fail and no connection was killed, which is why most teams do not alert on it. It is also the most reliable predictor that an 823 or 824 is coming.&lt;/p></description></item><item><title>SQL Server Error 825: read-retry succeeded and the disk is failing</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-error-825-read-retry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-error-825-read-retry/</guid><description>&lt;h1 id="sql-server-error-825-read-retry-succeeded-and-the-disk-is-failing">SQL Server Error 825: read-retry succeeded and the disk is failing&lt;/h1>
&lt;p>Error 825 is what SQL Server writes to the error log when a disk read failed on the first attempt but succeeded on a retry (attempt 2, 3, or 4). The query completes. The application sees no failure. But the storage underneath just told you it is failing.&lt;/p>
&lt;p>Most monitoring setups never surface Error 825. It is a severity-10 informational message, and typical SQL Server Agent alert configurations target severity 19 and above. The error sits quietly in the log until something harder arrives: an 823 (hard I/O error) or an 824 (logical consistency error). By then the page may already be unreadable.&lt;/p></description></item><item><title>SQL Server Error 9002: the transaction log for the database is full</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-error-9002-transaction-log-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-error-9002-transaction-log-full/</guid><description>&lt;h1 id="sql-server-error-9002-the-transaction-log-for-the-database-is-full">SQL Server Error 9002: the transaction log for the database is full&lt;/h1>
&lt;p>Your application is throwing write failures and the SQL Server error log shows: &amp;ldquo;The transaction log for database &amp;lsquo;X&amp;rsquo; is full due to &amp;lsquo;LOG_BACKUP&amp;rsquo;&amp;rdquo; (or ACTIVE_TRANSACTION, AVAILABILITY_REPLICA, REPLICATION, or another reason in quotes). Every INSERT, UPDATE, and DELETE against that database now fails with error 9002. Read-only queries may still work, which makes the outage look confusingly partial from the outside.&lt;/p></description></item><item><title>SQL Server failed login storm: brute force, credential drift, and service-account failures</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-failed-login-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-failed-login-storm/</guid><description>&lt;h1 id="sql-server-failed-login-storm-brute-force-credential-drift-and-service-account-failures">SQL Server failed login storm: brute force, credential drift, and service-account failures&lt;/h1>
&lt;p>A failed-login storm fills ERRORLOG with &amp;ldquo;Login failed for user&amp;rdquo; entries (Error 18456). Counting them is the first instinct and the wrong one. Aggregate rate cannot tell brute force from a batch job running on rotated secrets, a service account whose password just expired, an application pool recycling, or &lt;!-- TODO: verify --> a SQL Server 2025 replication secondary emitting benign noise every few minutes.&lt;/p></description></item><item><title>SQL Server HADR_SYNC_COMMIT waits: a synchronous secondary throttling primary commits</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-hadr-sync-commit-waits/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-hadr-sync-commit-waits/</guid><description>&lt;h1 id="sql-server-hadr_sync_commit-waits-a-synchronous-secondary-throttling-primary-commits">SQL Server HADR_SYNC_COMMIT waits: a synchronous secondary throttling primary commits&lt;/h1>
&lt;p>HADR_SYNC_COMMIT at the top of your wait statistics is a counterintuitive failure. The primary replica looks idle: CPU is low, local I/O is fast, batch requests look normal. Yet every write transaction stalls. The cause is not on the primary but on the synchronous secondary, the network between replicas, or the AG transport itself.&lt;/p>
&lt;p>In synchronous-commit mode, the primary cannot acknowledge a transaction commit until the secondary hardens the log record to disk. Every write transaction pays that round-trip tax. When the secondary or the path to it degrades, that tax becomes seconds of latency per commit. Under sustained write load, the delays lengthen lock hold times on the primary and cascade into blocking, worker thread growth, and eventually THREADPOOL waits. The application experiences this as a general slowdown, but the root cause is a single replica that cannot keep up.&lt;/p></description></item><item><title>SQL Server high compilations per second: plan cache pollution and CPU burn</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-high-compilations-per-second/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-high-compilations-per-second/</guid><description>&lt;h1 id="sql-server-high-compilations-per-second-plan-cache-pollution-and-cpu-burn">SQL Server high compilations per second: plan cache pollution and CPU burn&lt;/h1>
&lt;p>SQL Compilations/sec is climbing, CPU is pinned, and the application is reporting latency even though storage and locking look clean. Nothing is &amp;ldquo;broken&amp;rdquo; in the error log, but the engine is spending a large share of its CPU budget turning query text into execution plans instead of executing them.&lt;/p>
&lt;p>Compilation is expensive. Every new plan costs parse, optimization, and often a compile-time memory grant. When the ratio of SQL Compilations/sec to Batch Requests/sec climbs past roughly 10%, the plan cache is failing to do its job. Sustained above 20% with CPU pressure, you have an active problem.&lt;/p></description></item><item><title>SQL Server high recompilations: stale statistics and schema changes churning plans</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-high-recompilations/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-high-recompilations/</guid><description>&lt;h1 id="sql-server-high-recompilations-stale-statistics-and-schema-changes-churning-plans">SQL Server high recompilations: stale statistics and schema changes churning plans&lt;/h1>
&lt;p>CPU is climbing, &lt;code>SOS_SCHEDULER_YIELD&lt;/code> is creeping up, batch requests look normal, and your top waits are not obviously I/O or lock related. Check the SQL Statistics counters: if &lt;code>SQL Re-Compilations/sec&lt;/code> is running well above 10% of &lt;code>SQL Compilations/sec&lt;/code>, you have a plan stability problem, not a plan cache efficiency problem.&lt;/p>
&lt;p>Recompilations are not the same as first-time compilations. First-time compilations (or recompilations forced by plan cache eviction under memory pressure) belong in &lt;a href="https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-high-compilations-per-second/">SQL Server high compilations per second&lt;/a>. Recompilations mean SQL Server had a cached plan, decided it could no longer trust it, and spent CPU rebuilding it. Each recompile is CPU work, and if the recompiling statement sits inside a multi-statement batch or stored procedure, the cost cascades across dependent statements.&lt;/p></description></item><item><title>SQL Server high VLF count: transaction log fragmentation that slows recovery</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-vlf-count-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-vlf-count-high/</guid><description>&lt;h1 id="sql-server-high-vlf-count-transaction-log-fragmentation-that-slows-recovery">SQL Server high VLF count: transaction log fragmentation that slows recovery&lt;/h1>
&lt;p>A database restarts and takes 45 minutes to come back ONLINE. An AlwaysOn failover completes in seconds, but the new primary sits in RECOVERING for an hour while the redo queue drains. The error log shows nothing obviously wrong: no 823 or 824, no corruption, no missing files. CPU and I/O during recovery look modest, just stretched out. Users escalate.&lt;/p></description></item><item><title>SQL Server I/O stall high: per-file storage latency from dm_io_virtual_file_stats</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-io-stall-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-io-stall-high/</guid><description>&lt;h1 id="sql-server-io-stall-high-per-file-storage-latency-from-dm_io_virtual_file_stats">SQL Server I/O stall high: per-file storage latency from dm_io_virtual_file_stats&lt;/h1>
&lt;p>When &lt;code>PAGEIOLATCH_*&lt;/code> or &lt;code>WRITELOG&lt;/code> climbs to the top of &lt;code>sys.dm_os_wait_stats&lt;/code>, the question is whether storage is the bottleneck or whether SQL Server is doing too much physical I/O because the buffer pool is undersized. &lt;code>sys.dm_io_virtual_file_stats(NULL, NULL)&lt;/code> is the only DMV that answers this at file granularity from inside the engine.&lt;/p>
&lt;p>The values it returns are cumulative since the SQL Server service last started, so a single snapshot is almost useless. Compute deltas over a short window (15-60 seconds during the incident) and divide stall time by operation count to get current per-file latency. Lifetime averages on a long-running instance hide the bad minutes that matter.&lt;/p></description></item><item><title>SQL Server instance down: no response on port 1433 and where to look first</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-instance-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-instance-down/</guid><description>&lt;h1 id="sql-server-instance-down-no-response-on-port-1433-and-where-to-look-first">SQL Server instance down: no response on port 1433 and where to look first&lt;/h1>
&lt;p>Your availability probe just fired: TCP connect plus &lt;code>SELECT 1&lt;/code> against port 1433 has failed three or more times over at least 60 seconds. Before you restart anything, separate the two failure modes that get lumped together as &amp;ldquo;SQL Server is down&amp;rdquo;. They have different causes, different fixes, and different blast radii.&lt;/p>
&lt;p>Mode one: no TCP connect at all. The listener is not accepting connections on 1433 (or the named instance&amp;rsquo;s dynamic port). The service is stopped, the host is down, the network path is broken, or the listener is misconfigured, commonly after an AlwaysOn failover. The engine is not there to talk to.&lt;/p></description></item><item><title>SQL Server LCK_M waits high: lock contention and what the suffixes mean</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-lck-m-waits-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-lck-m-waits-high/</guid><description>&lt;h1 id="sql-server-lck_m-waits-high-lock-contention-and-what-the-suffixes-mean">SQL Server LCK_M waits high: lock contention and what the suffixes mean&lt;/h1>
&lt;p>When LCK_M_* wait types dominate &lt;code>sys.dm_os_wait_stats&lt;/code>, worker threads are suspended waiting for locks instead of doing work. The prefix is uniform; the suffix is the lock mode the waiter requested, and it is the diagnostic signal. &lt;code>LCK_M_S&lt;/code> means a reader is blocked. &lt;code>LCK_M_IX&lt;/code> means an intent-exclusive writer is queued behind a conflicting holder. &lt;code>LCK_M_SCH_M&lt;/code> is a DDL operation waiting on schema modification. The mode narrows the suspect list immediately.&lt;/p></description></item><item><title>SQL Server lock escalation: when row locks become a table lock and block everyone</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-lock-escalation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-lock-escalation/</guid><description>&lt;h1 id="sql-server-lock-escalation-when-row-locks-become-a-table-lock-and-block-everyone">SQL Server lock escalation: when row locks become a table lock and block everyone&lt;/h1>
&lt;p>A batch UPDATE or DELETE runs longer than usual. CPU looks fine, I/O looks fine, and then dozens of sessions start waiting on &lt;code>LCK_M_*&lt;/code> waits, all blocked by the batch session. By the time you log in, the worker thread pool is draining and the instance is heading toward &lt;code>THREADPOOL&lt;/code>. The root cause is not a hung transaction or a resource bottleneck. It is lock escalation: SQL Server traded thousands of row locks for a single table lock, and every other session that wants to touch that table is now serialized behind it.&lt;/p></description></item><item><title>SQL Server log autogrow stall: why every write pauses while the log file grows</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-log-autogrow-stall/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-log-autogrow-stall/</guid><description>&lt;h1 id="sql-server-log-autogrow-stall-why-every-write-pauses-while-the-log-file-grows">SQL Server log autogrow stall: why every write pauses while the log file grows&lt;/h1>
&lt;p>Applications report periodic, mysterious write stalls. Throughput drops briefly, then recovers. Repeat. The transaction log on the affected database is creeping upward, and there are no alarms on disk space. What you are seeing is almost certainly log autogrow stall: every write transaction in the database pauses while SQL Server expands and zero-initializes the transaction log file.&lt;/p></description></item><item><title>SQL Server log backups missing: the full-recovery log that grows forever</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-log-backup-chain-broken/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-log-backup-chain-broken/</guid><description>&lt;h1 id="sql-server-log-backups-missing-the-full-recovery-log-that-grows-forever">SQL Server log backups missing: the full-recovery log that grows forever&lt;/h1>
&lt;p>The application starts throwing write errors. The database is online, reads work, but every INSERT, UPDATE, and DELETE fails with error 9002: the transaction log is full. The log volume is at zero free space, or the log file has auto-grown to many times the size of the data files. When you ask when the last log backup ran, nobody knows.&lt;/p></description></item><item><title>SQL Server log_reuse_wait_desc: why the transaction log will not truncate</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-log-reuse-wait-desc/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-log-reuse-wait-desc/</guid><description>&lt;h1 id="sql-server-log_reuse_wait_desc-why-the-transaction-log-will-not-truncate">SQL Server log_reuse_wait_desc: why the transaction log will not truncate&lt;/h1>
&lt;p>The database is throwing error 9002, writes are failing, the log file has eaten the volume, and someone is about to add another log file or shrink the existing one. Before anyone touches file sizes, run one query:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-sql" data-lang="sql">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">SELECT&lt;/span> name, recovery_model_desc, log_reuse_wait_desc
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">FROM&lt;/span> sys.databases;
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>That third column is the root-cause field most teams skip. It tells you why log truncation could not clear any Virtual Log Files (VLFs) the last time SQL Server tried. Every fix for a full log flows from this value. Adding disk space without reading it treats the symptom: the log fills again, usually within hours, and now you also have VLF fragmentation to deal with.&lt;/p></description></item><item><title>SQL Server max server memory: setting it so the OS and buffer pool both survive</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-max-server-memory-configuration/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-max-server-memory-configuration/</guid><description>&lt;h1 id="sql-server-max-server-memory-setting-it-so-the-os-and-buffer-pool-both-survive">SQL Server max server memory: setting it so the OS and buffer pool both survive&lt;/h1>
&lt;p>SQL Server is built to consume memory until something stops it. On a dedicated box with no cap, the engine will eat nearly all physical RAM for the buffer pool and keep going until the OS pushes back. That is not a leak; it is the design. &lt;code>max server memory&lt;/code> is the one knob that turns that behavior into something safe to run next to other processes.&lt;/p></description></item><item><title>SQL Server Memory Grants Pending above zero: queries queued before they can run</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-memory-grants-pending/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-memory-grants-pending/</guid><description>&lt;h1 id="sql-server-memory-grants-pending-above-zero-queries-queued-before-they-can-run">SQL Server Memory Grants Pending above zero: queries queued before they can run&lt;/h1>
&lt;p>Memory Grants Pending is a SQLServer:Memory Manager counter that should almost always be zero. When it rises above zero, queries are fully parsed and optimized but cannot start execution because the query workspace memory pool is exhausted. Applications see queries that appear hung: connections stay open, latency climbs, and CPU and I/O may look idle.&lt;/p>
&lt;p>The counter pairs with the RESOURCE_SEMAPHORE wait type. The counter shows how many queries are queued right now. The wait type aggregates how long they spent in that queue. Both point at the same bottleneck: SQL Server cannot honor a memory grant request because the workspace memory pool has been drained by other queries, by one oversized grant from a bad plan, or by external OS pressure on max server memory.&lt;/p></description></item><item><title>SQL Server memory pressure spiral: PLE collapse, physical reads, and I/O saturation</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-memory-pressure-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-memory-pressure-spiral/</guid><description>&lt;h1 id="sql-server-memory-pressure-spiral-ple-collapse-physical-reads-and-io-saturation">SQL Server memory pressure spiral: PLE collapse, physical reads, and I/O saturation&lt;/h1>
&lt;p>The signature of a memory pressure spiral is a rapid Page Life Expectancy (PLE) drop combined with rising &lt;code>PAGEIOLATCH_*&lt;/code> waits and increasing I/O stall on data files. CPU sits low to moderate because the bottleneck is not compute. It is the buffer pool being drained, which forces physical reads, which saturate the storage path, which makes every query slower, which piles up more concurrent sessions, which pressures memory further.&lt;/p></description></item><item><title>SQL Server Page Life Expectancy dropping: the buffer pool under memory pressure</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-page-life-expectancy-dropping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-page-life-expectancy-dropping/</guid><description>&lt;h1 id="sql-server-page-life-expectancy-dropping-the-buffer-pool-under-memory-pressure">SQL Server Page Life Expectancy dropping: the buffer pool under memory pressure&lt;/h1>
&lt;p>Page Life Expectancy (PLE) measures the expected seconds a data page stays in the buffer pool before eviction. A sharp drop means the lazy writer is flushing pages faster than the workload can reuse them, and queries that previously hit cache now trigger physical reads.&lt;/p>
&lt;p>The signature of a real pressure event is a 50%+ drop from baseline coinciding with rising &lt;code>PAGEIOLATCH_*&lt;/code> waits and a falling buffer cache hit ratio. The old &amp;ldquo;300 seconds&amp;rdquo; rule is obsolete: it was calibrated for 32-bit servers with ~4GB buffer pools. On a modern instance with 256GB of memory, PLE above 100,000 seconds is normal, and 300 seconds would be a crisis.&lt;/p></description></item><item><title>SQL Server PAGEIOLATCH waits: the buffer pool waiting on slow storage</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-pageiolatch-waits/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-pageiolatch-waits/</guid><description>&lt;h1 id="sql-server-pageiolatch-waits-the-buffer-pool-waiting-on-slow-storage">SQL Server PAGEIOLATCH waits: the buffer pool waiting on slow storage&lt;/h1>
&lt;p>Your top wait type is &lt;code>PAGEIOLATCH_SH&lt;/code> or &lt;code>PAGEIOLATCH_EX&lt;/code>. Queries that used to be sub-50ms now take seconds. CPU may be low. There are no blocking chains. The buffer pool is waiting on disk.&lt;/p>
&lt;p>PAGEIOLATCH waits are the direct fingerprint of physical I/O. A worker needs an 8KB data page that is not in the buffer pool, takes an in-memory latch on the buffer descriptor, and waits for storage to return the page. When the read completes, the worker continues. Sustained PAGEIOLATCH time means the engine is spending wall-clock time waiting on disk.&lt;/p></description></item><item><title>SQL Server parameter sniffing: a good plan for one value, catastrophic for the next</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-parameter-sniffing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-parameter-sniffing/</guid><description>&lt;h1 id="sql-server-parameter-sniffing-a-good-plan-for-one-value-catastrophic-for-the-next">SQL Server parameter sniffing: a good plan for one value, catastrophic for the next&lt;/h1>
&lt;p>A stored procedure that ran in 30 ms for weeks is now taking 8 seconds per call. Query text unchanged. Indexes unchanged. Statistics look fine. The application is timing out. The same query with literal values returns instantly. With &lt;code>OPTION (RECOMPILE)&lt;/code> appended, it returns instantly. The cached plan is the problem.&lt;/p>
&lt;p>This is parameter sniffing. The optimizer compiled the plan for the parameter values it saw first, cached the result, and every subsequent call reuses a plan that is wrong for the values it is now processing. On a table with skewed data, a plan sized for 12 rows is catastrophic when fed 12 million.&lt;/p></description></item><item><title>SQL Server plan cache bloat: single-use ad-hoc plans wasting memory</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-plan-cache-bloat/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-plan-cache-bloat/</guid><description>&lt;h1 id="sql-server-plan-cache-bloat-single-use-ad-hoc-plans-wasting-memory">SQL Server plan cache bloat: single-use ad-hoc plans wasting memory&lt;/h1>
&lt;p>SQL Server plan cache bloat from single-use ad-hoc plans is a slow degradation. Queries return correct results. The instance stays up. But memory that should cache hot data pages holds thousands of compiled plans that will never execute again. Operators typically discover it when PLE drifts down for no obvious reason, when buffer cache hit ratio slips, or when compilations per second stays elevated relative to batch requests. By the time it surfaces in user-facing latency, the plan cache has been stealing buffer pool memory for weeks.&lt;/p></description></item><item><title>SQL Server RESOURCE_SEMAPHORE waits: queries stuck waiting for a memory grant</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-resource-semaphore-waits/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-resource-semaphore-waits/</guid><description>&lt;h1 id="sql-server-resource_semaphore-waits-queries-stuck-waiting-for-a-memory-grant">SQL Server RESOURCE_SEMAPHORE waits: queries stuck waiting for a memory grant&lt;/h1>
&lt;p>RESOURCE_SEMAPHORE is the wait type SQL Server records when a worker thread cannot get a query memory grant. Before a query runs a sort, hash, or certain joins, the optimizer estimates how much workspace memory it needs and asks the grant pool for it. When the pool is exhausted, parsed-and-optimized queries sit in a queue. To the application they look hung.&lt;/p></description></item><item><title>SQL Server restore readiness: why a backup you never test is not a backup</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-restore-readiness/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-restore-readiness/</guid><description>&lt;h1 id="sql-server-restore-readiness-why-a-backup-you-never-test-is-not-a-backup">SQL Server restore readiness: why a backup you never test is not a backup&lt;/h1>
&lt;p>The most common SQL Server backup monitoring answers one question: did the job succeed? A green checkmark means SQL Server wrote bytes to a file. It does not mean those bytes can be read back, decrypted, applied through a recovery sequence, and brought online within your RTO.&lt;/p>
&lt;p>The gap between &amp;ldquo;backup succeeded&amp;rdquo; and &amp;ldquo;we can recover&amp;rdquo; is where most DR failures actually live. Backups fail to restore because of media corruption the backup job never validated, because a TDE or backup-encryption certificate was not backed up alongside the data, because an ad-hoc operation broke the log chain and silently invalidated hours of log backups, or because the restore takes four times longer than anyone measured. None of those surface in backup-completion monitoring.&lt;/p></description></item><item><title>SQL Server runnable tasks backlog: the in-engine CPU queue OS metrics miss</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-runnable-tasks-backlog/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-runnable-tasks-backlog/</guid><description>&lt;h1 id="sql-server-runnable-tasks-backlog-the-in-engine-cpu-queue-os-metrics-miss">SQL Server runnable tasks backlog: the in-engine CPU queue OS metrics miss&lt;/h1>
&lt;p>Your monitoring says the host is at 45% CPU. Users say the database is slow. Both are right, because SQL Server does not use the OS scheduler for query execution. It runs its own cooperative scheduling layer, SQLOS, with one scheduler per logical CPU, its own run queues, and its own worker thread pool. OS CPU percent measures what the host kernel sees. It does not see tasks sitting in SQLOS queues waiting for a scheduler slot, and on virtualized hosts it does not see time the hypervisor stole from the guest.&lt;/p></description></item><item><title>SQL Server sleeping head blocker: the idle session holding a lock and an open transaction</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-sleeping-head-blocker/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-sleeping-head-blocker/</guid><description>&lt;h1 id="sql-server-sleeping-head-blocker-the-idle-session-holding-a-lock-and-an-open-transaction">SQL Server sleeping head blocker: the idle session holding a lock and an open transaction&lt;/h1>
&lt;p>The most dangerous blocking shape in SQL Server does not look active. Sessions pile up behind a head blocker, batch requests stall, and transactions per second drop. CPU and I/O are idle. The engine looks healthy on resource metrics, but queries are timing out.&lt;/p>
&lt;p>When you query &lt;code>sys.dm_exec_requests&lt;/code> for the blocker, you find nothing. The &lt;code>session_id&lt;/code> that everyone is waiting on does not appear in the requests DMV because the head blocker has no active request. It is sleeping. It is, however, holding locks under an open transaction that never committed, never rolled back, and never will on its own.&lt;/p></description></item><item><title>SQL Server SOS_SCHEDULER_YIELD waits: CPU scheduler pressure explained</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-sos-scheduler-yield-waits/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-sos-scheduler-yield-waits/</guid><description>&lt;h1 id="sql-server-sos_scheduler_yield-waits-cpu-scheduler-pressure-explained">SQL Server SOS_SCHEDULER_YIELD waits: CPU scheduler pressure explained&lt;/h1>
&lt;p>When SOS_SCHEDULER_YIELD dominates &lt;code>sys.dm_os_wait_stats&lt;/code>, the reflex is to assume CPU pressure and start hunting for bad queries. That reflex is right about half the time. The other half, you are chasing a signal that is doing exactly what it was designed to do: recording every time a worker voluntarily yielded its 4ms quantum because other runnable workers were queued.&lt;/p>
&lt;p>The wait type has existed in SQL Server since the SQLOS era and the 4ms quantum is fixed in every version. It cannot be tuned. What you can tune is your interpretation. A workload doing efficient set-based scans of pages already in memory will yield constantly and rack up enormous SOS_SCHEDULER_YIELD numbers without anything being wrong. A VM on an oversubscribed host reports the same wait type while the hypervisor silently steals CPU cycles SQL Server cannot see.&lt;/p></description></item><item><title>SQL Server suspect_pages: the durable record of storage corruption</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-suspect-pages/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-suspect-pages/</guid><description>&lt;h1 id="sql-server-suspect_pages-the-durable-record-of-storage-corruption">SQL Server suspect_pages: the durable record of storage corruption&lt;/h1>
&lt;p>The SQL Server error log recycles on every service restart, and by default only six archived logs are retained. When a storage fault produces errors 823, 824, or torn-page detection, the entries that document it may be gone before anyone investigates. &lt;code>msdb.dbo.suspect_pages&lt;/code> is where those events also land, and unlike the error log it survives restarts. Treat it as the durable forensic record of page-level I/O corruption on the instance.&lt;/p></description></item><item><title>SQL Server TDE and endpoint certificate expiry: silent AG and backup failures</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-tde-certificate-expiry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-tde-certificate-expiry/</guid><description>&lt;p>SQL Server does not alert when a certificate is about to expire. There is no performance counter, no error log entry at expiry time, and no DMV flag that flips the moment the date passes. The engine keeps using the certificate silently until something forces a re-evaluation, and at that point the failure mode depends entirely on what the certificate was protecting.&lt;/p>
&lt;p>Three consumers of the same self-signed certificate pattern behave in three different ways:&lt;/p></description></item><item><title>SQL Server TempDB file configuration: one data file per CPU and why it matters</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-tempdb-file-configuration/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-tempdb-file-configuration/</guid><description>&lt;h1 id="sql-server-tempdb-file-configuration-one-data-file-per-cpu-and-why-it-matters">SQL Server TempDB file configuration: one data file per CPU and why it matters&lt;/h1>
&lt;p>TempDB is the shared scratchpad every database on a SQL Server instance writes to. Temp tables, table variables, sort and hash spills, row version stores for RCSI and AlwaysOn readable secondaries, and internal worktables all land here. When TempDB serializes, every database on the instance serializes with it. The most common serialization point is not space or I/O. It is logical latch contention on a handful of allocation bitmap pages.&lt;/p></description></item><item><title>SQL Server TempDB full: the shared scratch database that halts every query</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-tempdb-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-tempdb-full/</guid><description>&lt;h1 id="sql-server-tempdb-full-the-shared-scratch-database-that-halts-every-query">SQL Server TempDB full: the shared scratch database that halts every query&lt;/h1>
&lt;p>Applications start failing with error 1105 (&amp;ldquo;could not allocate space for object in database &amp;rsquo;tempdb&amp;rsquo; because the &amp;lsquo;PRIMARY&amp;rsquo; filegroup is full&amp;rdquo;) or error 3958, and the failures are not limited to one database. Every query on the instance that needs a temp table, a sort or hash spill, a worktable, or a row version touches TempDB. When TempDB cannot allocate space, all of them fail at once.&lt;/p></description></item><item><title>SQL Server TempDB PAGELATCH contention: allocation page latch waits on PFS, GAM, and SGAM</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-tempdb-pagelatch-contention/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-tempdb-pagelatch-contention/</guid><description>&lt;h1 id="sql-server-tempdb-pagelatch-contention-allocation-page-latch-waits-on-pfs-gam-and-sgam">SQL Server TempDB PAGELATCH contention: allocation page latch waits on PFS, GAM, and SGAM&lt;/h1>
&lt;p>CPU is moderate, I/O latency is normal, the buffer pool is healthy, and throughput has dropped. Wait statistics show PAGELATCH_UP or PAGELATCH_EX dominating, and &lt;code>sys.dm_os_waiting_tasks&lt;/code> shows sessions waiting on pages like &lt;code>2:1:1&lt;/code>, &lt;code>2:1:2&lt;/code>, or &lt;code>2:1:3&lt;/code>.&lt;/p>
&lt;p>This is TempDB allocation page latch contention. Every temp table, spilled sort, or worktable allocation needs space in TempDB. SQL Server tracks free space and allocation state in bitmap pages: PFS (Page Free Space), GAM (Global Allocation Map), and SGAM (Shared Global Allocation Map). When sessions contend for the same allocation pages, they serialize on in-memory latches. Disk and CPU are not the bottleneck. The problem is logical contention on a fixed set of pages in the buffer pool.&lt;/p></description></item><item><title>SQL Server TempDB version store growth: long transactions under RCSI and snapshot isolation</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-tempdb-version-store-growth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-tempdb-version-store-growth/</guid><description>&lt;h1 id="sql-server-tempdb-version-store-growth-long-transactions-under-rcsi-and-snapshot-isolation">SQL Server TempDB version store growth: long transactions under RCSI and snapshot isolation&lt;/h1>
&lt;p>TempDB is filling and the version store is the consumer. You query &lt;code>tempdb.sys.dm_db_file_space_usage&lt;/code> and &lt;code>version_store_reserved_page_count&lt;/code> dominates the page budget, often after enabling Read Committed Snapshot Isolation (RCSI) or snapshot isolation, after turning a secondary replica into a readable one, or during an online index operation. Queries that depend on TempDB (sorts, hashes, temp tables, even row-versioned reads on secondaries) stall or fail when the volume runs out.&lt;/p></description></item><item><title>SQL Server THREADPOOL waits: worker thread exhaustion and refused connections</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-threadpool-waits/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-threadpool-waits/</guid><description>&lt;h1 id="sql-server-threadpool-waits-worker-thread-exhaustion-and-refused-connections">SQL Server THREADPOOL waits: worker thread exhaustion and refused connections&lt;/h1>
&lt;p>Your monitoring says the SQL Server host is fine: CPU at 15%, disk latency normal, memory steady. But the application is timing out, new connections hang, and the instance might as well be down. When you finally get in, the wait stats tell the story: THREADPOOL.&lt;/p>
&lt;p>THREADPOOL means every worker thread in the SQLOS pool is busy, and new requests are queuing for a thread that does not exist. From the client&amp;rsquo;s perspective this is equivalent to connection refusal. The server is not slow. It is not accepting work at all.&lt;/p></description></item><item><title>SQL Server transaction log percent used climbing toward full</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-log-space-used-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-log-space-used-high/</guid><description>&lt;h1 id="sql-server-transaction-log-percent-used-climbing-toward-full">SQL Server transaction log percent used climbing toward full&lt;/h1>
&lt;p>The &lt;code>Percent Log Used&lt;/code> counter on one of your databases is climbing and it is not coming back down. This is the leading gauge before Error 9002 (&amp;ldquo;The transaction log for database &amp;lsquo;X&amp;rsquo; is full&amp;rdquo;), at which point every write against that database fails. Reads may still work, which makes the outage look strange from the application side: queries succeed, inserts and updates throw errors.&lt;/p></description></item><item><title>SQL Server user connections climbing: connection pool leaks and retry storms</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-connection-count-climbing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-connection-count-climbing/</guid><description>&lt;h1 id="sql-server-user-connections-climbing-connection-pool-leaks-and-retry-storms">SQL Server user connections climbing: connection pool leaks and retry storms&lt;/h1>
&lt;p>A rising &lt;code>User Connections&lt;/code> counter is easy to misread. It does not mean that many queries are running. Most application connections are pooled and idle, so the count can climb for hours while CPU, I/O, and batch rate look almost normal.&lt;/p>
&lt;p>The useful split is shape and correlation. A slow upward trend without a matching rise in &lt;code>Batch Requests/sec&lt;/code> usually points to an application-tier connection pool leak or pool fragmentation. A sudden spike, especially after transient errors, points to a retry storm or a burst of new application instances. The dangerous endpoint is the same: enough concurrent active work consumes worker threads until new requests queue on &lt;code>THREADPOOL&lt;/code>, which is effectively connection refusal.&lt;/p></description></item><item><title>SQL Server wait statistics: reading sys.dm_os_wait_stats to find the real bottleneck</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-wait-statistics-explained/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-wait-statistics-explained/</guid><description>&lt;h1 id="sql-server-wait-statistics-reading-sysdm_os_wait_stats-to-find-the-real-bottleneck">SQL Server wait statistics: reading sys.dm_os_wait_stats to find the real bottleneck&lt;/h1>
&lt;p>Every time a SQL Server worker thread cannot proceed, the engine records what it was waiting for and for how long. The cumulative result lives in &lt;code>sys.dm_os_wait_stats&lt;/code>. When users say &amp;ldquo;the database is slow&amp;rdquo; and CPU, memory, and disk all look acceptable, wait statistics are usually where the answer is.&lt;/p>
&lt;p>The catch: the DMV is a cumulative counter since instance startup, it contains dozens of benign background waits that drown out the signal, and it tells you what the engine waited on, not which query did the waiting. Read it naively and you will chase the wrong bottleneck. Read it correctly and it decomposes performance into exactly which subsystem is contended.&lt;/p></description></item><item><title>SQL Server worker thread exhaustion: when the instance stops accepting work</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-worker-thread-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-worker-thread-exhaustion/</guid><description>&lt;h1 id="sql-server-worker-thread-exhaustion-when-the-instance-stops-accepting-work">SQL Server worker thread exhaustion: when the instance stops accepting work&lt;/h1>
&lt;p>The application reports that the database is down. CPU is low, disk I/O is low, memory looks fine, the sqlservr process is running. A TCP connection to port 1433 even succeeds. But queries hang, logins time out, and nothing completes. The instance is alive and refusing to work.&lt;/p>
&lt;p>This is worker thread exhaustion. Every active request in SQL Server needs a worker thread from a bounded pool. When every worker is occupied, usually suspended waiting on something, new requests cannot be scheduled at all. They queue on the THREADPOOL wait, which from the client&amp;rsquo;s perspective is indistinguishable from the server being down.&lt;/p></description></item><item><title>SQL Server WRITELOG waits: commit latency from a slow transaction log</title><link>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-writelog-waits-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/microsoft-sql-server/microsoft-sql-server-writelog-waits-high/</guid><description>&lt;h1 id="sql-server-writelog-waits-commit-latency-from-a-slow-transaction-log">SQL Server WRITELOG waits: commit latency from a slow transaction log&lt;/h1>
&lt;p>WRITELOG is the wait type SQL Server records when a worker thread blocks on a transaction log flush. Write-ahead logging requires the log block to be hardened to disk before the engine acknowledges a commit, so WRITELOG directly bounds write throughput. When it dominates your top-waits list, every write transaction is paying a latency tax at commit.&lt;/p>
&lt;p>The reported symptom is rarely &amp;ldquo;WRITELOG is high.&amp;rdquo; It is commit latency, write transaction timeouts, application retry storms, Availability Group replication lag, or batch requests/sec collapsing while CPU sits low. WRITELOG is the in-engine signature you find in &lt;code>sys.dm_os_wait_stats&lt;/code> or &lt;code>sys.dm_exec_session_wait_stats&lt;/code>. It tells you the bottleneck is in the log write path: the storage below the log file, the I/O stack between SQL Server and that storage, or the commit rate the Log Writer is being asked to service. Tuning queries will not fix it.&lt;/p></description></item><item><title>Squid</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/squid/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/squid/</guid><description/></item><item><title>Squid log files</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/squid-log-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/squid-log-files/</guid><description/></item><item><title>Squid Monitoring</title><link>https://www.netdata.cloud/monitoring-101/squid-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/squid-monitoring/</guid><description>&lt;h2 id="squid-monitoring">Squid Monitoring&lt;/h2>
&lt;h3 id="what-is-squid">What Is Squid?&lt;/h3>
&lt;p>Squid is a caching proxy for the web, supporting HTTP, HTTPS, FTP, and more. It reduces bandwidth and improves response times by caching and reusing frequently-requested web pages. Squid is a critical component in managing data collection and efficiency for web servers and web proxies.&lt;/p>
&lt;h3 id="monitoring-squid-with-netdata">Monitoring Squid With Netdata&lt;/h3>
&lt;p>Monitor your Squid infrastructure effectively with &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/squid/">Netdata&amp;rsquo;s Squid monitoring tool&lt;/a>. Netdata provides real-time monitoring with interactive visualizations, enabling you to keep tabs on all Squid instances whether they&amp;rsquo;re locally or remotely hosted. This allows you to instantly detect performance bottlenecks and operational issues.&lt;/p></description></item><item><title>Squid Monitoring</title><link>https://www.netdata.cloud/monitoring-101/squidlog-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/squidlog-monitoring/</guid><description>&lt;h2 id="squid-monitoring">Squid Monitoring&lt;/h2>
&lt;h3 id="what-is-squid">What Is Squid?&lt;/h3>
&lt;p>Squid is a caching and forwarding web proxy that optimizes web delivery and reduces bandwidth. By storing copies of frequently requested web content, Squid enhances response times and reduces server loads. With a rich feature set supporting HTTP, HTTPS, FTP, and more, Squid plays a crucial role in optimizing web proxies.&lt;/p>
&lt;h3 id="monitoring-squid-with-netdata">Monitoring Squid With Netdata&lt;/h3>
&lt;p>To effectively monitor Squid log files, Netdata offers an insightful solution through its go.d.plugin. The Squid monitoring tool from Netdata focuses on parsing access log files to give you real-time visibility into your Squid server operations. This ensures that you can track server responses, bandwidth usage, and client interaction in real-time.&lt;/p></description></item><item><title>Stale FDB/MAC tables: why endpoint location is wrong</title><link>https://www.netdata.cloud/guides/network/network-fdb-mac-staleness/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-fdb-mac-staleness/</guid><description>&lt;h1 id="stale-fdbmac-tables-why-endpoint-location-is-wrong">Stale FDB/MAC tables: why endpoint location is wrong&lt;/h1>
&lt;p>Your topology platform says endpoint &lt;code>aa:bb:cc:dd:ee:ff&lt;/code> is on switch port &lt;code>Gi1/0/24&lt;/code>. Your security team sends someone to that port. The endpoint is not there. It moved hours ago, or it went offline, or it vMotioned to a different host. The FDB entry was stale and the platform presented it as current.&lt;/p>
&lt;p>The Forwarding Database (FDB), also called the MAC address table or CAM table, maps MAC addresses to switch ports. Topology inference engines use FDB data, cross-referenced with ARP tables and CDP/LLDP neighbor data, to deduce where endpoints are physically connected. The inference is probabilistic. It degrades as input data freshness degrades.&lt;/p></description></item><item><title>Standard SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/standard-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/standard-snmp-traps/</guid><description/></item><item><title>Starent Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/starent-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/starent-networks-snmp-traps/</guid><description/></item><item><title>Starline Holdings SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/starline-holdings-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/starline-holdings-snmp-traps/</guid><description/></item><item><title>Starlink (SpaceX)</title><link>https://www.netdata.cloud/integrations/data-collection/networking/starlink-spacex/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/starlink-spacex/</guid><description/></item><item><title>Starlink (SpaceX) Monitoring</title><link>https://www.netdata.cloud/monitoring-101/starlink-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/starlink-monitoring/</guid><description>&lt;h2 id="starlink-monitoring">Starlink Monitoring&lt;/h2>
&lt;h3 id="what-is-starlink-spacex">What Is Starlink (SpaceX)?&lt;/h3>
&lt;p>Starlink, developed by SpaceX, is an ambitious project aimed at providing high-speed satellite internet across the globe. By deploying a constellation of small satellites in low Earth orbit, Starlink aims to deliver reliable internet connectivity even in the most remote areas. With its advanced technology, monitoring is crucial to maintain and optimize internet service management and performance.&lt;/p>
&lt;h3 id="monitoring-starlink-with-netdata">Monitoring Starlink With Netdata&lt;/h3>
&lt;p>To effectively monitor Starlink, Netdata utilizes an openmetrics (Prometheus) exporter. The &lt;a href="https://github.com/danopstech/starlink_exporter">Starlink Exporter&lt;/a> integrates seamlessly with Netdata, allowing for the collection and visualization of Starlink satellite internet metrics. Netdata excels in aggregating data from any Prometheus exporter, offering automated dashboards and alerts without the need for a dedicated Prometheus server or Grafana. This streamlined process ensures real-time and comprehensive monitoring of your Starlink connection.&lt;/p></description></item><item><title>Static Metadata</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/static-metadata/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/network-flows/enrichment-methods/static-metadata/</guid><description/></item><item><title>StatusPage</title><link>https://www.netdata.cloud/integrations/data-collection/applications/statuspage/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/statuspage/</guid><description/></item><item><title>StatusPage Monitoring</title><link>https://www.netdata.cloud/monitoring-101/statuspage-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/statuspage-monitoring/</guid><description>&lt;h2 id="statuspage-monitoring">StatusPage Monitoring&lt;/h2>
&lt;h3 id="what-is-statuspage">What Is StatusPage?&lt;/h3>
&lt;p>StatusPage is an effective tool for incident management and communication. It allows you to inform your users about the status of your services, incidents in real time, and updates about resolutions. With StatusPage, you have a centralized method to keep your operations transparent. Understanding how to monitor it effectively can significantly improve your incident response times and service communication.&lt;/p>
&lt;h3 id="monitoring-statuspage-with-netdata">Monitoring StatusPage With Netdata&lt;/h3>
&lt;p>Monitoring StatusPage with Netdata is both simple and powerful. Utilize the StatusPage Exporter, available &lt;a href="https://github.com/vladvasiliu/statuspage-exporter">here&lt;/a>, to transform StatusPage data into openmetrics compatible format. Netdata can ingest metrics from any Prometheus exporter, allowing you to monitor StatusPage without needing to set up a Prometheus server or Grafana. This means you get out-of-the-box dashboards, alerts, and visualization for immediate insights, all available on the same network without additional configurations.&lt;/p></description></item><item><title>Steam</title><link>https://www.netdata.cloud/integrations/data-collection/applications/steam/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/steam/</guid><description/></item><item><title>Steam Monitoring</title><link>https://www.netdata.cloud/monitoring-101/steam_a2s-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/steam_a2s-monitoring/</guid><description>&lt;h2 id="steam-monitoring">Steam Monitoring&lt;/h2>
&lt;h3 id="what-is-steam">What Is Steam?&lt;/h3>
&lt;p>Steam is a leading platform for digital distribution of video games and related media, a community space for gamers, and a tool for multiplayer gaming. With millions of concurrent users, monitoring Steam servers is crucial for maintaining seamless gaming experiences.&lt;/p>
&lt;h3 id="monitoring-steam-with-netdata">Monitoring Steam With Netdata&lt;/h3>
&lt;p>To efficiently monitor Steam, Netdata employs an openmetrics (Prometheus) exporter, specifically tailored to the Steam A2S-supported game servers. This allows you to gather insightful metrics on performance and availability. Netdata&amp;rsquo;s versatility shines through as it can ingest data from any Prometheus exporter, providing automated dashboards, alerts, and more—eliminating the need for a standalone Prometheus server or Grafana. Discover the ease and efficiency with which you can monitor Steam with &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata&amp;rsquo;s Live Demo&lt;/a>.&lt;/p></description></item><item><title>Stonesoft Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/stonesoft-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/stonesoft-corp-snmp-traps/</guid><description/></item><item><title>Storage Computer Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/storage-computer-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/storage-computer-corporation-snmp-traps/</guid><description/></item><item><title>Storage Networking Industry Association SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/storage-networking-industry-association-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/storage-networking-industry-association-snmp-traps/</guid><description/></item><item><title>Storage Technology Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/storage-technology-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/storage-technology-corporation-snmp-traps/</guid><description/></item><item><title>StoreCLI RAID</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/storecli-raid/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/storecli-raid/</guid><description/></item><item><title>StoreCLI RAID Monitoring</title><link>https://www.netdata.cloud/monitoring-101/storcli-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/storcli-monitoring/</guid><description>&lt;h2 id="storecli-raid-monitoring">StoreCLI RAID Monitoring&lt;/h2>
&lt;h3 id="what-is-storecli-raid">What Is StoreCLI RAID?&lt;/h3>
&lt;p>StoreCLI RAID is a powerful tool that manages and monitors your RAID (Redundant Array of Independent Disks) setup. It allows you to oversee the performance and health of your storage system by tracking RAID adapters, physical drives, and backup batteries using the StorCLI command line interface. For more details, refer to the official &lt;a href="https://docs.broadcom.com/doc/12352476">StoreCLI RAID documentation&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-storecli-raid-with-netdata">Monitoring StoreCLI RAID With Netdata&lt;/h3>
&lt;p>Netdata enhances the way you monitor StoreCLI RAID by enabling real-time insights to help you understand your storage system&amp;rsquo;s health and performance. Using Netdata, you can leverage a seamless experience to track your RAID arrays without executing binaries directly, improving both security and ease of use. Interested in trying it out? &lt;a href="https://app.netdata.cloud/?utm_source=website&amp;amp;utm_content=monitoring101">Sign up for a Free Trial&lt;/a>.&lt;/p></description></item><item><title>Storidge</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/storidge/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/storidge/</guid><description/></item><item><title>Storidge Monitoring</title><link>https://www.netdata.cloud/monitoring-101/storidge-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/storidge-monitoring/</guid><description>&lt;h2 id="storidge-monitoring">Storidge Monitoring&lt;/h2>
&lt;h3 id="what-is-storidge">What Is Storidge?&lt;/h3>
&lt;p>Storidge is an innovative solution designed to simplify storage management and deliver high-performance, scalable storage services within enterprise environments. By offering exceptional operational efficiency, Storidge aims to streamline the storage process and enhance data handling capabilities.&lt;/p>
&lt;h3 id="monitoring-storidge-with-netdata">Monitoring Storidge With Netdata&lt;/h3>
&lt;p>To monitor Storidge effectively, Netdata leverages an openmetrics (Prometheus) exporter. Netdata is a robust Storidge monitoring tool that can seamlessly ingest data from any Prometheus exporter, allowing users to visualize metrics through automated dashboards and receive real-time alerts. What sets Netdata apart is that you can achieve all of this without the need for a separate Prometheus server or Grafana setup, simplifying the monitoring journey for DevOps, SREs, developers, and IT admins.&lt;/p></description></item><item><title>Stormshield Formerly Netasq SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/stormshield-formerly-netasq-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/stormshield-formerly-netasq-snmp-traps/</guid><description/></item><item><title>STP Topology</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/stp-topology/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/stp-topology/</guid><description/></item><item><title>STP topology-change storms: reconvergence cascades explained</title><link>https://www.netdata.cloud/guides/network/network-stp-topology-change-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-stp-topology-change-storm/</guid><description>&lt;h1 id="stp-topology-change-storms-reconvergence-cascades-explained">STP topology-change storms: reconvergence cascades explained&lt;/h1>
&lt;p>A topology-change notification (TCN) is not itself a failure. STP generates one every time a non-edge port transitions up or down. That is normal during maintenance, link recovery, or device boot. The problem is what happens next. When TCNs fire repeatedly, or when a single TCN hits a large Layer 2 domain with thousands of MAC addresses, the protocol&amp;rsquo;s designed response becomes a self-inflicted traffic event.&lt;/p></description></item><item><title>Stratacom SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/stratacom-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/stratacom-snmp-traps/</guid><description/></item><item><title>Stratus Computer SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/stratus-computer-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/stratus-computer-snmp-traps/</guid><description/></item><item><title>Streamcore SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/streamcore-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/streamcore-snmp-traps/</guid><description/></item><item><title>strongSwan</title><link>https://www.netdata.cloud/integrations/data-collection/networking/strongswan/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/strongswan/</guid><description/></item><item><title>strongSwan Monitoring</title><link>https://www.netdata.cloud/monitoring-101/strongswan-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/strongswan-monitoring/</guid><description>&lt;h2 id="strongswan-monitoring">strongSwan Monitoring&lt;/h2>
&lt;h3 id="what-is-strongswan">What Is strongSwan?&lt;/h3>
&lt;p>strongSwan is an open-source implementation of IPSec, a key protocol used to secure virtual private networks (VPNs). It’s distinctive for its focus on encryption standards, scalability, and cross-platform capability. strongSwan enables developers and IT admins to create secure network connections and enforce security policies efficiently.&lt;/p>
&lt;h3 id="monitoring-strongswan-with-netdata">Monitoring strongSwan With Netdata&lt;/h3>
&lt;p>To effectively monitor strongSwan, Netdata utilizes an openmetrics exporter, specifically structured to integrate seamlessly with any Prometheus exporter, including the &lt;a href="https://github.com/jlti-dev/ipsec_exporter">strongSwan/IPSec/vici Exporter&lt;/a>. This powerful capability ensures users can collect VPN performance metrics and manage system health without requiring a centralized Prometheus server or Grafana setup. By incorporating Netdata, IT engineers can automatically generate dashboards and receive proactive alerts, significantly simplifying the tasks of monitoring and troubleshooting.&lt;/p></description></item><item><title>Stulz GmbH Klimatechnik SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/stulz-gmbh-klimatechnik-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/stulz-gmbh-klimatechnik-snmp-traps/</guid><description/></item><item><title>Sub10 Systems Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sub10-systems-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sub10-systems-ltd-snmp-traps/</guid><description/></item><item><title>Sunspec Solar Energy</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/sunspec-solar-energy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/sunspec-solar-energy/</guid><description/></item><item><title>Sunspec Solar Energy Monitoring</title><link>https://www.netdata.cloud/monitoring-101/sunspec-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/sunspec-monitoring/</guid><description>&lt;h2 id="sunspec-solar-energy-monitoring">Sunspec Solar Energy Monitoring&lt;/h2>
&lt;h3 id="what-is-sunspec-solar-energy">What Is Sunspec Solar Energy?&lt;/h3>
&lt;p>Sunspec Solar Energy refers to the standardized protocols and metrics developed by the SunSpec Alliance to ensure interoperability in monitoring and managing solar energy systems. These standards are critical for efficient solar energy management, enabling seamless communication between devices, improving data accuracy, and enhancing system reliability.&lt;/p>
&lt;h3 id="monitoring-sunspec-solar-energy-with-netdata">Monitoring Sunspec Solar Energy With Netdata&lt;/h3>
&lt;p>To monitor Sunspec Solar Energy effectively, Netdata leverages an openmetrics (Prometheus) exporter. This means that Netdata can gather and visualize metrics from the &lt;a href="https://github.com/inosion/prometheus-sunspec-exporter">Sunspec Solar Energy Exporter&lt;/a> without needing a separate Prometheus server or Grafana dashboard setup. Netdata&amp;rsquo;s powerful monitoring capabilities allow users to create automated dashboards and set up alerts, providing real-time insights into their solar energy systems.&lt;/p></description></item><item><title>Super Micro Computer Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/super-micro-computer-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/super-micro-computer-inc-snmp-traps/</guid><description/></item><item><title>Supervisor</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/supervisor/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/supervisor/</guid><description/></item><item><title>Supervisor Monitoring</title><link>https://www.netdata.cloud/monitoring-101/supervisord-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/supervisord-monitoring/</guid><description>&lt;h2 id="supervisor-monitoring">Supervisor Monitoring&lt;/h2>
&lt;h3 id="what-is-supervisor">What Is Supervisor?&lt;/h3>
&lt;p>Supervisor is a client/server system that allows its users to monitor and control a number of processes on UNIX-like operating systems. Its primary function is to ensure that processes start, restart, and run as they should, providing the user with feedback about the status and error messages of the monitored programs.&lt;/p>
&lt;h3 id="monitoring-supervisor-with-netdata">Monitoring Supervisor With Netdata&lt;/h3>
&lt;p>Netdata offers a comprehensive Supervisor monitoring tool that provides deep insights into your system&amp;rsquo;s processes. Using this tool allows you to gain real-time visibility into all processes managed by Supervisor, ensuring seamless operation and quick troubleshooting of any irregularities. For full documentation, visit our &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/supervisord/">Supervisor collector documentation&lt;/a>.&lt;/p></description></item><item><title>Suricata</title><link>https://www.netdata.cloud/integrations/data-collection/applications/suricata/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/suricata/</guid><description/></item><item><title>Suricata Monitoring</title><link>https://www.netdata.cloud/monitoring-101/suricata-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/suricata-monitoring/</guid><description>&lt;h2 id="suricata-monitoring">Suricata Monitoring&lt;/h2>
&lt;h3 id="what-is-suricata">What Is Suricata?&lt;/h3>
&lt;p>Suricata is an open-source network threat detection engine that features intrusion detection (IDS), intrusion prevention (IPS), and network security monitoring capabilities. It analyzes network traffic and identifies suspicious activities by utilizing data collection methods, such as deep packet inspection and pattern matching.&lt;/p>
&lt;h3 id="monitoring-suricata-with-netdata">Monitoring Suricata With Netdata&lt;/h3>
&lt;p>Monitoring Suricata with Netdata offers an unparalleled view into your network&amp;rsquo;s security apparatus. Netdata utilizes an openmetrics (Prometheus) exporter, the &lt;a href="https://github.com/corelight/suricata_exporter">Suricata Exporter&lt;/a>, to gather metrics efficiently. Unlike traditional setups requiring a Prometheus server or Grafana for display, Netdata handles it all seamlessly. It ingests data from any Prometheus exporter, automatically presenting intuitive dashboards, real-time alerts, and in-depth analyses without the complexity typically involved.&lt;/p></description></item><item><title>SUSE Linux</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/suse-linux/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/suse-linux/</guid><description/></item><item><title>Swapcom SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/swapcom-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/swapcom-snmp-traps/</guid><description/></item><item><title>Swichtec Power Systems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/swichtec-power-systems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/swichtec-power-systems-snmp-traps/</guid><description/></item><item><title>Symantec Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/symantec-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/symantec-corporation-snmp-traps/</guid><description/></item><item><title>Symbol Technologies Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/symbol-technologies-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/symbol-technologies-inc-snmp-traps/</guid><description/></item><item><title>Symmetricom SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/symmetricom-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/symmetricom-snmp-traps/</guid><description/></item><item><title>Symplex Communications Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/symplex-communications-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/symplex-communications-corp-snmp-traps/</guid><description/></item><item><title>Synaccess Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/synaccess-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/synaccess-networks-inc-snmp-traps/</guid><description/></item><item><title>Synamedia SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/synamedia-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/synamedia-snmp-traps/</guid><description/></item><item><title>Sync Research Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sync-research-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/sync-research-inc-snmp-traps/</guid><description/></item><item><title>Synernetics Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/synernetics-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/synernetics-inc-snmp-traps/</guid><description/></item><item><title>Synology ActiveBackup</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/synology-activebackup/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/synology-activebackup/</guid><description/></item><item><title>Synology ActiveBackup Monitoring</title><link>https://www.netdata.cloud/monitoring-101/synology_activebackup-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/synology_activebackup-monitoring/</guid><description>&lt;h2 id="synology-activebackup-monitoring">Synology ActiveBackup Monitoring&lt;/h2>
&lt;h3 id="what-is-synology-activebackup">What Is Synology ActiveBackup?&lt;/h3>
&lt;p>Synology ActiveBackup is a comprehensive data backup solution designed for efficient data protection and management. This tool allows IT professionals to back up critical data across multiple environments, ensuring file security and rapid recovery capabilities, crucial for any business continuity plan.&lt;/p>
&lt;h3 id="monitoring-synology-activebackup-with-netdata">Monitoring Synology ActiveBackup With Netdata&lt;/h3>
&lt;p>When it comes to monitor Synology ActiveBackup, Netdata provides an effortless and robust solution. Netdata uses an openmetrics (Prometheus) exporter to collect detailed metrics from Synology ActiveBackup. Unlike traditional methods requiring standalone Prometheus server installations and extensive configurations with Grafana, Netdata simplifies this by ingesting data directly from any Prometheus exporter, providing automated dashboards and alerting mechanisms. This streamlined process ensures you can focus more on analyzing the metrics rather than setting up complex monitoring infrastructure.&lt;/p></description></item><item><title>Synology Disk Station</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/synology-disk-station/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/synology-disk-station/</guid><description/></item><item><title>Synoptics SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/synoptics-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/synoptics-snmp-traps/</guid><description/></item><item><title>Synproxy</title><link>https://www.netdata.cloud/integrations/data-collection/networking/synproxy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/synproxy/</guid><description/></item><item><title>Synso Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/synso-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/synso-inc-snmp-traps/</guid><description/></item><item><title>Sysload</title><link>https://www.netdata.cloud/integrations/data-collection/applications/sysload/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/sysload/</guid><description/></item><item><title>Sysload Monitoring</title><link>https://www.netdata.cloud/monitoring-101/sysload-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/sysload-monitoring/</guid><description>&lt;h2 id="sysload-monitoring">Sysload Monitoring&lt;/h2>
&lt;h3 id="what-is-sysload">What Is Sysload?&lt;/h3>
&lt;p>Sysload is a monitoring solution designed to collect and analyze system load metrics. It provides detailed insights into the performance and resource management of your systems, essential for efficient system operations.&lt;/p>
&lt;h3 id="monitoring-sysload-with-netdata">Monitoring Sysload With Netdata&lt;/h3>
&lt;p>Netdata offers a powerful and seamless way to monitor Sysload using the openmetrics (Prometheus) exporter. This integration allows you to monitor Sysload effectively without needing a dedicated Prometheus server or Grafana setup. By ingesting data from any Prometheus exporter, Netdata automatically creates dashboards and alerts, allowing you to stay on top of your system&amp;rsquo;s performance metrics effortlessly. To get started, you only need to configure the Sysload Exporter and Netdata will handle the rest.&lt;/p></description></item><item><title>syslog</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/syslog/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/syslog/</guid><description/></item><item><title>Syslog from Network Devices</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/syslog-from-network-devices/syslog-from-network-devices/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/syslog-from-network-devices/syslog-from-network-devices/</guid><description/></item><item><title>Syslog parser backpressure: when one chatty device stalls the pipeline</title><link>https://www.netdata.cloud/guides/network/network-syslog-parser-backpressure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-syslog-parser-backpressure/</guid><description>&lt;h1 id="syslog-parser-backpressure-when-one-chatty-device-stalls-the-pipeline">Syslog parser backpressure: when one chatty device stalls the pipeline&lt;/h1>
&lt;p>A single device floods your syslog collector. The parser thread pool saturates, queues fill, and UDP datagrams start dropping at the kernel socket buffer. Critical messages from other devices, including BGP NOTIFICATIONS and hardware alarms, are silently lost. The dashboard shows a normal or slightly elevated syslog rate because dropped packets never reach the application layer.&lt;/p>
&lt;p>The collector process is still running. The network is fine. The failure is inside the ingestion pipeline, at the seam between the kernel socket buffer and the parser, where backpressure builds and has nowhere to go.&lt;/p></description></item><item><title>System Engineering International SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/system-engineering-international-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/system-engineering-international-snmp-traps/</guid><description/></item><item><title>System Load Average</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/system-load-average/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/system-load-average/</guid><description/></item><item><title>System Management Arts Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/system-management-arts-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/system-management-arts-inc-snmp-traps/</guid><description/></item><item><title>System Memory Fragmentation</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/system-memory-fragmentation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/system-memory-fragmentation/</guid><description/></item><item><title>System statistics</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/system-statistics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/system-statistics/</guid><description/></item><item><title>System thermal zone</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/system-thermal-zone/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/system-thermal-zone/</guid><description/></item><item><title>System Uptime</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/system-uptime/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/system-uptime/</guid><description/></item><item><title>system.ram</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/system.ram/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/system.ram/</guid><description/></item><item><title>Systemd Journal Logs</title><link>https://www.netdata.cloud/integrations/logs/systemd-journal-logs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/logs/systemd-journal-logs/</guid><description/></item><item><title>Systemd Services</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/systemd-services/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/systemd-services/</guid><description/></item><item><title>Systemd Units</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/systemd-units/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/systemd-units/</guid><description/></item><item><title>Systemd Units Monitoring</title><link>https://www.netdata.cloud/monitoring-101/systemdunits-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/systemdunits-monitoring/</guid><description>&lt;h2 id="systemd-units-monitoring">Systemd Units Monitoring&lt;/h2>
&lt;h3 id="what-is-systemd-units">What Is Systemd Units?&lt;/h3>
&lt;p>Systemd Units represent the entities managed by systemd, a powerful init system and service manager for Linux operating systems. It is at the core of various Linux distributions and offers a suite of functionalities for managing system services. Understanding and monitoring Systemd Units is crucial for ensuring the health and performance of your operating systems as they define how a service is started, stopped, and managed.&lt;/p></description></item><item><title>systemd-logind users</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/systemd-logind-users/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/systemd-logind-users/</guid><description/></item><item><title>systemd-logind users Monitoring</title><link>https://www.netdata.cloud/monitoring-101/logind-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/logind-monitoring/</guid><description>&lt;h2 id="systemd-logind-users-monitoring">systemd-logind users Monitoring&lt;/h2>
&lt;h3 id="what-is-systemd-logind-users">What Is systemd-logind users?&lt;/h3>
&lt;p>Systemd-logind is a system service in Linux operating systems that manages user logins. This service is part of the larger systemd suite, which is essential for resource management, session tracking, and user processes control. Monitoring systemd-logind users involves tracking active sessions and users, ensuring seamless operation across different systems.&lt;/p>
&lt;h3 id="monitoring-systemd-logind-users-with-netdata">Monitoring systemd-logind users With Netdata&lt;/h3>
&lt;p>Monitoring systemd-logind users is made efficient and straightforward with Netdata. Netdata’s real-time monitoring capabilities provide comprehensive insights into session activities and user states, making it an ideal systemd-logind users monitoring tool. By utilizing Netdata, you can gain instant visualization of key metrics and receive alerts on any unusual activities.&lt;/p></description></item><item><title>systemd-nspawn Containers</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/systemd-nspawn-containers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/systemd-nspawn-containers/</guid><description/></item><item><title>TACACS</title><link>https://www.netdata.cloud/integrations/data-collection/applications/tacacs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/tacacs/</guid><description/></item><item><title>TACACS Monitoring</title><link>https://www.netdata.cloud/monitoring-101/tacas-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/tacas-monitoring/</guid><description>&lt;h2 id="tacacs-monitoring">TACACS Monitoring&lt;/h2>
&lt;h3 id="what-is-tacacs">What Is TACACS?&lt;/h3>
&lt;p>The Terminal Access Controller Access-Control System (TACACS) is a protocol developed for network authentication and authorization management. It&amp;rsquo;s widely used to manage authentication tasks and provide detailed accounting information. TACACS plays a crucial role in securing network environments by efficiently tracking access, thus maintaining the integrity of IT infrastructures.&lt;/p>
&lt;h3 id="monitoring-tacacs-with-netdata">Monitoring TACACS With Netdata&lt;/h3>
&lt;p>To monitor TACACS, Netdata leverages a robust openmetrics (Prometheus) exporter. This setup allows Netdata to seamlessly ingest data from any Prometheus exporter. Users can enjoy automated dashboards and alerts, all without the necessity of maintaining a separate Prometheus server or Grafana instance. This integration simplifies the monitoring process while providing real-time insights into your TACACS system.&lt;/p></description></item><item><title>Tado smart heating solution</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/tado-smart-heating-solution/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/tado-smart-heating-solution/</guid><description/></item><item><title>Tado Smart Heating Solution Monitoring</title><link>https://www.netdata.cloud/monitoring-101/tado-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/tado-monitoring/</guid><description>&lt;h2 id="tado-smart-heating-solution-monitoring">Tado Smart Heating Solution Monitoring&lt;/h2>
&lt;h3 id="what-is-tado-smart-heating-solution">What Is Tado Smart Heating Solution?&lt;/h3>
&lt;p>Tado is a cutting-edge smart heating solution designed to optimize the efficiency and comfort of home heating and cooling management. By connecting thermostats to an intelligent, internet-based platform, Tado provides users with precise control over their home temperatures, adapting dynamically to the needs of the household.&lt;/p>
&lt;h3 id="monitoring-tado-smart-heating-solution-with-netdata">Monitoring Tado Smart Heating Solution With Netdata&lt;/h3>
&lt;p>Monitoring Tado smart heating solution is a critical aspect of ensuring your home stays comfortable and energy-efficient. For seamless monitoring, Netdata uses an openmetrics (Prometheus) exporter. The &lt;a href="https://github.com/eko/tado-exporter">Tado Exporter&lt;/a> is employed to collect various metrics necessary for maintaining an efficient Tado environment.&lt;/p></description></item><item><title>Tail F Systems AB SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tail-f-systems-ab-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tail-f-systems-ab-snmp-traps/</guid><description/></item><item><title>Tailyn Communication Company SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tailyn-communication-company-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tailyn-communication-company-snmp-traps/</guid><description/></item><item><title>Tait International Limited SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tait-international-limited-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tait-international-limited-snmp-traps/</guid><description/></item><item><title>Talari Networks SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/talari-networks-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/talari-networks-snmp-traps/</guid><description/></item><item><title>Talk to a member of our Sales team!</title><link>https://www.netdata.cloud/contact-sales/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/contact-sales/</guid><description/></item><item><title>Talk to a member of our Sales team!</title><link>https://www.netdata.cloud/request-enterprise/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/request-enterprise/</guid><description/></item><item><title>Tandberg Television SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tandberg-television-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tandberg-television-snmp-traps/</guid><description/></item><item><title>Tandem Computers SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tandem-computers-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tandem-computers-snmp-traps/</guid><description/></item><item><title>Tankerkoenig API</title><link>https://www.netdata.cloud/integrations/data-collection/applications/tankerkoenig-api/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/tankerkoenig-api/</guid><description/></item><item><title>Tankerkoenig API Monitoring</title><link>https://www.netdata.cloud/monitoring-101/tankerkoenig-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/tankerkoenig-monitoring/</guid><description>&lt;h2 id="tankerkoenig-api-monitoring">Tankerkoenig API Monitoring&lt;/h2>
&lt;h3 id="what-is-tankerkoenig-api">What Is Tankerkoenig API?&lt;/h3>
&lt;p>Tankerkoenig API is a powerful tool that provides fuel price data, allowing developers and organizations to collect real-time and historical data on fuel prices across various regions. It&amp;rsquo;s crucial for businesses that rely on fuel data for logistics, analysis, or reporting to stay informed and make data-driven decisions.&lt;/p>
&lt;h3 id="monitoring-tankerkoenig-api-with-netdata">Monitoring Tankerkoenig API With Netdata&lt;/h3>
&lt;p>Netdata offers an elegant solution for real-time monitoring of the Tankerkoenig API. By leveraging an openmetrics (Prometheus) exporter, Netdata is capable of seamlessly ingesting data from any Prometheus exporter to provide automated dashboards, alerts, and comprehensive insights, all without the need for a separate Prometheus server or Grafana. With the &lt;a href="https://github.com/lukasmalkmus/tankerkoenig_exporter">Tankerkoenig API exporter&lt;/a>, integrating and monitoring your Tankerkoenig API becomes more efficient and hassle-free.&lt;/p></description></item><item><title>Tasman Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tasman-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tasman-networks-inc-snmp-traps/</guid><description/></item><item><title>Tavve Software Co SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tavve-software-co-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tavve-software-co-snmp-traps/</guid><description/></item><item><title>tc QoS classes</title><link>https://www.netdata.cloud/integrations/data-collection/networking/tc-qos-classes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/tc-qos-classes/</guid><description/></item><item><title>TCP endpoints Monitoring</title><link>https://www.netdata.cloud/monitoring-101/tcpendpoints-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/tcpendpoints-monitoring/</guid><description>&lt;h2 id="what-is-tcp-endpoint">What is TCP endpoint?&lt;/h2>
&lt;p>A TCP endpoint is a combination of an IP address and a port number that identifies a specific process or service running on a computer or other network device. It is used to establish and manage an end-to-end connection between two applications, typically over the Internet. The TCP protocol provides a reliable, ordered delivery of data between the two endpoints.&lt;/p>
&lt;h2 id="monitoring-tcp-endpoint-with-netdata">Monitoring TCP endpoint with Netdata&lt;/h2>
&lt;p>The prerequisites for monitoring TCP endpoint is to have &lt;a href="https://learn.netdata.cloud/docs/cloud/get-started/">Netdata installed&lt;/a> on your system.&lt;/p></description></item><item><title>TCP/UDP Endpoints</title><link>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/tcp-udp-endpoints/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/tcp-udp-endpoints/</guid><description/></item><item><title>TCP/UDP Endpoints Monitoring</title><link>https://www.netdata.cloud/monitoring-101/portcheck-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/portcheck-monitoring/</guid><description>&lt;h2 id="tcpudp-endpoints-monitoring">TCP/UDP Endpoints Monitoring&lt;/h2>
&lt;h3 id="what-is-tcpudp-endpoints">What Is TCP/UDP Endpoints?&lt;/h3>
&lt;p>Understanding TCP/UDP Endpoints involves recognizing the critical roles these protocols play in network communications. TCP stands for Transmission Control Protocol, managing data delivery through acknowledgment and retransmission processes. Conversely, UDP, the User Datagram Protocol, is connectionless, meaning it caters to applications where speed trumps reliability.&lt;/p>
&lt;h3 id="monitoring-tcpudp-endpoints-with-netdata">Monitoring TCP/UDP Endpoints With Netdata&lt;/h3>
&lt;p>When it comes to efficient and real-time monitoring, Netdata offers a profound solution. The &lt;a href="https://app.netdata.cloud/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata platform&lt;/a> monitors TCP/UDP endpoints efficiently, helping you track the availability and responsiveness of your network services. Netdata’s &lt;code>portcheck&lt;/code> monitoring tool checks the availability of specific ports, which is crucial for assessing service health.&lt;/p></description></item><item><title>Tecnair S P A SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tecnair-s-p-a-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tecnair-s-p-a-snmp-traps/</guid><description/></item><item><title>Tecnopro S.A. SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tecnopro-s.a.-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tecnopro-s.a.-snmp-traps/</guid><description/></item><item><title>Tegile Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tegile-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tegile-systems-inc-snmp-traps/</guid><description/></item><item><title>Telco Systems Nac SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/telco-systems-nac-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/telco-systems-nac-snmp-traps/</guid><description/></item><item><title>Teldat S A SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/teldat-s-a-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/teldat-s-a-snmp-traps/</guid><description/></item><item><title>Telefonaktiebolaget Lm Ericsson SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/telefonaktiebolaget-lm-ericsson-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/telefonaktiebolaget-lm-ericsson-snmp-traps/</guid><description/></item><item><title>Telegram</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/telegram/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/telegram/</guid><description/></item><item><title>Telegram</title><link>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/telegram/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/telegram/</guid><description/></item><item><title>Telematics International Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/telematics-international-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/telematics-international-inc-snmp-traps/</guid><description/></item><item><title>Telesend Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/telesend-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/telesend-inc-snmp-traps/</guid><description/></item><item><title>Teleste Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/teleste-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/teleste-corporation-snmp-traps/</guid><description/></item><item><title>Telesystems Slw Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/telesystems-slw-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/telesystems-slw-inc-snmp-traps/</guid><description/></item><item><title>Televes S A SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/televes-s-a-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/televes-s-a-snmp-traps/</guid><description/></item><item><title>Television Systems Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/television-systems-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/television-systems-ltd-snmp-traps/</guid><description/></item><item><title>Telstrat International Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/telstrat-international-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/telstrat-international-ltd-snmp-traps/</guid><description/></item><item><title>Teltonika SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/teltonika-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/teltonika-snmp-traps/</guid><description/></item><item><title>Temperature, fan, and PSU monitoring: predicting hardware failure</title><link>https://www.netdata.cloud/guides/network/network-device-environment-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-device-environment-monitoring/</guid><description>&lt;h1 id="temperature-fan-and-psu-monitoring-predicting-hardware-failure">Temperature, fan, and PSU monitoring: predicting hardware failure&lt;/h1>
&lt;p>Environmental sensors on network devices are the earliest leading indicators of hardware failure. Temperature trends, fan state changes, and PSU status transitions often precede field-replaceable unit failures by hours or days. The data is not hard to collect, but the MIB landscape is fragmented across vendors, thresholds vary by platform, and inherited polling templates frequently target deprecated OIDs. A template that worked on a Catalyst 3560 can silently return nothing on a Catalyst 8500.&lt;/p></description></item><item><title>Tengine</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/tengine/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/tengine/</guid><description/></item><item><title>Tengine Monitoring</title><link>https://www.netdata.cloud/monitoring-101/tengine-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/tengine-monitoring/</guid><description>&lt;h2 id="tengine-monitoring">Tengine Monitoring&lt;/h2>
&lt;h3 id="what-is-tengine">What Is Tengine?&lt;/h3>
&lt;p>Tengine is a high-performance web server and reverse proxy, based on Nginx, engineered for greater scalability and enhanced security. It is widely used in large-scale websites to manage traffic, reduce load, and ensure improved web performance. Learn more about Tengine on &lt;a href="https://tengine.taobao.org/">their official site&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-tengine-with-netdata">Monitoring Tengine With Netdata&lt;/h3>
&lt;p>Monitoring Tengine is crucial for DevOps and IT administrators seeking to optimize web performance. Netdata offers a robust Tengine monitoring tool that provides real-time insights into your server’s performance, enabling prompt issue diagnosis and efficient resource management. By utilizing Netdata&amp;rsquo;s comprehensive monitoring capabilities, you can seamlessly monitor Tengine instances, from basic network metrics to advanced performance statistics.&lt;/p></description></item><item><title>Teracom Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/teracom-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/teracom-ltd-snmp-traps/</guid><description/></item><item><title>Teracom Telematica Ltda SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/teracom-telematica-ltda-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/teracom-telematica-ltda-snmp-traps/</guid><description/></item><item><title>Terms of Service</title><link>https://www.netdata.cloud/service-terms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/service-terms/</guid><description/></item><item><title>Terms of Use</title><link>https://www.netdata.cloud/terms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/terms/</guid><description/></item><item><title>Terra SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/terra-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/terra-snmp-traps/</guid><description/></item><item><title>Tesla vehicle</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/tesla-vehicle/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/tesla-vehicle/</guid><description/></item><item><title>Tesla Vehicle Monitoring</title><link>https://www.netdata.cloud/monitoring-101/tesla_vehicle-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/tesla_vehicle-monitoring/</guid><description>&lt;h2 id="tesla-vehicle-monitoring">Tesla Vehicle Monitoring&lt;/h2>
&lt;h3 id="what-is-tesla-vehicle-monitoring">What Is Tesla Vehicle Monitoring?&lt;/h3>
&lt;p>Tesla vehicle monitoring involves tracking various metrics essential for managing and optimizing the performance of electric vehicles. These metrics can include battery health, energy consumption, charging status, and more, providing valuable insights for efficient vehicle management.&lt;/p>
&lt;h3 id="monitoring-tesla-vehicles-with-netdata">Monitoring Tesla Vehicles With Netdata&lt;/h3>
&lt;p>To monitor Tesla vehicles, Netdata employs an openmetrics (Prometheus) exporter. This enables the collection of key metrics without the need for a dedicated Prometheus server or Grafana. With Netdata, you can ingest data from any Prometheus exporter, receiving automated dashboards, alerts, and anomaly detection tailored to Tesla vehicle metrics. By using the &lt;a href="https://github.com/wywywywy/tesla-prometheus-exporter">Tesla Prometheus Exporter&lt;/a>, you can seamlessly integrate and visualize important data, optimizing the use and management of your Tesla vehicles.&lt;/p></description></item><item><title>Tesla Wall Connector</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/tesla-wall-connector/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/tesla-wall-connector/</guid><description/></item><item><title>Tesla Wall Connector Monitoring</title><link>https://www.netdata.cloud/monitoring-101/tesla_wall_connector-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/tesla_wall_connector-monitoring/</guid><description>&lt;h2 id="tesla-wall-connector-monitoring">Tesla Wall Connector Monitoring&lt;/h2>
&lt;h3 id="what-is-tesla-wall-connector">What Is Tesla Wall Connector?&lt;/h3>
&lt;p>Tesla Wall Connector is an advanced charging station specifically designed for Tesla electric vehicles, allowing users to efficiently manage their vehicle&amp;rsquo;s charging. As electric vehicle adoption grows, ensuring that charging infrastructure is reliable and efficient becomes crucial. Monitoring the Tesla Wall Connector becomes an essential part of maintaining an optimal charging setup.&lt;/p>
&lt;h3 id="monitoring-tesla-wall-connector-with-netdata">Monitoring Tesla Wall Connector With Netdata&lt;/h3>
&lt;p>To effectively monitor the Tesla Wall Connector, Netdata utilizes an OpenMetrics (Prometheus) exporter. The &lt;a href="https://github.com/benclapp/tesla_wall_connector_exporter">Tesla Wall Connector Exporter&lt;/a> collects a wide range of metrics that can be processed by Netdata&amp;rsquo;s state-of-the-art monitoring solutions. With Netdata, you can ingest data from any Prometheus exporter, granting you access to automated dashboards, real-time alerts, and insightful analytics without needing a Prometheus server or Grafana configuration. View Netdata&amp;rsquo;s capabilities in this &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">live demo&lt;/a>.&lt;/p></description></item><item><title>Thales E Security SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/thales-e-security-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/thales-e-security-snmp-traps/</guid><description/></item><item><title>Thank you for your interest.</title><link>https://www.netdata.cloud/thanks-for-reaching-out/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/thanks-for-reaching-out/</guid><description/></item><item><title>Thanos</title><link>https://www.netdata.cloud/integrations/exporters/thanos/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/thanos/</guid><description/></item><item><title>The Advantage Group SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/the-advantage-group-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/the-advantage-group-snmp-traps/</guid><description/></item><item><title>The Freeradius Server Project SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/the-freeradius-server-project-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/the-freeradius-server-project-snmp-traps/</guid><description/></item><item><title>The SSD write cliff: latency spikes when the pre-erased block pool empties</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-ssd-write-cliff/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-ssd-write-cliff/</guid><description>&lt;h1 id="the-ssd-write-cliff-latency-spikes-when-the-pre-erased-block-pool-empties">The SSD write cliff: latency spikes when the pre-erased block pool empties&lt;/h1>
&lt;p>Write latency on an SSD jumps from sub-millisecond to hundreds of milliseconds or seconds. Applications time out. The kernel log may show command timeouts. But &lt;code>smartctl -H&lt;/code> says PASSED, every SMART attribute looks clean, and the drive has plenty of endurance left.&lt;/p>
&lt;p>The write cliff happens when an SSD exhausts its pool of pre-erased NAND blocks. Under normal conditions, the controller performs garbage collection (GC) in the background: it reads valid pages from partially invalidated blocks, writes them elsewhere, and erases the now-empty block to replenish the free pool. When the write rate outpaces background GC, the free pool empties. Every host write now requires a synchronous erase cycle before it can complete, and latency explodes by 100x to 1000x.&lt;/p></description></item><item><title>The zombie drive: bad sectors, read retries, and high iowait with idle CPU</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-zombie-drive-bad-sectors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-zombie-drive-bad-sectors/</guid><description>&lt;h1 id="the-zombie-drive-bad-sectors-read-retries-and-high-iowait-with-idle-cpu">The zombie drive: bad sectors, read retries, and high iowait with idle CPU&lt;/h1>
&lt;p>The system is grinding to a halt. SSH sessions lag, commands take seconds to return, and applications time out. You check &lt;code>top&lt;/code> or &lt;code>htop&lt;/code> and CPU is nearly idle, yet load average is climbing. &lt;code>iostat -x 1&lt;/code> tells the real story: &lt;code>await&lt;/code> on one disk is spiking into the hundreds of milliseconds, sometimes multiple seconds, while every other disk looks fine. The process stuck on that I/O is in D state (uninterruptible sleep), unkillable, waiting for a read that the drive cannot complete.&lt;/p></description></item><item><title>TiKV</title><link>https://www.netdata.cloud/integrations/exporters/tikv/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/tikv/</guid><description/></item><item><title>Timeplex SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/timeplex-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/timeplex-snmp-traps/</guid><description/></item><item><title>TimescaleDB</title><link>https://www.netdata.cloud/integrations/exporters/timescaledb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/timescaledb/</guid><description/></item><item><title>Timesten Performance Software SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/timesten-performance-software-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/timesten-performance-software-snmp-traps/</guid><description/></item><item><title>Timestep Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/timestep-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/timestep-corp-snmp-traps/</guid><description/></item><item><title>Timex</title><link>https://www.netdata.cloud/integrations/data-collection/networking/timex/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/timex/</guid><description/></item><item><title>Tintri Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tintri-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tintri-inc-snmp-traps/</guid><description/></item><item><title>Tomcat</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/tomcat/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/tomcat/</guid><description/></item><item><title>Tomcat 'appears to have started a thread but has failed to stop it': undeploy leak warnings</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-webapp-failed-to-stop-thread/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-webapp-failed-to-stop-thread/</guid><description>&lt;h1 id="tomcat-appears-to-have-started-a-thread-but-has-failed-to-stop-it-undeploy-leak-warnings">Tomcat &amp;lsquo;appears to have started a thread but has failed to stop it&amp;rsquo;: undeploy leak warnings&lt;/h1>
&lt;p>The exact log line is the symptom. When you undeploy or redeploy a Tomcat webapp and see this in &lt;code>catalina.out&lt;/code>:&lt;/p>
&lt;pre tabindex="0">&lt;code>SEVERE [main] org.apache.catalina.loader.WebappClassLoaderBase.clearReferencesThreads
The web application [myapp] appears to have started a thread named [AbandonedConnectionCleanupThread]
but has failed to stop it. This is very likely to create a memory leak.
&lt;/code>&lt;/pre>&lt;p>that is not generic noise. During Context stop, &lt;code>WebappClassLoaderBase.clearReferencesThreads()&lt;/code> walked every live thread in the JVM, found at least one whose context classloader is still the webapp&amp;rsquo;s &lt;code>WebappClassLoader&lt;/code>, and named it. The bracketed thread name is the breadcrumb back to the library or application code that started it.&lt;/p></description></item><item><title>Tomcat 401 flood on the Manager app: credential brute force and LockOutRealm</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-401-manager-brute-force/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-401-manager-brute-force/</guid><description>&lt;h1 id="tomcat-401-flood-on-the-manager-app-credential-brute-force-and-lockoutrealm">Tomcat 401 flood on the Manager app: credential brute force and LockOutRealm&lt;/h1>
&lt;p>A 401 flood is a sustained spike of HTTP 401 responses against &lt;code>/manager&lt;/code>, &lt;code>/host-manager&lt;/code>, or any realm-secured endpoint. The signature is concentration: many 401s from a small set of source IPs, or many 401s aimed at a single username. Scattered 401s from mistyped passwords are noise; concentration and sustain are the signal.&lt;/p>
&lt;p>Against the Manager app, a 401 flood is almost always credential brute force or stuffing and a precursor to compromise. A successful Manager login lets an attacker deploy WAR files (RCE), enumerate sessions, and undeploy applications. An internet-exposed Manager with weak or default credentials can be fully taken over within minutes of the flood beginning.&lt;/p></description></item><item><title>Tomcat 5xx error rate: separating server failures from crawler 404s</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-5xx-error-rate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-5xx-error-rate/</guid><description>&lt;h1 id="tomcat-5xx-error-rate-separating-server-failures-from-crawler-404s">Tomcat 5xx error rate: separating server failures from crawler 404s&lt;/h1>
&lt;p>Your Tomcat error rate alert fired. The dashboard shows a spike in &lt;code>errorCount&lt;/code>. The JVM is healthy, the thread pool has headroom, heap is fine. The access log tells a different story: 404s from a crawler storm, not 500s from a failing app. The page was a false alarm.&lt;/p>
&lt;p>This is the Tomcat error-rate trap. The JMX &lt;code>errorCount&lt;/code> attribute exposed via &lt;code>Catalina:type=GlobalRequestProcessor,name=&amp;quot;http-nio-8080&amp;quot;&lt;/code> counts every response with status &amp;gt;= 400. It does not distinguish 4xx from 5xx. A crawler hitting dead URLs inflates it identically to an application throwing unhandled exceptions. If you alert on this counter as a ratio of total requests, crawler noise will page you.&lt;/p></description></item><item><title>Tomcat accept queue overflow: acceptCount, somaxconn, and Recv-Q</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-accept-queue-overflow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-accept-queue-overflow/</guid><description>&lt;h1 id="tomcat-accept-queue-overflow-acceptcount-somaxconn-and-recv-q">Tomcat accept queue overflow: acceptCount, somaxconn, and Recv-Q&lt;/h1>
&lt;p>Clients start seeing &amp;ldquo;connection refused&amp;rdquo; or TCP timeouts, but the Tomcat JVM looks healthy. The thread pool has capacity. Heap is stable. JMX shows nothing wrong. The manager status page reports normal request counts. The application logs are quiet.&lt;/p>
&lt;p>The problem is in a place Tomcat cannot see: the OS-level TCP accept queue, also called the listen backlog. When this queue fills, the kernel refuses new connections by sending RST or silently dropping the completed handshake. No JMX counter tracks this. No Tomcat log records it. The only direct signal is &lt;code>ss&lt;/code> output showing &lt;code>Recv-Q&lt;/code> climbing toward &lt;code>Send-Q&lt;/code> on the listening socket.&lt;/p></description></item><item><title>Tomcat accepts connections but never responds: the TCP-connect trap</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-connector-not-responding/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-connector-not-responding/</guid><description>&lt;h1 id="tomcat-accepts-connections-but-never-responds-the-tcp-connect-trap">Tomcat accepts connections but never responds: the TCP-connect trap&lt;/h1>
&lt;p>A client opens a TCP connection to Tomcat on port 8080. The three-way handshake completes. The client sends an HTTP request. Nothing comes back. The connection hangs until the client times out. Meanwhile, the load balancer health check still reports the instance as healthy, because its check is a bare TCP connect that succeeds every time.&lt;/p>
&lt;p>A successful &lt;code>connect()&lt;/code> only proves the kernel accepted the socket. It says nothing about whether Tomcat has a worker thread available, whether the JVM is mid-garbage-collection, or whether any deployed application can serve a response. On NIO Tomcat, the default connector since 8.5, the OS accept queue and the connector poller can hold thousands of connections even when the worker thread pool is completely exhausted. A port check is nearly useless as a health signal.&lt;/p></description></item><item><title>Tomcat access log setup: adding %D and %T for per-request latency</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-access-log-response-time/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-access-log-response-time/</guid><description>&lt;h1 id="tomcat-access-log-setup-adding-d-and-t-for-per-request-latency">Tomcat access log setup: adding %D and %T for per-request latency&lt;/h1>
&lt;p>Tomcat&amp;rsquo;s &lt;code>common&lt;/code> and &lt;code>combined&lt;/code> access log patterns record the request line, status, and byte count. They do not record per-request latency. Without per-request timing in the access log, you are limited to the cumulative average that JMX exposes via &lt;code>processingTime / requestCount&lt;/code> on the &lt;code>GlobalRequestProcessor&lt;/code> MBean. That average hides the tail: a few 15-second requests averaged against a thousand 10-millisecond requests looks fine while real users time out.&lt;/p></description></item><item><title>Tomcat activeSessions growing without plateau: the session leak</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-session-count-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-session-count-growing/</guid><description>&lt;h1 id="tomcat-activesessions-growing-without-plateau-the-session-leak">Tomcat activeSessions growing without plateau: the session leak&lt;/h1>
&lt;p>The chart shows &lt;code>activeSessions&lt;/code> from &lt;code>Catalina:type=Manager,host=localhost,context=/&amp;lt;app&amp;gt;&lt;/code> climbing in a near-straight line. No plateau. You restarted Tomcat this morning, and by midafternoon the count is already back to where it was before the restart. Sessions are being created faster than they expire, and with the default 30-minute &lt;code>session-timeout&lt;/code>, expiry is slow enough that the slope looks gentle until you compare it to the request rate.&lt;/p></description></item><item><title>Tomcat average latency lies: why you need p95/p99 from the access log</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-average-vs-percentile-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-average-vs-percentile-latency/</guid><description>&lt;h1 id="tomcat-average-latency-lies-why-you-need-p95p99-from-the-access-log">Tomcat average latency lies: why you need p95/p99 from the access log&lt;/h1>
&lt;p>Your latency dashboard says Tomcat is healthy at 250ms. Your users say requests take 15 seconds. Both are right. The JMX average your dashboard consumes is structurally incapable of representing the tail.&lt;/p>
&lt;p>&lt;code>Catalina:type=GlobalRequestProcessor&lt;/code> exposes &lt;code>processingTime&lt;/code> (cumulative milliseconds since startup) and &lt;code>requestCount&lt;/code>. Divide the deltas over a window and you get a mean. A workload where p50 is 100ms but p99 is 15s produces an average of roughly 250ms. 1% of your users wait an eternity, but the chart looks fine. Stuck threads averaged into fast requests are the classic blind spot, and Tomcat gives you nothing from JMX that breaks the mean open.&lt;/p></description></item><item><title>Tomcat classloader leak on redeploy: why the old WebappClassLoader never dies</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-classloader-leak-on-redeploy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-classloader-leak-on-redeploy/</guid><description>&lt;h1 id="tomcat-classloader-leak-on-redeploy-why-the-old-webappclassloader-never-dies">Tomcat classloader leak on redeploy: why the old WebappClassLoader never dies&lt;/h1>
&lt;p>Every hot redeploy leaves the previous &lt;code>WebappClassLoader&lt;/code> in memory. After enough redeploys the JVM hits &lt;code>OutOfMemoryError: Metaspace&lt;/code>, or if &lt;code>-XX:MaxMetaspaceSize&lt;/code> is unset, the OS OOM-kills the process with no JVM error at all. Heap looks flat. GC looks fine. RSS climbs until the process dies.&lt;/p>
&lt;p>Each Tomcat webapp gets its own &lt;code>WebappClassLoader&lt;/code> for isolation. On undeploy, the classloader and every class it loaded should be collected. If anything still holds a reference to a class loaded by the webapp, a &lt;code>ThreadLocal&lt;/code> value, a JDBC driver registered with &lt;code>DriverManager&lt;/code>, a &lt;code>Timer&lt;/code> thread that was never cancelled, a log appender, or a static field in a shared library, the classloader is pinned. All its classes stay in Metaspace forever.&lt;/p></description></item><item><title>Tomcat connection refused: maxConnections and acceptCount both exhausted</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-connection-refused/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-connection-refused/</guid><description>&lt;h1 id="tomcat-connection-refused-maxconnections-and-acceptcount-both-exhausted">Tomcat connection refused: maxConnections and acceptCount both exhausted&lt;/h1>
&lt;p>Clients are getting &amp;ldquo;connection refused&amp;rdquo; or connections hanging until timeout. The Tomcat JVM is up, the HTTP port is bound, heap looks fine, and there is nothing in catalina.out. Manager status shows the connector alive. From Tomcat&amp;rsquo;s perspective, nothing is wrong. From the kernel&amp;rsquo;s perspective, the listening socket&amp;rsquo;s accept queue is full and new SYNs are being dropped or RST&amp;rsquo;d.&lt;/p>
&lt;p>This happens when two limits stack: the NIO poller has hit &lt;code>maxConnections&lt;/code> (default 8192 for NIO) and the OS accept queue, bounded by &lt;code>acceptCount&lt;/code> (default 100) and clamped by &lt;code>net.core.somaxconn&lt;/code>, is also full. No JMX counter exposes this state. Tomcat logs nothing. Detection is either client-side connection error monitoring or &lt;code>ss -tnl&lt;/code> showing Recv-Q stuck at the accept queue limit.&lt;/p></description></item><item><title>Tomcat context in FAILED state: the app is down but Tomcat looks healthy</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-webapp-failed-to-start/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-webapp-failed-to-start/</guid><description>&lt;h1 id="tomcat-context-in-failed-state-the-app-is-down-but-tomcat-looks-healthy">Tomcat context in FAILED state: the app is down but Tomcat looks healthy&lt;/h1>
&lt;p>Your monitoring says Tomcat is up. The JVM process is running, the HTTP connector accepts TCP connections, and a curl to port 8080 returns an HTTP response. But every request to the application returns 404, and users cannot reach any endpoint.&lt;/p>
&lt;p>This is the signature of a context in FAILED or STOPPED state. Tomcat continues running. The connectors keep listening. Other deployed contexts keep serving. But the failed context&amp;rsquo;s servlet mappings are inactive, so every request to it returns 404. From the outside, it looks like a missing route or an undeployed application.&lt;/p></description></item><item><title>Tomcat file descriptor usage: OpenFileDescriptorCount vs the ulimit</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-file-descriptor-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-file-descriptor-exhaustion/</guid><description>&lt;h1 id="tomcat-file-descriptor-usage-openfiledescriptorcount-vs-the-ulimit">Tomcat file descriptor usage: OpenFileDescriptorCount vs the ulimit&lt;/h1>
&lt;p>Tomcat&amp;rsquo;s JVM holds file descriptors for every open socket, every log file, every JAR on the classpath, the NIO selector itself, and various internal pipes. The pool is finite and process-scoped. When it runs out, Tomcat stops accepting connections and stops writing to logs in the same instant. The error is &lt;code>java.net.SocketException: Too many open files&lt;/code>, and by the time you see it the failure is already total.&lt;/p></description></item><item><title>Tomcat frequent Full GC: pause time, G1, and the 5% overhead rule</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-full-gc-frequency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-full-gc-frequency/</guid><description>&lt;h1 id="tomcat-frequent-full-gc-pause-time-g1-and-the-5-overhead-rule">Tomcat frequent Full GC: pause time, G1, and the 5% overhead rule&lt;/h1>
&lt;p>A Tomcat instance that was steady for hours or days starts pausing. Latency p99 climbs, request throughput drops in bursts, and the JVM is alive but barely making progress. The thread pool is not the bottleneck: the wall clock is being eaten by garbage collection. In the GC log you see &amp;ldquo;Pause Full&amp;rdquo; lines arriving every few seconds, each stopping the world for hundreds of milliseconds to multiple seconds.&lt;/p></description></item><item><title>Tomcat GC death spiral: full GCs dominating and throughput collapsing</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-gc-death-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-gc-death-spiral/</guid><description>&lt;p>Your Tomcat JVM is up, the connector port is bound, but requests crawl or hang. Throughput graphs stutter: brief bursts of activity separated by flatlines where no requests complete. CPU is pinned near saturation.&lt;/p>
&lt;p>The mechanism is a positive feedback loop. Live data on the heap has grown until each GC cycle reclaims almost nothing. The JVM compensates by running GC more frequently, which consumes more CPU, which leaves less wall clock for application threads, which causes requests to take longer, which causes more concurrent threads to stay active, which allocates more memory. The loop ends in &lt;code>OutOfMemoryError&lt;/code> or effective livelock where GC runs continuously and the application makes no progress.&lt;/p></description></item><item><title>Tomcat Ghostcat (CVE-2020-1938): AJP connector exposure and arbitrary file read</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-ghostcat-ajp-cve-2020-1938/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-ghostcat-ajp-cve-2020-1938/</guid><description>&lt;h1 id="tomcat-ghostcat-cve-2020-1938-ajp-connector-exposure-and-arbitrary-file-read">Tomcat Ghostcat (CVE-2020-1938): AJP connector exposure and arbitrary file read&lt;/h1>
&lt;p>Ghostcat (CVE-2020-1938) is a configuration-driven vulnerability in Apache Tomcat&amp;rsquo;s AJP connector. Before the 9.0.31 / 8.5.51 / 7.0.100 fixes, the AJP protocol carried no authentication. Any client that could reach port 8009 could craft AJP requests that made Tomcat read and return arbitrary files from the server, including &lt;code>WEB-INF/web.xml&lt;/code> and &lt;code>server.xml&lt;/code>. In some configurations the same primitive enables remote code execution through JSP inclusion.&lt;/p></description></item><item><title>Tomcat heap dump before restart: capturing evidence with jmap and jstack</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-heap-dump-before-restart/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-heap-dump-before-restart/</guid><description>&lt;h1 id="tomcat-heap-dump-before-restart-capturing-evidence-with-jmap-and-jstack">Tomcat heap dump before restart: capturing evidence with jmap and jstack&lt;/h1>
&lt;p>When Tomcat is failing, the instinct is to restart. That instinct is right for recovery and wrong for diagnosis. The JVM process holds the only copy of the evidence: the object graph that explains the heap exhaustion, the thread stack traces that explain the pool stall, the GC state that explains the death spiral. Once the process exits, that evidence is gone. You are left restarting blind with nothing but an &lt;code>OutOfMemoryError&lt;/code> line in &lt;code>catalina.out&lt;/code>.&lt;/p></description></item><item><title>Tomcat heap usage: watch the post-GC baseline, not the sawtooth peak</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-heap-post-gc-baseline/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-heap-post-gc-baseline/</guid><description>&lt;h1 id="tomcat-heap-usage-watch-the-post-gc-baseline-not-the-sawtooth-peak">Tomcat heap usage: watch the post-GC baseline, not the sawtooth peak&lt;/h1>
&lt;p>Every JVM heap graph looks the same: a jagged sawtooth that climbs steadily, drops sharply, and repeats. If you alert on &amp;ldquo;heap &amp;gt; 80%&amp;rdquo;, you will page on every pre-GC peak. That alert fires dozens of times per hour on a healthy Tomcat, training your team to ignore heap warnings until a real &lt;code>OutOfMemoryError&lt;/code> arrives and nobody saw it coming.&lt;/p></description></item><item><title>Tomcat hot redeploy vs clean restart: when autoDeploy leaks and when it doesn't</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-hot-redeploy-vs-restart/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-hot-redeploy-vs-restart/</guid><description>&lt;h1 id="tomcat-hot-redeploy-vs-clean-restart-when-autodeploy-leaks-and-when-it-doesnt">Tomcat hot redeploy vs clean restart: when autoDeploy leaks and when it doesn&amp;rsquo;t&lt;/h1>
&lt;p>When you deploy a new WAR, Tomcat gives you two paths. Let &lt;code>autoDeploy&lt;/code> watch &lt;code>webapps/&lt;/code> and swap the application in place, or stop the JVM and start a clean one. Both deliver the new code. They differ in what they leave behind, measured in Metaspace.&lt;/p>
&lt;p>The root cause is the &lt;code>WebappClassLoader&lt;/code>. Each Context gets its own. On a hot redeploy the old one is supposed to be garbage collected. Often it is not. A lingering &lt;code>ThreadLocal&lt;/code>, a &lt;code>DriverManager&lt;/code> registration, a timer thread, or a static field can pin the old classloader and every class it loaded. On a clean restart the JVM dies, and every classloader dies with it. There is nothing to pin.&lt;/p></description></item><item><title>Tomcat HTTP Status 503 Service Unavailable: the connector is out of threads</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-http-503-service-unavailable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-http-503-service-unavailable/</guid><description>&lt;h1 id="tomcat-http-status-503-service-unavailable-the-connector-is-out-of-threads">Tomcat HTTP Status 503 Service Unavailable: the connector is out of threads&lt;/h1>
&lt;p>A 503 in a Tomcat topology means something between the user and the servlet gave up. The instinct is to blame the connector thread pool, and that is often the right place to look, but the mechanism is more layered than it appears. Treating &amp;ldquo;503&amp;rdquo; and &amp;ldquo;out of threads&amp;rdquo; as the same condition leads to misdiagnosis.&lt;/p>
&lt;p>With the default NIO connector, Tomcat decouples worker threads from TCP connections. A request thread pool can be saturated while the connector keeps accepting connections, and the connection pool can be saturated while the JVM still passes a TCP health check. The path from a full thread pool to a visible 503 passes through two more buffers before any client sees a failure.&lt;/p></description></item><item><title>Tomcat java.lang.OutOfMemoryError: Java heap space: the heap is genuinely full</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-outofmemoryerror-java-heap-space/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-outofmemoryerror-java-heap-space/</guid><description>&lt;h1 id="tomcat-javalangoutofmemoryerror-java-heap-space-the-heap-is-genuinely-full">Tomcat java.lang.OutOfMemoryError: Java heap space: the heap is genuinely full&lt;/h1>
&lt;p>The string &lt;code>java.lang.OutOfMemoryError: Java heap space&lt;/code> in &lt;code>catalina.out&lt;/code> means the JVM could not satisfy an allocation because live (reachable) objects plus the requested size exceeded the maximum heap (&lt;code>-Xmx&lt;/code>). This is not a transient GC spike or a young-gen promotion failure. The heap is genuinely full of data the collector cannot reclaim.&lt;/p>
&lt;p>Tomcat&amp;rsquo;s heap holds HTTP sessions, request/response buffers, application objects, and in-process caches. When live data outgrows &lt;code>-Xmx&lt;/code>, GC frequency rises, each cycle reclaims less, and application threads get less time between collections. Eventually an allocation fails and the JVM throws &lt;code>OutOfMemoryError: Java heap space&lt;/code>. The process may limp on in a degraded state, serving some requests and failing others, until restarted.&lt;/p></description></item><item><title>Tomcat java.lang.OutOfMemoryError: Metaspace: the classloader leak hot redeploys cause</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-outofmemoryerror-metaspace/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-outofmemoryerror-metaspace/</guid><description>&lt;p>A &lt;code>java.lang.OutOfMemoryError: Metaspace&lt;/code> after a redeploy is the canonical classloader leak symptom in long-running Tomcat instances. The JVM still has heap headroom, GC looks healthy, and the process may have been up for weeks. Then a deploy lands and Tomcat dies, or in containerized setups the kernel OOM-kills it with no JVM-level error at all.&lt;/p>
&lt;p>Recovery is the same in every case: restart the JVM. That clears Metaspace and brings the application back, but it does not fix anything. If you hot-deploy again, the leak returns and Metaspace climbs one step higher per cycle until the next crash.&lt;/p></description></item><item><title>Tomcat java.net.BindException: Address already in use: the connector never starts</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-address-already-in-use/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-address-already-in-use/</guid><description>&lt;h1 id="tomcat-javanetbindexception-address-already-in-use-the-connector-never-starts">Tomcat java.net.BindException: Address already in use: the connector never starts&lt;/h1>
&lt;p>You restart Tomcat. The JVM comes up. The process check goes green. Then &lt;code>catalina.out&lt;/code> shows:&lt;/p>
&lt;pre tabindex="0">&lt;code>SEVERE [main] org.apache.catalina.util.LifecycleBase.handleSubClassException Failed to initialize component [Connector[HTTP/1.1-8080]]
java.net.BindException: Address already in use
&lt;/code>&lt;/pre>&lt;p>The JVM is alive, but the HTTP connector never bound its port. No traffic is served. Health checks that only verify the PID report healthy while every client gets connection refused.&lt;/p>
&lt;p>This is the &amp;ldquo;Connector Binding vs Process Alive&amp;rdquo; trap: the JVM stays up after a connector init failure, so process-based monitoring passes while nothing listens. The failure modes are narrow. A stale Tomcat still holding 8080 or 8443, a port conflict with another service, or a previous shutdown that did not release the socket. Each has a different fix, and conflating them produces the wrong remediation.&lt;/p></description></item><item><title>Tomcat java.net.SocketException: Too many open files: file descriptor exhaustion</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-too-many-open-files/</guid><description>&lt;h1 id="tomcat-javanetsocketexception-too-many-open-files-file-descriptor-exhaustion">Tomcat java.net.SocketException: Too many open files: file descriptor exhaustion&lt;/h1>
&lt;p>&lt;code>java.net.SocketException: Too many open files&lt;/code> in a Tomcat log is a hard cliff, not gradual degradation. The JVM goes from serving traffic to unable to accept a new TCP connection, open a log file, or load a JAR resource. The error is frequently misread as a disk fault (log writes break) or a network fault (&lt;code>accept()&lt;/code> fails), when the real constraint is the per-process file descriptor limit.&lt;/p></description></item><item><title>Tomcat JDBC connection leak: numActive stuck at max even when idle</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-jdbc-connection-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-jdbc-connection-leak/</guid><description>&lt;h1 id="tomcat-jdbc-connection-leak-numactive-stuck-at-max-even-when-idle">Tomcat JDBC connection leak: numActive stuck at max even when idle&lt;/h1>
&lt;p>The signature of a Tomcat JDBC connection leak is specific: &lt;code>numActive&lt;/code> sits at &lt;code>maxActive&lt;/code> and refuses to fall, even when request rate is near zero. New requests that need a database connection block on &lt;code>getConnection()&lt;/code> until &lt;code>maxWait&lt;/code> elapses, then throw a pool exhaustion error. The database reports a flock of idle &lt;code>Sleep&lt;/code> connections from Tomcat that never close.&lt;/p>
&lt;p>The cascade is what takes the service down. Each request blocking on &lt;code>getConnection()&lt;/code> also holds an HTTP worker thread. As leaked connections accumulate, more threads stall on the pool, the thread pool fills, and requests queue in the acceptor. The JVM is healthy, the database is healthy, CPU is low, and the service is effectively down.&lt;/p></description></item><item><title>Tomcat JDBC connection pool exhaustion: numActive at maxActive and threads blocking</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-jdbc-connection-pool-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-jdbc-connection-pool-exhaustion/</guid><description>&lt;h1 id="tomcat-jdbc-connection-pool-exhaustion-numactive-at-maxactive-and-threads-blocking">Tomcat JDBC connection pool exhaustion: numActive at maxActive and threads blocking&lt;/h1>
&lt;p>HTTP requests start timing out, &lt;code>currentThreadsBusy&lt;/code> climbs toward &lt;code>maxThreads&lt;/code>, CPU stays low, and the database looks healthy. The Tomcat JDBC pool has hit its ceiling: &lt;code>numActive&lt;/code> equals &lt;code>maxActive&lt;/code>, &lt;code>waitCount&lt;/code> is greater than zero, and every request that needs a database connection is parked in &lt;code>getConnection()&lt;/code>.&lt;/p>
&lt;p>The JDBC pool and the HTTP worker thread pool are coupled. A request holds a worker thread for its entire duration. When it blocks in &lt;code>getConnection()&lt;/code>, that worker stays occupied. Database pool exhaustion drives thread pool exhaustion, and thread pool exhaustion drives accept queue buildup and 503s. By the time users see errors, both resources are already exhausted.&lt;/p></description></item><item><title>Tomcat keepalive and connectionTimeout: coordinating timers with your reverse proxy</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-keepalive-connection-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-keepalive-connection-tuning/</guid><description>&lt;h1 id="tomcat-keepalive-and-connectiontimeout-coordinating-timers-with-your-reverse-proxy">Tomcat keepalive and connectionTimeout: coordinating timers with your reverse proxy&lt;/h1>
&lt;p>When Tomcat sits behind nginx, HAProxy, Apache httpd, or a cloud load balancer, the proxy-to-Tomcat connection is almost always persistent keep-alive. Both ends maintain idle-timeout timers that independently decide when to close. If those timers are not coordinated, the proxy periodically tries to reuse a connection Tomcat has already closed, and the user gets a 502 or a connection reset.&lt;/p></description></item><item><title>Tomcat Manager app exposed: default credentials, WAR upload, and RCE</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-manager-app-exposed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-manager-app-exposed/</guid><description>&lt;h1 id="tomcat-manager-app-exposed-default-credentials-war-upload-and-rce">Tomcat Manager app exposed: default credentials, WAR upload, and RCE&lt;/h1>
&lt;p>The Tomcat Manager and Host Manager web applications are administrative interfaces bundled with standalone Tomcat distributions. The Manager app deploys and undeploys WAR files, lists and invalidates HTTP sessions, reloads applications, and exposes server status. The deploy capability is a direct path to remote code execution: an attacker uploads a WAR containing a webshell or reverse-shell payload, Tomcat auto-deploys it, and the code runs inside the JVM with the filesystem and network access of the Tomcat process.&lt;/p></description></item><item><title>Tomcat maxConnections saturation: the NIO poller stops accepting</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-maxconnections-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-maxconnections-saturation/</guid><description>&lt;h1 id="tomcat-maxconnections-saturation-the-nio-poller-stops-accepting">Tomcat maxConnections saturation: the NIO poller stops accepting&lt;/h1>
&lt;p>Clients report connection timeouts or &amp;ldquo;connection refused&amp;rdquo; errors. The Tomcat JVM is running, the HTTP port is bound, GC looks normal, and the worker thread pool may be mostly idle. What you are looking at is maxConnections saturation: the NIO poller has filled to its configured ceiling and the acceptor has stopped registering new sockets.&lt;/p>
&lt;p>The default maxConnections for NIO is 8192 (since Tomcat 9.0.30; it was 10000 for NIO on older 8.5.x and early 9.x releases). When connectionCount approaches that number, Tomcat stops accepting new connections until existing ones close. New SYNs pile into the OS TCP backlog, bounded by acceptCount (default 100). When that queue fills, the kernel sends RST and clients see a hard refusal.&lt;/p></description></item><item><title>Tomcat MaxMetaspaceSize unset: the silent OS OOM-kill with no Java error</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-maxmetaspacesize-not-set/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-maxmetaspacesize-not-set/</guid><description>&lt;h1 id="tomcat-maxmetaspacesize-unset-the-silent-os-oom-kill-with-no-java-error">Tomcat MaxMetaspaceSize unset: the silent OS OOM-kill with no Java error&lt;/h1>
&lt;p>The Tomcat process is gone. &lt;code>catalina.out&lt;/code> ends mid-line or shows a normal shutdown request that never executed. There is no &lt;code>OutOfMemoryError&lt;/code>, no &lt;code>HeapDumpOnOutOfMemoryError&lt;/code> file, and no JFR recording. The systemd unit reports &lt;code>code=killed, status=9/KILL&lt;/code> or simply a vanished PID. Your heap dashboard looked healthy right up to the moment the process disappeared.&lt;/p>
&lt;p>This is the silent Metaspace kill. The JVM never throws a Java-level error because the kernel terminates it first. When &lt;code>-XX:MaxMetaspaceSize&lt;/code> is left at its default (effectively unlimited), class metadata lives in native memory that is not bounded by &lt;code>-Xmx&lt;/code>, not covered by heap alerts, and not reclaimed when the heap looks fine. Metaspace grows until process RSS hits the OS limit or the container cgroup limit, and the Linux OOM-killer ends the JVM with SIGKILL.&lt;/p></description></item><item><title>Tomcat maxThreads and minSpareThreads: sizing the executor correctly</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-maxthreads-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-maxthreads-tuning/</guid><description>&lt;h1 id="tomcat-maxthreads-and-minsparethreads-sizing-the-executor-correctly">Tomcat maxThreads and minSpareThreads: sizing the executor correctly&lt;/h1>
&lt;p>The &lt;code>&amp;lt;Executor&amp;gt;&lt;/code> thread pool, or the connector-level pool when no explicit executor is referenced, is the most critical bounded resource in a Tomcat instance. &lt;code>maxThreads&lt;/code> sets the ceiling on concurrent request processing. &lt;code>minSpareThreads&lt;/code> sets the floor of pre-warmed threads. Every active HTTP request occupies one worker thread for its entire processing duration, unless you use async servlets. When the pool is full, connections keep arriving but requests wait.&lt;/p></description></item><item><title>Tomcat Monitoring</title><link>https://www.netdata.cloud/monitoring-101/tomcat-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/tomcat-monitoring/</guid><description>&lt;h2 id="tomcat-monitoring">Tomcat Monitoring&lt;/h2>
&lt;h3 id="what-is-tomcat">What Is Tomcat?&lt;/h3>
&lt;p>Tomcat, a robust servlet container by Apache, is a critical component enabling Java-based web applications to run seamlessly. As a cornerstone in the Java EE ecosystem, it facilitates the execution of servlets and Java Server Pages (JSP) and supports the deployment of dynamic applications. For more details, visit the &lt;a href="https://tomcat.apache.org/">Apache Tomcat official website&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-tomcat-with-netdata">Monitoring Tomcat With Netdata&lt;/h3>
&lt;p>Monitoring Tomcat effectively requires a comprehensive toolset, and Netdata stands out by offering real-time, extensive insights into Tomcat&amp;rsquo;s performance metrics. By leveraging &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/tomcat/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata&amp;rsquo;s Tomcat monitoring tool&lt;/a>, you can gain visibility into various operational metrics like bandwidth, threads, and processing time. Netdata enables seamless integration with your existing systems, providing a holistic view of your infrastructure.&lt;/p></description></item><item><title>Tomcat monitoring checklist: the signals every production instance needs</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-monitoring-checklist/</guid><description>&lt;h1 id="tomcat-monitoring-checklist-the-signals-every-production-instance-needs">Tomcat monitoring checklist: the signals every production instance needs&lt;/h1>
&lt;p>Tomcat&amp;rsquo;s capacity model rests on a small number of bounded resources: a worker thread pool, a connection poller, the JVM heap, Metaspace, and the OS file descriptor table. Most production outages are one of these hitting its limit while the JVM process keeps running. A monitoring setup that only tracks CPU and memory will miss the dominant failure mode: thread pool exhaustion with the JVM at low CPU and a healthy process.&lt;/p></description></item><item><title>Tomcat monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-monitoring-maturity-model/</guid><description>&lt;h1 id="tomcat-monitoring-maturity-model-from-survival-to-expert">Tomcat monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most Tomcat monitoring setups start with a process check, a health endpoint probe, and maybe a 5xx alert. That catches crashes. It does not catch the failures that actually page teams: thread pool exhaustion, GC death spirals, accept queue overflow, classloader leaks after hot redeploys. This article maps the signals that matter across four maturity levels so you can audit what you have, identify what you are missing, and prioritize what to add next.&lt;/p></description></item><item><title>Tomcat NoClassDefFoundError after redeploy: stale classloaders serving half-loaded classes</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-noclassdeffounderror-after-redeploy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-noclassdeffounderror-after-redeploy/</guid><description>&lt;h1 id="tomcat-noclassdeffounderror-after-redeploy-stale-classloaders-serving-half-loaded-classes">Tomcat NoClassDefFoundError after redeploy: stale classloaders serving half-loaded classes&lt;/h1>
&lt;p>A hot redeploy on Tomcat succeeds, the new context reports STARTED, and within minutes requests start failing with &lt;code>java.lang.NoClassDefFoundError&lt;/code> for classes that are visibly present in the new WAR. Sometimes it is &lt;code>ClassNotFoundException&lt;/code>. Sometimes the error names a class that was renamed or removed between versions. Restarting the JVM clears it. The next redeploy brings it back.&lt;/p>
&lt;p>The root mechanism is the &lt;code>WebappClassLoader&lt;/code> lifecycle. Tomcat creates a new classloader for each web application context and replaces it on every redeploy. The old classloader should be garbage collected along with every class it loaded. When something holds a strong reference to it (a thread, a ThreadLocal, a JDBC driver registration, a logging appender, a static field, a shutdown hook), it stays alive. Threads still bound to it try to resolve classes against it. Those classes either no longer exist in the new context or conflict with what the old classloader already loaded, and the JVM throws &lt;code>NoClassDefFoundError&lt;/code> or &lt;code>ClassNotFoundException&lt;/code>.&lt;/p></description></item><item><title>Tomcat OutOfMemoryError: GC overhead limit exceeded: GC running but freeing nothing</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-outofmemoryerror-gc-overhead-limit-exceeded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-outofmemoryerror-gc-overhead-limit-exceeded/</guid><description>&lt;p>&lt;code>java.lang.OutOfMemoryError: GC overhead limit exceeded&lt;/code> means the heap is functionally full. GC has been running almost continuously and recovering almost nothing. This is the terminal stage of a GC death spiral, thrown just before a hard &lt;code>Java heap space&lt;/code> OOM.&lt;/p>
&lt;p>The exact rule the JVM applies: more than 98% of CPU time spent in GC and less than 2% of heap recovered, sustained across five consecutive collections. When both conditions hold, the JVM aborts rather than burn CPU indefinitely. Instantaneous heap usage can still read under &lt;code>-Xmx&lt;/code> when this fires. By that point, the live set is effectively pinned to the ceiling.&lt;/p></description></item><item><title>Tomcat OutOfMemoryError: unable to create new native thread: the JVM can't spawn threads</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-outofmemoryerror-unable-to-create-new-native-thread/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-outofmemoryerror-unable-to-create-new-native-thread/</guid><description>&lt;h1 id="tomcat-outofmemoryerror-unable-to-create-new-native-thread-the-jvm-cant-spawn-threads">Tomcat OutOfMemoryError: unable to create new native thread: the JVM can&amp;rsquo;t spawn threads&lt;/h1>
&lt;p>When Tomcat throws &lt;code>java.lang.OutOfMemoryError: unable to create new native thread&lt;/code>, do not reach for heap dumps or GC tuning. This is not a heap problem. The JVM called &lt;code>pthread_create&lt;/code> and the kernel refused. The heap can be 20% full when this fires.&lt;/p>
&lt;p>Heap OOMs need &lt;code>-Xmx&lt;/code> increases, leak hunting, or GC tuning. Native thread OOMs need &lt;code>ulimit -u&lt;/code> increases, cgroup &lt;code>pids.max&lt;/code> adjustments, &lt;code>-Xss&lt;/code> tuning, or finding the code that creates threads without bound.&lt;/p></description></item><item><title>Tomcat process not running: crashes, OOM-kills, and failed restarts</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-jvm-process-alive/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-jvm-process-alive/</guid><description>&lt;h1 id="tomcat-process-not-running-crashes-oom-kills-and-failed-restarts">Tomcat process not running: crashes, OOM-kills, and failed restarts&lt;/h1>
&lt;p>The Tomcat JVM is gone. &lt;code>pgrep -f org.apache.catalina.startup.Bootstrap&lt;/code> returns nothing, the service check is red, and your health probe is timing out. Before you restart anything, find out why it died, because the same cause will kill it again, often within minutes.&lt;/p>
&lt;p>&amp;ldquo;Process not running&amp;rdquo; conflates three distinct events: the JVM crashed inside its own runtime, the kernel OOM-killed it with SIGKILL, or the supervisor (systemd, init, container runtime) failed to bring it back. Each leaves evidence in a different place, and only one of them leaves anything in &lt;code>catalina.out&lt;/code>. This guide separates those three classes plus the &amp;ldquo;alive but nonfunctional&amp;rdquo; D-state trap.&lt;/p></description></item><item><title>Tomcat request processing time climbing: reading processingTime correctly</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-request-processing-time-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-request-processing-time-high/</guid><description>&lt;h1 id="tomcat-request-processing-time-climbing-reading-processingtime-correctly">Tomcat request processing time climbing: reading processingTime correctly&lt;/h1>
&lt;p>Your monitoring shows Tomcat request processing time trending upward. Before you chase the trend, understand what the &lt;code>processingTime&lt;/code> attribute actually represents.&lt;/p>
&lt;p>&lt;code>processingTime&lt;/code> on the &lt;code>GlobalRequestProcessor&lt;/code> MBean is a cumulative counter of wall-clock milliseconds across all requests processed since Tomcat started. It is not a per-request value, not a rate, and not a percentile. Reading it raw produces a number that only increases until the JVM restarts. To get anything useful, compute a delta over a time window and divide by the delta of &lt;code>requestCount&lt;/code> over the same window.&lt;/p></description></item><item><title>Tomcat request throughput dropping: requestCount rate below baseline</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-request-throughput-drop/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-request-throughput-drop/</guid><description>&lt;h1 id="tomcat-request-throughput-dropping-requestcount-rate-below-baseline">Tomcat request throughput dropping: requestCount rate below baseline&lt;/h1>
&lt;p>The Tomcat &lt;code>requestCount&lt;/code> counter on &lt;code>GlobalRequestProcessor&lt;/code> is climbing slower than expected. The drop tells you fewer requests are being processed, not why. The cause can be upstream of Tomcat (load balancer routed traffic away, health check started failing, CDN absorbing load) or inside Tomcat (worker thread pool exhausted, JVM in a GC death spiral, deadlock consuming the pool).&lt;/p>
&lt;p>The first diagnostic fork: is traffic still arriving at this host? If the answer is no, you have a routing, load balancer, or health check problem, and chasing thread dumps wastes time. If yes, Tomcat is failing to keep up, and the question becomes whether the failure is I/O bound (threads blocked on a slow backend) or CPU bound (GC thrashing or a compute-bound code path).&lt;/p></description></item><item><title>Tomcat RSS growing while heap looks flat: native memory and the OOM killer</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-rss-vs-heap-native-memory/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-rss-vs-heap-native-memory/</guid><description>&lt;h1 id="tomcat-rss-growing-while-heap-looks-flat-native-memory-and-the-oom-killer">Tomcat RSS growing while heap looks flat: native memory and the OOM killer&lt;/h1>
&lt;p>The Tomcat JVM dies overnight. &lt;code>systemctl status&lt;/code> shows exit code 137. Heap dashboards look flat the entire time. No &lt;code>hs_err_pid.log&lt;/code>, no heap dump, no &lt;code>OutOfMemoryError&lt;/code> in &lt;code>catalina.out&lt;/code>. The kernel, not the JVM, ended the process.&lt;/p>
&lt;p>The Linux OOM killer scores victims by RSS plus an adjustment (&lt;code>oom_score_adj&lt;/code>), not by Java heap. Everything the JVM keeps resident outside the heap counts: Metaspace, thread stacks, direct ByteBuffers, JIT code cache, JNI allocations, and glibc arenas. A Tomcat whose heap sawtooth looks healthy can still have RSS climbing toward the container limit. When RSS crosses &lt;code>memory.max&lt;/code>, the kernel sends SIGKILL. The JVM cannot trap it, log it, or dump anything.&lt;/p></description></item><item><title>Tomcat session timeout and maxActiveSessions: bounding session memory</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-session-timeout-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-session-timeout-tuning/</guid><description>&lt;h1 id="tomcat-session-timeout-and-maxactivesessions-bounding-session-memory">Tomcat session timeout and maxActiveSessions: bounding session memory&lt;/h1>
&lt;p>HTTP sessions are heap-resident state. Each active session holds references to its attribute objects, and those objects remain live until the session is invalidated or expires. Two configuration parameters bound this heap usage: the session timeout (how long a session can live) and maxActiveSessions (how many can exist simultaneously).&lt;/p>
&lt;p>At their defaults, neither bounds memory. The session timeout defaults to 30 minutes. maxActiveSessions defaults to -1, meaning no limit. Under sustained traffic, bot-driven session creation, or a traffic spike, sessions accumulate until the heap fills. The result is either an OutOfMemoryError or a GC death spiral where the JVM spends most of its CPU collecting and barely processing requests.&lt;/p></description></item><item><title>Tomcat sessions eating the heap: bots, getSession(true), and unbounded growth</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-session-memory-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-session-memory-leak/</guid><description>&lt;h1 id="tomcat-sessions-eating-the-heap-bots-getsessiontrue-and-unbounded-growth">Tomcat sessions eating the heap: bots, getSession(true), and unbounded growth&lt;/h1>
&lt;p>Post-GC heap baseline climbing steadily, Full GC frequency increasing, and eventually an OutOfMemoryError or a GC death spiral that leaves the JVM effectively unresponsive. Thread dumps look normal, CPU is dominated by GC threads, and a restart clears the problem temporarily before the cycle repeats within hours or days.&lt;/p>
&lt;p>A common and overlooked root cause is HTTP session accumulation. Under Tomcat&amp;rsquo;s default StandardManager, every active session lives in JVM heap. When sessions are created faster than they expire, they fill the heap, promote to Old Gen, and drive major GCs. Sessions are frequently the largest heap consumer in a Tomcat app, and unbounded session count is a classic memory leak vector.&lt;/p></description></item><item><title>Tomcat SEVERE Error deploying web application: reading startup failures in catalina.out</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-severe-error-deploying-web-application/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-severe-error-deploying-web-application/</guid><description>&lt;h1 id="tomcat-severe-error-deploying-web-application-reading-startup-failures-in-catalinaout">Tomcat SEVERE Error deploying web application: reading startup failures in catalina.out&lt;/h1>
&lt;p>The line you searched for looks something like this in &lt;code>catalina.out&lt;/code>:&lt;/p>
&lt;pre tabindex="0">&lt;code>SEVERE [...] org.apache.catalina.startup.HostConfig.deployWAR Error deploying web application [...]
java.lang.IllegalStateException: ContainerBase.addChild: start: org.apache.catalina.LifecycleException: Failed to start component [StandardEngine[Catalina].StandardHost[localhost].StandardContext[/yourapp]]
&lt;/code>&lt;/pre>&lt;p>That stack is a wrapper. The class that actually failed, the missing dependency, the listener that threw, or the &lt;code>NoSuchMethodError&lt;/code> from a library clash is not in &lt;code>catalina.out&lt;/code>. It is in &lt;code>localhost.YYYY-MM-DD.log&lt;/code> in the same &lt;code>$CATALINA_BASE/logs&lt;/code> directory. The Tomcat default &lt;code>logging.properties&lt;/code> routes &lt;code>org.apache.catalina.core.ContainerBase.[Catalina].[localhost]&lt;/code> to the &lt;code>2localhost.org.apache.juli.AsyncFileHandler&lt;/code>, which writes that file. The &lt;code>SEVERE&lt;/code> line in &lt;code>catalina.out&lt;/code> only tells you the deploy failed and that the context is now a &lt;code>FailedContext&lt;/code> placeholder rather than a &lt;code>StandardContext&lt;/code>.&lt;/p></description></item><item><title>Tomcat shutdown port 8005: the default SHUTDOWN command is an instant kill</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-shutdown-port-exposed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-shutdown-port-exposed/</guid><description>&lt;h1 id="tomcat-shutdown-port-8005-the-default-shutdown-command-is-an-instant-kill">Tomcat shutdown port 8005: the default SHUTDOWN command is an instant kill&lt;/h1>
&lt;p>The shutdown port is one of Tomcat&amp;rsquo;s oldest features. The stock &lt;code>server.xml&lt;/code> ships with &lt;code>&amp;lt;Server port=&amp;quot;8005&amp;quot; shutdown=&amp;quot;SHUTDOWN&amp;quot;&amp;gt;&lt;/code>, which opens a raw TCP listener. Anything that can reach that port and send the bytes &lt;code>SHUTDOWN&lt;/code> triggers an immediate, orderly shutdown. No authentication, no HTTP, no handshake. The string is well-known because it is the shipped default.&lt;/p>
&lt;p>This is a documented feature, not a CVE. But it is a configuration risk that operators routinely misjudge. If port 8005 is reachable from an untrusted network and still carries the default command, a single netcat line stops the JVM. The mitigation is straightforward: bind the listener to localhost, replace the command with a long random string, or disable the port entirely with &lt;code>port=&amp;quot;-1&amp;quot;&lt;/code>.&lt;/p></description></item><item><title>Tomcat Slowloris slow-client attack: connections climbing while bytes stay near zero</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-slowloris-slow-client/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-slowloris-slow-client/</guid><description>&lt;p>Your Tomcat HTTP connector&amp;rsquo;s connection count keeps climbing. The bytes received rate is nearly flat. CPU is normal, heap looks healthy, GC is quiet, and the thread pool is nowhere near saturated. Yet legitimate users are timing out, and new connections are starting to get refused.&lt;/p>
&lt;p>This is the Slowloris slow-client pattern. An attacker opens many TCP connections to Tomcat and dribbles data across them, one byte at a time, or sends a partial HTTP request and never finishes the headers. The connection stays open, consuming a slot in the NIO poller or, on older BIO connectors, a worker thread. The request never completes, so it never reaches your application code. From Tomcat&amp;rsquo;s perspective, the JVM is idle. From the user&amp;rsquo;s perspective, the site is down.&lt;/p></description></item><item><title>Tomcat stuck threads: StuckThreadDetectionValve and 'may be stuck' warnings</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-stuck-threads/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-stuck-threads/</guid><description>&lt;h1 id="tomcat-stuck-threads-stuckthreaddetectionvalve-and-may-be-stuck-warnings">Tomcat stuck threads: StuckThreadDetectionValve and &amp;lsquo;may be stuck&amp;rsquo; warnings&lt;/h1>
&lt;p>The warning shows up in &lt;code>catalina.out&lt;/code>: a worker thread &amp;ldquo;has been active for [N] milliseconds &amp;hellip; and may be stuck.&amp;rdquo; By the time you see it, the thread has already been blocked for the full threshold duration. If the valve is running with its default 600-second threshold, that is 10 minutes of a worker thread doing nothing useful. With &lt;code>maxThreads=200&lt;/code>, a handful of these threads eats real capacity.&lt;/p></description></item><item><title>Tomcat thread dumps: reading jstack to find what the pool is waiting on</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-thread-dump-jstack/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-thread-dump-jstack/</guid><description>&lt;h1 id="tomcat-thread-dumps-reading-jstack-to-find-what-the-pool-is-waiting-on">Tomcat thread dumps: reading jstack to find what the pool is waiting on&lt;/h1>
&lt;p>When &lt;code>currentThreadsBusy&lt;/code> reaches &lt;code>maxThreads&lt;/code>, Tomcat stops processing new requests even though the JVM is healthy. The thread pool gauge tells you the pool is full. It does not tell you why. The only diagnostic that shows exactly what every worker thread is doing at that moment is a JVM thread dump.&lt;/p>
&lt;p>&lt;code>jstack &amp;lt;pid&amp;gt;&lt;/code> (or the modern equivalent &lt;code>jcmd &amp;lt;pid&amp;gt; Thread.print&lt;/code>) is the primary first response for thread-related Tomcat failures. A single dump shows whether each &lt;code>http-nio-exec&lt;/code> thread is idle and parked in the task queue, blocked in a native socket read against a slow backend, waiting to acquire a database connection, or contending on a Java monitor. Many threads parked in the same stack frame is the stuck path. A cluster of threads &lt;code>BLOCKED&lt;/code> on the same lock is contention, or a deadlock.&lt;/p></description></item><item><title>Tomcat thread pool exhaustion: currentThreadsBusy at maxThreads and requests hanging</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-thread-pool-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-thread-pool-exhaustion/</guid><description>&lt;h1 id="tomcat-thread-pool-exhaustion-currentthreadsbusy-at-maxthreads-and-requests-hanging">Tomcat thread pool exhaustion: currentThreadsBusy at maxThreads and requests hanging&lt;/h1>
&lt;p>Tomcat is up: the JVM process is running, the HTTP port is open, health checks pass. But requests hang or return 503. The dashboard shows &lt;code>currentThreadsBusy&lt;/code> pinned at &lt;code>maxThreads&lt;/code> (default 200). CPU is low. Heap is stable. No &lt;code>OutOfMemoryError&lt;/code> in the logs.&lt;/p>
&lt;p>This is thread pool exhaustion. Every worker thread is occupied and no thread is available to process new requests. The JVM is healthy; the application is just not being served.&lt;/p></description></item><item><title>Tomcat threads blocked forever: the missing outbound timeout</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-backend-timeout-missing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-backend-timeout-missing/</guid><description>&lt;h1 id="tomcat-threads-blocked-forever-the-missing-outbound-timeout">Tomcat threads blocked forever: the missing outbound timeout&lt;/h1>
&lt;p>The JVM is healthy. Heap is fine, no Full GCs, CPU sits at 5%, and the process answers a TCP connect on 8080 within milliseconds. But every HTTP request hangs, the access log has gone quiet, and clients are timing out. When you finally grab a thread dump, most of the &lt;code>http-nio-8080-exec&lt;/code> threads are parked in the same stack frame: &lt;code>java.net.SocketInputStream.socketRead0&lt;/code>, blocked against a backend that will never reply.&lt;/p></description></item><item><title>Tomcat threads busy but CPU idle: telling a blocked backend from a GC spiral</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-threads-busy-low-cpu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-threads-busy-low-cpu/</guid><description>&lt;h1 id="tomcat-threads-busy-but-cpu-idle-telling-a-blocked-backend-from-a-gc-spiral">Tomcat threads busy but CPU idle: telling a blocked backend from a GC spiral&lt;/h1>
&lt;p>When users report that Tomcat is &amp;ldquo;hung&amp;rdquo; but the JVM is up, ports are open, and the operating system looks fine, the fastest split you can make is to look at one number alongside the thread pool state: JVM CPU. Two different failure modes produce &lt;code>currentThreadsBusy == maxThreads&lt;/code>, and they have opposite CPU signatures.&lt;/p>
&lt;p>If CPU is low while every worker thread is busy, those threads are parked on I/O: a slow database, a hung downstream HTTP service, an unresolvable DNS lookup, or a stalled NFS mount. The JVM is healthy; a backend is not. Restarting Tomcat is the wrong move, because the threads re-block the moment traffic returns.&lt;/p></description></item><item><title>Tomcat virtual threads (JDK 21+): why currentThreadsBusy reports -1</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-virtual-threads-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-virtual-threads-monitoring/</guid><description>&lt;h1 id="tomcat-virtual-threads-jdk-21-why-currentthreadsbusy-reports--1">Tomcat virtual threads (JDK 21+): why currentThreadsBusy reports -1&lt;/h1>
&lt;p>You upgraded to JDK 21, enabled virtual threads, and the one Tomcat gauge every dashboard and every alert leans on (&lt;code>currentThreadsBusy&lt;/code>) now reads &lt;code>-1&lt;/code>. &lt;code>maxThreads&lt;/code> still shows 200 but is meaningless. The thread pool utilization panel that used to be your primary saturation signal is blank or broken, and your SLO alerts are either firing constantly or silently dead.&lt;/p>
&lt;p>This is expected, by-design behavior, not a bug. When a connector runs on virtual threads (&lt;code>useVirtualThreads=&amp;quot;true&amp;quot;&lt;/code> on NioEndpoint, or &lt;code>StandardVirtualThreadExecutor&lt;/code>, on Tomcat 10.1.x/11.0.x with JDK 21+, or Tomcat 9.0.84+ with &lt;code>useVirtualThreads&lt;/code>), Tomcat&amp;rsquo;s classic bounded-thread-pool model stops applying. There is no fixed pool to measure. &lt;code>currentThreadsBusy&lt;/code> returns -1, &lt;code>maxThreads&lt;/code> is inherited from defaults and not honored, and the connection layer becomes your only proxy for active request load.&lt;/p></description></item><item><title>Tomcat weak TLS: disabling TLSv1.0/1.1 and hardening ciphers on 8443</title><link>https://www.netdata.cloud/guides/tomcat/tomcat-tls-weak-protocols/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/tomcat/tomcat-tls-weak-protocols/</guid><description>&lt;h1 id="tomcat-weak-tls-disabling-tlsv1011-and-hardening-ciphers-on-8443">Tomcat weak TLS: disabling TLSv1.0/1.1 and hardening ciphers on 8443&lt;/h1>
&lt;p>An HTTPS connector on 8443 that still negotiates TLSv1.0 or TLSv1.1 is active downgrade surface. Clients that can force TLSv1.0 can exploit protocol-level weaknesses, and ciphers that exist only to support legacy protocols weaken posture for every client. The fix has two parts: restrict the protocol list to TLSv1.2 and TLSv1.3, and prune the cipher list so weak suites are not offered.&lt;/p></description></item><item><title>Tor</title><link>https://www.netdata.cloud/integrations/data-collection/networking/tor/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/tor/</guid><description/></item><item><title>Tor Monitoring</title><link>https://www.netdata.cloud/monitoring-101/tor-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/tor-monitoring/</guid><description>&lt;h2 id="tor-monitoring">Tor Monitoring&lt;/h2>
&lt;h3 id="what-is-tor">What Is Tor?&lt;/h3>
&lt;p>Tor, short for The Onion Router, is a free and open-source software that allows anonymous communication over the internet. Created by the Tor Project, it directs internet traffic through a free, worldwide, volunteer overlay network consisting of more than seven thousand relays. This conceals a user&amp;rsquo;s location and usage from network surveillance or traffic analysis.&lt;/p>
&lt;h3 id="monitoring-tor-with-netdata">Monitoring Tor With Netdata&lt;/h3>
&lt;p>To monitor Tor effectively, it&amp;rsquo;s crucial to use a tool designed for real-time performance monitoring, and that&amp;rsquo;s where the &lt;a href="https://app.netdata.cloud/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata monitoring tool&lt;/a> comes in. Netdata offers seamless and robust monitoring capabilities, designed to gather and analyze the most relevant metrics for maintaining optimal performance.&lt;/p></description></item><item><title>Toshiba Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/toshiba-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/toshiba-corporation-snmp-traps/</guid><description/></item><item><title>TP Link</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/tp-link/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/tp-link/</guid><description/></item><item><title>Tp Link Systems Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tp-link-systems-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tp-link-systems-inc-snmp-traps/</guid><description/></item><item><title>Traefik</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/traefik/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/traefik/</guid><description/></item><item><title>Traefik /ping returns 200 while everything is broken: the health-check trap</title><link>https://www.netdata.cloud/guides/traefik/traefik-ping-health-check-trap/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-ping-health-check-trap/</guid><description>&lt;h1 id="traefik-ping-returns-200-while-everything-is-broken-the-health-check-trap">Traefik /ping returns 200 while everything is broken: the health-check trap&lt;/h1>
&lt;p>Traefik&amp;rsquo;s &lt;code>/ping&lt;/code> endpoint answers one narrow question: &amp;ldquo;Is the Traefik process alive enough to answer this request?&amp;rdquo; It does not answer &amp;ldquo;Can Traefik correctly route production traffic?&amp;rdquo;&lt;/p>
&lt;p>That distinction matters because Traefik is both a data-plane proxy and a control-plane configuration reconciler. The process can stay alive while its Docker socket or Kubernetes API watch has failed, its routing table is stale, every backend is failing health checks, or a certificate is approaching expiry. In all of those states, &lt;code>/ping&lt;/code> can still return &lt;code>200 OK&lt;/code>.&lt;/p></description></item><item><title>Traefik 401/403 spike: auth failures, ForwardAuth outages, and brute force</title><link>https://www.netdata.cloud/guides/traefik/traefik-401-403-spike/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-401-403-spike/</guid><description>&lt;h1 id="traefik-401403-spike-auth-failures-forwardauth-outages-and-brute-force">Traefik 401/403 spike: auth failures, ForwardAuth outages, and brute force&lt;/h1>
&lt;p>Your dashboard shows &lt;code>traefik_service_requests_total{code=&amp;quot;401&amp;quot;}&lt;/code> climbing steeply, or the entrypoint-level 4xx panel has turned red. A 401/403 surge is one of the most ambiguous signals a reverse proxy produces, because the same symptom maps to three different situations: your auth layer correctly rejecting bad credentials, your auth layer incorrectly rejecting everyone, or an attacker hammering a login endpoint.&lt;/p>
&lt;p>The trap: a ForwardAuth service that is up but degraded (session store exhausted, internal rate limit hit, broken ACL rule) rejects all requests with 401/403. To Traefik this looks like &amp;ldquo;auth working as designed.&amp;rdquo; To your users it is a full outage. Conversely, a credential-stuffing run produces the same metric shape but demands the opposite response: block, don&amp;rsquo;t fix.&lt;/p></description></item><item><title>Traefik 404 not found: requests arriving with no matching router</title><link>https://www.netdata.cloud/guides/traefik/traefik-404-not-found/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-404-not-found/</guid><description>&lt;h1 id="traefik-404-not-found-requests-arriving-with-no-matching-router">Traefik 404 not found: requests arriving with no matching router&lt;/h1>
&lt;p>A client reports that your service returns &amp;ldquo;404 page not found&amp;rdquo;. The backend is healthy, its logs show zero traffic, and the 404 body is a bare 19-byte plain-text response, not your application&amp;rsquo;s error page. That 404 did not come from your backend. It came from Traefik itself.&lt;/p>
&lt;p>When a request arrives at a Traefik entrypoint and matches no router, Traefik answers it directly with a 404. The request never touches a service, never runs through a middleware chain, and never reaches an upstream. This is a routing-layer problem with a completely different investigation path from a backend-originated 404, and mixing them up is one of the most common time-wasters in Traefik operations.&lt;/p></description></item><item><title>Traefik 502 Bad Gateway: when the backend is unreachable or returns garbage</title><link>https://www.netdata.cloud/guides/traefik/traefik-502-bad-gateway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-502-bad-gateway/</guid><description>&lt;h1 id="traefik-502-bad-gateway-when-the-backend-is-unreachable-or-returns-garbage">Traefik 502 Bad Gateway: when the backend is unreachable or returns garbage&lt;/h1>
&lt;p>A 502 from Traefik means the request got past routing and middleware, Traefik selected a service and tried to talk to a backend, and the backend hop failed. Traefik connected (or tried to connect) and got an invalid response, a reset connection, or nothing usable back. The usual mechanics: the backend crashed mid-response, sent something that is not valid HTTP for the negotiated protocol, or closed a connection Traefik wanted to reuse.&lt;/p></description></item><item><title>Traefik 503 Service Unavailable: no healthy backends left in the pool</title><link>https://www.netdata.cloud/guides/traefik/traefik-503-service-unavailable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-503-service-unavailable/</guid><description>&lt;h1 id="traefik-503-service-unavailable-no-healthy-backends-left-in-the-pool">Traefik 503 Service Unavailable: no healthy backends left in the pool&lt;/h1>
&lt;p>Traefik is returning &lt;code>503 Service Unavailable&lt;/code> for every request to a service. Clients are down. Traefik itself looks fine: the process is up, &lt;code>/ping&lt;/code> returns 200, other services route normally. This is the backend pool collapse failure mode: Traefik matched a router to a service, but every backend in that service&amp;rsquo;s load balancer pool has been marked down by health checks, so Traefik has nowhere to send the request.&lt;/p></description></item><item><title>Traefik 504 Gateway Timeout: the backend is alive but too slow</title><link>https://www.netdata.cloud/guides/traefik/traefik-504-gateway-timeout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-504-gateway-timeout/</guid><description>&lt;h1 id="traefik-504-gateway-timeout-the-backend-is-alive-but-too-slow">Traefik 504 Gateway Timeout: the backend is alive but too slow&lt;/h1>
&lt;p>Your Traefik instance is up. &lt;code>/ping&lt;/code> returns 200. The backend service is running, its health checks pass, and clients are getting &lt;code>504 Gateway Timeout&lt;/code>. The connection to the backend was established, but the response never arrived within the configured timeout.&lt;/p>
&lt;p>A 504 in Traefik means: the backend exists, Traefik reached it, the TCP connection (and usually the request) succeeded, but the response took too long. That makes a 504 fundamentally different from a 502 (invalid response or connection error from the backend) or a 503 (no healthy backends at all). Lumping all 5xx together leads directly to investigating the wrong layer.&lt;/p></description></item><item><title>Traefik 5xx error rate: telling Traefik-generated errors from backend errors</title><link>https://www.netdata.cloud/guides/traefik/traefik-5xx-error-rate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-5xx-error-rate/</guid><description>&lt;h1 id="traefik-5xx-error-rate-telling-traefik-generated-errors-from-backend-errors">Traefik 5xx error rate: telling Traefik-generated errors from backend errors&lt;/h1>
&lt;p>Your edge dashboard shows a 5xx spike on Traefik. Someone pages the backend team. The backend team finds nothing wrong in their application logs. An hour later it turns out Traefik had no healthy backends for one service, or a route disappeared after a config change, and the application was never in the request path at all.&lt;/p>
&lt;p>Treating all 5xx responses as backend failures is one of the most common time-wasters in Traefik operations. A 502 generated by Traefik because it could not reach an upstream is a different incident from a 502 the backend produced itself. A 503 because every backend failed health checks is a different incident from a 503 your application returns under load. The root causes, the owners, and the fixes are all different.&lt;/p></description></item><item><title>Traefik access log blocking: when logging stalls request handling</title><link>https://www.netdata.cloud/guides/traefik/traefik-access-log-blocking/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-access-log-blocking/</guid><description>&lt;h1 id="traefik-access-log-blocking-when-logging-stalls-request-handling">Traefik access log blocking: when logging stalls request handling&lt;/h1>
&lt;p>Latency spikes across every route on a Traefik instance. Backends report nothing wrong. Service-level latency looks normal, but entrypoint latency is elevated. Health checks pass. If you see this combination and access logging is enabled at high verbosity or high request rate, the log writer is a prime suspect.&lt;/p>
&lt;p>Traefik&amp;rsquo;s access log sits in the request path. With the default configuration (&lt;code>bufferingSize: 0&lt;/code>), the access log line for a request is written synchronously from the request-handling goroutine before the request fully completes. When log volume is extreme (thousands of requests per second with verbose field selection), or when the write destination stops draining (full filesystem, blocked pipe), the log write stalls and the goroutine handling that request stalls with it. Requests do not fail; they get slow. That distinction is what makes this failure mode easy to miss.&lt;/p></description></item><item><title>Traefik ACME challenge failed: HTTP-01, DNS-01, and TLS-ALPN-01 renewal errors</title><link>https://www.netdata.cloud/guides/traefik/traefik-acme-challenge-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-acme-challenge-failure/</guid><description>&lt;h1 id="traefik-acme-challenge-failed-http-01-dns-01-and-tls-alpn-01-renewal-errors">Traefik ACME challenge failed: HTTP-01, DNS-01, and TLS-ALPN-01 renewal errors&lt;/h1>
&lt;p>Your certificates are approaching expiry and Traefik&amp;rsquo;s logs contain lines like &amp;ldquo;unable to obtain ACME certificate&amp;rdquo; or &amp;ldquo;Error renewing certificate&amp;rdquo;. HTTPS still works for now because Traefik keeps serving the old certificate, but the clock is running: Let&amp;rsquo;s Encrypt certificates live 90 days, Traefik attempts renewal 30 days before expiry, and if you are seeing renewal errors today, they have likely been failing silently for weeks.&lt;/p></description></item><item><title>Traefik ACME lock contention: stuck distributed locks blocking renewal</title><link>https://www.netdata.cloud/guides/traefik/traefik-acme-lock-contention/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-acme-lock-contention/</guid><description>&lt;h1 id="traefik-acme-lock-contention-stuck-distributed-locks-blocking-renewal">Traefik ACME lock contention: stuck distributed locks blocking renewal&lt;/h1>
&lt;p>Every Traefik instance in your HA deployment is healthy. &lt;code>/ping&lt;/code> returns 200, traffic flows, backends are up. Yet &lt;code>traefik_tls_certs_not_after&lt;/code> keeps creeping toward the current time, and the logs on each instance repeat some variant of &amp;ldquo;unable to acquire lock&amp;rdquo;. Certificates are drifting toward expiry and nobody is renewing them.&lt;/p>
&lt;p>This is the multi-instance ACME failure mode: Traefik uses a distributed lock in a KV store (Consul or etcd) so that exactly one instance performs ACME issuance and renewal at a time. When the lock holder crashes mid-renewal, is force-killed, or loses its session to a network partition, the lock can be left behind. Every remaining instance refuses to renew because the lock appears held, and the cluster slides toward certificate expiry while looking healthy on every conventional health signal.&lt;/p></description></item><item><title>Traefik ACME rate limit: too many certificates already issued for this domain</title><link>https://www.netdata.cloud/guides/traefik/traefik-acme-rate-limit/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-acme-rate-limit/</guid><description>&lt;h1 id="traefik-acme-rate-limit-too-many-certificates-already-issued-for-this-domain">Traefik ACME rate limit: too many certificates already issued for this domain&lt;/h1>
&lt;p>You find it in the Traefik log, usually right after a deploy, a restart storm, or a cluster rebuild: an ACME order rejected with an HTTP 429 and the message &amp;ldquo;too many certificates already issued&amp;rdquo; for your domain. Traefik cannot get a new certificate, and it will keep retrying and keep getting rejected.&lt;/p>
&lt;p>Two very different situations sit behind this error. If your existing certificates are still valid and stored in &lt;code>acme.json&lt;/code>, traffic keeps flowing and you have days or weeks to fix the renewal path before expiry. If &lt;code>acme.json&lt;/code> was lost along with the certificates, you are rate limited and serving nothing valid on HTTPS, which is an outage on a timer.&lt;/p></description></item><item><title>Traefik acme.json permissions and corruption: renewal silently blocked</title><link>https://www.netdata.cloud/guides/traefik/traefik-acme-json-corruption/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-acme-json-corruption/</guid><description>&lt;h1 id="traefik-acmejson-permissions-and-corruption-renewal-silently-blocked">Traefik acme.json permissions and corruption: renewal silently blocked&lt;/h1>
&lt;p>Your certificates are approaching expiry and Traefik has not renewed them. There is no alert, no error metric, and depending on how the failure happened, possibly nothing useful in the logs either. The root cause in a large share of these incidents is the ACME storage file itself: &lt;code>acme.json&lt;/code> has the wrong permissions, or it has been corrupted by a partial write, a restore, or an out-of-memory event during a write.&lt;/p></description></item><item><title>Traefik backend connection pool: keep-alive, MaxIdleConnsPerHost, and reuse</title><link>https://www.netdata.cloud/guides/traefik/traefik-backend-connection-pool/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-backend-connection-pool/</guid><description>&lt;h1 id="traefik-backend-connection-pool-keep-alive-maxidleconnsperhost-and-reuse">Traefik backend connection pool: keep-alive, MaxIdleConnsPerHost, and reuse&lt;/h1>
&lt;p>Every request Traefik proxies needs a TCP connection to a backend. Whether that connection is freshly dialed or reused from a pool determines how much per-request overhead you pay, and it is one of the least monitored parts of the proxy. A misconfigured pool shows up as elevated latency on every request (TCP and possibly TLS handshake costs per request), as growing TIME_WAIT socket counts marching toward ephemeral port exhaustion, or as sporadic 502s with no backend outage to explain them.&lt;/p></description></item><item><title>Traefik cannot assign requested address: ephemeral port exhaustion</title><link>https://www.netdata.cloud/guides/traefik/traefik-cannot-assign-requested-address/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-cannot-assign-requested-address/</guid><description>&lt;h1 id="traefik-cannot-assign-requested-address-ephemeral-port-exhaustion">Traefik cannot assign requested address: ephemeral port exhaustion&lt;/h1>
&lt;p>Traefik starts returning 502 Bad Gateway for some or all proxied requests. Backend health checks are green. Backends respond fine when you hit them directly. The Traefik logs show the real error: &lt;code>dial tcp &amp;lt;backend-ip&amp;gt;:&amp;lt;port&amp;gt;: connect: cannot assign requested address&lt;/code>.&lt;/p>
&lt;p>This is ephemeral port exhaustion: the kernel has no free local ports left for Traefik to open a new outbound TCP connection to a backend. Almost every port in the ephemeral range is pinned by a socket in TIME_WAIT, left behind by a connection that closed up to a minute ago.&lt;/p></description></item><item><title>Traefik cascading backend failure: how a partial outage becomes a total one</title><link>https://www.netdata.cloud/guides/traefik/traefik-cascading-backend-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-cascading-backend-failure/</guid><description>&lt;h1 id="traefik-cascading-backend-failure-how-a-partial-outage-becomes-a-total-one">Traefik cascading backend failure: how a partial outage becomes a total one&lt;/h1>
&lt;p>A backend pool behind Traefik rarely fails all at once. It fails one server at a time, and each failure makes the next one more likely. Two of your eight backends go down, the remaining six absorb the redistributed traffic plus the retry traffic Traefik generates, one of the six starts timing out under the extra load, and now five are carrying the full weight. Within minutes the pool is empty and Traefik returns 503 to every client, even though Traefik itself is perfectly healthy.&lt;/p></description></item><item><title>Traefik certificate expired: when ACME renewal has been failing silently</title><link>https://www.netdata.cloud/guides/traefik/traefik-certificate-expired/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-certificate-expired/</guid><description>&lt;h1 id="traefik-certificate-expired-when-acme-renewal-has-been-failing-silently">Traefik certificate expired: when ACME renewal has been failing silently&lt;/h1>
&lt;p>Browsers are rejecting your site with a certificate error, and Traefik is the TLS terminator. The process is up, &lt;code>/ping&lt;/code> returns 200, traffic flows on port 80, and every HTTPS client reports an expired or expiring certificate. This is the signature of silent ACME renewal failure.&lt;/p>
&lt;p>The nasty part is the timeline. Traefik attempts renewal 30 days before expiry. For a 90-day Let&amp;rsquo;s Encrypt certificate, that means renewal has been attempted since day 60. If a certificate is 7 days from expiry, renewal has already been failing for roughly 53 days. There is no Prometheus metric for ACME failure. The only metric you get is the consequence: &lt;code>traefik_tls_certs_not_after&lt;/code> sliding toward the current time while nothing alerts.&lt;/p></description></item><item><title>Traefik circuit breaker: shedding load from a failing backend</title><link>https://www.netdata.cloud/guides/traefik/traefik-circuit-breaker/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-circuit-breaker/</guid><description>&lt;h1 id="traefik-circuit-breaker-shedding-load-from-a-failing-backend">Traefik circuit breaker: shedding load from a failing backend&lt;/h1>
&lt;p>A backend is degrading. It is not fully down, so health checks still pass or flap, and Traefik keeps sending it full traffic. If you also have the retry middleware attached, every failed request comes back two or three more times, and the retry amplification loop finishes what the original failure started. The circuit breaker middleware exists for exactly this situation: it watches error and latency ratios on a router, and when they cross a threshold you define, it stops forwarding requests and answers with a fast 503 until the backend recovers.&lt;/p></description></item><item><title>Traefik CLOSE_WAIT pile-up: backend connections that never close</title><link>https://www.netdata.cloud/guides/traefik/traefik-close-wait-connection-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-close-wait-connection-leak/</guid><description>&lt;h1 id="traefik-close_wait-pile-up-backend-connections-that-never-close">Traefik CLOSE_WAIT pile-up: backend connections that never close&lt;/h1>
&lt;p>You run &lt;code>ss -s&lt;/code> or glance at &lt;code>process_open_fds&lt;/code> and notice the Traefik process is holding thousands of sockets in CLOSE_WAIT. Traffic is still flowing, health checks pass, and &lt;code>/ping&lt;/code> returns 200, but the count grows every hour. Left alone, the leak consumes file descriptors until the process hits its limit, at which point every new connection fails and clients start seeing 502s.&lt;/p></description></item><item><title>Traefik config last reload success: monitoring configuration freshness</title><link>https://www.netdata.cloud/guides/traefik/traefik-config-last-reload-success/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-config-last-reload-success/</guid><description>&lt;h1 id="traefik-config-last-reload-success-monitoring-configuration-freshness">Traefik config last reload success: monitoring configuration freshness&lt;/h1>
&lt;p>Traefik&amp;rsquo;s most dangerous failure mode is the one that looks like nothing. The process is up, &lt;code>/ping&lt;/code> returns 200, existing routes keep serving traffic, and every dashboard is green. Meanwhile the configuration provider disconnected two hours ago, the three services you deployed since then are invisible, and the service you scaled down is still getting traffic at dead addresses. This is provider desync, and the only Traefik-native signal that exposes it is configuration freshness.&lt;/p></description></item><item><title>Traefik configuration reload storm: provider churn eating CPU</title><link>https://www.netdata.cloud/guides/traefik/traefik-config-reload-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-config-reload-storm/</guid><description>&lt;h1 id="traefik-configuration-reload-storm-provider-churn-eating-cpu">Traefik configuration reload storm: provider churn eating CPU&lt;/h1>
&lt;p>Traefik&amp;rsquo;s CPU is pinned, request latency is jittery, GC pauses are climbing, and traffic volume looks completely normal. The culprit is not the traffic. It is the control plane: &lt;code>rate(traefik_config_reloads_total[5m])&lt;/code> is sitting above 1 reload per second, and every reload rebuilds the entire routing table.&lt;/p>
&lt;p>In busy Kubernetes and Docker environments, every pod event (creation, deletion, readiness change) can trigger a configuration rebuild. At high churn rates Traefik spends more CPU rebuilding the routing table than it spends routing requests. This is most visible during cluster-wide rollouts and autoscaler events, but it also appears when health checks flap and pods re-register in a loop.&lt;/p></description></item><item><title>Traefik context canceled during reload: intermittent errors on config change</title><link>https://www.netdata.cloud/guides/traefik/traefik-context-canceled-reload/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-context-canceled-reload/</guid><description>&lt;h1 id="traefik-context-canceled-during-reload-intermittent-errors-on-config-change">Traefik context canceled during reload: intermittent errors on config change&lt;/h1>
&lt;p>You are seeing small bursts of failed requests: a handful of 499s or 502s, lasting a second or two, then nothing. The backends are healthy, health checks pass, latency is normal, and the errors do not repeat on any schedule. When you overlay the error timestamps on Traefik&amp;rsquo;s metrics, each burst lines up exactly with an increment of &lt;code>traefik_config_reloads_total&lt;/code>.&lt;/p>
&lt;p>During a dynamic configuration rebuild, Traefik&amp;rsquo;s &lt;code>cancelPrevState()&lt;/code> cancels the context attached to the previous router configuration. In-flight requests, middleware chains, or backend connections that still hold a reference to that context receive a context-canceled error mid-request and fail. It is rare, intermittent, and almost impossible to reproduce on demand, which is why it gets misdiagnosed as flaky backends or network blips.&lt;/p></description></item><item><title>Traefik dashboard and API exposed: your routing table on the public internet</title><link>https://www.netdata.cloud/guides/traefik/traefik-dashboard-api-exposed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-dashboard-api-exposed/</guid><description>&lt;h1 id="traefik-dashboard-and-api-exposed-your-routing-table-on-the-public-internet">Traefik dashboard and API exposed: your routing table on the public internet&lt;/h1>
&lt;p>If Traefik&amp;rsquo;s dashboard, API endpoints (&lt;code>/api/rawdata&lt;/code>, &lt;code>/api/http/services&lt;/code>, &lt;code>/api/http/routers&lt;/code>), or &lt;code>/debug/pprof&lt;/code> answer requests from the public internet, you are publishing a map of your infrastructure. The API returns every router rule, backend server URL, middleware configuration, and upstream health status. &lt;code>/debug/pprof&lt;/code> adds Go runtime profiling data. Automated scanners fingerprint Traefik continuously, and the dashboard API has a history of information-disclosure issues. The fix is usually a ten-minute configuration change, but only if you know the exposure exists. Many teams never test from outside their own network.&lt;/p></description></item><item><title>Traefik dashboard returns 404: reaching the API and dashboard correctly</title><link>https://www.netdata.cloud/guides/traefik/traefik-dashboard-404/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-dashboard-404/</guid><description>&lt;h1 id="traefik-dashboard-returns-404-reaching-the-api-and-dashboard-correctly">Traefik dashboard returns 404: reaching the API and dashboard correctly&lt;/h1>
&lt;p>You opened a browser to what you think is the dashboard address and got a bare &amp;ldquo;404 page not found&amp;rdquo;. Or the dashboard HTML loads but every panel is empty because the underlying &lt;code>/api&lt;/code> calls all return 404. This almost always comes down to one of three things: the API is not enabled, you are hitting the wrong port or path, or secure mode is on but no router was ever defined for the internal API service.&lt;/p></description></item><item><title>Traefik entrypoint vs service latency: measuring middleware overhead</title><link>https://www.netdata.cloud/guides/traefik/traefik-entrypoint-vs-service-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-entrypoint-vs-service-latency/</guid><description>&lt;h1 id="traefik-entrypoint-vs-service-latency-measuring-middleware-overhead">Traefik entrypoint vs service latency: measuring middleware overhead&lt;/h1>
&lt;p>A client reports that requests through Traefik are slow. You check &lt;code>traefik_service_request_duration_seconds&lt;/code> and the backends look fine: p95 is 80ms, well inside the SLO. Yet clients measure 400ms end to end. The missing time is being spent inside Traefik itself, and the service-level metric cannot see it.&lt;/p>
&lt;p>Traefik exposes request duration at two points in its pipeline: at the entrypoint, where the request arrives from the client, and at the service, where the request is handed to a backend. The difference between the two is Traefik&amp;rsquo;s own overhead: TLS termination, router matching, the middleware chain, compression, and request/response buffering. Because Traefik ships no per-middleware Prometheus metrics, this subtraction is the only metrics-based way to localise middleware cost without capturing Go profiles or traces.&lt;/p></description></item><item><title>Traefik file descriptor monitoring: process_open_fds, limits, and headroom</title><link>https://www.netdata.cloud/guides/traefik/traefik-file-descriptor-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-file-descriptor-monitoring/</guid><description>&lt;h1 id="traefik-file-descriptor-monitoring-process_open_fds-limits-and-headroom">Traefik file descriptor monitoring: process_open_fds, limits, and headroom&lt;/h1>
&lt;p>Traefik file descriptor exhaustion is a cliff-edge failure. The proxy works until it hits its OS limit on open files, then 100% of new connections fail instantly with &amp;ldquo;too many open files&amp;rdquo; errors. Existing connections keep working, so dashboards can look calm while every new client is dropped. There is no graceful degradation.&lt;/p>
&lt;p>The most obvious Traefik connection metric, &lt;code>traefik_open_connections&lt;/code>, does not measure the thing that kills you. It tracks only entrypoint connections, a subset of total FD usage. The comprehensive signal is the ratio &lt;code>process_open_fds / process_max_fds&lt;/code> from the Go process collector, which counts everything the process holds open: client sockets, backend sockets, provider connections, log files, ACME storage, and pipes.&lt;/p></description></item><item><title>Traefik GC pauses: when Go garbage collection shows up in request latency</title><link>https://www.netdata.cloud/guides/traefik/traefik-gc-pause-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-gc-pause-latency/</guid><description>&lt;h1 id="traefik-gc-pauses-when-go-garbage-collection-shows-up-in-request-latency">Traefik GC pauses: when Go garbage collection shows up in request latency&lt;/h1>
&lt;p>Your p99 latency graph has periodic spikes that do not line up with traffic, deployments, or backend slowness. The spikes hit every service at once, last tens to hundreds of milliseconds, and then everything goes back to normal. When a latency event affects all concurrent requests simultaneously and leaves backends untouched, the suspect list gets short, and Go garbage collection is near the top of it.&lt;/p></description></item><item><title>Traefik goroutine leak: go_goroutines climbing toward OOM</title><link>https://www.netdata.cloud/guides/traefik/traefik-goroutine-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-goroutine-leak/</guid><description>&lt;h1 id="traefik-goroutine-leak-go_goroutines-climbing-toward-oom">Traefik goroutine leak: go_goroutines climbing toward OOM&lt;/h1>
&lt;p>Your Traefik dashboard shows &lt;code>go_goroutines&lt;/code> climbing steadily over hours or days. Traffic is flat. Latency and error rates look normal. Yet the goroutine count keeps rising, RSS grows in lockstep, and you can extrapolate a straight line from today&amp;rsquo;s memory usage to the container limit and read off the day Traefik gets OOM-killed, dropping every in-flight connection with it.&lt;/p>
&lt;p>In Traefik this almost always has the same shape: a backend accepts the TCP connection but never sends response headers, and the forwarding transport has no timeout (or a very long one) on that wait. Each stuck request leaves a goroutine blocked indefinitely. The goroutine count is the earliest visible signal, long before users feel anything.&lt;/p></description></item><item><title>Traefik HA config drift: replicas serving different routing tables</title><link>https://www.netdata.cloud/guides/traefik/traefik-ha-config-drift/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-ha-config-drift/</guid><description>&lt;h1 id="traefik-ha-config-drift-replicas-serving-different-routing-tables">Traefik HA config drift: replicas serving different routing tables&lt;/h1>
&lt;p>The symptom: a route works, then it doesn&amp;rsquo;t, then it does. Refreshing the page flips between a working service and a 404 or 502. Testing from one network works; from another it fails. Behind this randomness is a simple mechanism: you run multiple Traefik replicas for availability, one has gone stale, and the load balancer in front of them sends your requests to whichever replica it picks.&lt;/p></description></item><item><title>Traefik health checks pass but requests fail: when the probe lies</title><link>https://www.netdata.cloud/guides/traefik/traefik-health-check-false-positive/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-health-check-false-positive/</guid><description>&lt;h1 id="traefik-health-checks-pass-but-requests-fail-when-the-probe-lies">Traefik health checks pass but requests fail: when the probe lies&lt;/h1>
&lt;p>Every backend for the service shows &lt;code>traefik_service_server_up = 1&lt;/code>. The dashboard is green. And clients are getting 502s and 504s on real requests. Traefik is not malfunctioning: it is faithfully reporting the results of the probe you configured. The problem is that the probe and the production traffic path are testing different things.&lt;/p>
&lt;p>This failure mode inverts your usual triage instinct. Normally &lt;code>server_up = 1&lt;/code> means &amp;ldquo;rule out the backend.&amp;rdquo; Here it means nothing of the sort: the health check answered a question, just not the one your users are asking. Three distinct mechanisms produce this state, and they have different fixes.&lt;/p></description></item><item><title>Traefik high CPU usage: TLS handshakes, regex rules, and config rebuilds</title><link>https://www.netdata.cloud/guides/traefik/traefik-high-cpu-usage/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-high-cpu-usage/</guid><description>&lt;h1 id="traefik-high-cpu-usage-tls-handshakes-regex-rules-and-config-rebuilds">Traefik high CPU usage: TLS handshakes, regex rules, and config rebuilds&lt;/h1>
&lt;p>Traefik is pegging cores and request latency is climbing. The process is alive, &lt;code>/ping&lt;/code> returns 200, and traffic is still flowing, but p99 latency is creeping up and TLS connections are slower to establish. You need to know which of the three usual suspects is burning the CPU: cryptographic work, middleware processing, or configuration rebuilds.&lt;/p>
&lt;p>Traefik CPU is dominated by TLS handshake computation (especially with RSA certificates), then middleware chain processing (gzip compression, JWT validation, regex-based routing rules), then configuration rebuilds whose cost scales with routers times middlewares times services. The diagnosis is almost entirely correlation: match the CPU curve against the TLS handshake rate, the config reload rate, and the request rate, and the cause usually names itself.&lt;/p></description></item><item><title>Traefik high request latency: isolating Traefik overhead from backend slowness</title><link>https://www.netdata.cloud/guides/traefik/traefik-high-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-high-latency/</guid><description>&lt;h1 id="traefik-high-request-latency-isolating-traefik-overhead-from-backend-slowness">Traefik high request latency: isolating Traefik overhead from backend slowness&lt;/h1>
&lt;p>Latency dashboards are red, p95 is climbing, and the first question from the incident channel is &amp;ldquo;is it the proxy or the app?&amp;rdquo; With Traefik in the path, that question has a precise answer if you compare the right two metrics and read the JSON access-log fields most operators skip.&lt;/p>
&lt;p>Traefik measures request duration at two points in its pipeline. &lt;code>traefik_entrypoint_request_duration_seconds&lt;/code> measures the request at the edge, including TLS, middleware processing, and backend time. &lt;code>traefik_service_request_duration_seconds&lt;/code> measures only the leg from Traefik to the backend and back. The difference between the two is Traefik&amp;rsquo;s own overhead. If entrypoint latency is high but service latency is normal, the time is being burned inside Traefik itself: TLS handshakes, middleware chain processing, gzip compression, response buffering, or an access-log write that is blocking request goroutines. If both are high, the backend is slow and Traefik is just the messenger.&lt;/p></description></item><item><title>Traefik intermittent 502 behind a load balancer: the timeout chain mismatch</title><link>https://www.netdata.cloud/guides/traefik/traefik-timeout-chain-502/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-timeout-chain-502/</guid><description>&lt;h1 id="traefik-intermittent-502-behind-a-load-balancer-the-timeout-chain-mismatch">Traefik intermittent 502 behind a load balancer: the timeout chain mismatch&lt;/h1>
&lt;p>Some percentage of requests behind your cloud load balancer return 502. Not all of them, just a low-grade drip: one in a few hundred, sometimes one in a few thousand. Backend logs show nothing. Traefik logs show a burst of 502s with no corresponding backend error. Load balancer health checks are green. Everything looks healthy, and users keep hitting errors.&lt;/p></description></item><item><title>Traefik memory usage growing: routing table size and heap trends</title><link>https://www.netdata.cloud/guides/traefik/traefik-memory-usage-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-memory-usage-growing/</guid><description>&lt;h1 id="traefik-memory-usage-growing-routing-table-size-and-heap-trends">Traefik memory usage growing: routing table size and heap trends&lt;/h1>
&lt;p>&lt;code>process_resident_memory_bytes&lt;/code> on your Traefik instance has been climbing for days. Request rates look normal, error rates are flat, and &lt;code>/ping&lt;/code> returns 200. The question is whether the proxy legitimately needs more memory because its routing table grew, or whether it is leaking and will get OOM-killed at 3 a.m. with zero warning.&lt;/p>
&lt;p>The failure mode is asymmetric. Go&amp;rsquo;s garbage collector absorbs growing allocation pressure gracefully, running more often and burning more CPU, until it cannot. Then the kernel OOM killer terminates the process instantly, dropping every connection in flight. There is no graceful degradation phase.&lt;/p></description></item><item><title>Traefik Monitoring</title><link>https://www.netdata.cloud/monitoring-101/traefik-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/traefik-monitoring/</guid><description>&lt;h2 id="traefik-monitoring">Traefik Monitoring&lt;/h2>
&lt;h3 id="what-is-traefik">What Is Traefik?&lt;/h3>
&lt;p>Traefik is a popular open-source edge router and load balancer that helps manage your API traffic seamlessly. It provides an easy way to access your services by automatically configuring routes and managing the flow of requests between clients and your microservices. With its capabilities to handle dynamic adaptive routing, it is a suitable choice for containerized environments and microservices architectures.&lt;/p>
&lt;h3 id="monitoring-traefik-with-netdata">Monitoring Traefik With Netdata&lt;/h3>
&lt;p>Monitoring Traefik with Netdata is a straightforward process that allows you to leverage real-time data visualization and comprehensive monitoring. Netdata’s &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/traefik/">Traefik monitoring tool&lt;/a> is designed specifically to monitor the performance and health of your Traefik instances. By utilizing Netdata, you can gain insight into traffic patterns, server load, and potential bottlenecks with minimal configuration.&lt;/p></description></item><item><title>Traefik monitoring checklist: the signals every production edge router needs</title><link>https://www.netdata.cloud/guides/traefik/traefik-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-monitoring-checklist/</guid><description>&lt;h1 id="traefik-monitoring-checklist-the-signals-every-production-edge-router-needs">Traefik monitoring checklist: the signals every production edge router needs&lt;/h1>
&lt;p>Traefik fails in ways that static proxies do not. It is both a data-plane proxy and a control-plane configuration reconciler, and its worst failure modes are silent: a dead provider connection, a stalled ACME renewal, a file descriptor limit creeping toward exhaustion. The process stays up, &lt;code>/ping&lt;/code> returns 200, and traffic keeps flowing on stale configuration while new deployments get 404s.&lt;/p></description></item><item><title>Traefik monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/traefik/traefik-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-monitoring-maturity-model/</guid><description>&lt;h1 id="traefik-monitoring-maturity-model-from-survival-to-expert">Traefik monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most Traefik outages are not exotic monitoring gaps. They are the same six failures repeating across teams: the process died, the FD limit was 1024, the cert expired, the provider went silent, the backend pool collapsed, or retries amplified a partial failure into a total one. A maturity model helps because it sequences coverage against the failures you are actually going to have, in the order you are going to have them.&lt;/p></description></item><item><title>Traefik open connections growing: spotting a connection leak</title><link>https://www.netdata.cloud/guides/traefik/traefik-open-connections-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-open-connections-growing/</guid><description>&lt;h1 id="traefik-open-connections-growing-spotting-a-connection-leak">Traefik open connections growing: spotting a connection leak&lt;/h1>
&lt;p>The &lt;code>traefik_open_connections&lt;/code> gauge is climbing. Request rate is flat. Every hour the number is higher than the last, and it never comes back down, even after the traffic peak passes. That divergence is the signature of a connection leak: connections are opened and never closed, and each one holds a file descriptor, a goroutine, and some memory.&lt;/p>
&lt;p>Left alone, this ends one way. Traefik hits its file descriptor limit and stops accepting new connections instantly. Existing connections keep working, so some health checks still pass, but every new client is refused. It is a cliff-edge failure with no graceful degradation, and the climb gives you hours or days of warning if you are watching the right signal.&lt;/p></description></item><item><title>Traefik out of memory: OOM kills and the crash loop that follows</title><link>https://www.netdata.cloud/guides/traefik/traefik-out-of-memory-oom/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-out-of-memory-oom/</guid><description>&lt;h1 id="traefik-out-of-memory-oom-kills-and-the-crash-loop-that-follows">Traefik out of memory: OOM kills and the crash loop that follows&lt;/h1>
&lt;p>Traefik was proxying traffic normally. Then the process vanished. In Kubernetes you see &lt;code>OOMKilled&lt;/code> in the pod&amp;rsquo;s last terminated state and a restart count that keeps climbing. On bare metal or Docker you see exit code 137 and a container that came back up, ran for a while, and died again.&lt;/p>
&lt;p>An OOM kill has no graceful degradation phase. Go&amp;rsquo;s garbage collector absorbs growing memory pressure by running more often, right up until the moment it cannot. Then the kernel OOM killer terminates the process instantly. Every in-flight connection drops at once: client connections, backend connections, WebSockets, gRPC streams. From the outside it looks like a total, simultaneous outage of everything behind the proxy.&lt;/p></description></item><item><title>Traefik request rate monitoring: entrypoint and service throughput baselines</title><link>https://www.netdata.cloud/guides/traefik/traefik-request-rate-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-request-rate-monitoring/</guid><description>&lt;h1 id="traefik-request-rate-monitoring-entrypoint-and-service-throughput-baselines">Traefik request rate monitoring: entrypoint and service throughput baselines&lt;/h1>
&lt;p>Traefik exposes two request counters that look similar but measure different things. &lt;code>traefik_entrypoint_requests_total&lt;/code> counts every request that arrives at a listener, whether or not Traefik finds a route for it. &lt;code>traefik_service_requests_total&lt;/code> counts only requests that matched a router, survived the middleware chain, and were proxied to a backend. The difference between the two is where most routing and middleware problems show up first.&lt;/p></description></item><item><title>Traefik retry amplification: how retries turn a slow backend into an outage</title><link>https://www.netdata.cloud/guides/traefik/traefik-retry-amplification/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-retry-amplification/</guid><description>&lt;h1 id="traefik-retry-amplification-how-retries-turn-a-slow-backend-into-an-outage">Traefik retry amplification: how retries turn a slow backend into an outage&lt;/h1>
&lt;p>A backend service starts degrading. Responses get slow, some connections fail. Traefik&amp;rsquo;s retry middleware re-sends the failed requests, usually to other backends in the pool. Backend load doubles or triples. The struggling service falls further behind, which produces more failures, which produces more retries. Within minutes, a partial degradation becomes a complete outage, and Traefik is doing a significant share of the damage while trying to help.&lt;/p></description></item><item><title>Traefik route not loading: silently ignored annotations and labels</title><link>https://www.netdata.cloud/guides/traefik/traefik-route-not-loading/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-route-not-loading/</guid><description>&lt;h1 id="traefik-route-not-loading-silently-ignored-annotations-and-labels">Traefik route not loading: silently ignored annotations and labels&lt;/h1>
&lt;p>You deployed a service, added the Traefik annotations or labels, and the route does not work. Requests to the hostname return 404. Traefik is up, other routes work, the provider is connected, and there is no error anywhere in the logs. The service simply does not exist as far as Traefik is concerned.&lt;/p>
&lt;p>This is silent rejection: Traefik ignores any annotation or label key it does not recognize. A wrong prefix, a misspelled key, a label at the wrong YAML level, or a value mangled by YAML parsing all produce the same outcome: the route is never created, no error metric is emitted, and nothing is logged at the default log level. The configuration you wrote and the configuration Traefik loaded diverge without either side complaining.&lt;/p></description></item><item><title>Traefik router priority conflicts: when overlapping rules match the wrong service</title><link>https://www.netdata.cloud/guides/traefik/traefik-router-priority-conflict/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-router-priority-conflict/</guid><description>&lt;h1 id="traefik-router-priority-conflicts-when-overlapping-rules-match-the-wrong-service">Traefik router priority conflicts: when overlapping rules match the wrong service&lt;/h1>
&lt;p>Traffic for &lt;code>api.example.com&lt;/code> is landing on your generic dashboard service instead of the API backend. Or you deployed a new router with what looks like a correct rule, and it never matches a single request. Traefik is up, &lt;code>/ping&lt;/code> returns 200, there are no errors in the logs, and the config reload metric keeps advancing. Everything looks healthy except the routing decision itself.&lt;/p></description></item><item><title>Traefik routing table explosion: thousands of routes and slow rebuilds</title><link>https://www.netdata.cloud/guides/traefik/traefik-routing-table-explosion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-routing-table-explosion/</guid><description>&lt;h1 id="traefik-routing-table-explosion-thousands-of-routes-and-slow-rebuilds">Traefik routing table explosion: thousands of routes and slow rebuilds&lt;/h1>
&lt;p>Your Traefik instance is still passing traffic, but something is off. Memory climbs week over week. Every Ingress change causes a CPU blip and a small latency spike. Restart counts are creeping up, and the last restart was an OOM kill. When you look at the provider, you find thousands of Ingress or IngressRoute objects that nobody ever pruned.&lt;/p>
&lt;p>This is the routing table explosion failure mode. Traefik rebuilds its entire routing table on every configuration change, holds the whole thing in memory, and pays a per-request matching cost for every router it knows about. The failure is slow and linear right up until the OOM killer makes it instant.&lt;/p></description></item><item><title>Traefik service retries climbing: reading the retry signal before it cascades</title><link>https://www.netdata.cloud/guides/traefik/traefik-service-retries-total/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-service-retries-total/</guid><description>&lt;h1 id="traefik-service-retries-climbing-reading-the-retry-signal-before-it-cascades">Traefik service retries climbing: reading the retry signal before it cascades&lt;/h1>
&lt;p>&lt;code>traefik_service_retries_total&lt;/code> is climbing for one of your services. Your dashboards show green: clients are getting 200s, the error rate looks flat, and &lt;code>/ping&lt;/code> is happy. That is exactly what makes this signal dangerous. Retries mask backend instability from clients while multiplying the load Traefik sends to the backends. By the time client-facing errors appear, the amplification loop may already be running.&lt;/p></description></item><item><title>Traefik service server up at zero: backend health checks are failing</title><link>https://www.netdata.cloud/guides/traefik/traefik-service-server-up-zero/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-service-server-up-zero/</guid><description>&lt;h1 id="traefik-service-server-up-at-zero-backend-health-checks-are-failing">Traefik service server up at zero: backend health checks are failing&lt;/h1>
&lt;p>&lt;code>traefik_service_server_up&lt;/code> is 0 for one or more URLs in a service. Traefik&amp;rsquo;s health checker has pulled those backends out of the load balancer rotation. If every URL in the service is at 0, Traefik has nowhere to send requests and is returning 503 to clients right now.&lt;/p>
&lt;p>This is a backend-side symptom, not a Traefik-side one. Traefik is doing what you configured it to do: stop sending traffic to servers that fail its checks. The useful question is whether the checks are telling the truth about the backends, and if so, why the backends are failing.&lt;/p></description></item><item><title>Traefik serving stale configuration: the silent provider desync</title><link>https://www.netdata.cloud/guides/traefik/traefik-provider-desync-stale-config/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-provider-desync-stale-config/</guid><description>&lt;h1 id="traefik-serving-stale-configuration-the-silent-provider-desync">Traefik serving stale configuration: the silent provider desync&lt;/h1>
&lt;p>A developer deploys a new service. The pods are running, the Ingress or IngressRoute exists, everything looks right in Kubernetes. But the service returns 404 from the edge. Meanwhile, a service deleted an hour ago is still receiving traffic and failing. You check Traefik: process is up, &lt;code>/ping&lt;/code> returns 200, existing routes work. Nothing is alerting.&lt;/p>
&lt;p>This is the silent provider desync, and it is the most commonly missed Traefik failure mode. The provider watch (Kubernetes API, Docker socket, Consul, file) has died or fallen behind. Traefik does not flush its routes when this happens. It keeps the last-known-good configuration and retries the provider connection in the background, with no loud failure signal. The routing table freezes at the moment of disconnection.&lt;/p></description></item><item><title>Traefik serving the default certificate: SNI and certificate selection</title><link>https://www.netdata.cloud/guides/traefik/traefik-default-certificate-served/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-default-certificate-served/</guid><description>&lt;h1 id="traefik-serving-the-default-certificate-sni-and-certificate-selection">Traefik serving the default certificate: SNI and certificate selection&lt;/h1>
&lt;p>A client hits your site and gets a certificate warning. The certificate presented is a self-signed cert with CN &amp;ldquo;TRAEFIK DEFAULT CERT&amp;rdquo; instead of the certificate you configured. Everything else looks fine: Traefik is up, &lt;code>/ping&lt;/code> returns 200, backends are healthy, routes work. The proxy is completely healthy and the only broken thing is which certificate it picked during the TLS handshake.&lt;/p></description></item><item><title>Traefik TLS certificate expiry monitoring: reading traefik_tls_certs_not_after</title><link>https://www.netdata.cloud/guides/traefik/traefik-tls-certs-not-after/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-tls-certs-not-after/</guid><description>&lt;h1 id="traefik-tls-certificate-expiry-monitoring-reading-traefik_tls_certs_not_after">Traefik TLS certificate expiry monitoring: reading traefik_tls_certs_not_after&lt;/h1>
&lt;p>&lt;code>traefik_tls_certs_not_after&lt;/code> is the only signal Traefik gives you about certificate expiry. There is no metric for ACME renewal failures, no counter for rate limit rejections, no gauge for &amp;ldquo;the challenge did not complete.&amp;rdquo; You get a Unix timestamp per certificate, and everything else has to be inferred from it or pulled from logs.&lt;/p>
&lt;p>Most teams wire it to a pager with a threshold like &amp;ldquo;alert at 7 days.&amp;rdquo; Then the first ACME renewal happens, the old certificate series stays in the output, and the pager fires for a certificate that is no longer serving anything. Or the alert fires for a staging certificate, or for a dormant cert left in the store from a decommissioned hostname, and the on-call learns to ignore it. This guide covers reading the metric correctly and building an alert chain that only wakes someone up when a certificate that is actually serving production traffic is about to expire.&lt;/p></description></item><item><title>Traefik TLS handshake latency: RSA cost, session resumption, and CPU</title><link>https://www.netdata.cloud/guides/traefik/traefik-tls-handshake-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-tls-handshake-latency/</guid><description>&lt;h1 id="traefik-tls-handshake-latency-rsa-cost-session-resumption-and-cpu">Traefik TLS handshake latency: RSA cost, session resumption, and CPU&lt;/h1>
&lt;p>Clients report slow first page loads or high connect times, but backend latency looks fine and error rates are flat. When the delay shows up before the first byte of the request is processed, the TLS handshake at the Traefik entrypoint is the prime suspect.&lt;/p>
&lt;p>Handshakes are the largest per-connection CPU cost for a terminating proxy, and they add latency to every new connection before any routing, middleware, or backend work happens. Traefik exposes no dedicated handshake-latency metric, so you have to infer it from latency decomposition and process CPU.&lt;/p></description></item><item><title>Traefik TLS version monitoring: catching TLS 1.0/1.1 legacy traffic</title><link>https://www.netdata.cloud/guides/traefik/traefik-tls-version-compliance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-tls-version-compliance/</guid><description>&lt;h1 id="traefik-tls-version-monitoring-catching-tls-1011-legacy-traffic">Traefik TLS version monitoring: catching TLS 1.0/1.1 legacy traffic&lt;/h1>
&lt;p>TLS 1.0 and 1.1 have been deprecated for years, and compliance regimes like PCI DSS require TLS 1.2 as a minimum. The risky part of setting a minimum TLS version is not the config change. It is finding out, before you enforce the floor, which clients still negotiate old protocol versions. Flip &lt;code>minVersion&lt;/code> blind and you will discover the legacy client population from user complaints.&lt;/p></description></item><item><title>Traefik too many open files: file descriptor exhaustion at the edge</title><link>https://www.netdata.cloud/guides/traefik/traefik-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-too-many-open-files/</guid><description>&lt;h1 id="traefik-too-many-open-files-file-descriptor-exhaustion-at-the-edge">Traefik too many open files: file descriptor exhaustion at the edge&lt;/h1>
&lt;p>Your Traefik logs are filling with &lt;code>accept4: too many open files&lt;/code> errors. New clients cannot connect. Requests that do get through may return 502 or 503. Meanwhile, some clients insist everything works fine, because their existing connections are still alive.&lt;/p>
&lt;p>This is file descriptor (FD) exhaustion. Traefik sits at the edge and holds at least two FDs per proxied connection: one for the client side, one for the backend side. Add provider connections (Docker socket, Kubernetes API watches), log files, and ACME storage handles, and the count climbs fast. When it hits the OS limit, every new &lt;code>accept4()&lt;/code> call fails instantly.&lt;/p></description></item><item><title>Traefik traffic dropped to zero: entrypoint request rate flatlined</title><link>https://www.netdata.cloud/guides/traefik/traefik-traffic-drop-to-zero/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-traffic-drop-to-zero/</guid><description>&lt;h1 id="traefik-traffic-dropped-to-zero-entrypoint-request-rate-flatlined">Traefik traffic dropped to zero: entrypoint request rate flatlined&lt;/h1>
&lt;p>The request rate on an entrypoint that normally serves traffic has fallen to zero, or close enough. Clients are timing out or getting connection errors. Sometimes Traefik&amp;rsquo;s process is still running and &lt;code>/ping&lt;/code> still returns 200, which makes this worse: your health checks are green while no traffic is flowing.&lt;/p>
&lt;p>There are two fundamentally different families of root cause, and telling them apart early is the whole game. Either traffic is not reaching Traefik at all (DNS, cloud load balancer, firewall, network partition upstream of the proxy), or traffic is arriving at the host but Traefik cannot accept it (a dead listener, file descriptor exhaustion, a port-bind failure after restart). The checks below are ordered to split those two families within the first few minutes.&lt;/p></description></item><item><title>Traefik traffic volume: request and response bytes for capacity planning</title><link>https://www.netdata.cloud/guides/traefik/traefik-traffic-volume-bytes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-traffic-volume-bytes/</guid><description>&lt;h1 id="traefik-traffic-volume-request-and-response-bytes-for-capacity-planning">Traefik traffic volume: request and response bytes for capacity planning&lt;/h1>
&lt;p>Request rate and error rate tell you how much traffic Traefik is handling and whether it is succeeding. They do not tell you how much data is moving. A proxy serving ten thousand small API calls per second and a proxy serving ten thousand large file downloads per second look identical on a requests-per-second dashboard, but they have completely different bandwidth, memory, and buffering profiles.&lt;/p></description></item><item><title>Traefik under scanning: exploit-path probing and request smuggling signals</title><link>https://www.netdata.cloud/guides/traefik/traefik-scanning-probing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/traefik/traefik-scanning-probing/</guid><description>&lt;h1 id="traefik-under-scanning-exploit-path-probing-and-request-smuggling-signals">Traefik under scanning: exploit-path probing and request smuggling signals&lt;/h1>
&lt;p>Any Traefik instance with a public entrypoint is being scanned right now. Requests for &lt;code>/.env&lt;/code>, &lt;code>/.git/config&lt;/code>, &lt;code>/wp-admin&lt;/code>, &lt;code>/phpMyAdmin&lt;/code>, and &lt;code>/actuator&lt;/code> arrive continuously from botnets and research scanners, and on a correctly configured instance they all get the same answer: a 404 generated by Traefik itself, because no router matches. That is normal background radiation, and paging on it will burn out your on-call in a week.&lt;/p></description></item><item><title>Trango Networks LLC SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/trango-networks-llc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/trango-networks-llc-snmp-traps/</guid><description/></item><item><title>Transition Engineering Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/transition-engineering-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/transition-engineering-inc-snmp-traps/</guid><description/></item><item><title>Trap and syslog flood from link flaps: surviving the storm</title><link>https://www.netdata.cloud/guides/network/network-trap-syslog-flood/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-trap-syslog-flood/</guid><description>&lt;h1 id="trap-and-syslog-flood-from-link-flaps-surviving-the-storm">Trap and syslog flood from link flaps: surviving the storm&lt;/h1>
&lt;p>A single bad SFP starts flapping. Within seconds, your trap receiver is processing hundreds of linkDown/linkUp pairs per second, your syslog pipeline is drowning in LINK-3-UPDOWN messages, and STP topology change notifications are cascading across the L2 domain. The kernel socket buffer on UDP 162 overflows, and the root-cause hardware alarm is as likely to be dropped as any other datagram in the flood.&lt;/p></description></item><item><title>Trapeze Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/trapeze-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/trapeze-networks-inc-snmp-traps/</guid><description/></item><item><title>Trend Micro Inc Formerly Tippingpoint Technologies SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/trend-micro-inc-formerly-tippingpoint-technologies-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/trend-micro-inc-formerly-tippingpoint-technologies-snmp-traps/</guid><description/></item><item><title>Trend Micro Incorporated SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/trend-micro-incorporated-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/trend-micro-incorporated-snmp-traps/</guid><description/></item><item><title>Tridium SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tridium-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tridium-snmp-traps/</guid><description/></item><item><title>Tripp Lite SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tripp-lite-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tripp-lite-snmp-traps/</guid><description/></item><item><title>Tripplite</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/tripplite/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/tripplite/</guid><description/></item><item><title>Tripplite PDU</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/tripplite-pdu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/tripplite-pdu/</guid><description/></item><item><title>Tripplite UPS</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/tripplite-ups/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/tripplite-ups/</guid><description/></item><item><title>Tu Braunschweig SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tu-braunschweig-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tu-braunschweig-snmp-traps/</guid><description/></item><item><title>Twilio</title><link>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/twilio/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/agent-dispatched-notifications/twilio/</guid><description/></item><item><title>Twitch</title><link>https://www.netdata.cloud/integrations/data-collection/applications/twitch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/twitch/</guid><description/></item><item><title>Twitch Monitoring</title><link>https://www.netdata.cloud/monitoring-101/twitch-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/twitch-monitoring/</guid><description>&lt;h2 id="twitch-monitoring">Twitch Monitoring&lt;/h2>
&lt;h3 id="what-is-twitch">What Is Twitch?&lt;/h3>
&lt;p>Twitch is a leading live streaming platform primarily for gamers, but it is also home to other types of creative content and interactive entertainment. With massive live audiences and real-time interaction possibilities, ensuring optimal performance and smooth streaming experiences for both broadcasters and viewers is crucial.&lt;/p>
&lt;h3 id="monitoring-twitch-with-netdata">Monitoring Twitch With Netdata&lt;/h3>
&lt;p>To monitor Twitch effectively, Netdata employs an openmetrics (Prometheus) exporter. The &lt;a href="https://github.com/damoun/twitch_exporter">Twitch exporter&lt;/a> collects essential metrics, providing a comprehensive overview of the performance and health of your Twitch streams. Netdata excels in leveraging this data by offering automated dashboards, alerts, and insights, all without the need for a dedicated Prometheus server or Grafana setup. This means you can monitor Twitch streams efficiently, gaining crucial insights with minimal setup, directly from the Netdata platform.&lt;/p></description></item><item><title>Tylink SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tylink-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/tylink-snmp-traps/</guid><description/></item><item><title>Typesense</title><link>https://www.netdata.cloud/integrations/data-collection/databases/typesense/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/typesense/</guid><description/></item><item><title>Typesense Monitoring</title><link>https://www.netdata.cloud/monitoring-101/typesense-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/typesense-monitoring/</guid><description>&lt;h2 id="typesense-monitoring">Typesense Monitoring&lt;/h2>
&lt;h3 id="what-is-typesense">What Is Typesense?&lt;/h3>
&lt;p>Typesense is an open-source, in-memory search engine that delivers fast, instant search results right out of the box. It is designed to help you build typo-tolerant and real-time search experiences without the complexity that comes with integrating with heavyweight traditional search engines. With robust features like multi-tenant service, extensive language support, and real-time search updates, Typesense is a preferred choice for many developers.&lt;/p>
&lt;h3 id="monitoring-typesense-with-netdata">Monitoring Typesense With Netdata&lt;/h3>
&lt;p>Netdata offers an innovative and user-friendly experience for monitoring Typesense. As a comprehensive monitoring tool, Netdata allows you to gain insights into the health and performance of your Typesense servers, with real-time data collection and visualization. With &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata&amp;rsquo;s live demo&lt;/a>, you can see Netdata&amp;rsquo;s capabilities in action.&lt;/p></description></item><item><title>U C Davis Ece Dept Tom SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/u-c-davis-ece-dept-tom-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/u-c-davis-ece-dept-tom-snmp-traps/</guid><description/></item><item><title>Ubiquiti UFiber OLT</title><link>https://www.netdata.cloud/integrations/data-collection/networking/ubiquiti-ufiber-olt/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/ubiquiti-ufiber-olt/</guid><description/></item><item><title>Ubiquiti UFiber OLT Monitoring</title><link>https://www.netdata.cloud/monitoring-101/ubiquity_ufiber-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/ubiquity_ufiber-monitoring/</guid><description>&lt;h2 id="ubiquiti-ufiber-olt-monitoring">Ubiquiti UFiber OLT Monitoring&lt;/h2>
&lt;h3 id="what-is-ubiquiti-ufiber-olt">What Is Ubiquiti UFiber OLT?&lt;/h3>
&lt;p>Ubiquiti UFiber OLT is a robust solution for fiber-optic communication networks, designed to connect with multiple endpoints. It&amp;rsquo;s essential for managing and maintaining high-performance fiber-to-the-premises installations efficiently.&lt;/p>
&lt;h3 id="monitoring-ubiquiti-ufiber-olt-with-netdata">Monitoring Ubiquiti UFiber OLT With Netdata&lt;/h3>
&lt;p>Monitoring Ubiquiti UFiber OLT with Netdata provides invaluable insights into network performance and reliability. Netdata uses an openmetrics (Prometheus) exporter to monitor Ubiquiti UFiber OLT. It can ingest data seamlessly from any Prometheus exporter, allowing for automated dashboards, alerts, and more, without the need for a Prometheus server or Grafana. This integration empowers users to have real-time monitoring capabilities with comprehensive metrics analysis.&lt;/p></description></item><item><title>Ubiquiti Unifi</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ubiquiti-unifi/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ubiquiti-unifi/</guid><description/></item><item><title>Ubiquiti Unifi Security Gateway</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ubiquiti-unifi-security-gateway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/ubiquiti-unifi-security-gateway/</guid><description/></item><item><title>Ubuntu</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/ubuntu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/ubuntu/</guid><description/></item><item><title>Ucopia Communications SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ucopia-communications-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ucopia-communications-snmp-traps/</guid><description/></item><item><title>UDMA_CRC_Error_Count rising: it's the cable, not the drive</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-udma-crc-error-count/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-udma-crc-error-count/</guid><description>&lt;h1 id="udma_crc_error_count-rising-its-the-cable-not-the-drive">UDMA_CRC_Error_Count rising: it&amp;rsquo;s the cable, not the drive&lt;/h1>
&lt;p>Rising UDMA CRC Error Count (SMART attribute 199) is one of the most misdiagnosed signals in production. Operators see the counter climbing, swap the drive, and the replacement develops the same errors because the cable, backplane port, or HBA was the actual cause. The original drive goes back as an RMA, the replacement fails the same way, and the real fault is still in the chassis.&lt;/p></description></item><item><title>Udp_RcvbufErrors: tuning kernel receive buffers for flow, trap, and syslog collectors</title><link>https://www.netdata.cloud/guides/network/network-udp-rcvbuf-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-udp-rcvbuf-errors/</guid><description>&lt;h1 id="udp_rcvbuferrors-tuning-kernel-receive-buffers-for-flow-trap-and-syslog-collectors">Udp_RcvbufErrors: tuning kernel receive buffers for flow, trap, and syslog collectors&lt;/h1>
&lt;p>&lt;code>Udp_RcvbufErrors&lt;/code> is incrementing on your flow collector. Flow charts show traffic declining during what is actually a traffic spike. The kernel is receiving datagrams from exporters but the socket receive buffer is full, so it drops them silently. No application-level counter moves. No error log fires. The dashboards lie downward while the real traffic goes upward.&lt;/p>
&lt;p>Flow collectors (NetFlow v5/v9, IPFIX, sFlow), SNMP trap receivers (UDP 162), and syslog receivers (UDP 514) all depend on UDP socket buffers. When the buffer overflows, the kernel increments &lt;code>Udp_RcvbufErrors&lt;/code> in &lt;code>/proc/net/snmp&lt;/code> and discards the datagram. The application never sees it.&lt;/p></description></item><item><title>UK Mod De S SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/uk-mod-de-s-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/uk-mod-de-s-snmp-traps/</guid><description/></item><item><title>Unauthorized NVIDIA GPU processes: detecting cryptomining and rogue workloads</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-unauthorized-process-cryptomining/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-unauthorized-process-cryptomining/</guid><description>&lt;h1 id="unauthorized-nvidia-gpu-processes-detecting-cryptomining-and-rogue-workloads">Unauthorized NVIDIA GPU processes: detecting cryptomining and rogue workloads&lt;/h1>
&lt;p>A GPU node that should be idle shows sustained utilization. Or a scheduler reports a job cannot land because memory is held, yet no known workload is running. &lt;code>nvidia-smi&lt;/code> shows a process named &lt;code>python3&lt;/code> burning 99% SM utilization on a box that has no Python workload scheduled. The question to answer fast: is this an expected workload you have lost track of, or something that should not be there at all?&lt;/p></description></item><item><title>Unbound</title><link>https://www.netdata.cloud/integrations/data-collection/networking/unbound/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/unbound/</guid><description/></item><item><title>Unbound Monitoring</title><link>https://www.netdata.cloud/monitoring-101/unbound-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/unbound-monitoring/</guid><description>&lt;h2 id="unbound-monitoring">Unbound Monitoring&lt;/h2>
&lt;h3 id="what-is-unbound">What Is Unbound?&lt;/h3>
&lt;p>Unbound is a validating, recursive, and caching DNS resolver designed to provide robust DNS security features. As a part of network infrastructure, it plays a crucial role in efficiently resolving domain names to IP addresses while keeping caches of previous queries to expedite future lookups. Learn more about &lt;a href="https://nlnetlabs.nl/projects/unbound/about/">Unbound&lt;/a>.&lt;/p>
&lt;h3 id="monitoring-unbound-with-netdata">Monitoring Unbound With Netdata&lt;/h3>
&lt;p>To effectively monitor Unbound, Netdata provides a comprehensive solution with its powerful &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/unbound/">Unbound monitoring tool&lt;/a>. Netdata&amp;rsquo;s real-time metrics and beautiful visualizations help DevOps, SREs, developers, IT admins, and engineers maintain optimal DNS performance. Explore the &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">Live Demo&lt;/a> to see Netdata&amp;rsquo;s capabilities in action.&lt;/p></description></item><item><title>Unexpected drive firmware or identity change: counter resets and identity drift</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-firmware-version-change/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-firmware-version-change/</guid><description>&lt;h1 id="unexpected-drive-firmware-or-identity-change-counter-resets-and-identity-drift">Unexpected drive firmware or identity change: counter resets and identity drift&lt;/h1>
&lt;p>A cumulative SMART counter decreased between polls. Power-On Hours went backward. The drive&amp;rsquo;s model, serial number, or firmware version changed since the last snapshot. SMART counters are designed to be monotonically increasing. Identity fields are set at the factory. When either changes without a documented reason, determine whether this is routine maintenance, a firmware update, or something requiring forensic review.&lt;/p></description></item><item><title>Unisys Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/unisys-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/unisys-corp-snmp-traps/</guid><description/></item><item><title>Unitrends Software Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/unitrends-software-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/unitrends-software-corp-snmp-traps/</guid><description/></item><item><title>Unsafe Shutdowns / unexpected power-loss count rising</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-unsafe-shutdowns/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-unsafe-shutdowns/</guid><description>&lt;h1 id="unsafe-shutdowns--unexpected-power-loss-count-rising">Unsafe Shutdowns / unexpected power-loss count rising&lt;/h1>
&lt;p>The unsafe shutdown counter is a cumulative lifetime metric tracking how many times a drive lost power without receiving a clean shutdown notification. On ATA drives, this is attribute 174 (&lt;code>Unexpect_Power_Loss_Ct&lt;/code>) on some SSDs or attribute 192 (&lt;code>Power-Off_Retract_Count&lt;/code>) on HDDs. On NVMe drives, it is the &amp;ldquo;Unsafe Shutdowns&amp;rdquo; field in the SMART/Health Information Log (Log Page 02h).&lt;/p>
&lt;p>A non-zero value is not inherently alarming. Drives accumulate unsafe shutdowns over their lifetime from kernel panics, hard resets, UPS failures, and factory burn-in. The operational question is never &amp;ldquo;is the count non-zero?&amp;rdquo; but &amp;ldquo;is it growing, and how fast?&amp;rdquo;&lt;/p></description></item><item><title>UPS (NUT)</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/ups-nut/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/ups-nut/</guid><description/></item><item><title>UPS (NUT) Monitoring</title><link>https://www.netdata.cloud/monitoring-101/upsd-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/upsd-monitoring/</guid><description>&lt;h2 id="ups-nut-monitoring">UPS (NUT) Monitoring&lt;/h2>
&lt;h3 id="what-is-ups-nut">What Is UPS (NUT)?&lt;/h3>
&lt;p>UPS (NUT) Monitoring involves overseeing the Uninterruptible Power Supplies that keep your systems operational during power outages. Utilizing the Network UPS Tools (NUT) framework, it enables monitoring and management of power devices from widely-used brands, ensuring that your infrastructure remains protected even when the power grid fails.&lt;/p>
&lt;h3 id="monitoring-ups-nut-with-netdata">Monitoring UPS (NUT) With Netdata&lt;/h3>
&lt;p>Netdata&amp;rsquo;s comprehensive and intuitive monitoring tool offers real-time insights into your power systems by seamlessly integrating with the UPS daemon via the NUT protocol. With Netdata, you can monitor critical parameters like load, battery status, temperature, and input/output voltages in real-time, empowering you to maintain optimal power conditions and prevent unexpected downtimes.&lt;/p></description></item><item><title>Ups Manufacturing SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ups-manufacturing-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/ups-manufacturing-snmp-traps/</guid><description/></item><item><title>uptime</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/uptime/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/uptime/</guid><description/></item><item><title>Uptimerobot</title><link>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/uptimerobot/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/uptimerobot/</guid><description/></item><item><title>Uptimerobot Monitoring</title><link>https://www.netdata.cloud/monitoring-101/uptimerobot-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/uptimerobot-monitoring/</guid><description>&lt;h2 id="uptimerobot-monitoring">Uptimerobot Monitoring&lt;/h2>
&lt;h3 id="what-is-uptimerobot">What Is Uptimerobot?&lt;/h3>
&lt;p>Uptimerobot is a popular service that provides website uptime monitoring to ensure that your applications and services are available and performing optimally. It&amp;rsquo;s an essential tool for DevOps, Site Reliability Engineers (SREs), and IT administrators seeking to maintain high availability and reduce downtime.&lt;/p>
&lt;h3 id="monitoring-uptimerobot-with-netdata">Monitoring Uptimerobot With Netdata&lt;/h3>
&lt;p>Monitoring Uptimerobot with Netdata allows you to maintain a reliable overview of your website&amp;rsquo;s uptime metrics. To monitor Uptimerobot, Netdata employs an OpenMetrics (Prometheus) exporter which can seamlessly ingest data from any Prometheus exporter. This provides users with automated dashboards and alerts, enabling real-time monitoring and troubleshooting without the need for a dedicated Prometheus server or Grafana.&lt;/p></description></item><item><title>User Groups</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/user-groups/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/user-groups/</guid><description/></item><item><title>Users</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/users/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/users/</guid><description/></item><item><title>Utstarcom Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/utstarcom-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/utstarcom-inc-snmp-traps/</guid><description/></item><item><title>Utstarcom Incorporated SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/utstarcom-incorporated-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/utstarcom-incorporated-snmp-traps/</guid><description/></item><item><title>uWSGI</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/uwsgi/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/uwsgi/</guid><description/></item><item><title>uWSGI all workers busy: reading the busy ratio before the queue fills</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-all-workers-busy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-all-workers-busy/</guid><description>&lt;h1 id="uwsgi-all-workers-busy-reading-the-busy-ratio-before-the-queue-fills">uWSGI all workers busy: reading the busy ratio before the queue fills&lt;/h1>
&lt;p>When every uWSGI worker shows &lt;code>status: &amp;quot;busy&amp;quot;&lt;/code>, the next incoming request does not wait in a place you can see. It lands in the kernel socket backlog, which uWSGI cannot reliably measure on standard Linux. If that backlog fills, the kernel drops connections silently. No log entry, no error counter, no uWSGI-level signal. The application looks alive but stops serving real traffic.&lt;/p></description></item><item><title>uWSGI avg_rt is not a real average: why the latency number lies</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-avg-rt-not-cumulative/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-avg-rt-not-cumulative/</guid><description>&lt;h1 id="uwsgi-avg_rt-is-not-a-real-average-why-the-latency-number-lies">uWSGI avg_rt is not a real average: why the latency number lies&lt;/h1>
&lt;p>uWSGI exposes a per-worker field called &lt;code>avg_rt&lt;/code> in its stats server JSON. Most monitoring tools label it &amp;ldquo;average response time&amp;rdquo; and graph it as a latency indicator. It is not a cumulative or lifetime average.&lt;/p>
&lt;p>The field is updated with the formula &lt;code>(old_avg_rt + current_request_time) / 2&lt;/code>, an exponential moving average with a smoothing factor of 0.5. The most recent request contributes 50% of the displayed value. The request before that contributes 25%. By the seventh request back, the contribution is under 1%.&lt;/p></description></item><item><title>uWSGI behind nginx: 502 Bad Gateway and 'upstream prematurely closed connection'</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-nginx-502-bad-gateway/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-nginx-502-bad-gateway/</guid><description>&lt;h1 id="uwsgi-behind-nginx-502-bad-gateway-and-upstream-prematurely-closed-connection">uWSGI behind nginx: 502 Bad Gateway and &amp;lsquo;upstream prematurely closed connection&amp;rsquo;&lt;/h1>
&lt;p>The error string &lt;code>upstream prematurely closed connection while reading response header from upstream&lt;/code> in the nginx error log means nginx had an established connection to a uWSGI worker, the worker accepted the request, and then the connection closed before nginx received a complete response header. The worker vanished mid-response.&lt;/p>
&lt;p>This is distinct from two other errors that also produce 5xx responses but have different root causes:&lt;/p></description></item><item><title>uWSGI behind nginx: 504 Gateway Timeout when uwsgi_read_timeout fires first</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-nginx-504-gateway-timeout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-nginx-504-gateway-timeout/</guid><description>&lt;h1 id="uwsgi-behind-nginx-504-gateway-timeout-when-uwsgi_read_timeout-fires-first">uWSGI behind nginx: 504 Gateway Timeout when uwsgi_read_timeout fires first&lt;/h1>
&lt;p>A 504 from nginx with a &lt;code>uwsgi_pass&lt;/code> upstream means nginx stopped waiting for uWSGI before the worker finished. The default &lt;code>uwsgi_read_timeout&lt;/code> is 60 seconds. If uWSGI has no &lt;code>harakiri&lt;/code> configured (the default), or if harakiri is set higher than &lt;code>uwsgi_read_timeout&lt;/code>, the worker keeps processing a request whose response will never be read. The client already has their 504. The worker burns CPU, holds database connections, and occupies a slot that could serve real traffic.&lt;/p></description></item><item><title>uWSGI cache subsystem: hit ratio drops and 'full' insert failures</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-cache-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-cache-full/</guid><description>&lt;h1 id="uwsgi-cache-subsystem-hit-ratio-drops-and-full-insert-failures">uWSGI cache subsystem: hit ratio drops and &amp;lsquo;full&amp;rsquo; insert failures&lt;/h1>
&lt;p>Response times are creeping upward on cached endpoints. Application logs show nothing. Worker busy ratio is normal. The cause is likely silent: uWSGI cache misses are increasing, and each miss forces the worker to compute or fetch the response instead of serving from shared memory.&lt;/p>
&lt;p>This article covers uWSGI&amp;rsquo;s built-in cache subsystem (&lt;code>cache2&lt;/code>), not external caches. If &lt;code>caches[]&lt;/code> is absent from the stats server output, the uWSGI cache is not enabled and these diagnostics do not apply.&lt;/p></description></item><item><title>uWSGI capacity planning: the leading indicators before saturation</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-capacity-saturation-indicators/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-capacity-saturation-indicators/</guid><description>&lt;h1 id="uwsgi-capacity-planning-the-leading-indicators-before-saturation">uWSGI capacity planning: the leading indicators before saturation&lt;/h1>
&lt;p>uWSGI does not degrade gracefully. Performance looks fine until a hard limit is reached, then the service drops requests, kills workers, or refuses connections. Capacity planning is about knowing where each cliff is, measuring how close you are, and adding capacity before the edge.&lt;/p>
&lt;p>Four resources saturate as cliffs: worker pool (concurrency), memory (per-worker RSS), socket backlog (kernel listen queue), and file descriptors. Each has a distinct failure mode, leading indicator, and measurement method. Some signals come from the uWSGI stats server JSON. Others require external tools because uWSGI&amp;rsquo;s internal measurement is broken or silent.&lt;/p></description></item><item><title>uWSGI chain reload: cycling workers one at a time for zero-downtime deploys</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-chain-reload-zero-downtime/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-chain-reload-zero-downtime/</guid><description>&lt;h1 id="uwsgi-chain-reload-cycling-workers-one-at-a-time-for-zero-downtime-deploys">uWSGI chain reload: cycling workers one at a time for zero-downtime deploys&lt;/h1>
&lt;p>The default uWSGI graceful reload sends all workers the shutdown signal at once. Each worker finishes its current request, exits, and the master forks a replacement. During the gap between old workers dying and new workers becoming ready to accept connections, serving capacity falls to zero. For applications with fast startup, this gap is a brief hiccup. For applications that load large models, warm connection pools, or run heavy imports on startup, the gap can stretch into seconds or minutes of complete unavailability.&lt;/p></description></item><item><title>uWSGI cheaper subsystem: dynamic worker scaling and the false 'missing workers' alert</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-cheaper-subsystem/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-cheaper-subsystem/</guid><description>&lt;h1 id="uwsgi-cheaper-subsystem-dynamic-worker-scaling-and-the-false-missing-workers-alert">uWSGI cheaper subsystem: dynamic worker scaling and the false &amp;lsquo;missing workers&amp;rsquo; alert&lt;/h1>
&lt;p>The uWSGI cheaper subsystem dynamically scales the worker pool up and down at runtime based on demand. The master process spawns additional workers when traffic increases and reaps them when it subsides, reducing memory consumption during idle periods and providing automatic capacity during bursts.&lt;/p>
&lt;p>When the cheaper subsystem is active, the number of alive workers fluctuates by design. Cheaped workers (those scaled down by the subsystem) appear in the stats server output with &lt;code>&amp;quot;status&amp;quot;:&amp;quot;cheap&amp;quot;&lt;/code> and &lt;code>&amp;quot;pid&amp;quot;:0&lt;/code>. Monitoring that expects a fixed worker count will fire false &amp;ldquo;missing workers&amp;rdquo; alerts every time the system scales down.&lt;/p></description></item><item><title>uWSGI connection pool cascade: downstream latency that stalls every worker</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-connection-pool-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-connection-pool-cascade/</guid><description>&lt;h1 id="uwsgi-connection-pool-cascade-downstream-latency-that-stalls-every-worker">uWSGI connection pool cascade: downstream latency that stalls every worker&lt;/h1>
&lt;p>Workers climbing toward all-busy. Response times creeping up. Throughput falling. No traffic spike, no deploy, no code change. The pattern built over minutes, not seconds.&lt;/p>
&lt;p>The mechanism is a reinforcing feedback loop. A downstream dependency (database, cache, external API) gets slower. Not dead, just slower. Each request now holds its downstream connection longer. The pool fills. The next request that needs a connection blocks waiting for one to be returned. That wait adds to the request&amp;rsquo;s wall-clock duration, which keeps the worker busy longer, which keeps the connection occupied longer. The loop tightens until every worker is blocked on pool checkout and throughput collapses.&lt;/p></description></item><item><title>uWSGI connection refused: clients turned away when the backlog overflows</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-connection-refused/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-connection-refused/</guid><description>&lt;h1 id="uwsgi-connection-refused-clients-turned-away-when-the-backlog-overflows">uWSGI connection refused: clients turned away when the backlog overflows&lt;/h1>
&lt;p>Clients connecting to your uWSGI application receive connection refused or TCP RST. Behind nginx, the error log shows &lt;code>connect() failed (111: Connection refused) while connecting to upstream&lt;/code>. The application process may still be running, the master PID may exist, and the stats server may respond, yet real traffic is being turned away.&lt;/p>
&lt;p>This symptom has two root causes that look identical to the client but require opposite fixes. Either the listen backlog has overflowed because workers are saturated and cannot call &lt;code>accept()&lt;/code> fast enough, or the listener itself is dead (master gone, wrong socket path, broken socket permissions). Resolving that diagnostic fork is the first task.&lt;/p></description></item><item><title>uWSGI dropped connections: reading TcpExtListenOverflows and ListenDrops</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-tcp-listen-overflows/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-tcp-listen-overflows/</guid><description>&lt;h1 id="uwsgi-dropped-connections-reading-tcpextlistenoverflows-and-listendrops">uWSGI dropped connections: reading TcpExtListenOverflows and ListenDrops&lt;/h1>
&lt;p>Users report intermittent &amp;ldquo;connection refused&amp;rdquo; or timeouts. Nginx logs show 502s. uWSGI logs show nothing: no errors, no exceptions, no harakiri events. The master is alive, workers are accepting requests, throughput looks normal on average. But clients are being turned away.&lt;/p>
&lt;p>The explanation is in two kernel counters that uWSGI cannot report on its own: &lt;code>TcpExtListenOverflows&lt;/code> and &lt;code>TcpExtListenDrops&lt;/code>. When all uWSGI workers are busy and the kernel&amp;rsquo;s accept queue (the listen backlog) fills, the kernel drops new connections before uWSGI&amp;rsquo;s &lt;code>accept()&lt;/code> call ever runs. The application has no visibility into this event. The only evidence lives in &lt;code>/proc/net/netstat&lt;/code>.&lt;/p></description></item><item><title>uWSGI Emperor healthy but vassal dead: monitoring each instance independently</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-emperor-vassal-down/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-emperor-vassal-down/</guid><description>&lt;h1 id="uwsgi-emperor-healthy-but-vassal-dead-monitoring-each-instance-independently">uWSGI Emperor healthy but vassal dead: monitoring each instance independently&lt;/h1>
&lt;p>The Emperor process is running, its stats endpoint responds, and it continues scanning config directories. But one application is down. Clients get connection refused or timeouts. Your monitoring says the Emperor is healthy because it is. Emperor liveness is not application liveness.&lt;/p>
&lt;p>In Emperor mode, the Emperor manages one vassal (a full uWSGI instance) per config file. Each vassal is an independent process tree with its own master, workers, and stats server. The Emperor monitors config files and spawns or reloads vassals, but a vassal can die and fail to restart on a broken config while the Emperor runs unaffected. Vassal states:&lt;/p></description></item><item><title>uWSGI file descriptor limits: raising ulimit -n and systemd LimitNOFILE</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-file-descriptor-limits/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-file-descriptor-limits/</guid><description>&lt;h1 id="uwsgi-file-descriptor-limits-raising-ulimit--n-and-systemd-limitnofile">uWSGI file descriptor limits: raising ulimit -n and systemd LimitNOFILE&lt;/h1>
&lt;p>The default per-process file descriptor limit on many Linux distributions is 1024. For a uWSGI instance running multiple workers, each holding connections, sockets, log files, and database handles, that ceiling is too low for production.&lt;/p>
&lt;p>File descriptor exhaustion in uWSGI is silent. When the limit is hit, &lt;code>accept()&lt;/code> and &lt;code>open()&lt;/code> calls fail with &lt;code>EMFILE&lt;/code>. New connections are rejected with no uWSGI-level error, no log entry, and no stats counter reflecting the problem. Clients see connection resets or timeouts. The master process stays alive. Worker status looks normal. The only evidence is at the OS level.&lt;/p></description></item><item><title>uWSGI harakiri death spiral: workers killed and respawned while throughput collapses</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-harakiri-death-spiral/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-harakiri-death-spiral/</guid><description>&lt;h1 id="uwsgi-harakiri-death-spiral-workers-killed-and-respawned-while-throughput-collapses">uWSGI harakiri death spiral: workers killed and respawned while throughput collapses&lt;/h1>
&lt;p>Workers are being killed by harakiri and respawned in a tight loop. Every request blocks past the configured timeout. The master sends SIGKILL to the worker, forks a replacement, and the new worker immediately accepts the next queued request, which also blocks. Throughput collapses to near zero while the worker pool appears &amp;ldquo;busy&amp;rdquo; at or near 100%.&lt;/p>
&lt;p>This is the harakiri death spiral: a composite failure where the root cause is almost never uWSGI itself. A downstream dependency (database, external API, DNS resolver) has become unresponsive or unreachable. Every request that touches that dependency hangs. Harakiri is working as designed, killing stuck workers, but the recycling provides no relief because the next request is equally doomed.&lt;/p></description></item><item><title>uWSGI harakiri not configured: stuck workers with no timeout and no recovery</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-harakiri-not-configured/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-harakiri-not-configured/</guid><description>&lt;h1 id="uwsgi-harakiri-not-configured-stuck-workers-with-no-timeout-and-no-recovery">uWSGI harakiri not configured: stuck workers with no timeout and no recovery&lt;/h1>
&lt;p>uWSGI workers are vanishing one at a time. The master process is alive, the stats server responds, and &lt;code>harakiri_count&lt;/code> reads zero across every worker. By every metric you thought mattered, the service looks healthy. Then you notice that half the workers have been in &lt;code>busy&lt;/code> status for the last ten minutes without processing a single new request. Requests are timing out at the proxy. The listen queue is filling. The application is effectively down, and nothing in uWSGI is attempting to recover it.&lt;/p></description></item><item><title>uWSGI HARAKIRI ON WORKER: requests killed for exceeding the timeout</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-harakiri-worker-killed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-harakiri-worker-killed/</guid><description>&lt;h1 id="uwsgi-harakiri-on-worker-requests-killed-for-exceeding-the-timeout">uWSGI HARAKIRI ON WORKER: requests killed for exceeding the timeout&lt;/h1>
&lt;pre tabindex="0">&lt;code>HARAKIRI ON WORKER N (pid: XXXX, try: 1) !!!
&lt;/code>&lt;/pre>&lt;p>You see it in the uWSGI error log, followed by the master reporting the worker died by signal 9. Your monitoring shows a spike in 502 or 504 responses from the reverse proxy. Something in the request path is hanging long enough to exceed the configured &lt;code>harakiri&lt;/code> limit, and the master process is killing workers to prevent total pool exhaustion.&lt;/p></description></item><item><title>uWSGI harakiri timeout: setting it against request duration and nginx timeouts</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-harakiri-timeout-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-harakiri-timeout-tuning/</guid><description>&lt;h1 id="uwsgi-harakiri-timeout-setting-it-against-request-duration-and-nginx-timeouts">uWSGI harakiri timeout: setting it against request duration and nginx timeouts&lt;/h1>
&lt;p>The harakiri timeout is uWSGI&amp;rsquo;s per-request watchdog: if a request runs longer than the configured threshold, the master kills the worker (SIGKILL by default) and respawns it. Without it, a hung worker stays hung indefinitely, consuming a slot in the pool until every worker is stuck and the service is dark.&lt;/p>
&lt;p>Harakiri has to sit in a narrow band: above the p99 of legitimate requests so you do not kill real traffic, but below the point where the upstream proxy gives up. If nginx&amp;rsquo;s &lt;code>uwsgi_read_timeout&lt;/code> fires first, you get a 504 while the uWSGI worker keeps processing a response nobody will read. If harakiri fires first, nginx sees an upstream disconnect and returns a 502. The capacity implications differ: a 504 means wasted work, a 502 means a recycled worker.&lt;/p></description></item><item><title>uWSGI harakiri-verbose: finding the blocked syscall behind a timeout</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-harakiri-verbose-diagnosis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-harakiri-verbose-diagnosis/</guid><description>&lt;h1 id="uwsgi-harakiri-verbose-finding-the-blocked-syscall-behind-a-timeout">uWSGI harakiri-verbose: finding the blocked syscall behind a timeout&lt;/h1>
&lt;p>When a uWSGI worker exceeds the harakiri timeout, the master kills it with SIGKILL and respawns a replacement. The default harakiri log line identifies which worker died and when, but not what the worker was doing when it got stuck. The request could be CPU-bound (a pathological regex, a tight loop), blocked on I/O (a database query that never returns), or deadlocked on an internal lock. Each requires a different fix.&lt;/p></description></item><item><title>uWSGI in gevent/async mode: why worker busy ratio stops meaning anything</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-gevent-async-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-gevent-async-monitoring/</guid><description>&lt;h1 id="uwsgi-in-geventasync-mode-why-worker-busy-ratio-stops-meaning-anything">uWSGI in gevent/async mode: why worker busy ratio stops meaning anything&lt;/h1>
&lt;p>Switching uWSGI from pre-fork to gevent mode lets a single worker multiplex dozens or hundreds of concurrent requests on an event loop. The tradeoff: the worker busy ratio, the primary capacity metric in pre-fork mode, becomes unreliable. A worker reports &amp;ldquo;busy&amp;rdquo; whenever its event loop is running, which is nearly always, regardless of actual request load.&lt;/p>
&lt;p>Teams that keep alerting on busy ratio after switching to gevent either see perpetual 100% utilization alarms or, worse, silence while real problems go undetected. The cheaper_busyness algorithm has a known incompatibility with gevent that causes workers to spawn under load but never scale back down.&lt;/p></description></item><item><title>uWSGI listen backlog and net.core.somaxconn: sizing the connection queue</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-listen-backlog-somaxconn/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-listen-backlog-somaxconn/</guid><description>&lt;h1 id="uwsgi-listen-backlog-and-netcoresomaxconn-sizing-the-connection-queue">uWSGI listen backlog and net.core.somaxconn: sizing the connection queue&lt;/h1>
&lt;p>The connection backlog is the buffer between arriving TCP connections and your uWSGI workers. It is set by two independent values that must agree: uWSGI&amp;rsquo;s &lt;code>--listen&lt;/code> option and the kernel&amp;rsquo;s &lt;code>net.core.somaxconn&lt;/code>. The effective backlog is the smaller of the two. If either is too small, the queue fills during brief traffic spikes or downstream slowdowns, and the kernel starts dropping connections silently.&lt;/p></description></item><item><title>uWSGI listen queue full: the backlog overflow that drops connections silently</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-listen-queue-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-listen-queue-full/</guid><description>&lt;h1 id="uwsgi-listen-queue-full-the-backlog-overflow-that-drops-connections-silently">uWSGI listen queue full: the backlog overflow that drops connections silently&lt;/h1>
&lt;p>Clients report connection timeouts or resets. Your load balancer shows 502s or 504s on requests to the uWSGI backend. The uWSGI master is running, workers are alive, the stats server responds, and application logs show no errors. This is listen queue overflow: the kernel accept backlog on the uWSGI listening socket is full, and the kernel is silently dropping new connections before &lt;code>accept()&lt;/code>.&lt;/p></description></item><item><title>uWSGI listen_queue always zero: why the stats field is broken on Linux</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-listen-queue-stats-unreliable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-listen-queue-stats-unreliable/</guid><description>&lt;h1 id="uwsgi-listen_queue-always-zero-why-the-stats-field-is-broken-on-linux">uWSGI listen_queue always zero: why the stats field is broken on Linux&lt;/h1>
&lt;p>You open the uWSGI stats server JSON during a traffic spike. &lt;code>listen_queue&lt;/code> reads &lt;code>0&lt;/code>. &lt;code>load&lt;/code> reads &lt;code>0&lt;/code>. &lt;code>listen_queue_errors&lt;/code> reads &lt;code>0&lt;/code>. But nginx is returning 502s, clients are seeing connection refused, and all workers are busy.&lt;/p>
&lt;p>The fields do not measure what you think. On standard Linux, &lt;code>listen_queue&lt;/code> is broken. &lt;code>load&lt;/code> is identical to &lt;code>listen_queue&lt;/code> (the source code has a &lt;code>TODO&lt;/code> comment admitting this). &lt;code>listen_queue_errors&lt;/code> is dead code that is never incremented. All three read &lt;code>0&lt;/code> regardless of actual socket backlog pressure.&lt;/p></description></item><item><title>uWSGI master process dead: total outage while the PID file lingers</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-master-process-dead/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-master-process-dead/</guid><description>&lt;h1 id="uwsgi-master-process-dead-total-outage-while-the-pid-file-lingers">uWSGI master process dead: total outage while the PID file lingers&lt;/h1>
&lt;p>Nginx returns 502. The uWSGI PID file exists at its expected path, so monitoring reports the service as &amp;ldquo;up.&amp;rdquo; But nothing is serving traffic. The master process is dead, and the stale PID file is lying to you.&lt;/p>
&lt;p>The master holds the listening socket, forks workers, enforces harakiri timeouts, handles graceful reloads, and serves the stats endpoint. When it dies, workers are gone. No connections are accepted. The outage is total.&lt;/p></description></item><item><title>uWSGI max-requests: worker recycling that masks leaks and crashes</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-max-requests-recycling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-max-requests-recycling/</guid><description>&lt;h1 id="uwsgi-max-requests-worker-recycling-that-masks-leaks-and-crashes">uWSGI max-requests: worker recycling that masks leaks and crashes&lt;/h1>
&lt;p>You set &lt;code>max-requests&lt;/code> because it bounds memory leaks in uWSGI workers. Workers recycle on schedule. But max-requests does not fix leaks. It bounds them by killing the worker before the leak becomes fatal.&lt;/p>
&lt;p>Two blind spots follow. First, peak RSS per worker is &lt;code>leak_rate x max_requests&lt;/code>, not zero. At 1MB leaked per request and &lt;code>max-requests 1000&lt;/code>, each worker reaches roughly 1GB before recycling. With 8 workers, that is 8GB consumed by leak tolerance alone. Second, the respawn activity from max-requests looks identical to crash churn in the stats server. If workers are also dying from segfaults, OOM kills, or harakiri timeouts, the respawn counter alone cannot tell you why.&lt;/p></description></item><item><title>uWSGI Monitoring</title><link>https://www.netdata.cloud/monitoring-101/uwsgi-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/uwsgi-monitoring/</guid><description>&lt;h2 id="uwsgi-monitoring">uWSGI Monitoring&lt;/h2>
&lt;h3 id="what-is-uwsgi">What Is uWSGI?&lt;/h3>
&lt;p>uWSGI is an application server used to manage and serve web applications, primarily written in the Python programming language. It is widely used in production environments to facilitate reliable communication between web servers and application code by acting as a middle layer.&lt;/p>
&lt;h3 id="monitoring-uwsgi-with-netdata">Monitoring uWSGI With Netdata&lt;/h3>
&lt;p>Monitoring uWSGI is crucial for maintaining optimal application performance. Netdata offers a comprehensive uWSGI monitoring tool that provides real-time insights into server and application health. By utilizing Netdata, you can monitor key metrics such as requests, transmitted data, and exceptions to ensure your web applications run smoothly.&lt;/p></description></item><item><title>uWSGI monitoring checklist: the signals every production app server needs</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-monitoring-checklist/</guid><description>&lt;h1 id="uwsgi-monitoring-checklist-the-signals-every-production-app-server-needs">uWSGI monitoring checklist: the signals every production app server needs&lt;/h1>
&lt;p>The stats server (enabled with &lt;code>--stats &amp;lt;socket&amp;gt;&lt;/code>) exports a JSON document with worker state, request counters, memory usage, and error rates. This checklist organizes those signals by maturity level, from survival to expert. Each level assumes the previous one is in place.&lt;/p>
&lt;p>uWSGI is in maintenance mode (bugfixes only, no new features), so the stats schema and signal semantics are stable. Three fields are not: &lt;code>listen_queue&lt;/code>, &lt;code>load&lt;/code>, and &lt;code>listen_queue_errors&lt;/code> are broken or dead code on standard Linux. This checklist flags every unreliable field and points to the external measurement that works instead.&lt;/p></description></item><item><title>uWSGI monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-monitoring-maturity-model/</guid><description>&lt;h1 id="uwsgi-monitoring-maturity-model-from-survival-to-expert">uWSGI monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>A four-level progression for uWSGI monitoring, from bare liveness checks to deep signal correlation. The levels are cumulative: you cannot skip to Level 3 by tracking RSS growth trends while ignoring harakiri rate. Each tier closes a specific class of blind spot that the previous tier could not see.&lt;/p>
&lt;p>The model assumes the uWSGI stats server is enabled with &lt;code>--stats &amp;lt;address&amp;gt;&lt;/code>. Without it, every level above survival is unreachable. HTTP access to the stats server requires the additional &lt;code>--stats-http&lt;/code> flag; otherwise use &lt;code>uwsgi --connect-and-read &amp;lt;addr&amp;gt;&lt;/code> for TCP sockets or &lt;code>socat - UNIX-CONNECT:&amp;lt;path&amp;gt;&lt;/code> for UNIX sockets.&lt;/p></description></item><item><title>uWSGI processes vs threads: sizing concurrency without wasting memory</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-processes-vs-threads/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-processes-vs-threads/</guid><description>&lt;h1 id="uwsgi-processes-vs-threads-sizing-concurrency-without-wasting-memory">uWSGI processes vs threads: sizing concurrency without wasting memory&lt;/h1>
&lt;p>The concurrency model you choose in uWSGI sets how many requests your server handles simultaneously, what that costs in memory, and what your monitoring signals mean once traffic arrives. Get it wrong and you either waste memory on idle processes or starve the kernel listen queue because you misread what &amp;ldquo;all workers busy&amp;rdquo; indicates.&lt;/p>
&lt;p>uWSGI supports three concurrency models. Pre-fork gives you one request per worker process. Threaded gives you N requests per worker via threads. Async (gevent or asyncio) gives you many requests per worker via an event loop. Each has a different memory profile, a different CPU characteristic, and a different definition of &amp;ldquo;busy.&amp;rdquo;&lt;/p></description></item><item><title>uWSGI reload blackout: a broken deploy leaves zero workers running</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-reload-blackout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-reload-blackout/</guid><description>&lt;h1 id="uwsgi-reload-blackout-a-broken-deploy-leaves-zero-workers-running">uWSGI reload blackout: a broken deploy leaves zero workers running&lt;/h1>
&lt;p>You deployed new code to uWSGI. The graceful reload killed the old workers, but throughput is zero. The master process is alive. Every worker it forks dies before accepting a connection.&lt;/p>
&lt;p>The reload mechanism worked. The new code cannot initialize. Workers crash during startup, the master respawns them, and they crash again. No worker stays alive long enough to serve a request.&lt;/p></description></item><item><title>uWSGI reload thundering herd: capacity drops to zero during a slow restart</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-graceful-reload-thundering-herd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-graceful-reload-thundering-herd/</guid><description>&lt;h1 id="uwsgi-reload-thundering-herd-capacity-drops-to-zero-during-a-slow-restart">uWSGI reload thundering herd: capacity drops to zero during a slow restart&lt;/h1>
&lt;p>You deploy a new version of your application. Seconds later, request throughput collapses to zero. Nginx returns 502s. The uWSGI master process is alive and the stats server responds, but no workers are accepting connections. Thirty seconds to several minutes later, throughput recovers. If this pattern aligns exactly with every deployment, you are hitting the reload thundering herd.&lt;/p></description></item><item><title>uWSGI reload-on-rss vs evil-reload-on-rss: recycling workers on memory limits</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-reload-on-rss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-reload-on-rss/</guid><description>&lt;h1 id="uwsgi-reload-on-rss-vs-evil-reload-on-rss-recycling-workers-on-memory-limits">uWSGI reload-on-rss vs evil-reload-on-rss: recycling workers on memory limits&lt;/h1>
&lt;p>Two uWSGI options recycle workers when their resident set size crosses a threshold: &lt;code>reload-on-rss&lt;/code> and &lt;code>evil-reload-on-rss&lt;/code>. Both produce the same effect on a monitoring dashboard. The &lt;code>respawn_count&lt;/code> counter increments, RSS drops to baseline, a new worker is forked. The memory sawtooth looks healthy. From the master process perspective, the two are indistinguishable.&lt;/p>
&lt;p>The difference is in what happens to the in-flight request when the kill occurs.&lt;/p></description></item><item><title>uWSGI respawn rate high: telling crashes apart from max-requests recycling</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-respawn-rate-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-respawn-rate-high/</guid><description>&lt;h1 id="uwsgi-respawn-rate-high-telling-crashes-apart-from-max-requests-recycling">uWSGI respawn rate high: telling crashes apart from max-requests recycling&lt;/h1>
&lt;p>You see &lt;code>respawn_count&lt;/code> climbing across your uWSGI workers. The question is whether workers are crashing or whether uWSGI is doing what you configured it to do.&lt;/p>
&lt;p>&lt;code>respawn_count&lt;/code> is a single monotonic counter per worker slot that increments for every reason a worker can die and come back: &lt;code>max-requests&lt;/code> recycling, &lt;code>reload-on-rss&lt;/code> memory recycling, harakiri timeout kills, segfaults, OOM kills, and manual &lt;code>kill -9&lt;/code>. The raw number tells you workers are churning. It does not tell you why.&lt;/p></description></item><item><title>uWSGI response time climbing: rising avg_rt and where it comes from</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-response-time-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-response-time-high/</guid><description>&lt;h1 id="uwsgi-response-time-climbing-rising-avg_rt-and-where-it-comes-from">uWSGI response time climbing: rising avg_rt and where it comes from&lt;/h1>
&lt;p>When avg_rt climbs in the uWSGI stats server, the instinct is to look for slow application code. That is often right, but avg_rt is more nuanced than a simple average response time. Understanding what it actually measures is the difference between a fast diagnosis and a misleading rabbit hole.&lt;/p>
&lt;p>avg_rt is an exponential moving average (EMA), not a cumulative average. When it climbs, something per-request is getting slower. The question is whether the cause is a downstream dependency (database pool exhaustion, slow external API), resource contention (CPU saturation, GC storms, GIL pressure in threaded Python), or cascading starvation (workers backing up because each request holds the worker longer). Correlating avg_rt with the worker busy ratio narrows this quickly.&lt;/p></description></item><item><title>uWSGI RSS vs VSZ: reading worker memory without being fooled by shared pages</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-rss-vs-vsz/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-rss-vs-vsz/</guid><description>&lt;h1 id="uwsgi-rss-vs-vsz-reading-worker-memory-without-being-fooled-by-shared-pages">uWSGI RSS vs VSZ: reading worker memory without being fooled by shared pages&lt;/h1>
&lt;p>You look at your uWSGI stats and see 8 workers each reporting 500 MB of RSS. The host has 4 GB of RAM. By the math, you should be deep into swap, but &lt;code>vmstat&lt;/code> shows zero swap activity and response times are fine.&lt;/p>
&lt;p>The gap is copy-on-write (COW) sharing. uWSGI&amp;rsquo;s default model loads the application once in the master process, then forks workers that inherit the master&amp;rsquo;s memory. Linux marks those inherited pages as shared and read-only. As long as workers do not write to them, the pages stay shared across all workers. But Linux RSS counts every shared page fully for every process that maps it. Sum worker RSS across 8 workers and you may overstate real memory consumption by 2-5x.&lt;/p></description></item><item><title>uWSGI spooler backlog: deferred tasks piling up on disk</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-spooler-backlog/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-spooler-backlog/</guid><description>&lt;h1 id="uwsgi-spooler-backlog-deferred-tasks-piling-up-on-disk">uWSGI spooler backlog: deferred tasks piling up on disk&lt;/h1>
&lt;p>The uWSGI spooler is an on-disk deferred-job queue. Application code serializes work units into files in a spool directory, and dedicated spooler processes consume them asynchronously. When producers enqueue work faster than the spooler drains it, or when the spooler crashes or stalls, files accumulate on disk. The backlog grows silently, deferred work latency increases, and in the worst case the disk fills or tasks are silently dropped.&lt;/p></description></item><item><title>uWSGI stats server unreachable: socket permissions and a stuck master</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-stats-server-unreachable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-stats-server-unreachable/</guid><description>&lt;h1 id="uwsgi-stats-server-unreachable-socket-permissions-and-a-stuck-master">uWSGI stats server unreachable: socket permissions and a stuck master&lt;/h1>
&lt;p>Your monitoring collector reports the uWSGI stats server as unreachable. Before debugging the stats server, answer one question: is uWSGI actually serving traffic?&lt;/p>
&lt;p>The uWSGI stats server runs inside the master process and serves raw JSON over a socket (UNIX or TCP) when enabled with &lt;code>--stats&lt;/code>. The master does not serve application requests. This separation matters: the stats server can be unreachable while every worker is processing traffic, and it can be reachable while every worker is stuck. Stats reachability and application availability are independent signals.&lt;/p></description></item><item><title>uWSGI stats server: enabling it and reading it without --stats-http</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-stats-server-setup/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-stats-server-setup/</guid><description>&lt;h1 id="uwsgi-stats-server-enabling-it-and-reading-it-without---stats-http">uWSGI stats server: enabling it and reading it without &amp;ndash;stats-http&lt;/h1>
&lt;p>The uWSGI stats server is the primary data source for every monitoring signal in the uWSGI playbook: worker busy ratio, harakiri count, avg_rt, respawn rate, RSS, exception counts, and more. Without it, you are blind to internal state. With it, you have complete visibility into worker pool health, saturation, and failure modes.&lt;/p>
&lt;p>The stats server must be explicitly enabled with the &lt;code>--stats&lt;/code> option. It is not on by default. Once enabled, it serves a JSON document containing the full internal state of the master process and all workers, including PIDs, worker status, request counters, memory usage, and per-core request data.&lt;/p></description></item><item><title>uWSGI stats socket exposure: internal state leaking on a public port</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-stats-socket-exposure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-stats-socket-exposure/</guid><description>&lt;h1 id="uwsgi-stats-socket-exposure-internal-state-leaking-on-a-public-port">uWSGI stats socket exposure: internal state leaking on a public port&lt;/h1>
&lt;p>The uWSGI stats server, enabled with &lt;code>--stats&lt;/code>, serves a JSON blob containing the complete internal state of the master and all workers: PIDs, memory usage, request counts, response times, configuration details, and in-flight request data. When this socket is reachable from an untrusted network, every field in that JSON is an information-disclosure vector.&lt;/p>
&lt;p>The stats server has no authentication. Access control is entirely network-level. Any client that can connect to the socket address receives the full dump with no challenge, no token, no ACL. The common pattern &lt;code>--stats :9191&lt;/code> (address with no IP specified) binds on every interface, which means the stats endpoint is open to anyone who can reach the host on that port.&lt;/p></description></item><item><title>uWSGI threaded mode and the GIL: why more threads don't add CPU parallelism</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-gil-cpu-bound/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-gil-cpu-bound/</guid><description>&lt;h1 id="uwsgi-threaded-mode-and-the-gil-why-more-threads-dont-add-cpu-parallelism">uWSGI threaded mode and the GIL: why more threads don&amp;rsquo;t add CPU parallelism&lt;/h1>
&lt;p>You doubled the thread count on each uWSGI worker, expecting throughput to scale with your multi-core box. CPU utilization barely moved. Response time did not improve. Memory looks fine. The queue still fills under load.&lt;/p>
&lt;p>The cause is almost certainly the CPython Global Interpreter Lock. In threaded mode, a uWSGI worker runs N OS threads, but all N threads share a single GIL within that process. For CPU-bound Python code, only one thread executes bytecode at a time regardless of how many threads you configured. Adding threads helps with I/O-bound workloads, where threads can yield the GIL while waiting on network or disk, but it does nothing for computation-heavy request handlers.&lt;/p></description></item><item><title>uWSGI throughput drop: requests per second falling with traffic steady</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-throughput-drop/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-throughput-drop/</guid><description>&lt;h1 id="uwsgi-throughput-drop-requests-per-second-falling-with-traffic-steady">uWSGI throughput drop: requests per second falling with traffic steady&lt;/h1>
&lt;p>Throughput in uWSGI is a derived metric: sum per-worker &lt;code>requests&lt;/code> counters across all workers, then compute the delta between polling intervals. Anything that disrupts those counters or changes how fast workers complete requests shows up as a throughput trend.&lt;/p>
&lt;p>Two caveats before you start:&lt;/p>
&lt;ol>
&lt;li>
&lt;p>&lt;strong>Throughput can rise during failure.&lt;/strong> If the app starts returning 500s immediately without processing, requests-per-second goes up while useful throughput goes down. Always pair throughput with exception rate.&lt;/p></description></item><item><title>uWSGI thundering herd: accept() contention and the thunder-lock fix</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-thundering-herd-thunder-lock/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-thundering-herd-thunder-lock/</guid><description>&lt;h1 id="uwsgi-thundering-herd-accept-contention-and-the-thunder-lock-fix">uWSGI thundering herd: accept() contention and the thunder-lock fix&lt;/h1>
&lt;p>You have a uWSGI deployment with multiple worker processes, and something does not add up. CPU usage is elevated. Workers toggle between idle and busy. But request throughput is low, response times are higher than expected, and adding more workers makes things worse instead of better. The system looks under load but is not actually doing much work.&lt;/p>
&lt;p>This is the uWSGI thundering herd problem. When a new connection arrives on the shared listening socket, every idle worker process wakes up and races to call &lt;code>accept()&lt;/code>. Only one wins. The rest burn CPU and kernel time on a context switch for nothing, then go back to sleep. Under low-to-moderate traffic with many workers, this contention consumes more CPU than the actual request processing.&lt;/p></description></item><item><title>uWSGI too many open files: file descriptor exhaustion and EMFILE</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-too-many-open-files/</guid><description>&lt;h1 id="uwsgi-too-many-open-files-file-descriptor-exhaustion-and-emfile">uWSGI too many open files: file descriptor exhaustion and EMFILE&lt;/h1>
&lt;p>When the uWSGI process hits its file descriptor limit (&lt;code>RLIMIT_NOFILE&lt;/code>), &lt;code>accept()&lt;/code> and &lt;code>open()&lt;/code> start returning &lt;code>EMFILE&lt;/code> (errno 24). New connections are silently rejected. Logging fails. Application code throws exceptions that look like disk or network failures. The uWSGI stats endpoint does not report file descriptor usage, so there is no uWSGI-level signal pointing to the real problem. Operators typically chase disk space, I/O, or network connectivity before realizing the process has simply run out of file descriptors.&lt;/p></description></item><item><title>uWSGI worker exceptions climbing: unhandled errors reaching the WSGI layer</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-worker-exceptions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-worker-exceptions/</guid><description>&lt;h1 id="uwsgi-worker-exceptions-climbing-unhandled-errors-reaching-the-wsgi-layer">uWSGI worker exceptions climbing: unhandled errors reaching the WSGI layer&lt;/h1>
&lt;p>The &lt;code>workers[].exceptions&lt;/code> counter is climbing, which means unhandled exceptions are propagating past your application code and reaching the uWSGI WSGI layer. Each increment typically corresponds to a 500 response delivered to the client. The counter is per-worker and monotonically increasing, so track the rate of change (delta over your polling interval), not the absolute value.&lt;/p>
&lt;p>Critical nuance: this counter systematically undercounts application errors. If your framework (Django, Flask, FastAPI, and most others) catches exceptions via middleware and returns a 500 response itself, uWSGI never sees the exception. The request completes &amp;ldquo;successfully&amp;rdquo; from uWSGI&amp;rsquo;s perspective, and the counter does not increment. When the counter does climb, exceptions are escaping the framework entirely &amp;ndash; a more severe condition than a framework-handled 500.&lt;/p></description></item><item><title>uWSGI worker killed by the OOM killer: mysterious respawns under memory pressure</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-worker-oom-killed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-worker-oom-killed/</guid><description>&lt;h1 id="uwsgi-worker-killed-by-the-oom-killer-mysterious-respawns-under-memory-pressure">uWSGI worker killed by the OOM killer: mysterious respawns under memory pressure&lt;/h1>
&lt;p>You see worker respawns in the uWSGI log that do not match your &lt;code>max-requests&lt;/code> cadence. The master reports workers dying from signal 9 (&lt;code>SIGKILL&lt;/code>), and either harakiri is not configured or the harakiri count is zero. Workers come back, serve traffic for a while, then die again. The interval shrinks over time.&lt;/p>
&lt;p>This is the signature of the Linux OOM killer targeting uWSGI workers. As total worker RSS grows beyond available RAM, the kernel swaps, performance degrades, and the OOM killer selects the largest process on the system. In a uWSGI deployment, that process is almost always a worker. The master respawns it, the new worker re-imports the application, RSS climbs back, and the cycle repeats.&lt;/p></description></item><item><title>uWSGI worker memory leak: RSS climbing across all workers</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-worker-memory-leak/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-worker-memory-leak/</guid><description>&lt;h1 id="uwsgi-worker-memory-leak-rss-climbing-across-all-workers">uWSGI worker memory leak: RSS climbing across all workers&lt;/h1>
&lt;p>Per-worker RSS climbs steadily. All workers track each other in lockstep. After a restart, RSS looks stable for hours or days, then the pattern repeats. With &lt;code>max-requests&lt;/code> or &lt;code>reload-on-rss&lt;/code> configured, you see a sawtooth: RSS rises, a worker recycles, RSS drops, then rises again. Without those directives, workers eventually hit the OOM killer or start swapping.&lt;/p>
&lt;p>The cause may be a genuine leak in application code, Python allocator fragmentation that never returns pages to the OS, a C extension bug, or Docker file-descriptor inflation. The first two are operationally identical: RSS grows monotonically and does not come back down without a process recycle.&lt;/p></description></item><item><title>uWSGI worker pool starvation: the silent outage where every worker is busy</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-worker-pool-starvation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-worker-pool-starvation/</guid><description>&lt;h1 id="uwsgi-worker-pool-starvation-the-silent-outage-where-every-worker-is-busy">uWSGI worker pool starvation: the silent outage where every worker is busy&lt;/h1>
&lt;p>Every worker shows &lt;code>status: &amp;quot;busy&amp;quot;&lt;/code>. The master process is alive. The stats server responds instantly. Your load balancer health check returns 200. But real users are seeing timeouts, connection refused errors, or hanging pages. This is worker pool starvation.&lt;/p>
&lt;p>The mechanism is a concurrency cliff. uWSGI&amp;rsquo;s pre-fork model assigns one request per worker at a time in the default configuration. When every worker is occupied with a slow or hanging request, new connections pile into the kernel listen queue (socket backlog). Once that queue fills, the kernel silently drops connections with no uWSGI log entry, no error counter, and no exception. The service is dead for real traffic while every surface-level health signal stays green.&lt;/p></description></item><item><title>uWSGI worker respawn loop: 'DAMN ! worker died :( trying respawn'</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-worker-respawn-loop/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-worker-respawn-loop/</guid><description>&lt;h1 id="uwsgi-worker-respawn-loop-damn--worker-died--trying-respawn">uWSGI worker respawn loop: &amp;lsquo;DAMN ! worker died :( trying respawn&amp;rsquo;&lt;/h1>
&lt;p>The uWSGI master logs &lt;code>DAMN ! worker N (pid: XXX) died, killed by signal S :( trying respawn ...&lt;/code> when it receives SIGCHLD for a worker that exited unexpectedly. The next line is typically &lt;code>Respawned uWSGI worker N (new pid: XXXX)&lt;/code>. In a respawn loop, these pairs repeat rapidly: the master forks a replacement, the replacement dies, and the cycle continues.&lt;/p></description></item><item><title>uWSGI worker segfault: SIGSEGV in a C extension and the respawn that follows</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-worker-segfault/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-worker-segfault/</guid><description>&lt;h1 id="uwsgi-worker-segfault-sigsegv-in-a-c-extension-and-the-respawn-that-follows">uWSGI worker segfault: SIGSEGV in a C extension and the respawn that follows&lt;/h1>
&lt;p>A worker dies with signal 11. The master respawns it. The uWSGI log records &lt;code>DAMN ! worker N (pid: XXXX) died, killed by signal 11 :( trying respawn ...&lt;/code> followed by &lt;code>Respawned uWSGI worker N (new pid: YYYY)&lt;/code>. From the outside, the service looks like it survived a momentary blip. It did not. A SIGSEGV means something in C-level code crashed: a compiled extension (numpy, lxml, a database driver, OpenSSL, protobuf), a uWSGI internal bug, or memory corruption. Python application code cannot normally produce SIGSEGV.&lt;/p></description></item><item><title>uWSGI worker stuck in busy: a hung request that never returns</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-stuck-worker/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-stuck-worker/</guid><description>&lt;h1 id="uwsgi-worker-stuck-in-busy-a-hung-request-that-never-returns">uWSGI worker stuck in busy: a hung request that never returns&lt;/h1>
&lt;p>One worker sits in &lt;code>busy&lt;/code> far longer than any legitimate request should take. The others cycle through &lt;code>idle&lt;/code> and &lt;code>busy&lt;/code> normally. The stuck worker is not crashing, not erroring, and not completing. It is holding a request open indefinitely and will stay that way until something kills it.&lt;/p>
&lt;p>This is the single-worker poisoning pattern. A specific request triggered a pathological code path: blocking I/O without a timeout, a regex catastrophe, an unbounded database query, or a deadlock in a C extension. The worker called &lt;code>accept()&lt;/code>, began processing, and never returned. Its request counter is frozen. Its &lt;code>running_time&lt;/code> stopped advancing. One slot of your worker pool is permanently consumed.&lt;/p></description></item><item><title>uWSGI write and read errors: broken pipes and clients that disconnect</title><link>https://www.netdata.cloud/guides/uwsgi/uwsgi-write-read-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/uwsgi/uwsgi-write-read-errors/</guid><description>&lt;h1 id="uwsgi-write-and-read-errors-broken-pipes-and-clients-that-disconnect">uWSGI write and read errors: broken pipes and clients that disconnect&lt;/h1>
&lt;p>write_errors and read_errors are per-core counters in the uWSGI stats server. write_errors increments when a socket write fails during response delivery, almost always because the client disconnected before the response finished (broken pipe). read_errors increments when the connection is lost during request body reading. Some of both are normal: users navigate away, mobile clients switch networks, browsers cancel pending requests. A sustained rise in write_errors, especially when correlated with rising response time or worker respawns, points to a real problem.&lt;/p></description></item><item><title>V Solution SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/v-solution-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/v-solution-snmp-traps/</guid><description/></item><item><title>Varnish</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/varnish/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/varnish/</guid><description/></item><item><title>Varnish 400 Bad Request spike: malformed requests, scanning, and smuggling</title><link>https://www.netdata.cloud/guides/varnish/varnish-400-bad-request-spike/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-400-bad-request-spike/</guid><description>&lt;h1 id="varnish-400-bad-request-spike-malformed-requests-scanning-and-smuggling">Varnish 400 Bad Request spike: malformed requests, scanning, and smuggling&lt;/h1>
&lt;p>A spike in &lt;code>MAIN.client_req_400&lt;/code> means Varnish is rejecting an unusual volume of client requests at the HTTP parsing layer, before VCL runs. A low steady rate of 400s is normal. Bots, scanners, and sloppy clients are constant background radiation on any public-facing proxy. What matters is the delta from your baseline.&lt;/p>
&lt;p>The threshold for concern is a sustained rate spike greater than 5x your established baseline. At that volume, the causes narrow to: active scanning or fuzzing, a broken client SDK pushing malformed requests, request-smuggling attempts probing parser discrepancies, or a configuration change that made Varnish reject traffic it previously tolerated.&lt;/p></description></item><item><title>Varnish backend connection reuse low: keepalive not working and slow TTFB</title><link>https://www.netdata.cloud/guides/varnish/varnish-backend-connection-reuse-low/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-backend-connection-reuse-low/</guid><description>&lt;h1 id="varnish-backend-connection-reuse-low-keepalive-not-working-and-slow-ttfb">Varnish backend connection reuse low: keepalive not working and slow TTFB&lt;/h1>
&lt;p>When Varnish reuses backend TCP connections through HTTP keepalive, cache misses skip connection setup: no TCP handshake, no optional TLS negotiation, no kernel connection-tracking overhead. When reuse collapses, every backend fetch pays that cost, and it shows up directly in time-to-first-byte.&lt;/p>
&lt;p>The reuse ratio is &lt;code>backend_reuse / (backend_reuse + backend_conn)&lt;/code> from &lt;code>varnishstat&lt;/code>. A ratio below 50% when your backend supports keepalive means the majority of fetches are opening fresh TCP connections. This adds latency on every cache miss, increases backend CPU and connection-tracking load, and accelerates file descriptor consumption on the Varnish child process.&lt;/p></description></item><item><title>Varnish backend is sick: health probes, all-backends-sick, and grace</title><link>https://www.netdata.cloud/guides/varnish/varnish-backend-sick/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-backend-sick/</guid><description>&lt;h1 id="varnish-backend-is-sick-health-probes-all-backends-sick-and-grace">Varnish backend is sick: health probes, all-backends-sick, and grace&lt;/h1>
&lt;p>A sick Varnish backend receives zero traffic. When all backends in a director go sick, Varnish either serves stale content via grace or returns 503 on every cache miss. The health transition is governed by a threshold/window probe model, not a simple pass/fail. That distinction matters when you are reading &lt;code>backend.list&lt;/code> output during an incident.&lt;/p>
&lt;h2 id="what-this-means">What this means&lt;/h2>
&lt;p>Varnish marks a backend healthy or sick based on a sliding window of probe results. A probe is an HTTP request Varnish sends to the backend at a configurable interval. The backend is healthy if at least &lt;code>threshold&lt;/code> out of the last &lt;code>window&lt;/code> probes returned the expected response (default 200) within the timeout. The defaults are stable across Varnish 6.x through 9.x:&lt;/p></description></item><item><title>Varnish backend probe configuration: threshold, window, interval, and initial</title><link>https://www.netdata.cloud/guides/varnish/varnish-backend-probe-configuration/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-backend-probe-configuration/</guid><description>&lt;h1 id="varnish-backend-probe-configuration-threshold-window-interval-and-initial">Varnish backend probe configuration: threshold, window, interval, and initial&lt;/h1>
&lt;p>A Varnish backend probe is a periodic HTTP request that determines whether a backend receives traffic. Its parameters control detection latency, flapping behavior, and false health states.&lt;/p>
&lt;p>Tighter windows detect failures faster but flap on network hiccups. Longer intervals reduce probe load but extend the detection blind spot. The default &lt;code>.initial&lt;/code> prevents false-sick states at startup, but a freshly loaded VCL marks backends unhealthy until the first real probe succeeds.&lt;/p></description></item><item><title>Varnish backend TTFB high: the leading indicator of thread pool death</title><link>https://www.netdata.cloud/guides/varnish/varnish-backend-ttfb-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-backend-ttfb-high/</guid><description>&lt;h1 id="varnish-backend-ttfb-high-the-leading-indicator-of-thread-pool-death">Varnish backend TTFB high: the leading indicator of thread pool death&lt;/h1>
&lt;p>Backend time-to-first-byte (TTFB) is the time from Varnish sending a backend request to receiving the first response header byte. Varnish uses a thread-per-request model with a bounded pool. Every backend fetch holds a worker thread for the entire fetch duration: TCP connect, wait for first byte, transfer body, process headers. When TTFB rises, each fetch holds a thread longer, and the pool fills faster.&lt;/p></description></item><item><title>Varnish backend_fail, backend_unhealthy, and backend_busy: three different backend problems</title><link>https://www.netdata.cloud/guides/varnish/varnish-backend-conn-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-backend-conn-failures/</guid><description>&lt;h1 id="varnish-backend_fail-backend_unhealthy-and-backend_busy-three-different-backend-problems">Varnish backend_fail, backend_unhealthy, and backend_busy: three different backend problems&lt;/h1>
&lt;p>Three &lt;code>varnishstat&lt;/code> counters track backend problems, and operators routinely conflate them. &lt;code>backend_fail&lt;/code>, &lt;code>backend_unhealthy&lt;/code>, and &lt;code>backend_busy&lt;/code> each fire at a different point in the backend connection decision flow, have different root causes, and need different fixes.&lt;/p>
&lt;p>The most dangerous confusion is also the subtlest: a backend can be completely offline with &lt;code>backend_fail&lt;/code> sitting at zero. If the backend is probe-sick, Varnish never attempts a connection, so &lt;code>backend_fail&lt;/code> never increments. Only &lt;code>backend_unhealthy&lt;/code> reveals this state. An operator who monitors &lt;code>backend_fail&lt;/code> alone sees a healthy system while the backend is dark.&lt;/p></description></item><item><title>Varnish ban list growing: O(n) lookups and the lurker falling behind</title><link>https://www.netdata.cloud/guides/varnish/varnish-ban-list-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-ban-list-growing/</guid><description>&lt;h1 id="varnish-ban-list-growing-on-lookups-and-the-lurker-falling-behind">Varnish ban list growing: O(n) lookups and the lurker falling behind&lt;/h1>
&lt;p>Varnish&amp;rsquo;s ban list is a linear chain of invalidation rules. Every cache lookup tests the requested object against all active bans before serving it. When the list grows, each lookup pays O(n) cost. Hit latency rises. CPU climbs. But hit ratio stays flat, backend request rate stays flat, and error counts stay flat. This makes a ban list explosion hard to detect with standard monitoring, because most teams alert on hit rate and error rate, not hit latency.&lt;/p></description></item><item><title>Varnish ban lurker not keeping up: contention and ban_lurker_sleep</title><link>https://www.netdata.cloud/guides/varnish/varnish-ban-lurker-not-keeping-up/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-ban-lurker-not-keeping-up/</guid><description>&lt;h1 id="varnish-ban-lurker-not-keeping-up-contention-and-ban_lurker_sleep">Varnish ban lurker not keeping up: contention and ban_lurker_sleep&lt;/h1>
&lt;p>The ban lurker is Varnish&amp;rsquo;s background invalidation worker. When it falls behind, the ban list grows and every cache lookup must test the object against all uncompleted bans. The signature pattern is &lt;code>MAIN.bans&lt;/code> climbing steadily while &lt;code>MAIN.bans_lurker_contention&lt;/code> rises and &lt;code>MAIN.bans_lurker_tested&lt;/code> stays low or stalls entirely.&lt;/p>
&lt;p>This is a slow-motion degradation. Hit rate often looks fine, but per-request latency on cache hits creeps upward because each lookup scans the full ban list. A ban list in the tens of thousands turns each cache hit into an O(n) scan.&lt;/p></description></item><item><title>Varnish cache hit ratio dropped: hit rate collapse and backend overload</title><link>https://www.netdata.cloud/guides/varnish/varnish-cache-hit-ratio-dropped/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-cache-hit-ratio-dropped/</guid><description>&lt;h1 id="varnish-cache-hit-ratio-dropped-hit-rate-collapse-and-backend-overload">Varnish cache hit ratio dropped: hit rate collapse and backend overload&lt;/h1>
&lt;p>When cache hit ratio drops, every missed request that previously served from memory now hits the backend. If your backends are sized for cached traffic, not the raw request rate, they saturate quickly. Backend TTFB rises, worker threads are held longer, the thread pool fills, and sessions drop.&lt;/p>
&lt;p>The critical diagnostic signal is temporal ordering. Hit rate collapse always precedes backend degradation when the cache is the root cause. If backend TTFB rises first and hit rate falls second, the backend is the problem. If hit rate drops first and backend metrics follow, the cache stopped being effective and the backend is collateral damage. This distinction determines whether you fix VCL, storage, and invalidation logic, or investigate the origin.&lt;/p></description></item><item><title>Varnish cache stampede: a popular object expires and the herd hits the backend</title><link>https://www.netdata.cloud/guides/varnish/varnish-cache-stampede-thundering-herd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-cache-stampede-thundering-herd/</guid><description>&lt;p>A popular cached object hits its TTL and expires. In the next second, hundreds of concurrent requests for that object all miss simultaneously. Varnish forwards all of them to the backend, which was sized for the small fraction of traffic that normally leaks through, not a synchronized burst of identical requests. The backend slows. Worker threads pile up waiting for responses. If the backend cannot recover quickly, thread exhaustion follows and sessions start dropping.&lt;/p></description></item><item><title>Varnish cache_hitpass / cache_hitmiss climbing: uncacheable content bleeding to the backend</title><link>https://www.netdata.cloud/guides/varnish/varnish-hitpass-hitmiss-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-hitpass-hitmiss-high/</guid><description>&lt;h1 id="varnish-cache_hitpass--cache_hitmiss-climbing-uncacheable-content-bleeding-to-the-backend">Varnish cache_hitpass / cache_hitmiss climbing: uncacheable content bleeding to the backend&lt;/h1>
&lt;p>&lt;code>cache_hitmiss&lt;/code> and &lt;code>cache_hitpass&lt;/code> are climbing. Backend request rate is growing. Cache hit rate looks stable. This is Varnish&amp;rsquo;s learned &amp;ldquo;do not cache&amp;rdquo; mechanism silently routing traffic past the cache.&lt;/p>
&lt;p>When Varnish fetches an object and determines it cannot be cached, it caches that decision itself. For the next 120 seconds (the builtin.vcl default), every request for that URL bypasses the cache and goes straight to the backend. Unlike normal cache misses, hit-for-miss requests are not coalesced: concurrent requests for the same URL each generate an independent backend fetch. The result is silent, compounding backend load that your hit-rate alert probably cannot see.&lt;/p></description></item><item><title>Varnish child panic: Child died signal, core dumps, and the crash loop</title><link>https://www.netdata.cloud/guides/varnish/varnish-child-panic/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-child-panic/</guid><description>&lt;h1 id="varnish-child-panic-child-died-signal-core-dumps-and-the-crash-loop">Varnish child panic: Child died signal, core dumps, and the crash loop&lt;/h1>
&lt;p>&lt;code>Child (NNNN) died signal=N&lt;/code> means the Varnish child (worker) process crashed. The management process supervises and restarts it automatically, so a single crash is self-recovering. But each restart empties the cache, and if the child keeps crashing, caching drops to zero and backends absorb full uncached traffic.&lt;/p>
&lt;p>The immediate question: one-off crash or crash loop? If the child recovers and the cache warms back up, it is a TICKET. If &lt;code>MAIN.uptime&lt;/code> never stabilizes and stays far below &lt;code>MGT.uptime&lt;/code>, the child is dying before the cache can warm. That is a PAGE.&lt;/p></description></item><item><title>Varnish Error 503 Backend fetch failed: what the error page actually means</title><link>https://www.netdata.cloud/guides/varnish/varnish-503-backend-fetch-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-503-backend-fetch-failed/</guid><description>&lt;h1 id="varnish-error-503-backend-fetch-failed-what-the-error-page-actually-means">Varnish Error 503 Backend fetch failed: what the error page actually means&lt;/h1>
&lt;p>The &amp;ldquo;Error 503 Backend fetch failed&amp;rdquo; page is the default synthetic error Varnish serves when it cannot get a usable response from any backend. It displays &amp;ldquo;Guru Meditation&amp;rdquo; and an XID identifier, but nothing about the actual failure cause. The error page is a symptom, not a diagnosis.&lt;/p>
&lt;p>The root cause is always in the &lt;code>FetchError&lt;/code> tag in the shared memory log. Every Varnish-synthesised 503 is preceded by a backend transaction that logged a specific FetchError string: a timeout, a premature close, a protocol violation, or a health-probe failure that left no healthy backend to try. Reading that string is the single most important diagnostic step, and the one most operators skip.&lt;/p></description></item><item><title>Varnish ESI errors: broken pages and workspace pressure from Edge Side Includes</title><link>https://www.netdata.cloud/guides/varnish/varnish-esi-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-esi-errors/</guid><description>&lt;h1 id="varnish-esi-errors-broken-pages-and-workspace-pressure-from-edge-side-includes">Varnish ESI errors: broken pages and workspace pressure from Edge Side Includes&lt;/h1>
&lt;p>Pages assembled with Edge Side Includes (ESI) can fail in ways that look like application bugs but are cache-layer problems. When &lt;code>MAIN.esi_errors&lt;/code> increments, users receive partial pages with missing fragments or HTTP 500 responses from workspace exhaustion. &lt;code>MAIN.esi_warnings&lt;/code> counts ESI tags that Varnish skipped, silently dropping content from otherwise valid responses.&lt;/p>
&lt;p>Each &lt;code>&amp;lt;esi:include&amp;gt;&lt;/code> tag in a backend response triggers a full sub-request through Varnish&amp;rsquo;s VCL pipeline. A page with five includes means six VCL cycles: the parent plus five children. Each sub-request consumes a worker thread for its full lifecycle, allocates workspace memory for headers and VCL processing, and can itself contain ESI tags that spawn further sub-requests. The recursion is bounded by &lt;code>max_esi_depth&lt;/code>, &lt;!-- TODO: verify default is 5 in current Varnish --> but hitting that bound produces an error, not graceful degradation.&lt;/p></description></item><item><title>Varnish fetch_failed: backend connected but the fetch broke</title><link>https://www.netdata.cloud/guides/varnish/varnish-fetch-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-fetch-failed/</guid><description>&lt;h1 id="varnish-fetch_failed-backend-connected-but-the-fetch-broke">Varnish fetch_failed: backend connected but the fetch broke&lt;/h1>
&lt;p>The &lt;code>MAIN.fetch_failed&lt;/code> counter is climbing in varnishstat. Clients are seeing intermittent 503 responses. Backend health probes report healthy, and &lt;code>MAIN.backend_fail&lt;/code> is zero. The TCP connection to the backend succeeded, but the fetch itself broke.&lt;/p>
&lt;p>The problem is not connectivity. It is what happens after: malformed response headers, a truncated body, broken chunked encoding, a failed gzip decompression, or a timeout between bytes. Varnish established the connection, began the HTTP transaction, and something failed before a complete response was received.&lt;/p></description></item><item><title>Varnish grace masking a backend outage: the ticking-clock incident</title><link>https://www.netdata.cloud/guides/varnish/varnish-grace-masking-backend-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-grace-masking-backend-failure/</guid><description>&lt;h1 id="varnish-grace-masking-a-backend-outage-the-ticking-clock-incident">Varnish grace masking a backend outage: the ticking-clock incident&lt;/h1>
&lt;p>Your origin servers went down 12 minutes ago. The dashboard shows a 97% cache hit rate, zero 503s, and normal response times.&lt;/p>
&lt;p>This is grace masking. Varnish is serving cached objects past their TTL because no healthy backend can refresh them. Clients see stale 200s. Monitoring sees healthy cache traffic. The outage stays invisible until grace expires on enough popular objects, at which point 503s cascade.&lt;/p></description></item><item><title>Varnish Guru Meditation: reading the XID and tracing the failing request</title><link>https://www.netdata.cloud/guides/varnish/varnish-guru-meditation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-guru-meditation/</guid><description>&lt;h1 id="varnish-guru-meditation-reading-the-xid-and-tracing-the-failing-request">Varnish Guru Meditation: reading the XID and tracing the failing request&lt;/h1>
&lt;p>The Guru Meditation page is Varnish&amp;rsquo;s default 503 response. It fires when Varnish cannot serve from cache and cannot fetch from a backend. The page carries a transaction ID (XID) that maps directly to the shared memory log entries for that failed request. The XID is the fastest path from &amp;ldquo;users are seeing 503s&amp;rdquo; to &amp;ldquo;here is the FetchError, the backend, and the VCL subroutine where the failure happened.&amp;rdquo;&lt;/p></description></item><item><title>Varnish losthdr: HTTP headers silently dropped past http_max_hdr</title><link>https://www.netdata.cloud/guides/varnish/varnish-losthdr/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-losthdr/</guid><description>&lt;h1 id="varnish-losthdr-http-headers-silently-dropped-past-http_max_hdr">Varnish losthdr: HTTP headers silently dropped past http_max_hdr&lt;/h1>
&lt;p>MAIN.losthdr increments every time Varnish drops an HTTP header because the request or response exceeded the http_max_hdr limit (default 64). No error reaches the client in many cases. The request appears to succeed, but the header is gone, and Varnish does not reveal which one without querying the shared memory log.&lt;/p>
&lt;p>The consequences depend on which header was lost. A dropped Vary header causes cache poisoning: the wrong content variant gets cached and served to subsequent users for that URL. A dropped Authorization header means the backend never sees credentials, causing authentication failures or, depending on backend behavior, authentication bypass. A dropped Cache-Control header changes how the response is cached. None produce a Varnish error. They produce incorrect behavior that is difficult to trace unless you know this counter exists.&lt;/p></description></item><item><title>Varnish management CLI exposure: a -T bound to the world is full control</title><link>https://www.netdata.cloud/guides/varnish/varnish-management-cli-exposure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-management-cli-exposure/</guid><description>&lt;h1 id="varnish-management-cli-exposure-a--t-bound-to-the-world-is-full-control">Varnish management CLI exposure: a -T bound to the world is full control&lt;/h1>
&lt;p>The Varnish management CLI (&lt;code>varnishadm&lt;/code>, the &lt;code>-T&lt;/code> flag, default port 6082) is not a monitoring endpoint. It is a full administrative control plane. Through it, an authenticated user can load arbitrary VCL, inject bans, change every runtime parameter, and stop the cache child process. If VCL inline-C is enabled, loading a VCL is code execution on the Varnish host, running as the cache child user.&lt;/p></description></item><item><title>Varnish Monitoring</title><link>https://www.netdata.cloud/monitoring-101/varnish-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/varnish-monitoring/</guid><description>&lt;h2 id="varnish-monitoring">Varnish Monitoring&lt;/h2>
&lt;h3 id="what-is-varnish">What Is Varnish?&lt;/h3>
&lt;p>Varnish is a robust open-source HTTP accelerator, often used as a web cache. It efficiently stores copies of web pages and serves them to users quickly, reducing the time to render a requested web page. Whether supporting large-scale web operations or aiding e-commerce sites, Varnish effectively manages load and improves site performance.&lt;/p>
&lt;h3 id="monitoring-varnish-with-netdata">Monitoring Varnish With Netdata&lt;/h3>
&lt;p>Netdata offers an easy-to-use Varnish monitoring tool that keeps you informed on metrics essential to your infrastructure&amp;rsquo;s performance and efficiency. By presenting data on client sessions, cache hits and misses, thread activities, and more, Netdata provides a comprehensive view of your Varnish instances&amp;rsquo; health in real time. Explore more through our &lt;a href="https://app.netdata.cloud/spaces/netdata-demo/?utm_source=website&amp;amp;utm_content=monitoring101">Live Demo&lt;/a> or &lt;a href="https://app.netdata.cloud/?utm_source=website&amp;amp;utm_content=monitoring101">sign up for a free trial&lt;/a>.&lt;/p></description></item><item><title>Varnish monitoring checklist: the signals every production cache needs</title><link>https://www.netdata.cloud/guides/varnish/varnish-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-monitoring-checklist/</guid><description>&lt;h1 id="varnish-monitoring-checklist-the-signals-every-production-cache-needs">Varnish monitoring checklist: the signals every production cache needs&lt;/h1>
&lt;p>Varnish Cache is a reverse HTTP proxy that serves cached content from memory. It uses a dual-process architecture: a management process (root-owned, handles VCL compilation, child supervision, and the CLI) and a worker/child process (drops privileges, handles all cache operations). The management process restarts the child automatically on crash, so a child crash is not always a full outage. Repeated child restarts indicate a systemic problem.&lt;/p></description></item><item><title>Varnish monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/varnish/varnish-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-monitoring-maturity-model/</guid><description>&lt;h1 id="varnish-monitoring-maturity-model-from-survival-to-expert">Varnish monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most Varnish deployments start with the same three questions: is the process alive, is the hit ratio acceptable, are backends reachable. That is enough to catch a full outage. It is not enough to catch the slow failures that actually dominate production incidents: thread pool exhaustion with idle CPU, ban list growth turning every cache hit into a linear scan, transient storage OOM with no SMA counter movement, or grace mode silently masking a dead backend until the stale content expires.&lt;/p></description></item><item><title>Varnish n_lru_nuked vs n_expired: healthy eviction or an undersized cache</title><link>https://www.netdata.cloud/guides/varnish/varnish-lru-nuked-vs-expired/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-lru-nuked-vs-expired/</guid><description>&lt;h1 id="varnish-n_lru_nuked-vs-n_expired-healthy-eviction-or-an-undersized-cache">Varnish n_lru_nuked vs n_expired: healthy eviction or an undersized cache&lt;/h1>
&lt;p>Every Varnish cache evicts objects. The operational question is whether those evictions are healthy housekeeping or evidence that storage is too small for the working set. Two counters hold the answer. &lt;code>MAIN.n_expired&lt;/code> tracks objects that aged out past their TTL. &lt;code>MAIN.n_lru_nuked&lt;/code> tracks objects forcibly evicted from storage because a new object needed the space.&lt;/p>
&lt;p>The distinction matters because nuking is not inherently a problem. A right-sized cache continuously nukes the unpopular tail of its working set, which is correct LRU behavior. The problem arrives when nuking removes objects that would still serve hits, driving cache hit rate down and pushing more traffic to backends.&lt;/p></description></item><item><title>Varnish n_vcl accumulation: cold VCLs never discarded</title><link>https://www.netdata.cloud/guides/varnish/varnish-vcl-reload-accumulation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-vcl-reload-accumulation/</guid><description>&lt;h1 id="varnish-n_vcl-accumulation-cold-vcls-never-discarded">Varnish n_vcl accumulation: cold VCLs never discarded&lt;/h1>
&lt;p>&lt;code>MAIN.n_vcl&lt;/code> is climbing steadily in varnishstat. Each deploy loads a new VCL, and the count never comes back down. &lt;code>varnishadm vcl.list&lt;/code> shows dozens of VCLs in &amp;ldquo;available&amp;rdquo; state, most of them cold. Process RSS is creeping upward. This is VCL accumulation: automated reload pipelines that load new VCLs on every deploy but never run &lt;code>vcl.discard&lt;/code> on old ones.&lt;/p>
&lt;p>Cold VCLs (no longer serving traffic) have released some runtime resources per the VCL temperature system, but the compiled shared object remains loaded until explicitly discarded. A deployment pipeline that only runs &lt;code>vcl.load&lt;/code> and &lt;code>vcl.use&lt;/code> without &lt;code>vcl.discard&lt;/code> leaves every old VCL resident indefinitely.&lt;/p></description></item><item><title>Varnish not caching: Set-Cookie, Vary, and Cache-Control killing your hit rate</title><link>https://www.netdata.cloud/guides/varnish/varnish-not-caching-set-cookie/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-not-caching-set-cookie/</guid><description>&lt;h1 id="varnish-not-caching-set-cookie-vary-and-cache-control-killing-your-hit-rate">Varnish not caching: Set-Cookie, Vary, and Cache-Control killing your hit rate&lt;/h1>
&lt;p>Your hit rate is low and &lt;code>backend_req&lt;/code> tracks &lt;code>client_req&lt;/code> almost one-to-one. Varnish is up, storage has room, there are no error spikes, but &lt;code>cache_hitpass&lt;/code> or &lt;code>cache_hitmiss&lt;/code> dominates the cache outcome counters. The cache is running but not caching.&lt;/p>
&lt;p>The root cause is almost never Varnish itself. It is the interaction between backend response headers and the built-in VCL rules that determine cacheability. Three header problems account for the vast majority of cases: &lt;code>Set-Cookie&lt;/code> on every response, &lt;code>Cache-Control: private&lt;/code> or &lt;code>no-cache&lt;/code> on public content, and an overly broad &lt;code>Vary&lt;/code> header that fragments the cache into non-shareable variants.&lt;/p></description></item><item><title>Varnish pass vs miss: why s_pass and cache_miss are not the same thing</title><link>https://www.netdata.cloud/guides/varnish/varnish-pass-vs-miss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-pass-vs-miss/</guid><description>&lt;h1 id="varnish-pass-vs-miss-why-s_pass-and-cache_miss-are-not-the-same-thing">Varnish pass vs miss: why s_pass and cache_miss are not the same thing&lt;/h1>
&lt;p>A Varnish &lt;code>cache_miss&lt;/code> and an &lt;code>s_pass&lt;/code> both result in a backend fetch, and that surface similarity is where the confusion starts. Teams see backend request rates climbing, glance at the counters, and reach for the wrong fix. A miss is a cache lookup that found nothing and will try to cache the response. A pass is a deliberate VCL decision to bypass the cache entirely via &lt;code>return(pass)&lt;/code>. The difference matters because the two paths have different storage behavior, different tuning levers, and different operational consequences.&lt;/p></description></item><item><title>Varnish purge vs ban vs xkey: choosing an invalidation method that scales</title><link>https://www.netdata.cloud/guides/varnish/varnish-purge-vs-ban-xkey/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-purge-vs-ban-xkey/</guid><description>&lt;p>Varnish offers three mechanisms for invalidating cached objects: purge, ban, and the xkey VMOD (surrogate-key invalidation). Each has a fundamentally different cost model. Purge is O(1) and immediate but works only on a single exact hash. Bans are expression-based and flexible but accumulate in a list that every cache lookup must evaluate. xkey operates on secondary key indexes and avoids list growth entirely, but the open-source implementation has known scaling limits.&lt;/p></description></item><item><title>Varnish req.* vs obj.* bans: why req-level bans never leave the list</title><link>https://www.netdata.cloud/guides/varnish/varnish-req-vs-obj-bans/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-req-vs-obj-bans/</guid><description>&lt;h1 id="varnish-req-vs-obj-bans-why-req-level-bans-never-leave-the-list">Varnish req.* vs obj.* bans: why req-level bans never leave the list&lt;/h1>
&lt;p>req.* bans persist on the ban list because the ban lurker thread has no request context and cannot evaluate them asynchronously. The list grows, lookup latency degrades across every request, and the fix is the same in nearly every case: rewrite req-based bans as obj-based bans so the lurker can process them in the background. For the broader failure pattern, see &lt;a href="https://www.netdata.cloud/guides/varnish/varnish-ban-list-growing/">Varnish ban list growing&lt;/a>.&lt;/p></description></item><item><title>Varnish sc_rapid_reset: the HTTP/2 Rapid Reset DDoS (CVE-2023-44487)</title><link>https://www.netdata.cloud/guides/varnish/varnish-http2-rapid-reset/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-http2-rapid-reset/</guid><description>&lt;h1 id="varnish-sc_rapid_reset-the-http2-rapid-reset-ddos-cve-2023-44487">Varnish sc_rapid_reset: the HTTP/2 Rapid Reset DDoS (CVE-2023-44487)&lt;/h1>
&lt;p>&lt;code>sc_rapid_reset&lt;/code> incrementing on a Varnish instance means the HTTP/2 Rapid Reset rate limiter has closed at least one session for exceeding its configured reset budget. This is the primary detection signal for CVE-2023-44487, the HTTP/2 Rapid Reset DDoS disclosed in October 2023. The attack exploits an asymmetry in HTTP/2: a client can open hundreds of streams per connection and immediately cancel each one with an RST_STREAM frame, forcing the server to do real work (stream setup, header parsing, resource allocation) at near-zero cost to the attacker.&lt;/p></description></item><item><title>Varnish sess_dropped vs req_dropped: HTTP/1 connection drops and HTTP/2 stream drops</title><link>https://www.netdata.cloud/guides/varnish/varnish-sess-dropped-req-dropped/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-sess-dropped-req-dropped/</guid><description>&lt;h1 id="varnish-sess_dropped-vs-req_dropped-http1-connection-drops-and-http2-stream-drops">Varnish sess_dropped vs req_dropped: HTTP/1 connection drops and HTTP/2 stream drops&lt;/h1>
&lt;p>Two Varnish counters track the worst outcome a cache can produce: a client gets nothing. &lt;code>MAIN.sess_dropped&lt;/code> counts HTTP/1 sessions dropped because the worker thread queue was full. &lt;code>MAIN.req_dropped&lt;/code> counts HTTP/2 streams and other request types dropped for the same reason. Both mean Varnish refused to serve a client because no thread was available and the bounded queue was at its limit.&lt;/p></description></item><item><title>Varnish sess_fail: session accept failures at the front door</title><link>https://www.netdata.cloud/guides/varnish/varnish-sess-fail/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-sess-fail/</guid><description>&lt;h1 id="varnish-sess_fail-session-accept-failures-at-the-front-door">Varnish sess_fail: session accept failures at the front door&lt;/h1>
&lt;p>When &lt;code>MAIN.sess_fail&lt;/code> starts incrementing in Varnish, new TCP connections are failing at the accept() call. Varnish never brings them into the worker pipeline. Clients experience connection resets, timeouts, or refused connections. Unlike &lt;code>sess_dropped&lt;/code> (where the connection was accepted but the thread queue was full), &lt;code>sess_fail&lt;/code> means Varnish never got far enough to process the request.&lt;/p>
&lt;p>The counter is an aggregate. Since Varnish 6.1 &lt;!-- TODO: verify whether sub-counters were also backported to 6.0 LTS releases -->, it decomposes into sub-counters that isolate the cause: &lt;code>sess_fail_emfile&lt;/code> for file descriptor exhaustion, &lt;code>sess_fail_econnaborted&lt;/code> for client-side aborts, &lt;code>sess_fail_enomem&lt;/code> for memory pressure, and several others. Reading only the aggregate counter is a common diagnostic mistake.&lt;/p></description></item><item><title>Varnish session close reasons: reading the sc_* counters</title><link>https://www.netdata.cloud/guides/varnish/varnish-session-close-reasons/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-session-close-reasons/</guid><description>&lt;h1 id="varnish-session-close-reasons-reading-the-sc_-counters">Varnish session close reasons: reading the sc_* counters&lt;/h1>
&lt;p>The MAIN.sc_* counters in Varnish record why every client session ended. Each close reason is a DIAG-level counter visible through &lt;code>varnishstat&lt;/code>. The distribution across these counters is one of the fastest triage signals available: it tells you whether sessions are ending normally or whether the client, the network, or Varnish itself is causing problems.&lt;/p>
&lt;p>The sc_* family has grown across versions. Some counters shifted accounting between releases, and several error counters are benign under certain traffic patterns. Treating every nonzero error counter as an incident leads to alert fatigue; treating them all as noise means missing real problems.&lt;/p></description></item><item><title>Varnish shared-memory log overruns: losing the evidence during an incident</title><link>https://www.netdata.cloud/guides/varnish/varnish-shm-log-overrun/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-shm-log-overrun/</guid><description>&lt;h1 id="varnish-shared-memory-log-overruns-losing-the-evidence-during-an-incident">Varnish shared-memory log overruns: losing the evidence during an incident&lt;/h1>
&lt;p>Cache hit rate dropped, backend request rate spiked, users saw elevated 503s. You need to reconstruct what happened from Varnish request logs: which URLs triggered the problem, what backend errors were returned, how long fetches took. You reach for &lt;code>varnishncsa&lt;/code> output or &lt;code>varnishlog&lt;/code> traces and find gaps, partial records, or nothing at all.&lt;/p>
&lt;p>Varnish served traffic throughout the incident. The logging subsystem quietly discarded data when you needed it most.&lt;/p></description></item><item><title>Varnish slow responses: hit latency, miss latency, and why varnishstat has no timing</title><link>https://www.netdata.cloud/guides/varnish/varnish-slow-response-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-slow-response-latency/</guid><description>&lt;h1 id="varnish-slow-responses-hit-latency-miss-latency-and-why-varnishstat-has-no-timing">Varnish slow responses: hit latency, miss latency, and why varnishstat has no timing&lt;/h1>
&lt;p>Users report slow responses through Varnish. You open varnishstat and look for a latency counter. There isn&amp;rsquo;t one. Varnish&amp;rsquo;s counter subsystem maintains event counts and gauges: &lt;code>cache_hit&lt;/code>, &lt;code>cache_miss&lt;/code>, &lt;code>threads&lt;/code>, &lt;code>backend_fail&lt;/code>. None of them measure how long a request took.&lt;/p>
&lt;p>This is architectural. Varnish writes per-request timing into the shared memory log (VSL) as Timestamp tags with per-phase resolution. The tools to read them are varnishlog, varnishncsa, and varnishhist. If you are diagnosing slow responses using varnishstat alone, you have no timing data.&lt;/p></description></item><item><title>Varnish SMA c_fail: allocation failures and malloc fragmentation</title><link>https://www.netdata.cloud/guides/varnish/varnish-storage-allocation-failure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-storage-allocation-failure/</guid><description>&lt;h1 id="varnish-sma-c_fail-allocation-failures-and-malloc-fragmentation">Varnish SMA c_fail: allocation failures and malloc fragmentation&lt;/h1>
&lt;p>&lt;code>SMA.&amp;lt;name&amp;gt;.c_fail&lt;/code> counts storage allocation failures after LRU eviction. Under normal conditions it stays at zero. When it climbs, memory management is wrong, and the cause is not always &amp;ldquo;the cache is full.&amp;rdquo;&lt;/p>
&lt;p>The deceptive variant is the &lt;code>malloc&lt;/code> stevedore. Over weeks of object churn, the underlying allocator (jemalloc or glibc malloc) fragments the heap. Varnish reports free bytes in &lt;code>SMA.&amp;lt;name&amp;gt;.g_space&lt;/code>, but those bytes are scattered across many small free regions that cannot satisfy a contiguous allocation. Varnish evicts objects, retries, and fails. &lt;code>c_fail&lt;/code> increments while &lt;code>g_space&lt;/code> insists there is room.&lt;/p></description></item><item><title>Varnish storage full: the LRU nuke storm and cache thrashing</title><link>https://www.netdata.cloud/guides/varnish/varnish-storage-full-lru-nuking/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-storage-full-lru-nuking/</guid><description>&lt;h1 id="varnish-storage-full-the-lru-nuke-storm-and-cache-thrashing">Varnish storage full: the LRU nuke storm and cache thrashing&lt;/h1>
&lt;p>Varnish cache hit ratio is declining and backend request rate is climbing. Storage utilization is near 100%. &lt;code>MAIN.n_lru_nuked&lt;/code> is incrementing steadily. You are in a storage exhaustion cascade: the cache is too small for the working set, and every new object forces Varnish to evict an existing one via LRU. Evicted objects get re-requested, miss the cache, trigger a backend fetch, and the cycle repeats.&lt;/p></description></item><item><title>Varnish storage sizing: malloc vs file, and why -s malloc,ALL-RAM kills you</title><link>https://www.netdata.cloud/guides/varnish/varnish-storage-sizing-malloc-file/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-storage-sizing-malloc-file/</guid><description>&lt;h1 id="varnish-storage-sizing-malloc-vs-file-and-why--s-mallocall-ram-kills-you">Varnish storage sizing: malloc vs file, and why -s malloc,ALL-RAM kills you&lt;/h1>
&lt;p>The &lt;code>-s malloc&lt;/code> startup parameter controls how much memory Varnish allocates for cached object storage. It does not control total process memory. That distinction is the single most expensive misunderstanding in Varnish operations.&lt;/p>
&lt;p>When you set &lt;code>-s malloc,32G&lt;/code> on a 32 GB machine, you have allocated 32 GB for object storage and left nothing for the operating system, the Varnish process itself, worker thread stacks, per-request workspace, transient storage, and allocator fragmentation. The kernel OOM killer sees a process consuming all available memory and terminates it. Varnish restarts, the management process reloads the child, the cache warms from zero, and the cycle repeats if the configuration has not changed.&lt;/p></description></item><item><title>Varnish thread pool exhaustion: workers all busy, queue full, sessions dropped</title><link>https://www.netdata.cloud/guides/varnish/varnish-thread-pool-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-thread-pool-exhaustion/</guid><description>&lt;h1 id="varnish-thread-pool-exhaustion-workers-all-busy-queue-full-sessions-dropped">Varnish thread pool exhaustion: workers all busy, queue full, sessions dropped&lt;/h1>
&lt;p>Varnish is dropping client connections and CPU looks idle. The process is running, the management CLI responds, and backends are technically up, but users see connection resets or stream resets. This is thread pool exhaustion: all worker threads are busy, the request queue is full, and Varnish has started refusing traffic.&lt;/p>
&lt;p>Varnish&amp;rsquo;s concurrency model is thread-per-request with a bounded pool. Each pool holds between &lt;code>thread_pool_min&lt;/code> and &lt;code>thread_pool_max&lt;/code> threads, and Varnish runs &lt;code>thread_pools&lt;/code> pools (default 2). When a request arrives and no idle worker is available, it enters a bounded queue (&lt;code>thread_queue_limit&lt;/code>, default 20 per pool). When the queue is full, the session or stream is dropped.&lt;/p></description></item><item><title>Varnish thread pool tuning: thread_pool_min, thread_pool_max, and thread_pools</title><link>https://www.netdata.cloud/guides/varnish/varnish-thread-pool-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-thread-pool-tuning/</guid><description>&lt;h1 id="varnish-thread-pool-tuning-thread_pool_min-thread_pool_max-and-thread_pools">Varnish thread pool tuning: thread_pool_min, thread_pool_max, and thread_pools&lt;/h1>
&lt;p>Varnish uses a thread-per-request concurrency model with a bounded pool. Every client request occupies a worker thread for its entire lifecycle, from accept through response delivery. When the pool is full and the overflow queue overflows, sessions are dropped.&lt;/p>
&lt;p>The three parameters that control this behavior, &lt;code>thread_pools&lt;/code>, &lt;code>thread_pool_min&lt;/code>, and &lt;code>thread_pool_max&lt;/code>, determine how Varnish responds to load spikes, slow backends, and idle periods. The dominant factor in thread pool sizing is backend response time, not request rate. A cache miss that holds a thread for 2 seconds consumes the same pool capacity as thousands of sub-millisecond cache hits. Slow backends are the single most common cause of thread pool exhaustion, and they make &lt;code>thread_pool_max&lt;/code> the parameter that matters most under stress.&lt;/p></description></item><item><title>Varnish thread_queue_len above zero: requests waiting for a worker</title><link>https://www.netdata.cloud/guides/varnish/varnish-thread-queue-len/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-thread-queue-len/</guid><description>&lt;h1 id="varnish-thread_queue_len-above-zero-requests-waiting-for-a-worker">Varnish thread_queue_len above zero: requests waiting for a worker&lt;/h1>
&lt;p>&lt;code>MAIN.thread_queue_len&lt;/code> is the point-in-time count of client sessions sitting in Varnish&amp;rsquo;s bounded worker queue, waiting for an idle thread. In normal operation it is zero. Any sustained nonzero value means every worker thread across all pools is busy and incoming requests are piling up in the last buffer before Varnish starts dropping them.&lt;/p>
&lt;p>The thread pool has a cliff-edge failure curve: performance is fine until the pool is saturated, then requests queue, then the queue fills, then sessions are dropped. The distance between &amp;ldquo;queue length is 1&amp;rdquo; and &amp;ldquo;clients are getting connection resets&amp;rdquo; can be seconds if the queue limit is small (the default &lt;code>thread_queue_limit&lt;/code> is 20 per pool).&lt;/p></description></item><item><title>Varnish threads_failed: the OS refusing to create worker threads</title><link>https://www.netdata.cloud/guides/varnish/varnish-threads-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-threads-failed/</guid><description>&lt;h1 id="varnish-threads_failed-the-os-refusing-to-create-worker-threads">Varnish threads_failed: the OS refusing to create worker threads&lt;/h1>
&lt;p>When &lt;code>MAIN.threads_failed&lt;/code> is nonzero, the Linux kernel is refusing Varnish&amp;rsquo;s &lt;code>pthread_create()&lt;/code> calls. This is an OS-level resource limit preventing the worker thread pool from growing, not a Varnish configuration problem.&lt;/p>
&lt;p>Do not confuse this with &lt;code>MAIN.threads_limited&lt;/code>, which increments when Varnish declines to create a thread because &lt;code>thread_pool_max&lt;/code> has been reached. &lt;code>threads_limited&lt;/code> means Varnish chose not to create the thread. &lt;code>threads_failed&lt;/code> means Varnish tried and the OS said no. The remediation differs: &lt;code>threads_limited&lt;/code> requires raising a Varnish parameter; &lt;code>threads_failed&lt;/code> requires fixing OS-level limits.&lt;/p></description></item><item><title>Varnish threads_limited climbing: hitting thread_pool_max</title><link>https://www.netdata.cloud/guides/varnish/varnish-threads-limited/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-threads-limited/</guid><description>&lt;h1 id="varnish-threads_limited-climbing-hitting-thread_pool_max">Varnish threads_limited climbing: hitting thread_pool_max&lt;/h1>
&lt;p>&lt;code>MAIN.threads_limited&lt;/code> increments each time Varnish wanted to create a new worker thread but &lt;code>thread_pool_max&lt;/code> prevented it. Movement during a traffic burst is expected. A sustained nonzero rate during steady-state operation means the configured thread pool ceiling is a binding constraint on concurrency, and if the bounded queue behind the pool fills, Varnish will drop sessions.&lt;/p>
&lt;p>The default &lt;code>thread_pool_max&lt;/code> is 5000 threads per pool. With the default &lt;code>thread_pools&lt;/code> value of 2, the absolute ceiling is 10,000 worker threads. For most workloads that ceiling is high enough that the counter stays at zero. The problem arises when backends are slow: each thread is held longer per request, effective concurrency drops, and the pool fills at a request rate that would be trivial if backends responded in single-digit milliseconds.&lt;/p></description></item><item><title>Varnish Too many open files: EMFILE, sess_fail_emfile, and the FD ceiling</title><link>https://www.netdata.cloud/guides/varnish/varnish-too-many-open-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-too-many-open-files/</guid><description>&lt;h1 id="varnish-too-many-open-files-emfile-sess_fail_emfile-and-the-fd-ceiling">Varnish Too many open files: EMFILE, sess_fail_emfile, and the FD ceiling&lt;/h1>
&lt;p>&lt;code>sess_fail_emfile&lt;/code> is climbing and new client connections are being refused. The Varnish child process has hit its file descriptor ceiling. This is a cliff-edge failure: once the FD limit is exhausted, Varnish cannot accept new connections, cannot open backend connections, and can panic if internal operations that require FDs fail. There is no graceful degradation and no queuing.&lt;/p>
&lt;p>The default &lt;code>ulimit -n&lt;/code> of 1024 on most Linux distributions is far too low for production Varnish. The Varnish package&amp;rsquo;s default systemd unit sets &lt;code>LimitNOFILE=131072&lt;/code>, but inherited or misconfigured environments often leave the lower default in place.&lt;/p></description></item><item><title>Varnish transient storage OOM: the unbounded memory path that kills the process</title><link>https://www.netdata.cloud/guides/varnish/varnish-transient-storage-oom/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-transient-storage-oom/</guid><description>&lt;h1 id="varnish-transient-storage-oom-the-unbounded-memory-path-that-kills-the-process">Varnish transient storage OOM: the unbounded memory path that kills the process&lt;/h1>
&lt;p>Varnish dies. The OOM killer takes it. No panic, no crash, no VCL error in the logs. Process RSS was far above your configured &lt;code>-s malloc,8G&lt;/code> storage, and the kernel reclaimed memory the only way it knows how. This is the transient storage OOM path: the number-one missed signal in Varnish operations.&lt;/p>
&lt;p>The problem is structural. Varnish has two storage paths: your configured cache storage (malloc or file, capped at a known size) and transient storage. Transient storage holds objects that will never be cached: pass responses, hit-for-pass, hit-for-miss, and piped traffic. By default, transient storage uses malloc with no upper limit. A pass storm, a batch of large uncacheable responses, or a slow leak over weeks can push process RSS past the configured storage size and into system memory the kernel will not let you keep.&lt;/p></description></item><item><title>Varnish unauthorized PURGE/BAN: cache invalidation as an attack surface</title><link>https://www.netdata.cloud/guides/varnish/varnish-unauthorized-purge-ban/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-unauthorized-purge-ban/</guid><description>&lt;h1 id="varnish-unauthorized-purgeban-cache-invalidation-as-an-attack-surface">Varnish unauthorized PURGE/BAN: cache invalidation as an attack surface&lt;/h1>
&lt;p>Varnish treats PURGE and BAN as ordinary HTTP methods. It has no built-in VCL that handles, gates, or rejects them. If your custom VCL handles these methods but does not check an ACL first, anyone who can reach your Varnish listener can invalidate cached objects. Unauthenticated Varnish PURGE has been reported as a valid security finding against major platforms through bug bounty programs.&lt;/p></description></item><item><title>Varnish VCL compilation failed: reload rejected and the old VCL still live</title><link>https://www.netdata.cloud/guides/varnish/varnish-vcl-compilation-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-vcl-compilation-failed/</guid><description>&lt;h1 id="varnish-vcl-compilation-failed-reload-rejected-and-the-old-vcl-still-live">Varnish VCL compilation failed: reload rejected and the old VCL still live&lt;/h1>
&lt;p>You deployed a new VCL file, ran &lt;code>systemctl reload varnish&lt;/code> or &lt;code>varnishadm vcl.load&lt;/code>, and the reload was rejected. The VCC-compiler printed an error with a line and column number, and the management process refused to load the new configuration. The previous VCL is still serving traffic.&lt;/p>
&lt;p>This is not an outage. Varnish reloads are hitless: the new VCL is compiled and loaded into memory before activation. When compilation fails, the old VCL stays in place and no client sees an interruption. But your change did not land. Until you fix the compilation error and retry, the running configuration is whatever was active before the failed deploy.&lt;/p></description></item><item><title>Varnish vcl_fail: VCL runtime errors during request processing</title><link>https://www.netdata.cloud/guides/varnish/varnish-vcl-fail/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-vcl-fail/</guid><description>&lt;p>&lt;code>MAIN.vcl_fail&lt;/code> is incrementing. The VCL loaded and compiled successfully at &lt;code>vcl.load&lt;/code> time, but something is breaking at runtime while real requests flow through the VCL subroutines.&lt;/p>
&lt;p>The &lt;code>vcl_fail&lt;/code> counter (Varnish 6.0 and later) counts failures that prevented VCL from completing. Each increment means a request could not finish its normal VCL lifecycle. On the client side, the request is diverted to &lt;code>vcl_synth&lt;/code> with a 503 status. On the backend side, the fetch fails. In both cases the user gets an error response instead of the content they asked for.&lt;/p></description></item><item><title>Varnish workspace overflow types: client, backend, thread, and session</title><link>https://www.netdata.cloud/guides/varnish/varnish-workspace-overflow-types/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-workspace-overflow-types/</guid><description>&lt;h1 id="varnish-workspace-overflow-types-client-backend-thread-and-session">Varnish workspace overflow types: client, backend, thread, and session&lt;/h1>
&lt;p>Varnish allocates small, fixed-size memory regions called workspaces to hold HTTP headers, VCL string operations, and intermediate request state during processing. Each stage of the request lifecycle draws from a different workspace type, with its own size parameter and overflow counter. When any workspace runs out of room, Varnish cannot finish processing the request and returns HTTP 500 to the client. This is a 500, not a 503: the backend did not fail. Varnish itself ran out of working memory.&lt;/p></description></item><item><title>Varnish workspace_client overflow: 500 errors from oversized headers and cookies</title><link>https://www.netdata.cloud/guides/varnish/varnish-workspace-client-overflow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/varnish/varnish-workspace-client-overflow/</guid><description>&lt;h1 id="varnish-workspace_client-overflow-500-errors-from-oversized-headers-and-cookies">Varnish workspace_client overflow: 500 errors from oversized headers and cookies&lt;/h1>
&lt;p>HTTP 500 responses from Varnish when the backend is healthy, the object is in cache, the thread pool is not saturated, and the 503 backend-fetch path is not involved. The 500s come from Varnish itself, and they cluster on requests carrying large &lt;code>Cookie&lt;/code> headers, long URLs, deep &lt;code>Via&lt;/code> or &lt;code>X-Forwarded-For&lt;/code> chains, or responses where VCL has added many synthetic headers.&lt;/p></description></item><item><title>Vault</title><link>https://www.netdata.cloud/integrations/all/vault/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/all/vault/</guid><description/></item><item><title>Vault PKI</title><link>https://www.netdata.cloud/integrations/data-collection/applications/vault-pki/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/vault-pki/</guid><description/></item><item><title>Vault PKI Monitoring</title><link>https://www.netdata.cloud/monitoring-101/vault_pki-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/vault_pki-monitoring/</guid><description>&lt;h2 id="vault-pki-monitoring">Vault PKI Monitoring&lt;/h2>
&lt;h3 id="what-is-vault-pki">What Is Vault PKI?&lt;/h3>
&lt;p>Vault PKI refers to the Public Key Infrastructure provided by HashiCorp Vault, a powerful system designed to manage sensitive information like secrets, encryption keys, and certificates. It&amp;rsquo;s an essential component in ensuring secure communication, especially when dealing with distributed systems. Vault PKI helps automate certificate management, making it a crucial element for security.&lt;/p>
&lt;h3 id="monitoring-vault-pki-with-netdata">Monitoring Vault PKI With Netdata&lt;/h3>
&lt;p>To effectively monitor Vault PKI, Netdata leverages an openmetrics exporter called the &lt;a href="https://github.com/aarnaud/vault-pki-exporter">Vault PKI Exporter&lt;/a>. This exporter collects critical metrics on your Vault PKI setup, and Netdata can seamlessly ingest this data without relying on a separate Prometheus server or Grafana dashboards. With Netdata, you get automated dashboards, real-time alerts, and comprehensive visibility into your Vault PKI metrics and beyond, enabling proactive monitoring and troubleshooting.&lt;/p></description></item><item><title>vCenter '503 Service Unavailable': the vSphere Client will not load</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-503-service-unavailable/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-503-service-unavailable/</guid><description>&lt;h1 id="vcenter-503-service-unavailable-the-vsphere-client-will-not-load">vCenter &amp;lsquo;503 Service Unavailable&amp;rsquo;: the vSphere Client will not load&lt;/h1>
&lt;p>A 503 from the vSphere Client means the reverse HTTP proxy (&lt;code>rhttpproxy&lt;/code>) accepted the TLS connection but could not reach the backend it routes to. The proxy itself is healthy. One of its dependents, typically &lt;code>vpxd&lt;/code>, &lt;code>vmware-vapi-endpoint&lt;/code>, &lt;code>vmware-stsd&lt;/code> (STS), or the HTML5 client backend (&lt;code>vsphere-ui&lt;/code>), is stopped, still starting, or crash-looping. The error string often reads &amp;ldquo;Initialization of one of the components failed.&amp;rdquo;&lt;/p></description></item><item><title>vCenter 'Cannot complete login due to an incorrect user name or password': SSO failures</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-cannot-login/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-cannot-login/</guid><description>&lt;h1 id="vcenter-cannot-complete-login-due-to-an-incorrect-user-name-or-password-sso-failures">vCenter &amp;lsquo;Cannot complete login due to an incorrect user name or password&amp;rsquo;: SSO failures&lt;/h1>
&lt;p>The &amp;ldquo;Cannot complete login due to an incorrect user name or password&amp;rdquo; string is the exact message operators see in the vSphere Client, in PowerCLI sessions, and in API responses when SSO authentication fails. The text is misleading: the cause is rarely a typo. For a single user it is usually a credential or permission problem. For every account at once it is an SSO/STS infrastructure failure.&lt;/p></description></item><item><title>vCenter /storage/db full: vPostgres stops and the whole management plane dies</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-storage-db-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-storage-db-full/</guid><description>&lt;h1 id="vcenter-storagedb-full-vpostgres-stops-and-the-whole-management-plane-dies">vCenter /storage/db full: vPostgres stops and the whole management plane dies&lt;/h1>
&lt;p>The vSphere Client returns 503 Service Unavailable. PowerCLI sessions hang and time out. DRS has stopped evaluating, vMotion orchestration is gone, and provisioning fails. Running VMs on the ESXi hosts continue to operate, but the management plane is gone.&lt;/p>
&lt;p>The root cause is almost certainly the &lt;code>/storage/db&lt;/code> partition on the vCenter Server Appliance (VCSA). This is where vPostgres keeps its data files. At 95% utilization on any partition, VMware automatically shuts down &lt;code>vmware-vpxd&lt;/code> to protect the database from corruption. At 100%, vPostgres cannot extend a data file or write a WAL record and crashes. Once vPostgres is down, &lt;code>vpxd&lt;/code> has no database and cannot restart.&lt;/p></description></item><item><title>vCenter /storage/log full: the log-bomb disk death spiral</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-storage-log-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-storage-log-full/</guid><description>&lt;h1 id="vcenter-storagelog-full-the-log-bomb-disk-death-spiral">vCenter /storage/log full: the log-bomb disk death spiral&lt;/h1>
&lt;p>You log in to the vSphere Client and get a 503, or the UI hangs mid-task. You SSH into the VCSA and run &lt;code>df -h&lt;/code>: &lt;code>/storage/log&lt;/code> is at 100%. The root filesystem may still have plenty of free space, which is why a generic disk-space alert missed it. The VCSA has many dedicated partitions, and they fill independently.&lt;/p>
&lt;p>A single failing service can write gigabytes of logs per hour. STS authentication failures, database connection errors, alarm flapping, or a misbehaving SDK client flooding vpxd with errors will take &lt;code>/storage/log&lt;/code> from 40% to 100% within hours. Once the partition is full, services that try to log crash. vmon restarts them. The restart itself generates more log lines as the service hits the same fault and tries to log it again. The loop is self-reinforcing, and clearing space temporarily makes the next iteration worse because the service can write again, refilling the partition faster.&lt;/p></description></item><item><title>vCenter /storage/seat full: stats, events, alarms, and tasks outgrowing their partition</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-storage-seat-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-storage-seat-full/</guid><description>&lt;h1 id="vcenter-storageseat-full-stats-events-alarms-and-tasks-outgrowing-their-partition">vCenter /storage/seat full: stats, events, alarms, and tasks outgrowing their partition&lt;/h1>
&lt;p>The &lt;code>/storage/seat&lt;/code> partition on the vCenter Server Appliance (VCSA) holds the vPostgres tables for Stats, Events, Alarms, and Tasks: &lt;code>vpx_event&lt;/code>, &lt;code>vpx_event_arg&lt;/code>, &lt;code>vpx_task&lt;/code>, and the &lt;code>vpxd_hist_stat*&lt;/code> rollup tables. In modern VCSA it is a dedicated mount, so it can fill while &lt;code>/storage/db&lt;/code>, &lt;code>/storage/log&lt;/code>, and &lt;code>/&lt;/code> all show healthy utilization. Operators checking only &lt;code>/&lt;/code> or the VAMI dashboard&amp;rsquo;s &amp;ldquo;VCDB&amp;rdquo; usage will miss it until vpxd refuses to start.&lt;/p></description></item><item><title>vCenter appliance undersized: inventory outgrowing the deployment size</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-vcsa-undersized/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-vcsa-undersized/</guid><description>&lt;h1 id="vcenter-appliance-undersized-inventory-outgrowing-the-deployment-size">vCenter appliance undersized: inventory outgrowing the deployment size&lt;/h1>
&lt;p>A VCSA deployed as &amp;ldquo;Small&amp;rdquo; three years ago is not the same as a fresh &amp;ldquo;Small&amp;rdquo; today if your inventory has grown. The VCSA&amp;rsquo;s resource footprint is not fixed by CPU and RAM alone. It is dominated by the size of the in-memory inventory cache that vpxd maintains, and that cache grows with managed-object count: VMs, hosts, datastores, dvSwitch portgroups, tags, alarms, permissions, and snapshots.&lt;/p></description></item><item><title>vCenter certificate expired: the STS signing cert outage nobody saw coming</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-certificate-expired/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-certificate-expired/</guid><description>&lt;h1 id="vcenter-certificate-expired-the-sts-signing-cert-outage-nobody-saw-coming">vCenter certificate expired: the STS signing cert outage nobody saw coming&lt;/h1>
&lt;p>vCenter is down. Not &amp;ldquo;slow&amp;rdquo; or &amp;ldquo;degraded.&amp;rdquo; Down. The vSphere Client shows a white screen or a 503. PowerCLI sessions fail to connect. API calls return authentication errors. ESXi hosts show as disconnected in bulk. Every integration that depends on vCenter (NSX, vRA, SRM, backup products) has lost connectivity simultaneously. VMs on the hosts are still running, but you cannot manage, migrate, or orchestrate anything.&lt;/p></description></item><item><title>vCenter clock skew: the NTP offset that breaks tokens and disconnects hosts</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-ntp-clock-skew/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-ntp-clock-skew/</guid><description>&lt;h1 id="vcenter-clock-skew-the-ntp-offset-that-breaks-tokens-and-disconnects-hosts">vCenter clock skew: the NTP offset that breaks tokens and disconnects hosts&lt;/h1>
&lt;p>Sudden SSO login failures. ESXi hosts flipping to &amp;ldquo;Not Responding&amp;rdquo; without a network cause. Certificate validation errors on certificates you know are valid. All three point to clock skew between vCenter, ESXi, and the identity infrastructure they depend on.&lt;/p>
&lt;p>SAML token validation, certificate validation, Kerberos, HA heartbeats, and log correlation all assume clocks agree. When they diverge, the failures look like unrelated problems instead of one root cause. Nobody thinks to check the clock.&lt;/p></description></item><item><title>vCenter CPU and memory pressure: vpxd heap, swap, and the undersized appliance</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-cpu-memory-pressure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-cpu-memory-pressure/</guid><description>&lt;h1 id="vcenter-cpu-and-memory-pressure-vpxd-heap-swap-and-the-undersized-appliance">vCenter CPU and memory pressure: vpxd heap, swap, and the undersized appliance&lt;/h1>
&lt;p>vCenter slows down before it fails. The vSphere Client takes 30 seconds to load a VM summary, API calls time out from backup and monitoring tools, hosts flap to &amp;ldquo;not responding&amp;rdquo; even though they are healthy, and DRS recommendations stop appearing. Inside the appliance, &lt;code>top&lt;/code> shows vpxd consuming most of the memory and a small but non-zero amount of swap. From the hypervisor, the VCSA VM either looks busy or, worse, looks idle while sitting at 50% CPU ready time.&lt;/p></description></item><item><title>vCenter database bloat: SEAT tables, statistics level, and the failed purge job</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-database-bloat/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-database-bloat/</guid><description>&lt;h1 id="vcenter-database-bloat-seat-tables-statistics-level-and-the-failed-purge-job">vCenter database bloat: SEAT tables, statistics level, and the failed purge job&lt;/h1>
&lt;p>vCenter Server&amp;rsquo;s embedded PostgreSQL database (vPostgres) grows continuously. Under normal conditions, internal purge jobs keep the SEAT tables (Stats, Events, Alarms, Tasks) bounded by the configured retention windows. When retention is misconfigured, the statistics level is too high, or the purge job stops running, those tables grow without bound. The result is a slow vCenter, a filling &lt;code>/storage/seat&lt;/code> or &lt;code>/storage/db&lt;/code> partition, and eventually a vpxd crash when the database can no longer write.&lt;/p></description></item><item><title>vCenter HA (VCHA) replication broken: protection that is not protecting</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-vcha-replication/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-vcha-replication/</guid><description>&lt;h1 id="vcenter-ha-vcha-replication-broken-protection-that-is-not-protecting">vCenter HA (VCHA) replication broken: protection that is not protecting&lt;/h1>
&lt;p>VCHA makes vCenter look protected. The VAMI dashboard shows a green cluster, the active node serves traffic, and the passive node exists as a standby. When PostgreSQL streaming replication between the active and passive node stops, the passive node holds a stale database. A failover loses every transaction written since replication broke.&lt;/p>
&lt;p>VCHA&amp;rsquo;s health surface is shallow. The VAMI summary can report healthy while &lt;code>pg_stat_replication&lt;/code> shows &lt;code>NOT_REPLICATING&lt;/code>. The passive VM is powered on. The witness is unreachable, so automated failover cannot reach quorum. None of this surfaces until an operator triggers failover and discovers the passive is hours behind, or until WAL accumulation on the active fills &lt;code>/storage/db&lt;/code> and vpxd stops.&lt;/p></description></item><item><title>vCenter Server Appliance</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/vcenter-server-appliance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/vcenter-server-appliance/</guid><description/></item><item><title>vCenter Server Appliance Monitoring</title><link>https://www.netdata.cloud/monitoring-101/vcsa-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/vcsa-monitoring/</guid><description>&lt;h2 id="vcenter-server-appliance-monitoring">vCenter Server Appliance Monitoring&lt;/h2>
&lt;h3 id="what-is-vcenter-server-appliance">What Is vCenter Server Appliance?&lt;/h3>
&lt;p>vCenter Server Appliance (vCSA) is a powerful, preconfigured Linux-based virtual machine optimized for running VMware vCenter Server and associated services. It is a vital component in managing virtualized environments, providing centralized management of virtualized hosts and virtual machines from a single console.&lt;/p>
&lt;h3 id="monitoring-vcenter-server-appliance-with-netdata">Monitoring vCenter Server Appliance With Netdata&lt;/h3>
&lt;p>Netdata offers a seamless and efficient way to monitor vCenter Server Appliance. As a comprehensive monitoring tool, it provides real-time insights and detailed metrics that are crucial for maintaining the health and performance of your vCSA environment. With Netdata’s &lt;a href="https://www.netdata.cloud/">free and open-source platform&lt;/a>, you get simple configurations, interactive visualizations, and a low-overhead monitoring solution.&lt;/p></description></item><item><title>vCenter service will not start: vmon dependency order and the max-restart wall</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-service-dependency-deadlock/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-service-dependency-deadlock/</guid><description>&lt;h1 id="vcenter-service-will-not-start-vmon-dependency-order-and-the-max-restart-wall">vCenter service will not start: vmon dependency order and the max-restart wall&lt;/h1>
&lt;p>When vCenter services refuse to start, the failure is rarely the service you see stuck STOPPED. vmon (the VMware service lifecycle manager) starts each child service in dependency order and supervises it with a per-service restart policy. If a low-level dependency like vPostgres or STS is down or slow, every service layered on top of it fails its health check, retries within vmon&amp;rsquo;s bounded budget, and then silently stops trying.&lt;/p></description></item><item><title>vCenter shows all hosts Disconnected / Not Responding: vCenter-side vs host-side</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-hosts-disconnected/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-hosts-disconnected/</guid><description>&lt;p>The vSphere Client shows many or all ESXi hosts as &amp;ldquo;Not Responding&amp;rdquo; or &amp;ldquo;Disconnected&amp;rdquo;. Do not start debugging individual hosts yet. In most mass-disconnect events, the hosts are fine and vCenter is the problem.&lt;/p>
&lt;p>The decisive signal is &lt;code>HostSystem.runtime.connectionState&lt;/code> across the inventory. When many hosts flip to &lt;code>notResponding&lt;/code> at the same time, the cause is almost always vCenter-side: vpxd overload, a VCSA network problem, a vpxd restart, an expired certificate breaking trust, or STS/auth intermittency. Individual host failures do not fan out across the inventory in seconds.&lt;/p></description></item><item><title>vCenter slow and unresponsive: vpxd overload and the task queue backlog</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-slow-unresponsive/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-slow-unresponsive/</guid><description>&lt;h1 id="vcenter-slow-and-unresponsive-vpxd-overload-and-the-task-queue-backlog">vCenter slow and unresponsive: vpxd overload and the task queue backlog&lt;/h1>
&lt;p>vCenter is sluggish. The vSphere Client hangs on login, PowerCLI calls time out, and tasks that normally finish in seconds sit in &lt;code>Running&lt;/code> for minutes. ESXi hosts start flipping to &lt;code>Not Responding&lt;/code> even though the hosts themselves are healthy and VMs keep running. This is the vpxd overload cascade, and it is one of the most commonly misdiagnosed vCenter incidents.&lt;/p></description></item><item><title>vCenter SSO intermittently failing: STS heap, flaky AD, and near-expiry certs</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-sso-degraded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcenter-sso-degraded/</guid><description>&lt;h1 id="vcenter-sso-intermittently-failing-sts-heap-flaky-ad-and-near-expiry-certs">vCenter SSO intermittently failing: STS heap, flaky AD, and near-expiry certs&lt;/h1>
&lt;p>Users report &amp;ldquo;vCenter logged me out randomly.&amp;rdquo; PowerCLI jobs fail with a token validation error, then succeed on the second attempt. One host shows &amp;ldquo;Not Responding&amp;rdquo; for 90 seconds, then comes back. Ten minutes later a different host does the same. The web client sometimes loads, sometimes returns a 503. STS itself sits at &amp;ldquo;Started/Green&amp;rdquo; in vmon the whole time.&lt;/p></description></item><item><title>vCenter vPostgres 'too many clients already': connection exhaustion</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vpostgres-connections-exhausted/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vpostgres-connections-exhausted/</guid><description>&lt;h1 id="vcenter-vpostgres-too-many-clients-already-connection-exhaustion">vCenter vPostgres &amp;rsquo;too many clients already&amp;rsquo;: connection exhaustion&lt;/h1>
&lt;p>vCenter operations time out. Power-on tasks hang in &amp;ldquo;Running&amp;rdquo; state. The vSphere Client is sluggish or fails to load. PowerCLI sessions return errors. Backup jobs fail mid-run. SSH into the VCSA and tail the vpxd or vPostgres logs, and you find the string that anchors the incident: &lt;code>FATAL: sorry, too many clients already&lt;/code>.&lt;/p>
&lt;p>The embedded vPostgres database is rejecting new connections because it has reached &lt;code>max_connections&lt;/code>. VMware tunes this value per deployment size, and a tuning script rewrites it based on the appliance&amp;rsquo;s memory allocation.&lt;!-- TODO: verify exact pg_tuning script name and whether it runs on every vmware-vpostgres restart or only at appliance boot/upgrade --> At the limit, anything needing a database connection fails immediately. This is a cliff-edge failure, not graceful degradation.&lt;/p></description></item><item><title>vCenter vPostgres autovacuum and WAL: dead tuples and an unbounded pg_wal</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vpostgres-vacuum-wal/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vpostgres-vacuum-wal/</guid><description>&lt;p>The embedded PostgreSQL database in vCenter Server Appliance (vPostgres) stores every managed object, event, task, alarm, and performance stat in the environment. Two internal mechanisms keep that database from consuming its own disk: autovacuum, which reclaims space from deleted and updated rows, and WAL (Write-Ahead Log) management, which controls how transaction logs are written, checkpointed, and recycled.&lt;/p>
&lt;p>When either mechanism falls behind, the symptoms are subtle at first. Queries slow down. vpxd task execution takes longer. The &lt;code>/storage/db&lt;/code> partition creeps upward. By the time an outside-in probe notices (vSphere Client timing out, vpxd crashing, services refusing to start), the database has been degraded for 30-60 minutes.&lt;/p></description></item><item><title>vCenter vpxd crash loop: the core service that keeps restarting</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vpxd-crash-loop/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vpxd-crash-loop/</guid><description>&lt;h1 id="vcenter-vpxd-crash-loop-the-core-service-that-keeps-restarting">vCenter vpxd crash loop: the core service that keeps restarting&lt;/h1>
&lt;p>vpxd is the C++ core of vCenter Server. It holds the entire managed inventory in memory, dispatches every management task to ESXi hosts via hostd, runs DRS, executes statistics rollups, and serves every SDK client (vSphere Client, PowerCLI, Veeam, NSX Manager, Aria Operations, custom automation). When vpxd dies, vCenter is functionally down: no provisioning, no vMotion orchestration, no DRS, no HA reconfiguration. VMs already running on hosts keep running, and FDM still restarts them after a host failure, because HA does not depend on vpxd.&lt;/p></description></item><item><title>Veeam Software SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/veeam-software-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/veeam-software-snmp-traps/</guid><description/></item><item><title>Velocloud Edge</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/velocloud-edge/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/velocloud-edge/</guid><description/></item><item><title>Vendor API 429 throttling: Meraki, Cato, and PAN-OS rate limits</title><link>https://www.netdata.cloud/guides/network/network-vendor-api-429-throttling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-vendor-api-429-throttling/</guid><description>&lt;h1 id="vendor-api-429-throttling-meraki-cato-and-pan-os-rate-limits">Vendor API 429 throttling: Meraki, Cato, and PAN-OS rate limits&lt;/h1>
&lt;p>Meraki, Cato, or PAN-OS API-polled devices go dark in your dashboard while ICMP and SNMP to the same devices return healthy responses. Every device sourced from the same vendor API flatlines at the same timestamp. The cause: your collector exhausted the vendor API rate-limit budget and is now receiving HTTP 429 instead of data.&lt;/p>
&lt;p>This pattern is frequently misdiagnosed because the symptom (devices appearing &amp;ldquo;down&amp;rdquo;) sits two layers above the cause (rate limit exhausted). It surfaces most often during incidents, when teams tighten polling intervals for faster data, or silently when multiple tools share a single API key without coordination.&lt;/p></description></item><item><title>Vendor API latency and pagination: monitoring pull-mode collection</title><link>https://www.netdata.cloud/guides/network/network-vendor-api-latency-pagination/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-vendor-api-latency-pagination/</guid><description>&lt;h1 id="vendor-api-latency-and-pagination-monitoring-pull-mode-collection">Vendor API latency and pagination: monitoring pull-mode collection&lt;/h1>
&lt;p>For SD-WAN controllers, cloud-managed networking, and modern firewall platforms, vendor API pull-mode collection is now the primary telemetry path. Operators depend on HTTPS calls to Meraki, Cato, PAN-OS, RESTCONF, and gRPC endpoints to learn tunnel state, license validity, session counts, and topology. Each API has its own authentication model, rate-limit budget, and pagination semantics. A collector that ignores these constraints will silently lose data, get throttled, or report healthy when the payload is empty.&lt;/p></description></item><item><title>Vendor API silent data gap: HTTP 200 with an empty payload</title><link>https://www.netdata.cloud/guides/network/network-vendor-api-silent-gap/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/network/network-vendor-api-silent-gap/</guid><description>&lt;h1 id="vendor-api-silent-data-gap-http-200-with-an-empty-payload">Vendor API silent data gap: HTTP 200 with an empty payload&lt;/h1>
&lt;p>Your SD-WAN controller dashboard shows flat lines. The Meraki organization API has not updated in twenty minutes. The PAN-OS firewall telemetry stopped at 03:00. Your collector logs show zero errors, every request returned HTTP 200, and no 5xx or timeout appears anywhere. But the data is gone.&lt;/p>
&lt;p>The API endpoint is reachable, the TCP connection succeeds, the HTTP status code says OK, and the response body is empty, null, or contains an error wrapped inside a success envelope. Your collector accepted the response as valid because it checked the status code and nothing else. Many API adapters treat a 200 with an empty payload as &amp;ldquo;no data to report&amp;rdquo; rather than &amp;ldquo;the API is broken.&amp;rdquo; Charts go flat, but no error fires. If the API is your only telemetry source for an SD-WAN overlay or a cloud-managed firewall estate, you are blind without knowing it.&lt;/p></description></item><item><title>Venturi Wireless SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/venturi-wireless-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/venturi-wireless-snmp-traps/</guid><description/></item><item><title>Verax Systems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/verax-systems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/verax-systems-snmp-traps/</guid><description/></item><item><title>Verifying rndc reload: catching a failed zone load before your clients do</title><link>https://www.netdata.cloud/guides/bind-dns/bind-dns-rndc-reload-verification/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/bind-dns/bind-dns-rndc-reload-verification/</guid><description>&lt;h1 id="verifying-rndc-reload-catching-a-failed-zone-load-before-your-clients-do">Verifying rndc reload: catching a failed zone load before your clients do&lt;/h1>
&lt;p>&lt;code>rndc reload&lt;/code> queues a zone reload and returns immediately with a success message. It does not confirm that the zone actually loaded. If the zone file has a syntax error, a missing include, or a permission problem, the error goes into the BIND log and nowhere else. The operator sees success. The zone is either stale (on reload failure, BIND continues serving the previous version) or answering SERVFAIL/REFUSED (on restart or initial load failure).&lt;/p></description></item><item><title>Verilink Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/verilink-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/verilink-corp-snmp-traps/</guid><description/></item><item><title>Veritas Software Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/veritas-software-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/veritas-software-corp-snmp-traps/</guid><description/></item><item><title>Veritas Technologies LLC SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/veritas-technologies-llc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/veritas-technologies-llc-snmp-traps/</guid><description/></item><item><title>VerneMQ</title><link>https://www.netdata.cloud/integrations/data-collection/databases/vernemq/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/vernemq/</guid><description/></item><item><title>VerneMQ Monitoring</title><link>https://www.netdata.cloud/monitoring-101/vernemq-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/vernemq-monitoring/</guid><description>&lt;h2 id="vernemq-monitoring">VerneMQ Monitoring&lt;/h2>
&lt;h3 id="what-is-vernemq">What Is VerneMQ?&lt;/h3>
&lt;p>&lt;a href="https://vernemq.com">VerneMQ&lt;/a> is a high-performance, distributed MQTT broker implemented in Erlang/OTP. It&amp;rsquo;s designed to handle large numbers of concurrent clients and is particularly well-suited for applications requiring low latency. Utilizing the full power of Erlang, VerneMQ offers a robust and scalable solution for message brokering.&lt;/p>
&lt;h3 id="monitoring-vernemq-with-netdata">Monitoring VerneMQ With Netdata&lt;/h3>
&lt;p>Netdata provides a comprehensive VerneMQ monitoring tool, which is part of its suite of monitoring solutions. By integrating VerneMQ monitoring, you can track important performance metrics in real-time, allowing you to proactively manage system health and diagnose potential issues before they become critical.&lt;/p></description></item><item><title>Versa Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/versa-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/versa-networks-inc-snmp-traps/</guid><description/></item><item><title>Vertica</title><link>https://www.netdata.cloud/integrations/data-collection/databases/vertica/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/vertica/</guid><description/></item><item><title>Vertica Monitoring</title><link>https://www.netdata.cloud/monitoring-101/vertica-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/vertica-monitoring/</guid><description>&lt;h2 id="vertica-monitoring">Vertica Monitoring&lt;/h2>
&lt;h3 id="what-is-vertica">What Is Vertica?&lt;/h3>
&lt;p>Vertica is an advanced analytics database platform designed to handle large volumes of data, providing fast query performance and real-time analytics. It&amp;rsquo;s typically used in data-intensive environments where speed and efficiency are crucial.&lt;/p>
&lt;h3 id="monitoring-vertica-with-netdata">Monitoring Vertica With Netdata&lt;/h3>
&lt;p>To monitor Vertica, Netdata leverages an openmetrics (Prometheus) exporter, specifically the &lt;a href="https://github.com/vertica/vertica-prometheus-exporter">vertica-prometheus-exporter&lt;/a>. Netdata can ingest data from any Prometheus exporter, enabling comprehensive monitoring without the need for a Prometheus server or Grafana. This approach offers users automated dashboards, alerts, and extensive insights, facilitating seamless database performance management.&lt;/p></description></item><item><title>Vertical Networks Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vertical-networks-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vertical-networks-inc-snmp-traps/</guid><description/></item><item><title>Vertiv Formerly Emerson Computer Power SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vertiv-formerly-emerson-computer-power-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vertiv-formerly-emerson-computer-power-snmp-traps/</guid><description/></item><item><title>Vertiv Formerly Emerson Energy Systems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vertiv-formerly-emerson-energy-systems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vertiv-formerly-emerson-energy-systems-snmp-traps/</guid><description/></item><item><title>Vertiv Formerly Geist Manufacturing Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vertiv-formerly-geist-manufacturing-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vertiv-formerly-geist-manufacturing-inc-snmp-traps/</guid><description/></item><item><title>Vertiv Liebert AC</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/vertiv-liebert-ac/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/vertiv-liebert-ac/</guid><description/></item><item><title>Vertiv Tech Co Ltd Formerly Emerson Network Power Co Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vertiv-tech-co-ltd-formerly-emerson-network-power-co-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vertiv-tech-co-ltd-formerly-emerson-network-power-co-ltd-snmp-traps/</guid><description/></item><item><title>Vertiv Watchdog</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/vertiv-watchdog/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/vertiv-watchdog/</guid><description/></item><item><title>Viavideo Communications Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/viavideo-communications-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/viavideo-communications-inc-snmp-traps/</guid><description/></item><item><title>VictoriaMetrics</title><link>https://www.netdata.cloud/integrations/exporters/victoriametrics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/victoriametrics/</guid><description/></item><item><title>Videoframe Systems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/videoframe-systems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/videoframe-systems-snmp-traps/</guid><description/></item><item><title>Videoserver Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/videoserver-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/videoserver-inc-snmp-traps/</guid><description/></item><item><title>Vigintos Elektronika SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vigintos-elektronika-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vigintos-elektronika-snmp-traps/</guid><description/></item><item><title>Viptela Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/viptela-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/viptela-inc-snmp-traps/</guid><description/></item><item><title>Virtual Machines</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/virtual-machines/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/virtual-machines/</guid><description/></item><item><title>vm.loadavg</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/vm.loadavg/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/vm.loadavg/</guid><description/></item><item><title>vm.stats.sys.v_intr</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/vm.stats.sys.v_intr/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/vm.stats.sys.v_intr/</guid><description/></item><item><title>vm.stats.sys.v_soft</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/vm.stats.sys.v_soft/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/vm.stats.sys.v_soft/</guid><description/></item><item><title>vm.stats.sys.v_swtch</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/vm.stats.sys.v_swtch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/vm.stats.sys.v_swtch/</guid><description/></item><item><title>vm.stats.vm.v_pgfaults</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/vm.stats.vm.v_pgfaults/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/vm.stats.vm.v_pgfaults/</guid><description/></item><item><title>vm.stats.vm.v_swappgs</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/vm.stats.vm.v_swappgs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/vm.stats.vm.v_swappgs/</guid><description/></item><item><title>vm.swap_info</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/vm.swap_info/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/vm.swap_info/</guid><description/></item><item><title>vm.vmtotal</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/vm.vmtotal/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/vm.vmtotal/</guid><description/></item><item><title>VMware Aria</title><link>https://www.netdata.cloud/integrations/exporters/vmware-aria/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/vmware-aria/</guid><description/></item><item><title>Vmware ESX</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/vmware-esx/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/vmware-esx/</guid><description/></item><item><title>Vmware Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vmware-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vmware-inc-snmp-traps/</guid><description/></item><item><title>VMware vCenter Server</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/vmware-vcenter-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/vmware-vcenter-server/</guid><description/></item><item><title>VMware vCenter Server Monitoring</title><link>https://www.netdata.cloud/monitoring-101/vsphere-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/vsphere-monitoring/</guid><description>&lt;h2 id="vmware-vcenter-server-monitoring">VMware vCenter Server Monitoring&lt;/h2>
&lt;h3 id="what-is-vmware-vcenter-server">What Is VMware vCenter Server?&lt;/h3>
&lt;p>VMware vCenter Server is a centralized platform for managing VMware vSphere environments. It allows administrators to automate and deliver a virtual infrastructure with confidence. vCenter Server provides essential vSphere and ESXi host management capabilities for IT teams.&lt;/p>
&lt;h3 id="monitoring-vmware-vcenter-server-with-netdata">Monitoring VMware vCenter Server With Netdata&lt;/h3>
&lt;p>Netdata offers a real-time monitoring solution for VMware vCenter Server with its robust &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/vsphere/?utm_source=website&amp;amp;utm_content=monitoring101">vSphere collector&lt;/a>. This powerful tool allows you to track host and virtual machine (VM) performance statistics, providing valuable insights into your virtual environments.&lt;/p></description></item><item><title>Vocaltec Communications Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vocaltec-communications-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vocaltec-communications-ltd-snmp-traps/</guid><description/></item><item><title>Voice Print International Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/voice-print-international-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/voice-print-international-inc-snmp-traps/</guid><description/></item><item><title>Volubill SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/volubill-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/volubill-snmp-traps/</guid><description/></item><item><title>Vpnet SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vpnet-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vpnet-snmp-traps/</guid><description/></item><item><title>VSCode</title><link>https://www.netdata.cloud/integrations/data-collection/applications/vscode/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/vscode/</guid><description/></item><item><title>VSCode Monitoring</title><link>https://www.netdata.cloud/monitoring-101/vscode-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/vscode-monitoring/</guid><description>&lt;h2 id="vscode-monitoring">VSCode Monitoring&lt;/h2>
&lt;h3 id="what-is-vscode">What Is VSCode?&lt;/h3>
&lt;p>Visual Studio Code (VSCode) is a highly popular source-code editor developed by Microsoft. It includes key features like support for debugging, syntax highlighting, intelligent code completion, snippets, and code refactoring. It&amp;rsquo;s vital for developers who require a streamlined and extensible environment to accelerate productivity and efficiency in coding tasks.&lt;/p>
&lt;h3 id="monitoring-vscode-with-netdata">Monitoring VSCode With Netdata&lt;/h3>
&lt;p>Monitoring VSCode is crucial for maintaining optimal performance and ensuring a seamless coding experience. Netdata facilitates this by utilizing an openmetrics (Prometheus) exporter to collect comprehensive metrics. Whether you are a developer or an IT professional, the flexibility of Netdata allows you to ingest data from any Prometheus exporter. Thus, you can access automated dashboards, real-time alerts, and comprehensive visualizations without needing to set up a standalone Prometheus server or Grafana. This streamlined monitoring solution increases efficiency and helps detect issues promptly in your development environment.&lt;/p></description></item><item><title>vSphere 'Virtual machine disks consolidation is needed': clearing the warning without stunning the VM</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-snapshot-consolidation-needed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-snapshot-consolidation-needed/</guid><description>&lt;h1 id="vsphere-virtual-machine-disks-consolidation-is-needed-clearing-the-warning-without-stunning-the-vm">vSphere &amp;lsquo;Virtual machine disks consolidation is needed&amp;rsquo;: clearing the warning without stunning the VM&lt;/h1>
&lt;p>The &amp;ldquo;Virtual machine disks consolidation is needed&amp;rdquo; warning means the VMkernel left delta VMDKs on the datastore after a snapshot delete that did not fully commit. The VM is still running, but its writes are going through delta files that were never meant to persist.&lt;/p>
&lt;p>The warning is set by &lt;code>VirtualMachine.runtime.consolidationNeeded&lt;/code> in the vCenter inventory. It is distinct from the Snapshot Manager view: a VM can have this flag set while showing zero snapshots in the manager, because the flag tracks orphaned files on the datastore, not the snapshot tree.&lt;/p></description></item><item><title>vSphere active vs consumed vs granted memory: why the percentage lies</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-active-vs-consumed-memory/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-active-vs-consumed-memory/</guid><description>&lt;h1 id="vsphere-active-vs-consumed-vs-granted-memory-why-the-percentage-lies">vSphere active vs consumed vs granted memory: why the percentage lies&lt;/h1>
&lt;p>The &amp;ldquo;Memory Usage&amp;rdquo; percentage on a vSphere host summary is one of the most misread signals in infrastructure monitoring. An 85% number that pages you at 3 a.m. may represent a healthy host with no reclamation at all. The same number on a different host may mean VMs are being actively swapped to disk. The percentage alone tells you nothing useful about either state.&lt;/p></description></item><item><title>vSphere CPU co-stop high (%CSTP): the SMP vCPU co-scheduling penalty</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-cpu-co-stop-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-cpu-co-stop-high/</guid><description>&lt;h1 id="vsphere-cpu-co-stop-high-cstp-the-smp-vcpu-co-scheduling-penalty">vSphere CPU co-stop high (%CSTP): the SMP vCPU co-scheduling penalty&lt;/h1>
&lt;p>&lt;code>%CSTP&lt;/code> in esxtop is the time a vCPU in a multi-vCPU VM sits halted because the ESXi scheduler is waiting to co-schedule the VM&amp;rsquo;s other vCPUs. In a healthy environment it is essentially zero. Sustained above a few percent on modern ESXi means a sizing or topology problem, not a performance problem you can tune away.&lt;/p>
&lt;p>The classic shape: you give a database 16 vCPUs and it gets slower. The guest OS reports low CPU utilization because the vCPUs are not doing work. They are parked in COSTOP waiting for their siblings. From inside the VM this is invisible. The application runs slowly while the OS reports idle capacity.&lt;/p></description></item><item><title>vSphere CPU limit hit (%MLMTD): the forgotten MHz cap that silently throttles a VM</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-cpu-limit-maxlimited/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-cpu-limit-maxlimited/</guid><description>&lt;h1 id="vsphere-cpu-limit-hit-mlmtd-the-forgotten-mhz-cap-that-silently-throttles-a-vm">vSphere CPU limit hit (%MLMTD): the forgotten MHz cap that silently throttles a VM&lt;/h1>
&lt;p>A VM is slow. The application team reports degraded throughput. You check the usual suspects: guest CPU utilization is high, but that is expected for a busy workload. Host CPU utilization is moderate, nowhere near saturated. %RDY, the standard vSphere CPU contention signal, is low. Everything looks healthy from the hypervisor&amp;rsquo;s perspective, yet the VM is underperforming.&lt;/p></description></item><item><title>vSphere CPU ready time high (%RDY): VMs starved while the guest looks idle</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-cpu-ready-time-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-cpu-ready-time-high/</guid><description>&lt;h1 id="vsphere-cpu-ready-time-high-rdy-vms-starved-while-the-guest-looks-idle">vSphere CPU ready time high (%RDY): VMs starved while the guest looks idle&lt;/h1>
&lt;p>A database VM takes twice as long to run its nightly batch. Application latency pings fire. You SSH into the guest, run &lt;code>top&lt;/code>, and CPU utilization sits at 25%. Memory is fine. Disk I/O looks normal. Nothing inside the VM explains the slowdown.&lt;/p>
&lt;p>This is the classic signature of CPU ready time in vSphere. The guest OS has no visibility into hypervisor scheduling decisions. When the ESXi CPU scheduler cannot find a free physical CPU for a runnable vCPU, the vCPU waits in the READY state. The guest never learns it was descheduled, so from inside the VM everything looks idle while the hypervisor sees a starved VM.&lt;/p></description></item><item><title>vSphere datastore full: 'No space left on device', paused VMs, and power-on failures</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-datastore-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-datastore-full/</guid><description>&lt;h1 id="vsphere-datastore-full-no-space-left-on-device-paused-vms-and-power-on-failures">vSphere datastore full: &amp;lsquo;No space left on device&amp;rsquo;, paused VMs, and power-on failures&lt;/h1>
&lt;p>A vSphere datastore hitting 100% is a cliff-edge failure. Below 100%, VM performance is unaffected. At 100%, every VM that needs to write to the datastore stops: running VMs pause with the &amp;ldquo;There is no more space for virtual disk&amp;rdquo; dialog, thin-provisioned VMDKs cannot extend, snapshot deltas cannot grow, and power-on operations fail because the per-VM &lt;code>.vswp&lt;/code> swap file cannot be created.&lt;/p></description></item><item><title>vSphere datastore IOPS and throughput: spotting storage saturation before latency bites</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-datastore-iops-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-datastore-iops-saturation/</guid><description>&lt;h1 id="vsphere-datastore-iops-and-throughput-spotting-storage-saturation-before-latency-bites">vSphere datastore IOPS and throughput: spotting storage saturation before latency bites&lt;/h1>
&lt;p>vSphere exposes IOPS and throughput counters at the datastore and virtual-disk level, but most teams only look at storage metrics after VMs are already slow. By the time DAVG or GAVG spike, the device queue is saturated and every VM on the datastore is paying for it. IOPS and throughput are leading indicators: they tell you how hard you are pushing the backend before the backend pushes back.&lt;/p></description></item><item><title>vSphere datastore latency high: reading GAVG, DAVG, and KAVG</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-datastore-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-datastore-latency-high/</guid><description>&lt;h1 id="vsphere-datastore-latency-high-reading-gavg-davg-and-kavg">vSphere datastore latency high: reading GAVG, DAVG, and KAVG&lt;/h1>
&lt;p>High datastore latency is the single most common cause of &amp;ldquo;everything is slow&amp;rdquo; in vSphere. Applications time out, guest iowait climbs, and in severe cases VMs lose heartbeats. Storage I/O traverses guest OS, virtual SCSI adapter, VMkernel SCSI stack, storage driver, fabric, and array. A single &amp;ldquo;latency is high&amp;rdquo; reading does not tell you where the time is going.&lt;/p>
&lt;p>Three counters slice that path into layers: GAVG is what the guest sees, DAVG is what the array reports, KAVG is what the VMkernel adds in between. The relationship among the three is the diagnostic.&lt;/p></description></item><item><title>vSphere dropped packets (%DRPRX/%DRPTX): ring buffers, CPU, and uplink backpressure</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-dropped-packets/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-dropped-packets/</guid><description>&lt;h1 id="vsphere-dropped-packets-drprxdrptx-ring-buffers-cpu-and-uplink-backpressure">vSphere dropped packets (%DRPRX/%DRPTX): ring buffers, CPU, and uplink backpressure&lt;/h1>
&lt;p>%DRPRX and %DRPTX in esxtop are usually the first sign a VM is losing packets inside the host. They should be zero at steady state. When they are not, the guest retransmits, latency climbs, and for latency-sensitive workloads (IP-based storage, databases, replicated queues) the impact can be severe well before the drop rate looks alarming.&lt;/p>
&lt;p>The distinction that matters: %DRPRX and %DRPTX count drops at the virtual switch port, between the vSwitch and the guest OS driver. They are not physical NIC drops. The uplink vmnic can report zero drops via &lt;code>esxcli network nic stats get&lt;/code> while %DRPRX is non-zero on the VM attached to it. Treating them as the same counter is the most common diagnostic mistake.&lt;/p></description></item><item><title>vSphere DRS not balancing: affinity rules and reservations blocking placement</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-drs-not-balancing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-drs-not-balancing/</guid><description>&lt;h1 id="vsphere-drs-not-balancing-affinity-rules-and-reservations-blocking-placement">vSphere DRS not balancing: affinity rules and reservations blocking placement&lt;/h1>
&lt;p>DRS is enabled and Fully Automated, yet one host sits at 90% CPU while another idles at 30%. Vmotions are not happening, the recommendations queue is empty or full of unapplied entries, and the DRS score is poor. In most cases DRS is doing exactly what it was told: it cannot find a migration that satisfies every constraint.&lt;/p>
&lt;p>DRS gates every placement decision behind a chain of compatibility and capacity checks before it compares hosts by load. A single VM-Host anti-affinity rule, a reservation that fully commits a host&amp;rsquo;s CPU or memory, or a missing vMotion network can each silently zero out the set of legal destinations. When every candidate fails one check, DRS emits no recommendation and the cluster drifts. Cluster averages hide this: a cluster &amp;ldquo;averaging 60% CPU&amp;rdquo; can mask one host at 90% and another at 30%. The per-host skew that DRS is supposed to fix is invisible in aggregate, and so is the constraint that prevents the fix.&lt;/p></description></item><item><title>vSphere DRS thrashing: vMotion churn with no stable placement</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-drs-thrashing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-drs-thrashing/</guid><description>&lt;h1 id="vsphere-drs-thrashing-vmotion-churn-with-no-stable-placement">vSphere DRS thrashing: vMotion churn with no stable placement&lt;/h1>
&lt;p>DRS thrashing is when the Distributed Resource Scheduler cannot find stable placement and keeps migrating the same VMs between hosts. Each migration costs vMotion bandwidth and inflicts a brief stun on the VM. The cluster looks balanced on paper, but the migrations never stop.&lt;/p>
&lt;p>This is almost always the DRS cost-benefit model oscillating because something underneath is unstable: bursty CPU or memory pressure, an aggressive migration threshold, conflicting affinity rules, or capacity too tight to ever be balanced.&lt;/p></description></item><item><title>vSphere ESXi hardware health: ECC errors, fan failure, and thermal throttling</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-host-hardware-health/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-host-hardware-health/</guid><description>&lt;h1 id="vsphere-esxi-hardware-health-ecc-errors-fan-failure-and-thermal-throttling">vSphere ESXi hardware health: ECC errors, fan failure, and thermal throttling&lt;/h1>
&lt;p>Hardware faults do not surface through hypervisor performance counters. They appear either as a Purple Screen of Death (PSOD) that kills every VM on the host instantly, or as slow degradation that looks like software misconfiguration until someone checks temperature sensors. Every DIMM error, failed fan, and degraded disk is the hypervisor&amp;rsquo;s problem.&lt;/p>
&lt;p>Three failure classes dominate production incidents on ESXi: ECC memory errors that progress from correctable to uncorrectable, fan failures that trigger thermal throttling and silently cap CPU frequency, and predictive disk failures that kick off RAID rebuilds consuming storage I/O. All three share a common operational trap: without the vendor CIM provider VIB installed, most sensors report &amp;ldquo;unknown&amp;rdquo; or are absent entirely, leaving you blind to the root cause.&lt;/p></description></item><item><title>vSphere forgotten snapshot growing: the delta VMDK time bomb</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-old-snapshot-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-old-snapshot-growing/</guid><description>&lt;h1 id="vsphere-forgotten-snapshot-growing-the-delta-vmdk-time-bomb">vSphere forgotten snapshot growing: the delta VMDK time bomb&lt;/h1>
&lt;p>A VM has a snapshot that should have been deleted days ago. The delta VMDK grows with every guest write. The datastore slowly fills. The guest has no idea anything is wrong. By the time someone notices, consolidation is a multi-hour, I/O-intensive operation that stuns the VM, and the datastore is hours from full. Forgotten snapshots are one of the most common preventable vSphere incidents.&lt;/p></description></item><item><title>vSphere HA 'Insufficient resources to satisfy configured failover level': admission control</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-ha-insufficient-resources/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-ha-insufficient-resources/</guid><description>&lt;h1 id="vsphere-ha-insufficient-resources-to-satisfy-configured-failover-level-admission-control">vSphere HA &amp;lsquo;Insufficient resources to satisfy configured failover level&amp;rsquo;: admission control&lt;/h1>
&lt;p>The error &amp;ldquo;Insufficient resources to satisfy configured failover level for vSphere HA&amp;rdquo; is admission control refusing a VM power-on, vMotion, or reservation change because granting it would leave the cluster without enough spare capacity to honor the configured HA failover policy. Admission control is doing its job: protecting the restart guarantee after a host failure.&lt;/p>
&lt;p>The cluster may physically hold more capacity than admission control lets you commit. A cluster with 500 GHz of CPU and 2 TB of RAM may only let you deploy against roughly 70% of that, with the rest held in reserve so HA can restart protected VMs after a host failure. Operators who bought hardware expecting to use all of it hit this wall during provisioning and reach for the disable switch.&lt;/p></description></item><item><title>vSphere HA host isolation and split-brain: when isolation response goes wrong</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-ha-host-isolation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-ha-host-isolation/</guid><description>&lt;h1 id="vsphere-ha-host-isolation-and-split-brain-when-isolation-response-goes-wrong">vSphere HA host isolation and split-brain: when isolation response goes wrong&lt;/h1>
&lt;p>vCenter shows one or more ESXi hosts as &amp;ldquo;Not Responding.&amp;rdquo; VMs on those hosts may have been restarted on surviving hosts by HA. Or they may still be running on the unreachable host with no way to manage them. In the worst case, the same VM is now running in two places at once, and nobody noticed until storage corruption or application errors surfaced.&lt;/p></description></item><item><title>vSphere host 'Not Responding': dead, isolated, or is hostd hung?</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-host-not-responding/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-host-not-responding/</guid><description>&lt;h1 id="vsphere-host-not-responding-dead-isolated-or-is-hostd-hung">vSphere host &amp;lsquo;Not Responding&amp;rsquo;: dead, isolated, or is hostd hung?&lt;/h1>
&lt;p>When vCenter shows an ESXi host as &amp;ldquo;Not Responding&amp;rdquo;, it has stopped receiving heartbeats from that host and cannot reach it on the management plane. This is distinct from &amp;ldquo;Disconnected&amp;rdquo; (a deliberate state set by an admin, or caused by license expiry) and from &amp;ldquo;Maintenance&amp;rdquo; (an intentional state for patching). The greyed-out host icon is one of the most paged-on vSphere symptoms because the same UI state covers at least four different underlying conditions.&lt;/p></description></item><item><title>vSphere host swapping (SWCUR/SWW/s): hypervisor swap and the memory death spiral</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-host-swapping/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-host-swapping/</guid><description>&lt;h1 id="vsphere-host-swapping-swcurswws-hypervisor-swap-and-the-memory-death-spiral">vSphere host swapping (SWCUR/SWW/s): hypervisor swap and the memory death spiral&lt;/h1>
&lt;p>When &lt;code>SWR/s&lt;/code> is sustained above zero on an ESXi host, the VMkernel is actively reading VM memory pages back from &lt;code>.vswp&lt;/code> files on the datastore. That is not a warning state. It is an active performance emergency. Every swapped-in page costs roughly 100x DRAM latency, and the swap I/O itself competes with VM disk I/O on the same datastore, producing a double penalty that degrades every VM on the host simultaneously.&lt;/p></description></item><item><title>vSphere memory ballooning (MCTLSZ): the host is reclaiming guest RAM</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-memory-ballooning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-memory-ballooning/</guid><description>&lt;h1 id="vsphere-memory-ballooning-mctlsz-the-host-is-reclaiming-guest-ram">vSphere memory ballooning (MCTLSZ): the host is reclaiming guest RAM&lt;/h1>
&lt;p>You open esxtop, switch to the memory view, and a VM&amp;rsquo;s MCTLSZ column is no longer zero. A few hundred megabytes or several gigabytes, the VMkernel has inflated the vmmemctl balloon driver inside that guest and is forcing the guest OS to hand back memory it thought it owned. From the host&amp;rsquo;s perspective this is gentle reclamation. From the guest&amp;rsquo;s and the application&amp;rsquo;s perspective, it is often the start of a silent performance decline.&lt;/p></description></item><item><title>vSphere memory compression: the reclamation tier between balloon and swap</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-memory-compression/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-memory-compression/</guid><description>&lt;h1 id="vsphere-memory-compression-the-reclamation-tier-between-balloon-and-swap">vSphere memory compression: the reclamation tier between balloon and swap&lt;/h1>
&lt;p>Memory compression is the third tier in ESXi&amp;rsquo;s four-tier memory reclamation hierarchy. It sits between ballooning, which asks the guest OS to return pages, and host-level swapping, which writes pages to .vswp files on a datastore and destroys latency. Compression is ESXi buying time: keep the page in RAM but in a smaller form, at the cost of CPU cycles and added access latency.&lt;/p></description></item><item><title>vSphere memory reclamation cascade: balloon to compress to swap in minutes</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-memory-pressure-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-memory-pressure-cascade/</guid><description>&lt;h1 id="vsphere-memory-reclamation-cascade-balloon-to-compress-to-swap-in-minutes">vSphere memory reclamation cascade: balloon to compress to swap in minutes&lt;/h1>
&lt;p>Hosts can sit at 75% consumed memory for hours with zero reclamation activity. A workload spike, a large VM power-on, or a memory leak can push consumed memory past physical capacity in minutes. ESXi&amp;rsquo;s response is not graceful: balloon inflates across VMs, compression activates, .vswp files open on the production datastore, swap-in begins, and datastore latency rises as swap I/O competes with VM disk I/O. By the time swap-in rate is above zero, the host has exhausted every gentler mechanism and is reading VM memory pages back from disk. That is a production emergency, not a tuning exercise.&lt;/p></description></item><item><title>vSphere monitoring checklist: the signals every host, VM, and vCenter needs</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-monitoring-checklist/</guid><description>&lt;h1 id="vsphere-monitoring-checklist-the-signals-every-host-vm-and-vcenter-needs">vSphere monitoring checklist: the signals every host, VM, and vCenter needs&lt;/h1>
&lt;p>Send this to someone standing up vSphere monitoring for the first time, or rebuilding an alerting setup that pages too often and misses real incidents. It lists the signals worth collecting across the hypervisor plane (ESXi hosts and VMs) and the management plane (vCenter Server Appliance).&lt;/p>
&lt;p>vSphere does not fail like a generic Linux box. CPU contention is invisible from inside the guest. Memory goes from fine to catastrophic in minutes once host swapping starts. A datastore at 99% full looks identical to one at 5% full from inside a VM, until every VM on it halts. And vCenter can degrade for weeks before anyone notices, because DRS, HA, and the API quietly keep working until they don&amp;rsquo;t. Generic CPU/disk/network dashboards miss most of this.&lt;/p></description></item><item><title>vSphere monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-monitoring-maturity-model/</guid><description>&lt;h1 id="vsphere-monitoring-maturity-model-from-survival-to-expert">vSphere monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>vSphere monitoring setups cluster into four maturity levels. Where you sit determines which incidents you catch early, which ones you discover after users complain, and which ones you never see coming until the management plane is already down. Use this as an assessment tool: read the signal lists for your current level and the one above, identify the gaps, and prioritize closing them.&lt;/p></description></item><item><title>vSphere NUMA locality low: wide VMs paying the remote-memory tax</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-numa-locality-low/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-numa-locality-low/</guid><description>&lt;h1 id="vsphere-numa-locality-low-wide-vms-paying-the-remote-memory-tax">vSphere NUMA locality low: wide VMs paying the remote-memory tax&lt;/h1>
&lt;p>Low NUMA locality is a silent performance killer in vSphere. Guest CPU utilization looks normal, memory usage is healthy, and disk and network latency are within baseline. But latency-sensitive workloads - databases, in-memory caches, analytics engines - run 10-30% slower than they should. The problem is below the guest, in the physical memory topology.&lt;/p>
&lt;p>When a VM spans multiple NUMA nodes (a &amp;ldquo;wide VM&amp;rdquo;), some memory accesses traverse the interconnect (QPI/UPI on Intel, Infinity Fabric on AMD) to reach memory owned by a different socket or chiplet. Each remote access costs roughly 1.5-2x the latency of a local access, adding approximately 50-100ns depending on platform and interconnect generation.&lt;/p></description></item><item><title>vSphere physical uplink saturation: one flow uses one NIC, and storage shares the wire</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-uplink-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-uplink-saturation/</guid><description>&lt;h1 id="vsphere-physical-uplink-saturation-one-flow-uses-one-nic-and-storage-shares-the-wire">vSphere physical uplink saturation: one flow uses one NIC, and storage shares the wire&lt;/h1>
&lt;p>A vSphere host with two 10GbE uplinks in a NIC team looks like 20Gbps. For aggregate multi-flow traffic across many VMs, it is. But a single TCP flow between one VM and one remote host uses exactly one of those uplinks. The team does not split that flow across both links. That is why a VM doing a bulk transfer can saturate one uplink while the other sits idle.&lt;/p></description></item><item><title>vSphere PSOD (purple screen of death): diagnosing an ESXi host crash</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-host-psod/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-host-psod/</guid><description>&lt;h1 id="vsphere-psod-purple-screen-of-death-diagnosing-an-esxi-host-crash">vSphere PSOD (purple screen of death): diagnosing an ESXi host crash&lt;/h1>
&lt;p>A purple screen of death (PSOD) is the ESXi VMkernel&amp;rsquo;s deliberate halt. When the kernel detects an unrecoverable condition, an uncorrectable machine check, a driver panic, or a corrupted data structure, it stops the host on purpose, paints the purple diagnostic screen, and writes a core dump if a target is configured. Every VM on that host dies instantly. There is no graceful shutdown and no live migration off the host.&lt;/p></description></item><item><title>vSphere SCSI sense codes and path failover: intermittent fabric faults in vmkernel.log</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-scsi-sense-path-failover/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-scsi-sense-path-failover/</guid><description>&lt;h1 id="vsphere-scsi-sense-codes-and-path-failover-intermittent-fabric-faults-in-vmkernellog">vSphere SCSI sense codes and path failover: intermittent fabric faults in vmkernel.log&lt;/h1>
&lt;p>When storage connectivity degrades, the first hard evidence usually lands in &lt;code>/var/log/vmkernel.log&lt;/code> as SCSI sense codes. Latency counters (DAVG, KAVG, GAVG) tell you I/O is slow; the sense codes tell you why. ESXi logs a status tuple and optional sense data on every command that fails or is retried, and the Native Multipathing Plugin (NMP) records the resulting path state transitions.&lt;/p></description></item><item><title>vSphere storage latency cliff: the 'everything is slow' incident that hits every VM at once</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-everything-is-slow-storage/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-everything-is-slow-storage/</guid><description>&lt;h1 id="vsphere-storage-latency-cliff-the-everything-is-slow-incident-that-hits-every-vm-at-once">vSphere storage latency cliff: the &amp;rsquo;everything is slow&amp;rsquo; incident that hits every VM at once&lt;/h1>
&lt;p>Every VM on a datastore degrades at the same moment. Application pages fire across unrelated services. The root cause is one shared resource below ESXi. This is the storage latency cliff.&lt;/p>
&lt;p>The pattern is distinctive once you have seen it. Latency on the affected datastore jumps from a few milliseconds to tens or hundreds of milliseconds. Queue depth climbs above zero. Every workload on that datastore suffers, regardless of which host it runs on. Guest operating systems report high I/O wait, but CPU and memory are fine. Applications report timeouts, but the application itself is healthy.&lt;/p></description></item><item><title>vSphere storage queue depth: QUED, ACTV, and DSNRO saturation</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-storage-queue-depth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-storage-queue-depth/</guid><description>&lt;h1 id="vsphere-storage-queue-depth-qued-actv-and-dsnro-saturation">vSphere storage queue depth: QUED, ACTV, and DSNRO saturation&lt;/h1>
&lt;p>Storage latency on an otherwise healthy all-flash array is usually a queue problem, not an array problem. The VMkernel stacks I/O requests in queues at every layer of the path between the guest and the physical device. When those queues fill, requests wait, and the latency the guest sees climbs independently of what the array is actually doing.&lt;/p>
&lt;p>The two numbers that tell you whether you are in this state are ACTV (in-flight I/O at the device) and QUED (I/O waiting in the VMkernel). Sustained QUED greater than zero with KAVG climbing is the signature of device queue saturation. The fix is rarely &amp;ldquo;make the array faster&amp;rdquo;; it is almost always &amp;ldquo;give the device more outstanding-I/O capacity, or stop funneling everything through a single path.&amp;rdquo;&lt;/p></description></item><item><title>vSphere thin-provisioned VMDK growth: space that never comes back without UNMAP</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-thin-provisioning-growth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-thin-provisioning-growth/</guid><description>&lt;h1 id="vsphere-thin-provisioned-vmdk-growth-space-that-never-comes-back-without-unmap">vSphere thin-provisioned VMDK growth: space that never comes back without UNMAP&lt;/h1>
&lt;p>A thin-provisioned VMDK starts small and grows as the guest writes data. It never shrinks on its own. When a guest deletes a 100 GB database dump, the space inside the guest filesystem becomes free, but the VMDK file on the datastore stays at its high-water mark. Over months, the datastore fills with blocks the guest considers empty. The symptom is a datastore that creeps toward full while the guests report plenty of free space inside.&lt;/p></description></item><item><title>vSphere Topology</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/vsphere-topology/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/topologies/vsphere-topology/</guid><description/></item><item><title>vSphere vCPU oversizing: why adding vCPUs made the VM slower</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcpu-oversizing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vcpu-oversizing/</guid><description>&lt;h1 id="vsphere-vcpu-oversizing-why-adding-vcpus-made-the-vm-slower">vSphere vCPU oversizing: why adding vCPUs made the VM slower&lt;/h1>
&lt;p>You gave the database VM 16 vCPUs because it was slow at 8. Now it is slower. Host CPU sits at 65%. Guest OS reports 20% utilization. Nothing in the usual dashboards explains the degradation.&lt;/p>
&lt;p>This is the vCPU oversizing spiral. Adding vCPUs to a VM that does not need them does not add capacity. It adds scheduling overhead. The ESXi CPU scheduler uses relaxed co-scheduling for multi-vCPU VMs, which means it must find enough simultaneously available physical CPUs to co-schedule a VM&amp;rsquo;s vCPUs together. The more vCPUs you give a VM, the harder that search becomes, and the longer each vCPU spends in the READY state before it can execute.&lt;/p></description></item><item><title>vSphere vMotion slow or failing: memory dirty rate, bandwidth, and convergence</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vmotion-slow-failing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vmotion-slow-failing/</guid><description>&lt;h1 id="vsphere-vmotion-slow-or-failing-memory-dirty-rate-bandwidth-and-convergence">vSphere vMotion slow or failing: memory dirty rate, bandwidth, and convergence&lt;/h1>
&lt;p>A vMotion task that should take minutes instead crawls for ten, twenty, or thirty minutes. The vSphere Client progress bar hangs in the high nineties. Sometimes the migration completes with a multi-second stun. Sometimes it fails outright: &amp;ldquo;The migration was cancelled because the amount of changing memory for the VM was greater than the available network bandwidth.&amp;rdquo;&lt;/p>
&lt;p>Almost every slow or failing vMotion traces back to one relationship: the VM is dirtying memory faster than the vMotion network can transmit it. The iterative pre-copy loop cannot converge. ESXi either stuns the vCPUs to force convergence, hurting guest performance, or cancels the migration.&lt;/p></description></item><item><title>vSphere vMotion stun time: the switchover pause that drops connections</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vmotion-stun-time/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vmotion-stun-time/</guid><description>&lt;h1 id="vsphere-vmotion-stun-time-the-switchover-pause-that-drops-connections">vSphere vMotion stun time: the switchover pause that drops connections&lt;/h1>
&lt;p>vMotion promises live migration with no disruption. For most VMs, the final switchover pause is short enough that in-guest applications and network clients never notice. But &amp;ldquo;live&amp;rdquo; is not &amp;ldquo;instantaneous.&amp;rdquo; During the final switchover, the VM is stunned: its vCPUs stop executing while the last set of dirty memory pages and device state transfers from the source host to the destination. When that pause stretches past a second, TCP stacks reset, databases miss heartbeats, and clustered applications fail over.&lt;/p></description></item><item><title>vSphere VMware Tools heartbeat red: guest crash vs extreme CPU starvation</title><link>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vm-heartbeat-red/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/vmware-vsphere/vmware-vsphere-vm-heartbeat-red/</guid><description>&lt;h1 id="vsphere-vmware-tools-heartbeat-red-guest-crash-vs-extreme-cpu-starvation">vSphere VMware Tools heartbeat red: guest crash vs extreme CPU starvation&lt;/h1>
&lt;p>A red &lt;code>guestHeartbeatStatus&lt;/code> on a powered-on vSphere VM means VMware Tools has stopped sending heartbeats to the VMkernel. The common assumption is a guest OS crash. The actual cause may be extreme CPU starvation, a hung guest, a crashed Tools process, or a transient vCenter API flap.&lt;/p>
&lt;p>The risk is the automation attached to the signal. With vSphere HA VM Monitoring enabled, a sustained red heartbeat plus no I/O causes HA to reset the VM. If the root cause is CPU starvation, the reset does not fix it: the VM returns, gets descheduled again, loses heartbeats again, and HA resets again. A misread cascades into restart loops.&lt;/p></description></item><item><title>Vutlan Sro Formerly Sky Control Sro SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vutlan-sro-formerly-sky-control-sro-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/vutlan-sro-formerly-sky-control-sro-snmp-traps/</guid><description/></item><item><title>Wago Kontakttechnik GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wago-kontakttechnik-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wago-kontakttechnik-gmbh-snmp-traps/</guid><description/></item><item><title>Warning and Critical Composite Temperature Time: past overheating that already did damage</title><link>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-composite-temperature-time/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/smartctl-disk-monitoring/smartctl-composite-temperature-time/</guid><description>&lt;h1 id="warning-and-critical-composite-temperature-time-past-overheating-that-already-did-damage">Warning and Critical Composite Temperature Time: past overheating that already did damage&lt;/h1>
&lt;p>Two fields in the NVMe SMART/Health Information Log (Log Page 02h) record cumulative thermal exposure that a point-in-time temperature reading cannot show: Warning Composite Temperature Time and Critical Composite Temperature Time. They count the minutes the drive has spent above vendor-defined warning and critical temperature thresholds over its entire lifetime.&lt;/p>
&lt;p>A non-zero value means the controller firmware accumulated a count. Even if Composite Temperature currently reads 40 degrees Celsius, these fields reveal thermal events that happened between polls, hours, days, or months ago. Thermal damage to NAND is cumulative and irreversible: sustained high temperature accelerates cell wear and degrades data retention. A drive that spent 200 minutes above its critical threshold may have aged more in those minutes than in months of normal operation. The question is not whether you can undo it (you cannot) but whether you can detect it and account for it in replacement planning.&lt;/p></description></item><item><title>Warp10</title><link>https://www.netdata.cloud/integrations/data-collection/databases/warp10/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/warp10/</guid><description/></item><item><title>Warp10 Monitoring</title><link>https://www.netdata.cloud/monitoring-101/warp10-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/warp10-monitoring/</guid><description>&lt;h2 id="warp10-monitoring">Warp10 Monitoring&lt;/h2>
&lt;h3 id="what-is-warp10">What Is Warp10?&lt;/h3>
&lt;p>Warp10 is a robust open-source platform tailored for managing time-series data. It allows users to collect, store, and analyze massive amounts of time-stamped information. With its unique approach to time-series management, Warp10 offers unparalleled performance and efficiency, making it an ideal choice for scenarios requiring fast data collection and real-time analytics.&lt;/p>
&lt;h3 id="monitoring-warp10-with-netdata">Monitoring Warp10 With Netdata&lt;/h3>
&lt;p>To monitor Warp10 effectively, Netdata leverages an openmetrics (Prometheus) exporter approach. This setup allows seamless integration, where Netdata collects data from any Prometheus exporter. Unlike traditional solutions requiring multiple software components, Netdata provides automated dashboards, real-time alerts, and comprehensive insights without needing a full-stack setup like Prometheus server or Grafana.&lt;/p></description></item><item><title>Watchguard</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/watchguard/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/watchguard/</guid><description/></item><item><title>Watchguard Technologies Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/watchguard-technologies-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/watchguard-technologies-inc-snmp-traps/</guid><description/></item><item><title>Wavefront</title><link>https://www.netdata.cloud/integrations/exporters/wavefront/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/exporters/wavefront/</guid><description/></item><item><title>Waystream AB Formerly Packetfront Network Products AB SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/waystream-ab-formerly-packetfront-network-products-ab-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/waystream-ab-formerly-packetfront-network-products-ab-snmp-traps/</guid><description/></item><item><title>Web server log files</title><link>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/web-server-log-files/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/web-servers-and-proxies/web-server-log-files/</guid><description/></item><item><title>Web server log files Monitoring</title><link>https://www.netdata.cloud/monitoring-101/weblog-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/weblog-monitoring/</guid><description>&lt;h2 id="web-server-log-files-monitoring">Web Server Log Files Monitoring&lt;/h2>
&lt;h3 id="what-is-web-server-log-files-monitoring">What Is Web Server Log Files Monitoring?&lt;/h3>
&lt;p>Monitoring web server log files is crucial for understanding the performance and health of a web server. Web logs contain detailed information about HTTP requests and responses, including metadata like client IP, request URI, response time, and status codes. Monitoring these files can help diagnose issues in real-time, optimize server performance, and improve user experience.&lt;/p>
&lt;h3 id="monitoring-web-server-log-files-with-netdata">Monitoring Web Server Log Files With Netdata&lt;/h3>
&lt;p>Netdata provides a comprehensive solution for web server log monitoring with its &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/web_log/">web_log&lt;/a> module. This $name monitoring tool automatically detects log files from popular web servers like Nginx and Apache and parses them to provide detailed metrics. With Netdata, you can easily monitor web server logs in real time, visualize data, and receive actionable insights.&lt;/p></description></item><item><title>Web Servers Log Monitoring</title><link>https://www.netdata.cloud/monitoring-101/webserverslog-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/webserverslog-monitoring/</guid><description>&lt;h2 id="why-monitor-web-servers">Why monitor Web servers?&lt;/h2>
&lt;p>Web servers are among the most important components in modern IT infrastructures. They host the websites, web services, and web applications that we use on a daily basis. Social networking, media streaming, software as a service (SaaS), and other activities wouldn’t be possible without the use of web servers. And with the advent of cloud computing and the movement of more services online, web servers and their monitoring are only becoming more important. Given the extensive usage of Web servers, Sysadmins and SREs should monitor web servers as a key aspect for performance.&lt;/p></description></item><item><title>Webhook</title><link>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/webhook/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/notifications/centralized-cloud-notifications/webhook/</guid><description/></item><item><title>Webinar: Actionable Network Device Monitoring With AI</title><link>https://www.netdata.cloud/webinars/network-device-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/webinars/network-device-monitoring/</guid><description/></item><item><title>Webinar: Cloud MCP Server – AI-Powered Monitoring</title><link>https://www.netdata.cloud/webinars/netdata-cloud-mcp-server/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/webinars/netdata-cloud-mcp-server/</guid><description/></item><item><title>Webinar: Kubernetes Throttling? It Doesn't Have to Suck!</title><link>https://www.netdata.cloud/webinars/kubernetes-cpu-throttling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/webinars/kubernetes-cpu-throttling/</guid><description/></item><item><title>Webinar: Live Database &amp; Network Diagnostics</title><link>https://www.netdata.cloud/webinars/live-functions-database-network-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/webinars/live-functions-database-network-monitoring/</guid><description/></item><item><title>Webinar: Maximize Uptime With Powerful Monitoring</title><link>https://www.netdata.cloud/webinars/maximize-uptime-monitoring-solutions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/webinars/maximize-uptime-monitoring-solutions/</guid><description/></item><item><title>Webinar: Netdata AI Now Talks to Your Other Tools to Find Root Cause Faster</title><link>https://www.netdata.cloud/webinars/netdata-ai-mcp-client-root-cause/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/webinars/netdata-ai-mcp-client-root-cause/</guid><description/></item><item><title>Webinar: Network Topology Maps and NetFlow Analysis</title><link>https://www.netdata.cloud/webinars/network-topology-maps-netflow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/webinars/network-topology-maps-netflow/</guid><description/></item><item><title>Webinar: OpenTelemetry Monitoring with Netdata</title><link>https://www.netdata.cloud/webinars/opentelemetry-monitoring-with-netdata/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/webinars/opentelemetry-monitoring-with-netdata/</guid><description/></item><item><title>Webinar: Real-Time Windows Server Monitoring</title><link>https://www.netdata.cloud/webinars/windows-server-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/webinars/windows-server-monitoring/</guid><description/></item><item><title>Webinar: The Runbook is Dead, the Agent is Alive</title><link>https://www.netdata.cloud/webinars/runbook-is-dead-agent-is-alive/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/webinars/runbook-is-dead-agent-is-alive/</guid><description/></item><item><title>Webscreen Technology Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/webscreen-technology-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/webscreen-technology-ltd-snmp-traps/</guid><description/></item><item><title>Websense Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/websense-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/websense-inc-snmp-traps/</guid><description/></item><item><title>Wellfleet SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wellfleet-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wellfleet-snmp-traps/</guid><description/></item><item><title>Westek Technology Ltd John Tucker SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/westek-technology-ltd-john-tucker-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/westek-technology-ltd-john-tucker-snmp-traps/</guid><description/></item><item><title>Westell Inc Formerly Kentrox SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/westell-inc-formerly-kentrox-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/westell-inc-formerly-kentrox-snmp-traps/</guid><description/></item><item><title>Westermo Teleindustri AB SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/westermo-teleindustri-ab-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/westermo-teleindustri-ab-snmp-traps/</guid><description/></item><item><title>Western Digital Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/western-digital-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/western-digital-corporation-snmp-traps/</guid><description/></item><item><title>Western Digital Mycloud EX2 Ultra</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/western-digital-mycloud-ex2-ultra/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/western-digital-mycloud-ex2-ultra/</guid><description/></item><item><title>Western Multiplex SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/western-multiplex-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/western-multiplex-snmp-traps/</guid><description/></item><item><title>Western Telematic Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/western-telematic-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/western-telematic-inc-snmp-traps/</guid><description/></item><item><title>Whois domain expiry Monitoring</title><link>https://www.netdata.cloud/monitoring-101/whois-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/whois-monitoring/</guid><description>&lt;h2 id="what-is-whois-domain-expiry">What is Whois domain expiry?&lt;/h2>
&lt;p>WHOIS is a query and response protocol that is widely used for querying databases that store the registered users or assignees of an Internet resource, such as a domain name, an IP address block or an autonomous system, but is also used for a wider range of other information. The protocol stores and delivers database content in a human-readable format.&lt;/p>
&lt;p>Among other things WHOIS can be used to query for domain expiry.&lt;/p></description></item><item><title>Why NVIDIA GPU utilization is misleading: 100% doesn't mean saturated</title><link>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-utilization-misleading/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/nvidia-gpu/nvidia-gpu-utilization-misleading/</guid><description>&lt;h1 id="why-nvidia-gpu-utilization-is-misleading-100-doesnt-mean-saturated">Why NVIDIA GPU utilization is misleading: 100% doesn&amp;rsquo;t mean saturated&lt;/h1>
&lt;p>A GPU dashboard pinned at 100% looks like proof that the hardware is fully used. It is not. During an incident, that distinction matters: a job can miss its throughput target while every utilization chart says the GPU is completely busy.&lt;/p>
&lt;p>NVIDIA&amp;rsquo;s common GPU utilization metric is a time-domain measurement. It reports the fraction of a sampling period during which at least one kernel was executing. It does not report how many streaming multiprocessors were active, how many warps were resident, whether tensor cores were used, or how much memory bandwidth was consumed.&lt;/p></description></item><item><title>Wiesemann Theis GmbH SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wiesemann-theis-gmbh-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wiesemann-theis-gmbh-snmp-traps/</guid><description/></item><item><title>Wind River Systems SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wind-river-systems-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wind-river-systems-snmp-traps/</guid><description/></item><item><title>Windows</title><link>https://www.netdata.cloud/integrations/deploy/operating-systems/windows/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/deploy/operating-systems/windows/</guid><description/></item><item><title>Windows Event Logs</title><link>https://www.netdata.cloud/integrations/logs/windows-event-logs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/logs/windows-event-logs/</guid><description/></item><item><title>Windows Monitoring</title><link>https://www.netdata.cloud/monitoring-101/windows-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/windows-monitoring/</guid><description>&lt;h2 id="_note-learn-more-about-netdatas-native-windows-agent-herehttpswwwnetdatacloudsolutionswindows-monitoring_">&lt;em>&lt;strong>Note: Learn more about Netdata&amp;rsquo;s &lt;a href="https://www.netdata.cloud/solutions/windows-monitoring/">Native Windows Agent here&lt;/a>.&lt;/strong>&lt;/em>&lt;/h2>
&lt;h2 id="effective-windows-server-monitoring">Effective Windows Server Monitoring&lt;/h2>
&lt;p>If you are a Windows System Administrator or developer you know how important it is to monitor your Windows Servers and make sure they&amp;rsquo;re up and running, smoothly.&lt;/p>
&lt;p>And you also know that sometimes things go south and your servers go kaput leaving you in the dark as to what really went wrong.&lt;/p>
&lt;ul>
&lt;li>Was it that rogue process that ate up all the CPU cycles?&lt;/li>
&lt;li>Did your server hit a memory bottleneck and start swapping like mad?&lt;/li>
&lt;li>Or maybe there was a disk error or a network glitch that you missed?&lt;/li>
&lt;/ul>
&lt;p>Effective Windows server monitoring requires the following:&lt;/p></description></item><item><title>Windows Network Protocols</title><link>https://www.netdata.cloud/integrations/data-collection/networking/windows-network-protocols/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/windows-network-protocols/</guid><description/></item><item><title>Windows Services</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/windows-services/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/windows-services/</guid><description/></item><item><title>Wired For Management SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wired-for-management-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wired-for-management-snmp-traps/</guid><description/></item><item><title>WireGuard</title><link>https://www.netdata.cloud/integrations/data-collection/networking/wireguard/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/wireguard/</guid><description/></item><item><title>WireGuard Monitoring</title><link>https://www.netdata.cloud/monitoring-101/wireguard-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/wireguard-monitoring/</guid><description>&lt;h2 id="wireguard-monitoring">WireGuard Monitoring&lt;/h2>
&lt;h3 id="what-is-wireguard">What Is WireGuard?&lt;/h3>
&lt;p>WireGuard is a high-performance VPN technology designed for ease of use and simple configuration while ensuring secure communications. Unlike traditional VPNs, WireGuard operates at the network layer and utilizes state-of-the-art cryptography. This makes it stand out as a modern and efficient solution for secure network connectivity.&lt;/p>
&lt;h3 id="monitoring-wireguard-with-netdata">Monitoring WireGuard With Netdata&lt;/h3>
&lt;p>Monitoring WireGuard with Netdata provides deep insights into the VPN device and peer traffic. Netdata&amp;rsquo;s comprehensive monitoring dashboard displays real-time performance metrics that are vital for maintaining optimal system and network operations. To monitor WireGuard effectively, &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/wireguard/">Netdata&amp;rsquo;s WireGuard integration&lt;/a> is an indispensable tool that automatically detects instances to ensure seamless data collection.&lt;/p></description></item><item><title>Wireless network interfaces</title><link>https://www.netdata.cloud/integrations/data-collection/networking/wireless-network-interfaces/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/networking/wireless-network-interfaces/</guid><description/></item><item><title>Wisi SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wisi-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wisi-snmp-traps/</guid><description/></item><item><title>Wordperfect Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wordperfect-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wordperfect-corp-snmp-traps/</guid><description/></item><item><title>World Wide Packets SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/world-wide-packets-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/world-wide-packets-snmp-traps/</guid><description/></item><item><title>Wuhan Research Institute Of Posts And Telecommunications SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wuhan-research-institute-of-posts-and-telecommunications-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wuhan-research-institute-of-posts-and-telecommunications-snmp-traps/</guid><description/></item><item><title>Wyse Technology SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wyse-technology-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/wyse-technology-snmp-traps/</guid><description/></item><item><title>X.509 certificate</title><link>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/x.509-certificate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/synthetic-testing/x.509-certificate/</guid><description/></item><item><title>X.509 Certificate Monitoring</title><link>https://www.netdata.cloud/monitoring-101/x509check-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/x509check-monitoring/</guid><description>&lt;h2 id="x509-certificate-monitoring">X.509 Certificate Monitoring&lt;/h2>
&lt;h3 id="what-is-an-x509-certificate">What Is an X.509 Certificate?&lt;/h3>
&lt;p>X.509 certificates are critical to internet security, setting the foundation for a secure online experience. These certificates verify identities through digital signatures, establishing secure communication channels via protocols like SSL/TLS.&lt;/p>
&lt;h3 id="monitoring-x509-certificates-with-netdata">Monitoring X.509 Certificates with Netdata&lt;/h3>
&lt;p>With Netdata, monitoring X.509 certificates becomes an intuitive and streamlined process. Netdata&amp;rsquo;s &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/x509check/">X.509 certificate monitoring tool&lt;/a> stands out by providing real-time insights into your certificates’ expiration times and revocation statuses.&lt;/p></description></item><item><title>X.509 certificates Monitoring</title><link>https://www.netdata.cloud/monitoring-101/x509-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/x509-monitoring/</guid><description>&lt;h2 id="what-are-x509-certificates">What are X.509 certificates?&lt;/h2>
&lt;p>X.509 is an International Telecommunication Union (ITU) standard defining the format of public key certificates. X.509 certificates are used in many Internet protocols, including TLS/SSL, which is the basis for HTTPS, the secure protocol for browsing the web. They are also used in offline applications, like electronic signatures.&lt;/p>
&lt;h2 id="monitoring-x509-certificates-with-netdata">Monitoring X.509 certificates with Netdata&lt;/h2>
&lt;p>The prerequisites for monitoring x509certificates with Netdata are to have x509certificates and &lt;a href="https://learn.netdata.cloud/docs/cloud/get-started/">Netdata installed&lt;/a> on your system.&lt;/p></description></item><item><title>Xedia Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/xedia-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/xedia-corporation-snmp-traps/</guid><description/></item><item><title>Xen XCP-ng</title><link>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/xen-xcp-ng/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/containers-and-vms/xen-xcp-ng/</guid><description/></item><item><title>Xerox SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/xerox-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/xerox-snmp-traps/</guid><description/></item><item><title>Xiaomi Mi Flora</title><link>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/xiaomi-mi-flora/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/hardware-and-sensors/xiaomi-mi-flora/</guid><description/></item><item><title>Xiaomi Mi Flora Monitoring</title><link>https://www.netdata.cloud/monitoring-101/xiaomi_mi_flora-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/xiaomi_mi_flora-monitoring/</guid><description>&lt;h2 id="xiaomi-mi-flora-monitoring">Xiaomi Mi Flora Monitoring&lt;/h2>
&lt;h3 id="what-is-xiaomi-mi-flora">What Is Xiaomi Mi Flora?&lt;/h3>
&lt;p>Xiaomi Mi Flora is an IoT device designed for plant care, providing insights into essential metrics such as soil moisture, temperature, and sunlight exposure. Ideal for horticulturists and plant enthusiasts, this nifty gadget ensures optimal plant growth and health by relaying real-time data to your chosen monitoring platform.&lt;/p>
&lt;h3 id="monitoring-xiaomi-mi-flora-with-netdata">Monitoring Xiaomi Mi Flora With Netdata&lt;/h3>
&lt;p>Netdata offers a seamless solution to monitor Xiaomi Mi Flora using an openmetrics (Prometheus) exporter. By leveraging an exporter specifically designed for Mi Flora, such as the &lt;a href="https://github.com/xperimental/flowercare-exporter">MiFlora/Flower Care Exporter&lt;/a>, Netdata ingests data effortlessly, eliminating the need for dedicated Prometheus servers or complex Grafana setups. With Netdata, users gain access to automated dashboards and instant alerts, optimizing their experience with Xiaomi Mi Flora monitoring tools.&lt;/p></description></item><item><title>Xiotech Corporation SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/xiotech-corporation-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/xiotech-corporation-snmp-traps/</guid><description/></item><item><title>Xircom SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/xircom-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/xircom-snmp-traps/</guid><description/></item><item><title>Xirrus Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/xirrus-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/xirrus-inc-snmp-traps/</guid><description/></item><item><title>Xylogics Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/xylogics-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/xylogics-inc-snmp-traps/</guid><description/></item><item><title>Xytronix Research Design Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/xytronix-research-design-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/xytronix-research-design-inc-snmp-traps/</guid><description/></item><item><title>YOURLS URL Shortener</title><link>https://www.netdata.cloud/integrations/data-collection/applications/yourls-url-shortener/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/yourls-url-shortener/</guid><description/></item><item><title>YOURLS URL Shortener Monitoring</title><link>https://www.netdata.cloud/monitoring-101/yourls-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/yourls-monitoring/</guid><description>&lt;h2 id="yourls-url-shortener-monitoring">YOURLS URL Shortener Monitoring&lt;/h2>
&lt;h3 id="what-is-yourls-url-shortener">What Is YOURLS URL Shortener?&lt;/h3>
&lt;p>YOURLS, or Your Own URL Shortener, is a self-hosted URL shortening service built on open-source principles. It provides the flexibility to create and manage custom short URLs, perfect for anyone wanting full control over their URL shortening service without relying on third-party platforms.&lt;/p>
&lt;h3 id="monitoring-yourls-with-netdata">Monitoring YOURLS With Netdata&lt;/h3>
&lt;p>Monitoring YOURLS is crucial to ensure its uptime, reliability, and performance. Netdata makes it easy to monitor YOURLS by utilizing an OpenMetrics (Prometheus) exporter, specifically the &lt;a href="https://github.com/just1not2/prometheus-exporter-yourls">YOURLS exporter&lt;/a>. With Netdata, you can ingest data from any Prometheus exporter, enabling automated dashboards and real-time alerts without the need for a standalone Prometheus server or Grafana setup. This seamless integration ensures you have the insights you need to keep YOURLS running optimally.&lt;/p></description></item><item><title>YugabyteDB</title><link>https://www.netdata.cloud/integrations/data-collection/databases/yugabytedb/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/databases/yugabytedb/</guid><description/></item><item><title>YugabyteDB Monitoring</title><link>https://www.netdata.cloud/monitoring-101/yugabytedb-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/yugabytedb-monitoring/</guid><description>&lt;h2 id="yugabytedb-monitoring">YugabyteDB Monitoring&lt;/h2>
&lt;h3 id="what-is-yugabytedb">What Is YugabyteDB?&lt;/h3>
&lt;p>&lt;a href="https://www.yugabyte.com/yugabytedb">YugabyteDB&lt;/a> is a high-performance, distributed database built for cloud-native applications. It offers the scalability and resilience of NoSQL databases while maintaining the ACID transactions and functionality of traditional relational databases. As a result, it&amp;rsquo;s a popular choice for organizations needing to scale rapidly without compromising on data integrity.&lt;/p>
&lt;h3 id="monitoring-yugabytedb-with-netdata">Monitoring YugabyteDB With Netdata&lt;/h3>
&lt;p>Netdata is a comprehensive real-time monitoring solution ideal for monitoring YugabyteDB. With its ability to provide a plethora of performance metrics out-of-the-box, you can seamlessly track the health and performance of your YugabyteDB instances. Whether you need to dig deep into database operations or monitor general server health, Netdata has you covered.&lt;/p></description></item><item><title>Zao Light Communication SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/zao-light-communication-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/zao-light-communication-snmp-traps/</guid><description/></item><item><title>Zebra Printer</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/zebra-printer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/zebra-printer/</guid><description/></item><item><title>Zenitel Norway AS SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/zenitel-norway-as-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/zenitel-norway-as-snmp-traps/</guid><description/></item><item><title>Zerto</title><link>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/zerto/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/cloud-and-devops/zerto/</guid><description/></item><item><title>Zerto Monitoring</title><link>https://www.netdata.cloud/monitoring-101/zerto-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/zerto-monitoring/</guid><description>&lt;h2 id="zerto-monitoring">Zerto Monitoring&lt;/h2>
&lt;h3 id="what-is-zerto">What Is Zerto?&lt;/h3>
&lt;p>Zerto is a disaster recovery and data protection platform designed to orchestrate backup and recovery processes across complex IT environments. It provides seamless integration for data backup, recovery, and efficient management of your IT infrastructure. By ensuring data integrity and continuity, businesses can minimize downtime and ensure smooth operations even in the face of unexpected disruptions.&lt;/p>
&lt;h3 id="monitoring-zerto-with-netdata">Monitoring Zerto With Netdata&lt;/h3>
&lt;p>When it comes to monitor Zerto performance and reliability, Netdata provides an unparalleled advantage with its openmetrics (Prometheus) exporter capabilities. By leveraging the &lt;a href="https://github.com/claranet/zerto-exporter">Zerto Exporter&lt;/a>, Netdata enables the seamless ingestion of data from any Prometheus exporter. Users benefit from automated dashboards that offer real-time insights, alerts, and more—all without the need for a Prometheus server or Grafana.&lt;/p></description></item><item><title>Zeus Technology Ltd SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/zeus-technology-ltd-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/zeus-technology-ltd-snmp-traps/</guid><description/></item><item><title>zfs</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/zfs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/zfs/</guid><description/></item><item><title>ZFS Adaptive Replacement Cache</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/zfs-adaptive-replacement-cache/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/zfs-adaptive-replacement-cache/</guid><description/></item><item><title>ZFS ARC and the OOM killer: applications killed while the cache will not shrink fast enough</title><link>https://www.netdata.cloud/guides/zfs/zfs-arc-oom-killer/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-arc-oom-killer/</guid><description>&lt;h1 id="zfs-arc-and-the-oom-killer-applications-killed-while-the-cache-will-not-shrink-fast-enough">ZFS ARC and the OOM killer: applications killed while the cache will not shrink fast enough&lt;/h1>
&lt;p>Your database or application process just got OOM-killed on a ZFS host that &amp;ldquo;had plenty of memory.&amp;rdquo; &lt;code>dmesg&lt;/code> shows the OOM killer invoked and a victim chosen, while the ARC was holding tens of gigabytes at the time. The machine was not out of memory. It was out of memory the kernel could get back fast enough.&lt;/p></description></item><item><title>ZFS ARC hit ratio low: cache misses, cold caches, and working sets that outgrew RAM</title><link>https://www.netdata.cloud/guides/zfs/zfs-arc-hit-ratio-low/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-arc-hit-ratio-low/</guid><description>&lt;h1 id="zfs-arc-hit-ratio-low-cache-misses-cold-caches-and-working-sets-that-outgrew-ram">ZFS ARC hit ratio low: cache misses, cold caches, and working sets that outgrew RAM&lt;/h1>
&lt;p>Your monitoring says the ZFS ARC hit ratio dropped, or reads got slow and someone traced it to cache misses. The number itself tells you almost nothing until you know three things: which hit ratio you are looking at (demand vs prefetch), what the workload&amp;rsquo;s baseline is, and whether the ARC is actually at its intended size.&lt;/p></description></item><item><title>ZFS ARC shrinking below c_max: reading memory pressure before latency hits</title><link>https://www.netdata.cloud/guides/zfs/zfs-arc-size-shrinking/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-arc-size-shrinking/</guid><description>&lt;h1 id="zfs-arc-shrinking-below-c_max-reading-memory-pressure-before-latency-hits">ZFS ARC shrinking below c_max: reading memory pressure before latency hits&lt;/h1>
&lt;p>You have a ZFS box that &amp;ldquo;got slow&amp;rdquo; and nobody can say why. The disks are fine, the pool is ONLINE, no scrub is running, and &lt;code>zpool iostat&lt;/code> shows nothing obviously wrong. Then you look at &lt;code>/proc/spl/kstat/zfs/arcstats&lt;/code> and the ARC is sitting at a fraction of its configured maximum, the target size &lt;code>c&lt;/code> keeps ratcheting downward, and &lt;code>arc_no_grow&lt;/code> is 1.&lt;/p></description></item><item><title>ZFS ARC using all memory: the Linux default that eats your RAM</title><link>https://www.netdata.cloud/guides/zfs/zfs-arc-using-all-memory/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-arc-using-all-memory/</guid><description>&lt;h1 id="zfs-arc-using-all-memory-the-linux-default-that-eats-your-ram">ZFS ARC using all memory: the Linux default that eats your RAM&lt;/h1>
&lt;p>You run &lt;code>free -h&lt;/code> on a ZFS host and see 58 of 64 GiB &amp;ldquo;used&amp;rdquo;, with almost nothing in buffers/cache. Your application processes account for maybe 8 GiB. Something is eating the machine, and you start hunting for a leak.&lt;/p>
&lt;p>There is no leak. The missing memory is the ZFS ARC (Adaptive Replacement Cache), ZFS&amp;rsquo;s primary read cache in kernel memory. On Linux the ARC lives outside the kernel page cache, so &lt;code>free&lt;/code> and &lt;code>top&lt;/code> report it as used slab, not as reclaimable cache. Operators who do not know this tune swappiness, add swap, or kill innocent processes to &amp;ldquo;free&amp;rdquo; memory that was never in danger.&lt;/p></description></item><item><title>ZFS cannot destroy dataset is busy: clones, holds, and mounted filesystems</title><link>https://www.netdata.cloud/guides/zfs/zfs-cannot-destroy-dataset-busy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-cannot-destroy-dataset-busy/</guid><description>&lt;h1 id="zfs-cannot-destroy-dataset-is-busy-clones-holds-and-mounted-filesystems">ZFS cannot destroy dataset is busy: clones, holds, and mounted filesystems&lt;/h1>
&lt;p>You run &lt;code>zfs destroy&lt;/code> on a dataset or snapshot and get one of two errors back:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-text" data-lang="text">&lt;span style="display:flex;">&lt;span>cannot destroy &amp;#39;tank/data&amp;#39;: dataset is busy
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>or:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-text" data-lang="text">&lt;span style="display:flex;">&lt;span>cannot destroy &amp;#39;tank/data@daily-2026-07-20&amp;#39;: snapshot has dependent clones
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Both mean the same thing at a mechanical level: something still references the object you are trying to remove, and ZFS refuses to guess which reference you are willing to lose. The frustrating part is that the error tells you almost nothing about what holds the reference. The job here is to find the dependent, understand why it exists, and remove it in the right order.&lt;/p></description></item><item><title>ZFS cannot import pool: missing devices, a damaged label, and root-pool boot failure</title><link>https://www.netdata.cloud/guides/zfs/zfs-cannot-import-pool/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-cannot-import-pool/</guid><description>&lt;h1 id="zfs-cannot-import-pool-missing-devices-a-damaged-label-and-root-pool-boot-failure">ZFS cannot import pool: missing devices, a damaged label, and root-pool boot failure&lt;/h1>
&lt;p>&lt;code>zpool import&lt;/code> fails and the pool refuses to come online. The error is one of a small set: &lt;code>cannot import 'tank': no such pool available&lt;/code>, &lt;code>one or more devices is currently unavailable&lt;/code>, &lt;code>cannot import 'tank': insufficient replicas&lt;/code>, or a complaint about an unsupported version or feature. Each message points at a different stage of the import process, and each has a different recovery path.&lt;/p></description></item><item><title>ZFS capacity cliff: why the pool falls off a performance edge near 80-90% full</title><link>https://www.netdata.cloud/guides/zfs/zfs-pool-capacity-cliff/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-pool-capacity-cliff/</guid><description>&lt;h1 id="zfs-capacity-cliff-why-the-pool-falls-off-a-performance-edge-near-80-90-full">ZFS capacity cliff: why the pool falls off a performance edge near 80-90% full&lt;/h1>
&lt;p>Your ZFS pool was fine last month. This week write latency is spiking, TXG syncs are stretching past their timeout, and applications are stalling on writes while reads still feel snappy. You check &lt;code>zpool list&lt;/code> and see CAP at 87%. Nothing failed. No disk died. The pool just crossed a threshold it was never going to tell you about.&lt;/p></description></item><item><title>ZFS capacity planning: runway estimation before the pool fills</title><link>https://www.netdata.cloud/guides/zfs/zfs-capacity-runway-planning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-capacity-runway-planning/</guid><description>&lt;h1 id="zfs-capacity-planning-runway-estimation-before-the-pool-fills">ZFS capacity planning: runway estimation before the pool fills&lt;/h1>
&lt;p>ZFS does not fail gracefully at full. Long before &lt;code>ENOSPC&lt;/code>, the metaslab allocator starts working harder to find free space, write latency climbs, TXG syncs stretch, and the pool slides into the capacity-fragmentation cliff described in &lt;a href="https://www.netdata.cloud/guides/zfs/zfs-how-it-works-in-production/">how ZFS actually works in production&lt;/a>. By the time applications see errors, you are already in emergency territory, and the recovery options (destroying snapshots under I/O pressure, expanding a pool mid-incident) are the worst versions of themselves.&lt;/p></description></item><item><title>ZFS checksum errors (CKSUM): the definitive signal of silent corruption</title><link>https://www.netdata.cloud/guides/zfs/zfs-checksum-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-checksum-errors/</guid><description>&lt;h1 id="zfs-checksum-errors-cksum-the-definitive-signal-of-silent-corruption">ZFS checksum errors (CKSUM): the definitive signal of silent corruption&lt;/h1>
&lt;p>A non-zero number in the CKSUM column of &lt;code>zpool status&lt;/code> means a block read from that device did not match the checksum ZFS stored for it. The data came back wrong, and ZFS can prove it. This is the only signal in your storage stack that definitively says &amp;ldquo;silent corruption happened here&amp;rdquo;, and zero is the only acceptable value in production.&lt;/p></description></item><item><title>ZFS checksum errors on multiple devices: suspect RAM or the controller, not the disks</title><link>https://www.netdata.cloud/guides/zfs/zfs-cksum-errors-multiple-devices/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-cksum-errors-multiple-devices/</guid><description>&lt;h1 id="zfs-checksum-errors-on-multiple-devices-suspect-ram-or-the-controller-not-the-disks">ZFS checksum errors on multiple devices: suspect RAM or the controller, not the disks&lt;/h1>
&lt;p>&lt;code>zpool status&lt;/code> shows non-zero CKSUM counters, and not on one disk. Two, four, or every device in the vdev has them. The instinct is to start RMA-ing drives. Stop. Independent disks do not fail in the same way at the same time. When checksum errors appear on multiple unrelated devices simultaneously, the failure is almost always upstream of the disks: something is corrupting data between the application and the platters, and every disk downstream of it is faithfully recording the damage.&lt;/p></description></item><item><title>ZFS deadman events: hung I/O and a stalled pool sync</title><link>https://www.netdata.cloud/guides/zfs/zfs-deadman-hung-io/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-deadman-hung-io/</guid><description>&lt;h1 id="zfs-deadman-events-hung-io-and-a-stalled-pool-sync">ZFS deadman events: hung I/O and a stalled pool sync&lt;/h1>
&lt;p>A ZFS deadman event means the kernel has watched an I/O operation sit incomplete for at least five minutes, or a pool sync sit incomplete for at least ten minutes, and has given up waiting quietly. The event shows up as &lt;code>FM_EREPORT_ZFS_DEADMAN&lt;/code> in &lt;code>zpool events&lt;/code>, and it is one of the few ZFS signals that justifies waking someone up immediately.&lt;/p></description></item><item><title>ZFS dedup memory exhaustion: when the DDT outgrows ARC and the pool crawls</title><link>https://www.netdata.cloud/guides/zfs/zfs-dedup-memory-exhaustion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-dedup-memory-exhaustion/</guid><description>&lt;h1 id="zfs-dedup-memory-exhaustion-when-the-ddt-outgrows-arc-and-the-pool-crawls">ZFS dedup memory exhaustion: when the DDT outgrows ARC and the pool crawls&lt;/h1>
&lt;p>The pool is ONLINE. &lt;code>zpool status -x&lt;/code> says everything is healthy. Disk latencies look mediocre but not dead. Yet every write takes tens to hundreds of milliseconds, reads that used to come from cache now hit disk, and application latency is uniformly awful in both directions. Nothing is DEGRADED, no scrub is running, and capacity looks unremarkable.&lt;/p></description></item><item><title>ZFS deleted files but no space freed: snapshots holding the blocks</title><link>https://www.netdata.cloud/guides/zfs/zfs-deleted-files-no-space-freed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-deleted-files-no-space-freed/</guid><description>&lt;h1 id="zfs-deleted-files-but-no-space-freed-snapshots-holding-the-blocks">ZFS deleted files but no space freed: snapshots holding the blocks&lt;/h1>
&lt;p>You deleted 100 GB of files. &lt;code>df&lt;/code> shows the same free space as before. &lt;code>zpool list&lt;/code> has not moved. Nothing is broken: ZFS is copy-on-write, and a snapshot taken before the delete still references every block you just removed. Until that snapshot is destroyed, the blocks stay allocated and the pool gains nothing.&lt;/p>
&lt;p>This becomes an incident when the pool is already past 90%: the operator deletes files, gets nothing back, panics, and starts destroying snapshots at random. Mass snapshot destruction on a nearly-full pool triggers heavy asynchronous block freeing that competes for the same I/O bandwidth the pool is already short on. That is how a capacity annoyance turns into a write stall.&lt;/p></description></item><item><title>ZFS device FAULTED - too many errors: a disk ejected from the pool</title><link>https://www.netdata.cloud/guides/zfs/zfs-vdev-faulted-too-many-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-vdev-faulted-too-many-errors/</guid><description>&lt;h1 id="zfs-device-faulted---too-many-errors-a-disk-ejected-from-the-pool">ZFS device FAULTED - too many errors: a disk ejected from the pool&lt;/h1>
&lt;p>You ran &lt;code>zpool status&lt;/code> and found a device FAULTED with &amp;ldquo;too many errors&amp;rdquo;, and the pool has dropped to DEGRADED. ZFS did this on purpose: the device&amp;rsquo;s READ, WRITE, or CKSUM error counters crossed a threshold, and ZFS took the disk out of service to stop it from corrupting or stalling the pool further.&lt;/p>
&lt;p>FAULTED is a verdict, not a glitch. ZFS only faults a device after repeated I/O or checksum failures, so the counters you see are the tail end of a problem that has been building. Your job is to find out whether the disk is dying or something between ZFS and the disk (cable, backplane, controller, power) is at fault, then replace or repair before the redundancy you have left disappears.&lt;/p></description></item><item><title>ZFS device UNAVAIL or REMOVED: a disk that fell off the bus</title><link>https://www.netdata.cloud/guides/zfs/zfs-device-unavail-removed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-device-unavail-removed/</guid><description>&lt;h1 id="zfs-device-unavail-or-removed-a-disk-that-fell-off-the-bus">ZFS device UNAVAIL or REMOVED: a disk that fell off the bus&lt;/h1>
&lt;p>You ran &lt;code>zpool status&lt;/code> and one of your devices is no longer ONLINE. It shows UNAVAIL or REMOVED, and the pool has flipped to DEGRADED or worse. This is not a ZFS software problem. Something between the kernel and the disk broke: the drive was pulled, the cable or backplane dropped it, the controller lost it, or a hot-plug event went badly.&lt;/p></description></item><item><title>ZFS dirty data throttling: the write delay that masquerades as slow disks</title><link>https://www.netdata.cloud/guides/zfs/zfs-dirty-data-throttling/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-dirty-data-throttling/</guid><description>&lt;h1 id="zfs-dirty-data-throttling-the-write-delay-that-masquerades-as-slow-disks">ZFS dirty data throttling: the write delay that masquerades as slow disks&lt;/h1>
&lt;p>Your applications report write latency spiking from microseconds to tens or hundreds of milliseconds. Throughput falls off a cliff under sustained write load. &lt;code>iostat&lt;/code> shows the disks are not busy. SMART is clean. The NVMe drives benchmark fine. Everything points at the storage, and nothing is wrong with the storage.&lt;/p>
&lt;p>This is the most misdiagnosed latency source in ZFS: the dirty data write throttle. When dirty (uncommitted) data in RAM crosses a threshold, ZFS deliberately injects artificial delay into every write syscall to slow writers down. If dirty data reaches the hard limit, writes stall completely until the syncing transaction group finishes. The system is working exactly as designed: it protects the pool from memory exhaustion by making applications wait.&lt;/p></description></item><item><title>ZFS encryption key not loaded: cannot mount an encrypted dataset</title><link>https://www.netdata.cloud/guides/zfs/zfs-encryption-key-not-loaded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-encryption-key-not-loaded/</guid><description>&lt;h1 id="zfs-encryption-key-not-loaded-cannot-mount-an-encrypted-dataset">ZFS encryption key not loaded: cannot mount an encrypted dataset&lt;/h1>
&lt;p>You rebooted a host, imported a pool, or tried to mount a dataset, and ZFS refused with a variation of &lt;code>cannot mount 'tank/secure': encryption key not loaded&lt;/code>. The pool itself is ONLINE, &lt;code>zpool status&lt;/code> is clean, and the data is intact. The dataset is simply locked: until the encryption key is loaded into the kernel, ZFS cannot decrypt the dataset&amp;rsquo;s metadata well enough to mount it.&lt;/p></description></item><item><title>ZFS I/O queue depth: telling backend saturation apart from a hang</title><link>https://www.netdata.cloud/guides/zfs/zfs-io-queue-depth-saturation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-io-queue-depth-saturation/</guid><description>&lt;h1 id="zfs-io-queue-depth-telling-backend-saturation-apart-from-a-hang">ZFS I/O queue depth: telling backend saturation apart from a hang&lt;/h1>
&lt;p>The pool is slow. Applications are timing out on writes, or reads that used to take milliseconds now take seconds. The first question that decides everything else is this: is the storage backend working hard and falling behind, or has something actually stopped moving? Both look identical from the application side. Both show up as &amp;ldquo;high latency.&amp;rdquo; The fix for one (more IOPS, better devices, workload shaping) is completely different from the fix for the other (a dead disk, a stuck controller, a hung I/O that needs intervention).&lt;/p></description></item><item><title>ZFS L2ARC ineffective: a cache that burns SSD endurance for nothing</title><link>https://www.netdata.cloud/guides/zfs/zfs-l2arc-ineffective/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-l2arc-ineffective/</guid><description>&lt;h1 id="zfs-l2arc-ineffective-a-cache-that-burns-ssd-endurance-for-nothing">ZFS L2ARC ineffective: a cache that burns SSD endurance for nothing&lt;/h1>
&lt;p>You added an SSD as an L2ARC device expecting faster reads. Months later, read latency has not moved, the SSD&amp;rsquo;s wear indicator is climbing, and the ARC is smaller than it should be. The L2ARC is being written to constantly and read from almost never. You are paying for the cache in RAM and drive endurance and getting nothing back.&lt;/p></description></item><item><title>ZFS monitoring checklist: the signals every production pool needs</title><link>https://www.netdata.cloud/guides/zfs/zfs-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-monitoring-checklist/</guid><description>&lt;h1 id="zfs-monitoring-checklist-the-signals-every-production-pool-needs">ZFS monitoring checklist: the signals every production pool needs&lt;/h1>
&lt;p>Most ZFS incidents are gaps in the basics: a pool DEGRADED for three weeks because nobody paged on it, a pool at 94% capacity discovered when writes started stalling, a disk accumulating checksum errors that a scrub would have caught months earlier. ZFS tells you almost everything you need to know, but only if you collect the right signals continuously instead of running &lt;code>zpool status&lt;/code> by hand after something breaks.&lt;/p></description></item><item><title>ZFS monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/zfs/zfs-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-monitoring-maturity-model/</guid><description>&lt;h1 id="zfs-monitoring-maturity-model-from-survival-to-expert">ZFS monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most ZFS incidents are not monitoring failures. They are monitoring-coverage failures. The pool was watched, but at the wrong level: the team had capacity graphs and no scrub result alerting, or pool state alerts and no TXG sync visibility, and the failure mode that actually fired lived exactly in the gap.&lt;/p>
&lt;p>This article is a reference model with four levels: Survival, Operational, Mature, and Expert. Each level adds signals that catch failure classes the previous level structurally cannot see. Use it two ways: as an audit of what you monitor today, and as a roadmap for what to add next. The levels are cumulative. Skipping a level buys you alert noise, not insight.&lt;/p></description></item><item><title>ZFS No space left on device: ENOSPC, the slop reserve, and the pool you cannot delete from</title><link>https://www.netdata.cloud/guides/zfs/zfs-no-space-left-on-device/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-no-space-left-on-device/</guid><description>&lt;h1 id="zfs-no-space-left-on-device-enospc-the-slop-reserve-and-the-pool-you-cannot-delete-from">ZFS No space left on device: ENOSPC, the slop reserve, and the pool you cannot delete from&lt;/h1>
&lt;p>Your application just failed with &lt;code>No space left on device&lt;/code>. &lt;code>df&lt;/code> shows free space. &lt;code>zpool list&lt;/code> shows FREE above zero. You try to delete files to make room, and &lt;code>rm&lt;/code> fails with the same error: &lt;code>rm: cannot remove 'file': No space left on device&lt;/code>. The pool is in the worst state a ZFS pool can be in: full enough that even freeing space requires space you do not have.&lt;/p></description></item><item><title>ZFS one slow disk in a vdev: the dying drive that drags the whole pool</title><link>https://www.netdata.cloud/guides/zfs/zfs-slow-disk-in-vdev/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-slow-disk-in-vdev/</guid><description>&lt;h1 id="zfs-one-slow-disk-in-a-vdev-the-dying-drive-that-drags-the-whole-pool">ZFS one slow disk in a vdev: the dying drive that drags the whole pool&lt;/h1>
&lt;p>The pool is ONLINE. &lt;code>zpool status -x&lt;/code> says all pools are healthy. Error counters are zero or near zero. And yet write latency has doubled, TXG syncs are stretching, and applications are complaining about storage. When one disk in a vdev starts dying, it usually does not fail cleanly. It develops bad sectors, and the drive firmware starts retrying reads and remapping blocks internally. Each retry adds milliseconds to individual I/Os. The disk never errors out, so ZFS never faults it.&lt;/p></description></item><item><title>ZFS periodic write latency spikes: the TXG sync storm pattern</title><link>https://www.netdata.cloud/guides/zfs/zfs-write-latency-spikes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-write-latency-spikes/</guid><description>&lt;h1 id="zfs-periodic-write-latency-spikes-the-txg-sync-storm-pattern">ZFS periodic write latency spikes: the TXG sync storm pattern&lt;/h1>
&lt;p>Your applications write fast for a few seconds, then every write stalls for 10 to 60 seconds, then everything is fast again. The stalls recur on a rough cycle, and between them the system looks completely healthy. Reads are fine. The pool is ONLINE. &lt;code>zpool status -x&lt;/code> says all pools are healthy. Disk-level tools show the devices mostly idle, except for periodic bursts of intense write activity.&lt;/p></description></item><item><title>ZFS permanent errors have been detected in the following files: recovering from data loss</title><link>https://www.netdata.cloud/guides/zfs/zfs-permanent-errors-detected/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-permanent-errors-detected/</guid><description>&lt;h1 id="zfs-permanent-errors-have-been-detected-in-the-following-files-recovering-from-data-loss">ZFS permanent errors have been detected in the following files: recovering from data loss&lt;/h1>
&lt;p>You ran &lt;code>zpool status -v&lt;/code> and at the bottom, under the errors section, you see:&lt;/p>
&lt;pre tabindex="0">&lt;code>errors: Permanent errors have been detected in the following files:

 tank/data@daily-2026-07-18:backups/db.dump
 tank/data/logs/app.log
&lt;/code>&lt;/pre>&lt;p>This is the most severe per-file integrity message ZFS produces. It means one or more data blocks failed checksum validation and ZFS could not repair them from redundancy. The data in those blocks is gone. This is not a warning, not a transient condition, and not something a reboot or &lt;code>zpool clear&lt;/code> will fix. Irrecoverable data loss has already occurred.&lt;/p></description></item><item><title>ZFS pool DEGRADED: redundancy lost and one failure from data loss</title><link>https://www.netdata.cloud/guides/zfs/zfs-pool-degraded-state/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-pool-degraded-state/</guid><description>&lt;h1 id="zfs-pool-degraded-redundancy-lost-and-one-failure-from-data-loss">ZFS pool DEGRADED: redundancy lost and one failure from data loss&lt;/h1>
&lt;p>Your monitoring fired, or a routine &lt;code>zpool status&lt;/code> run shows the pool in DEGRADED state. One or more devices have failed or gone unavailable, but the pool is still serving I/O because a mirror partner or RAIDZ parity is covering the gap. Applications see no errors.&lt;/p>
&lt;p>That is what makes DEGRADED dangerous. The pool is stable in this state and will run there indefinitely, so teams sit on it. But redundancy in the affected vdev group is gone or reduced: on a two-way mirror or RAIDZ1, the next failure in that group is unrecoverable data loss. On RAIDZ2 you have one fault of margin left, not two.&lt;/p></description></item><item><title>ZFS pool FAULTED: when the pool can no longer serve I/O</title><link>https://www.netdata.cloud/guides/zfs/zfs-pool-faulted/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-pool-faulted/</guid><description>&lt;h1 id="zfs-pool-faulted-when-the-pool-can-no-longer-serve-io">ZFS pool FAULTED: when the pool can no longer serve I/O&lt;/h1>
&lt;p>&lt;code>zpool status&lt;/code> shows the pool state as FAULTED, the vdev tree shows devices FAULTED or UNAVAIL, and the status message says &amp;ldquo;insufficient replicas for the pool to continue functioning.&amp;rdquo; Applications cannot read or write anything on the pool. Unlike DEGRADED, where redundancy is still covering for a failed device, FAULTED means ZFS has lost more devices than the vdev topology can tolerate, or cannot open enough devices to guarantee data integrity.&lt;/p></description></item><item><title>ZFS pool fragmentation high: write amplification you cannot defragment away</title><link>https://www.netdata.cloud/guides/zfs/zfs-pool-fragmentation-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-pool-fragmentation-high/</guid><description>&lt;h1 id="zfs-pool-fragmentation-high-write-amplification-you-cannot-defragment-away">ZFS pool fragmentation high: write amplification you cannot defragment away&lt;/h1>
&lt;p>You ran &lt;code>zpool list&lt;/code> and the FRAG column on your production pool reads 55%, 65%, maybe higher. Writes feel slower than they used to, scrub takes longer every month, and nobody can say when it started. Now you are searching for the ZFS equivalent of &lt;code>defrag&lt;/code> and discovering there is not one.&lt;/p>
&lt;p>The FRAG percentage measures how scattered your pool&amp;rsquo;s free space is across metaslab spacemaps. It is not file fragmentation, and it says nothing about how fragmented already-written data is. It tells you how hard the allocator will have to work for every future write. On a pool with high fragmentation, logically sequential writes land in scattered physical locations, turning sequential workloads into random I/O at the device layer.&lt;/p></description></item><item><title>ZFS pool history: the forensic log almost nobody watches</title><link>https://www.netdata.cloud/guides/zfs/zfs-pool-history-audit/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-pool-history-audit/</guid><description>&lt;h1 id="zfs-pool-history-the-forensic-log-almost-nobody-watches">ZFS pool history: the forensic log almost nobody watches&lt;/h1>
&lt;p>Every ZFS pool keeps a built-in audit log. Every property change, every snapshot creation and destruction, every import, export, key operation, and delegation change is recorded by &lt;code>zpool history&lt;/code>, with timestamps, and with the user and hostname if you ask for them. Almost nobody monitors it.&lt;/p>
&lt;p>This is a problem in two directions. Forensically, when something goes wrong (a dataset destroyed at 3 a.m., &lt;code>sync&lt;/code> suddenly disabled on a database dataset, a pool exported that nobody admits to exporting) the answers were in the history all along, often already rotated out by the time anyone looks. On the security side, mass snapshot destruction is exactly what ransomware or an attacker covering tracks looks like, and an unauthorised &lt;code>sync=disabled&lt;/code> or delegation change is a privilege escalation signal. Both are visible in one place, and that place is not wired into anything by default.&lt;/p></description></item><item><title>ZFS pool I/O is currently suspended: a hung pool and blocked I/O</title><link>https://www.netdata.cloud/guides/zfs/zfs-pool-suspended-io/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-pool-suspended-io/</guid><description>&lt;h1 id="zfs-pool-io-is-currently-suspended-a-hung-pool-and-blocked-io">ZFS pool I/O is currently suspended: a hung pool and blocked I/O&lt;/h1>
&lt;p>You ran a &lt;code>zpool&lt;/code> command, or an application tried to touch a dataset, and you got this back:&lt;/p>
&lt;pre tabindex="0">&lt;code>cannot open &amp;#39;tank&amp;#39;: pool I/O is currently suspended
&lt;/code>&lt;/pre>&lt;p>Or &lt;code>zpool status&lt;/code> shows the pool in state SUSPENDED with an action line telling you to reconnect devices and run &lt;code>zpool clear&lt;/code>. Every process touching the pool is stuck in uninterruptible sleep. New SSH sessions that touch the mountpoint hang. The system itself may still be responsive, but everything that depends on the pool is frozen.&lt;/p></description></item><item><title>ZFS pool I/O latency high: reading zpool iostat -l before blaming the disks</title><link>https://www.netdata.cloud/guides/zfs/zfs-pool-io-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-pool-io-latency-high/</guid><description>&lt;h1 id="zfs-pool-io-latency-high-reading-zpool-iostat--l-before-blaming-the-disks">ZFS pool I/O latency high: reading zpool iostat -l before blaming the disks&lt;/h1>
&lt;p>Applications are stalling on storage calls, &lt;code>zpool iostat -l&lt;/code> shows ugly latency numbers, and someone has already said &amp;ldquo;the disks are dying.&amp;rdquo; Before you open a hardware ticket, look at which latency column is actually elevated. ZFS reports total wait, disk wait, and queue wait separately, and the split between them tells you whether the problem is the devices or something inside ZFS itself.&lt;/p></description></item><item><title>ZFS pool ONLINE with non-zero errors: why zpool status -x lies</title><link>https://www.netdata.cloud/guides/zfs/zfs-online-with-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-online-with-errors/</guid><description>&lt;h1 id="zfs-pool-online-with-non-zero-errors-why-zpool-status--x-lies">ZFS pool ONLINE with non-zero errors: why zpool status -x lies&lt;/h1>
&lt;p>Your monitoring script runs &lt;code>zpool status -x&lt;/code>, gets back &amp;ldquo;all pools are healthy&amp;rdquo;, and moves on. Meanwhile, one disk in a mirror has 4,000 checksum errors, the pool is silently correcting every bad block from the good side, and the failing disk is weeks from dropping off the bus. You find out when the second disk in the vdev fails and the resilver uncovers corruption it can no longer reconstruct.&lt;/p></description></item><item><title>ZFS pool was previously in use from another system: MMP and multihost protection</title><link>https://www.netdata.cloud/guides/zfs/zfs-pool-in-use-another-system/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-pool-in-use-another-system/</guid><description>&lt;h1 id="zfs-pool-was-previously-in-use-from-another-system-mmp-and-multihost-protection">ZFS pool was previously in use from another system: MMP and multihost protection&lt;/h1>
&lt;p>You run &lt;code>zpool import tank&lt;/code> and get:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-text" data-lang="text">&lt;span style="display:flex;">&lt;span>cannot import &amp;#39;tank&amp;#39;: pool was previously in use from another system.
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>Last accessed by host2 (hostid=0x1a2b3c4d) at Wed Jul 22 03:12:44 2026
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>The pool can be imported, use &amp;#39;zpool import -f&amp;#39; to import the pool.
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Or, on a pool with multihost protection enabled, the harder version:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-text" data-lang="text">&lt;span style="display:flex;">&lt;span>cannot import &amp;#39;tank&amp;#39;: pool is imported on host2 (hostid: 0x1a2b3c4d)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Both messages exist to stop one specific disaster: two hosts importing the same pool at the same time. ZFS has no distributed locking. If two systems write to the same vdevs concurrently, they allocate the same free blocks, overwrite each other&amp;rsquo;s metadata, and destroy the pool. Your job is to determine whether the fence is protecting you from a live peer, or blocking you on stale state.&lt;/p></description></item><item><title>ZFS Pools</title><link>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/zfs-pools/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/storage-and-filesystems/zfs-pools/</guid><description/></item><item><title>ZFS Pools Monitoring</title><link>https://www.netdata.cloud/monitoring-101/zfspool-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/zfspool-monitoring/</guid><description>&lt;h2 id="zfs-pools-monitoring">ZFS Pools Monitoring&lt;/h2>
&lt;h3 id="what-is-zfs-pools">What Is ZFS Pools?&lt;/h3>
&lt;p>ZFS Pools are a high-performance, robust storage platform that combines a file system and logical volume manager designed to simplify data management and scaling. They are central to the ZFS ecosystem, enabling advanced data integrity, scalability, and reliability features.&lt;/p>
&lt;h3 id="monitoring-zfs-pools-with-netdata">Monitoring ZFS Pools With Netdata&lt;/h3>
&lt;p>To monitor ZFS Pools effectively, using a comprehensive tool like Netdata is crucial. Netdata provides real-time insights into the health and performance of ZFS Pools, including key metrics such as space utilization, fragmentation, and health states. By deploying &lt;a href="https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin/modules/zfspool/?utm_source=website&amp;amp;utm_content=monitoring101">Netdata&amp;rsquo;s ZFS Pools monitoring tool&lt;/a>, you gain access to detailed visualizations and alerts to help you proactively manage your ZFS systems.&lt;/p></description></item><item><title>ZFS READ and WRITE errors: transport-level device failures in zpool status</title><link>https://www.netdata.cloud/guides/zfs/zfs-read-write-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-read-write-errors/</guid><description>&lt;h1 id="zfs-read-and-write-errors-transport-level-device-failures-in-zpool-status">ZFS READ and WRITE errors: transport-level device failures in zpool status&lt;/h1>
&lt;p>You ran &lt;code>zpool status&lt;/code> and the READ or WRITE column on one or more devices is not zero. The pool may still show ONLINE. No application has complained yet. This is where ZFS lulls operators into inaction: redundancy is absorbing the failures, so nothing is visibly broken, but the counters are telling you a device, cable, controller, or power path is misbehaving.&lt;/p></description></item><item><title>ZFS replacing a failed disk: zpool replace, autoreplace, and hot spares</title><link>https://www.netdata.cloud/guides/zfs/zfs-replace-failed-disk/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-replace-failed-disk/</guid><description>&lt;h1 id="zfs-replacing-a-failed-disk-zpool-replace-autoreplace-and-hot-spares">ZFS replacing a failed disk: zpool replace, autoreplace, and hot spares&lt;/h1>
&lt;p>A disk in your pool has failed or is failing. The pool shows DEGRADED in &lt;code>zpool status&lt;/code>, which means redundancy is gone in that vdev group and the next failure in the same group is data loss. The pool keeps serving I/O, but you are on the clock.&lt;/p>
&lt;p>This runbook covers getting a replacement disk in and resilvered: identifying the physical device correctly, choosing between manual &lt;code>zpool replace&lt;/code>, &lt;code>autoreplace&lt;/code>, and hot spares, what sequential resilver changes, and the extra steps a root pool needs before the new disk can boot the machine.&lt;/p></description></item><item><title>ZFS resilver in progress: the reduced-redundancy window after a disk replace</title><link>https://www.netdata.cloud/guides/zfs/zfs-resilver-in-progress/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-resilver-in-progress/</guid><description>&lt;h1 id="zfs-resilver-in-progress-the-reduced-redundancy-window-after-a-disk-replace">ZFS resilver in progress: the reduced-redundancy window after a disk replace&lt;/h1>
&lt;p>You replaced a failed disk, &lt;code>zpool status&lt;/code> now shows &lt;code>scan: resilver in progress&lt;/code>, and the pool is DEGRADED. This is normal recovery, but it is not a safe state. Until the resilver completes, the affected vdev is running with reduced redundancy, and every hour in this window is an hour where one more failure can take the pool down.&lt;/p></description></item><item><title>ZFS resilver slow or stalled: multi-day rebuilds and the second-failure race</title><link>https://www.netdata.cloud/guides/zfs/zfs-resilver-slow-stuck/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-resilver-slow-stuck/</guid><description>&lt;h1 id="zfs-resilver-slow-or-stalled-multi-day-rebuilds-and-the-second-failure-race">ZFS resilver slow or stalled: multi-day rebuilds and the second-failure race&lt;/h1>
&lt;p>A disk was replaced, &lt;code>zpool status&lt;/code> shows &lt;code>resilver in progress&lt;/code>, and the ETA says three days. Or worse: the scanned byte count has not moved in twenty minutes. Either way, the pool is running with reduced redundancy, and every hour the resilver takes is an hour where the next disk failure becomes a data-loss event.&lt;/p>
&lt;p>Two facts frame everything below. First, ZFS deliberately throttles resilver I/O so production traffic wins. A slow resilver is often the system working as designed; the throttle-vs-risk trade-off is a decision you make, not an accident. Second, on large RAIDZ pools of spinning disks, a multi-day resilver is normal arithmetic, and a resilver that is decelerating frequently means a second device in the same vdev is also failing. That second case is the one that kills pools.&lt;/p></description></item><item><title>ZFS scrub not running: 'CKSUM 0' means nothing without regular scrubs</title><link>https://www.netdata.cloud/guides/zfs/zfs-scrub-not-running/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-scrub-not-running/</guid><description>&lt;h1 id="zfs-scrub-not-running-cksum-0-means-nothing-without-regular-scrubs">ZFS scrub not running: &amp;lsquo;CKSUM 0&amp;rsquo; means nothing without regular scrubs&lt;/h1>
&lt;p>The pool looks healthy. &lt;code>zpool status&lt;/code> shows every device ONLINE, and the READ, WRITE, and CKSUM columns are all zero. No scrub errors have ever been reported. So the data is safe, right?&lt;/p>
&lt;p>Not necessarily. Zero checksum errors means zero errors &lt;em>detected&lt;/em>, not zero corruption. ZFS only discovers silent corruption when it reads a block and verifies its checksum, and most blocks on a typical pool are read rarely or never. The scrub is the only mechanism that systematically reads every allocated block and checks it. A pool that has not scrubbed in six months has unknown integrity.&lt;/p></description></item><item><title>ZFS scrub repaired errors: correctable rot versus permanent data loss</title><link>https://www.netdata.cloud/guides/zfs/zfs-scrub-found-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-scrub-found-errors/</guid><description>&lt;h1 id="zfs-scrub-repaired-errors-correctable-rot-versus-permanent-data-loss">ZFS scrub repaired errors: correctable rot versus permanent data loss&lt;/h1>
&lt;p>A scrub finishes and the &lt;code>scan:&lt;/code> line in &lt;code>zpool status&lt;/code> reads something like &lt;code>scrub repaired 8.09M in 04:12:33 with 0 errors on Sun Jul 20 04:12:34 2026&lt;/code>. Two numbers in that line decide your next 24 hours: how much data ZFS had to fix, and whether anything was lost for good. Operators misread this line in both directions: panicking over a large repaired count that cost zero data, or shrugging at a small error count that means files are permanently damaged.&lt;/p></description></item><item><title>ZFS scrub versus resilver: why one preempts the other</title><link>https://www.netdata.cloud/guides/zfs/zfs-scrub-vs-resilver/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-scrub-vs-resilver/</guid><description>&lt;h1 id="zfs-scrub-versus-resilver-why-one-preempts-the-other">ZFS scrub versus resilver: why one preempts the other&lt;/h1>
&lt;p>You kicked off the weekly scrub, checked back an hour later, and &lt;code>zpool status&lt;/code> no longer says &amp;ldquo;scrub in progress&amp;rdquo;. It says &amp;ldquo;resilver in progress&amp;rdquo;, the pool is DEGRADED, and &lt;code>zpool scrub -s&lt;/code> refuses to cancel anything. Or the reverse: you replaced a disk, the resilver is running, and your monitoring keeps alerting that the pool has not scrubbed in 30 days.&lt;/p></description></item><item><title>ZFS silent data corruption: how bit rot happens and why scrubs catch it</title><link>https://www.netdata.cloud/guides/zfs/zfs-silent-data-corruption/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-silent-data-corruption/</guid><description>&lt;h1 id="zfs-silent-data-corruption-how-bit-rot-happens-and-why-scrubs-catch-it">ZFS silent data corruption: how bit rot happens and why scrubs catch it&lt;/h1>
&lt;p>A disk can return the wrong bytes without returning an error. The drive reports success, the controller reports success, and the application gets corrupt data. Conventional filesystems have no way to notice: they trust the storage stack, and at scale the storage stack lies often enough that corruption is a when, not an if.&lt;/p>
&lt;p>ZFS is designed around the assumption that every layer below it will eventually return wrong data. Every block carries a checksum, and every read verifies it. That is why ZFS operators see corruption events that ext4 or XFS operators never see: not because ZFS systems corrupt more, but because ZFS is the only layer looking.&lt;/p></description></item><item><title>ZFS SLOG device failed: sync latency 100x worse while the pool stays ONLINE</title><link>https://www.netdata.cloud/guides/zfs/zfs-slog-device-failed/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-slog-device-failed/</guid><description>&lt;h1 id="zfs-slog-device-failed-sync-latency-100x-worse-while-the-pool-stays-online">ZFS SLOG device failed: sync latency 100x worse while the pool stays ONLINE&lt;/h1>
&lt;p>NFS clients are timing out. Database commit latency has jumped from microseconds to tens or hundreds of milliseconds. &lt;code>zpool status -x&lt;/code> says &amp;ldquo;all pools are healthy.&amp;rdquo; Reads are fine, capacity is fine, scrubs are clean.&lt;/p>
&lt;p>This is the SLOG failure trap. A SLOG (separate intent log) failure does not degrade pool redundancy, so the pool state stays ONLINE. But the ZIL has silently fallen back to writing on the main pool vdevs, and every synchronous write now pays the full latency of your data disks instead of the fast log device. For sync-heavy workloads, that is a 10x to 100x latency regression with zero change in pool health state.&lt;/p></description></item><item><title>ZFS SLOG endurance: the SSD that wears out from concentrated sync writes</title><link>https://www.netdata.cloud/guides/zfs/zfs-slog-endurance-wearout/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-slog-endurance-wearout/</guid><description>&lt;h1 id="zfs-slog-endurance-the-ssd-that-wears-out-from-concentrated-sync-writes">ZFS SLOG endurance: the SSD that wears out from concentrated sync writes&lt;/h1>
&lt;p>A SLOG (Separate Intent Log) device is usually the smallest, fastest SSD in the box, and it is almost always the first one to die. Every synchronous write in the pool, from every dataset and every application, lands on this one device before the application gets its acknowledgment. That concentration is the point of a SLOG, and it is also why the device burns through write endurance far faster than the data vdevs around it.&lt;/p></description></item><item><title>ZFS slow pool import: space-map loading, ZIL replay, and import hangs</title><link>https://www.netdata.cloud/guides/zfs/zfs-slow-pool-import/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-slow-pool-import/</guid><description>&lt;h1 id="zfs-slow-pool-import-space-map-loading-zil-replay-and-import-hangs">ZFS slow pool import: space-map loading, ZIL replay, and import hangs&lt;/h1>
&lt;p>You ran &lt;code>zpool import tank&lt;/code> (or the system is booting root-on-ZFS) and nothing has happened for ten minutes. No output, no prompt, no error. The instinct is that the import has hung and to reach for a reboot or a forced import. Both are usually wrong.&lt;/p>
&lt;p>A slow import is usually ZFS replaying state it must replay before the pool is safe to use: loading space maps for metaslabs, replaying the ZIL after an unclean shutdown, and working through pending async destroy work left by deleted datasets or snapshots. On a large, fragmented, heavily snapshotted pool this can legitimately take minutes to, in extreme cases, an hour or more.&lt;/p></description></item><item><title>ZFS snapshot destroy slow: async destroy, the freeing property, and I/O contention</title><link>https://www.netdata.cloud/guides/zfs/zfs-snapshot-destroy-slow/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-snapshot-destroy-slow/</guid><description>&lt;h1 id="zfs-snapshot-destroy-slow-async-destroy-the-freeing-property-and-io-contention">ZFS snapshot destroy slow: async destroy, the freeing property, and I/O contention&lt;/h1>
&lt;p>You ran &lt;code>zfs destroy&lt;/code> on a large snapshot or a dataset full of snapshots. The command returned in seconds, but &lt;code>zpool list&lt;/code> shows the space did not come back, write latency is climbing, and applications are starting to complain. Or the destroy itself hung, and now every &lt;code>zfs&lt;/code> and &lt;code>zpool&lt;/code> command against that pool is stuck in D state.&lt;/p></description></item><item><title>ZFS snapshot space consumption: the 40%-of-the-pool blindspot</title><link>https://www.netdata.cloud/guides/zfs/zfs-snapshot-space-consumption/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-snapshot-space-consumption/</guid><description>&lt;h1 id="zfs-snapshot-space-consumption-the-40-of-the-pool-blindspot">ZFS snapshot space consumption: the 40%-of-the-pool blindspot&lt;/h1>
&lt;p>A ZFS pool reports 90% capacity. The operator deletes 500 GB of old files, and the pool does not get any emptier. This is one of the most common ZFS incidents, and it is not a bug. It is copy-on-write working exactly as designed.&lt;/p>
&lt;p>ZFS never overwrites blocks in place, so a snapshot retains references to every block that existed when it was taken. When you delete a file in the live dataset, the block is freed only if no snapshot still references it. If snapshots exist, the space stays allocated until the last referencing snapshot is destroyed. Teams routinely discover, mid-incident, that snapshots are holding 40% or more of the pool and that the default tooling never showed them.&lt;/p></description></item><item><title>ZFS special vdev full: the hidden bottleneck while the pool looks empty</title><link>https://www.netdata.cloud/guides/zfs/zfs-special-vdev-full/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-special-vdev-full/</guid><description>&lt;h1 id="zfs-special-vdev-full-the-hidden-bottleneck-while-the-pool-looks-empty">ZFS special vdev full: the hidden bottleneck while the pool looks empty&lt;/h1>
&lt;p>The pool is at 55% capacity, fragmentation is moderate, and &lt;code>zpool status -x&lt;/code> says all pools are healthy. Yet anything touching lots of small files or metadata has gone from fast to miserable, and no pool-level dashboard explains why.&lt;/p>
&lt;p>On pools with a special allocation class, the cause is often the special vdev itself: the (hopefully mirrored) fast SSDs that hold all pool metadata and, where &lt;code>special_small_blocks&lt;/code> is set, small data blocks. When the special vdev fills, ZFS raises no alert. New metadata and small blocks silently land on the main pool vdevs, so metadata I/O now runs at bulk-storage latency while every pool-level capacity signal still shows headroom. The pool is ONLINE, capacity looks fine, and the bottleneck sits one level down in the vdev tree where most monitoring never looks.&lt;/p></description></item><item><title>ZFS synchronous write latency high: fsync, NFS, and database commits stalling</title><link>https://www.netdata.cloud/guides/zfs/zfs-sync-write-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-sync-write-latency-high/</guid><description>&lt;h1 id="zfs-synchronous-write-latency-high-fsync-nfs-and-database-commits-stalling">ZFS synchronous write latency high: fsync, NFS, and database commits stalling&lt;/h1>
&lt;p>Your database commit latency just jumped from 2ms to 40ms. Or NFS clients are reporting sluggish writes while the pool itself looks healthy: &lt;code>zpool status -x&lt;/code> says all pools are healthy, read latency is fine, and throughput has not collapsed. Applications that call &lt;code>fsync()&lt;/code>, open files with &lt;code>O_SYNC&lt;/code>, or run over NFS are stalling, while everything else seems normal.&lt;/p></description></item><item><title>ZFS TXG sync time high: the most diagnostic write-path signal operators ignore</title><link>https://www.netdata.cloud/guides/zfs/zfs-txg-sync-time-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-txg-sync-time-high/</guid><description>&lt;h1 id="zfs-txg-sync-time-high-the-most-diagnostic-write-path-signal-operators-ignore">ZFS TXG sync time high: the most diagnostic write-path signal operators ignore&lt;/h1>
&lt;p>Applications are stalling on writes. &lt;code>zpool status -x&lt;/code> says all pools are healthy. Disk latency looks mostly fine, or at least inconsistent with the severity of the application impact. The signal that explains what is actually happening sits in one file most operators never open: &lt;code>/proc/spl/kstat/zfs/&amp;lt;pool&amp;gt;/txgs&lt;/code>, specifically the &lt;code>stime&lt;/code> field, which records how long each transaction group took to commit to stable storage.&lt;/p></description></item><item><title>ZFS zfs_arc_max: capping the ARC without starving read performance</title><link>https://www.netdata.cloud/guides/zfs/zfs-arc-max-tuning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-arc-max-tuning/</guid><description>&lt;h1 id="zfs-zfs_arc_max-capping-the-arc-without-starving-read-performance">ZFS zfs_arc_max: capping the ARC without starving read performance&lt;/h1>
&lt;p>On Linux, ZFS does not ship with a sane upper bound on the ARC for you. Left alone, the ARC aggressively consumes available memory, and because ARC memory is managed outside the kernel page cache, it shows up as &amp;ldquo;used&amp;rdquo; in &lt;code>free&lt;/code> even though it is reclaimable. On a shared host, the outcome is familiar: a database or application process gets OOM-killed, and nothing in the storage layer looks wrong.&lt;/p></description></item><item><title>ZFS ZIL commit stalls and errors: reading the intent-log kstats</title><link>https://www.netdata.cloud/guides/zfs/zfs-zil-commit-stalls/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zfs/zfs-zil-commit-stalls/</guid><description>&lt;h1 id="zfs-zil-commit-stalls-and-errors-reading-the-intent-log-kstats">ZFS ZIL commit stalls and errors: reading the intent-log kstats&lt;/h1>
&lt;p>A database that normally commits in single-digit milliseconds suddenly takes seconds per commit. NFS clients time out. &lt;code>zpool status&lt;/code> shows the pool ONLINE, read latency looks fine, and &lt;code>zpool status -x&lt;/code> says all pools are healthy. The breakage is confined to synchronous writes, which points at one subsystem: the ZFS Intent Log.&lt;/p>
&lt;p>The ZIL guarantees synchronous write semantics. Every &lt;code>fsync&lt;/code>, &lt;code>O_SYNC&lt;/code> write, and NFS commit goes to the ZIL before the application gets its acknowledgment. When the ZIL backs up, because the SLOG is slow or failed, the pool cannot absorb log writes, or the write pipeline is throttled, every synchronous writer stalls at once. Reads from the ARC keep working, which is why this failure mode is confusing on first contact.&lt;/p></description></item><item><title>Zhone Technologies Inc SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/zhone-technologies-inc-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/zhone-technologies-inc-snmp-traps/</guid><description/></item><item><title>Zhongxing Telecom Co Ltd Abbr Zte SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/zhongxing-telecom-co-ltd-abbr-zte-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/zhongxing-telecom-co-ltd-abbr-zte-snmp-traps/</guid><description/></item><item><title>Zoho Corporation Formerly Advent Network Management SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/zoho-corporation-formerly-advent-network-management-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/zoho-corporation-formerly-advent-network-management-snmp-traps/</guid><description/></item><item><title>ZooKeeper</title><link>https://www.netdata.cloud/integrations/data-collection/applications/zookeeper/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/applications/zookeeper/</guid><description/></item><item><title>ZooKeeper "Cannot open channel to N at election address": the blocked election port</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-cannot-open-channel-at-election-address/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-cannot-open-channel-at-election-address/</guid><description>&lt;h1 id="zookeeper-cannot-open-channel-to-n-at-election-address-the-blocked-election-port">ZooKeeper &amp;ldquo;Cannot open channel to N at election address&amp;rdquo;: the blocked election port&lt;/h1>
&lt;p>The log line is &lt;code>Cannot open channel to &amp;lt;id&amp;gt; at election address /host:3888&lt;/code>. It is emitted by &lt;code>QuorumCnxManager.connectOne()&lt;/code> when &lt;code>Socket.connect()&lt;/code> to a peer&amp;rsquo;s leader election port fails with &lt;code>ConnectException&lt;/code> (refused) or &lt;code>SocketTimeoutException&lt;/code> (timed out). The error is harmless during steady state and fatal during an election.&lt;/p>
&lt;p>ZooKeeper ensembles use two inter-server TCP ports. Port 2888 (the quorum port) carries the ZAB proposal/ACK/commit stream between followers and the active leader. Port 3888 (the leader election port) is touched only when &lt;code>FastLeaderElection&lt;/code> needs pairwise TCP channels to every voting peer. If 2888 is reachable but 3888 is not, the ensemble runs fine until the leader is lost, at which point no new leader can be elected.&lt;/p></description></item><item><title>ZooKeeper "Client session timed out, have not heard from server": the heartbeat miss</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-client-session-timed-out/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-client-session-timed-out/</guid><description>&lt;h1 id="zookeeper-client-session-timed-out-have-not-heard-from-server-the-heartbeat-miss">ZooKeeper &amp;ldquo;Client session timed out, have not heard from server&amp;rdquo;: the heartbeat miss&lt;/h1>
&lt;p>A &lt;code>Client session timed out, have not heard from server in &amp;lt;ms&amp;gt;ms for session id 0x..., closing socket connection and attempting reconnect&lt;/code> log line is the client&amp;rsquo;s &lt;code>SendThread&lt;/code> reporting that no PING response arrived inside its heartbeat window. This is a client-side symptom of a heartbeat miss. It is not the server expiring the session, and it is not yet &lt;code>SessionExpired&lt;/code>. The client closes its socket and tries another ensemble member.&lt;/p></description></item><item><title>ZooKeeper "Detected pause in JVM or host machine (eg GC)": the pause-monitor warning</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-detected-pause-in-jvm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-detected-pause-in-jvm/</guid><description>&lt;h1 id="zookeeper-detected-pause-in-jvm-or-host-machine-eg-gc-the-pause-monitor-warning">ZooKeeper &amp;ldquo;Detected pause in JVM or host machine (eg GC)&amp;rdquo;: the pause-monitor warning&lt;/h1>
&lt;p>The log line looks like this:&lt;/p>
&lt;pre>&lt;code>Detected pause in JVM or host machine (eg GC): pause of approximately 5234ms
&lt;/code>&lt;/pre>
&lt;p>ZooKeeper&amp;rsquo;s JvmPauseMonitor emits that line when the process froze longer than its configured threshold. The monitor thread sleeps for a fixed interval, wakes, and measures how long the sleep actually took. Anything beyond the expected sleep plus the warn threshold gets logged. If the JVM was not in a visible GC at that moment, the line ends with &amp;ldquo;No GCs detected&amp;rdquo;, which is the operator&amp;rsquo;s cue that something else on the host stole CPU.&lt;/p></description></item><item><title>ZooKeeper "fsync-ing the write ahead log took too long": the disk warning behind most write stalls</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-fsync-warning-adversely-affect-latency/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-fsync-warning-adversely-affect-latency/</guid><description>&lt;h1 id="zookeeper-fsync-ing-the-write-ahead-log-took-too-long-the-disk-warning-behind-most-write-stalls">ZooKeeper &amp;ldquo;fsync-ing the write ahead log took too long&amp;rdquo;: the disk warning behind most write stalls&lt;/h1>
&lt;p>The warning:&lt;/p>
&lt;pre tabindex="0">&lt;code>fsync-ing the write ahead log in SyncThread:0 took 1234ms which will adversely affect operation latency...
&lt;/code>&lt;/pre>&lt;p>fires when fsync on the transaction log exceeds &lt;code>fsync.warningthresholdms&lt;/code> (default 1000ms). The wording is deliberate: every write in ZooKeeper blocks on a quorum of fsyncs. If fsync takes a second, every write takes a second. If fsync takes 10 seconds, you are one missed heartbeat away from a leader election.&lt;/p></description></item><item><title>ZooKeeper "Packet len is out of range": jute.maxbuffer and oversized znodes</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-packet-len-out-of-range/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-packet-len-out-of-range/</guid><description>&lt;h1 id="zookeeper-packet-len-is-out-of-range-jutemaxbuffer-and-oversized-znodes">ZooKeeper &amp;ldquo;Packet len is out of range&amp;rdquo;: jute.maxbuffer and oversized znodes&lt;/h1>
&lt;p>&lt;code>Packet len &amp;lt;N&amp;gt; is out of range!&lt;/code> looks like a network framing problem. It is not. The ZooKeeper client is telling you the server&amp;rsquo;s serialized response exceeded the client&amp;rsquo;s maximum deserialization buffer, and the client closed the connection rather than read a truncated packet.&lt;/p>
&lt;p>On the server side, the same condition produces a terser log line: &lt;code>Len error&lt;/code>. That appears when a client attempts a write whose payload exceeds the server&amp;rsquo;s configured buffer limit. Both sides are governed by one Java system property: &lt;code>jute.maxbuffer&lt;/code>, which defaults to &lt;code>0xfffff&lt;/code> (1048575 bytes, just under 1 MB).&lt;/p></description></item><item><title>ZooKeeper "Too many connections from /IP - max is 60": maxClientCnxns rejecting clients</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-too-many-connections-max-is/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-too-many-connections-max-is/</guid><description>&lt;h1 id="zookeeper-too-many-connections-from-ip---max-is-60-maxclientcnxns-rejecting-clients">ZooKeeper &amp;ldquo;Too many connections from /IP - max is 60&amp;rdquo;: maxClientCnxns rejecting clients&lt;/h1>
&lt;p>&lt;code>WARN ... Error accepting new connection: Too many connections from /1.2.3.4 - max is 60&lt;/code> is ZooKeeper&amp;rsquo;s &lt;code>maxClientCnxns&lt;/code> limiter refusing a new TCP connection from a specific source IP. By the time it appears in the server log, the client has already been denied.&lt;/p>
&lt;p>The first trap: &lt;code>maxClientCnxns&lt;/code> is enforced per source IP, not as a total. A single ZooKeeper node can hold thousands of healthy sessions while still refusing every new connection from one IP. &lt;code>zk_num_alive_connections&lt;/code>, the metric most teams watch, is a total. It can look completely normal while clients behind a shared host IP are being silently turned away.&lt;/p></description></item><item><title>ZooKeeper "Unable to load database on disk": corrupt snapshot on startup</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-unable-to-load-database-on-disk/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-unable-to-load-database-on-disk/</guid><description>&lt;h1 id="zookeeper-unable-to-load-database-on-disk-corrupt-snapshot-on-startup">ZooKeeper &amp;ldquo;Unable to load database on disk&amp;rdquo;: corrupt snapshot on startup&lt;/h1>
&lt;p>The startup error &lt;code>Unable to load database on disk&lt;/code> from &lt;code>FileTxnSnapLog&lt;/code> means ZooKeeper cannot reconstruct its in-memory data tree from the on-disk snapshot and transaction log. The node refuses to join the ensemble and exits before serving traffic. You typically see this only on the next restart after the corruption happened, often days later.&lt;/p>
&lt;p>The failure is nasty because the running process looks fine until it does not. The corruption was already on disk; the restart made it impossible to ignore. A node that was serving requests an hour ago can refuse to come back after an unclean shutdown, an OOMKill, or a full &lt;code>dataDir&lt;/code> disk.&lt;/p></description></item><item><title>ZooKeeper "X is not executed because it is not in the whitelist": four-letter-word commands blocked</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-command-not-in-whitelist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-command-not-in-whitelist/</guid><description>&lt;h1 id="zookeeper-x-is-not-executed-because-it-is-not-in-the-whitelist-four-letter-word-commands-blocked">ZooKeeper &amp;ldquo;X is not executed because it is not in the whitelist&amp;rdquo;: four-letter-word commands blocked&lt;/h1>
&lt;p>The error string is exact. When you run &lt;code>echo mntr | nc localhost 2181&lt;/code> against a ZooKeeper 3.5.3+ server that has not been configured for it, the server replies:&lt;/p>
&lt;pre tabindex="0">&lt;code>mntr is not executed because it is not in the whitelist.
&lt;/code>&lt;/pre>&lt;p>Same shape for &lt;code>ruok&lt;/code>, &lt;code>isro&lt;/code>, &lt;code>stat&lt;/code>, &lt;code>conf&lt;/code>, &lt;code>envi&lt;/code>, &lt;code>cons&lt;/code>, &lt;code>wchs&lt;/code>, and the rest of the four-letter-word (4lw) command set. Only &lt;code>srvr&lt;/code> works out of the box, because the bundled &lt;code>zkServer.sh&lt;/code> status check depends on it.&lt;/p></description></item><item><title>ZooKeeper authentication failures: SASL/Digest auth_failed_count climbing</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-auth-failed-count/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-auth-failed-count/</guid><description>&lt;h1 id="zookeeper-authentication-failures-sasldigest-auth_failed_count-climbing">ZooKeeper authentication failures: SASL/Digest auth_failed_count climbing&lt;/h1>
&lt;p>The &lt;code>zk_auth_failed_count&lt;/code> counter exposed by ZooKeeper&amp;rsquo;s &lt;code>mntr&lt;/code> four-letter command increments every time a client fails authentication under the Digest or SASL schemes. In a stable, locked-down production ensemble this counter is effectively flat between restarts. When it moves, a client is connecting with credentials the server rejects.&lt;/p>
&lt;p>The metric is per-server and cumulative since process start. It does not break out by auth scheme, source IP, or principal, so the counter alone tells you something is wrong but not who or why. You resolve the &amp;ldquo;who and why&amp;rdquo; by reading the ZooKeeper log, the surrounding metrics (&lt;code>zk_ensemble_auth_fail&lt;/code> &lt;!-- TODO: verify exact mntr name; ZooKeeper sources often expose `zk_ensemble_auth_failures` -->, &lt;code>zk_connection_rejected&lt;/code> &lt;!-- TODO: verify exact name -->), and the deployment timeline.&lt;/p></description></item><item><title>ZooKeeper autopurge not configured: snapshots and logs filling the disk over months</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-autopurge-not-configured/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-autopurge-not-configured/</guid><description>&lt;h1 id="zookeeper-autopurge-not-configured-snapshots-and-logs-filling-the-disk-over-months">ZooKeeper autopurge not configured: snapshots and logs filling the disk over months&lt;/h1>
&lt;p>&lt;code>autopurge.purgeInterval&lt;/code> defaults to &lt;code>0&lt;/code>, meaning snapshots and transaction logs accumulate forever. On a quiet ensemble the growth is slow enough that nobody notices for months, then the &lt;code>dataLogDir&lt;/code> partition hits 100%, ZooKeeper cannot fsync the next write, and the process dies. The leader throws an &lt;code>IOException&lt;/code> on the transaction log and the ensemble loses a member, or quorum if more than one node fills simultaneously.&lt;/p></description></item><item><title>ZooKeeper avg_latency hides write stalls: why the headline number lies</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-avg-latency-hiding-write-stalls/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-avg-latency-hiding-write-stalls/</guid><description>&lt;h1 id="zookeeper-avg_latency-hides-write-stalls-why-the-headline-number-lies">ZooKeeper avg_latency hides write stalls: why the headline number lies&lt;/h1>
&lt;p>The dashboard says &lt;code>zk_avg_latency&lt;/code> is 1.2 ms. Clients are timing out on writes. Both can be true. On a read-heavy ZooKeeper ensemble, the headline latency number can look healthy while the write path is stalled.&lt;/p>
&lt;p>Two properties cause this. First, &lt;code>zk_avg_latency&lt;/code>, &lt;code>zk_min_latency&lt;/code>, and &lt;code>zk_max_latency&lt;/code> aggregate reads and writes into one number. Reads are served from local memory and complete in microseconds. Writes require a quorum round-trip plus a transaction log fsync before acknowledgment. When reads dominate the request mix, a severe write stall is diluted by thousands of cheap reads and disappears into the average.&lt;/p></description></item><item><title>ZooKeeper connection drops spiking: sessions dying in bursts</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-connection-drop-spike/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-connection-drop-spike/</guid><description>&lt;h1 id="zookeeper-connection-drops-spiking-sessions-dying-in-bursts">ZooKeeper connection drops spiking: sessions dying in bursts&lt;/h1>
&lt;p>A burst in &lt;code>zk_connection_drop_count&lt;/code> means connections to a ZooKeeper server are closing in a tight window, not one at a time. When the burst pushes &lt;code>zk_stale_sessions_expired&lt;/code> up simultaneously, you are looking at a session expiration storm in progress or one about to land on dependent services.&lt;/p>
&lt;p>Occasional single drops across a large fleet are background noise. A sustained drop rate above roughly 0.1% of total connections per minute is where the signal stops being normal churn. Bursts that fire on a rhythm (every few minutes, hourly, at the same minute past the hour) almost always point to a JVM garbage collection cycle or a scheduled job that briefly saturates the leader.&lt;/p></description></item><item><title>ZooKeeper data size growing: using ZooKeeper as a database is an anti-pattern</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-approximate-data-size-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-approximate-data-size-growing/</guid><description>&lt;h1 id="zookeeper-data-size-growing-using-zookeeper-as-a-database-is-an-anti-pattern">ZooKeeper data size growing: using ZooKeeper as a database is an anti-pattern&lt;/h1>
&lt;p>ZooKeeper is a distributed coordination service, not a datastore. It holds the entire znode tree in JVM heap, replicates every mutation through ZAB, and periodically serializes the full tree to a snapshot on disk. The design point is small, hot, strongly consistent metadata: leader election handles, service discovery registrations, distributed locks, configuration pointers. When teams treat it as a general-purpose key-value store and let &lt;code>zk_approximate_data_size&lt;/code> climb, they inherit the operational characteristics of an in-memory database without the tooling, schema, or compaction strategies of one.&lt;/p></description></item><item><title>ZooKeeper data tree digest mismatch: detecting corruption before it spreads</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-digest-mismatch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-digest-mismatch/</guid><description>&lt;h1 id="zookeeper-data-tree-digest-mismatch-detecting-corruption-before-it-spreads">ZooKeeper data tree digest mismatch: detecting corruption before it spreads&lt;/h1>
&lt;p>When &lt;code>zk_digest_mismatches_count&lt;/code> increments on a ZooKeeper node, the in-memory data tree on that node has diverged from the checksum ZooKeeper expects. This is a data-integrity alarm, not a performance signal. Clients reading from that node may be receiving wrong answers, and if the divergence came from a ZAB replication bug rather than local corruption, the same divergence may be propagating to other ensemble members.&lt;/p></description></item><item><title>ZooKeeper dataLogDir sharing a disk with snapshots: the #1 fsync-latency footgun</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-datalogdir-not-separated/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-datalogdir-not-separated/</guid><description>&lt;h1 id="zookeeper-datalogdir-sharing-a-disk-with-snapshots-the-1-fsync-latency-footgun">ZooKeeper dataLogDir sharing a disk with snapshots: the #1 fsync-latency footgun&lt;/h1>
&lt;p>You are chasing intermittent ZooKeeper write-latency spikes that appear to have no cause. Average latency is fine most of the time. Then, every few minutes, p99 update latency jumps by an order of magnitude, &lt;code>zk_outstanding_requests&lt;/code> briefly climbs, and clients on tight timeouts see a flicker of connection churn. By the time you SSH in, the cluster looks healthy again.&lt;/p></description></item><item><title>ZooKeeper follower doing a SNAP sync: full snapshot transfer and its blast radius</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-follower-snap-sync/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-follower-snap-sync/</guid><description>&lt;h1 id="zookeeper-follower-doing-a-snap-sync-full-snapshot-transfer-and-its-blast-radius">ZooKeeper follower doing a SNAP sync: full snapshot transfer and its blast radius&lt;/h1>
&lt;p>A follower that fell too far behind the leader does not catch up transaction by transaction. Once its last-seen zxid is older than the leader&amp;rsquo;s retained transaction log, the leader ships the entire data tree as a snapshot. This is a SNAP sync, the most expensive recovery path a healthy ensemble runs short of a leader election.&lt;/p></description></item><item><title>ZooKeeper follower sync time climbing: a follower approaching ejection</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-follower-sync-time-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-follower-sync-time-high/</guid><description>&lt;h1 id="zookeeper-follower-sync-time-climbing-a-follower-approaching-ejection">ZooKeeper follower sync time climbing: a follower approaching ejection&lt;/h1>
&lt;p>When &lt;code>zk_avg_follower_sync_time&lt;/code> or &lt;code>zk_max_follower_sync_time&lt;/code> starts climbing, a follower is taking longer to process proposals from the leader. The metric measures how close a follower is to being ejected from the quorum.&lt;/p>
&lt;p>The hard ceiling is &lt;code>syncLimit x tickTime&lt;/code>. With defaults of &lt;code>syncLimit=5&lt;/code> and &lt;code>tickTime=2000ms&lt;/code>, that ceiling is 10 seconds. When a follower&amp;rsquo;s sync time approaches that limit, the leader closes the connection, stops pushing updates, and the follower must re-enter leader discovery. In a 3-node ensemble, that drops you to minimum quorum with zero remaining fault tolerance.&lt;/p></description></item><item><title>ZooKeeper GC pause cascade: how a Stop-the-World freeze expires sessions and re-elects the leader</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-gc-pause-cascade/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-gc-pause-cascade/</guid><description>&lt;h1 id="zookeeper-gc-pause-cascade-how-a-stop-the-world-freeze-expires-sessions-and-re-elects-the-leader">ZooKeeper GC pause cascade: how a Stop-the-World freeze expires sessions and re-elects the leader&lt;/h1>
&lt;p>A ZooKeeper ensemble loses connections in bursts. Sessions expire en masse. The leader changes without a network cause. Request latency spikes on a rhythm that does not match disk I/O. The cause is usually a JVM Stop-the-World pause, and the fix is on the JVM, not the network.&lt;/p>
&lt;p>The mechanism: a JVM STW pause freezes the entire ZooKeeper process. No heartbeats go out. No requests are processed. No quorum ACKs flow. The TCP listener still accepts sockets, so the server looks alive from the outside, but the process does nothing with them. Clients miss their heartbeat window and their sessions expire. Followers miss the leader&amp;rsquo;s heartbeat and, if the pause runs past &lt;code>syncLimit * tickTime&lt;/code> (default 10 seconds with &lt;code>tickTime=2000&lt;/code> and &lt;code>syncLimit=5&lt;/code>), they declare the leader dead and start a new election. When the JVM resumes, the queued backlog drains as a latency spike and disconnected clients reconnect in a herd.&lt;/p></description></item><item><title>ZooKeeper heap usage climbing: catching the GC death spiral before it starts</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-heap-pressure-climbing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-heap-pressure-climbing/</guid><description>&lt;h1 id="zookeeper-heap-usage-climbing-catching-the-gc-death-spiral-before-it-starts">ZooKeeper heap usage climbing: catching the GC death spiral before it starts&lt;/h1>
&lt;p>ZooKeeper keeps its entire data tree on the JVM heap: every znode, its data bytes, ACLs, children lists, watch registrations, and session state. When heap climbs, it is the leading indicator for the OOM that eventually kills the ensemble. The signal that matters is not the sawtooth peaks from young-generation GC, but the rising post-GC trough that means the live set itself is growing.&lt;/p></description></item><item><title>ZooKeeper KeeperErrorCode = ConnectionLoss: the transient disconnect every client hits</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-connectionloss/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-connectionloss/</guid><description>&lt;h1 id="zookeeper-keepererrorcode--connectionloss-the-transient-disconnect-every-client-hits">ZooKeeper KeeperErrorCode = ConnectionLoss: the transient disconnect every client hits&lt;/h1>
&lt;p>&lt;code>KeeperErrorCode = ConnectionLoss&lt;/code> is the error every ZooKeeper client eventually logs. It means the TCP connection between the client and the server it was talking to broke before the operation&amp;rsquo;s response arrived. It does not mean the operation failed, and it does not mean the session is gone. The outcome of the in-flight operation is unknown, and the correct response is an idempotent retry.&lt;/p></description></item><item><title>ZooKeeper KeeperErrorCode = NoAuth: ACL denials on protected znodes</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-noauth-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-noauth-error/</guid><description>&lt;h1 id="zookeeper-keepererrorcode--noauth-acl-denials-on-protected-znodes">ZooKeeper KeeperErrorCode = NoAuth: ACL denials on protected znodes&lt;/h1>
&lt;p>&lt;code>KeeperErrorCode = NoAuth for /path&lt;/code> appears in client logs when a ZooKeeper operation is rejected because the calling session lacks the ACL permission required for that operation on that znode. The matching server-side line is &lt;code>Permission denied&lt;/code>. This is not a transient connectivity issue. The request reached a server, the server evaluated the znode&amp;rsquo;s ACL, and the session did not match.&lt;/p></description></item><item><title>ZooKeeper KeeperErrorCode = NodeExists: create failing on an already-created znode</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-nodeexists-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-nodeexists-error/</guid><description>&lt;h1 id="zookeeper-keepererrorcode--nodeexists-create-failing-on-an-already-created-znode">ZooKeeper KeeperErrorCode = NodeExists: create failing on an already-created znode&lt;/h1>
&lt;p>The error:&lt;/p>
&lt;pre tabindex="0">&lt;code>org.apache.zookeeper.KeeperException$NodeExistsException: KeeperErrorCode = NodeExists for /some/path
&lt;/code>&lt;/pre>&lt;p>A &lt;code>create()&lt;/code> hit a znode path that already exists. ZooKeeper classifies this as a state exception, not a system fault. The cluster refused to overwrite an existing node, exactly as specified. The question is whether the caller expected that path to be free.&lt;/p>
&lt;p>Most production hits follow one of two patterns. The first is a retry after &lt;code>ConnectionLoss&lt;/code>: the original &lt;code>create()&lt;/code> committed but the response was lost in flight. The client does not know whether the operation succeeded, retries, and receives &lt;code>NodeExists&lt;/code>. The second is a genuine race: two candidates trying to create the same ephemeral leader or lock node, where exactly one wins and the other should lose gracefully. Both are expected. The incident starts when a client mishandles the exception, or when an orphaned ephemeral blocks the legitimate owner indefinitely.&lt;/p></description></item><item><title>ZooKeeper KeeperErrorCode = NoNode: operating on a path that doesn't exist</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-nonode-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-nonode-error/</guid><description>&lt;h1 id="zookeeper-keepererrorcode--nonode-operating-on-a-path-that-doesnt-exist">ZooKeeper KeeperErrorCode = NoNode: operating on a path that doesn&amp;rsquo;t exist&lt;/h1>
&lt;p>A client logs &lt;code>KeeperErrorCode = NoNode for /some/path&lt;/code>. The server returned &lt;code>Code.NONODE&lt;/code> (integer -101), which the Java client surfaces as &lt;code>KeeperException.NoNodeException&lt;/code>. The failed operation was a &lt;code>getData&lt;/code>, &lt;code>getChildren&lt;/code>, &lt;code>exists&lt;/code>, &lt;code>setData&lt;/code>, &lt;code>delete&lt;/code>, or &lt;code>create&lt;/code> against a znode that is not currently in the data tree.&lt;/p>
&lt;p>&lt;code>NoNode&lt;/code> is not a server fault. It is the API contract enforced correctly: ZooKeeper refuses to operate on a missing path. The operator&amp;rsquo;s job is to find out why the path is missing. Three cases cover almost every incident: the path was never created (usually a missing parent), the path was deleted by the server (an ephemeral tied to an expired session, or a container/TTL node auto-cleaned), or a deploy changed the znode layout clients expect.&lt;/p></description></item><item><title>ZooKeeper KeeperErrorCode = Session expired: ephemeral nodes gone, clients evicted</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-session-expired/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-session-expired/</guid><description>&lt;h1 id="zookeeper-keepererrorcode--session-expired-ephemeral-nodes-gone-clients-evicted">ZooKeeper KeeperErrorCode = Session expired: ephemeral nodes gone, clients evicted&lt;/h1>
&lt;p>The exact error clients log is &lt;code>KeeperErrorCode = Session expired&lt;/code>. On the ZooKeeper side you see &lt;code>Expiring session 0x... timeout of Nms exceeded&lt;/code>. Once that line lands, the client&amp;rsquo;s ZooKeeper handle is dead and every piece of state it owned through that session is gone: ephemeral znodes deleted, watches invalidated, ACLs no longer enforceable. The client cannot reconnect on the same handle. It must build a new ZooKeeper object, negotiate a new session, and recreate every ephemeral node it relied on.&lt;/p></description></item><item><title>ZooKeeper leader election storm: an ensemble that keeps re-electing</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-leader-election-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-leader-election-storm/</guid><description>&lt;h1 id="zookeeper-leader-election-storm-an-ensemble-that-keeps-re-electing">ZooKeeper leader election storm: an ensemble that keeps re-electing&lt;/h1>
&lt;p>&lt;code>zk_looking_count&lt;/code> is incrementing on multiple nodes. The ZooKeeper log fills with repeated &lt;code>LEADING&lt;/code> and &lt;code>FOLLOWING&lt;/code> transitions. Every few seconds or minutes, the ensemble elects a new leader, and that leader quickly loses quorum. Writes are intermittent, and downstream systems that depend on ZooKeeper for coordination (Kafka controllers, HBase region assignment, distributed locks) experience cascading failures.&lt;/p>
&lt;p>This is a leader election storm. Unlike a single failover where remaining nodes elect a stable replacement, here every potential leader hits the same wall. The root cause is shared across ensemble members: fsync stalls on overloaded storage, GC pauses exceeding the quorum timeout, or intermittent network failures between server pairs. Because every candidate experiences the same problem, no leader holds the role long enough for the cluster to recover.&lt;/p></description></item><item><title>ZooKeeper Monitoring</title><link>https://www.netdata.cloud/monitoring-101/zookeeper-monitoring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/monitoring-101/zookeeper-monitoring/</guid><description>&lt;h2 id="zookeeper-monitoring">ZooKeeper Monitoring&lt;/h2>
&lt;h3 id="what-is-zookeeper">What Is ZooKeeper?&lt;/h3>
&lt;p>ZooKeeper is a high-performance coordination service for distributed applications, developed as a project of the Apache Software Foundation. It provides a reliable, centralized service for maintaining configuration information, naming, providing distributed synchronization, and providing group services. Whether you’re using it for service discovery or to build resilient distributed locks, understanding its operations can ensure high availability and performance of your applications.&lt;/p>
&lt;h3 id="monitoring-zookeeper-with-netdata">Monitoring ZooKeeper With Netdata&lt;/h3>
&lt;p>Monitoring ZooKeeper effectively is crucial for ensuring the health and performance of your distributed systems. The Netdata agent makes this process straightforward by automatically detecting ZooKeeper instances running on known TCP sockets such as &lt;code>127.0.0.1:2181&lt;/code>. To get started, you&amp;rsquo;ll need to &lt;a href="https://zookeeper.apache.org/doc/current/zookeeperAdmin.html#sc_4lw">add &lt;code>mntr&lt;/code> to ZooKeeper&amp;rsquo;s 4lw.commands.whitelist&lt;/a>. Once set up, Netdata provides real-time monitoring capabilities, making it the ideal ZooKeeper monitoring tool.&lt;/p></description></item><item><title>ZooKeeper monitoring checklist: the signals every production ensemble needs</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-monitoring-checklist/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-monitoring-checklist/</guid><description>&lt;h1 id="zookeeper-monitoring-checklist-the-signals-every-production-ensemble-needs">ZooKeeper monitoring checklist: the signals every production ensemble needs&lt;/h1>
&lt;p>A reference checklist for engineers running production ZooKeeper ensembles. Signals are organized into four maturity levels: survival, operational, mature, and expert. Each level adds visibility for failure modes the previous level cannot see.&lt;/p>
&lt;p>The levels are cumulative. Level 2 assumes Level 1 is covered. Skipping to Level 4 without Levels 1 through 3 leaves gaps in the signals that actually page you during incidents: disk stalls, GC cascades, and quorum loss.&lt;/p></description></item><item><title>ZooKeeper monitoring maturity model: from survival to expert</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-monitoring-maturity-model/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-monitoring-maturity-model/</guid><description>&lt;h1 id="zookeeper-monitoring-maturity-model-from-survival-to-expert">ZooKeeper monitoring maturity model: from survival to expert&lt;/h1>
&lt;p>Most ZooKeeper monitoring stops too early. Teams run &lt;code>ruok&lt;/code> as their only health check, collect &lt;code>zk_avg_latency&lt;/code> without understanding it is cumulative since the last reset, and never look at fsync latency until a write stall cascades into a Kafka outage. The ensemble looks healthy in dashboards until it suddenly does not, and the postmortem reveals the signals were there all along, uncollected.&lt;/p></description></item><item><title>ZooKeeper OutOfMemoryError: Java heap space - the OOM that kills the whole ensemble at once</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-heap-exhaustion-oom/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-heap-exhaustion-oom/</guid><description>&lt;h1 id="zookeeper-outofmemoryerror-java-heap-space---the-oom-that-kills-the-whole-ensemble-at-once">ZooKeeper OutOfMemoryError: Java heap space - the OOM that kills the whole ensemble at once&lt;/h1>
&lt;p>You grep the ZooKeeper log and find &lt;code>java.lang.OutOfMemoryError: Java heap space&lt;/code>. The process is gone. A minute later another node dies with the same error, then the third. The whole ensemble went down inside a single window, not as a rolling failure. That simultaneity is the signature, not a cascade.&lt;/p>
&lt;p>ZooKeeper holds the entire data tree on the JVM heap: every znode, its data, ACL references, children lists, stat structures, plus session state, watch tables, and request queues. Every ensemble member holds the same tree. Whatever fills the heap on one node fills it on all of them at roughly the same rate, so when the tree finally exceeds the heap they OOM near-simultaneously. This is a single-cause total outage.&lt;/p></description></item><item><title>ZooKeeper outstanding requests growing: the request pipeline is backing up</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-outstanding-requests-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-outstanding-requests-growing/</guid><description>&lt;h1 id="zookeeper-outstanding-requests-growing-the-request-pipeline-is-backing-up">ZooKeeper outstanding requests growing: the request pipeline is backing up&lt;/h1>
&lt;p>&lt;code>zk_outstanding_requests&lt;/code> counts requests queued in the server&amp;rsquo;s request processor pipeline that have not yet completed. In steady state it sits at or near zero. When it climbs, something downstream in the pipeline has stopped draining faster than clients are submitting.&lt;/p>
&lt;p>The queue is a leading indicator. Requests pile up before &lt;code>zk_avg_latency&lt;/code> reacts, and well before the server starts dropping client traffic. If you wait for latency alerts, you have already lost the lead time needed to keep dependent services like Kafka, HBase, and Solr from feeling the stall.&lt;/p></description></item><item><title>ZooKeeper pending syncs growing: followers can't keep up with the write rate</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-pending-syncs-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-pending-syncs-growing/</guid><description>&lt;h1 id="zookeeper-pending-syncs-growing-followers-cant-keep-up-with-the-write-rate">ZooKeeper pending syncs growing: followers can&amp;rsquo;t keep up with the write rate&lt;/h1>
&lt;p>&lt;code>zk_pending_syncs&lt;/code> is the leader&amp;rsquo;s count of in-flight sync operations to followers. In steady state it sits at zero. Sustained non-zero values mean at least one follower cannot absorb the proposal stream as fast as the leader generates it. This is ZooKeeper&amp;rsquo;s replication-lag signal, distinct from generic latency metrics because it points directly at the write pipeline&amp;rsquo;s fan-out side.&lt;/p></description></item><item><title>ZooKeeper proposals not committing: proposal_count outpacing commit_count</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-proposal-commit-divergence/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-proposal-commit-divergence/</guid><description>&lt;h1 id="zookeeper-proposals-not-committing-proposal_count-outpacing-commit_count">ZooKeeper proposals not committing: proposal_count outpacing commit_count&lt;/h1>
&lt;p>In a healthy ZooKeeper ensemble, every proposal the leader broadcasts is committed a few milliseconds later, after a quorum of followers acknowledges it. The two counters &lt;code>zk_proposal_count&lt;/code> and &lt;code>zk_commit_count&lt;/code> track the front and back of that pipeline, and on a cluster with active writers their rates should track each other closely.&lt;/p>
&lt;p>When &lt;code>zk_proposal_count&lt;/code> keeps climbing but &lt;code>zk_commit_count&lt;/code> stalls, the leader is generating proposals but cannot reach quorum ACK. Writes are not landing. Clients with operation timeouts start failing, ephemeral nodes are not being created, and dependent systems (Kafka controller, HBase master, Solr overseer) start logging &amp;ldquo;operation timeout&amp;rdquo; or &amp;ldquo;session expired&amp;rdquo;. A worse signal is both metrics staying flat while writers are active: the write pipeline is fully blocked, or the leader is isolated from the quorum.&lt;/p></description></item><item><title>ZooKeeper quorum ack latency high: followers slow to acknowledge proposals</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-quorum-ack-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-quorum-ack-latency-high/</guid><description>&lt;h1 id="zookeeper-quorum-ack-latency-high-followers-slow-to-acknowledge-proposals">ZooKeeper quorum ack latency high: followers slow to acknowledge proposals&lt;/h1>
&lt;p>&lt;code>zk_p99_quorum_ack_latency&lt;/code> is climbing on the leader and every write in the ensemble is paying for it. This metric measures the time from the leader sending a PROPOSE message to receiving quorum acknowledgments from followers. It is a leader-only signal.&lt;/p>
&lt;!-- TODO: verify availability. quorum_ack_latency and its percentile variants require the new metrics framework (3.6+). On older versions the metric will not exist. -->
&lt;p>Every ZooKeeper write blocks until a quorum of followers ACK. An ACK means the follower has written the proposal to its transaction log and fsync&amp;rsquo;d it to persistent storage. So quorum ack latency bundles two things into one number: the inter-node network round trip, and the follower&amp;rsquo;s fsync plus processing time. When this metric is elevated, the slowest follower in the quorum is bottlenecking writes for every client connected to the ensemble.&lt;/p></description></item><item><title>ZooKeeper quorum loss: no leader elected and every write is failing</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-quorum-loss-no-writes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-quorum-loss-no-writes/</guid><description>&lt;h1 id="zookeeper-quorum-loss-no-leader-elected-and-every-write-is-failing">ZooKeeper quorum loss: no leader elected and every write is failing&lt;/h1>
&lt;p>Every write to your ZooKeeper ensemble is timing out. Clients report &lt;code>ConnectionLoss&lt;/code> and &lt;code>SessionExpired&lt;/code>. Downstream systems that depend on ZK for coordination, such as Kafka controller elections or HBase region assignment, are cascading into failure. On the surviving ZK nodes, &lt;code>ruok&lt;/code> still returns &lt;code>imok&lt;/code>. The process is alive; the ensemble is not.&lt;/p>
&lt;p>Quorum loss is ZooKeeper&amp;rsquo;s worst-case availability scenario. When fewer than &lt;code>floor(N/2)+1&lt;/code> voting members can communicate, no leader can be elected and every write fails. Surviving nodes sit in &lt;code>LOOKING&lt;/code> state, unable to make progress through ZAB.&lt;/p></description></item><item><title>ZooKeeper read latency high: memory reads that should never be slow</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-read-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-read-latency-high/</guid><description>&lt;h1 id="zookeeper-read-latency-high-memory-reads-that-should-never-be-slow">ZooKeeper read latency high: memory reads that should never be slow&lt;/h1>
&lt;p>Reads in ZooKeeper are local heap lookups. A &lt;code>getData&lt;/code>, &lt;code>getChildren&lt;/code>, or &lt;code>exists&lt;/code> call should return in well under a millisecond because no quorum is involved. The connected server walks its in-memory data tree and replies. When &lt;code>zk_p99_readlatency&lt;/code> sits above 50ms for minutes at a time, something on that JVM is competing with request processing.&lt;/p>
&lt;p>Do not treat read latency like write latency. Write latency is dominated by transaction-log fsync and quorum ACK. Read latency has none of that machinery, so when it climbs the cause is almost always local: a Stop-the-World GC pause, a pathologically deep or wide znode tree pushing heap pressure, or a watch-delivery backlog stealing CPU from the request thread. Correlating &lt;code>zk_p99_readlatency&lt;/code> with &lt;code>zk_jvm_pause_time_ms&lt;/code>, &lt;code>zk_znode_count&lt;/code>, and &lt;code>zk_watch_count&lt;/code> is what separates a 30-second diagnosis from an hour of guessing.&lt;/p></description></item><item><title>ZooKeeper request throttling: globalOutstandingLimit and TCP backpressure</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-throttled-ops-global-outstanding-limit/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-throttled-ops-global-outstanding-limit/</guid><description>&lt;h1 id="zookeeper-request-throttling-globaloutstandinglimit-and-tcp-backpressure">ZooKeeper request throttling: globalOutstandingLimit and TCP backpressure&lt;/h1>
&lt;p>&lt;code>zk_throttled_ops&lt;/code> incrementing in production is a saturation alarm, not a tuning knob. By the time this counter moves, the request pipeline is already full: the server has stopped reading from client sockets because the global outstanding request queue has reached &lt;code>globalOutstandingLimit&lt;/code> (default 1000), TCP backpressure is propagating to every connected client, and any client that cannot absorb the added latency is on its way to a session expiration.&lt;/p></description></item><item><title>ZooKeeper server stuck in LOOKING: a node that never rejoins the quorum</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-server-stuck-looking/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-server-stuck-looking/</guid><description>&lt;p>A single ZooKeeper node sits in LOOKING long after the ensemble has settled, or flaps between LOOKING and FOLLOWING. The rest of the ensemble holds a stable leader with healthy write throughput. Restart the orphaned node and it drops back into LOOKING. Restart the leader and it briefly rejoins, then falls out again.&lt;/p>
&lt;p>This is not quorum loss. Quorum loss is every node entering LOOKING at once because no majority can form. That case is covered in &lt;a href="https://www.netdata.cloud/guides/zookeeper/zookeeper-quorum-loss-no-writes/">ZooKeeper quorum loss: no leader elected and every write is failing&lt;/a>. This article covers the narrower symptom: one permanently-orphaned node while the rest of the ensemble serves traffic.&lt;/p></description></item><item><title>ZooKeeper session count climbing: leaks and duplicate sessions</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-session-count-climbing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-session-count-climbing/</guid><description>&lt;h1 id="zookeeper-session-count-climbing-leaks-and-duplicate-sessions">ZooKeeper session count climbing: leaks and duplicate sessions&lt;/h1>
&lt;p>&lt;code>zk_global_sessions&lt;/code> climbing while your client fleet is stable is a slow ZooKeeper failure mode. The ensemble keeps serving reads and writes, latency looks fine, quorum is intact, but the session table keeps growing. Each entry costs heap and periodic heartbeat processing. Eventually you hit a GC death spiral, an OOM, or a &lt;code>maxClientCnxns&lt;/code>-shaped outage.&lt;/p>
&lt;p>The signal is simple to read but easy to misinterpret. ZooKeeper exposes two distinct populations: global sessions, which the leader echoes across the ensemble, and local sessions (only present when &lt;code>localSessionsEnabled=true&lt;/code>, added in ZK 3.5, default &lt;code>false&lt;/code>), which live on a single follower and upgrade to global when the client creates an ephemeral node. If you only watch &lt;code>zk_global_sessions&lt;/code> you may be looking at a subset of the actual session population, and if you only watch &lt;code>zk_num_alive_connections&lt;/code> you cannot tell a leak from a deployment event.&lt;/p></description></item><item><title>ZooKeeper session expiration storm: the ephemeral-node thundering herd</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-session-expiration-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-session-expiration-storm/</guid><description>&lt;h1 id="zookeeper-session-expiration-storm-the-ephemeral-node-thundering-herd">ZooKeeper session expiration storm: the ephemeral-node thundering herd&lt;/h1>
&lt;p>A session expiration storm is the worst-case thundering herd in a coordination service. Many clients lose contact long enough for the ensemble to declare their sessions dead. The cluster then deletes every ephemeral node owned by those sessions and fires every watch attached to those nodes. Every disconnected client reconnects at the same time, recreates its ephemeral nodes, and re-registers watches, hammering an ensemble that is already stressed.&lt;/p></description></item><item><title>ZooKeeper slow startup: snapshot load and txnlog replay taking minutes</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-slow-recovery-on-restart/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-slow-recovery-on-restart/</guid><description>&lt;h1 id="zookeeper-slow-startup-snapshot-load-and-txnlog-replay-taking-minutes">ZooKeeper slow startup: snapshot load and txnlog replay taking minutes&lt;/h1>
&lt;p>A ZooKeeper node you just restarted is not answering client requests. The process is up, the port is listening, but &lt;code>mntr&lt;/code> hangs or returns nothing useful, and dependent services are logging connection failures and session timeouts. Your dashboard shows &lt;code>zk_uptime&lt;/code> climbing past one, two, five minutes with no leader participation, or your health checks have already paged because the node looks dead.&lt;/p></description></item><item><title>ZooKeeper snapshot errors: recovery safety at risk</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-snapshot-errors/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-snapshot-errors/</guid><description>&lt;h1 id="zookeeper-snapshot-errors-recovery-safety-at-risk">ZooKeeper snapshot errors: recovery safety at risk&lt;/h1>
&lt;p>&lt;code>zk_snapshot_error_count&lt;/code> increments when ZooKeeper fails to serialize the in-memory data tree to disk, or fails to load a snapshot during startup. The node keeps serving reads and writes from its in-memory copy. The cost shows up on the next restart, when recovery cannot find a usable snapshot and either fails outright or replays stale state.&lt;/p>
&lt;p>A single increment is often transient. A backup job, a brief disk-full condition, or I/O contention from a colocated batch process can cause one snapshot to fail. The next snapshot cycle (taken every &lt;code>snapCount&lt;/code> transactions, default 100,000) succeeds and silently recovers. If you page on every increment, you burn through your team&amp;rsquo;s attention on events that self-resolve.&lt;/p></description></item><item><title>ZooKeeper split-brain: two nodes both reporting leader</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-split-brain-two-leaders/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-split-brain-two-leaders/</guid><description>&lt;h1 id="zookeeper-split-brain-two-nodes-both-reporting-leader">ZooKeeper split-brain: two nodes both reporting leader&lt;/h1>
&lt;p>Your monitoring polls each ZooKeeper node independently and two of them return &lt;code>zk_server_state leader&lt;/code>. Two servers in the same ensemble both believe they are the active leader. Treat this as an unconditional page: each side can accept writes that diverge from the other.&lt;/p>
&lt;p>ZAB (ZooKeeper Atomic Broadcast) is designed to make sustained split-brain impossible. A leader only becomes durable after a quorum of followers (floor(N/2)+1) has acknowledged the NEW_LEADER proposal. In a clean partition, the minority side cannot reach quorum and must block writes. So when two nodes both report leader for more than a brief convergence window, something has broken the quorum accounting: an asymmetric partition, a disk-full recovery bug, an authentication bypass, or a stale-read protocol corner case.&lt;/p></description></item><item><title>ZooKeeper stale reads from followers: zxid lag and read-after-write surprises</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-stale-reads-follower-lag/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-stale-reads-follower-lag/</guid><description>&lt;h1 id="zookeeper-stale-reads-from-followers-zxid-lag-and-read-after-write-surprises">ZooKeeper stale reads from followers: zxid lag and read-after-write surprises&lt;/h1>
&lt;p>A follower can pass every standard health check, serve reads at sub-millisecond latency, and still return data dozens of transactions behind what the leader just committed. No error is logged. No alert fires. The ensemble reports a leader, all followers are synced, and quorum is intact.&lt;/p>
&lt;p>This is not a bug. ZooKeeper offers sequential consistency for reads, not linearizability. Reads are served locally from each server&amp;rsquo;s in-memory data tree, and that tree may lag behind the leader&amp;rsquo;s committed state by anywhere from a few transactions to far more under load. For configuration data or service discovery where eventual consistency is tolerable, this is fine. For distributed locks, leader election fencing, or read-after-write logic, it can cause duplicate work, lock violations, or lost updates.&lt;/p></description></item><item><title>ZooKeeper stale requests dropped: requests aging out of the pipeline</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-stale-requests-dropped/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-stale-requests-dropped/</guid><description>&lt;h1 id="zookeeper-stale-requests-dropped-requests-aging-out-of-the-pipeline">ZooKeeper stale requests dropped: requests aging out of the pipeline&lt;/h1>
&lt;p>&lt;code>zk_stale_requests_dropped&lt;/code> incrementing is a late signal in a ZooKeeper saturation cascade. By the time a request ages out of the pipeline and is dropped, the server has already exhausted queue headroom, engaged throttling, and held the request long enough that the client gave up or the connection died. Treat any non-zero rate as an incident, and treat it as proof that an earlier signal was missed.&lt;/p></description></item><item><title>ZooKeeper synced_followers below ensemble size: degraded fault tolerance</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-synced-followers-degraded/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-synced-followers-degraded/</guid><description>&lt;h1 id="zookeeper-synced_followers-below-ensemble-size-degraded-fault-tolerance">ZooKeeper synced_followers below ensemble size: degraded fault tolerance&lt;/h1>
&lt;p>&lt;code>zk_synced_followers&lt;/code> reports how many followers are currently synced with the leader. In a healthy ensemble it equals &lt;code>ensemble_size - 1&lt;/code> (voting members only; observers are excluded). When it drops, a follower is disconnected or lagging, and your fault tolerance margin has shrunk.&lt;/p>
&lt;p>This metric is emitted only by the leader. Followers, observers, and standalone nodes do not report it. If your collector scrapes a fixed node or only followers, you have a blind spot: identify the leader dynamically, or scrape every node and keep only the values from the node reporting &lt;code>leader&lt;/code> in &lt;code>zk_server_state&lt;/code>.&lt;/p></description></item><item><title>ZooKeeper TLS handshake failures: unsuccessful handshakes and non-mTLS connections</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-tls-handshake-failures/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-tls-handshake-failures/</guid><description>&lt;h1 id="zookeeper-tls-handshake-failures-unsuccessful-handshakes-and-non-mtls-connections">ZooKeeper TLS handshake failures: unsuccessful handshakes and non-mTLS connections&lt;/h1>
&lt;p>When ZooKeeper is configured for TLS, the TLS layer becomes a new failure surface between clients and the ensemble. A spike in &lt;code>zk_unsuccessful_handshake&lt;/code> or &lt;code>zk_tls_handshake_exceeded&lt;/code> means clients are attempting TLS connections that never complete. If your environment requires mutual TLS, any non-zero value in &lt;code>zk_non_mtls_remote_conn_count&lt;/code> means a client bypassed mTLS entirely.&lt;/p>
&lt;p>These failures are noisy in a particular way. The client sees a connection timeout or a refused handshake, while the server logs an &amp;ldquo;Unsuccessful handshake&amp;rdquo; entry that does not always explain why. Certificate expiry, cipher mismatch, hostname verification failure, and low host entropy all produce similar surface symptoms. The fix is rarely the server itself; it is usually a client configuration error, an expired credential, or a monitoring tool sending plaintext to a secure port.&lt;/p></description></item><item><title>ZooKeeper transaction log disk full: the crash with no graceful degradation</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-disk-full-txnlog/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-disk-full-txnlog/</guid><description>&lt;h1 id="zookeeper-transaction-log-disk-full-the-crash-with-no-graceful-degradation">ZooKeeper transaction log disk full: the crash with no graceful degradation&lt;/h1>
&lt;p>ZooKeeper has no graceful degradation path for a full &lt;code>dataLogDir&lt;/code> partition. When the WAL append fails, the server throws an IOException and dies. There is no read-only fallback, no throttling, and no &lt;code>mntr&lt;/code> warning that precedes the crash. The same applies to the snapshot directory when the next snapshot write or pre-allocation fails.&lt;/p>
&lt;p>The most common root cause is broken or disabled autopurge. With &lt;code>autopurge.purgeInterval&lt;/code> defaulting to &lt;code>0&lt;/code> (disabled) and &lt;code>autopurge.snapRetainCount&lt;/code> defaulting to &lt;code>3&lt;/code>, an ensemble that has never been explicitly configured will accumulate transaction logs and snapshots forever. Disk consumption is silent and cliff-edge. By the time &lt;code>ruok&lt;/code> fails, the process is already gone.&lt;/p></description></item><item><title>ZooKeeper unexpected leader election: finding why the leader dropped</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-unexpected-leader-election/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-unexpected-leader-election/</guid><description>&lt;h1 id="zookeeper-unexpected-leader-election-finding-why-the-leader-dropped">ZooKeeper unexpected leader election: finding why the leader dropped&lt;/h1>
&lt;p>An unexpected ZooKeeper leader election is an availability event. While the ensemble is in LOOKING state, no writes are processed. Systems that depend on ZK for coordination queue or fail their mutations, and &lt;code>zk_sum_leader_unavailable_time&lt;/code> climbs. Treat every unplanned election like a database failover: the cluster recovered, but you still need the root cause.&lt;/p>
&lt;p>The cause is on the old leader or on the path to it. Followers call an election when they have not heard from the leader within &lt;code>syncLimit * tickTime&lt;/code> (default 5 * 2000ms = 10 seconds). The question is always: what stopped the old leader from sending heartbeats, or what stopped a quorum of followers from receiving them.&lt;/p></description></item><item><title>ZooKeeper unrecoverable error: when a node's integrity is compromised</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-unrecoverable-error/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-unrecoverable-error/</guid><description>&lt;h1 id="zookeeper-unrecoverable-error-when-a-nodes-integrity-is-compromised">ZooKeeper unrecoverable error: when a node&amp;rsquo;s integrity is compromised&lt;/h1>
&lt;p>You got paged because &lt;code>zk_unrecoverable_error_count&lt;/code> incremented. This is one of the few &lt;code>mntr&lt;/code> counters you never want to see move. It tracks errors ZooKeeper cannot recover from internally: data corruption, invariant violations, or resource exhaustion that puts the node into a state it cannot safely keep serving from.&lt;/p>
&lt;p>ZooKeeper is designed to fail fast. When it hits an unrecoverable condition it does not limp along. The critical thread logs a severe error, notifies the supervision listener, and the process exits. By the time the counter increments, the JVM is usually already gone or about to be, and the question is no longer &amp;ldquo;is this node healthy&amp;rdquo; but &amp;ldquo;do I trust its on-disk state at all&amp;rdquo;.&lt;/p></description></item><item><title>ZooKeeper watch storm: thousands of notifications when one hot znode changes</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-watch-count-storm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-watch-count-storm/</guid><description>&lt;h1 id="zookeeper-watch-storm-thousands-of-notifications-when-one-hot-znode-changes">ZooKeeper watch storm: thousands of notifications when one hot znode changes&lt;/h1>
&lt;p>A single znode changes. Seconds later, &lt;code>zk_packets_sent&lt;/code> spikes to many times its normal rate, CPU on the ZooKeeper process surges, and &lt;code>zk_outstanding_requests&lt;/code> begins climbing. Clients report latency spikes, connection timeouts, or session expirations. If many clients watch the same znode, you are looking at a watch storm.&lt;/p>
&lt;p>The mechanism: ZooKeeper maintains a watch table mapping znode paths to registered watchers. When a watched znode changes, the server queues a notification for every client that registered a watch on that path. For a path with 10,000 watchers, a single &lt;code>setData&lt;/code> call produces 10,000 notification packets. Notification serialization and queuing happen inside the request processing pipeline, competing with all other request handling.&lt;/p></description></item><item><title>ZooKeeper with no authentication: the open-by-default coordination store</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-no-authentication-open-access/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-no-authentication-open-access/</guid><description>&lt;h1 id="zookeeper-with-no-authentication-the-open-by-default-coordination-store">ZooKeeper with no authentication: the open-by-default coordination store&lt;/h1>
&lt;p>ZooKeeper ships with no authentication by default. Any TCP client that can reach port 2181 can open a session, read any znode, and write any znode whose ACL has not been explicitly restricted. The default ACL on znodes created without an explicit ACL is &lt;code>OPEN_ACL_UNSAFE&lt;/code>: &lt;code>world:anyone&lt;/code> with full &lt;code>cdrwa&lt;/code> (create, read, write, delete, admin) permissions.&lt;/p>
&lt;p>This matters because ZooKeeper is the coordination store for systems that treat its contents as authoritative: Kafka broker registrations and controller elections (pre-KRaft), HBase region assignment and master election, HDFS NameNode HA fencing state. An unauthenticated writer in those subtrees can silently corrupt cluster state, force leader changes, or trigger cascading failovers, and the writes succeed because the ACL permits them. There is no second layer that catches them.&lt;/p></description></item><item><title>ZooKeeper write latency high: read zk_updatelatency, not just avg_latency</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-write-latency-high/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-write-latency-high/</guid><description>&lt;h1 id="zookeeper-write-latency-high-read-zk_updatelatency-not-just-avg_latency">ZooKeeper write latency high: read zk_updatelatency, not just avg_latency&lt;/h1>
&lt;p>The dashboard says ZooKeeper is fine. &lt;code>zk_avg_latency&lt;/code> is 2ms. Clients are timing out anyway: distributed locks expiring mid-acquisition, Kafka controllers flapping, HBase regions bouncing. The signal you are missing is &lt;code>zk_updatelatency&lt;/code>, the write-specific latency family that 3.6+ exposes separately from the misleading aggregate.&lt;/p>
&lt;p>The trap is structural. &lt;code>zk_avg_latency&lt;/code>, &lt;code>zk_min_latency&lt;/code>, and &lt;code>zk_max_latency&lt;/code> from &lt;code>mntr&lt;/code> combine reads and writes into one cumulative statistic. Reads are local in-memory lookups, typically sub-millisecond. Writes require a leader round-trip, a ZAB proposal, a quorum ACK, a commit, and an fsync. When read volume dominates, a healthy average masks pathological write latency. These are also server-cumulative statistics since the last &lt;code>srst&lt;/code> reset, not sliding windows: a single fsync stall from three hours ago still inflates &lt;code>zk_max_latency&lt;/code>.&lt;/p></description></item><item><title>ZooKeeper znode count growing unbounded: the silent heap killer</title><link>https://www.netdata.cloud/guides/zookeeper/zookeeper-znode-count-growing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/guides/zookeeper/zookeeper-znode-count-growing/</guid><description>&lt;h1 id="zookeeper-znode-count-growing-unbounded-the-silent-heap-killer">ZooKeeper znode count growing unbounded: the silent heap killer&lt;/h1>
&lt;p>The symptom is familiar: a ZooKeeper ensemble that ran cleanly for months suddenly enters a GC death spiral. Heap climbs, full GC pauses stretch from milliseconds to seconds, sessions expire, and the JVM OOMs. The process restarts, the data tree reloads from snapshot, and the cycle repeats. All members OOM at roughly the same time because they carry the same in-memory data tree.&lt;/p></description></item><item><title>ZRAM</title><link>https://www.netdata.cloud/integrations/data-collection/operating-systems/zram/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/data-collection/operating-systems/zram/</guid><description/></item><item><title>Zyxel Communications Corp SNMP Traps</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/zyxel-communications-corp-snmp-traps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/snmp-traps/zyxel-communications-corp-snmp-traps/</guid><description/></item><item><title>Zyxel Switch</title><link>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/zyxel-switch/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.netdata.cloud/integrations/network-performance-monitoring/device-metrics/zyxel-switch/</guid><description/></item></channel></rss>