The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / consul / consul-gossip-queue-backlog ▌

Operations Guides

Consul serf queue backlog: an agent falling behind on gossip

A Consul agent whose consul.serf.queue.Event, consul.serf.queue.Intent, or consul.serf.queue.Query metric stays above zero is no longer keeping up with gossip. In steady state all three should read zero. Transient spikes during bulk joins, leaves, and rolling restarts are expected and self-drain within a few gossip intervals. The problem begins when the spike does not drain.

Serf’s queue depth is the observable proxy for the node’s internal health score. A sustained non-zero value means the agent is receiving gossip faster than its event loop can process. Downstream effects start subtle: failure detection latency rises, join and leave intents propagate slowly, and the node’s own probe replies arrive late at peers. Then the feedback loop engages.

The slow node gets marked suspect and then failed by peers that no longer hear from it within the probe window. Each state transition generates fresh gossip: suspicion, failure, and later refute messages when the node recovers and re-announces itself. That additional gossip lands back on the overloaded node’s queue, deepening the backlog. If the affected agent is a server, the cost compounds: CPU cycles spent draining gossip are cycles not available for Raft heartbeats and FSM applies.

What this means

Consul’s Serf layer maintains three named queues that buffer inbound gossip before the agent’s event loop processes it: Event for user events and member transitions, Intent for join and leave intents, and Query for serf queries. Beneath those, memberlist maintains its own broadcast queue exposed as consul.memberlist.queue.broadcasts. A backlog in any layer delays everything that depends on timely membership convergence.

The healthy baseline is zero across all queues, all the time. When you see non-zero values, ask two questions: is it transient, and is it growing? A spike that returns to zero within 30 to 60 seconds during a known event is the protocol working as designed. A value that plateaus above zero or climbs monotonically is the protocol losing ground.

Sustained backlog is self-reinforcing. A node that falls behind cannot answer probes within the configured timeout. Peers suspect it, and suspicion and failure transitions are themselves gossip messages distributed to every member. Lifeguard (the Serf health-score mechanismLifeguard was first introduced in Consul 0.7 and further refined in Consul 1.0.0 with updated memberlist LAN gossip tuning) dampens this by letting an overloaded node advertise a degraded health score so peers back off, but it mitigates rather than eliminates the loop. Encryption amplifies the cost: every gossip message is encrypted on send and decrypted on receive, so a CPU-starved node pays the crypto tax on backlog it cannot drain.

flowchart TD
    A[Agent receives gossip faster than it processes] --> B[serf.queue.Event / .Intent / .Query above 0]
    B --> C[Probe replies arrive late or miss timeout]
    C --> D[Peers mark node suspect then failed]
    D --> E[Transitions generate more gossip]
    E --> A
    B --> F[Node serfHealth check goes critical]
    F --> E

Common causes

CauseWhat it looks likeFirst thing to check
Agent CPU starvationQueue depth tracks CPU utilization; spikes during health check bursts or GC pausestop -p $(pgrep consul) and consul.runtime.gc_pause_ns trend
Cluster too large for gossip paramsAll agents show elevated queues during normal operation, not just oneCompare agent count to the 5,000-per-pool guidance and current gossip_interval
Gossip encryption without CPU headroomBacklog appears only on nodes with encryption enabled; CPU split between crypto and processingconsul keyring -list and per-core CPU during gossip bursts
Slow or lossy network to many peersOne node or one rack shows backlog while others are cleanUDP packet loss on the gossip port, netstat -s for receive errors
Mass recovery after partitionMany nodes rejoin simultaneously, all queues spike togetherCorrelate timing with AZ recovery or deploy event

Quick checks

Read-only and safe to run on any agent.

# Current serf queue depths (look for sustained non-zero)
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep -E 'serf.queue'

# Memberlist broadcast queue (the layer beneath serf)
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep -E 'memberlist.queue'

# Member status from this node's gossip view
curl -s http://127.0.0.1:8500/v1/agent/members | python3 -c "
import sys,json
for m in json.load(sys.stdin):
    print(m['Name'], m['Status'])
"

# Agent CPU pressure proxies
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep -E 'runtime.gc_pause|runtime.num_goroutines'

# Serf debug block from consul info
consul info | grep -A 20 serf_lan

# Gossip encryption state (all nodes should report the same key)
consul keyring -list

# UDP receive errors on the gossip interface
netstat -s | grep -iE 'receive.*error|packet receive'

Some telemetry exporters (Prometheus format) replace dots with underscores in metric names, for example consul_serf_queue_Event. Both forms refer to the same underlying value.

How to diagnose it

  1. Confirm the backlog is sustained, not transient. Sample the queue metrics every few seconds for at least a minute. If values return to zero between samples, the cause was a bounded event. Move to prevention rather than remediation.

  2. Determine scope. Is this one node, one rack, one datacenter, or every agent? Query the metrics endpoint across a representative sample. A single affected node points to local CPU, network, or disk. A cluster-wide pattern points to gossip parameter tuning or cluster size.

  3. Correlate with CPU. Pull consul.runtime.gc_pause_ns and OS-level CPU for the consul process on affected nodes. Gossip processing is CPU-bound. If the node is at or near its CPU limit, the backlog is a symptom of starvation, not a gossip problem.

  4. Check member status divergence. Compare consul members output from several nodes. If the affected node appears failed or suspect on peers while reporting itself alive, the feedback loop is active. The backlog is now generating its own load.

  5. Verify encryption state. Run consul keyring -list. If nodes report different key counts or a rotation is mid-flight, some nodes are encrypting or decrypting against keys others have discarded, adding dropped-message overhead.

  6. Compare agent count to the gossip pool guidance. HashiCorp’s scale documentation recommends a maximum of 5,000 client agents per gossip pool. Near or above that ceiling with default parameters, the backlog reflects protocol saturation rather than any single node’s health.

  7. Inspect gossip parameters. Defaults for LAN are gossip_interval 200ms, probe_interval 1s, retransmit_mult 4, suspicion_mult 4. These are memberlist defaults that predate Consul 1.18 If your cluster has grown without revisiting these values, the parameter set may no longer match the workload.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
consul.serf.queue.IntentPrimary indicator of join/leave backlog; the queue most sensitive to churnSustained non-zero for more than 60 seconds outside a known event
consul.serf.queue.EventBacklog of member transitions and user eventsGrowing trend alongside .Intent
consul.serf.queue.QueryBacklog of serf query handlingRarely elevated alone; investigate if so
consul.memberlist.queue.broadcastsTransport layer beneath serf; if this grows, serf cannot drainAny sustained non-zero value
consul.runtime.gc_pause_nsGo runtime pauses block the gossip event loopSpikes above 50ms correlating with queue growth
OS CPU for consul processGossip processing and crypto are CPU-boundSustained above 80% on a node with backlog
serfHealth check statusThe node’s own gossip participation checkTransitions to critical indicate the feedback loop is engaged
Serf member status divergencePeers seeing the node as suspect or failedDisagreement across consul members outputs from different nodes

There is no officially documented numeric threshold for any serf queue metric. HashiCorp’s guidance is qualitative: consistently high values indicate the gossip pool cannot keep up with churn. Treat any sustained non-zero value as the alert condition and let duration and trend determine severity.

Fixes

The right fix depends on whether the backlog is local to one node or systemic. Do not restart the consul process as a first response. A restart clears the queue but does not address the cause, and the backlog returns as soon as the node rejoins and receives catch-up gossip.

CPU starvation on a single node

If the backlog correlates with high CPU on one agent, the gossip layer is competing with something else on the host. Common offenders: script-based health checks spawning subprocesses, GC pauses on a large heap, or CPU limits in a container scheduler.

  • Raise the CPU limit or move the agent to a less contended host.
  • Audit health check scripts for subprocess cost. A check that shells out to curl or jq on a tight interval multiplies fast.
  • If GC pauses are the driver (consul.runtime.gc_pause_ns spiking), raise GOGC to reduce GC frequency at the cost of higher peak heap, or reduce heap pressure by trimming catalog bloat.

Cluster too large for current gossip parameters

If every agent shows elevated queues during normal operation, the protocol is saturated. The 5,000-agent-per-pool ceiling is the documented soft limit for default parameters.

  • Increase gossip_interval to spread message emission over more time. This trades convergence speed for processing headroom.
  • Increase gossip_nodes to raise the fan-out per round, reducing total rounds needed at the cost of larger packets.
  • Adjust suspicion_mult to give slow nodes more time before peers declare them failed, dampening the feedback loop.
  • For clusters well above 5,000 agents, consider network segments or Consul on Kubernetes, which reduces the need for a client agent on every node and shrinks the gossip pool.

Test any gossip parameter change on a staging cluster first. These values affect failure detection latency cluster-wide.

Encryption overhead

Gossip encryption is enabled with the encrypt config field. Every message is encrypted on send and decrypted on receive. On CPU-constrained nodes, the crypto cost competes directly with gossip processing.

  • Ensure encrypted nodes have CPU headroom. A node sized for plaintext gossip may need more CPU once encryption is on.
  • Verify all nodes share the same key with consul keyring -list. A mid-rotation state where some nodes hold old and new keys adds processing overhead on every message.
  • Complete key rotations promptly. Lingering dual-key states increase per-message work.

Mass recovery storms

If the backlog follows an AZ recovery, partition heal, or mass rolling restart, the cause is legitimate catch-up load. The cluster is processing a bounded burst of join events and anti-entropy syncs.

  • Monitor Raft commit time during the storm. If commit time stays below the election timeout, the cluster will self-recover. Intervention risks making it worse.
  • If commit time approaches the election timeout, reduce write load temporarily: disable non-critical health checks or stagger the remaining rejoining nodes.
  • Do not tune gossip parameters reactively during a storm. Changes made under load are hard to validate and may persist after the storm drains.

Prevention

  • Alert on sustained non-zero serf queue depth. A 60-second window above zero outside known maintenance windows is a reliable early signal.
  • Track agent count against the 5,000-per-pool guidance as a capacity-planning metric, not just an incident trigger.
  • Include gossip parameters in your cluster documentation. If you have grown past the size where defaults were chosen, schedule a tuning review.
  • Monitor consul.runtime.gc_pause_ns and CPU on every agent, not just servers. Client agents run gossip too.
  • Periodically verify gossip key state with consul keyring -list, especially after any rotation.
  • Treat the feedback loop as a design constraint. If your failure detection is tight enough that a slow node gets marked failed before it can recover, your suspicion_mult may be too aggressive for your infrastructure’s worst-case CPU latency.

How Netdata helps

  • Per-second collection of consul.serf.queue.Event, .Intent, and .Query exposes sustained backlog that minute-granular scraping misses. A queue that spikes and drains between samples is invisible at lower resolution.
  • ML anomaly detection flags the transition from transient spike to sustained elevation without requiring a hand-tuned threshold, which matters because HashiCorp publishes no official numeric threshold for these metrics.
  • Correlating serf queue depth with OS CPU, GC pause duration, and goroutine count on the same timeline isolates CPU starvation from protocol saturation in a single view.
  • Member status and serfHealth signals sit alongside the queue metrics, so the feedback loop (queue grows, node marked suspect, more gossip generated) is visible as a coordinated pattern rather than isolated alerts.
  • The memberlist broadcast queue metric, when exposed, appears in the same dashboard, letting you see whether the backlog originates at the serf layer or beneath it.