The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nats / nats-connection-churn ▌

Operations Guides

NATS connection churn: a stable connection count hiding constant reconnects

Your NATS dashboard shows 4,000 client connections. It showed 4,000 an hour ago, and 4,000 yesterday. Everything looks stable. Meanwhile, clients are connecting and disconnecting hundreds of times per minute. Every reconnect burns CPU on protocol handshakes (and TLS handshakes, if enabled), the server logs fill with connect and disconnect events, and your auth system processes a constant stream of authentication attempts.

This is connection churn, and it is one of the most commonly missed NATS failure modes because the metric everyone charts, the current connections gauge, is designed to hide it. A client that disconnects and reconnects within one scrape interval leaves the gauge unchanged. The churn is real; your chart just cannot see it.

The detection metric is one most teams never chart: the delta of total_connections, the server’s lifetime cumulative connection counter. When connections is flat but total_connections is climbing fast, clients are flapping. This article covers how to detect the pattern, how to find the clients responsible, and how to fix the usual causes.

What this means

NATS exposes two connection figures on the /varz monitoring endpoint:

  • connections: the number of active client connections right now. A gauge.
  • total_connections: every connection established since the server started. A cumulative counter.

In a healthy deployment, clients are long-lived. connections stays near its steady-state value and total_connections advances slowly, only when clients deploy, scale, or genuinely restart. When something goes wrong on the client side, the two decouple: total_connections climbs while connections stays flat, because each disconnect is immediately replaced by a reconnect.

Each cycle has a real cost: a TCP handshake, optionally a TLS handshake, the NATS protocol handshake, authentication, and subscription re-establishment. A few hundred flapping clients can keep a core busy on handshakes alone while delivering no useful messages. Churn is also a leading indicator of worse things: it frequently accompanies the slow consumer disconnect-reconnect spiral, and reconnect bursts can push the server into max_connections or file descriptor exhaustion.

flowchart LR
  C[Client] -->|connect + auth| S[NATS server]
  S -->|disconnect: slow consumer, timeout, LB drop| C
  C -->|immediate reconnect, no backoff| S
  S --> V[varz: connections gauge flat]
  S --> T[varz: total_connections climbing fast]
  T --> D[churn = rate of total_connections]

The detection formula:

churn rate = delta(total_connections) / interval

If that rate is well above your baseline while connections is steady, you have churn. Baseline is workload-dependent: a batch system with short-lived workers legitimately has a higher connection rate than a fleet of daemons. What matters is deviation from your own norm.

One caveat before you alert on the raw counter: total_connections resets to zero on server restart, and the Prometheus exporter (gnatsd_varz_total_connections) exposes it as a gauge carrying the running total, not a Prometheus counter. rate() still works on it, but a server restart produces a reset that can look like a spike or a dip in increase(). Correlate with uptime before treating a counter discontinuity as a churn event.

Common causes

CauseWhat it looks likeFirst thing to check
Crash-looping clientsSteady churn correlated with one service or deployment; client pods/processes restartingClient-side restart counts (orchestrator), client logs
No reconnect backoff in the clientTight disconnect-reconnect loop after any error; churn spikes during any server or network blipClient library reconnect configuration
Slow consumer disconnect-reconnect spiralChurn plus rising slow_consumers counter; same clients flagged repeatedlyslow_consumer_stats breakdown, /connz?sort=pending
Load balancer health-check flappingChurn from LB source IPs; short-lived connections that never subscribe/connz?state=closed, group by source IP
Network event recovery stormSharp churn burst after a partition heals; all clients reconnect at onceNetwork device/cloud event timeline, preceding connections drop
Leaf node link cyclingChurn on leaf connections specifically; edge locations losing and regaining hub connectivity/leafz state and RTT
Auth or credential problemsConnect attempts fail after TCP establish; high churn with auth errors in server logsServer logs for authorization violations

Quick checks

All checks are read-only against the monitoring port (default 8222).

# Snapshot the three connection figures
curl -s http://localhost:8222/varz | jq '{active: .connections, total: .total_connections, max: .max_connections}'

# Measure churn directly: two samples 30s apart, compute the delta
A=$(curl -s http://localhost:8222/varz | jq .total_connections)
sleep 30
B=$(curl -s http://localhost:8222/varz | jq .total_connections)
echo "new connections in 30s: $((B - A))  (rate: $(( (B - A) / 30 ))/s)"

# Rule out a server restart resetting the counter mid-measurement
curl -s http://localhost:8222/varz | jq .uptime

# Check whether churn is tied to slow consumer events
curl -s http://localhost:8222/varz | jq '{slow_consumers, slow_consumer_stats}'

# Look at recently closed connections: who is leaving, and why
curl -s 'http://localhost:8222/connz?state=closed' | jq '.connections[:20] | .[] | {cid, name, ip, reason, subscriptions}'

# Find connections currently building write backlog (pre-slow-consumer)
curl -s 'http://localhost:8222/connz?sort=pending&limit=10' | jq '.connections[] | {cid, name, ip, pending_bytes, rtt}'

Notes on the closed-connection check: the server holds a bounded window of recently closed connections (configurable with max_closed_clients; the default is the last 10,000). Under heavy churn, that window can cover only a few minutes, so sample it promptly when you see the churn rate spike, not an hour later.

If you scrape via the Prometheus exporter, the equivalent query is:

# New connections per second, per server
rate(gnatsd_varz_total_connections[5m])

Remember the metric is a gauge that resets on restart; ignore samples spanning an uptime reset.

How to diagnose it

  1. Confirm churn exists. Take two /varz samples 30 to 60 seconds apart. If total_connections advances by far more than your expected deploy/scale activity while connections is flat, churn is confirmed. Note the rate: single digits per second is a slow leak, hundreds per second is an active storm.

  2. Check the clock. Read uptime. A server that restarted 5 minutes ago will show a “climbing” total_connections purely from normal client reconnection after the restart. Churn conclusions are only valid on a server with stable uptime. See NATS crash loop: unexpected uptime resets and repeated restarts if uptime itself is the problem.

  3. Correlate with slow consumers. Read slow_consumer_stats from /varz (the clients/routes/gateways breakdown is available since nats-server v2.10.0). If the clients component is rising in step with churn, clients are being disconnected for falling behind and reconnecting into the same backlog. That is the slow consumer death spiral, and the fix belongs on the consumer, not the connection layer. If routes or gateways are non-zero, the churn involves inter-server links and the blast radius is cluster-wide; treat it as more urgent.

  4. Identify who is flapping. Pull /connz?state=closed and group the results by ip, name, and reason. One service name or one source subnet dominating the list points at the culprit. Short-lived connections with zero or near-zero subscriptions and LB source IPs point at health-check traffic rather than real clients.

  5. Characterize the cycle timing. Compare the start and stop fields in the closed-connection data. Loops of a few seconds suggest no reconnect backoff or immediate failure after connect (auth, permissions, protocol error). Loops of tens of seconds to minutes suggest a client that connects fine, falls behind, and gets disconnected as a slow consumer.

  6. Check the client side. For the implicated service: process restart counts, client logs around reconnects, and the reconnect/backoff configuration of its NATS client library. Server-side data tells you who and how often; only client-side data tells you why the loop started.

  7. Quantify the cost. Compare cpu on /varz against message throughput (in_msgs, out_msgs). CPU elevated without a corresponding message rate means the server is spending cycles on handshakes and teardown rather than routing. With TLS enabled this gap is especially pronounced.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
total_connections delta (churn rate)The only server-side signal that sees flappingSustained positive rate above baseline with flat connections
connections vs max_connectionsReconnect bursts can hit the hard wallUtilization above 85%, or spikes toward the limit during churn events
slow_consumers rate and slow_consumer_stats breakdownTells you whether churn is caused by backpressure disconnects, and whether routes/gateways are involvedAny sustained positive rate; non-zero routes or gateways component
uptimeDistinguishes churn from a counter reset after restartUnexpected resets; churn analysis invalid around resets
stalled_clients and stale_connections (available in /varz since nats-server v2.12.0)Write-path distress and half-dead sockets that precede disconnectsAny non-zero value sustained over 5 minutes
cpu vs message rateExposes handshake overhead from churnCPU rising while in_msgs/out_msgs stay flat
Server log connect/disconnect and auth event rateEach churn cycle generates log events and possibly auth attemptsLog volume growth disproportionate to traffic

Fixes

Crash-looping clients

The server is a victim here, not the cause. Fix the client crash first: read the client logs, fix the panic or config error, then let churn settle. If many instances of the same service crash-loop together, stop the rollout at the orchestrator level rather than letting the whole fleet hammer the server. Tradeoff: pausing a rollout leaves you on the old version, but it stops the connection storm immediately.

Missing reconnect backoff

NATS client libraries reconnect automatically, and most include jittered backoff, but custom wrappers and misconfigured options can disable it or set the wait to near zero. Set a minimum reconnect wait and keep jitter enabled so a population of disconnected clients does not reconnect in lockstep. There is no server-side fix for a client that reconnects instantly; this must be changed in the client configuration and redeployed.

Slow consumer disconnect-reconnect spiral

This is the nastiest variant: the server correctly disconnects a client that cannot keep up, the client reconnects, resubscribes, immediately falls behind again, and the loop repeats, incrementing slow_consumers each cycle. Do not “fix” this by raising server buffers; that only delays the disconnect and grows memory. Fix the consumer’s throughput: remove synchronous I/O from the message handler, scale the consumer out, or move the workload to JetStream pull consumers where the consumer controls its own rate. If you genuinely need to adjust how long the server tolerates a slow writer, see NATS write_deadline and buffer sizing. For identifying which connections are building backlog before they are disconnected, see NATS pending bytes growing.

Load balancer health-check flapping

If closed connections are short-lived, carry no subscriptions, and come from LB addresses, your health checks are the churn. Where the load balancer supports it, point health checks at the HTTP monitoring port’s /healthz endpoint instead of opening a fresh TCP connection to the client port on every probe; see NATS /healthz explained for what the variants check. If the LB can only do TCP connects, lengthen the probe interval and exclude LB source ranges from your churn alerting so real client churn stays visible.

Reconnect storms after network recovery

When a partition heals, every disconnected client reconnects at once. This is a burst, not a loop: churn spikes and then decays. The operational risks are CPU saturation from simultaneous handshakes (especially with TLS) and running into max_connections or the OS file descriptor limit during the peak. Keep at least 20% headroom below max_connections precisely to absorb these storms; see NATS Maximum Connections Exceeded. If the “storm” repeats cyclically, the underlying network fault is still active and that is what needs fixing.

A leaf link under heavy hub-to-leaf load can enter a cycle where the write path to the leaf saturates, protocol liveness responses queue behind data, the leaf declares the connection stale, and reconnects into the same backlog. Treat this like a route-class slow consumer: reduce the traffic volume over the leaf or fix the bottleneck on the receiving side, not the reconnect timing.

Prevention

  • Chart the churn rate permanently. Add rate of total_connections (per server) to your standard NATS dashboard next to the connections gauge. This is the single highest-value change; every other item on this list is easier once you can see the signal.
  • Alert on deviation, not absolutes. Alert when the churn rate exceeds your rolling baseline by a wide margin for more than a few minutes. Absolute thresholds break across deployments of different sizes.
  • Gate alerts on uptime. Suppress churn alerts for servers with uptime under a few minutes; post-restart reconnection looks identical to churn.
  • Enforce reconnect backoff in shared client wrappers. If your organization wraps NATS client setup in a library, make minimum backoff and jitter non-overridable defaults.
  • Correlate churn with slow consumer signals. A churn alert that also shows rising slow_consumers and per-connection pending_bytes points straight at a backpressure problem and skips an hour of guessing.
  • Keep connection headroom. Size max_connections and OS file descriptor limits so a full-fleet reconnect storm fits with room to spare.

How Netdata helps

  • Netdata collects both connections and total_connections from /varz on every server, so the churn rate and the steady-state gauge are visible on the same dashboard without building custom scrapes.
  • Because collection is per-second, short-lived connect-disconnect cycles that a 30 or 60 second scrape interval would average away still show up in the cumulative counter’s slope.
  • Netdata charts slow_consumers alongside connection metrics, making the churn-plus-backpressure spiral a one-screen correlation instead of two separate queries.
  • Uptime is collected as a metric, so counter resets from server restarts are easy to distinguish from genuine churn when reviewing an incident timeline.
  • Per-server views across a cluster let you see whether churn is isolated to one node (client affinity, LB backend issue) or uniform (client-side bug, network-wide event).
  • ML-based anomaly detection on the total_connections rate catches churn that deviates from your normal deploy-driven baseline without hand-tuned thresholds.