The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nats / nats-stalled-clients-stale-connections ▌

Operations Guides

NATS stalled clients and stale connections: half-dead sockets and write-path distress

Two counters in /varz tell you about connections that are unhealthy but not yet dead: stalled_clients and stale_connections. Neither one means a connection has been dropped. That is the point: both describe states that precede visible failure. Stalled clients are on the way to becoming slow consumers; stale connections are sockets that look open but whose peer has stopped responding.

The operational risk is twofold. Stalled clients signal write-path backpressure: the server cannot flush data to a client fast enough, which is the precursor to slow consumer disconnection and, in core NATS, silent message loss. Stale connections are quieter: they consume file descriptors and per-connection memory while doing no useful work, and they usually indicate a network half-partition or a hung client process that will not recover on its own.

This article covers what each counter means, how to find the specific connections behind them, and how to fix the underlying causes. For the broader NATS signal taxonomy, see the NATS monitoring checklist.

What this means

Every client TCP connection gets dedicated read and write loops plus a per-connection pending buffer on the write side. When a client cannot consume messages fast enough, that buffer grows. The server enforces backpressure with write_deadline: if a write to the connection cannot complete within the deadline (default 10 seconds; tunable in server config), the connection is in write-path distress.

The stalled_clients counter in /varz tracks connections that have entered this stalled state. It is a pre-slow-consumer signal: the buffer is under pressure, but the server has not yet flagged the connection as a slow consumer and disconnected it.

Separately, the server runs a ping/pong health check over each connection. If a client fails to respond to protocol pings, the connection is declared stale, counted in stale_connections, and closed. The key property of a stale connection before closure is that the TCP socket still looks alive from the server’s perspective: the peer is not responding at the protocol level, but TCP has not noticed, or has not been allowed to notice. This is the classic half-dead socket. Common causes: a hung client process (GC storm, deadlock, stopped container), a stateful middlebox that silently dropped the flow state, or a network partition where neither side ever sees a RST.

flowchart TD
  A[Client connection] --> B{Write buffer draining?}
  B -- yes --> C[Healthy]
  B -- no, write_deadline hit --> D[stalled_clients]
  D --> E{Buffer keeps growing?}
  E -- yes --> F[slow consumer: disconnect, messages dropped]
  E -- no --> C
  A --> G{Responds to PING?}
  G -- yes --> C
  G -- no --> H[stale_connections: half-dead socket]
  H --> I[FDs and memory held with no useful work]

A useful mental split: stalled_clients is the server pushing data out and failing; stale_connections is the server checking liveness and failing. Different root causes, different blast radii. Never treat them as one combined “bad connections” number.

Both are cumulative counters since server start, so what matters is the rate of change, not the absolute value. A positive rate sustained over more than 5 minutes warrants investigation. A static non-zero value from last Tuesday’s deploy is history, not signal.

Common causes

CauseWhat it looks likeFirst thing to check
Hung or frozen client processstale_connections growing; client RTT inflating before it went silentIs the client process alive and scheduling? Check CPU, GC pauses, container state on the client host
Network half-partition or middlebox dropping flow statestale_connections on a subset of clients from one network segment; no RSTs seenDo the stale clients share a path (AZ, VPN, LB, NAT)? Check /connz for their IPs
Slow subscriber falling behindstalled_clients growing, then slow_consumers starts incrementing/connz?sort=pending for connections with high pending_bytes
Write buffer saturation from burst or fan-out spikestalled_clients correlates with an out_msgs spikeCompare in_msgs/out_msgs rate against the stall timing
Route or gateway write-path distressstalled_clients alongside route pending_size growth/routez pending_size; cluster impact is much larger than a client stall
Undersized or mis-tuned write_deadlinetransient stalls during normal bursts that self-clearDoes stall rate correlate with known traffic bursts, and does slow_consumers stay flat?

Quick checks

All of these are read-only against the monitoring HTTP port (default 8222).

# 1. Current counter values
curl -s http://localhost:8222/varz | jq '{stale_connections, stalled_clients}'

# 2. Two snapshots 60s apart give you a rate, not a point value
curl -s http://localhost:8222/varz | jq '{stale_connections, stalled_clients, uptime}'
sleep 60
curl -s http://localhost:8222/varz | jq '{stale_connections, stalled_clients, uptime}'

# 3. Clients with the largest write backlog (pre-slow-consumer)
curl -s "http://localhost:8222/connz?sort=pending&limit=10" | jq '.connections[] | {cid, name, ip, pending_bytes, subscriptions, rtt}'

# 4. Has the slow consumer consequence started firing?
curl -s http://localhost:8222/varz | jq '{slow_consumers, slow_consumer_stats}'

# 5. If clustered: are routes backing up? (higher blast radius than clients)
curl -s http://localhost:8222/routez | jq '.routes[] | {rid, ip, pending_size, rtt}'

# 6. Churn: stable connections with fast-growing total_connections means reconnect loops
curl -s http://localhost:8222/varz | jq '{connections, total_connections}'

# 7. RTT outliers: high-RTT clients are prime slow consumer candidates
curl -s "http://localhost:8222/connz?sort=rtt&limit=10" | jq '.connections[] | {cid, name, ip, rtt}'

Two safety notes. /connz is expensive on servers with tens of thousands of connections; always use limit and sort parameters rather than pulling the full list repeatedly. And /varz counters reset on server restart, so a sudden drop to zero is a restart, not a recovery. Check uptime before celebrating.

How to diagnose it

  1. Establish which counter is moving. Poll /varz twice, 60 seconds apart. If stalled_clients is growing, you have a write-path problem; go to step 2. If stale_connections is growing, you have a liveness problem; go to step 4. If both are moving, treat the write-path side first, since it precedes message loss.

  2. Identify the stalled connections. Use /connz?sort=pending and look at the top entries by pending_bytes. Note whether the worst offenders are application clients or, if the sort output and /routez suggest it, routes or gateways. A route or gateway with growing pending bytes is a cluster-wide problem, not a client problem. See NATS slow consumer breakdown: clients vs routes vs gateways for the blast-radius analysis.

  3. Confirm the trajectory. High pending bytes that clear on the next poll are burst absorption, which is normal. Pending bytes that grow across polls are a consumer falling behind. Check whether slow_consumers has started incrementing: once it does, NATS is disconnecting the connection and, in core NATS, dropping messages for it. At that point this article hands off to NATS slow consumer detected.

  4. For stale connections, find the common factor. The question is why the peer stopped answering. Pull the affected connections from /connz and look for patterns in ip, name, or account. A cluster of stale clients from one network segment points at a middlebox or partition. A single stale client with an inflated RTT before going silent points at a hung process on that host.

  5. Check the client side. For a hung client, verify process state on the client host: CPU starvation, a long GC pause, or a deadlock in the message handler will all stop protocol responses while leaving TCP open. For suspected network issues, check whether the path involves a NAT, VPN, or L4 load balancer with an idle-flow timeout shorter than the NATS ping interval.

  6. Quantify the resource cost. Each stale connection holds a file descriptor plus per-connection buffers and goroutines. If the stale count is large, compare connections against max_connections and check the process FD count against ulimit -n. This is how a quiet stale-connection problem becomes an FD exhaustion cliff later.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
/varz stalled_clientsPre-slow-consumer write-path distressAny positive rate sustained > 5 min
/varz stale_connectionsHalf-dead sockets holding FDs and memoryAny positive rate sustained > 5 min
/connz pending_bytes (top-N)Leading indicator for specific clientsSustained growth on the same CIDs across polls
/varz slow_consumers and slow_consumer_statsThe consequence firing; breakdown tells you clients vs routes vs gatewaysAny increment; route/gateway entries are urgent
/routez pending_sizeRoute-level backpressure, cluster-wide impactAny sustained non-zero value
/connz rttHigh-RTT clients become slow consumers firstRTT rising from baseline on busy subscribers
/varz connections vs total_connectionsChurn detection; stalls and staleness often drive reconnect loopstotal_connections growing fast while connections is flat
/varz memStale connections and stalled buffers both hold memoryGrowth correlated with stale/stalled counts
OS file descriptor usage on nats-serverStale connections silently consume the FD budgetFD count approaching ulimit -n

Fixes

Fix the slow consumer behind the stalls

The stalled state is the server correctly applying backpressure; the fault is on the connection that cannot keep up. The durable fixes are on the consumer side: move blocking work (database writes, synchronous I/O) out of the message handler, scale the subscriber horizontally (queue groups for core NATS), or reduce fan-out onto the struggling consumer. For JetStream workloads, prefer pull consumers, which control their own consumption rate and are not subject to push-based slow consumer flagging.

Raising write_deadline is a common workaround for transient bursts. Understand the tradeoff: it postpones the stall and the eventual slow consumer disconnect, but it does not fix a consumer that is structurally slower than its message rate. It only trades earlier detection for larger in-memory backlogs.

Reap the stale connections and fix what created them

A connection that has failed ping/pong is closed by the server once detected, which releases its FD and buffers. If stale_connections keeps growing, the interesting question is why clients keep entering that state. For hung client processes, fix the client: the server cannot make a deadlocked process answer pings. For network paths, look for idle-flow timeouts on NAT devices, firewalls, or load balancers between client and server, and make sure the client library’s ping interval is comfortably shorter than any middlebox timeout.

Do not restart the NATS server as a first response to either counter. The server is reporting these states accurately; restarting destroys the evidence and usually re-creates the condition when the same clients reconnect.

Reduce reconnect churn

If stalls are driving disconnect-reconnect loops (fast-growing total_connections with a flat connections count), the churn itself adds CPU and memory pressure. Fixing the slow consumer stops the loop. If you need a temporary pressure valve, shedding the worst offending subscriptions or pausing a non-critical publisher is safer than anything server-side.

Prevention

  • Alert on rate, not value. Both counters are cumulative. Alert on a positive delta sustained over 5 minutes, and treat a reset to zero as a restart event worth correlating with unexpected uptime resets.
  • Monitor the precursor, not just the consequence. Track top-N pending_bytes via /connz?sort=pending alongside stalled_clients. Pending bytes rise before the stall counter moves, which is before the slow consumer disconnect.
  • Keep client ping intervals shorter than middlebox idle timeouts. Any NAT, VPN, or L4 proxy in the path with an idle timeout shorter than the protocol ping interval will manufacture stale connections.
  • Separate client from route/gateway distress. Write-path problems on routes and gateways have cluster-wide blast radius. Watch /routez pending_size and slow_consumer_stats so the breakdown is never a surprise.
  • Size FD headroom for the stale-connection case. Keep total connections well under both max_connections and ulimit -n, because stale connections consume slots while doing nothing.
  • Avoid middleboxes in the data path where possible. NATS clients use persistent TCP connections; L7 proxies and aggressive L4 idle reaping are recurring sources of half-dead sockets.

How Netdata helps

Netdata’s NATS collector polls the server’s HTTP monitoring endpoints and charts the signals around this problem, which shortens the correlation work:

  • Slow consumers, including the per-type breakdown (clients, routes, gateways, leafs), so you can see immediately whether write-path distress is a client issue or a cluster issue.
  • Connection count and churn (connections and total_connections rates), which makes disconnect-reconnect spirals visible next to the slow consumer events that cause them.
  • Throughput (in/out messages and bytes), so you can line up stall onset with fan-out spikes and check for out_msgs dropping relative to in_msgs.
  • Server memory and CPU, letting you confirm whether stale connections and stalled buffers are translating into real resource pressure.
  • Per-second granularity, which matters here because pending-buffer spikes and brief stall windows are easy to miss at 60-second scrape intervals.

One gap: Netdata’s NATS collector does not currently chart stalled_clients or stale_connections from /varz. Poll those two fields directly with the curl commands above, or wire them into your own scrape job, and use Netdata’s surrounding signals (slow consumers, churn, memory) to corroborate what you find.