The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nats / nats-how-it-works-in-production ▌

Operations Guides

How NATS actually works in production: a mental model for operators

Most NATS incidents come from a mismatch between how operators think NATS works and how it actually works. Teams coming from Kafka or RabbitMQ carry a queue-centric mental model: messages go somewhere, wait there, and can be retrieved later. Core NATS does not work like that. The mismatch shows up at 3 a.m. as silent message loss, slow consumer cascades, or a Raft election storm nobody saw coming.

This article builds the mental model you need before touching any NATS runbook: what the server is doing internally, where it competes for resources, and which behaviors are by design rather than symptoms.

What NATS is and why the model matters

NATS is a subject-based message router written in Go. That sentence carries more operational weight than it appears to. It is not a broker that stores messages. It is not a queue that holds work. It receives messages on subjects, matches them against registered subscriptions, and writes them to the interested connections, in real time.

The consequence that bites teams hardest: in core NATS, if no subscriber is connected when a message is published, the message is silently discarded. No error, no log, no metric. This is fire-and-forget by design. If your architecture assumes at-least-once delivery, that guarantee has to come from JetStream, not from core NATS.

This one fact re-frames most “NATS is losing messages” reports: the server is usually doing exactly what it was designed to do.

How the routing engine works

The core of the server is an in-memory subject tree, a trie that maps subject strings to sets of subscriptions. Subjects are dot-delimited tokens (for example orders.us.east.created), and the trie supports two wildcards: * matches a single token, > matches one or more trailing tokens. Every inbound message is matched against this trie and fanned out directly to all matching subscribers. There is no intermediate queue between publisher and subscriber.

flowchart LR
    P[Publisher] -->|publish on subject| T[Subject trie]
    T -->|match and fan out| S1[Subscriber A write buffer]
    T -->|match and fan out| S2[Subscriber B write buffer]
    T -->|route interest exists| R[Route to peer server]
    T -->|no match| X[Message dropped silently]
    R --> T2[Peer subject trie]
    T2 --> S3[Subscriber C write buffer]

Two properties of this design matter operationally:

  • Fan-out is amplification. One inbound message becomes N outbound writes. A high out_msgs / in_msgs ratio is normal for fan-out workloads, but it also means any latency in the delivery path is multiplied across subscribers.
  • The trie is shared state. Subscriptions consume memory in the trie and in per-connection tracking. A subscription leak in one client application inflates the routing table, and in a cluster the leak propagates across routes to every server.

Per-connection machinery: where slow consumers come from

Every TCP connection to the server, whether an application client, a cluster route, a gateway, or a leaf node, gets dedicated read and write goroutines plus a per-connection pending write buffer. The write side is governed by write_deadline, which controls how long the server waits for a connection to drain before giving up. The default is 10 seconds (DEFAULT_FLUSH_DEADLINE) and applies uniformly to clients, routes, gateways, and leaf nodes when the per-section override is not set.

When a connection cannot consume as fast as the server produces, the pending buffer grows. When it exceeds the threshold, the server marks the connection as a slow consumer and, by default, disconnects it. Two things about this mechanism are routinely misunderstood:

  1. A slow consumer event is the server correctly enforcing backpressure. The fault is on the slow connection: a subscriber doing synchronous I/O in its message handler, a GC pause, a congested network path. Investigating the server first wastes time.
  2. It applies to every connection type, not just clients. A slow consumer on a route or gateway means inter-server message delivery is backing up, with cluster-wide blast radius. The /varz slow_consumer_stats object exposes the breakdown by connection type (clients, routes, gateways, leafs), alongside the aggregate slow_consumers counter.

The leading indicator is per-connection pending bytes, visible via /connz (clients) and /routez (routes, as pending_size). Pending bytes grow before the server declares distress, so watching the precursor lets you act before the disconnect-reconnect spiral begins. That spiral is the classic slow consumer cascade: subscriber falls behind, server disconnects it, client auto-reconnects and resubscribes, the backlog hits immediately, and the cycle repeats as connection churn.

Clustering: routes, gateways, and leaf nodes

NATS scales out through three distinct connection types, and each one is, internally, just another client with its own goroutines and buffers:

MechanismTopologyInterest behaviorFailure mode to know
RoutesFull mesh between servers in one cluster (N servers means N-1 routes per server)Subscription interest propagates across the cluster so messages route only where subscribers existRoute slow consumers; route drops mean partition; reconnects trigger full interest re-propagation
GatewaysConnections between independent clusters (supercluster)Optimistic send at startup, then converges to interest-only mode per accountSlow consumers; a reconnected gateway can cause a temporary bandwidth spike; zero traffic on an idle gateway can be normal
Leaf nodesLightweight connection from an edge server to a hub clusterOne logical connection carrying multiplexed traffic for potentially many accountsA single leaf disconnect can cut many logical channels at once; the edge loses messaging until reconnect

Clustering does not change the core semantics. Messages still route in real time, there is still no intermediate queue, and every hop has its own buffer that can back up. A route count below N-1 in a cluster means a partition, and in a 2-node cluster losing the single route is a complete partition.

JetStream: what gets bolted on

JetStream is a persistence layer on top of core NATS, not a replacement for it. Enabling it adds subsystems that compete for the same host resources:

  • WAL-based storage. Each stream uses file or memory storage. File storage is a write-ahead log with block-level organization, plus an in-memory index mapping message sequence numbers to file block positions. JetStream disk I/O is sequential writes, historical reads, compaction, and snapshots.
  • Consumer state tracking. Delivery cursors, acknowledgment state, and redelivery timers, per consumer. With explicit acks, unacknowledged messages accumulate up to MaxAckPending; at that limit, delivery stalls silently. The stream looks healthy while the consumer is effectively down. The default is 1000 for explicit-ack consumers that do not set MaxAckPending themselves (the server applies this only when the ack policy is not none).
  • Raft consensus, at two levels. In clustered JetStream there is a meta Raft group for cluster-wide metadata, and each replicated stream (and its consumers) has its own Raft group. A per-stream group can fail independently while the meta group is healthy.
  • Background work. Compaction, retention enforcement, snapshot/restore, and catch-up replication for lagging peers.

Operationally, JetStream converts fire-and-forget into a system where storage, consumer lag, and Raft stability are the primary failure surfaces. Retention interacts with consumer health in ways that surprise people: with interest retention, a message is deleted only after all consumers acknowledge it, so one stalled consumer blocks retention for the whole stream and storage fills. Conversely, a stream with interest retention and zero consumers deletes every message on arrival: effectively /dev/null while looking perfectly healthy.

Where resource pressure shows up

NATS competes for a small set of host resources, each with a characteristic degradation shape:

ResourceWhat consumes itDegradation shape
File descriptorsOne per client connection, plus routes, gateways, leaf nodes, JetStream file handles, and listenersCliff-edge. At the OS ulimit -n, accept() fails and no new connections are possible. A default ulimit of 1024 is catastrophically low
MemorySubject trie, per-connection buffers (tens of KB each, more with pending data), JetStream caches, Raft stateSawtooth from Go GC is normal; monotonic growth without GC recovery is not. Ends in an OOM kill
CPUSubject matching per message, TLS handshakes, JetStream indexing and compaction, RaftTLS dominates CPU on workloads that are otherwise cheap per message; GC pauses show as spikes and can trigger Raft election timeouts
Disk I/OJetStream only: WAL writes, reads, compactionLatency-sensitive. Slow Raft log appends cause heartbeat delays and elections. Network-attached storage with variable latency is the top cause of Raft election storms
NetworkFan-out amplification, route traffic, gateway replicationGradual. TCP backpressure becomes delivery latency becomes slow consumer events

The interplay worth internalizing: slow consumers buffer in memory, memory pressure lengthens GC pauses, GC pauses delay Raft heartbeats, and delayed heartbeats cause elections that pause writes. A JetStream outage can have its root cause three steps away from where the alert fired.

The failure archetypes to recognize

Recognizing the shape early is most of the battle:

  • Slow consumer cascade. Disconnect-reconnect churn on clients, routes, or gateways. Check the slow consumer breakdown by connection type first; the remediation for a slow route is completely different from a slow client.
  • File descriptor exhaustion. Sudden inability to accept new connections while existing ones work fine. Check ulimit -n before anything else; it is a cliff-edge with no graceful degradation.
  • JetStream Raft instability. Network latency, disk I/O stalls, or GC pauses cause election timeouts. Leader flapping means write failures and stalled consumers, and simultaneous elections across many stream groups can make the cluster effectively unavailable.
  • Memory pressure. Slow consumer buffering, subscription trie growth, JetStream cache growth. Alert on the trend, not spikes.
  • Silent message loss in core NATS. Publishers send to subjects with zero subscribers and messages vanish. The only signal is out_msgs running far below in_msgs after adjusting for expected fan-out.

Signals to watch in production

These signals map directly onto the machinery above. The monitoring port (default 8222) must be explicitly enabled with -m 8222 or http_port: 8222; if it is not, none of these endpoints exist.

SignalWhy it mattersWarning sign
/varz slow_consumersThe single most important core NATS signal; break it down by connection type to size the blast radiusAny positive rate of change; slow consumers on routes or gateways are urgent
/connz?sort=pending pending_bytes and /routez pending_sizeLeading indicator before slow consumer disconnectionsSustained growth on any connection; any sustained non-zero pending on routes
/varz in_msgs vs out_msgsFan-out ratio and the only detection for zero-subscriber lossout_msgs flat or far below in_msgs against expected fan-out
/varz connections vs max_connections, plus OS fd countConnection saturation is cliff-edgeAbove roughly 85% of max_connections, or total FDs near ulimit
/varz routes vs cluster size minus onePartition detectionAny route missing for more than a minute; route flapping
/varz mem trendBuffering, trie growth, JetStream cache pressureMonotonic growth over hours without GC recovery
/jsz api.errors and api.inflightJetStream write-path healthSustained error rate; high inflight points at Raft or disk I/O
/jsz?consumers=true num_pending and num_ack_pendingWhether consumers actually keep up; ack_pending at MaxAckPending means stalled deliverySustained growth in either
/jsz meta_cluster leader and replica current/offline/lagMeta Raft stabilityFrequent leader changes; any offline or non-current peer

One instrumentation caution: /subsz is expensive on servers with many subscriptions and can degrade the server itself. Track the subscription count from /varz, not the list from /subsz.

How Netdata helps

The mental model above tells you what to correlate; the value of monitoring is seeing those correlations on one timeline instead of assembling them from curl during an incident:

  • Slow consumer events against connection churn and pending bytes, so you can tell a one-off client hiccup from an active cascade, and see backpressure building before disconnects start.
  • in_msgs against out_msgs over time, which makes zero-subscriber message loss visible as a fan-out asymmetry instead of an application-side mystery.
  • Route count and route health alongside throughput, so a cluster partition shows up immediately rather than as unexplained delivery gaps.
  • Process RSS and CPU next to connection and subscription counts, which distinguishes expected growth from buffer accumulation or a subscription leak.
  • JetStream aggregates (storage, API errors) against OS-level disk latency and iowait, the correlation that separates storage exhaustion from a disk I/O stall, two problems with very different fixes.