The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nats / nats-monitoring-maturity-model ▌

Operations Guides

NATS monitoring maturity model: from survival to expert

Most NATS outages are not caused by a lack of metrics. The server exposes a rich monitoring API on port 8222. The failures happen because teams collect the wrong tier of signals for the failures they actually experience. A /healthz check tells you the process is alive. It tells you nothing about a consumer stalled at MaxAckPending, a route connection backing up with pending bytes, or a Raft meta cluster electing a new leader every ninety seconds.

This article lays out a four-level maturity model for NATS monitoring: Survival, Operational, Mature, and Expert. Each level assumes the previous one is solid. The goal is not to collect everything immediately. It is to know which level you are at, which level your reliability requirements demand, and which specific gaps explain your last incident.

Use this as a self-assessment. For each level, every signal listed is something you can collect today from the NATS HTTP monitoring endpoints, server logs, or the operating system. No signal here requires vendor tooling.

flowchart TD
  L1["Level 1 - Survival: is the process alive?"]
  L2["Level 2 - Operational: is traffic flowing correctly?"]
  L3["Level 3 - Mature: are internals degrading before failure?"]
  L4["Level 4 - Expert: what will break next week?"]
  L1 --> L2 --> L3 --> L4

Level 1: Survival

Survival monitoring answers one question: is the server process alive and minimally functional. Every production NATS deployment needs this from day one, including dev and staging environments. If you are below this level, you are flying blind.

The signals:

  • Health probe. curl -s http://localhost:8222/healthz?js-server-only=true. The query parameter matters. Bare /healthz on a JetStream-enabled server performs full JetStream health checks, including asset recovery, and will fail for minutes after a restart on large stores. Using bare /healthz as a page-level probe causes false pages during normal recovery. Use js-server-only=true for paging and bare /healthz as a ticket-level signal.
  • Uptime. /varz -> uptime. An unexpected reset means a crash or forced restart. More than three restarts in 30 minutes is a crash loop.
  • Active connections. /varz -> connections. If clients are expected and this is zero, something upstream is broken.
  • Memory. /varz -> mem (RSS). Monotonic growth over hours without GC recovery points at a leak or unbounded buffering. NATS is Go, so a sawtooth pattern is normal; a rising floor is not.
  • Slow consumers. /varz -> slow_consumers. Any positive rate of change means the server is disconnecting or dropping for clients that cannot keep up. In core NATS, each event can mean dropped messages.

That is five signals. They fit in any alerting system. The two classic mistakes at this level are using bare /healthz for paging and alerting on absolute connection counts instead of ratios.

Level 2: Operational

Operational monitoring answers: is traffic flowing, and is it flowing to the right places. This is the level a competent production team should reach within the first month of running NATS seriously.

Everything in Level 1, plus:

  • Message and byte throughput. /varz -> in_msgs, out_msgs, in_bytes, out_bytes. All four are cumulative counters, so compute rates. Two derived views matter most. The fan-out ratio out_msgs / in_msgs tells you how many subscribers receive each message. The asymmetry between in_msgs rising and out_msgs flat is the only signal you get for zero-subscriber message loss in core NATS: messages published to a subject with no subscribers are silently dropped, with no error, no log, and no dedicated metric.
  • Connections versus limit. /varz -> connections against max_connections (default 65536). Alert as a ratio, above 85%, not as an absolute number. Separately, remember the OS file descriptor limit is a different wall, and it is often hit first. Default ulimit -n of 1024 is catastrophically low for production NATS.
  • Connection churn. The delta on total_connections relative to a stable connections count. High churn with a flat active count means clients are flapping, and each cycle costs CPU, auth work, and often slow consumer events.
  • Route health. /varz -> routes. In a full-mesh cluster of N servers, each server should have N-1 routes. Gate the alert on expected_routes > 0 and current below expected for more than 60 seconds. Gating on routes > 0 misses the worst case, which is total route loss.
  • JetStream status. /jsz -> disabled. If JetStream should be enabled and this reports true, persistence is down. Gate on uptime over 600 seconds and sustain for five minutes to avoid cold-start false positives.
  • JetStream API errors. /jsz -> api.errors against api.total. Alert on a sustained positive error rate or an error ratio above roughly 5%. Note that idempotent create operations generate benign errors, so correlate with logs before declaring a server fault.
  • Critical consumer health. For your small set of known-critical durable consumers, track num_pending and num_ack_pending from /jsz?consumers=true or nats consumer info. A consumer with num_ack_pending pinned at MaxAckPending has stalled delivery completely, and no server-level metric will show it. This is the most commonly skipped signal in all of NATS operations.
  • Auth failure rate. From server logs (Authorization Violation, Authentication Timeout). A spike means either a credential rotation problem or probing.

A practical test for whether you are at Level 2: during your last incident, could you answer “are messages being silently dropped” and “is any critical consumer stalled” from your dashboards alone, in under a minute.

Level 3: Mature

Mature monitoring adds the internals and the leading indicators. The defining shift at this level is from lagging signals (slow consumer events, error counters) to precursor signals (pending bytes, lag trends, election rates). Mature teams get paged before users notice, not after.

Everything in Level 2, plus:

  • Per-connection pending bytes. /connz?sort=pending exposes pending_bytes per client; /routez exposes pending_size per route. Pending bytes grow before the server declares a connection slow. Monitoring the precursor lets you identify the specific client or route that is backing up before disconnection cascades. Route pending bytes are especially dangerous: a route slow consumer means inter-server delivery is failing, with cluster-wide blast radius. On high-connection-count servers, scraping all connections is expensive, so sample the top N by pending instead.
  • Client and route RTT. /connz and /routez expose rtt per connection. Sustained route RTT increase is an early warning for cluster instability and, in JetStream clusters, for Raft election trouble.
  • Gateway and leaf node health. /gatewayz for superclusters, /leafz for edge topologies. A missing configured gateway is a partition between clusters. Leaf drops isolate edge locations. Zero gateway traffic can be normal once interest-only mode converges, so alert on connection existence and convergence state, not traffic.
  • Raft leader distribution and elections. /jsz -> meta_cluster.leader and meta_cluster.replicas[] with current, offline, and lag. Leader changes more than about once per hour, peers stuck current=false or offline=true, and skewed leader distribution across nodes are all leading indicators of JetStream write failures. Per-stream Raft groups are separate from the meta group and can fail independently; /raftz exposes them.
  • Stream replica state. For replicated streams, replica lag and offline fields tell you whether failover would lose data. A replica perpetually one or two messages behind looks healthy but means every failover loses the most recent writes.
  • JetStream storage versus limits. /jsz -> storage, memory, reserved_storage, reserved_memory. Alert above 80 to 90 percent and project time to exhaustion from the growth trend. Behavior at the limit depends on retention and discard policy: DiscardOld silently evicts data, DiscardNew rejects publishes (which surfaces in api.errors).
  • Subscription count trend. /varz -> subscriptions. Use the count, never the full /subsz list, which can lock a server with millions of subscriptions. Growth without matching connection growth is a subscription leak, and in a cluster it bloats the subject trie on every node.
  • Stale connections and stalled clients. /varz -> stale_connections and stalled_clients. Half-dead clients hold file descriptors and memory; stalled clients are the write-path precursor to slow consumer events.
  • TLS certificate expiry. /varz -> tls_cert_not_after. Alert at 30 days, escalate at 7. Expiry kills every TLS client, route, gateway, and leaf connection at once. The field is available in /varz since nats-server v2.14.0; older versions do not expose it, so fall back to checking the certificate files with openssl.

Level 4: Expert

Expert monitoring is what teams build after their third major incident, when they realize the interesting failures were invisible at Level 3. These signals are about prediction and about closing the gaps where the server looks healthy but the system is broken.

Everything in Level 3, plus:

  • Connection churn analysis. Correlate slow consumer events with reconnect storms. The classic death spiral is: subscriber falls behind, server disconnects it, client auto-reconnects, backlog hits immediately, disconnect again. Churn metrics plus the slow_consumer_stats breakdown (clients versus routes versus gateways versus leafs) tell you whether you are in that spiral and how wide the blast radius is.
  • Fan-out ratio trend. out_msgs / in_msgs per unit time. A shift in this ratio means the subscriber population changed. If you expect three subscribers and the ratio drops to one, two are silently gone.
  • JetStream WAL fsync latency. Not directly exposed by NATS. Infer it from OS-level disk latency on the JetStream storage path (iostat, iowait). This is the single strongest predictor of Raft instability: slow WAL writes delay heartbeats, delayed heartbeats trigger elections. Network-attached storage with variable latency is the number one cause of Raft election storms.
  • Go GC pause time. Via pprof or the metrics endpoint. Pauses above roughly 10ms on a JetStream leader can trigger Raft election timeouts. Correlate GC pauses with election events before blaming the network.
  • JetStream API inflight. /jsz -> api.inflight. Sustained high inflight plus rising API errors points at Raft consensus delay or disk saturation, distinct from storage exhaustion.
  • The $SYS event stream. Subscribing to $SYS.> gives real-time server advisories: connects, disconnects, slow consumers, auth violations. This is the richest signal source NATS offers, but it requires a dedicated consumer and processing pipeline, and access to the system account must be tightly controlled.
  • Consumer high-water marks. Track how close each consumer regularly gets to MaxAckPending, not just the current value. A consumer routinely at 90 percent is one slow processing cycle away from a full stall.
  • Monitoring self-impact. /varz -> http_req_stats shows scrape rates per endpoint. Aggressive scraping of expensive endpoints (/subsz, /connz with subscription detail, /jsz?consumers=true on large deployments) can degrade the server you are trying to observe. Keep poll intervals at 10 seconds or more, and remember the endpoints are point-in-time snapshots: sub-second anomalies between scrapes are invisible.

Common gaps at every level

A few patterns repeat across teams regardless of level. Monitoring “is it up” but not “is it working”: a server can pass /healthz while JetStream is disabled, consumers are stalled, or all subscribers disconnected. Ignoring the in_msgs/out_msgs asymmetry, which is the only zero-subscriber loss signal in core NATS. Treating slow consumer events as a server problem when the server is correctly enforcing backpressure and the fault is on the consuming side. Monitoring aggregate JetStream storage but not consumer lag: a stream with 100 million pending messages is effectively down while every storage metric looks fine. And using absolute thresholds that break across deployment sizes instead of ratios against configured limits.

How Netdata helps

The maturity model maps onto signal collection, and the level you can operate at is bounded by what your collector actually gathers:

  • Netdata’s NATS collector polls the HTTP monitoring endpoints and charts the Level 1 and Level 2 signals per server: health, uptime, connections versus max_connections, message and byte rates, slow consumers, memory, CPU, and JetStream aggregates including api.errors and storage.
  • Per-second collection granularity catches the transient pending-buffer spikes and churn bursts that 10-15 second scrapes miss between snapshots.
  • Correlating NATS charts with host-level disk latency and iowait in the same dashboard is how you operationalize the Expert-level WAL fsync signal: JetStream API inflight rising alongside disk latency confirms an I/O stall rather than Raft misbehavior.
  • Uptime resets, connection drops, and slow consumer rate changes lined up on one timeline make the slow consumer death spiral recognizable in minutes instead of after postmortem log digging.
  • Known gaps to cover with other tooling: per-consumer lag, per-connection pending bytes, Raft leader fields, TLS expiry, and stale/stalled connection counts are not currently exposed as Netdata NATS metrics, so pair Netdata with log alerts or targeted scripts for Levels 3 and 4.

This is currently the only guide in the NATS section. The NATS guides hub at NATS operations guides will collect the troubleshooting and deep-dive articles as they are published.