The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nats / nats-gateway-disconnected ▌

Operations Guides

NATS gateway disconnected: cross-cluster traffic cut in a supercluster

Subscribers in cluster A have stopped receiving messages published in cluster B. Publishers see no errors. Every server’s /healthz returns ok, CPU and memory look normal, and yet a whole class of traffic has silently stopped flowing between two sites. In a NATS supercluster, this is the signature of a gateway disconnection: the outbound gateway connection to the remote cluster is missing, so no cross-cluster forwarding happens at all.

Gateways are the links between independent NATS clusters. Unlike cluster routes (full mesh between servers in one cluster), each server maintains gateway connections keyed by remote cluster name, and all cross-cluster delivery for an account flows through them. When one goes down, the failure is clean and quiet: local delivery keeps working, remote delivery stops, and nothing in the server-wide error counters necessarily moves.

This guide covers how to confirm a gateway is actually disconnected (versus merely idle), what typically cuts the connection, and how to restore and protect cross-cluster traffic.

What this means

Each NATS server in a supercluster tracks its gateway connections in two maps on the /gatewayz monitoring endpoint: outbound_gateways and inbound_gateways, both keyed by remote cluster name. Outbound is what your cluster initiated; inbound is what the remote cluster initiated toward you. Every configured remote cluster should appear. A configured gateway missing from outbound_gateways means cross-cluster delivery to that cluster is failing.

Two things make gateway incidents confusing:

  • Zero traffic is not a symptom. Gateways run in interest-only mode: messages for a subject only cross the gateway when the remote cluster has registered interest. If no client in cluster B subscribes to a subject, cluster A correctly sends nothing. Flat gateway traffic counters can be completely normal.
  • The server looks healthy. /healthz passes, local pub/sub works, and JetStream (if present) keeps serving local consumers. The only broken thing is the path between clusters.

After a gateway reconnects, there is a brief flood phase before interest-only mode re-converges, during which more messages are forwarded than strictly needed. Expect a temporary bandwidth spike on recovery. A gateway stuck out of interest-only convergence for a long time is itself a warning sign.

Common causes

CauseWhat it looks likeFirst thing to check
Cross-datacenter network partitionOutbound gateway missing to one remote cluster; other gateways fine; flow logs between sites show dropsNetwork reachability from this server to the remote cluster’s gateway port
Remote cluster down or restartingOutbound missing, inbound also empty; remote cluster’s servers unreachable on monitoring portRemote cluster /healthz and process uptime on its servers
Gateway URL misconfigurationGateway never connects after config change or deploy; repeated connect attempts in the log with no successful registrationThe gateways block: URLs, ports, and that all servers in a cluster share the same gateway name
Gateway TLS misconfigurationConnection attempts fail at handshake; TLS errors on the gateway port in the logCertificate validity, CA bundle on both sides, whether one side rotated certs before the other
Version-specific stuck reconnect (older servers)Outbound connection exists in a half-formed state or never re-registers after packet lossServer version; see the fixes section for the affected range

Quick checks

All read-only, run against the local monitoring port (default 8222):

# Which remote clusters does this server see, outbound and inbound?
curl -s http://localhost:8222/gatewayz | jq '{outbound: (.outbound_gateways | keys), inbound: (.inbound_gateways | keys)}'

# Detail per outbound gateway: is the configured remote cluster present and connected?
curl -s http://localhost:8222/gatewayz | jq '.outbound_gateways'

# Is this server itself healthy and not freshly restarted?
curl -s http://localhost:8222/varz | jq '{uptime, connections, routes}'

# Throughput: did out_msgs drop while in_msgs continues? Cross-cluster fan-out gone.
curl -s http://localhost:8222/varz | jq '{in_msgs, out_msgs}'

# Gateway slow consumer events (a gateway under backpressure can precede a drop)
curl -s http://localhost:8222/varz | jq '{slow_consumers, slow_consumer_stats}'

Interpretation notes:

  • Compare .outbound_gateways | keys against the remote clusters in your gateways configuration. The sets should match. A missing configured gateway is the incident.
  • A gateway that is present but whose traffic counters are flat is not proof of a problem. Check whether the remote cluster actually has interest (subscriptions) for the subjects you expect to flow.
  • slow_consumer_stats breaks slow consumers down by connection type. A non-zero gateway count means the gateway connection itself was falling behind, which has a much bigger blast radius than a slow client and often precedes or accompanies gateway instability. The per-type breakdown is in /varz and is available since nats-server v2.10.0.

How to diagnose it

flowchart TD
  A[Cross-cluster delivery stopped] --> B{Remote cluster in outbound_gateways?}
  B -- Yes, present --> C{Traffic flat?}
  C -- Yes --> D[Check remote interest. Likely normal: no subscribers]
  C -- No --> E[Gateway up. Look at remote consumers or subject mismatch]
  B -- Missing --> F{Remote cluster reachable?}
  F -- No --> G[Network partition or remote cluster down]
  F -- Yes --> H{TLS or config errors in log?}
  H -- Yes --> I[Fix certs / CA bundle / gateway URLs]
  H -- No --> J[Stuck reconnect: check server version]
  1. Confirm the gap. Run the /gatewayz key listing above on more than one server in the local cluster. If the remote cluster name is absent from outbound_gateways everywhere, the gateway link is down. If it is absent on one server only, scope the problem to that server.
  2. Rule out the zero-traffic false alarm. If the gateway entry exists but counters are flat, verify that a subscriber actually exists on the remote side for the subject in question. Interest-only mode means no interest, no traffic. Many “gateway down” reports end here.
  3. Check reachability. From the local server, test TCP connectivity to the remote cluster’s gateway listen address and port. Failure here points to a cross-DC partition, firewall change, or the remote cluster being down. Also check the remote cluster’s own /healthz and uptime: a remote cluster mid-restart or crash-looping explains a missing gateway with no local fault.
  4. Read the server log around the drop. Look for gateway connection attempts without a corresponding successful registration, and for TLS handshake failures on the gateway port. TLS failures after a certificate change on either side are a classic cause: if the remote cluster rotated its certificate and your CA bundle only trusts the old CA, the connection is severed at handshake.
  5. Verify gateway configuration symmetry. Every server in a cluster must use the same gateway name, and gateway URLs must be reachable from every gateway node in both directions. A name mismatch or an advertised address that peers cannot route to (for example, a pod-internal IP advertised instead of the reachable external address) will keep connections from establishing.
  6. Check the server version for the stuck-reconnect bug. On nats-server v2.10.14 and earlier, an outbound gateway connection could get stuck indefinitely during packet loss: the PING timers that detect a dead connection only started after the first INFO response from the remote, so a lost handshake left the connection hanging forever. The fix shipped in v2.11.0. If you are on an affected version and the gateway is stuck rather than cleanly refused, this is likely your cause.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
/gatewayz outbound/inbound keysDirect view of which remote clusters are connectedConfigured remote cluster missing for more than 60s
Gateway traffic counters (in_msgs/out_msgs per gateway)Shows whether cross-cluster flow matches interestFlat when remote interest is known to exist
Interest-only convergence stateGateways should settle into interest-only mode after connectStuck out of interest-only mode, or repeated flood phases (reconnect flapping)
slow_consumer_stats.gatewaysGateway-level backpressure, high blast radiusAny sustained increase
out_msgs vs in_msgs (server-wide)Cross-cluster fan-out disappearing lowers out_msgsout_msgs drops while in_msgs is steady
Server uptime across the superclusterSimultaneous resets indicate a shared event (network, cert rotation, deploy)Correlated restarts in both clusters
TLS certificate expiry (tls_cert_not_after in /varz where exposed, or external cert checks)Expired or mistrusted certs sever gateway handshakesUnder 30 days, or asymmetric rotation across clusters

Fixes

Network partition or remote cluster down

Restore connectivity or bring the remote cluster back; the gateway reconnects on its own once the path is up. Do not restart local servers as a first move: it does nothing for a partition and adds a client reconnect storm on top. If the partition is recurrent, treat it as a network engineering problem (flow logs, inter-DC link health) rather than a NATS problem.

Remote cluster mid-restart

Wait for the remote cluster to finish recovering. Note the flood phase after reconnection: expect a short bandwidth spike while interest re-converges, and do not mistake it for a leak.

Gateway URL or name misconfiguration

Fix the gateways block so URLs point at reachable gateway listen addresses, and ensure the gateway name is identical on every server in the cluster. In Kubernetes deployments, make sure the address each server advertises for gateway connections is the address peers can actually reach (a load balancer or external IP), not a pod-internal IP. Operators have hit a failure mode where a configured or discovered gateway URL contains a pod-internal IP, so peers cannot reach it and reconnect attempts keep failing. Setting an explicit reachable advertise address and doing a rolling restart of all clusters resolved it. In current nats-server releases, gateway { advertise } overrides the host and port announced in gateway INFO; the soliciting server tries every URL in a remote gateway’s list in one attempt before waiting one second for the next attempt.

TLS misconfiguration

Align the CA bundle on both sides so each cluster trusts the other’s current certificate. When rotating certificates across a supercluster, use a bundle containing both the old and new CAs and rotate one cluster at a time; rotating the remote side first with a CA bundle that only trusts the old CA severs the gateway immediately. Track expiry in advance (see the monitoring table) so this never becomes the incident cause.

Stuck outbound connection on v2.10.14 or earlier

Upgrade to v2.11.0 or later, where gateway PING timers start before the first INFO response and dead half-open connections are detected and retried. As an interim measure on an affected version, restarting the stuck server forces a fresh connection attempt, but treat that as a workaround, not the fix, and schedule the upgrade. If reconnect storms after recovery are a concern, nats-server v2.12.0 and later support gateway { connect_backoff: true } for exponential connection-attempt backoff; the option applies to both explicit and discovered gateway entries. gateway { connect_retries }, however, limits only implicit gateways.

Prevention

  • Alert on topology, not traffic. Page-worthy condition: a configured remote cluster missing from /gatewayz outbound connections, sustained more than 60 seconds. Do not alert on zero gateway traffic; that is interest-only mode working as designed.
  • Monitor gateway slow consumers separately. slow_consumer_stats.gateways should be zero. Any rate of change means cross-cluster delivery is backing up before it breaks.
  • Coordinate certificate rotation across clusters. One rotation calendar, shared CA bundles during transitions, and expiry alerting well before 30 days.
  • Keep servers current. The stuck-outbound-gateway bug class is fixed in maintained versions; running old patch releases across a supercluster is an avoidable risk.
  • Test the failure. Periodically verify, in staging, what your alerting does when a gateway drops and that dashboards distinguish “gateway down” from “no cross-cluster interest.”

How Netdata helps

  • Netdata polls the NATS monitoring endpoints and charts in_msgs/out_msgs rates per server, so the moment cross-cluster fan-out disappears (out_msgs dropping while in_msgs holds) shows up as a visible divergence across both clusters’ dashboards.
  • Slow consumer counters, including the per-type breakdown, are collected over time, letting you see gateway-level backpressure building before a disconnection rather than after.
  • Uptime tracking across all supercluster nodes makes correlated restarts obvious, separating a shared event (cert rotation, network maintenance) from a single-node fault.
  • Because Netdata collects at per-second granularity, the brief flood-phase bandwidth spike after a gateway reconnect is distinguishable from a sustained traffic anomaly, reducing false alarms during recovery.
  • Correlating NATS signals with host network metrics (interface errors, retransmits, link drops on the inter-DC path) shortens the “is it NATS or is it the network” loop that gateway incidents always trigger.