The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nats / nats-no-responders-available ▌

Operations Guides

NATS no responders available for request: request-reply into the void

Your service just started throwing nats: no responders available for request. The client got an answer back from the server almost instantly, and the answer was: nobody is listening on that subject. This is the request-reply counterpart to NATS’s silent message loss: instead of the request vanishing and the client waiting out its timeout, the server short-circuits the call and fails fast with a 503 status in the reply headers.

The fast failure is a feature. Core NATS is fire-and-forget: a message published to a subject with zero matching subscriptions is dropped with no error, no log line, and no dedicated metric. The no-responders mechanism exists so that request-reply callers at least find out. The server is telling you something specific: at the moment your request was routed, the subject had no subscribers. Your job is to figure out why.

The most important distinction up front: no-responders means the responder is absent. A timeout means the responder is present but did not reply in time. The fixes are completely different, so do not conflate them.

What this means

When a NATS client issues a request, it publishes a message on the target subject with an inbox reply subject, then waits for a response. On NATS Server 2.2.0 and later, with a client library that supports headers, the server checks the subject tree when the request arrives. If no subscription matches, the server immediately publishes an empty reply with a Status: 503 header. The client library surfaces this as a typed error: nats.ErrNoResponders in Go, NATSNoRespondersException in C#, NoRespondersError in Python, and equivalents in the other major clients.

Before 2.2.0, or with a client that does not negotiate header support, the same condition blocks until the client’s request timeout fires, and you get a timeout error indistinguishable from a slow responder. If you are seeing timeouts rather than explicit no-responders errors, check server and client versions before assuming the responder exists.

Because the check runs against the subject tree at routing time, no-responders is a statement about subscription interest, not about whether your responder process is alive. A responder that crashed, was never started, subscribed to a different subject, or sits behind a broken route all produce the same error.

flowchart TD
  A[Request published on subject S] --> B{Any subscription matches S?}
  B -- No --> C[Server replies with Status 503 header]
  C --> D[Client raises no-responders error immediately]
  B -- Yes --> E{Responder replies before client timeout?}
  E -- Yes --> F[Normal response]
  E -- No --> G[Client raises timeout error]
  D --> H[Responder absent: check process, subject spelling, routes]
  G --> I[Responder slow or hung: check its processing path]

Common causes

CauseWhat it looks likeFirst thing to check
Responder crashed or never startedNo-responders on every request; subscription count lower than expectedIs the responder process running and connected? Check /varz subscriptions
Subject typo or case mismatchNo-responders from one caller only; other services fineCompare subject strings character by character; subjects are case-sensitive and foo..bar matches nothing
Responder on a partitioned cluster nodeRequests from some servers get 503, others succeedRoute count on each server: /varz routes field vs expected N-1
JetStream API not ready (clustered)Synchronous Publish() fails with no-responders during leader elections/jsz meta_cluster leader stability, api.errors rate
JetStream not enabled at allnats stream ls or a publish returns no-responders instead of a clear error/jsz disabled field, or healthz?js-enabled-only=true
Cross-account service importClient gets a timeout, NOT no-responders, even though the service is downWhether the subject is reached via an account import (see below)
New JetStream cluster anomalyPersistent no-responders on KV writes or stream publishes on a freshly created clusterKnown open issue on some 2.10.x versions; see JetStream section below

One gotcha worth internalizing: when a service is reached through an account export/import, the server creates an internal subscription for the import. From the routing engine’s perspective there is always interest on the subject, so the no-responders shortcut never fires. Cross-account service calls can only fail by timeout. If your architecture uses account imports for services, you lose the fast-fail signal entirely and must treat timeouts as your absence detector; NATS maintainers describe service imports as always-on from the owning account’s perspective.

Quick checks

# Is the responder connected and subscribed? List connections with their subscriptions.
# Empty subscriptions_list across all connections means nobody is subscribed.
curl -s "http://localhost:8222/connz?subs=1" | jq '.connections[] | {cid, name, ip, subscriptions_list}'

# Total subscription count. Lower than expected for your deployment?
curl -s http://localhost:8222/varz | jq .subscriptions

# Cluster routes: in an N-node cluster each server should have N-1.
# A missing route means interest from one node is invisible to another.
curl -s http://localhost:8222/varz | jq .routes

# Message flow asymmetry: in_msgs rising while out_msgs stays flat means
# publishes are landing on subjects with no subscribers.
curl -s http://localhost:8222/varz | jq '{in_msgs, out_msgs}'

# JetStream: is it enabled and is the API answering?
curl -s http://localhost:8222/jsz | jq '{disabled, api_total: .api.total, api_errors: .api.errors}'

# JetStream meta cluster: is the leader stable? Frequent leader changes
# correlate with transient no-responders on JetStream publish.
curl -s http://localhost:8222/jsz | jq '.meta_cluster | {leader, replicas: [.replicas[]? | {name, current, offline}]}'

# Server health (basic readiness, avoids JetStream recovery false positives):
curl -s http://localhost:8222/healthz?js-server-only=true

All of these are read-only HTTP GETs against the monitoring port (default 8222, enabled with -m 8222 or http_port). Avoid /subsz on servers with very high subscription counts; it can be expensive. The subscription count from /varz is the safe check.

How to diagnose it

  1. Confirm the error type. Look at the actual client error. ErrNoResponders / 503 means absent responder. A timeout error means present but unresponsive responder. If you cannot tell from logs, reproduce with a minimal client call against the same subject.

  2. Verify the subject string. Subjects are case-sensitive, token-delimited by dots, and a double dot (foo..bar) creates an empty token that no normal subscription matches. Compare the publisher’s subject against the responder’s subscribe call character by character. This is the most common cause in practice, especially after config refactors or environment variable changes that assemble subjects from parts.

  3. Check whether the responder is connected at all. Use the /connz?subs=1 query above. If the responder’s connection is absent, the problem is the responder application: crashed, stuck in a restart loop, failing authentication, or connecting to a different server URL than you think. If connections are churning (compare total_connections growth against the stable connections count in /varz), the responder may be flapping: subscribing, dying, reconnecting.

  4. If clustered, check routes. Interest propagates across routes. If the responder is connected to server B and the requester to server A, and the A-B route is down, server A sees zero interest and returns 503. Compare /varz routes on each node against the expected N-1. Brief route drops during rolling restarts are normal; sustained missing routes mean partition.

  5. If the failing call is a JetStream publish, check JetStream separately. The synchronous Publish() in the Go client is implemented as an internal request to the JetStream API subject. During meta-group or stream leader elections in clustered JetStream, that API is briefly without a responder, and publishes fail with no-responders. The Go client retries this path (250 ms wait, 2 retries by default, tunable via RetryWait and RetryAttempts). If you see bursts of no-responders that self-resolve in under a second and correlate with leader changes in /jsz, this is your cause.

  6. If the call crosses accounts, stop expecting 503. A service import masks absence. Diagnose with a direct subscription check in the exporting account instead, and treat client timeouts as the failure signal.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
in_msgs vs out_msgs ratio (/varz)The only server-level signal for zero-subscriber loss in core NATSin_msgs rising, out_msgs flat or fan-out ratio below expected
subscriptions (/varz)Drops when responders disconnectCount below expected baseline for your deployment
Connection churn (total_connections delta)Flapping responders produce intermittent no-respondersStable connections with fast-growing total_connections
routes (/varz)Missing route hides remote interest, producing 503sCurrent routes < N-1 sustained > 60 s
api.errors and leader changes (/jsz)JetStream no-responders track elections and API failuresError rate rising, meta leader changing more than once per 5 minutes
Client-side no-responders error rateThe actual symptom, per serviceAny sustained rate; bursts correlated with deploys point at responder startup ordering

Fixes

Responder down or not started

Start or fix the responder. Then address the ordering problem: if requesters can start before responders are subscribed, add readiness gating in the requester (retry with backoff on ErrNoResponders, or block startup on a probe request). Client libraries do not retry request-reply for you on this error; the retry policy is application code.

Subject mismatch

Fix the string, then remove the class of bug: define subjects once in a shared constant or config value consumed by both publisher and subscriber. Validate assembled subjects at startup (non-empty, no empty tokens). If multiple environments share a cluster, include the environment prefix in a single place rather than string-concatenating at each call site.

Cluster partition

Restore the route: fix the network path, firewall rule, or DNS entry between the two servers. Verify with /routez that num_routes returns to N-1 and that per-route pending_size drains. Clients with multiple server URLs will fail over, but interest-based routing still requires the full mesh for cross-server delivery.

Transient JetStream no-responders during elections

If bursts correlate with leader changes, the real problem is Raft instability, not the publish path. Check route RTT, CPU, and disk latency on JetStream nodes; elections triggered by slow heartbeats will keep producing these blips until the underlying resource issue is fixed. Tuning the client’s RetryWait/RetryAttempts upward buys tolerance but does not fix the elections.

A similar 2.10.17 report involved a freshly created 3-node JetStream cluster with persistent no-responders errors on KeyValue and stream publishes; it was closed in 2024 after the operators stopped changing stream replica counts and deleting streams during startup, not by a 2.11 server fix. Check for concurrent stream creation, replica changes, and deletions before treating it as a server defect.

Cross-account service imports

You cannot restore the 503 behavior; the import subscription prevents it by design. Set client request timeouts deliberately (not the library default) so absence fails in bounded time, and monitor the exporting side’s subscription count as your absence signal.

Prevention

  • Shared subject definitions. One source of truth for every request-reply subject, imported by both sides. Most no-responders incidents in steady state are typos or drift.
  • Startup ordering and retries. Treat ErrNoResponders as retryable with capped backoff during deploys, fatal after the deploy window. Treat timeouts as a separate signal (responder slow), not as “service missing”.
  • Monitor the asymmetry. Alert on out_msgs falling below the expected fan-out of in_msgs. This catches the silent-loss version of the same failure before any requester notices.
  • Alert on subscription count drops for subjects that must always have a responder.
  • Watch routes and Raft leader stability so partitions and election storms are caught before they surface as application-level 503s.
  • Decide where you need persistence. If a request must not be lost when the responder is down, core request-reply is the wrong primitive regardless of the 503 feature. Use JetStream work queues and monitor consumer lag.

How Netdata helps

  • Netdata polls the NATS HTTP monitoring endpoints and charts in_msgs vs out_msgs per server, which makes the publish-without-subscribers asymmetry visible as a divergence instead of something you discover from application logs.
  • Subscription count and connection count are tracked continuously, so a responder disconnect shows up as a step down you can correlate with the first no-responders error timestamp from client logs.
  • Connection churn (the total_connections delta) exposes flapping responders that produce intermittent, hard-to-reproduce 503s.
  • Route count per server is charted against the expected full mesh, so partition-induced no-responders are visible from any node’s dashboard.
  • On JetStream servers, API error rate and meta-cluster state from /jsz let you confirm that a burst of publish-side no-responders lines up with a leader election rather than an application bug.