The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nats / nats-retention-policy-confusion ▌

Operations Guides

NATS JetStream retention: limits vs interest vs workqueue and the /dev/null stream

JetStream’s retention policy decides when a stored message is deleted, and the failure modes are asymmetric: get it wrong in one direction and storage grows until publishes are rejected; get it wrong in the other and every message you publish is deleted on arrival while the server reports itself perfectly healthy.

This article covers the three policies (limits, interest, workqueue), the exact conditions under which each one deletes a message, the silent /dev/null failure mode, and the checks that confirm your streams are retaining what you think they are. It assumes a working mental model of streams, consumers, and acks. If not, start with how NATS actually works in production.

What the retention policy actually controls

A stream’s retention policy answers one question: what event makes a message eligible for deletion?

It does not control delivery, acks, redelivery, or flow control. Those live on the consumer. Retention only governs the lifetime of the stored copy. Two other settings interact with it constantly:

  • The stream limits (max_msgs, max_bytes, max_age) act as upper bounds regardless of policy. Even an interest or workqueue stream will evict or reject messages when these limits are hit.
  • The discard policy (DiscardOld or DiscardNew) decides what happens at the limit: evict the oldest messages, or reject new publishes (visible as JetStream API errors and publish failures; see NATS insufficient storage).

So the real behavior of any stream is: retention policy decides deletion in the normal case, limits decide deletion or rejection in the overflow case.

The three policies

flowchart TD
  P[Message published to stream] --> Q{Retention policy}
  Q -->|limits| L[Kept until max age, bytes, or msgs hit]
  L --> D{Discard policy at limit}
  D -->|DiscardOld| D1[Oldest messages evicted]
  D -->|DiscardNew| D2[New publishes rejected]
  Q -->|interest| I{Any consumer interest on the subject}
  I -->|none| I1[Deleted immediately]
  I -->|one or more| I2[Kept until ALL matching consumers ack]
  Q -->|workqueue| W{Any one consumer acked}
  W -->|yes| W1[Deleted]
  W -->|no consumers yet| W2[Retained in stream]

limits (the default)

Messages are kept until a configured limit is reached: maximum age, maximum byte size, or maximum message count. Consumers have no influence on retention at all. A limits stream with zero consumers still stores everything it receives, up to its limits.

This is the policy that behaves the way people coming from Kafka or RabbitMQ assume everything behaves. It is also the policy most likely to fill your disk if you sized the limits optimistically, because retention is completely decoupled from whether anyone is consuming.

interest

A message is kept only while there is consumer interest in it, and deleted once all consumers whose subject filter matches the message have acknowledged it. Two consequences follow directly:

  1. A single stalled consumer blocks retention for the entire stream. If one consumer stops acking, its unacked messages stay, and because deletion requires all interested consumers to ack, nothing it touched can be cleaned up. This is a leading cause of the “consumers stalled, stream grew, publishes rejected” death spiral in NATS insufficient storage.
  2. Zero consumers means zero interest, which means immediate deletion. This is the /dev/null failure mode covered below.

One nuance: when the last consumer for a subject set is deleted (manually or via its inactive threshold), the server may defer the actual deletion of the now-uninterested messages for performance reasons rather than purging them synchronously. Treat the outcome as contractual, not the timing.

workqueue

A message is deleted as soon as any one consumer acknowledges it. This is the classic task-queue semantic: each message is processed exactly once, by one worker, and then gone.

Workqueue is frequently confused with interest because both tie deletion to acks. The difference is the quantifier: interest waits for all matching consumers, workqueue waits for any one. If you want fan-out (multiple independent consumers each seeing every message), workqueue is wrong. Workqueue streams enforce this: consumers must have non-overlapping subject filters, and creating overlapping consumers is rejected.

One edge case that bites operators: messages that exhaust their consumer’s MaxDeliver redelivery attempts are never acked, so on a workqueue stream they are never deleted by the ack path. They remain in the stream and must be removed explicitly via the JetStream API. A slow accumulation of poison messages on a workqueue stream is normal, not a bug, and needs an operational answer (alert, dead-letter, periodic purge).

The /dev/null stream

The single most dangerous retention configuration is an interest stream with no consumers:

  • Publisher sends a message to the stream’s subject.
  • The server evaluates interest. There are no consumers, so there is no interest.
  • The message is deleted immediately.
  • The publish itself can succeed from the client’s point of view. Nothing errors. No slow consumer fires. Server health is green.

The stream is functionally /dev/null while the operator believes data is being persisted. This is most commonly hit in two situations:

  • Ordering mistake during provisioning. The stream is created, publishers are deployed, and the consumer application is deployed later. Everything published in the gap is gone.
  • Ephemeral consumer expiry. The only consumers on the stream were ephemeral and were auto-deleted after their inactivity timeout. Interest dropped to zero, and the stream silently became a sinkhole.

The tell in the metrics: publish rate (in_msgs) is healthy, the stream’s stored message count stays at or near zero, and out_msgs is far below in_msgs. This is the JetStream cousin of the zero-subscriber silent loss pattern in core NATS: the server is doing exactly what it was told to do, and only the asymmetry between what went in and what is stored gives it away.

Note the asymmetry between policies. On a limits stream, no consumers means messages accumulate. On a workqueue stream, messages that have not been acked by anyone are retained, so a workqueue stream with temporarily absent workers keeps its backlog (bounded by the stream limits). On an interest stream, no consumers means messages vanish on arrival. Two of these three “no consumer” states are safe for your data and one is not.

Choosing a policy

  • Event log, replay, audit, “we might need this later”: limits. Retention is a capacity decision, not a consumer-health decision. Size max_bytes / max_age deliberately, and decide up front whether overflow should evict old data (DiscardOld) or backpressure publishers (DiscardNew).
  • Fan-out with delivery guarantee to a known set of consumers: interest. You get “keep it until everyone got it” semantics, at the cost of two operational obligations: never let interested consumer count hit zero, and never let one consumer stall indefinitely. Both need monitoring, not hope.
  • Task distribution to a worker pool: workqueue. Each message handled once, storage freed on first ack. Accept that poison messages past MaxDeliver pile up and plan a dead-letter or purge path.

If you find yourself wanting “interest but don’t delete when consumers are briefly gone,” that is not a retention policy; that is limits retention with durable consumers.

Verifying a stream retains what you think it does

All of these checks are read-only.

# Confirm the configured policy and limits
nats stream info ORDERS --json | jq '.config | {retention, discard, max_msgs, max_bytes, max_age}'

# See what is actually stored right now
nats stream info ORDERS --json | jq '.state | {messages, bytes, first_seq, last_seq, consumer_count}'

# All streams at a glance: stored messages vs consumers
nats stream report

# Server-side aggregate: storage used vs reserved, and the API error counter
curl -s http://localhost:8222/jsz | jq '{storage, reserved_storage, accounts, api_errors: .api.errors}'

# Per-stream detail via the monitoring port
curl -s 'http://localhost:8222/jsz?streams=true' | jq '.account_details[].stream_detail[] | {name: .name, messages: .state.messages, bytes: .state.bytes, consumers: .state.consumer_count}'

The specific cross-checks that catch retention misconfiguration:

  • Interest stream with consumer_count: 0: verify this is intended. If publishers are active and messages is zero, you are in the /dev/null state.
  • messages growing while first_seq never advances: retention is not freeing anything. On interest, find the stalled consumer (check num_ack_pending per consumer via nats consumer info or /jsz?consumers=true). On limits, your limits are simply too generous for the publish rate.
  • Interest or workqueue stream, storage growing despite healthy consumers: messages past MaxDeliver never get acked and never get deleted. Look at num_redelivered trends and dead-letter them.
  • Gap between last_seq - first_seq + 1 and messages: normal during compaction or after deletions; persistent large gaps on interest streams can indicate deferred deletion after a consumer removal.

One caution on tooling: on servers before v2.11, nats stream view creates an ephemeral consumer to read messages. On an interest-retention stream that ephemeral consumer counts as interest, and its lifecycle can interact with deletion. Prefer nats stream info and direct fetches for inspection on older servers, and upgrade when you can.

Version-specific bugs worth knowing

Retention policy behavior has had real bugs. If you are pinned to an older server, check whether you are exposed before trusting the semantics described above. These are drawn from public nats-server issues:

  • Interest + filtered consumers (fixed in v2.7.3): with staggered filtered consumers, acked messages could fail to be removed, causing unbounded growth on interest streams.
  • Interest + nats stream view (fixed in v2.11.0): the CLI’s ephemeral viewer consumer could trigger interest-based deletion of messages that had exceeded max delivery.
  • Interest + stream sourcing across clusters (fixed in v2.14.0): if the destination cluster went down, the replicated consumer disappeared, the source stream saw zero interest, and deleted messages. On earlier versions, prefer limits retention for sourced streams. The v2.14.0 reliable sourcing/mirroring change makes the source/mirror consumer durable and acknowledgment flow-control based.
  • Workqueue with R3 replicas losing max-delivered messages (fixed in v2.14.0): reported on v2.12.3/v2.12.4 in nats-server issue #7817 and fixed by PR #7845; the fix is included in v2.14.0.

The general lesson: interest and workqueue retention depend on consumer state, and consumer state is exactly the kind of distributed state that has edge cases. Limits retention has the fewest moving parts.

Signals to watch in production

SignalWhy it mattersWarning sign
Stream state.messages and state.bytes (/jsz?streams=true)Ground truth for what retention keptInterest stream at zero messages with active publishers; or steady growth that never plateaus
Stream state.consumer_countZero consumers on an interest stream is the /dev/null conditionDrops to 0 on any interest or workqueue stream you expect to be live
/jsz storage vs reserved_storageRetention failure shows up here as exhaustionStorage approaching reserved while consumers look healthy
Consumer num_ack_pending / num_redelivered (/jsz?consumers=true)Stalled or crash-looping consumers block interest/workqueue deletionAck pending pinned at MaxAckPending, redelivered climbing
/jsz api.errors rateDiscardNew rejections and storage-full publishes surface hereSustained error rate alongside growing storage
in_msgs vs out_msgs (/varz)Catches silent loss: messages accepted but delivered nowhereLarge asymmetry not explained by fan-out

How Netdata helps

Retention problems are correlation problems: no single metric says “your policy is wrong,” but two or three together say it clearly. Netdata’s NATS collector polls the server’s HTTP monitoring endpoints and charts the counters that matter here.

  • JetStream stored messages and bytes over time makes the /dev/null stream visible as a flat line at zero while publish throughput is healthy, and makes retention failure visible as an unbroken climb.
  • JetStream storage vs configured limits shows the trajectory toward exhaustion, so a mis-scoped limits stream pages you as a trend, not as a publish-rejection incident.
  • JetStream API error rates catch the DiscardNew and storage-full rejections that are the first externally visible symptom of retention not keeping up.
  • Message throughput asymmetry (in vs out) surfaces silent loss patterns, where messages are accepted and immediately deleted or delivered to nobody.
  • Consumer and stream counts expose the drop-to-zero-consumers event that turns an interest stream into a sinkhole.

Correlating these on one dashboard is what shortens diagnosis: “interest stream, consumer count hit zero at 14:02, stored messages flat since 14:02, publishes fine” is a one-minute conclusion instead of a multi-hour archaeology session.