The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nats / nats-meta-cluster-leader-flapping ▌

Operations Guides

NATS JetStream meta cluster leader flapping: admin operations that keep failing

Stream and consumer administration is failing, but existing JetStream traffic mostly continues. Creating a stream times out, consumer updates return errors, and retries sometimes succeed only to fail again a minute later.

This pattern points at the JetStream meta Raft group. The meta group manages cluster-wide JetStream metadata: stream and consumer create, update, and delete operations. Existing streams use separate Raft groups for replicated data, so the data plane can remain available while the administrative plane is unstable.

A single leader change during maintenance is not an incident by itself. More than one meta leader change in five minutes is concerning, especially when peers report current=false, offline=true, or persistent replication lag.

What this means

Every clustered JetStream deployment has one meta Raft group for cluster metadata, plus separate Raft groups for replicated streams and consumers. This distinction drives the diagnosis:

  • An unstable meta leader breaks administrative operations across JetStream.
  • An unstable stream Raft group affects writes and replication for that stream only.
  • Both can flap together when the root cause is CPU pressure, disk latency, network instability, or a wider Raft election storm.

During a meta election, an administrative request can reach a server that no longer leads the meta group, wait for consensus that never completes, or fail while leadership moves. Client retries then inflate JetStream API traffic and make the failure look larger than the original event.

flowchart LR
  A[Admin request] --> B[Meta Raft leader]
  B --> C{Quorum commits metadata}
  C -->|Stable leadership| D[Create update delete succeeds]
  B -->|Heartbeat or election failure| E[Leader change]
  E --> F[Admin request times out or errors]
  G[Existing stream Raft groups] --> H[Data plane often continues]
  E -. May also trigger .-> G

The signal that matters is not just the current leader name. It is the rate of leader changes, whether all expected peers are present, and whether each peer is current.

Common causes

CauseWhat it looks likeFirst thing to check
Network latency or partitionRoute count drops, route RTT rises, or route pending bytes grow. Leader changes correlate with route events.Compare route count with the expected N-1 routes and inspect /routez.
Meta leader overloadThe current or recent leader has high CPU, memory pressure, or GC pauses. Elections follow resource spikes.Check /varz CPU, memory, uptime, and throughput on every node.
JetStream disk I/O stallapi.inflight stays high, API errors rise, and OS-level iowait or disk latency increases.Check disk latency and look for backups, snapshots, or noisy neighbors.
Lagging or offline meta peerOne or more meta_cluster.replicas[] entries report current=false, offline=true, or sustained lag.Inspect /jsz meta cluster state from more than one server.
Raft election stormMany stream groups change leaders along with the meta group. CPU and API errors spike across nodes.Inspect /raftz and server logs for repeated leader changes.
Clock or timer instabilityElections occur without a clear route, disk, or CPU cause.Verify time synchronization on all JetStream nodes.
Expected maintenanceOne controlled restart or deployment causes one leader change, then the cluster stabilizes.Compare the event time with deployment and restart history.

Do not treat every API error as proof of meta instability. Idempotent create or bind operations can produce benign errors, and client retries inflate both api.total and api.errors. Confirm the leader is actually changing before treating this as a Raft incident.

Quick checks

Run these from a host that can reach the NATS monitoring port. They are read-only. The default monitoring port is 8222; use the port configured for your deployment.

# 1. Show the current meta leader and peer state
curl -s http://localhost:8222/jsz | jq '.meta_cluster | {leader, replicas: [.replicas[]? | {name, current, offline, lag}]}'

# 2. Poll for leader changes over two minutes
for i in $(seq 1 12); do
  date -u +%H:%M:%S
  curl -s http://localhost:8222/jsz | jq -r '.meta_cluster.leader // "no-meta-cluster"'
  sleep 10
done

# 3. Check JetStream API pressure and errors
curl -s http://localhost:8222/jsz | jq '{api_total: .api.total, api_errors: .api.errors, api_inflight: .api.inflight}'

# 4. Check route count, route RTT, and route backpressure
curl -s http://localhost:8222/routez | jq '{num_routes, routes: [.routes[]? | {remote_id, ip, rtt, pending_size}]}'

# 5. Check server CPU, memory, uptime, and connection pressure
curl -s http://localhost:8222/varz | jq '{cpu, cores, mem, uptime, connections, slow_consumers}'

# 6. Check whether JetStream itself is enabled
curl -s http://localhost:8222/jsz | jq '{disabled, streams, consumers, memory, storage}'

# 7. Inspect all Raft groups when the problem may extend beyond metadata
curl -s http://localhost:8222/raftz | jq .

# 8. Check disk latency and utilization
iostat -x 1 5

# 9. Look for repeated election events in the server log
grep -Ei 'JetStream cluster new metadata leader|Self is new JetStream cluster metadata leader|stepping down' /var/log/nats/nats-server.log | tail -50

Notes on these checks:

  • Expected route count is N-1 for a full mesh of N nodes. Use num_routes from /routez (expected to be N-1 for a full mesh) and meta_cluster.cluster_size from /jsz; current /varz does not expose a separate cluster-size field.
  • On standalone JetStream, meta_cluster is null because there is no clustered metadata Raft group.
  • /raftz output can be large on clusters with many streams. Use it after /jsz indicates a wider Raft problem.
  • iostat requires the sysstat package on most Linux distributions.
  • The log path varies by deployment; adjust it to where your nats-server writes logs.
  • Run the /jsz check on more than one node. A single server can have a stale or incomplete view while it is catching up.

How to diagnose it

1. Confirm that administration, not the whole server, is failing

Use a safe read operation through the JetStream API:

# Read stream state through the JetStream API
nats stream report

If this fails while /healthz?js-server-only=true remains healthy and existing publish or consume paths continue, the evidence points at the administrative plane. If basic server health is failing too, widen the investigation beyond the meta group.

2. Measure the leader change rate

Poll meta_cluster.leader for at least five to ten minutes. One change can be a clean failover or a maintenance event. Repeated changes, particularly more than one in five minutes, indicate active instability.

Record:

  • Old and new leader names.
  • Time of each change.
  • Whether the same node repeatedly loses leadership.
  • Whether changes align with deployments, backups, network events, or CPU spikes.

3. Check whether the meta group has all expected peers

Inspect every meta_cluster.replicas[] entry:

  • current=false means the peer is behind.
  • offline=true means the peer is unreachable or not participating.
  • lag shows replication lag in entries.

A peer briefly behind during a burst can recover. A peer offline for more than 60 seconds, or one whose lag never resolves, reduces the group’s fault tolerance and can prevent stable consensus.

4. Separate meta flapping from per-stream flapping

Use /raftz to inspect individual Raft groups. If only the meta group is changing leaders, existing replicated streams may continue to serve normally. If many stream groups are electing at the same time, you have a broader election storm and should expect publish failures or replication stalls too.

This distinction determines the blast radius. Do not assume the data plane is safe just because messages are flowing on one stream.

5. Correlate API errors with inflight requests

A rising error count with low inflight can be client-side or benign retry noise. Rising errors with sustained high api.inflight is stronger evidence that JetStream cannot complete requests, usually because consensus or disk writes are slow.

Calculate the error ratio over a fixed interval rather than alerting on cumulative counters:

# Compare API totals over a 60 second interval
curl -s http://localhost:8222/jsz | jq '{total: .api.total, errors: .api.errors}'
sleep 60
curl -s http://localhost:8222/jsz | jq '{total: .api.total, errors: .api.errors}'

A sustained error ratio above 5 percent indicates a systemic issue. Aggressive client retries can amplify both counters, so cross-check against inflight before concluding the server is at fault.

6. Inspect the nodes that recently held leadership

For the current and previous leaders, check:

  • CPU saturation and abrupt CPU spikes.
  • Memory growth and GC pauses.
  • Disk latency, iowait, and filesystem activity.
  • Route RTT and route pending bytes.
  • Unexpected uptime resets.

The server that loses leadership is often the node under pressure, but the root cause can also be another peer that cannot vote or append Raft entries quickly enough.

7. Look for a wider election storm

Search logs for alternating leader events and check whether stream groups are also unstable. A broad election storm can become self-reinforcing: resource pressure causes elections, elections increase CPU and disk work, and that additional work causes more elections.

If many groups are involved, reduce load and fix the shared resource problem before attempting administrative changes.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Meta leader nameTracks elections in the metadata control plane.More than one change in five minutes.
Meta peer currentShows whether each peer has caught up with the leader.Any peer remains current=false.
Meta peer offlineShows whether a peer is participating in consensus.Any peer remains offline for more than 60 seconds.
Meta peer lagShows replication lag in Raft entries.Lag persists or grows instead of returning to zero.
JetStream API inflightIndicates requests waiting on JetStream processing, consensus, or storage.Sustained elevation above baseline.
JetStream API error ratioConfirms that administration or publish requests are failing.Errors above 5 percent of API requests over a sustained interval.
Route countVerifies the full cluster mesh.Fewer than N-1 routes for more than 60 seconds.
Route RTT and pending sizeDetects network delay and inter-server backpressure.Sustained RTT increase or nonzero route pending size.
Server CPU and memoryFinds leader overload and GC pressure.CPU above 90 percent for five minutes or monotonic memory growth.
Disk latency and iowaitFinds storage stalls that delay Raft and JetStream writes.Persistent latency or iowait elevation during elections.
Server uptimeDetects restarts that trigger elections.Unexpected resets or repeated restarts across nodes.

Fixes

Repair route and network instability

Restore the expected full-mesh route count and investigate latency between JetStream nodes. Check network device events, firewall changes, DNS resolution, and cross-zone or cross-region paths.

Do not restart NATS nodes just because a route briefly disconnected. Routes reconnect automatically, and unnecessary restarts create more elections.

Relieve pressure on the current leader

If elections follow CPU or memory spikes, reduce pressure on the affected node. Pause noncritical publishers or administrative jobs, investigate noisy neighbors, and add CPU or memory capacity where the node is undersized.

Evaluate memory as a trend. Go garbage collection creates a sawtooth pattern, so a temporary rise followed by a drop is normal. Monotonic growth without recovery is not.

Remove the disk bottleneck

If api.inflight, disk latency, and iowait rise together, look for filesystem backups, snapshots, disk degradation, or other tenants sharing the storage device. Pause avoidable I/O work where possible.

For a persistent bottleneck, plan a migration to lower-latency local storage. Network-attached storage with variable latency is a common contributor to Raft instability. Treat storage migration as a planned change, not an in-incident experiment.

Restore unhealthy meta peers

Bring offline or lagging peers back before making cluster membership changes. Verify that routes are healthy, JetStream is enabled, disk latency is normal, and the peer has caught up.

Avoid decommissioning or replacing a node while the meta group is unstable. Membership operations during an active quorum or election problem can make recovery harder.

Break a wider election storm carefully

First reduce incoming load and correct the shared resource problem. If the storm remains self-sustaining, a rolling restart may reduce pressure, but it is disruptive and can trigger more stream and meta elections.

If you use it as a last resort:

  1. Restart one node at a time.
  2. Wait for routes to return.
  3. Verify JetStream is enabled.
  4. Verify meta peers are present, current, and not offline.
  5. Check per-stream Raft health before continuing.

Do not restart the whole cluster at once.

Repair time synchronization

If no resource or network cause explains repeated elections, verify clock synchronization on all JetStream nodes. Correct the time source, then watch whether election frequency returns to normal.

Be cautious with election timeout tuning

Operational guidance differs on whether Raft timing should be tuned, and behavior can be version-dependent. Do not assume there is a supported election timeout configuration for your exact server version. Verify the setting and its semantics against the deployed version before changing it. Fixing network, CPU, or disk delay is usually safer than masking it with longer timeouts.

Prevention

  • Leader change alerting: Alert when the meta leader changes more than once in five minutes. This threshold separates normal isolated failovers from active flapping.
  • Peer health alerting: Alert when a meta peer stays current=false or offline=true for more than 60 seconds. A reduced peer set leaves less room for another failure.
  • API correlation: Track API errors together with api.inflight. Error count alone includes benign retries; errors plus inflight better indicate stalled consensus or storage.
  • Route health monitoring: Compare current routes with the expected N-1 count and monitor route RTT and pending size. Route instability is a common upstream cause of elections.
  • Disk headroom: Keep disk utilization below 80 percent and preserve IOPS headroom. JetStream and Raft both become unstable when storage latency rises.
  • Leader distribution review: Check whether one server leads a disproportionate number of stream groups. Concentrated leadership increases the impact of a single overloaded node.
  • Deployment correlation: Record restarts, configuration changes, backups, and infrastructure events alongside leader changes. This shortens the search for the trigger.
  • Health check separation: Use /healthz?js-server-only=true for basic server readiness and treat bare /healthz JetStream failures carefully during recovery. This avoids confusing asset recovery with a process outage.

How Netdata helps

  • Netdata’s NATS monitoring can surface JetStream API errors, inflight requests, storage usage, stream counts, and consumer counts, helping confirm that administrative pressure is rising.
  • Server CPU, memory, uptime, throughput, connection, and slow-consumer signals help identify whether the node losing leadership is overloaded or restarting.
  • Host-level disk utilization, latency, and iowait can be correlated with election times to distinguish a storage stall from a network problem.
  • Route and connection signals help show whether the meta election follows inter-server connectivity loss rather than a JetStream-specific fault.
  • Comparing nodes side by side helps identify whether one peer is consistently behind while the rest of the cluster remains healthy.
  • Netdata collects the meta_cluster structure from /jsz but may not expose leader, peer status, or peer lag as dedicated metrics. Keep a separate /jsz poll or log-based watch for those fields.