The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / consul / consul-raft-peers-below-quorum ▌

Operations Guides

Consul lost quorum: Raft peers below the majority needed to elect a leader

A Consul cluster has lost quorum when the number of voting Raft peers drops below the majority required to elect a leader. For a 3-server cluster that means fewer than 2 voters. For a 5-server cluster, fewer than 3. With no leader, every write fails: service registrations, health-check state updates, KV writes, session creation, ACL token creation. Reads served in stale mode still return data, but that data is frozen at the moment quorum was lost.

The defining signal is an empty /v1/status/leader response combined with a voter count from /v1/operator/raft/configuration below the quorum threshold. Quorum is computed from voters only. Non-voter read replicas (an Enterprise feature) and servers mid-snapshot-join do not count. The Voter column of /v1/operator/raft/configuration is the number that matters, not the count of consul members rows.

The recovery path depends on whether a leader still exists anywhere. If at least one voter can reach a leader, you can remove a dead peer through the normal Raft API. If no leader can be elected, the only path is the manual peers.json recovery procedure, which is destructive and must be executed with care.

What this means

Raft requires a majority of voters to commit entries and to elect a leader. When voter count drops below quorum:

  • No new leader can be elected, because every election fails to reach a majority.
  • No follower can be promoted or re-added, because adding or removing peers requires a leader to commit the configuration change.
  • All writes return “no cluster leader” errors.
  • Stale reads continue to work, but they return state frozen at quorum loss.

The cluster cannot self-heal. The dead peers cannot be re-added because adding peers requires a leader, and no leader can be elected because the dead peers are gone.

The crucial distinction is between “below expected count but quorum maintained” and “below quorum.” A 5-server cluster with 4 voters has lost redundancy but is fully functional. A 5-server cluster with 2 voters is in a catastrophic failure. Count voters, not servers.

Common causes

CauseWhat it looks likeFirst thing to check
Permanent server lossTwo or more voter entries point at hosts that no longer exist or cannot run Consulconsul operator raft list-peers -stale against actual host inventory
Autopilot dead-server cleanup without replacementVoter count dropped after a CleanupDeadServers cycle, with no replacement server addedAutopilot config (min_quorum, cleanup_dead_servers) and recent dead-server removal events
Accidental consul operator raft remove-peerA peer disappeared after a manual operator action, often during “cleanup” or migrationOperator command history and audit logs
Stale peers after a failed migrationVoter addresses do not match current server addresses; old IPs lingerCompare the raft/configuration server list with consul members and current node IPs
Network partition isolating votersServers are up individually but cannot reach each other on the Raft RPC portPairwise connectivity checks between every surviving server pair

Quick checks

# Confirm leaderless state from a server
curl -s http://127.0.0.1:8500/v1/status/leader
# Empty string ("") means no leader

# Inspect the Raft peer set, tolerating leaderless state
curl -s "http://127.0.0.1:8500/v1/operator/raft/configuration?stale" \
  | jq '.Servers[] | {Node, Address, Voter, Leader}'

# Same thing via CLI; -stale permits reading without a leader
consul operator raft list-peers -stale

# Count voters explicitly
curl -s "http://127.0.0.1:8500/v1/operator/raft/configuration?stale" \
  | jq '[.Servers[] | select(.Voter == true)] | length'

# Cross-check with gossip membership (alive does NOT mean voter)
consul members

# Check the peers telemetry gauge and election state
curl -s http://127.0.0.1:8500/v1/agent/metrics \
  | grep -E 'consul\.raft\.peers|consul\.raft\.state\.(leader|candidate)'

# Look for election failures and consensus errors
journalctl -u consul --since '30 min ago' \
  | grep -iE 'raft|election|quorum|leader'

# Verify pairwise RPC reachability between server pairs (port 8300)
for peer in 10.0.0.2 10.0.0.3; do
  nc -zv -w 2 "$peer" 8300
done

Only count voters. consul members shows nodes as alive even when they are not in the Raft configuration. Gossip and Raft are independent subsystems and routinely diverge, which is a common source of confusion during quorum incidents.

The ?stale query parameter on /v1/operator/raft/configuration is essential here. Without it, the endpoint tries to consult a leader, which by definition does not exist. ?stale lets the endpoint return the cached configuration directly.

The CLI flag is a Boolean HTTP flag: -stale (or the equivalent -stale=true); it sets AllowStale, so the command can read the local/stale Raft configuration without a leader.

How to diagnose it

  1. Confirm leaderless state. Hit /v1/status/leader from each surviving server. An empty response on all of them confirms no leader exists cluster-wide. A non-empty response on one server but empty on others suggests a stale read or an asymmetric partition, not a true leaderless state.
  2. Enumerate voters using ?stale. This is the only reliable way to read the configuration when the cluster has no leader. Note each entry’s Node, Address, Voter, and Leader fields.
  3. Compare voter count to expected quorum. Quorum is (N / 2) + 1. For 3 servers, that is 2. For 5 servers, that is 3. If the voter count is at or above quorum but /v1/status/leader is empty, suspect a transient election, slow disk, or partition. Do not treat that as permanent quorum loss.
  4. Classify each missing voter. For each voter in the expected set that is absent from the configuration, determine whether the underlying host is permanently gone (terminated, decommissioned, disk failed) or temporarily unreachable (rebooting, network blip, snapshot-join in progress). The recovery path differs sharply between the two.
  5. Check Autopilot. If CleanupDeadServers is enabled (the default), Autopilot may have removed dead voters automatically. A missing or zero min_quorum setting lets Autopilot remove voters down to whatever it considers dead, including below quorum. This is one of the most common ways clusters silently lose redundancy during rolling restarts.
  6. Verify pairwise connectivity. Network partitions present exactly like quorum loss. Confirm that every surviving server can reach every other surviving server on the Raft RPC port (8300) and the Serf LAN port (8301). Asymmetric partitions are common and easy to miss.
  7. Decide on a recovery path using the decision tree below.
flowchart TD
  A["/v1/status/leader empty?"] -->|Yes| B["Read raft/configuration?stale"]
  B --> C{"Voter count >= quorum?"}
  C -->|Yes| D["Transient: disk, network, or election.
Do not remove peers."] C -->|No| E["Classify missing voters"] E --> F{"Any voter can reach a leader?"} F -->|Yes| G["consul operator raft remove-peer,
then add replacements."] F -->|No| H{"Partition suspected?"} H -->|Yes| I["Fix the network first.
Raft self-heals on heal."] H -->|No| J["Manual peers.json recovery.
Destructive: data loss possible."]

Metrics and signals to monitor

SignalWhy it mattersWarning sign
/v1/operator/raft/configuration with ?stale / list-peers -staleAuthoritative local Raft view of votersVoter count below the quorum threshold
/v1/operator/raft/configuration voter countAuthoritative voter roster with addresses and leader flagVoter count below (N/2)+1, or voters missing from the list
/v1/status/leaderBinary “leader exists” signalEmpty string beyond the election timeout window
consul.raft.state.candidate (counter)Counts election startsIncreases with no successful election
consul.raft.leader.lastContact (leader-side timer)Time since leader last contacted followersTrending toward the election timeout before leader loss
consul.serf.lan.members alive countGossip view of membership, independent of RaftDiverges from the Raft voter list
Autopilot min_quorumPrevents dead-server removal below the thresholdUnset or zero in a cluster relying on CleanupDeadServers
Disk write latency on the Raft volumeSlow fsync causes election timeouts that look like quorum lossawait sustained above 10ms on the data-dir volume

Fixes

If a leader still exists somewhere

If at least one voter can still reach a leader, the safe path is consul operator raft remove-peer. The change is committed through Raft itself, preserving log consistency.

# Remove a permanently-dead peer by address
consul operator raft remove-peer -address="10.0.0.3:8300"

Only remove peers you are certain are permanently gone. Premature removal of a peer that is still reachable can cause split-brain if that peer later rejoins with divergent state.

After removal, add replacement servers one at a time. Each new server must complete the snapshot-join before it is promoted to voter. Autopilot will promote it automatically once server_stabilization_time is satisfied. Do not promote manually unless you understand the consequences.

If no leader can be elected: peers.json recovery

When the cluster truly has no leader and cannot elect one, consul operator raft remove-peer does not work: the command requires a leader to commit the change. The only recovery path is the manual peers.json procedure, which rewrites the Raft configuration directly on disk.

Warning: this procedure is destructive. It implicitly commits outstanding Raft log entries, including uncommitted ones, and can cause data loss. Multiple servers being lost is the usual reason you are in this state, which means committed entries may already be incomplete. Treat this as a last resort.

Procedure:

  1. Pick a single surviving server to seed the new configuration. Its Raft log becomes the authoritative state for the recovered cluster. Choose the server with the most complete log if you can determine it.
  2. Stop Consul on every server. The cluster must be cold during recovery.
  3. On the chosen seed server, locate the Raft data directory (commonly <data_dir>/raft/).
  4. Read the node ID for each surviving voter you want in the new configuration. The ID lives in the node-id file at the root of the data directory (for example, <data_dir>/node-id), not inside the raft subdirectory.
  5. Write a peers.json file into the Raft directory. The format is a JSON array of objects, one per surviving voter:
[
  {"id": "<node-id-of-seed>", "address": "10.0.0.1:8300", "non_voter": false},
  {"id": "<node-id-of-second-survivor>", "address": "10.0.0.2:8300", "non_voter": false}
]

Only include servers that have valid Raft data on disk. Including a server with an empty data directory causes Consul to refuse to start, with an error indicating it will not recover a cluster with no initial state.

  1. Start Consul on the listed servers. Consul ingests peers.json, recovers the Raft configuration, and deletes the file on successful startup. Keep all listed servers stopped until the file is in place, and ensure every surviving server receives the same configuration before startup.
  2. Confirm a leader exists via /v1/status/leader.
  3. Once quorum is healthy, add any additional replacement servers through the normal consul join flow and let Autopilot promote them.

For Raft protocol v3 (the current default), ReadConfigJSON accepts the id, address, and optional non_voter fields shown above; non_voter defaults to false when omitted. Older protocol v2 clusters used the flat address-string format, and current v3 clusters parse only the object format.

If the cause is a network partition

Do not run remove-peer or peers.json recovery until you understand the partition geometry. A partition that heals will let Raft reconcile automatically: the minority side discards its state and resyncs from the majority. Manual peer manipulation during a transient partition can permanently split the cluster or discard committed entries.

Confirm the partition is real and persistent before treating it as permanent server loss. Asymmetric partitions are common: server A may reach B but not C, while B reaches C but not A. Test every pair, not just paths to the leader.

Prevention

  • Set autopilot.min_quorum. The default of 0 lets Autopilot prune without a quorum floor. Set it to the number of voters you must retain for quorum—3 for both 3- and 5-server clusters; use a higher value only if you want a stricter maintenance floor. This single change prevents the most common cause of cascading quorum loss during rolling restarts.
  • Wait for Voter: true between restarts. During rolling restarts, confirm each restarted server has rejoined as a voter in consul operator raft list-peers before restarting the next one. Autopilot can remove a temporarily-down server before it rejoins, and if the next server fails before the first re-promotes, quorum is lost.
  • Use odd server counts. A 4-server cluster has the same failure tolerance as a 3-server cluster (tolerates 1 loss) but loses quorum on the second failure just as fast. Use 3 or 5.
  • Spread servers across failure domains. Three servers in two availability zones guarantees that a single AZ failure can lose quorum. Spread servers so that no single AZ, rack, or power domain holds a majority.
  • Keep the Raft data directory on dedicated fast storage. Slow disks cause election timeouts that look identical to quorum loss in the metrics. SSDs are not optional for Consul servers.
  • Rehearse peers.json recovery. The procedure is destructive and easy to get wrong under pressure. Run it in a staging cluster at least once before you need it for real.
  • Page on voter count, not just leader absence. A cluster can lose a voter and continue operating with no alerts until the next failure takes it below quorum. Alert on voter count below expected, not just on an empty /v1/status/leader.

How Netdata helps

  • Per-second consul.raft.peers collection catches a voter drop the moment it happens, before the next failure pushes the cluster below quorum. A 1-minute polling interval can miss the entire window between Autopilot removal and the next failure.
  • Correlate consul.raft.peers with consul.raft.state.candidate and /v1/status/leader on a single timeline. The shape of the divergence tells you whether you are dealing with permanent quorum loss, a transient election, or a network partition.
  • ML anomaly detection on consul.raft.leader.lastContact and disk write latency surfaces the slow-disk pattern that precedes many quorum losses. Elections triggered by fsync latency look identical to elections triggered by server loss in raw metrics, but the preceding disk-latency anomaly distinguishes them.
  • Cross-server views of consul.serf.lan.members against the Raft configuration surface the gossip/Raft divergence that signals a phantom node or a stuck snapshot-join before it becomes a quorum incident.