The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / consul / consul-leader-election-storm ▌

Operations Guides

Consul leader election storm: repeated elections and rolling write outages

Your Consul cluster has a leader. Writes are still failing. The leader keeps changing. Applications see intermittent “no cluster leader” errors, watches reconnect, DNS returns stale results, and every few seconds a different server wins an election only to lose it again.

consul.raft.state.leader and consul.raft.state.candidate are counters: the former increments when a server becomes leader, and the latter when it starts an election. consul.server.isLeader is the separate 1/0 gauge.

This is a leader election storm. Unlike a clean one-time failover, the cluster never stabilizes. Each election blocks all writes for one to several seconds. When elections recur faster than the recovery window, the cluster is effectively write-unavailable while technically always having “a leader.” A naive alert on “no leader” stays silent. The real signal is recurrence: leadership transitions accumulating over time, each one a brief but real outage.

What this means

Each Raft leader election is a write outage window. During the election, no server accepts writes. Existing reads may continue in stale mode, but consistent reads and writes block. When the election completes, writes resume until the next election starts.

In a storm, this cycle repeats. The cluster oscillates between “briefly functional with a new leader” and “no leader, election in progress.” At one election every 30 seconds, most operations eventually succeed with elevated latency. At several elections per minute, most writes time out.

The root mechanism is almost always the same: something is preventing the leader from sending heartbeats to followers fast enough. Raft followers start an election if they do not receive a heartbeat within the election timeout. If the underlying cause (slow disk, CPU starvation, network latency) affects all servers equally, each new leader hits the same wall and loses leadership again.

flowchart TD
    A[Slow disk / CPU / network] --> B[Leader cannot fsync Raft log fast enough]
    B --> C[Heartbeat to followers delayed]
    C --> D[Follower election timeout fires]
    D --> E[Follower starts election]
    E --> F[New leader elected]
    F --> G[Same underlying cause persists]
    G --> B
    F --> H[Brief write outage during each election]

The operational threshold: more than 2 elections in 10 minutes outside a maintenance window indicates a systemic problem.

Common causes

CauseWhat it looks likeFirst thing to check
Slow disk I/O on leaderconsul.raft.commitTime elevated, disk await above 10ms, elections may correlate with snapshot creationiostat -x 1 on the leader’s Raft data volume
CPU starvationServer CPU pinned at 100%, gossip probe timeouts, elections without preceding high lastContactCheck cgroup CPU limits if containerized
Network latency or packet loss between serversconsul.raft.leader.lastContact trending up before each electionPairwise mtr between all server pairs
Go GC pauses on large heapconsul.runtime.gc_pause_ns spikes correlate with elections, heap is multi-GBGC pause telemetry, heap profile via pprof
Asymmetric network partitionOnly some followers show high lastContact, one server repeatedly wins then losesconsul members from each server, compare views
Version-specific regressionStorm started after upgrade, recovery takes minutes instead of secondsCheck Consul version against known regressions

Slow disk is the most common cause by a wide margin. The Raft log is persisted with fsync on every write. If the disk cannot complete fsync fast enough, heartbeats queue behind log writes. EBS gp2 volumes with exhausted burst credits, network-attached storage, and spinning disks are the usual suspects.

Quick checks

Run these read-only checks to confirm the storm and identify the current leader.

# Confirm current leader identity (empty string means no leader)
curl -s http://127.0.0.1:8500/v1/status/leader

# Check leadership state on this server
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep "raft.state.leader"

# Watch for election log lines in real time
journalctl -u consul -f | grep -E "entering leader state|heartbeat timeout reached, starting election"

# Full Raft peer configuration
curl -s http://127.0.0.1:8500/v1/operator/raft/configuration | python3 -m json.tool

# Last contact times (reported on the leader about each follower)
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep "raft.leader.lastContact"

# Raft commit time (leader only)
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep "raft.commitTime"

# Disk I/O on the Raft data volume
iostat -x 1 5

# GC pause duration
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep "gc_pause"

Do not restart servers during an active storm. Restarting removes one voter from the peer set and may push the cluster below quorum, converting a storm into total unavailability.

How to diagnose it

  1. Confirm recurrence. Check whether consul.raft.state.leader is changing repeatedly. A single transition is a failover. Multiple transitions within 10 minutes is a storm. Cross-check with server logs for repeated “entering leader state” lines.

  2. Identify the current leader. Query /v1/status/leader. Note which server holds leadership and whether it changes between checks.

  3. Check disk I/O on the leader. Run iostat -x 1 5 on the server that is currently leader. Look at await (write latency) and %util. If await is above 10ms sustained, or %util is pinned at 100%, disk I/O is the likely cause. Also check whether you are on EBS gp2 with exhausted burst credits.

  4. Check lastContact for all followers. The metric is leader-side and reports contact with followers. If all followers show rising lastContact before each election, the leader is struggling. If only one follower shows it, the problem is the network path to that specific follower.

  5. Check CPU and GC. If disk I/O is healthy, check whether the server is CPU-starved (pinned at 100%) or experiencing GC pauses. Both prevent the Raft goroutine from sending heartbeats on time.

  6. Check pairwise network connectivity. Run consul info on each server and compare the Serf LAN sections. Asymmetric partitions, where server A sees B but not C, are a common cause of repeated elections.

  7. Correlate with recent changes. Did the storm start after a deploy, a scaling event, a Consul upgrade, or a storage change? Version-specific regressions in Raft behavior exist and should be ruled out.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
consul.raft.state.leader (counter)Counts leadership acquisitions.More than 2 transitions in 10 minutes outside maintenance
consul.raft.leader.lastContact (leader-side timer)Time since leader last contacted a followerSustained above 200ms, or trending toward the election timeout
consul.raft.commitTimeEnd-to-end Raft write latency, leader onlySustained above 100ms, or approaching heartbeat timeout
Disk write latency (await)Leading indicator for Raft instabilitySustained above 10ms on the Raft data volume
consul.runtime.gc_pause_nsGC pauses block the Raft goroutineSpikes above 100ms correlating with elections
consul.raft.state.candidate (counter)Counts election startsAny non-zero value in production
Write API error rateUser-visible impactBursts of “no cluster leader” errors

The consul.raft.commitTime signal is only reported on the leader. When leadership changes, this metric disappears and reappears on a different server. Your monitoring must track which server is leader to interpret it correctly.

Fixes

Slow disk I/O

This is the most common cause. Investigate it first.

Immediate: Identify and reduce write load. If a health check thundering herd or runaway registration loop is generating excessive Raft writes, shedding that load reduces fsync pressure. Check the consul.catalog.register latency-timer sample rate for churn.

Short-term: If the leader’s disk is slow due to burst credit exhaustion on EBS gp2, credits will refill over time but the storm will continue until they do. You cannot wait this out during an active incident. Migrate the Raft data directory to faster storage: provisioned IOPS volumes, local NVMe, or a dedicated SSD volume.

Permanent: Never colocate the Raft data directory with other I/O-heavy workloads. HashiCorp production guidance calls for dedicated SSD storage with sufficient IOPS headroom. Rotational disks and network-attached storage are not suitable for Raft persistence.

CPU starvation

If the server is containerized, check CPU limits. Consul servers under-provisioned on CPU cannot process Raft, gossip, and RPC concurrently. The Raft goroutine gets starved and misses heartbeat deadlines.

Fix: Increase CPU allocation. HashiCorp’s production starting point is 8–16 CPU cores. If you cannot resize immediately, reduce non-critical load by disabling non-essential health checks or reducing DNS recursion.

Network latency between servers

Consul servers must maintain low-latency connectivity. The election timeout assumes sub-second round trips between servers. If inter-server latency approaches the election timeout, followers will repeatedly trigger elections.

Fix: Verify network connectivity between all server pairs, not just to and from the leader. Asymmetric partitions are common. HashiCorp’s production starting point is average RTT below 50ms and 99th-percentile RTT below 100ms. Verify inter-server RTT is within tolerance, especially if servers span availability zones.

Go GC pauses on large heaps

Consul servers with multi-GB heaps can experience stop-the-world GC pauses long enough to trigger election timeouts. The consul.runtime.gc_pause_ns metric will show spikes correlating with elections.

Fix: Reduce heap pressure. Identify what is consuming memory (large KV values, catalog bloat, watch accumulation) and address it. Tuning GOGC upward reduces GC frequency at the cost of higher peak memory. This is a stopgap, not a solution.

Version-specific regressions

Consul 1.13.x introduced a regression where evicting a single server node, even a non-leader follower, caused the cluster to cycle through leaders for 3-15 minutes before stabilizing. This was not present in 1.12.x where recovery took 2-10 seconds. The regression was reported against 1.13.1 through 1.14.0-beta1.

If your storm started immediately after an upgrade and the recovery behavior matches this pattern, check your Consul version against the known regression.

The regression was fixed in Consul 1.13.4; the corresponding Raft update was already present in Consul 1.14.0 (issue 15231).

Asymmetric partitions

If only some followers consistently lose contact with the leader, you may have an asymmetric network partition. Server A can reach B, B can reach C, but A cannot reach C. This causes elections that succeed from one perspective but fail from another.

Fix: Run consul members from each server and compare the views. Use mtr or ping to test all server pairs in both directions. Firewall changes, security group updates, and routing table drift are common causes.

Prevention

Monitor disk write latency as a page-level signal. This is the single most important preventive measure. Disk write latency (await) on the Raft data volume should stay below 10ms. Alert before it approaches the election timeout. Most storms are preventable if disk latency is caught early.

Alert on election count, not just leader absence. A naive “is there a leader” check will not catch a storm. Alert when leadership transitions exceed 2 in 10 minutes outside a maintenance window.

Suppress expected elections during rolling upgrades. During rolling server restarts, some elections are expected. Your alerting should account for maintenance windows or use suppression during planned operations.

Run 3 or 5 servers, never an even number. Even server counts can split quorum evenly during a partition (2-2 with 4 servers), preventing any leader election from succeeding. Always use odd numbers of voting servers.

Track consul.raft.commitTime trends over days. Commit time creeping upward is the early warning before a storm. If commit time p99 trends toward a significant fraction of the election timeout, investigate before it becomes an incident.

Validate Consul upgrades in staging. Version-specific regressions in Raft behavior exist. Test upgrades with realistic load before applying to production. Monitor election behavior specifically during and after upgrades.

How Netdata helps

  • Per-second metric resolution captures election transitions that coarser polling misses. A storm with elections every few seconds is visible as a steep staircase on consul.raft.state.leader rather than a single blurred sample.
  • Correlate consul.raft.leader.lastContact with disk latency on the same dashboard. When both spike together, the root cause is disk I/O. When lastContact spikes without disk latency, the cause is network or CPU.
  • Anomaly detection on consul.raft.commitTime surfaces gradual degradation before it crosses a static threshold.
  • Per-container CPU and memory metrics identify CPU starvation or GC pressure on Consul servers running in containers, where the problem is invisible to host-level metrics.