The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / consul / consul-snapshot-size-growth ▌

Operations Guides

Consul snapshot size growing: state bloat, slow restores, and commit spikes

Consul snapshots grow with the size of the FSM state: every registered service instance, every health check, every KV entry, every ACL token, every session. None of those individual writes looks expensive, so a team that registers a few new services per week will not see a spike in any single metric. What they get instead is a slow, monotonic climb in snapshot size that compounds across months until one day a snapshot save takes long enough to push Raft commit times past the heartbeat timeout, or a rejoined follower sits in snapshot restore long enough to fall out of the replication loop entirely.

This is a slow-burn failure. No single operation looks expensive, and the cost shows up at the worst possible moment: during leadership churn, a rolling restart, or a failover when you most need followers to catch up quickly. The mental model for how Consul’s Raft and catalog interact is covered in How Consul actually works in production; this article assumes that background and focuses on the snapshot-growth symptom.

The fix is not reactive. By the time a snapshot save is slow enough to page you, the catalog or KV bloat has been building for weeks. Track snapshot size as a PLAN-level weekly trend, correlate it with catalog and KV growth, and catch the slope before the cliff.

What this means

A snapshot is the serialized form of the in-memory FSM: catalog, KV, sessions, ACLs, intentions, prepared queries. Consul writes snapshots periodically (controlled by raft_snapshot_interval, default 30s in modern versions) once the Raft log accumulates enough entries since the last snapshot (controlled by raft_snapshot_threshold, default 16384).

Every server restores the same snapshot when it joins, rejoins after a restart, or falls behind the leader’s trailing log window and must be caught up via InstallSnapshot RPC.

Three costs scale linearly, or worse, with snapshot size:

  • Memory during creation. Snapshot save serializes in-memory state and temporarily inflates the working set. On a large FSM, this spike drives GC pause time up at the exact moment Raft needs the CPU.
  • Disk I/O during save and restore. Snapshots are large sequential writes on save and large sequential reads on restore. On shared or burst-credit-bounded storage (EBS gp2), a multi-gigabyte snapshot can saturate the volume for seconds, competing with Raft log fsyncs.
  • Restore time on rejoin. A follower that restarts must restore the latest snapshot before it can accept AppendEntries. Restore is single-threaded and CPU-bound. On a large snapshot this can take minutes, during which the follower is a non-participant in consensus.

The dangerous failure mode is when snapshot creation or restore takes long enough to interfere with Raft timing. Snapshot save holds disk I/O that Raft needs for log writes. Snapshot restore keeps a follower silent long enough that the leader’s trailing logs (raft_trailing_logs, default 10000 since Consul 1.5.3) are exhausted, forcing another full snapshot install in a loop.

flowchart TD
    A[Catalog + KV grow slowly] --> B[Snapshots grow proportionally]
    B --> C[Snapshot save spikes memory and disk IO]
    B --> D[Snapshot restore takes minutes on rejoin]
    C --> E[raft.commitTime rises during save]
    D --> F[Follower falls behind trailing log window]
    E --> G[Leader election during snapshot save]
    F --> H[InstallSnapshot loop - follower never catches up]
    G --> I[Write outage]
    H --> I

Common causes

CauseWhat it looks likeFirst thing to check
Catalog bloat from unbounded service/check growthSnapshot size climbs week over week; service instance counts grow faster than infrastructureconsul catalog services count vs inventory
KV store used as a databaseSnapshot grows; KV apply rate is a large fraction of total Raft apply rateEnumerate KV keyspace size and per-prefix bytes
Zombie service registrations from ephemeral workloadsSnapshot Register section dominates; service count does not match running instancesconsul snapshot inspect per-type breakdown
Missing deregister_critical_service_afterCritical checks accumulate; deregister rate near zeroCount of checks per service vs expected
Infrequent or failed compactionRaft log directory grows alongside snapshots; raft_snapshot_threshold rarely hitdu -sh <data_dir>/raft/ and snapshot file count
Slow disk amplifying the snapshot costiostat await rises during snapshot save; commit time spikes correlate with snapshot intervalsDisk write latency on the Raft volume during a snapshot

Quick checks

These are safe, read-only commands. Run them on a server. Adjust the data directory path to match your data_dir config; /opt/consul/data/ is a common default but not universal. Snapshots live under <data_dir>/raft/snapshots/.

# Snapshot size on disk
ls -lh /opt/consul/data/raft/snapshots/
du -sh /opt/consul/data/raft/

# Per-type breakdown of the latest snapshot (KVS, Register, Index, sessions, etc.)
consul snapshot inspect /opt/consul/data/raft/snapshots/<latest-snapshot-file>
# -kvdetails gives KV prefix usage; -kvdepth and -kvfilter control the breakdown

# Catalog and KV counts (rough proxy for FSM size)
curl -s http://127.0.0.1:8500/v1/catalog/services | jq 'length'
curl -s 'http://127.0.0.1:8500/v1/kv/?keys' | jq 'length'

# Raft apply rate and commit time (leader only reports commitTime).
# Note: /v1/agent/metrics returns JSON; grep is a quick filter, not a parsed read.
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep -E 'raft.commitTime|raft.apply'

# Disk write latency during a snapshot window
iostat -x 1 5

# Memory during snapshot save - watch alloc_bytes while a save runs
watch -n 1 "curl -s http://127.0.0.1:8500/v1/agent/metrics | grep runtime.alloc_bytes"

How to diagnose it

  1. Establish the slope. Snapshot size alone is not the signal; the slope is. Collect snapshot file sizes weekly and plot them. If snapshot size is growing faster than your legitimate service count, you have bloat, not growth.

  2. Break the snapshot down by type. Use consul snapshot inspect to see which section of the FSM dominates. A healthy cluster typically has catalog and KV as the largest sections. If Register history or ACL tokens dominate, that is the leak source.

  3. Correlate snapshot size with live state counts. Snapshot size should track roughly with consul.catalog.services, service instance count, KV key count, and session count. If snapshot size grows while these counts stay flat, something is accumulating inside the FSM that is not visible in the live API.

  4. Measure restore time on a test follower. In a staging cluster or during a planned restart, time how long a fresh follower takes from join to voter. If this number has grown from seconds to minutes, you are inside the danger zone for the InstallSnapshot loop. The threshold for concern is any restore that approaches the time it takes the leader to write raft_trailing_logs entries.

  5. Watch commit time during a snapshot save. Snapshots run on the leader. Pull consul.raft.commitTime alongside consul.runtime.alloc_bytes during a snapshot interval. If commit time spikes upward during snapshot save and the snapshot is large, disk contention or memory pressure is the mechanism.

  6. Verify disk latency on the Raft volume. Slow disk is the multiplier that turns a large snapshot into an outage. iostat -x 1 should show await well under 10ms; HashiCorp’s production server starting point is 7500+ IOPS and 250+ MB/s. Anything higher turns snapshot save into a sustained Raft stall.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Snapshot file sizeDirect measure of FSM state sizeMonotonic growth over weeks without corresponding service growth
consul.raft.fsm.lastRestoreDuration (gauge)How long a rejoining follower is unavailableTrending up; any value approaching the time to write raft_trailing_logs entries
consul.raft.commitTimeEnd-to-end write latency, leader onlySpikes that correlate with snapshot intervals
consul.runtime.alloc_bytesWorking set sizeSpike of 1.5x or more during snapshot save
consul.raft.leader.lastContact (leader-side timer)Leader freshness versus followersSpikes during snapshot save on the leader
consul.state.services and service instance countCatalog bloat proxyGrowing faster than infrastructure inventory
KV key count and total KV bytesKV bloat proxyGrowth without corresponding application changes
consul.raft.apply counterTotal Raft write rateApply rate dominated by KV or catalog register ops
Disk write latency (await) on Raft volumeMultiplier that turns large snapshots into outagesSustained above 10ms, or burst credit exhaustion on cloud volumes

The single most important signal is the one most teams do not collect: snapshot size over time. Catalog and KV state grow a few entries per week and nothing looks expensive in isolation, but over months the snapshot balloons, memory grows, restore times creep up, and one day a snapshot save takes long enough to cause a leader election during creation. Track snapshot size as a PLAN-level weekly trend, not a page.

Fixes

Catalog and KV bloat: remove the source

The only durable fix for snapshot growth is removing the state that should not be there. Identify the dominant section of the snapshot via consul snapshot inspect, then go after it:

  • KV used as a database. Move high-throughput or large-value workloads out of Consul KV. Consul KV is for configuration and coordination, not primary application state. Each KV write is a Raft commit replicated to every server and serialized into every snapshot.
  • Zombie service registrations. If ephemeral workloads (CI runners, batch jobs, short-lived pods) register services without deregistering, the catalog accumulates dead instances. Add deregister_critical_service_after to health checks so the catalog self-cleans.
  • Unbounded check count. Each registered check is a row in the FSM. Services that register dozens of per-instance checks (one per dependency, one per endpoint) multiply catalog size. Consolidate checks where possible.

Removing bloat does not shrink existing snapshots immediately. The next snapshot after cleanup will be smaller, but the on-disk snapshot files from before remain until they are rotated out by newer snapshots. Plan cleanup with enough lead time before the next rolling restart.

Raft tuning: break the InstallSnapshot loop

If followers are already stuck in a restore loop because the snapshot is large and write rate is high, the immediate lever is raft_trailing_logs. Increasing it from the default 10000 to a higher value (50000 is a common operational choice, not an officially published universal value) gives the leader more log entries to send via incremental AppendEntries before falling back to a full snapshot install.

Since Consul 1.10, raft_snapshot_threshold, raft_snapshot_interval, and raft_trailing_logs are reloadable at runtime via consul reload (which sends SIGHUP), with no rolling restart required:

# Apply raft tuning changes from config without a restart
consul reload

This does not shrink the snapshot. It widens the window in which a follower can catch up. The real fix is still reducing snapshot size or improving disk I/O.

Disk I/O: remove the multiplier

A large snapshot on slow disk is an outage waiting to happen. The operational guidance is consistent across the Consul community: dedicated SSD for the Raft data directory, nothing else on the volume. If you are on EBS gp2 with exhausted burst credits, the fix is either io1/io2 with provisioned IOPS or a larger gp3 baseline. The signal that matters is write latency (await), not throughput.

Memory headroom for snapshot save

Snapshot save temporarily inflates the working set. Size server memory so that steady-state heap leaves headroom for the snapshot spike and GC overhead. If steady-state heap is already high, a snapshot save can push you into GC thrashing, which then pushes Raft commit times past the heartbeat timeout. A useful rule: plan for snapshot creation to roughly double the effective memory pressure for the duration of the save; validate this heuristic against your state mix.

Prevention

  • Track snapshot size weekly. This is a PLAN-level signal. Plot it next to catalog service count, KV key count, and total KV bytes. The slope tells you whether you have legitimate growth or bloat.
  • Track restore duration on every follower restart. Each rolling restart is a free restore-time sample. Log how long each server takes from join to voter and alert on the trend.
  • Set deregister_critical_service_after on ephemeral service checks. This is the single highest-leverage prevention for catalog bloat in dynamic environments.
  • Enforce a KV value size budget. Alert on any KV value above a threshold (a few KB is a reasonable ceiling for configuration data). Large values in KV are almost always a misuse.
  • Do not colocate the Raft data directory with anything else. No logs, no other databases, no shared volumes. The Raft volume is a single-tenant resource.
  • Run periodic consul snapshot inspect on the latest snapshot and record the per-type breakdown. A sudden change in the dominant section (for example, Register overtaking KVS) is an early signal of a new bloat source.
  • Capacity-plan against snapshot size, not just service count. Server memory and disk headroom should be sized for the projected snapshot size, not the current one.

How Netdata helps

  • Per-second metric resolution exposes the commit-time spike during snapshot save that minute-granular monitoring misses. A snapshot save that pushes consul.raft.commitTime from 20ms to 400ms for 8 seconds is invisible at 1-minute aggregation.
  • Correlated dashboards let you put consul.raft.commitTime, consul.runtime.alloc_bytes, disk await, and consul.raft.leader.lastContact on one view. The shape of a snapshot-induced stall is distinctive: memory and disk latency rise together, commit time follows, last-contact spikes on followers.
  • ML anomaly detection flags the slow trend in snapshot size and catalog counts before they cross a static threshold. Snapshot growth is exactly the kind of slow monotonic drift that threshold-based alerting handles poorly.
  • Disk I/O per device at per-second resolution catches the EBS burst-credit exhaustion pattern that turns a large snapshot into a Raft stall. The disk metric is the multiplier; without it you will diagnose the snapshot as the cause when the disk is the actual bottleneck.
  • Composite alerts on the catalog-bloat pattern (snapshot size rising while service count is flat) catch the silent-catastrophe signature before it surfaces as a restore loop or an election during snapshot save.