The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / consul / consul-kv-write-latency-high ▌

Operations Guides

Consul KV write latency high: every write is a Raft commit

A KV PUT that used to return in 20ms is now taking 500ms, 2s, or timing out entirely. Applications stall on locks, configuration writes queue, and leader election alerts may be firing. Consul’s KV is not a standalone subsystem. Every KV write is a Raft log entry that must be replicated to a quorum of servers, fsynced to disk on each, and applied to the in-memory state machine before the API call returns. KV write latency is a direct reflection of Raft commit health.

The metric that captures this is consul.kvs.apply, a timer reported only on the leader. It measures end-to-end write latency for KV operations. When it tracks closely with consul.raft.commitTime, the problem is the Raft pipeline itself: disk I/O, network replication, or FSM apply cost. When KV latency is far above commit time, something between the client and Raft is queueing: RPC admission control, connection limits, or goroutine saturation.

Sustained consul.kvs.apply above 200ms warrants investigation. Above 1 second, Raft health itself is questionable and you may be one disk spike away from leader elections.

What this means

Consul’s KV store lives inside the Raft finite state machine. There is no separate write path. A PUT /v1/kv/<key> does the following:

  1. The request lands on whichever server the client contacted. If that server is not the leader, it forwards the write.
  2. The leader appends the KV mutation as a new Raft log entry.
  3. The leader replicates that entry to all followers in parallel and waits for a quorum to acknowledge.
  4. Each server in the quorum fsyncs the log entry to disk in the Raft data directory.
  5. Once committed, the entry is applied to the in-memory FSM. Consul blocks the write response until both commit and apply are complete.
  6. The API call returns to the client.

Every step is on the critical path. A slow disk on any quorum member inflates step 4. A saturated network between leader and a follower inflates step 3. A large KV value or a complex transaction inflates step 5 because the FSM must deserialize and index the new state. And because all state mutations share this pipeline, KV writes compete for Raft throughput with service registrations, health check updates, session operations, and ACL changes.

flowchart TD
    A[Client KV PUT] --> B{On leader?}
    B -- No --> C[Forward to leader]
    C --> D
    B -- Yes --> D[Append Raft log entry]
    D --> E[Replicate to followers]
    E --> F[Fsync to disk: leader + quorum]
    F --> G[Apply to FSM: KV state store]
    G --> H[Return to client]
    F -.->|slow disk| X[Inflated commitTime]
    E -.->|network latency| X
    G -.->|large values / deep trees| Y[Inflated consul.fsm.kvs]
    D -.->|RPC admission / conn limits| Z[Queueing before Raft]

Common causes

CauseWhat it looks likeFirst thing to check
Disk I/O latency on the Raft volumeconsul.raft.commitTime tracks KV latency closely; iostat await elevated; possible leader electionsiostat -x 1 on the leader’s Raft volume
Raft pipeline saturation from all writesKV latency tracks commit time; consul.raft.apply rate elevated; health check or catalog churn visibleconsul.raft.apply counter trend
Large KV valuesKV latency and consul.fsm.kvs both elevated; snapshot size growing; specific keys responsibleEnumerate key sizes and request sizes; the limit applies to the KV request body, default 512KB
RPC admission or connection queueingKV latency far above consul.raft.commitTime; goroutines or FDs near limit; writes slow but stale reads fastCompare consul.kvs.apply to consul.raft.commitTime directly
Leader instabilityIntermittent spikes correlated with leader transitions; “no cluster leader” errorsconsul.raft.state.leader increments and lastContact
Deep or wide key treesPrefix scans and blocking queries slow; writes to high-churn prefixes expensiveGET /v1/kv/?keys and check tree shape

Quick checks

Run these on the leader first. All are read-only except the latency probe, which writes a small value to an isolated diagnostic key.

# Identify the leader
curl -s http://127.0.0.1:8500/v1/status/leader

# Time a single KV write (writes a small probe value to a diagnostic key)
time curl -s -X PUT -d 'probe' http://127.0.0.1:8500/v1/kv/__diag__/latency-probe

# Time a stale read (should be sub-ms locally on the leader)
time curl -s http://127.0.0.1:8500/v1/kv/__diag__/latency-probe

# Time a consistent read (costs a leader verification round-trip)
time curl -s "http://127.0.0.1:8500/v1/kv/__diag__/latency-probe?consistent"

# Check consul.kvs.apply telemetry on the leader
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep -A3 "kvs.apply"

# Check consul.raft.commitTime on the leader
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep -A3 "raft.commitTime"

# Check lastContact on a follower (run on a follower server)
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep -A3 "raft.leader.lastContact"

# Enumerate KV key count
curl -s "http://127.0.0.1:8500/v1/kv/?keys" | python3 -c "import sys,json; print('Keys:', len(json.load(sys.stdin)))"

# Disk I/O on the Raft volume (identify the backing device first)
iostat -x 1 5

The critical comparison is between consul.kvs.apply and consul.raft.commitTime. If they track each other, the Raft pipeline is the bottleneck. If KV latency is meaningfully higher, look upstream at RPC handling, connection saturation, or goroutine pressure.

How to diagnose it

1. Confirm whether the problem is Raft or KV-specific.

Pull both consul.kvs.apply and consul.raft.commitTime from the leader. If commit time is the dominant cost (KV latency tracks commit time within a small margin), the Raft pipeline is saturated or slow. If KV latency exceeds commit time by a wide margin, the delay is in RPC admission, connection handling, or goroutine pressure before the write reaches Raft.

2. Check disk I/O on the leader and all quorum members.

Raft commit time is dominated by fsync latency. Run iostat -x 1 on each server and watch the await column for the device backing the Raft data directory (typically <data_dir>/raft/). Sustained write latency above 10ms on the Raft volume is a problem. Common causes: EBS gp2 burst credit exhaustion, shared or network-attached storage, spinning disks, noisy neighbors on the volume. This is the single most common root cause of high KV write latency in production.

3. Check the Raft apply rate and what is driving it.

KV writes are not the only thing flowing through Raft. Service registrations, health check state changes, sessions, and ACL operations all share the pipeline. Check consul.raft.apply and the consul.catalog.register / consul.catalog.deregister operation rates. If catalog churn is high, health check flapping or a registration storm is consuming Raft throughput that KV writes need. See the related guides on catalog bloat, registration storms, and gossip flapping.

4. Inspect the KV store for large values and key tree shape.

Large KV request bodies inflate the Raft log entry, slow the fsync, increase FSM apply cost, and grow snapshot size. The default limits.kv_max_value_size caps the request body at 512KB, but even values in the tens of KB written frequently degrade performance. Enumerate keys and check sizes:

# List keys sorted by value size, largest first
curl -s "http://127.0.0.1:8500/v1/kv/?recurse" | \
  python3 -c "
import sys,json,base64
items=json.load(sys.stdin)
def raw(i):
    try: return len(base64.b64decode(i.get('Value') or ''))
    except Exception: return 0
for i in sorted(items, key=raw, reverse=True)[:20]:
    print(f'{raw(i):>8} bytes  ' + i['Key'])
"

Also check whether any single prefix has an unusually deep or wide tree. Blocking queries on high-churn prefixes and prefix scans over wide trees are expensive even when individual writes are fast.

5. Check leader stability and lastContact on followers.

If KV latency spikes correlate with leader transitions, the leader is unstable. Check consul.raft.state.leader for increments (each increment is a transition to leader) and consul.raft.leader.lastContact on followers. Rising lastContact toward the election timeout indicates the leader is struggling to heartbeat. This often co-occurs with disk I/O problems but can also be caused by Go GC pauses on large heaps.

6. Check RPC admission and resource saturation.

If KV latency exceeds commit time, check goroutine count (consul.runtime.num_goroutines) and file descriptor usage on the server. RPC rate limiting, connection pool exhaustion, and goroutine accumulation from leaked blocking queries all add latency before the write reaches Raft.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
consul.kvs.applyEnd-to-end KV write latency, leader only. Primary symptom metric.Sustained above 200ms (TICKET) or 1s (PAGE)
consul.raft.commitTimeTime to commit a Raft entry to disk on the leader. Raft health proxy.Tracking upward with kvs.apply, or above 100ms sustained
consul.fsm.kvsTime to apply a committed KV operation to the state machine.Spikes on large values or complex transactions
consul.raft.leader.lastContactTime since follower heard from leader. Predictive of elections.Trending toward election timeout
consul.raft.apply (counter)Raft log entries applied per interval. Write throughput.Sudden 3x+ spike without new services (churn source)
Disk write latency (await)fsync is the dominant cost in Raft commit.Sustained above 10ms on the Raft volume
consul.runtime.num_goroutinesProxy for concurrent load and leak detection.Monotonic growth without load increase
Raft snapshot sizeTracks total state size including KV.Growing week over week without explanation
consul.raft.state.leader (counter)Increments each time this server becomes leader. Transitions indicate elections.Multiple transitions per 10 minutes outside maintenance

Fixes

Disk I/O is the bottleneck

This is the most common cause. Move the Raft data directory to a dedicated, faster volume. For cloud deployments, use provisioned IOPS (io1 or io2 on AWS) rather than gp2, which can exhaust burst credits and hit a sudden latency cliff. Never colocate the Raft volume with other I/O-heavy workloads like application logs or databases. On bare metal or VMs, dedicated SSDs are the baseline expectation, not an optimization.

If you cannot immediately migrate storage, reduce write pressure on Raft by throttling the heaviest write sources: flapping health checks, excessive service registration churn, or applications writing to KV at high frequency.

Raft pipeline is saturated by non-KV writes

KV latency is a victim of overall Raft throughput. If catalog churn is the driver, address the source. See the related guides on catalog bloat, registration storms, and gossip flapping. Temporary mitigations include increasing health check intervals, reducing the number of registered checks, and ensuring deregister_critical_service_after is set to prevent stale entries from accumulating.

Large KV values or KV-as-database usage

Consul KV is designed for configuration, coordination, and small metadata. It is not a general-purpose datastore. If applications are writing large values (approaching even tens of KB) or writing at high frequency, move that workload to an appropriate datastore. If you cannot move it immediately:

  • Reduce request size by storing only references or pointers in KV.
  • Clean up ephemeral or temporary keys that accumulate over time.
  • Consider batching related writes using Consul transactions (PUT /v1/txn), though each transaction is still a single Raft commit.

RPC admission or connection saturation

If KV latency exceeds commit time, the queueing is before Raft. Check and raise limits.rpc_rate and limits.rpc_max_burst if they are too restrictive for your cluster size. Raise file descriptor limits via ulimit -n or LimitNOFILE in systemd. Investigate goroutine leaks from abandoned blocking queries or watches. Each blocking query holds a goroutine and a connection for its entire duration, and leaks compound over weeks.

Leader instability

If leader elections are recurring, the root cause is almost always disk I/O or resource starvation on the leader, not a Raft configuration problem. Fix the disk first. As a last resort during an active incident, you can transfer leadership to a known-healthier server with consul operator raft transfer-leader, but this only buys time if the underlying resource issue persists on the new leader.

Prevention

  • Monitor disk write latency on all server Raft volumes as a PAGE-level signal. This is the leading indicator. Do not wait for leader elections to tell you the disk is slow.
  • Track consul.kvs.apply and consul.raft.commitTime together. Their relationship is the fastest diagnostic. Alert when either crosses the TICKET threshold.
  • Track Raft snapshot size over time. Growing snapshots indicate state growth (KV, catalog, or checks) that will eventually degrade commit and restore performance.
  • Establish a KV usage policy. Define maximum value sizes, maximum key counts, and prohibited use cases. Enforce it in code review.
  • Audit goroutine count and FD usage weekly. Slow leaks from blocking queries and watches compound over weeks and inflate KV latency through RPC queueing.
  • Review limits.kv_max_value_size if you need to cap request size. The default is the Raft suggested maximum of 512KB. Lowering it can prevent accidental large writes from destabilizing the cluster, but test carefully: the configuration reference warns that improper tuning can affect leadership stability.

How Netdata helps

  • Per-second consul.kvs.apply and consul.raft.commitTime surface the correlation between KV latency and Raft commit latency at the resolution needed to catch transient spikes that minute-aggregated metrics miss.
  • Anomaly detection on Raft commit time catches gradual drift that static thresholds miss, before KV latency becomes user-visible.
  • Disk I/O metrics per device (await, %util, read/write latency) alongside Consul metrics confirm or rule out disk saturation in a single view.
  • Leader election events correlated with KV latency spikes distinguish “Raft is slow” from “leadership is unstable” without manual log correlation.
  • Goroutine count and file descriptor trends reveal slow leaks from blocking queries that inflate KV latency through RPC queueing rather than Raft itself.