The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / consul / consul-raft-last-contact-high ▌

Operations Guides

Consul raft lastContact rising: followers drifting toward an election

consul.raft.leader.lastContact is the predictive signal before a Raft election fires. It measures the elapsed time since the leader last successfully contacted each follower server. When this value rises, a follower is drifting toward the election timeout. Cross that threshold and the follower starts a new election: writes stall, clients see “no cluster leader” errors, and downstream consumers retry in a thundering herd.

Most Consul metrics tell you something already broke. lastContact tells you something is about to break. A healthy cluster holds this value in the low tens of milliseconds. Trending above 200ms warrants investigation. Sustaining above 500ms means an election is imminent. The election timeout sits at approximately 1000ms with raft_multiplier set to 1 (the production recommendation). Higher multipliers raise the timeout proportionally.

Default raft_multiplier is 5 (since Consul 0.7); setting it to 1 gives highest-performance mode equivalent to pre-0.7 timing

The diagnostic pattern matters more than any single reading. All followers trending upward means the leader cannot send heartbeats on time. One follower drifting while others stay flat means the network path to that specific follower is degrading. These two scenarios have different root causes and different fixes.

What this means

consul.raft.leader.lastContact is a timer reported in milliseconds. It captures the gap between the current moment and the last successful heartbeat or AppendEntries acknowledgment from each follower. The leader sends heartbeats from the same goroutine loop that processes Raft log writes, so anything blocking that goroutine (disk fsync, GC pause, CPU starvation) delays heartbeats and inflates lastContact.

flowchart TD
    A["Healthy: lastContact under 50ms"] --> B["Drifting: 200-500ms, trending up"]
    B --> C["Imminent: over 500ms sustained"]
    C --> D{"Reaches election timeout?"}
    D -->|Yes| E["Follower starts election"]
    D -->|No, heartbeats resume| B
    E --> F{"Stable new leader?"}
    F -->|Yes| A
    F -->|No| G["Leader thrashing: repeated elections"]
    G --> C

Key properties:

  • Moves between servers on leadership change. In Prometheus, the time series may briefly vanish during the transition as the old leader stops emitting and the new leader starts.
  • Zero values are the leader itself. Filter them in Prometheus with consul_raft_leader_lastContact != 0.
  • Per-follower granularity is the whole point. A single follower with high lastContact is a network story. All followers with high lastContact is a leader story.

The relationship between lastContact and the election timeout is what makes this metric predictive. When lastContact on any follower approaches the election timeout, that follower will request a new election. With raft_multiplier set to 1, the election timeout is approximately 1000ms. The 500ms page threshold gives you roughly half the timeout window to react before an election fires.

Autopilot adds a second threshold. Its LastContactThreshold defaults to 200ms. When a server exceeds this, Autopilot marks it unhealthy, and CleanupDeadServers may remove it from the Raft configuration. A lastContact spike can therefore cascade from “slow follower” to “removed peer” to “reduced quorum margin” in a single cleanup cycle.

Confirmed: hashicorp/raft#524 is open (no library fix yet); affects Consul >= 1.13.0; root cause is a check introduced in Raft v1.3.7; workaround is to reset the returning node

Common causes

CauseWhat it looks likeFirst thing to check
Leader disk I/O saturationAll followers drifting; commitTime elevated; disk await high on data volumeiostat -x 1 5 on the leader
Network degradation to one followerSingle follower drifting while others stay lowPairwise latency between leader and the affected follower
Leader GC pausesAll followers drifting in bursts; gc_pause_ns spikes on leaderGC pause metrics on the leader
Leader CPU starvationAll followers drifting; CPU pinned near 100%Process CPU usage and container limits
Autopilot premature peer removallastContact exceeds 200ms; server disappears from Raft peersAutopilot configuration and cleanup dead servers setting

Quick checks

All safe, read-only operations. Run on the current leader unless noted. Adjust scheme and port if your API listener uses TLS or a non-default port.

# Identify the current leader
curl -s http://127.0.0.1:8500/v1/status/leader

# Check lastContact values
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep lastContact

# Check commit time (correlates with leader pressure)
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep commitTime

# List Raft peers and voter status
consul operator raft list-peers

# Check disk I/O latency on the leader's data volume
iostat -x 1 5

# Check GC pause duration on the leader
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep gc_pause

# Check Autopilot configuration (LastContactThreshold, CleanupDeadServers)
curl -s http://127.0.0.1:8500/v1/operator/autopilot/configuration

# Check gossip membership (gossip and Raft can disagree)
consul members

# Pairwise network check to a specific follower
ping -c 10 <follower-address>

How to diagnose it

Step 1: Determine scope. Pull lastContact from the current leader. Are all followers trending upward, or is it isolated to one?

  • All followers high: The leader cannot send heartbeats fast enough. The problem is on the leader.
  • One follower high: The network path between the leader and that follower is degraded, or the follower itself cannot process incoming RPCs.

Step 2: If all followers are high, check leader-side causes.

Run iostat -x 1 5 on the leader. Look at await (write latency) and %util on the volume hosting the Raft data directory. Sustained await above 10ms is a red flag. The Raft goroutine that sends heartbeats also handles log writes. When fsync blocks, heartbeats are delayed.

Check consul.raft.commitTime on the leader. If commitTime is also elevated, the Raft write pipeline is saturated. Disk I/O is the most common cause. If commitTime is normal but lastContact is high, suspect GC pauses or CPU starvation.

Check consul.runtime.gc_pause_ns. Stop-the-world GC pauses block the Raft goroutine directly. Even tens of milliseconds of GC pause can push lastContact past the 200ms Autopilot threshold if they coincide with heartbeat intervals.

Step 3: If one follower is high, check the network path.

Run pairwise latency checks between the leader and the affected follower. Use ping, mtr, or tcpdump on port 8300 (the Raft RPC port). Look for packet loss, asymmetric routing, or latency spikes that do not appear on paths to other followers.

Check the follower’s own resource usage. A follower that is CPU-starved or disk-saturated may not process AppendEntries RPCs fast enough, even if the network is fine.

Step 4: Check for Autopilot interaction.

Review the Autopilot configuration. If CleanupDeadServers is enabled and LastContactThreshold is at the default 200ms, a lastContact spike can trigger peer removal. Check consul operator raft list-peers to verify all expected voters are present. A missing peer after a lastContact spike suggests Autopilot cleanup.

Step 5: Check for version-specific issues.

If the metric seems absent rather than elevated, verify your Consul version. Earlier releases had telemetry library bugs that could suppress metric emission under certain conditions.

Confirmed: hashicorp/consul#15231 describes elections taking 3–15 minutes after a node eviction starting from 1.13.x; resolution status unclear

Metrics and signals to monitor

SignalWhy it mattersWarning sign
consul.raft.leader.lastContactPredictive election indicatorTrending above 200ms (ticket); above 500ms sustained (page)
consul.raft.commitTimeLeader write pipeline healthElevated alongside lastContact on all followers confirms leader pressure
consul.raft.state.leaderElection countMore than 2 transitions per 10 minutes indicates thrashing
consul.runtime.gc_pause_nsGC stop-the-world pausesSpikes coinciding with lastContact jumps identify GC as root cause
Disk I/O await on data volumeRaft log fsync latencySustained above 10ms; leading indicator for disk-driven elections
Per-follower lastContact breakdownIsolates follower-specific driftOne follower high while others normal signals network path problem
consul.raft.peers (gauge)Voting peer countDrop below expected cluster size means Autopilot removed a peer

Fixes

Leader disk I/O saturation

This is the single most common cause of lastContact spikes. The Raft goroutine sends heartbeats from the same loop that handles log writes. When fsync on the data directory blocks, heartbeats stop until the write completes.

Immediate actions:

  • Identify the source of write churn. Health check flapping and KV write storms generate excessive Raft log entries. Check consul.catalog.register and consul.raft.apply rates for abnormal spikes.
  • If running on AWS EBS gp2, check burst credit balance. Exhausted credits cause sudden latency cliffs with no warning in Consul metrics.

Structural fixes:

  • Move the Raft data directory to a dedicated SSD volume. Never colocate with application logs, other databases, or noisy neighbors.
  • Use provisioned IOPS rather than burst-dependent volumes. The cost difference is minor compared to the cost of a leader election storm.

Network path degradation to one follower

When only one follower shows high lastContact, the leader is fine. The problem is between the leader and that specific follower.

  • Check pairwise latency between all server pairs, not just leader to follower. Asymmetric partitions (A reaches B, but B cannot reach A) are common and easy to miss if you only test one direction.
  • Verify firewall rules on port 8300 (Raft RPC). A security group change can block Raft traffic while leaving gossip (port 8301) untouched, creating a split view where the node looks alive in gossip but is unreachable for consensus.
  • Use mtr to identify where packets are lost or delayed in the path.
  • Check the follower’s own CPU and disk. A follower that cannot process AppendEntries fast enough will show high lastContact even with a healthy network.

Leader GC pauses

Large Go heaps with high allocation rates produce stop-the-world pauses that block the Raft goroutine. Each pause inflates lastContact on all followers simultaneously.

  • Check consul.runtime.alloc_bytes and consul.runtime.heap_objects trends. Growing heaps with no corresponding service growth indicate a leak.
  • Correlate goroutine count with heap growth. Blocking query leaks and watch handler accumulation are common drivers of heap pressure.
  • Consider tuning GOGC. The default of 100 may cause frequent pauses on large heaps. Increasing it reduces pause frequency at the cost of higher baseline memory usage.

Autopilot cleanup cascade

When LastContactThreshold (default 200ms) is exceeded, Autopilot may remove the server from the Raft peer set. This reduces quorum margin and, in Consul 1.13+, has been associated with election livelock when the removed server reconnects and requests votes with higher terms.

  • Review whether CleanupDeadServers is appropriate for your environment. In networks with periodic latency spikes, aggressive cleanup can destabilize the cluster.
  • Consider raising LastContactThreshold if your baseline network latency between servers is naturally elevated (cross-AZ, cross-region).
  • Monitor consul.raft.peers count. An unexpected drop after a lastContact spike confirms Autopilot removal rather than a crash.

Leader CPU starvation

In containerized deployments, insufficient CPU limits cause scheduling delays that propagate to the Raft goroutine.

  • Check container CPU limits against actual usage. Consul servers need burst capacity for Raft processing, TLS handshakes, and gossip encryption.
  • CPU pinned at 100% causes gossip probe timeouts and heartbeat delays simultaneously. lastContact and gossip member health degrade together.
  • Separate the server workload from other containers on the same host. Noisy neighbors steal CPU cycles at the worst possible moment.

Prevention

  • Monitor lastContact at per-second resolution. Sub-minute drift that coarser aggregation smooths over is exactly the drift that precedes an election.
  • Alert on trend, not just threshold. A static threshold of 500ms catches the moment before disaster. A trend alert at 200ms catches the drift that leads there. Use both, but weight the trend.
  • Track disk I/O write latency as a first-class signal. The await metric on the Raft data volume is the leading indicator for disk-driven elections. Alert on sustained values above 10ms.
  • Keep Raft data on dedicated SSDs. This is the single most impactful infrastructure decision for Consul stability.
  • Monitor GC pause distributions, not just heap size. Go’s GC can keep heap stable while spending significant time in stop-the-world pauses. Track consul.runtime.gc_pause_ns percentiles.
  • Verify Autopilot thresholds match your network. The default 200ms LastContactThreshold assumes low-latency LAN. Cross-AZ or cross-region deployments may need a higher threshold.
  • Pair lastContact with election count monitoring. If consul.raft.state.leader increments alongside lastContact spikes, you are already in a thrashing pattern.

How Netdata helps

  • Per-second resolution catches the sub-minute drift that precedes an election. The trend from 50ms to 200ms can develop in seconds; coarser aggregation smooths it over.
  • Correlation across metrics distinguishes disk saturation from GC pressure from CPU starvation. When all followers drift, lastContact alongside commitTime, GC pauses, and disk await in a single view narrows the cause without switching tools.
  • Anomaly detection flags gradual creep (for example, 30ms to 150ms over an hour) that static thresholds miss but that still precedes an election.
  • Multi-node correlation handles the metric’s migration during leadership changes. When leadership moves, the metric follows the new leader.
  • Disk I/O on the same node completes the root cause picture. The await on the leader’s data volume is one correlation away.