The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / consul / consul-client-rpc-failed ▌

Operations Guides

Consul client rpc failed: agents alive but the catalog is going stale

A client agent’s consul.client.rpc.failed counter ticks up. The agent process is running, gossip reports the node as alive, local health checks execute on schedule, and consul members lists the node as healthy. The server cluster looks fine: leadership is stable, Raft commit times are normal, and there is no election noise. Nothing on the standard dashboard is red.

The catalog is going stale anyway.

Consul runs two independent network paths between an agent and the server cluster. Gossip membership flows over LAN Serf (TCP and UDP port 8301). State updates flow over the server RPC pipeline (TCP port 8300). The first can be perfectly healthy while the second is broken, and most dashboards only watch the first. Anti-entropy sync runs on an adaptive interval (about one minute for clusters of 1–128 agents, increasing with cluster size) and pushes local agent state to the server catalog over RPC. When RPC fails, the catalog stops receiving updates from that agent but keeps serving whatever it last knew. Consumers (DNS, HTTP API, load balancer integrations, service mesh sidecars) keep getting answers, just progressively wrong ones.

For the broader architecture, see the mental model. For the full signal list, see the monitoring checklist. This article is the narrow playbook for elevated consul.client.rpc.failed.

What this means

consul.client.rpc.failed is a counter that increments whenever a client agent attempts an RPC against a server and the call fails. The companion counter consul.client.rpc counts every attempt. A healthy agent shows consul.client.rpc.failed flat or near zero, with consul.client.rpc ticking at the anti-entropy cadence plus any watches or DNS forwarding.

When the failed counter climbs, the agent cannot push the following to the catalog:

  • New service registrations and deregistrations
  • Health check transitions (passing to critical, or the reverse)
  • Coordinate updates used by network coordinates
  • KV writes routed through the agent

The agent keeps running checks locally, so a process that crashed on the host is correctly marked critical in the agent’s local state, but that critical state never reaches the catalog. From the catalog’s perspective the instance still looks healthy. Load balancers and service mesh sidecars that read from the catalog keep sending traffic to a dead endpoint.

flowchart LR
  A[Client agent
checks run locally] -->|gossip 8301 OK| B[Server cluster
alive in members] A -->|RPC 8300 FAIL| C[Catalog
stale state] A -->|anti-entropy| C C -->|serves stale| D[DNS / API / LB
route to dead instance]

The asymmetry is the trap. Gossip works, so the node appears alive. The leader is stable, so server dashboards are green. The failure is silent and visible only in two places: the consul.client.rpc.failed counter on the affected agent, and drift between what the agent believes locally and what the catalog exposes.

Common causes

CauseWhat it looks likeFirst thing to check
Firewall blocks 8300 while 8301 stays openOne agent or one subnet suddenly failing RPC after a security-group change; consul members still lists the node as alivenc -zv <server-ip> 8300 from the agent host
Server RPC handler exhausted (file descriptors)Many agents across different subnets fail simultaneously; server logs show accept errors; server open-FD count near its process limitFD usage on servers vs ulimit -n
TLS certificate mismatch with verify_server_hostnameFailures appear after a cert rotation; agent logs show x509: certificate is valid for X, not server.<dc>.consul.<domain>Agent logs for TLS error strings
ACL token lacks permission for the RPC methodFailures scoped to one method (often Coordinate.Update); agent logs show Permission deniedAgent logs for ACL denied strings
Agent has no known serversknown_servers: 0 in consul info; logs show No known Consul servers; common after servers are replaced with new IPsconsul info on the affected agent

A subtle variant: if the agent’s own rate limiter (the limits block) is configured too aggressively, the counter that climbs is consul.client.rpc.exceeded, not consul.client.rpc.failed. The RPC is rejected locally before it leaves the agent, so there is no connection error in the logs. Treat sustained non-zero consul.client.rpc.exceeded as the same class of problem: the agent is not getting state to servers.

Quick checks

Run these on the affected client agent first, then on a server. All are read-only and safe during incidents.

# How many servers does this agent know about?
consul info | grep -A2 "known_servers"

# Is the agent alive in gossip from the server side?
consul members

# Is the RPC port actually reachable from this agent?
nc -zv <server-ip> 8300

# Pull the RPC counters from the agent's telemetry
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep -E "consul.client.rpc"

# Agent self view: known servers and config
curl -s http://127.0.0.1:8500/v1/agent/self | jq '.Stats.consul | {known_servers, server}'

# Check recent anti-entropy outcomes in the agent log
journalctl -u consul --since "10 min ago" | grep -i anti_entropy

# Tail agent logs for RPC, TLS, or ACL errors
journalctl -u consul -f --since "10 min ago" | grep -iE "rpc|tls|x509|permission|anti_entropy"

On a server, run:

# Leader identity and Raft peer count
consul operator raft list-peers

# Server FD consumption
ls /proc/$(pgrep -x consul)/fd | wc -l
cat /proc/$(pgrep -x consul)/limits | grep "Max open files"

# How many client RPC connections is this server holding?
ss -tnp '( sport = :8300 )' | wc -l

How to diagnose it

  1. Confirm the failure is RPC, not gossip. From the affected agent, consul members must show the server nodes as alive. If gossip itself is partitioned, you are looking at a different problem. See Consul gossip flapping or Consul serf queue backlog.
  2. Localize the scope. One agent points at host-level network, cert, or ACL. Many agents in the same subnet point at a firewall or route change. Many agents across subnets point at the server side.
  3. Verify the RPC port end to end. From the agent, nc -zv <server-ip> 8300. A timeout or refused connection narrows the cause to network or server handler.
  4. Inspect the agent’s known server list. consul info shows known_servers. Zero known servers means the agent has lost its server discovery path. This happens when servers are replaced simultaneously with new IPs and the agent’s cached addresses are stale.
  5. Read the agent logs for the failure class. The error string tells you which cause you are dealing with:
    • i/o deadline reached or connection refused: network or server handler.
    • x509: certificate is valid for ...: TLS hostname mismatch.
    • Permission denied: ACL.
    • No known Consul servers: empty server list.
  6. Check server-side capacity. If many agents fail at once, the servers are the bottleneck. Check FD usage, goroutine count, and open FD count against the configured process limit. Check whether limits.rpc_max_conns_per_client is configured (default 100; introduced in Consul 1.6.3) and whether per-source-IP connection caps are being hit.
  7. Check anti-entropy success and drift. Look for sustained anti_entropy errors in the agent log, then compare /v1/agent/services locally with /v1/catalog/service/<name> in the catalog; persistent divergence confirms that the RPC break is causing catalog drift.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
consul.client.rpc.failed rateDirect counter of failed agent-to-server RPCsAny sustained non-zero rate on any agent
consul.client.rpc.exceeded rateAgent-side rate limiter drops RPCs before they leaveAny sustained non-zero rate; correlates with limits misconfiguration
consul.client.rpc total rateBaseline RPC attempt rate; denominator for failure ratioSudden drop suggests the agent stopped attempting (no known servers)
Agent anti_entropy log errors and local-versus-catalog stateConfirms catalog drift from the agent sideSustained errors, or a service/check present locally but missing or stale in catalog
consul members LAN member statesGossip health; should remain stable while RPC failsIf the affected node is not alive, you have a gossip problem, not just RPC
Open FD count and consul.runtime.num_goroutines on serversServer-side capacity that gates RPC acceptFD count growing toward the process limit; goroutine growth sustained
known_servers in consul infoWhether the agent can find any server at allZero
Raft leadership and consul.raft.state.leaderRuling out server-side consensus failureMultiple elections; see Consul leader election storm

For the fuller signal list, see Consul monitoring maturity model.

Fixes

Firewall asymmetry (8301 open, 8300 blocked)

The most common cause. A security-group, host firewall, or network policy change keeps 8301 open for gossip but drops 8300 for RPC.

  • Confirm with nc -zv <server-ip> 8300 from the agent host.
  • Compare against the working baseline. If only the agent’s subnet is affected, the change is scoped to that path.
  • Open TCP 8300 from client agents to server nodes in the direction the connection is initiated.
  • Document the ports together so the next change does not split them. Consul requires 8300 for server RPC and 8301 (TCP and UDP) for LAN gossip as a pair.

Server RPC handler exhaustion

When servers run low on file descriptors, new RPC connections are refused. The client sees i/o deadline reached or connection reset by peer.

  • Check the Consul process open-FD count against its configured limit (/proc/$PID/fd versus /proc/$PID/limits). For write-heavy clusters, HashiCorp says the default ulimit of 1024 must be increased; no universal floor is published.
  • Check for connection leaks: ss -tnp '( sport = :8300 )' connection count vs agent count. A server holding tens of thousands of connections from a handful of agents indicates a leak in those agents.
  • Raise LimitNOFILE (systemd) or ulimit -n and restart the server during a maintenance window. This is disruptive and drops in-flight RPC connections.
  • Investigate the leak separately. Common sources are consul-template blocking queries, leaked watches, and service mesh xDS streams.

TLS certificate mismatch

With tls.internal_rpc.verify_server_hostname = true (the nested form added in Consul 1.12; verify_server_hostname is the deprecated pre-1.12 top-level equivalent), the agent validates that the server certificate is valid for server.<datacenter>.consul.<domain>. A cert valid for a different name produces x509: certificate is valid for X, not server.dc1.consul.example.com.

  • Pull the cert the agent is presenting or trusting and inspect SANs: openssl x509 -in <cert> -noout -text | grep -A1 "Subject Alternative Name".
  • Confirm the cert was issued with the Consul-internal RPC naming convention, not a generic service DNS name.
  • Re-issue and redistribute the cert. Rotating certs on servers without also rotating on agents (or vice versa) is the typical trigger.
  • Verify verify_incoming and verify_outgoing are consistent across the cluster. Mixed modes produce confusing partial failures.

ACL permission denied

The agent has a token but the token lacks permission for a specific RPC method. The classic signature is Coordinate.Update failing repeatedly because the token cannot write coordinates.

  • Pull the agent’s token and check the attached policies.
  • Confirm the token grants service:write (or the equivalent for what the agent registers), node:write, and coordinate-write permissions.
  • If the failure started after an ACL policy change, roll back the policy first, then tighten deliberately.
  • Distinguish from token replication lag in federated DCs. In a secondary DC, a recently created token may not have replicated yet.

No known servers

The agent has nobody to dial. Logs show No known Consul servers and consul info reports known_servers: 0.

  • Common cause: all servers were replaced with new IPs (new subnets, ASG replacement, redeploy) and the agent’s cached server list is stale.
  • Trigger a re-join: add retry_join pointing at the new server addresses, or restart the agent so it re-discovers via gossip.
  • Brief blips during rolling server restarts are normal. Persistent zero is not.

Agent-side rate limiter

If the climbing counter is consul.client.rpc.exceeded rather than consul.client.rpc.failed, the agent’s limits configuration is dropping RPCs locally.

  • Inspect the limits block on the affected agent.
  • Compare the configured RPC rate and burst against the agent’s actual workload (number of services, checks, watches).
  • Tune upward or remove the limit if it was set defensively without considering anti-entropy spikes after recovery.

Prevention

  • Alert on sustained non-zero consul.client.rpc.failed on every agent, not just servers. Server-side dashboards do not show this signal. It only exists on the agent.
  • Alert on consul.client.rpc.exceeded separately. Different cause, different fix, same user-visible symptom.
  • Treat ports 8300 and 8301 as a pair in firewall policy. Any change to one must review the other.
  • Monitor server FD usage with low thresholds. Page at 80% of ulimit -n, plan at 60%. FD exhaustion is cliff-edge and cascades into RPC refusal across the fleet.
  • Run cert rotation as a coordinated procedure, not a server-only task. Test that agent-trusted certs validate server.<dc>.consul.<domain> before rollout.
  • Periodically compare local agent state with catalog state. For a sample service, query both /v1/agent/services on the agent and /v1/catalog/service/<name> on a server. Persistent divergence indicates the RPC pipeline is not keeping up even if consul.client.rpc.failed is quiet.

How Netdata helps

  • Per-second collection of consul.client.rpc.failed, consul.client.rpc, and consul.client.rpc.exceeded on every agent lets you see the failed counter climb before anti-entropy lag becomes user-visible.
  • Correlate agent RPC failures with server-side open-FD count and goroutine count, and Raft commit time on the same timeline to localize the cause in seconds.
  • Composite alerts pair consul.client.rpc.failed > 0 with healthy consul.serf.lan.members to surface the silent catalog staleness signature directly, rather than waiting for downstream consumer complaints.
  • Per-agent dashboards make it cheap to spot the difference between one host with a bad cert and a whole subnet behind a bad firewall rule.