The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / consul
CONSUL · OPERATIONS PLAYBOOK

Consul runs three clusters at once: a Raft brain, a gossip nervous system, and a catalog that can silently lie to you

A service-discovery and service-mesh control plane where every write funnels through one elected Raft leader, membership is tracked by an entirely separate gossip layer, and each agent reconciles its own state into a shared catalog that can drift from reality without anything looking broken. We trace how those three subsystems interact, where they disagree, and what to do when they do.

"

Consul makes service discovery look simple — a leader, some agents, a catalog — right up until the moment its three independent subsystems disagree about what "healthy" means.

The defaults work. Until the data_dir volume's write latency creeps up, Raft heartbeats slip, followers time out and call elections, and the cluster thrashes — the single most common Consul incident in the industry. Until a client agent's RPC pipe to the servers breaks while its gossip stays perfectly healthy, so anti-entropy stops syncing and the catalog silently drifts while every dashboard shows green. Until a Connect leaf certificate expires and mTLS handshakes fail across the mesh with no early warning at all. Until a botched keyring rotation leaves half the pool on a key the other half cannot decrypt and the gossip pool splits in two. Until the default 1024 file-descriptor limit is hit and the server refuses RPC, xDS, and DNS in the same instant.

These guides are written for engineers who already run Consul, not for people learning what service discovery is. The goal is the mental model of how the three subsystems actually behave and interact under load, the failure patterns that keep recurring, the monitoring story that catches them before they page anyone, and the runbooks you wish someone had handed you before your last incident. The recurring trap: "is there a leader?" and "is the node alive in gossip?" both answer yes while the cluster is write-dead or the catalog is stale.

How Consul actually runs in production

Consul is not one system. It is three concurrent subsystems — Raft consensus (the brain, where all writes go through one elected leader), Serf gossip (the nervous system, SWIM-based membership on a LAN pool and a WAN pool), and the catalog reconciled by anti-entropy (the truth, held authoritatively on the servers). Connect layers a certificate authority and an xDS stream on top. Most production failures live between these layers, not inside any one of them: a node can be alive in gossip and useless for consensus.

01
clients + applications
Apps, <code>consul-template</code>, and Envoy sidecars query Consul over DNS (port 8600) and the HTTP API (port 8500). Watches use HTTP long-polling (blocking queries), and each one holds a goroutine and a connection on the agent for up to its wait timeout.
CLIENT
02
client agents + anti-entropy
Every node runs an agent. Agents run the node's health checks locally, register its services, and periodically reconcile local state into the catalog via anti-entropy over RPC. When that RPC pipe breaks the agent still looks alive while the catalog quietly goes stale.
AGENT
03
Serf gossip (LAN + WAN)
SWIM-based membership and failure detection over the LAN pool (port 8301) and the WAN pool (port 8302). It carries the automatic <code>serfHealth</code> check for every node and is completely independent of Raft — gossip health is not consensus health.
GOSSIP
04
agent to server RPC
Agents forward writes and consistent reads to the servers over RPC (port 8300). <code>consul.client.rpc.failed</code> is the health of this pipe. Firewalls that block 8300 while leaving 8301 open, an overloaded server RPC handler, or a TLS mismatch all break it silently.
RPC
05
Raft consensus (the leader)
One elected leader takes every write and replicates it to a quorum of voters. <code>consul.raft.commitTime</code> folds disk write latency, replication, and FSM apply into one number; lose the leader or the quorum and all writes stop.
RAFT
06
the FSM: catalog, KV, sessions
The replicated state machine on the servers holds the authoritative catalog, health state, key-value store, sessions, and intentions. Every registration and every KV write is a Raft entry — so catalog churn and KV abuse both land directly on the consensus pipeline.
STATE
07
storage: raft.db + snapshots
Raft appends fsync-heavy log entries to <code>raft.db</code> in the data directory, then truncates the log only after a successful snapshot. Disk write latency (await) on this volume is the hidden driver of leader stability, and a failing snapshot grows the log until the disk fills.
STORE
08
Connect mesh: CA + xDS
The built-in CA issues short-lived leaf certificates; servers stream endpoints, certs, and intentions to each Envoy sidecar over a long-lived gRPC xDS stream. The server-side footprint scales with sidecar count, not service count, and an expired cert is a silent, total outage.
MESH

Why this matters: "Consul is down" or "service discovery is broken" can come from a leaderless Raft, a gossip partition, a broken agent-to-server RPC pipe silently staling the catalog, slow disk driving elections, an expired mesh certificate, or file-descriptor exhaustion. The symptom rhymes but each subsystem has a different signal — and gossip health, consensus health, and agent health are three separate questions.

The failures you'll actually see

Most Consul incidents fall into a small set of recurring patterns. Recognise the shape, and triage gets dramatically faster.

CRITICAL

The leaderless cluster

With no elected leader every write fails: no service registration, no KV write, no session or token creation, no health state reaching the catalog. Reads may still be served, stale. Operators see rpc error making call: No cluster leader everywhere. A 1-5s blip during a rolling restart is expected; sustained leaderlessness past 15-30s almost always means quorum has been lost.

  • rpc error making call: No cluster leader on every write
  • /v1/status/leader returns empty (check several servers — a loser can report a stale leader)
  • /v1/operator/raft/configuration shows voters below the majority
  • KV writes, service registrations, and token creation all timing out at once
Investigate
CRITICAL

Leader thrashing

The cluster keeps electing: consul.raft.state.leader increments repeatedly, each election is a brief write outage, and elections faster than the recovery window make the cluster write-unavailable while technically always having a leader. The number-one driver is slow disk on the Raft volume — fsync latency creeps up, heartbeats slip, followers call elections. More than two elections in ten minutes outside maintenance is systemic.

  • Repeated 'entering leader state' and 'heartbeat timeout reached, starting election' in logs
  • consul.raft.state.leader incrementing more than once outside a rolling upgrade
  • consul.raft.commitTime and disk write latency (await) climbing together
  • Brief, repeating write failures that clear and then return
Investigate
ACTIVE

Silent catalog staleness

The failure domain nobody monitors. A client agent is alive in gossip and running its checks locally, but its RPC pipe to the servers (port 8300) is broken — so anti-entropy cannot push updates and the catalog silently drifts from reality. Consumers route to dead instances while every dashboard is green: servers fine, gossip fine, agent process up. consul.client.rpc.failed is the only signal that catches it.

  • consul.client.rpc.failed sustained non-zero on one or more agents
  • Catalog listing instances that are actually gone, or missing ones that exist
  • 'service not found' and routing to dead endpoints with no alarm anywhere
  • Gossip membership and the servers all reporting perfectly healthy
Investigate
CRITICAL

Service discovery goes dark

Applications stop resolving services: Consul's DNS interface (port 8600) returns SERVFAIL or NXDOMAIN and every consumer that discovers through DNS gets connection errors, timeouts, and cross-mesh failovers. Causes range from overloaded servers (DNS requires catalog lookups) to catalog inconsistency returning empty results for services that exist, to an agent that cannot reach any server. Downstream dnsmasq or systemd-resolved caching can mask or compound it.

  • SERVFAIL / NXDOMAIN from dig @127.0.0.1 -p 8600 <svc>.service.consul
  • Apps reporting connection errors and discovery timeouts fleet-wide
  • consul.dns.domain_query latency spiking or queries outright failing
  • A service known to exist returning empty DNS results
Investigate
IMMINENT

The certificate cliff

Everything in the Connect mesh works until the instant a certificate expires — then mTLS handshakes fail and services simply cannot talk. Leaf certs have short TTLs and should auto-rotate; a CA root has a long TTL but its rotation must be planned. Expiry alone is a countdown with no early warning, so you monitor expiry AND renewal success as separate signals. Clock skew silently makes valid certs 'not yet valid' or 'expired', and when Vault is the CA backend its downtime blocks all new issuance.

  • Envoy sidecars failing mTLS handshakes; services unable to connect
  • Leaf certificates within 1h of expiry, or the CA root within 24h
  • Leaf certs NOT rotating — past 50% of TTL and still unrenewed
  • Clock skew across nodes, or Vault (the CA backend) unreachable
Investigate
ACTIVE

File descriptor exhaustion

A slow bleed toward a hard cliff. Each RPC connection, gRPC/xDS stream, DNS listener, and health-check connection costs a file descriptor; the default ulimit of 1024 is catastrophically low for a server. At the limit the server refuses new connections, xDS streams drop, DNS stops answering, and gossip probes time out — triggering false failure detection on top of everything else. Slow, non-reversing growth points to a connection or stream leak.

  • too many open files in the Consul server logs
  • FD usage above 70% (ticket) and climbing toward the limit (page)
  • New RPC and gRPC connections and DNS queries being refused
  • Gossip probes timing out and nodes marked failed under FD pressure
Investigate
Choosing a tool

Best Consul Monitoring Tools - Ranked & Reviewed

A ranked review of the tools teams actually shortlist here, what each one is genuinely good at, and how the pricing behaves as you scale.

Consul monitoring maturity levels

Consul observability works in four practical levels. Each is a complete operation, not a stepping stone. Pick the level that matches how much your cluster matters. Most production clusters should land at the second level.

Level 1: Survival

Know that something is wrong

Survival monitoring is the floor. With these signals you can answer one question: is the cluster able to accept writes at all? You will not learn why it broke, but you will learn that it broke before your applications do. Survival is enough for dev clusters and non-critical discovery.

  • Leader exists /v1/status/leader non-empty — with none, every write fails.
  • All servers alive in gossip /v1/agent/members shows every server alive, none failed.
  • Agent up + HTTP API responsive The consul process running and /v1/agent/self answering.
  • Raft data directory has free space A full data_dir stops Raft writes or crashes the server.
  • Raft peer count == cluster size Voters equal to the number of servers you deployed.

Level 2: Operational

Diagnose most incidents on your own

Operational monitoring is what most production clusters should target. Survival tells you something is wrong; operational tells you what. With this coverage your team can usually diagnose leader instability, gossip failures, stale catalogs, slow DNS, and resource pressure without escalating.

  • consul.raft.commitTime The single best indicator of Raft health; page above ~500ms sustained.
  • consul.raft.leader.lastContact Followers drifting toward an election before it fires.
  • Leader election count More than one outside a rolling upgrade means thrashing.
  • DNS query latency Maps directly to app startup and failover speed.
  • Health check distribution Passing / warning / critical counts across the catalog.
  • consul.client.rpc.failed per agent The agent-to-server pipe; sustained non-zero staleness kills the catalog.
  • File descriptor usage Cliff-edge: at the limit the server refuses everything.
  • Server memory RSS Runway to the OOM, not just the current value.

Level 3: Mature

Catch problems before they become incidents

Mature monitoring catches problems before they wake anyone up. Goroutines leaking, gossip queues backing up, catalog churn climbing, snapshots bloating, xDS streams reconnecting, a leaf cert quietly failing to rotate. None of these page you on day one. They become page-out incidents on day thirty.

  • Goroutine count trend Monotonic growth is the most reliable leak indicator in any Go program.
  • Serf queue depth (Event/Intent/Query) Non-zero sustained means an agent falling behind on gossip.
  • Catalog register / deregister churn Every registration is a Raft write; churn without deploys is the problem.
  • KV apply latency (consul.kvs.apply) Tracks commit time because every write is a Raft entry.
  • xDS stream count vs sidecar count A high drain rate means proxies running on stale config.
  • Connect cert expiry + renewal success Expiry is a countdown; renewal success is the early warning.
  • Snapshot size trend State bloat that means slow restores and memory spikes.
  • Disk write latency (await) on the Raft volume The number-one hidden cause of leader elections.

Level 4: Expert

Reactive instrumentation after real incidents

Expert signals enter your stack the day after a specific incident proved you needed them. Raft log divergence, per-agent anti-entropy timing, blocking-query distribution, cross-DC failover behaviour, GC disturbing Raft. Most teams never need every signal here. Add the ones your incident history says you do.

  • Raft lastLog.index divergence across servers A corrupt or lagging follower before it ever wins an election.
  • Per-agent anti-entropy sync timing How far each agent's local state trails the catalog.
  • Blocking-query distribution Leaked watches piling up goroutines and connections.
  • Prepared-query cross-DC failover rate Failover masking a complete local-DC service outage.
  • Cross-DC RPC + WAN gossip latency A saturated inter-DC link versus a genuine partition.
  • consul.acl.resolveToken latency Token cache thrashing on every authenticated request.
  • consul.runtime.gc_pause_ns Stop-the-world pauses long enough to delay Raft heartbeats.
  • Session invalidation rate Distributed locks releasing across the cluster at once.

Operating mistakes worth avoiding

The traps Consul teams keep falling into. Each has a clear, well-known fix. Most teams only learn it after an incident.

Colocating the Raft data directory with other I/O

Slow disk is the single most common cause of Consul leader instability. Raft log writes are fsync-heavy, so when the <code>data_dir</code> volume's write latency (await) rises, heartbeats and commits slow, followers start elections, and the cluster thrashes. Monitor disk WRITE LATENCY proactively — not throughput or utilization — and follow HashiCorp's long-standing guidance: a dedicated SSD with nothing else on the volume, never colocated with logs or another database.

Trusting "is there a leader?" as your write-health check

A leader can hold leadership while its commit index is stalled — a full disk, an FSM apply deadlock, a wedged state machine — so the leader-exists check stays green while every write times out or silently fails. Watch commit-index progression and <code>consul.raft.apply</code> rate alongside leadership, not just leadership on its own.

Never monitoring consul.client.rpc.failed on agents

This is the failure domain nobody watches. An agent can be alive in gossip and running its checks locally while its RPC pipe to the servers is broken, so anti-entropy cannot sync and the catalog silently drifts from reality — consumers route to dead instances with no alarm anywhere. Alert on sustained non-zero rates on EVERY agent, not just the servers.

Leaking blocking queries and watches

Consul's watch mechanism uses HTTP long-polling; each blocking query holds a goroutine and a connection for up to its wait timeout (default 5 minutes). <code>consul-template</code> fleets and app watchers can open thousands, and when clients disconnect without ending them cleanly they accumulate until the server runs out of goroutines or file descriptors, then cascades. Watch the goroutine trend against expected blocking-query load.

Leaving the file-descriptor limit at 1024

The default ulimit is catastrophically low for a Consul server. Each RPC connection, xDS stream, DNS listener, and health-check connection costs an FD, and at the limit the server refuses new connections and gossip probes start timing out — triggering false failure detection on top of the outage. HashiCorp recommends at least 65536 for servers; raise both the OS ulimit and the systemd limit.

Confusing gossip health with consensus health and service health

These are three different questions. A node alive in gossip can be useless for Raft consensus, and the automatic <code>serfHealth</code> check going critical marks a node failing in gossip and deregisters ALL of its services at once — which looks like a service-level outage when it is really node-level. Read <code>consul members</code>, Raft configuration, and service health as separate signals.

Using the KV store as a database

Every KV write is a Raft entry replicated to every server and included in every snapshot. High-frequency writes, large values, or deep key trees back up the whole Raft pipeline — rising <code>kvs.apply</code> and <code>commitTime</code>, growing snapshots and memory — and degrade service discovery and health for everyone. Consul KV is for configuration, coordination, and small metadata; move heavy or large-value workloads to a real datastore.

Assuming a passing health check proves anything

"Passing" only means the check returned OK, not that it detects real failure — a check testing the wrong thing passes right through an outage. A check flapping passing to critical also drives a catalog write on every transition, feeding registration storms and making downstream load balancers and meshes thrash. Verify checks test the real dependency, and tune interval, timeout, and <code>DeregisterCriticalServiceAfter</code> so a brief blip does not deregister an instance.

Consul runbooks in this section

Each guide is a focused runbook for one symptom or topic. Pick one when you have an incident, or use the categories to learn the area.

WHERE TO GO NEXT

Setting up Consul monitoring, or putting out a fire?

If you're starting from scratch, the monitoring checklist is the path of least regret. If you're mid-incident, jump straight to the symptom that matches what you're seeing.