The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / consul / consul-acl-resolution-latency ▌

Operations Guides

Consul ACL resolution latency: token cache thrashing on every request

Every authenticated Consul request pays an ACL resolution tax. When that tax is sub-millisecond, nobody notices. When the token cache cannot hold the working set, every request becomes a miss and that tax multiplies across DNS lookups, HTTP API calls, RPC forwarding, and Connect intention evaluation. The cluster keeps answering, but slowly, and the slowness is uniform across every authenticated path.

The primary signal is consul.acl.ResolveToken (a summary timer, milliseconds). Healthy cached resolution is sub-millisecond; a cold miss against the authoritative datacenter costs low single-digit milliseconds. Sustained values above ~10ms are TICKET-worthy: ACL resolution sits on the hot path of every authenticated operation. It surfaces as elevated DNS latency, slow HTTP API responses, and stretched Connect sidecar handshakes long before any single subsystem fails.

Consul keeps compiled token authorizers in fixed-size 2Q LRU caches keyed on the token’s SecretID. When the number of unique active tokens exceeds the cache capacity, or when entries expire before reuse, the cache thrashes and every resolution falls through to the authoritative datacenter or the local server’s state store. The fixes are TTL tuning, down_policy selection, and reducing unique-token cardinality, not throwing more hardware at the servers.

What this means

ACL resolution runs inline on every request that carries a token. That includes:

  • HTTP API calls with X-Consul-Token (the token query parameter form is deprecated)
  • DNS queries from agents configured with an ACL token
  • RPC forwarding from client agents to servers
  • Connect intention evaluation, which resolves source and destination identities
  • Prepared queries that execute with a token

When consul.acl.ResolveToken is elevated, all of these slow down together. The symptom is rarely “ACL is broken.” It is “everything is a bit slow, and the slowness tracks request volume.”

Two failure-direction settings determine what happens when resolution itself fails (as opposed to being slow):

  • fail-closed: acl.down_policy variants that deny on failure reject legitimate traffic when the ACL subsystem cannot reach the authoritative source. This surfaces as 403s and connection resets.
  • fail-open / extend-cache: stale cached authorizations continue to be served. This keeps traffic flowing but may allow tokens that have since been revoked.

The default acl.down_policy is extend-cache. For multi-datacenter setups the async-cache value, introduced in Consul 1.2.1, performs asynchronous refreshes when a cached entry’s TTL expires, preventing a thundering herd of blocking RPCs from hitting the primary datacenter simultaneously when TTLs lapse.

flowchart TD
    A[Authenticated request] --> B{Token in cache?}
    B -- yes --> C[Return cached identity
sub-ms] B -- no --> D[Resolve against authoritative DC] D --> E{Primary DC reachable?} E -- yes --> F[RPC round-trip
low single-digit ms] E -- degraded WAN --> G[Thundering herd
all agents refresh at once] G --> H[Latency multiplies
10ms+ sustained] F --> I[Cache result until TTL] C --> J[Request proceeds] I --> J H --> J

Common causes

CauseWhat it looks likeFirst thing to check
Token cardinality exceeds cache capacityconsul.acl.token.cache_miss rate near or above cache_hit rate; resolution latency tracks unique-token countUnique token count vs cache size
TTLs too short (defaults are 30s)Periodic latency spikes every TTL window; miss rate spikes in syncacl.policy_ttl, acl.role_ttl, acl.token_ttl
token_ttl forgotten when policy_ttl was raisedMiss pattern persists even after raising policy TTLAll three TTL settings together
Deeply nested policy or role inheritanceMiss latency is high even with a warm cache; cost scales with inheritance depthPolicy and role structure for the slowest tokens
ACL replication lag in secondary DCsMisses in secondary DC are slow because they wait on the primaryGET /v1/acl/replication lag
WAN link degradation without async-cacheCoordinated miss storms when TTLs lapse across many agents at onceacl.down_policy value and WAN latency
Token leakage (accumulating stale tokens)Token count grows without bound; Raft commit latency risesToken inventory and Consul version
/acl/login burst during scale-upLogin calls hang for tens of seconds; goroutines pile up on serversWhether auth-method login is used for Connect injection

Quick checks

# Resolution latency (summary timer, milliseconds)
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep -i 'acl.*resolve'

# Cache hit and miss counters
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep -i 'acl.token.cache'

# ACL replication status (secondary DCs only)
curl -s http://127.0.0.1:8500/v1/acl/replication | jq .

# Total token count (requires management token)
curl -s -H "X-Consul-Token: $CONSUL_HTTP_TOKEN" http://127.0.0.1:8500/v1/acl/tokens | jq length

# Current TTL and down_policy configuration
consul info | grep -iE 'acl|ttl|policy'

# Check for ACL-denied events in the recent logs
journalctl -u consul --since '1 hour ago' | grep -iE 'permission denied|acl not found' | wc -l

Metric names render differently depending on format. The native JSON telemetry from /v1/agent/metrics uses dot notation (consul.acl.ResolveToken); the Prometheus-format endpoint converts dots to underscores. Consul exposes consul.acl.token.cache_hit and consul.acl.token.cache_miss counters, plus consul.acl.token.upsert as a timer. Older Consul versions emitted consul.acl.ResolveTokenToIdentity as a separate timer; which Consul 1.12.0 stopped reporting in favor of ResolveToken, so dashboards built before that version may silently stop graphing.

How to diagnose it

  1. Confirm the latency is ACL-bound. Pull consul.acl.ResolveToken percentiles over the incident window. If the p99 tracks overall HTTP and DNS latency spikes, ACL resolution is on the critical path. If HTTP latency is high but ResolveToken is flat, look elsewhere (Raft, catalog bloat, DNS).

  2. Compute the cache hit ratio. Compare consul.acl.token.cache_hit to consul.acl.token.cache_miss over a representative window. A healthy deployment should see hits dominate by at least an order of magnitude. A ratio approaching 1:1 means the cache cannot hold the working set.

  3. Check the unique token population. List tokens and compare against your expected service and operator count. If the count is far higher than expected, look for token leakage (automation creating tokens without cleanup) or a known bug. Issue #22613 reports login-derived token accumulation on Consul 1.17.3; it remained open at review time, so compare your version and release notes before attributing growth to it.

  4. Verify all three TTLs. Operators frequently raise acl.policy_ttl and forget acl.token_ttl (both default to 30s). If only one is raised, the other still expires every 30s and drives the miss pattern. Check all three: acl.policy_ttl, acl.role_ttl, acl.token_ttl.

  5. In secondary DCs, check replication lag. GET /v1/acl/replication shows lag and status. Lag above a few seconds means misses in the secondary DC must round-trip to the primary, multiplying latency. Healthy steady-state replication should be sub-second for tokens.

  6. For federated deployments, check down_policy and WAN latency. If acl.down_policy is not async-cache, a TTL lapse across many agents produces a synchronized refresh storm. Switching to async-cache serves stale entries while refreshing asynchronously.

  7. If login-derived tokens are involved, check for /acl/login contention. Issue #15157 reports intermittent extreme /acl/login latency during pod scale-up on Consul 1.12.6. The reported workaround is to use pre-generated static or service-identity tokens for Connect injection rather than auth-method login at pod startup.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
consul.acl.ResolveToken (p50, p99)Direct measurement of the per-request ACL taxp99 sustained above 10ms; p50 above 1ms
consul.acl.token.cache_hit / cache_miss ratioIndicates whether the cache holds the working setHit ratio below ~90%, or miss rate approaching hit rate
consul.acl.token.upsert rateToken creation rate; leaks and bursts surface hereSustained creation outside deploy windows
consul.acl.replication lag (secondary DCs)Determines miss cost in secondariesLag above 1s, or trending upward
HTTP 403 rateSurfaces fail-closed behavior and token revocation eventsSpike not correlated with an intentional policy change
Raft commit time (consul.raft.commitTime)Token leakage can inflate Raft commit latencyCommit time rising alongside token count
Goroutine count/acl/login contention piles up server goroutinesSpike during scale-up events that does not subside
WAN gossip healthUnderlying connectivity for cross-DC resolutionRemote DC members missing or suspect

Fixes

Raise the three TTLs together

The default for acl.policy_ttl, acl.role_ttl, and acl.token_ttl is 30 seconds each. For most production deployments this is conservative. Raising all three to a few minutes reduces the miss rate dramatically with minimal security cost, because revoked tokens are still enforced: a revocation invalidates the cache entry promptly regardless of TTL.

The common mistake is raising only policy_ttl. If token_ttl stays at 30s, the token cache entry still expires every 30s and you see no improvement. Set all three explicitly.

acl {
  policy_ttl = "5m"
  role_ttl   = "5m"
  token_ttl  = "5m"
}

Tradeoff: longer TTLs mean a revoked token stays usable for longer on agents that cannot reach the authoritative source when combined with extend-cache or async-cache. For most workloads this is acceptable. For high-security environments, keep TTLs short and instead attack the miss rate through token cardinality.

Switch to async-cache for multi-DC

In federated deployments, set acl.down_policy = "async-cache". When a cached entry’s TTL expires, Consul serves the stale entry and refreshes asynchronously. This eliminates the thundering-herd refresh storm where every agent simultaneously issues a blocking RPC to the primary DC on TTL lapse.

acl {
  down_policy = "async-cache"
}

This is especially valuable when the WAN link is the bottleneck. Without async-cache, a degraded WAN link turns every TTL boundary into a coordinated latency spike.

Reduce unique token cardinality

The ACL caches are fixed in the source and not user-tunable via the documented agent configuration. If your unique active token population exceeds the cache capacity, no amount of TTL tuning fully eliminates thrashing. Options:

  • Reuse tokens across instances of the same service. Per-instance tokens multiply cardinality. A service-identity token shared across instances of the same workload reduces the working set.
  • Prefer service-identity tokens with broader policies over many narrowly-scoped tokens, when your security model allows it.
  • Audit and clean up leaked tokens. Automation that creates tokens without revoking them grows the population indefinitely and eventually inflates Raft commit latency.

Address replication lag in secondary DCs

If secondary-DC misses are slow because replication lags, the fix is on the replication path, not the cache. Check the replication token validity and permissions, WAN bandwidth, and primary DC ACL load. Replication lag above a few seconds in steady state is abnormal. In a degraded WAN scenario, async-cache keeps secondary-DC traffic moving while replication catches up.

Handle known token-leakage bugs

If you are running a Consul version affected by a token-leakage bug (where login-derived tokens are not properly revoked on agent shutdown), the token population grows without bound and eventually inflates Raft commit latency. As of review, issue #22613 remained open and proposed PR #23197 had not merged, so no fixed release can be named from those records alone. Check the current issue and release notes, then clean up accumulated stale tokens.

Prevention

  • Monitor the hit ratio, not just latency. Latency is a lagging indicator. A declining cache_hit to cache_miss ratio gives you lead time before resolution latency crosses thresholds.
  • Alert on consul.acl.ResolveToken p99 above 10ms sustained. Treat this as a TICKET. Escalate to PAGE if it coincides with elevated 403s or replication-lag alarms.
  • Track total token count over time. Steady growth without corresponding service growth indicates a leak; review issue #22613 and your release notes before attributing growth to that specific bug.
  • Set all three TTLs explicitly in config. Do not rely on the 30s defaults. Document the chosen values so future operators do not assume only policy_ttl matters.
  • In federated deployments, default to async-cache. It is the correct down_policy for any topology where the WAN link is not guaranteed to be fast and reliable.
  • Version-track the ACL subsystem. Consul 1.12.0 stopped reporting consul.acl.ResolveTokenToIdentity; its values moved into consul.acl.ResolveToken, so older dashboards may break silently. Consul 1.4.0 restructured ACL config into the nested acl {} stanza; pre-1.4 flat keys are deprecated.

How Netdata helps

  • Per-second resolution of consul.acl.ResolveToken exposes TTL-boundary miss storms that minute-granularity monitoring misses. The periodic spike pattern is the signature of a too-short TTL.
  • Correlate cache hit and miss counters with resolution latency on one timeline. A rising miss rate that precedes the latency spike confirms cache thrashing rather than a server-side bottleneck.
  • Cross-signal correlation with Raft commit time and goroutine count distinguishes ACL-bound latency from Raft-bound latency. If commitTime rises in lockstep with token count, token leakage is the likely driver.
  • Anomaly detection on the hit ratio flags declining cache effectiveness before latency crosses a static threshold.
  • Replication-lag monitoring in secondary DCs surfaces the multi-DC miss-cost amplifier before failover events turn it into an outage.
  • HTTP 403 rate alongside ACL metrics clarifies whether elevated latency is also causing fail-closed denials, which changes severity from TICKET to PAGE.