The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / consul / consul-acl-not-found ▌

Operations Guides

Consul "ACL not found": requests rejected after a token or policy change

“ACL not found” appears as HTTP 403 responses carrying “ACL not found” or “token does not exist: ACL not found”. It surfaces in agent and server logs, sidecar injector output, and as failed service registrations, health check updates, or KV writes. The blast radius depends on which token is missing: a single application token breaks one service; a replication or agent token can break an entire datacenter’s authorization pipeline.

The message is frequently misread. “ACL not found” does not mean the token has the wrong permissions. It means the Consul server receiving the request has no record of the token’s SecretID. The token was deleted, never created in this datacenter, or has not yet arrived via ACL replication. The fix path is completely different from “Permission denied”, which indicates the token is known but lacks a specific right.

The most common trigger is a credential rotation, a policy migration, or a failover in a federated setup where ACL replication is asynchronous and rate-limited.

What this means

Consul resolves every authenticated API request by looking up the presented token in the ACL subsystem. When the local server has no record of the SecretID, the request is rejected before policy evaluation runs. A token can be valid in the primary datacenter and still produce “ACL not found” in a secondary that has not yet replicated it.

The authoritative distinction:

  • “ACL not found” - the token does not exist in this server’s view of the ACL store. The SecretID is unknown.
  • “Permission denied” - the token exists, but its attached policies, roles, or service identities do not grant the requested operation.

Mixing these wastes time. Re-issuing policies or expanding token scope does nothing when the token was destroyed, never replicated, or replaced during a rotation and some client still holds the old SecretID.

The error string is identical whether the ACL system is healthy and the token is genuinely gone, or whether the ACL system was never bootstrapped. A cluster where ACLs were never initialized will reject every token-bearing request with “ACL not found” because there is no token store to consult.

flowchart TD
    A[Request with X-Consul-Token] --> B{Token in local ACL store?}
    B -- no --> C{ACL system bootstrapped?}
    C -- no --> D["ACL not found: ACL system not initialized"]
    C -- yes --> E{Replication lag?}
    E -- lagging --> F["ACL not found: token not yet replicated"]
    E -- up to date --> G["ACL not found: token deleted or never created"]
    B -- yes --> H{Policies grant operation?}
    H -- no --> I["Permission denied"]
    H -- yes --> J[Request allowed]

Common causes

CauseWhat it looks likeFirst thing to check
Token deleted or replaced during rotationA specific app or agent starts failing at the rotation timestamp; old SecretID in logsconsul acl token read with the accessor ID; compare against the new token
ACL replication lag in secondary DCErrors only in the secondary; primary is healthy; lag visible in /v1/acl/replicationReplicatedTokenIndex vs the primary’s latest token index
Replication token expired or lost permissionsSecondary DC replication status shows errors; lag grows unboundedReplication token validity and its ACL policy scope
ACL system not bootstrappedEvery token-bearing request fails, including the bootstrap attemptWhether consul acl bootstrap has been run in this DC
Agent token mismatch after reinstallOne agent fails to register services or push health checksAgent’s configured token against a valid token on the server
Unauthenticated flood with garbage tokensSustained 403 spike across many distinct unknown SecretIDs; no recent ACL changeSource IPs in access logs; whether the anonymous token is involved
Nomad workload identity raceErrors appear during allocation stop/startWhether deregistration runs before Consul has registered the token Confirmed: the deregistration ownership fix shipped in Nomad 1.9.1 (backported to Nomad Enterprise 1.8.6 and 1.7.14)

Quick checks

Run these read-only checks from a Consul server with a privileged token. None mutate state.

# Confirm a leader exists; ACL writes need a leader
curl -s http://127.0.0.1:8500/v1/status/leader

# Replication status in a secondary DC
# Key fields: Enabled, ReplicatedIndex, ReplicatedTokenIndex, LastSuccess, LastError
curl -s http://127.0.0.1:8500/v1/acl/replication | jq .

# Read a specific token by accessor ID
consul acl token read --accessor-id <accessor-id>

# Search the token list for a known SecretID (requires listing all tokens)
consul acl token list -format json | jq '.[] | select(.SecretID == "<secret-id>")'

# ACL resolution latency and token upsert counters
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep -E 'acl.resolveToken|acl.token'

# Count recent ACL-not-found vs permission-denied events
journalctl -u consul --since '1 hour ago' | grep -c 'ACL not found'
journalctl -u consul --since '1 hour ago' | grep -c 'Permission denied'

# Confirm the ACL system is bootstrapped (token count should be non-zero)
consul acl token list -format json | jq 'length'

If /v1/acl/replication returns Enabled: false on a secondary DC that should be replicating, replication is off. It must be explicitly enabled in the config and the replication token must be valid.

How to diagnose it

  1. Classify the failure scope. A single token failing points to rotation or deletion. Many distinct tokens failing points to replication or bootstrap. Every token failing points to an ACL subsystem that is down or not bootstrapped.

  2. Confirm the token exists in the primary DC. Run consul acl token read --accessor-id <id> from a primary server. If the token is gone, the root cause is upstream: find what deleted it and reissue. If the token is present, the issue is propagation or local resolution.

  3. Check replication status in every secondary DC. Compare ReplicatedTokenIndex against the primary’s latest token index. Recommended thresholds: under 1s healthy, 1-5s degraded, over 5s critical, over 30s page-worthy. Replication is rate-limited, so a token created in the primary may not be visible downstream for several seconds under load. LastError populated or LastSuccess stale means the stream is broken, not merely slow.

  4. Distinguish rotation aftermath from an attack. Pull access or audit logs and group the failing SecretIDs. A rotation aftermath shows a small set of recently-rotated tokens failing from known service identities. An unauthenticated flood shows many distinct unknown SecretIDs from a small set of source IPs, often hitting endpoints that should require auth.

  5. Check the replication token. In a secondary DC, ACL replication runs under a specific token. If that token was deleted, expired, or had its policy narrowed, replication silently stops advancing. Verify the token exists in the primary and still has sufficient ACL scope. The documented example policy grants acl = "read", not acl:write

  6. Correlate with leader changes. ACL replication depends on a healthy primary. A leader election in the primary during a token creation burst can stall replication. Cross-reference consul.raft.state.leader transitions against the spike in “ACL not found” errors.

  7. Check acl_down_policy. With the default extend-cache, secondaries can still resolve cached tokens when the authoritative source is unreachable. If someone changed this under pressure, behavior shifts immediately. Confirm the configured value before assuming replication is the only path. When the primary is unreachable, extend-cache allows cached tokens to be used regardless of TTL; in normal operation, tokens may be stale up to acl.token_ttl (default 30s)

Metrics and signals to monitor

SignalWhy it mattersWarning sign
consul.acl.resolveToken latencyACL resolution is on the hot path of every authenticated requestSustained mean above 10ms suggests cache thrashing or policy complexity
ACL replication lag (secondary DC, from /v1/acl/replication)Directly measures token and policy freshness downstreamLag above 5s, LastError populated, or Enabled: false
consul.acl.token.upsert rateToken creation burst can overwhelm replicationSpike more than 10x baseline without a change ticket
403 rate split by messageDistinguishes missing tokens from wrong permissionsSudden increase in “ACL not found” specifically
consul.raft.state.leader transitionsLeader churn in the primary stalls ACL replicationMore than 2 transitions per 10 minutes outside maintenance
WAN gossip member count per DCReplication rides WAN healthA remote DC’s servers missing from the WAN pool
consul.client.rpc.failed on secondary serversRPC failures break the replication streamSustained non-zero rate on servers that should be replicating

Fixes

Token deleted or replaced during rotation

Reissue the token and update every client that still holds the old SecretID. Find all the places the old token lives: environment variables, Kubernetes secrets, Vault agent templates, consul-template configs, sidecar injector annotations. Until every consumer is updated, requests with the stale SecretID will keep failing.

If you cannot update consumers immediately, you can re-create a token with the same SecretID via the HTTP API by specifying the SecretID field on creation. This is confirmed by the official API docs. This is a stopgap, not a strategy. Rotate properly afterward.

ACL replication lag in a secondary DC

If replication is enabled and the replication token is valid, lag usually resolves on its own as the rate limiter catches up. Do not restart servers to fix lag. Restarts reset caches and can make the problem worse during the warmup window.

If replication is stuck, verify in order: WAN connectivity on the gossip and RPC ports, the replication token exists in the primary with sufficient scope, the primary has a stable leader, and the primary is not saturated by a token creation burst.

Replication token expired or lost permissions

Rotate the replication token in the primary, then update the secondary’s configuration with the new token and reload. Until the new token is in place, the secondary cannot pull ACL changes. Plan this outside an incident if possible; doing it under fire means every new token created in the primary is invisible downstream until you finish.

ACL system not bootstrapped

Run consul acl bootstrap in the affected DC. This produces the initial management token. Until bootstrap completes, every token-bearing request returns “ACL not found” because there is no token store. This is most common immediately after enabling ACLs on an existing cluster or after a disaster recovery restore.

Unauthenticated flood

If the spike is external noise rather than a real rotation, the fix is network-level: restrict API access, rotate any exposed tokens, and verify the anonymous token has only the intended minimal permissions. Do not confuse this with a replication problem, or you will chase replication metrics while a scanning client is the actual source.

Nomad workload identity race

If you are running Consul 1.19.x or later with Nomad workload identities, deregistration may use a token Consul has not yet registered. The fix is on the Nomad side: upgrade to a version where deregistration is owned by the Nomad client rather than the workload. The fix shipped in Nomad 1.9.1 with backports to Nomad Enterprise 1.8.6 and 1.7.14

Prevention

  • Treat token rotation as a multi-DC event. Plan for the replication lag window. Stage rotations so the primary creates the new token, replication catches up, and only then do consumers switch.
  • Monitor ACL replication lag in every secondary DC. “Works in primary” is not a sufficient health check. Treat lag above 5s as a ticket and above 30s as page-worthy.
  • Alert on the 403 split. Track “ACL not found” and “Permission denied” as separate counters. A spike in one and not the other tells you immediately which failure mode you are in.
  • Lock down the replication token scope. Give it the minimum it needs. Anything narrower breaks replication silently; anything broader is unnecessary exposure.
  • Document acl_down_policy decisions. The default extend-cache buys time during primary unreachability. Changing it under pressure without understanding the tradeoff causes secondary outages.
  • Version-pin and test Consul and Nomad together. The workload identity race is a cross-product issue. Catch it in staging before it surfaces as “ACL not found” in production.

How Netdata helps

  • Per-second ACL resolution latency exposes cache thrash and policy-complexity spikes as they happen, instead of smoothed-over aggregates that hide the burst.
  • Replication lag charts per secondary DC alongside WAN gossip health and Raft leadership let you distinguish replication lag from connectivity loss without switching tools.
  • Split 403 counters on a single dashboard separate “ACL not found” from “Permission denied”, which is the single most useful distinction when triaging this error.
  • Raft leader transition annotations overlaid on ACL error spikes show whether a primary leadership change caused the replication stall.
  • Token upsert rate next to replication lag reveals whether a credential-rotation burst is the upstream cause of downstream authorization failures.