The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / consul / consul-intention-denied-connection ▌

Operations Guides

Consul intention denied: service-to-service traffic blocked by policy

Service A can no longer reach service B. Envoy sidecars log 403 responses or connection resets. The Consul server cluster looks healthy: leader stable, Raft committing, gossip intact, certificates within their lifetime. Traffic that worked an hour ago is now blocked.

This is almost always policy, not infrastructure. Consul Connect intentions are the mesh authorization layer, evaluated at connection establishment and enforced by Envoy RBAC filters. When an intention denies a connection, the data plane is doing exactly what it was configured to do. The diagnostic job is to determine whether the denial is correct (policy working as designed) or a misconfiguration: wrong identity, missing allow, precedence mistake, or stale cache.

The classic trigger is switching acl.default_policy from allow to deny without authoring allow intentions for every legitimate service pair first. Every missing intention becomes a block.

What this means

Intentions are authorization policies bound to service identities in the Connect mesh. When a source service opens a connection to a destination, the destination’s Envoy sidecar evaluates the intention set for that (source, destination) pair before the connection completes. L4 (TCP) intentions are evaluated once per connection through Envoy’s network RBAC filter. L7 (HTTP) intentions are evaluated per request through the HTTP RBAC filter.

The acl.default_policy setting controls what happens when no explicit intention matches:

  • allow: missing intention means traffic is permitted. Permissive, but no default-deny posture.
  • deny: missing intention means traffic is blocked. Zero-trust default, but requires explicit allow intentions for every legitimate pair.

Consul evaluates intentions by specificity: the most specific match wins, and at equal specificity deny takes precedence over allow. A specific deny for (web, billing) overrides a wildcard allow for (*, billing). Wildcards are supported in L4 source and destination names, but not in L7 permission rules.

Intentions are cached locally on the Consul agent and in the Envoy proxy. Updates propagate to proxies via xDS streaming, not instantaneously. If an agent or proxy is disconnected from servers, cached intentions continue to be enforced until connectivity is restored and a new configuration arrives. Envoy’s default initial_fetch_timeout is 15 seconds; this bounds waiting for the first configuration, not the latency of subsequent pushes.

The consul intention check <src> <dst> command evaluates the L4 decision for a pair and returns Allowed or Denied. It requires destination service:read; any intention carrying L7 Permissions is treated as deny by this L4-only check. consul intention match instead requires intentions:read and lists rules without evaluating them.

Common causes

CauseWhat it looks likeFirst thing to check
default_policy switched to denyMany unrelated service pairs fail simultaneously; no intention writes, only an ACL config changeAgent config acl.default_policy; recent config deploys
Missing allow intentionSpecific pair fails; consul intention check returns Deniedconsul intention check <src> <dst>
Wildcard vs specific precedenceA wildcard allow is shadowed by a more specific denyconsul intention match <dst> to list all matching intentions
Service identity mismatchDenials after a rename or redeploy; service registered under a different Name/v1/agent/services ServiceName vs intention source
CA or leaf certificate failuremTLS handshake failures rather than clean 403s; certs near expiry/v1/connect/ca/roots and Envoy /certs
ACL permissions blocking intention readsProxy cannot fetch intentions; xDS errors in server logsServer logs for ACL errors on intention endpoints
Stale proxy cacheRecent intention change not reflected in enforcementxDS stream health; sidecar uptime

Quick checks

Run these read-only checks before changing anything.

# Check ACL status on the agent
consul info | grep -i acl

# Evaluate the L4 decision for a specific pair
consul intention check <src-service> <dst-service>

# List all intentions matching a destination, ordered by precedence
consul intention match <dst-service>

# Verify CA roots are active and not expired
curl -s http://127.0.0.1:8500/v1/connect/ca/roots | jq '{ActiveRootID, Roots: [.roots[] | select(.Active == true) | {Name, NotAfter}]}'

# Inspect intention enforcement and authorization metrics
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep -iE "intention|connect.authorize"

# Confirm the service identity a proxy is actually presenting
curl -s http://127.0.0.1:8500/v1/agent/services | jq '.[].Service'

# Check xDS stream health for the proxy's server
curl -s http://127.0.0.1:8500/v1/agent/metrics | grep -i xds

consul info does not reliably expose acl.default_policy; the running agent’s configuration is exposed through /v1/agent/self in its version-dependent DebugConfig, and the configuration file or source of truth should be checked as well.

The consul intention CLI tree is deprecated since 1.9 in favor of service-intentions config entries written via consul config write. The check and match commands still function for diagnosis, but new policies should be authored as config entries.

Envoy access-log formats vary by deployment. For HTTP RBAC denials, Envoy records response_code_details as rbac_access_denied_matched_policy[policy_name]; its RBAC filter also exposes <stat_prefix>.rbac.denied and shadow_denied counters. For TCP RBAC, inspect the Envoy RBAC statistics and connection logs configured by your deployment.

How to diagnose it

Work through the decision tree before touching policy. The goal is to classify the denial as correct, missing, shadowed, or identity-broken.

flowchart TD
    A[Sidecar logs denied or reset] --> B{Server cluster healthy?}
    B -- No --> C[Fix Raft, gossip, or certs first]
    B -- Yes --> D{Many pairs failing at once?}
    D -- Yes --> E[Check acl.default_policy]
    E --> F{Recently switched to deny?}
    F -- Yes --> G[Author allow intentions before switching]
    F -- No --> H[Check CA root and leaf cert expiry]
    D -- No --> I[Run intention check for the pair]
    I --> J{Denied by check?}
    J -- Yes --> K[Review specificity and wildcards]
    J -- No --> L[Verify service Name identity]
  1. Confirm the server cluster is healthy. Check leader existence, Raft commit progress, and gossip membership. If the control plane is broken, intention delivery is broken, and denials may be stale-cache artifacts rather than current policy. Fix the cluster first.

  2. Scope the blast radius. If many unrelated service pairs fail at once, suspect a global change: default_policy, CA root expiry, or a broad ACL token revocation. If only specific pairs fail, narrow to intention configuration or identity for those pairs.

  3. Run the intention check for the failing pair. consul intention check <src> <dst> returns the L4 decision. If it returns Denied, the policy itself blocks the pair. If it returns Allowed but traffic still fails, the problem is identity, certificates, or cache, not the intention set.

  4. Inspect the matching intentions for the destination. consul intention match <dst> lists all intentions that apply, ordered by precedence. Look for a specific deny shadowing a wildcard allow. At equal specificity, deny wins.

  5. Verify the service identity. Intentions match on the service Name, not the service ID. Check /v1/agent/services on the source agent and confirm the ServiceName matches the intention source. A service redeployed under a different name will never match the old intention. Consul 1.9’s UI had a topology-view defect that could create intentions with the service ID instead of the name (issue #9304); PR #9316 fixed the UI behavior after 1.9.0. Verify the installed UI before relying on topology-driven authoring.

  6. Check certificates. If the symptom is mTLS handshake failures rather than clean RBAC denials, inspect CA roots and leaf expiry. Leaf certificates typically have short TTLs (72 hours by default). If rotation has stalled, Envoy holds expired certificates and connections fail. Consul 1.15.0 and 1.15.1 shipped a known leaf-certificate rotation race that could cause mesh communication loss after the 72-hour LeafCertTTL; it was fixed in 1.15.2.

  7. Check ACL permissions on the proxy token. The proxy’s token needs service:read on the destination service to evaluate intentions. If the token lacks permission, the proxy cannot read the intention set and may fail closed. Look for “ACL not found” or “permission denied” in server logs around intention endpoints.

  8. Rule out cache staleness. If you recently changed an intention and the proxy has not picked it up, check xDS stream health. If the stream is dropped or the proxy is disconnected from servers, it enforces the last-known-good configuration. Reconnecting or restarting the sidecar forces a refresh. Note that restarting the sidecar causes a brief connection drop for all traffic through that proxy.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Envoy rbac.denied counter by listener/stat prefixAuthorization denials observed in the data planeNew source/destination pattern or spike after a policy deploy
consul.intention.apply timerServer-side time to apply intention policy changesSpike after policy writes; confirms policy-change load
CA root and leaf certificate expirymTLS fails as a cliff at expiryLeaf under 24h, root under 7 days
xDS stream count and errorsIntention propagation path to proxiesStreams dropping, reconnect churn
consul.acl.ResolveToken latencyEvery intention read requires ACL resolutionp99 above 10ms, 403 spikes without policy changes
Service identity (Name field) consistencyIntention source must match registered NameMismatch after rename or redeploy

Envoy and Consul names vary by telemetry format (Prometheus, StatsD, DogStatsD). Confirm rendered names in your deployment; the underlying signals above are Envoy RBAC statistics, consul.intention.apply, and consul.acl.ResolveToken.

Fixes

Group the fix by the root cause you identified. Each has a tradeoff between speed and safety.

default_policy switched to deny

The fastest restore is reverting acl.default_policy to allow. This immediately unblocks all traffic but removes the default-deny posture; treat it as a rollback, not a fix. The safer path is to author allow intentions for every legitimate service pair first, validate each with consul intention check, then switch the default. This is a staged migration, not a flag flip.

Missing or wrong allow intention

Create the intention as a service-intentions config entry via consul config write (preferred post-1.9). For legacy deployments, consul intention create -allow <src> <dst> still works but is deprecated. Verify with consul intention check <src> <dst> after writing. A broad wildcard allow (* to dst) restores traffic fast but weakens least-privilege; prefer specific source matches.

Precedence mistakes (wildcard vs specific)

If a specific deny shadows a wildcard allow, you must either remove the specific deny or add a more specific allow. List the full match set with consul intention match <dst> to see precedence ordering. At equal specificity, deny wins. A deny for (web, billing) beats an allow for (*, billing) because the specific source is more specific than the wildcard.

Service identity mismatch

Update either the service registration (so the Name matches what the intention expects) or the intention source (so it matches the registered Name). The Name field is what matters, not the ID. After fixing, verify with consul intention check using the correct source name.

CA or certificate failure

If the CA root is expired or the signing backend (for example, Vault) is unreachable, no new leaf certificates can be issued. Existing leaves continue to work until they expire, then all mTLS fails. Check /v1/connect/ca/roots for active, non-expired roots. If rotation has stalled and Envoy holds expired leaves, restarting the downstream job or the Consul client forces a certificate refresh. Switching CA providers mid-incident is disruptive and should be a planned operation, not an urgent decision.

ACL permissions blocking intention reads

Grant the proxy’s ACL token service:read on the destination service. Without it, the proxy cannot evaluate intentions and may fail closed. Verify by checking server logs for ACL errors on intention endpoints after the token change propagates.

Stale proxy cache

If the xDS stream is healthy and the intention change is committed, the proxy should refresh on the next configuration push. If it does not, restarting the sidecar forces a full configuration fetch. This drops existing connections through that proxy, so apply it during a maintenance window if possible. Investigate why the stream was not delivering updates rather than relying on restarts as a fix.

Prevention

  • Stage default_policy changes. Author all allow intentions first, validate each pair with consul intention check, then switch. Never flip allow to deny cold.
  • Monitor denial rate as a baseline. Alert on new source/destination pairs, not just absolute count. A new pair after a deploy is the earliest signal of a missing intention.
  • Treat intention changes like firewall changes. Peer review, canary, and rollback plan. A single typo in a source name blocks traffic.
  • Track service identity consistency. Validate that the ServiceName in registrations matches what intentions expect, as part of deployment checks.
  • Monitor certificate expiry with long runway. Alert on leaf certificates at 24h and CA roots at 7 days. Monitor renewal success rate, not just expiry time.
  • Stay current on version behavior. The consul intention CLI is deprecated since 1.9; config entries are the forward path. Know which API surface your automation targets.

How Netdata helps

  • Per-second visibility into consul.connect.authorize exposes a denial spike the moment a policy change lands. Use this as the first check when paged on a service-to-service failure.
  • Correlate denial rate with xDS stream health and ACL resolution latency on a single timeline to localize whether the cause is policy, propagation, or permissions.
  • Certificate expiry dashboards show CA and leaf TTL with enough runway to fix rotation before it becomes mTLS failures.
  • Anomaly detection on denial patterns separates expected baseline denials from new pairs introduced by a deploy or config change.
  • Envoy sidecar connection metrics confirm whether denials are impacting application traffic or are noise at the mesh boundary.