The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / envoy / envoy-nr-no-route ▌

Operations Guides

Envoy NR no route: 404s and 503s after a bad xDS route push

The NR response flag in Envoy access logs means no route matched the request. The HTTP connection manager walked its route table (delivered via RDS or static config) and found no match for the incoming Host header, path, or other match criteria. In production, sustained NR is a configuration error. A spike right after an xDS route push points to a bad or rejected config update.

NR is usually paired with HTTP 404. Operators sometimes see it paired with 503, which is a different root cause: on Envoy versions before 1.18.0, NR was also set when a route matched but the referenced upstream cluster did not exist. Envoy 1.18.0 (April 2021) introduced the NC (No Cluster) flag for that case, so on modern Envoy, NR with 503 suggests you are either on an older version or actually looking at NC.

The response is an Envoy-generated local reply. No upstream connection is attempted, so cluster health, connection pools, and circuit breakers are irrelevant for NR-flagged requests. The problem is entirely in the route configuration layer.

What this means

NR fires at the router filter level, before cluster selection. The HTTP connection manager receives the request, walks the configured route table, and finds no match. Envoy generates a local response and writes the access log entry with the NR flag.

Three failure modes produce this:

  1. Route configuration not delivered: RDS failed to push the route table. The HCM has active listeners but no route configuration attached. Every request to that listener gets NR. Check control_plane.connected_state and whether the route configuration appears in config_dump.

  2. Route exists but does not match the request: The route table is present, but the virtual host domains or route matchers do not align with the incoming request’s Host header, path prefix, or headers. The most common variant is a host:port mismatch where the route matches example.com but the client sends Host: example.com:443.

  3. Route matches but cluster is missing (pre-1.18.0 behavior): The route entry exists and matches, but the cluster it references is not in the cluster manager. On Envoy before 1.18.0, this produces NR with a 503 status code. On Envoy 1.18.0 and later, this produces NC instead. If you see NR with 503 on modern Envoy, verify you are not confusing this with NC.

The response_code_details access log field provides finer-grained information. The key values are route_not_found (no route matched), route_configuration_not_found (no RDS config at all), and cluster_not_found (route matched but cluster absent).

Verify your access log format includes %RESPONSE_CODE_DETAILS%.

flowchart TD
    A["Access log: NR flag"] --> B{"Route config present
in config_dump?"} B -->|"No"| C["RDS not delivered
check connected_state
check update_rejected"] B -->|"Yes"| D{"Route domains
match Host header?"} D -->|"No"| E["Host:port or domain
mismatch in route config"] D -->|"Yes"| F{"Envoy >= 1.18.0
and status is 503?"} F -->|"Yes"| G["Check for NC flag
instead of NR
cluster may be missing"] F -->|"No"| H["Route matcher
too narrow
path or header criteria"]

Common causes

CauseWhat it looks likeFirst thing to check
RDS route config rejected (NACK)NR spike correlates with config push time; update_rejected incrementing; control_plane.connected_state still 1config_dump for the affected listener’s route configuration
Host header or domain mismatchNR only for requests from certain clients or paths; access log shows Host value not in route domainsCompare request Host header to virtual host domains in route config
Route config not yet deliveredNR on all requests to a listener after startup or reconnect; listeners or clusters stuck in warmingcontrol_plane.connected_state and listener_manager.total_listeners_warming
Cluster removed mid-request (NC, not NR)Status 503 but flag is NC on Envoy 1.18.0+; incremental xDS eventual consistency issues/clusters?format=json for the referenced cluster name
Route path matcher too narrowNR only for specific paths or methods; other routes work fineRoute match config in config_dump

Quick checks

# Check if control plane is connected
curl -s http://localhost:9901/stats | grep 'control_plane.connected_state'

# Check for rejected config updates (NACKs)
curl -s http://localhost:9901/stats | grep -E '(update_rejected|update_failure|listener_create_failure)'

# Check for clusters or listeners stuck in warming
curl -s http://localhost:9901/stats | grep -E '(warming_clusters|total_listeners_warming)'

# Dump the active route configuration
curl -s http://localhost:9901/config_dump | jq '.configs[] | select(.["@type"]=="type.googleapis.com/envoy.admin.v3.RoutesConfigDump")'

# Verify a specific cluster exists in the active config
curl -s http://localhost:9901/config_dump | jq '.configs[] | select(.["@type"]=="type.googleapis.com/envoy.admin.v3.ClustersConfigDump") | .dynamic_active_clusters[]?.cluster.name'

# List active cluster names and health status
curl -s http://localhost:9901/clusters?format=json | jq '.cluster_statuses[].name'

# Confirm Envoy version (determines NR vs NC split behavior)
curl -s http://localhost:9901/server_info | jq '.version'

All commands are read-only and safe during incidents. In Istio sidecar mode, replace port 9901 with 15000.

How to diagnose it

  1. Confirm NR in access logs. Verify your access log format includes %RESPONSE_FLAGS%. In Istio, this is in the default format but may be stripped in custom configurations. Look for lines where the flag field contains NR.

  2. Check the HTTP status code paired with NR. If it is 404, the route table has no match for the request. If it is 503, verify whether you are actually seeing NC instead of NR (Envoy 1.18.0 and later), or whether you are on an older Envoy version where NR covered both cases.

  3. Check response_code_details if your access log format includes it. route_configuration_not_found means RDS did not deliver the route config at all. route_not_found means the route table exists but no route matched. cluster_not_found means the route matched but the cluster is absent.

  4. Correlate with the config push timeline. Compare the timestamp of the NR spike with update_success and update_rejected counters. If update_rejected incremented at the same time, the control plane pushed invalid config and Envoy NACKed it. Envoy keeps the old config, which may be missing routes for newly deployed services. The operator sees “deployment completed” on the control plane side while Envoy silently ignored it.

  5. Dump the route configuration. Use the /config_dump admin endpoint to verify the route table is present and matches what the control plane intended to push. Check that the virtual host domains include the Host header values clients are sending.

  6. Check for warming or stuck resources. Non-zero cluster_manager.warming_clusters or listener_manager.total_listeners_warming during steady state means Envoy received config but cannot activate it. A cluster stuck in warming does not receive traffic, and routes referencing it may fail.

  7. Verify the cluster exists. If the route matched but produced NR with 503 (or NC), check /clusters?format=json for the cluster name the route references. The cluster may have been removed by EDS or never created by CDS.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
%RESPONSE_FLAGS% containing NR in access logsNR means no route matched. Any sustained NR rate is a configuration error.Spike immediately after an xDS push
%RESPONSE_CODE_DETAILS% in access logsDistinguishes route_not_found from route_configuration_not_found and cluster_not_foundUnexpected value for matched requests
control_plane.connected_stateIf 0, Envoy is running stale config. New routes never arrive.Disconnection sustained more than 5 minutes
update_rejected (per resource type)Non-zero means Envoy NACKed a config push. Intended routes may not be active.Incrementing after a deployment
listener_manager.listener_create_failureListener config rejected by EnvoyAny non-zero value
cluster_manager.warming_clustersClusters received but not activated. Routes to warming clusters fail.Non-zero sustained during steady state
cluster.<name>.update_emptyControl plane sent an update with no endpointsSustained for a cluster that should have endpoints

Response flags are access-log only. Envoy does not expose aggregate NR or NC counters via the stats endpoint or Prometheus. Monitoring them requires access log processing, either through a log pipeline with counters or custom stats via a Lua or Wasm filter.

Fixes

Route config NACKed (update_rejected incrementing)

If Envoy rejected the new route config, the old config remains active. The fix is on the control plane side. Identify why the config was rejected by checking Envoy’s stderr logs for the NACK reason, fix the invalid configuration, and re-push.

Do not restart Envoy as a first fix. A restart triggers a fresh RDS request, and if the control plane sends the same invalid config, Envoy will NACK it again and continue running with stale routes.

If the NR impact is severe and the control plane fix will take time, roll back the control plane to the last-known-good configuration version. This restores the previous route tables.

Host header or domain mismatch

If response_code_details shows route_not_found and the route table exists in config_dump, the issue is in domain matching. Compare the Host header in the access log for NR-flagged requests against the domains list in the virtual host configuration.

Common fixes:

  • Add port variants to the virtual host domains list (for example, both example.com and example.com:443).
  • Configure strip_any_host_port or strip_matching_host_port on the HTTP connection manager to normalize the Host header before route matching.
  • Verify that the path prefix or regex matcher in the route entry covers the request paths that are failing.

Route config not delivered (RDS failure)

If config_dump shows no route configuration for the affected listener, the RDS subscription failed. Check that control_plane.connected_state is 1, that the route configuration resource name in the listener matches what the control plane expects to push, and Envoy’s logs for RDS subscription errors or timeouts.

If the control plane is connected but not delivering the route config, the issue is in the control plane’s resource selection logic. Restarting Envoy will trigger a fresh RDS request but does not fix a control plane that is not responding to it.

Cluster referenced by route is missing

If the route matched but the cluster does not exist, you will see NR with 503 on Envoy before 1.18.0, or NC on Envoy 1.18.0 and later. Check /clusters?format=json for the cluster name. If the cluster was removed via CDS or never created, ensure the control plane pushes the cluster that the route references. If the cluster was removed intentionally, update the route to reference the correct cluster or remove the route entirely.

For incremental xDS (Delta xDS) deployments, clusters can disappear transiently due to resource-tracking eventual consistency during incremental updates. Sustained NR or NC after the config should have converged suggests a subscription or resource-tracking problem in the control plane.

Prevention

Canary config validation. Before a full fleet rollout of route changes, validate the new config statically with envoy --mode validate -c /path/to/config.yaml and push to a small subset of instances first. Watch their access logs for NR or NC before proceeding to the full fleet. Exit code 0 means the config is valid; non-zero means it is not.

Alert on NR and NC in access logs. Since these are not aggregate stats, instrument your log pipeline to count response flags. Any non-zero NR or NC rate should page. These flags indicate configuration errors, not transient failures.

Monitor update_rejected and listener_create_failure. These counters are zero in healthy operation. Non-zero values mean Envoy silently rejected config from the control plane. Alert on any increment.

Include response_code_details in access logs. The default access log format may not include %RESPONSE_CODE_DETAILS%. Adding it gives you immediate root-cause granularity for NR events without requiring a config_dump at incident time.

Review route config changes. Require that route additions and modifications include the full set of domains, path matchers, and cluster references. A mismatched domain or a typo in a cluster name produces NR or NC that is invisible until clients hit the broken route.

How Netdata helps

  • Per-second metric resolution lets you pinpoint the exact second update_rejected or listener_create_failure increments, which you can align with the access log timestamp for the first NR-flagged request.
  • Anomaly detection on control_plane.connected_state, warming_clusters, and update counters surfaces silent NACKs that would otherwise go unnoticed until clients report failures.
  • Correlation across signals lets you overlay xDS update health with downstream 5xx rates, confirming whether a config push caused the error spike without jumping between tools.
  • Access log processing through log-based collectors can surface response flag distributions as metrics, giving you NR and NC visibility that Envoy’s native stats endpoint does not provide.
  • Composite dashboards showing Envoy, the control plane, and upstream services together help distinguish route-not-found from upstream-unhealthy within the same incident view.