The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / envoy / envoy-downstream-4xx-auth-spike ▌

Operations Guides

Envoy downstream 4xx spike: 401s, 403s, and 404s from the client side

A spike in http.<stat_prefix>.downstream_rq_4xx is usually a client-side story, not an Envoy story. The proxy is reporting that clients sent bad, unauthorized, or unroutable requests. The single counter lumps 400s, 401s, 403s, and 404s together, and Envoy does not expose per-status-code downstream counters. You will not find downstream_rq_401 or downstream_rq_403 in the stats dump.

The diagnostic path: take the aggregate signal seriously, but break it apart using access logs, the ext_authz and RBAC filter stats, and response flags. A 401/403 spike points at authentication infrastructure, a credential or certificate rotation, or a brute-force attempt. A 404 spike appearing right after an xDS push points at route misconfiguration. These have different owners and different runbooks.

This page covers how to decompose the spike, which Envoy signals separate auth failures from routing failures, and what to fix first.

What this means

downstream_rq_4xx is a monotonic counter on the HTTP connection manager. It increments for every 4xx response Envoy sends to a downstream client, regardless of who produced the response: the upstream app, an Envoy filter (ext_authz, RBAC, rate limit), or Envoy itself (no route, invalid header). It does not tell you which code or which component.

Two instrumentation gaps make this counter noisier than it looks:

  • No per-code downstream breakdown. Upstream stats have upstream_rq_401, upstream_rq_403, and so on, but downstream only has the class-level counter. To split 401s from 403s from 404s you need access logs or a log pipeline.
  • downstream_rq_4xx does not capture 431 (Request Header Fields Too Large). A header-too-large rejection increments downstream_cx_protocol_error but not the 4xx counter, so a client misconfiguration storm can be invisible here.

Because absolute 4xx volume is dominated by ordinary client behavior (legitimate 404s, scanner noise, expired browser tokens), alert on rate-of-change against a rolling baseline, not on an absolute count. A sudden 5x step on a normally flat counter is the signal that matters.

flowchart TD
    A["downstream_rq_4xx spike"] --> B["Group access logs by\nRESPONSE_CODE + RESPONSE_FLAGS"]
    B --> C{"Dominant code?"}
    C -->|"401 / 403"| D["Check ext_authz + rbac stats"]
    C -->|"404"| E["Check route config + xDS"]
    D --> F{"Flag UAEX or ext_authz.error?"}
    F -->|"UAEX, denied rising"| G["Auth policy regression"]
    F -->|"error rising, fail-closed"| H["Auth backend outage"]
    F -->|"failure_mode_allowed rising"| I["Fail-open active"]
    E --> J{"NR flag + recent push?"}
    J -->|"update_rejected or connected_state=0"| K["Stale or rejected config"]
    J -->|"VHDS convergence window"| L["Transient 404"]

Common causes

CauseWhat it looks likeFirst thing to check
Auth policy regression403 spike, response flag UAEX, ext_authz.denied or rbac.denied climbing in locksteprecent auth policy or RBAC deployment
ext_authz outage, fail-closed403 spike, ext_authz.error climbing, failure_mode_allowed flat at zeroauth service health, ext_authz.error rate
ext_authz outage, fail-openNo 4xx spike (traffic passes), ext_authz.failure_mode_allowed climbingfailure_mode_allowed counter, any nonzero is a security hole
Credential or cert rotation401 spike, possibly ssl.fail_verify_error climbing on the auth path/certs expiry, SDS connection, recent rotation
Brute force or credential stuffing401/403 spike concentrated on a few source IPs, small uniform bodiesaccess log source IP and path distribution
Route misconfiguration after xDS404 spike with NR flag, onset aligned with a config pushcontrol_plane.connected_state, update_rejected, config_dump
VHDS timing mismatchtransient 404s with NR during convergenceRDS vs VHDS update ordering, warming_clusters
RBAC header bypass (CVE-2026-26308)Denials drop while malicious requests succeed; version unpatchedEnvoy version against fixed releases

Quick checks

Run these read-only. None of them change Envoy state. Adjust the admin port for your deployment: 9901 for standalone Envoy, 15000 for Istio sidecars.

# 1. Confirm the aggregate 4xx rate and ratio (two samples 10s apart)
curl -s http://localhost:9901/stats/prometheus | grep 'downstream_rq_4xx\|downstream_rq_total'

# 2. Pull ext_authz filter stats
curl -s http://localhost:9901/stats | grep 'ext_authz'

# 3. Pull RBAC filter stats
curl -s http://localhost:9901/stats | grep 'rbac'

# 4. Check whether Envoy is rejecting config (silent NACKs)
curl -s http://localhost:9901/stats | grep -E 'update_rejected|listener_create_failure|connected_state'

# 5. Check cert runway
curl -s http://localhost:9901/certs | jq '.certificates[] | {subject: .cert_chain[].subject, days: .days_until_expiration}'

# 6. Confirm 5xx is flat (this is a client-side story, not an upstream outage)
curl -s http://localhost:9901/stats | grep 'downstream_rq_5xx'

# 7. Inspect the current route tables to validate 404s are routing, not missing
curl -s http://localhost:9901/config_dump | jq '.configs[] | select(."@type" | test("RouteConfiguration"))'

How to diagnose it

  1. Decompose the aggregate. Grep access logs for the spike window and group by %RESPONSE_CODE% and %RESPONSE_FLAGS%. The flags are access-log only and are not exposed as aggregate stats, so this step is mandatory. Without it, you are guessing at which code dominates.

  2. Split 401/403 from 404. They have different root causes and different owners. 401/403 is an auth story. 404 with NR is a routing story. Chasing both at once wastes the first 15 minutes of an incident.

  3. For 401/403, identify the denying component. If the response flag is UAEX, the ext_authz filter produced the denial. Cross-check against the filter stats:

    • ext_authz.denied rising with ext_authz.ok flat means the auth service is actively denying more requests. Look for a policy change.
    • ext_authz.error rising with ext_authz.denied roughly flat means the auth service is unreachable or erroring. Envoy is fabricating the 403 via status_on_error (default 403). This is an outage of the auth backend, not a policy problem.
    • ext_authz.failure_mode_allowed rising means Envoy is fail-open. There is no 4xx spike because traffic is passing unauthenticated. Treat any nonzero value here as a security incident.
    • rbac.denied rising without UAEX means the in-proxy RBAC filter is the source, typically after a policy deployment.
  4. For 404, correlate with config timing. A 404 spike with the NR flag means Envoy has no route for the Host and path. The two questions are whether a config push just happened and whether Envoy actually accepted it:

    • If control_plane.connected_state is 0, Envoy is on stale config and may be missing routes that the control plane thinks it pushed.
    • If update_rejected is climbing, Envoy is connected but NACKing the new config. The control plane reports a successful deploy; Envoy silently kept the old routes. This is the classic “deployment completed but nothing changed” failure.
    • If you use VHDS, transient 404s can appear when RDS updates land before the corresponding virtual host updates. The window is short but real during convergence.
  5. Cross-check upstream. Compare downstream_rq_4xx against the cluster’s upstream_rq_4xx. If the upstream counter is climbing too, the app is genuinely producing these codes and Envoy is just forwarding them. If only the downstream counter moves, the response originated in Envoy or in a filter.

  6. For abuse patterns, read the access log distribution, not the counters. A brute-force or credential-stuffing spike is concentrated on a handful of source IPs and a small set of paths, with uniform small bodies. Aggregate counters cannot reveal this shape.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
http.<stat_prefix>.downstream_rq_4xxaggregate client-side error loadrate-of-change spike against rolling baseline
ext_authz.denied / ext_authz.okratio of auth denials to allowsdenied ratio jumping after a policy change
ext_authz.errorauth service unreachable or erroringany sustained nonzero rate
ext_authz.failure_mode_allowedfail-open activeany nonzero value is a security hole
rbac.denied / rbac.shadow_deniedin-proxy policy denials and dry-rundenied spike after policy deploy; shadow_denied catching legit traffic before enforcement
response flag UAEX (access log)ext_authz produced the denialsustained nonzero rate
response flag NR (access log)no route matchedany nonzero rate in production is a config error
control_plane.connected_staterunning on stale config0 sustained
update_rejected / listener_create_failureEnvoy NACKing configany nonzero
ssl.fail_verify_errorcertificate verification failuresspike correlates with 401 bursts
downstream_rq_5xxconfirms the issue is client-sideshould stay flat during a 4xx spike

Fixes

Auth policy regression (403 with UAEX)

Roll back the policy change first, then debug the rules. If you use shadow RBAC, check rbac.shadow_denied before enforcing: a policy that denies legitimate traffic shows up there without breaking users. For ext_authz, confirm the auth service is returning the denial you expect by hitting it directly with a representative request. Increasing retry counts or timeouts here does not help; the service is answering, just with deny.

ext_authz outage, fail-closed (403 with ext_authz.error)

Restore the auth service. Envoy is generating the 403s locally via status_on_error because the backend is unreachable or returning errors. Note the stats gap: when the auth service returns 5xx, Envoy treats it as an error and the response code becomes status_on_error regardless of what the auth body said. You cannot distinguish “auth server denied” from “auth server 500” using the 4xx counter alone. While restoring service, do not flip to failure_mode_allow: true as a quick fix unless you intend to run unauthenticated.

ext_authz outage, fail-open (no 4xx spike, failure_mode_allowed climbing)

This is the dangerous variant. Traffic is flowing, the 4xx counter looks fine, and unauthenticated requests are passing. The fix is to restore the auth service, not to celebrate the flat error rate. Any nonzero failure_mode_allowed during an unplanned window is a security incident.

Credential or certificate rotation (401 spike)

Check /certs for expiry and confirm SDS is connected. If ssl.fail_verify_error is climbing on the path to the auth service, the rotation did not propagate. Re-trigger rotation or roll back to the previous credential set. Coordinate with the auth service owner: a 401 spike often means the verifier and the credential issuer disagree on the new key.

Route misconfiguration (404 with NR)

The fix is the config, not Envoy. First confirm whether Envoy accepted the push:

  • update_rejected nonzero: fix the rejected config on the control plane side. Envoy is protecting itself by keeping the old routes.
  • connected_state = 0: restore control plane connectivity. Envoy picks up the routes on reconnect.
  • VHDS ordering: ensure virtual host updates arrive after the RDS updates that reference them.

Roll back the route change if the new config is wrong. Do not paper over NR with catch-all routes; that hides the misconfiguration and breaks routing observability.

Brute force or abuse (401/403 concentrated on few IPs)

This is a rate-limiting or WAF problem, not an Envoy config bug. If you run local or global rate limiting in Envoy, watch ratelimit.over_limit. Otherwise, handle it at the edge. Adding the offending IPs to a deny list via RBAC is a valid short-term mitigation; verify the RBAC rule matches headers correctly given the duplicate-header concatenation behavior described below.

RBAC header bypass (CVE-2026-26308)

If your Envoy is older than the fixed releases (1.37.1, 1.36.5, 1.35.8, 1.34.13) and you rely on RBAC exact-match header rules, treat unexplained denial-rate drops as a possible bypass. Upgrade. The bug is that Envoy concatenates duplicate header values into a comma-separated string before matching, so a request carrying x-role: admin,user can evade a rule keyed on x-role: admin.

Prevention

  • Alert on rate-of-change, not absolute count. A flat 4xx baseline with normal scanner noise will trip any reasonable absolute threshold. Alert on deviation from a rolling baseline for the specific stat_prefix.
  • Build a 4xx-by-code signal from access logs. Because downstream has no per-code counter, ship access logs to a pipeline that counts by %RESPONSE_CODE% and %RESPONSE_FLAGS%. This is the only way to get 401/403/404 breakdowns in real time.
  • Run RBAC in shadow mode before enforcing. Watch rbac.shadow_denied for a full traffic cycle. If it catches legitimate traffic, fix the policy before it becomes denials.
  • Monitor failure_mode_allowed as a security signal. Any unplanned nonzero value is a page, not a ticket.
  • Track update_rejected alongside connected_state. A connected Envoy that NACKs every push is as stale as a disconnected one, and quieter.
  • Keep Envoy current on RBAC CVEs. Header-matching bypasses are silent; the only reliable signal is version hygiene.

How Netdata helps

  • Per-second downstream_rq_4xx and downstream_rq_total expose the step change the moment it starts, and ML anomaly detection flags the rate-of-change deviation without hand-tuned absolute thresholds.
  • Correlating ext_authz.denied, ext_authz.error, and ext_authz.failure_mode_allowed against the 4xx spike separates “policy is denying more” from “auth backend is down” from “fail-open is active” within a single chart view.
  • RBAC filter stats (rbac.denied, rbac.shadow_denied) sit next to the 4xx counter, so a policy rollout that spikes denials reads cleanly against the deploy timeline.
  • control_plane.connected_state and update_rejected plotted against a 404 spike make the stale-config or NACK diagnosis immediate, especially when aligned with the config-push window.
  • ssl.fail_verify_error alongside 401 bursts points a credential or cert rotation problem at the auth path rather than at the application.