The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / envoy / envoy-ext-authz-failure-mode-allowed ▌

Operations Guides

Envoy ext_authz failure_mode_allowed: unauthenticated traffic when auth is down

The Envoy ext_authz filter calls an external authorization service on every request that matches the filter chain. That call is on the request critical path: nothing is forwarded upstream until the auth service returns a decision. When the auth service is unreachable, returns an HTTP 5xx, or exceeds its timeout, Envoy either lets the request through without an auth decision or rejects it. That choice is controlled by failure_mode_allow, and every time the fail-open branch fires, Envoy increments http.<stat_prefix>.ext_authz.failure_mode_allowed.

A non-zero rate on failure_mode_allowed means requests are being forwarded upstream without an authorization decision. In any deployment where ext_authz is the security perimeter, that is a live security incident, not a performance problem. The opposite setting (the default) is fail-closed, which rejects every request with status_on_error while the auth service is unavailable. Both modes have valid use cases, but they fail in opposite directions, and the operator response to each is completely different. Know which mode every protected route is running before the auth service has an outage, not during one.

What failure_mode_allow actually does

failure_mode_allow is a boolean field on the v3 ExtAuthz filter config. The default is false (fail-closed).

  • failure_mode_allow: false (default, fail-closed). When the auth service fails, Envoy rejects the request using the status code configured in status_on_error. The default for status_on_error is 403 Forbidden. The body the auth service might have returned is dropped; the client sees the status_on_error code with an empty body. No traffic reaches the upstream until the auth service is healthy again.
  • failure_mode_allow: true (fail-open). When the auth service fails, Envoy forwards the request upstream as though it had been allowed. The counter http.<stat_prefix>.ext_authz.failure_mode_allowed increments. The upstream has no way of knowing the request was not authorized unless you also propagate the x-envoy-auth-failure-mode-allowed header.

The word “fails” needs to be precise. It does not mean the same thing as an explicit denial.

What counts as an ext_authz error

failure_mode_allow only triggers on errors, not on explicit denials. This distinction is the source of most operator confusion around the filter.

  • A 200 OK from the auth service means “allow”. Envoy forwards the request. Neither error nor failure_mode_allowed increments.
  • A 403 Forbidden (or whatever denied status the auth service returns) means “deny”. Envoy returns the denial to the client and increments ext_authz.denied. The response flag UAEX is set in access logs. This is not an error. failure_mode_allow is irrelevant here.
  • An HTTP 5xx from the auth service means the auth service could not make a decision. Envoy treats this as an error. Fail-closed returns status_on_error. Fail-open forwards the request and increments failure_mode_allowed.
  • A network failure (TCP reset, connect timeout, stream idle timeout, deadline exceeded) means Envoy could not reach the auth service. Same handling as a 5xx.
Auth service outcomefailure_mode_allow: false (default)failure_mode_allow: true
200 OK (allow)Forward upstreamForward upstream
4xx denial (e.g., 403)Return denial to client, flag UAEX, ext_authz.denied++Return denial to client, flag UAEX, ext_authz.denied++
5xx from auth serviceReturn status_on_error (default 403), ext_authz.error++Forward upstream, ext_authz.error++ and ext_authz.failure_mode_allowed++
Network failure to auth serviceReturn status_on_error, ext_authz.error++Forward upstream, ext_authz.error++ and ext_authz.failure_mode_allowed++

Note the asymmetry in the fail-closed column: a real denial from the auth service and an infrastructure error both surface to the client as a 403 (by default). From the client side they are indistinguishable. From the operator side they are completely different signals. Use ext_authz.denied versus ext_authz.error to separate them, and use the UAEX access log flag to confirm denials.

flowchart TD
    A[Client request] --> B[ext_authz filter]
    B --> C[Call auth service]
    C --> D{Response}
    D -->|200 OK allow| F[Forward upstream
ext_authz.ok++] D -->|4xx denial| E[Return denial
flag UAEX
ext_authz.denied++] D -->|5xx or network error| G{failure_mode_allow?} G -->|false default| H[Return status_on_error
default 403
ext_authz.error++] G -->|true| I[Forward upstream
ext_authz.error++
ext_authz.failure_mode_allowed++]

The signals

Every ext_authz-protected HTTP connection manager emits a small, fixed set of counters under http.<stat_prefix>.ext_authz. The stat_prefix comes from the HTTP connection manager config, not the cluster name.

  • http.<stat_prefix>.ext_authz.ok. Successful allow decisions.
  • http.<stat_prefix>.ext_authz.denied. Explicit denials returned by the auth service.
  • http.<stat_prefix>.ext_authz.error. Auth service failures, including both network errors and 5xx responses. In a fail-closed deployment, a non-zero rate here means traffic is being rejected by status_on_error. In a fail-open deployment, every error is also a failure_mode_allowed increment.
  • http.<stat_prefix>.ext_authz.failure_mode_allowed. The fail-open counter. Counts requests let through without an auth decision because the auth service failed and failure_mode_allow: true. This is the security-hole signal. It should be flat zero in steady state. Any increment is an event.
  • http.<stat_prefix>.ext_authz.latency. The auth service call latency. This is added directly to the request critical path for every authorized request, and it shows up in downstream_rq_time dollar-for-dollar. A slow auth service is a slow proxy.

The response flag UAEX in the access log marks ext_authz-denied requests. It does not mark fail-open traffic, because fail-open traffic was not denied. If you want fail-open visibility in logs, surface the x-envoy-auth-failure-mode-allowed header upstream and log it there, or build a log pipeline that joins the failure_mode_allowed counter against access log volume.

Where fail-open shows up in production

The auth service does not have to be “down” in the binary sense for fail-open to fire. Any condition that causes Envoy to treat the auth call as an error will trip it:

  • Auth service deploy or rolling restart. Brief burst of errors as endpoints churn. Generally self-resolves, but in fail-open mode every request during the window is unauthorized.
  • Auth service CPU saturation or GC pause. Tail latency spikes first, then per-try timeouts start firing, then errors.
  • Auth service OOM or crash. Sustained error burst, sustained fail-open window.
  • mTLS certificate rotation between Envoy and the auth service. TLS handshake failures count as ext_authz errors. If SDS is slow to push a new cert, expect a window of fail-open traffic.
  • Connection pool exhaustion to the auth cluster. The auth service is up but Envoy cannot get a connection. Same fail-open behavior as a hard outage.
  • Network partition to the auth service zone. Connect timeouts, then errors, then fail-open.
  • Downstream-client-driven bypass. CVE-2024-23324 describes a case where downstream clients can craft requests that cause Envoy to send invalid gRPC check requests to ext_authz, circumventing the auth check when failure_mode_allow: true. Affected versions include 1.26.x prior to 1.26.7, 1.27.0 through 1.27.2, 1.28.0, and 1.29.0; fixed in 1.26.7, 1.27.3, 1.28.1, and 1.29.1. The upstream advisory lists no known workarounds beyond upgrading. If you run fail-open on an affected version, the fail-open posture is an exploitable bypass, not just an availability tradeoff. Treat fail-open as a setting that needs both monitoring and a current patch level.

The header signal

When failure_mode_allow: true and the fail-open path fires, Envoy adds the header x-envoy-auth-failure-mode-allowed: true to the request headers forwarded upstream — but only when header addition is enabled. In Envoy 1.26.0 through 1.27.x this was automatic (controlled by the runtime feature flag envoy.reloadable_features.http_ext_auth_failure_mode_allow_header_add, default enabled). From 1.28.0 onward, the explicit config field failure_mode_allow_header_add: true must also be set; its default is false. Upstreams can use this header to apply degraded-mode behavior: read-only mode, aggressive rate limiting, feature gating, audit logging at a higher tier.

The header is the only in-band signal the upstream gets that the request was not authorized. Without it, fail-open traffic is indistinguishable from authorized traffic at the application layer.

Quick checks

All read-only. The admin port is 9901 by default and 15000 in Istio sidecar mode. Adjust accordingly.

# Inspect the four core ext_authz counters for a stat prefix
curl -s http://localhost:9901/stats | grep 'ext_authz'

# Confirm which failure_mode_allow is configured on each filter
curl -s http://localhost:9901/config_dump | \
  grep -E 'failure_mode_allow|stat_prefix|status_on_error'

# Spot fail-open traffic in real time (counters are monotonic)
watch -n 1 'curl -s http://localhost:9901/stats | grep ext_authz.failure_mode_allowed'

# Check the auth cluster's health, since auth cluster failure drives fail-open
curl -s http://localhost:9901/stats | grep -E \
  'cluster.<auth_cluster>.(upstream_cx_connect_fail|membership_healthy|upstream_rq_5xx)'

# Verify latency contribution from the auth call
curl -s http://localhost:9901/stats/prometheus | grep 'ext_authz'

If you are running fail-closed and triaging an outage, the same ext_authz.error counter is the one to watch. It tells you Envoy is actively rejecting traffic because the auth service is unavailable.

Fail-open versus fail-closed: the policy decision

Neither mode is universally correct. The choice is a policy decision with security and availability implications that should be made explicitly per route, not inherited from a tutorial config.

Fail-open (failure_mode_allow: true) is appropriate when:

  • The protected service has layered controls. Authz is one of several gates, not the only one.
  • The upstream can tolerate unauthenticated traffic for short windows without data integrity risk.
  • Total outage of the protected service is more costly than partial exposure.
  • You have aggressive monitoring on failure_mode_allowed and an audited runbook for any non-zero rate.

Fail-open is dangerous when:

  • ext_authz is the only perimeter. There is no second gate.
  • The upstream assumes all traffic is authenticated and exposes data or mutations based on that assumption.
  • Sensitive data, financial operations, or compliance-regulated workloads are involved.
  • Nobody is watching failure_mode_allowed and it has been incrementing for hours.

Fail-closed (failure_mode_allow: false, default) is appropriate when:

  • Unauthenticated traffic is unacceptable under any condition.
  • The service has strict security or compliance requirements.
  • A full outage of the protected service is preferable to an unauthorized request succeeding.

Fail-closed is dangerous when:

  • The auth service is a hard dependency with no SLO headroom. An auth service hiccup takes down everything behind Envoy.
  • The blast radius of total outage is larger than the blast radius of partial exposure.
  • You are not monitoring ext_authz.error. Fail-closed outages look like a generic 5xx spike on the protected service if you do not separate the cause.

In both modes, the auth service becomes the most critical dependency in the request path. Treat its SLO, capacity, and monitoring with the same rigor as the data plane.

Signals to watch in production

SignalWhy it mattersWarning sign
http.<stat_prefix>.ext_authz.failure_mode_allowed rateCounts requests forwarded without an auth decision in fail-open modeAny non-zero value. Flat zero is the only healthy state.
http.<stat_prefix>.ext_authz.error rateAuth service is failing (network or 5xx). In fail-closed mode, this is the counter that says traffic is being rejected.Sustained non-zero
http.<stat_prefix>.ext_authz.denied rateAuth service is explicitly denying requests. Sudden spike may indicate policy change, credential rotation, or attack.Sudden change from baseline
http.<stat_prefix>.ext_authz.latency (P99)Auth call latency is added to every authorized request’s critical path. A slow auth service is a slow proxy.Upward trend, especially P99 drifting while P50 is stable
cluster.<auth_cluster>.upstream_cx_connect_failAuth hosts not accepting connections. Often the leading indicator before errors start.Non-zero sustained
cluster.<auth_cluster>.membership_healthy ratioAuth cluster host health. A drop here typically precedes an ext_authz error burst.Ratio below baseline
Response flag UAEX in access logsConfirms explicit ext_authz denials (not errors, not fail-open)Spike correlated with policy or credential changes

For fail-open deployments, the alerting rule is straightforward: page on any sustained non-zero rate of failure_mode_allowed. For fail-closed deployments, the equivalent rule pages on sustained non-zero ext_authz.error combined with rising downstream_rq_503 (or whatever status_on_error is configured to).

Prevention

  • Make the mode explicit in config review. failure_mode_allow is a single boolean with outsized impact. Every protected route should have a deliberate choice, documented in the route’s runbook.
  • Monitor the counter, not just the auth service. A healthy auth service with a misconfigured filter can still produce fail-open traffic if the filter is wired wrong. The counter is ground truth.
  • Track auth service latency as a first-class SLO. Because ext_authz is on the critical path, auth service P99 directly determines protected service P99. Latency budget overruns on auth should be treated as capacity incidents.
  • Capacity-plan the auth cluster like a data-plane tier. The auth service is not control-plane infrastructure that can be slow. It is on every request’s hot path.
  • Re-evaluate fail-open after every CVE. Fail-open magnifies any vulnerability that lets a downstream client influence whether the auth check happens. Keep Envoy current if you run fail-open.
  • Propagate the header. If you run fail-open, configure upstreams to honor x-envoy-auth-failure-mode-allowed and degrade gracefully. Logging it also gives you an application-layer audit trail of the fail-open window.

How Netdata helps

The failure_mode_allowed counter should be flat zero in steady state. Per-second collection and anomaly detection surface the first increment without waiting for a scrape interval, and correlating auth cluster health metrics alongside the ext_authz counters shortens the diagnosis path.

  • Per-second collection of ext_authz.ok, ext_authz.denied, ext_authz.error, and ext_authz.failure_mode_allowed lets you see the exact second the auth service started failing and the exact second fail-open traffic started flowing.
  • ML anomaly detection on failure_mode_allowed flags the first increment in a flat-zero series without needing a fixed threshold.
  • Correlating ext_authz.error against the auth cluster’s upstream_cx_connect_fail, membership_healthy, and upstream_rq_5xx in a single view shortens the path from “auth is failing” to “auth cluster host 3 is ejecting”.
  • Tracking ext_authz.latency alongside downstream_rq_time makes the cost of a slow auth service visible as direct proxy overhead rather than a mysterious latency regression.
  • The same dashboard can surface the auth cluster’s circuit breaker state, so connection pool exhaustion to the auth service (a common root cause of fail-open bursts) is visible alongside the failure_mode_allowed counter that it triggers.