The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / envoy / envoy-response-flags ▌

Operations Guides

Envoy response flags: decoding UO, NR, UF, UT, UC and the rest

A 503 from Envoy can mean a dozen different things. The HTTP status code tells you what the client saw; it does not tell you why Envoy generated that response. The %RESPONSE_FLAGS% access log field separates a circuit breaker trip from a missing route, a dead upstream, or a client that hung up.

Response flags are the most precise debugging signal Envoy emits, and the most operationally misunderstood. They are not aggregate stats; they appear only in access logs. Most teams discover this the first time they try to alert on UO and find no Prometheus counter for it.

This is a reference for the response flags you will see in production: what each means, where in the request lifecycle it fires, what status code it typically pairs with, and what to check next. It covers UF, UO, NR, NC, URX, UT, UC, DC, LR, RL, UAEX, IH, SI, DPE, DI. Newer flags exist in the Envoy source for DNS resolution failure, overload manager actions, and drop overload; if you see a flag not covered here, source/common/stream_info/utility.h in the Envoy repository is the canonical list.

What response flags are (and what they are not)

%RESPONSE_FLAGS% is an access log formatter operator. Envoy sets zero or more flags on each request to record why it handled it the way it did. The flags explain the cause behind a response code, not the response code itself.

Three properties matter operationally.

Access log only. Response flags are not exposed as aggregate stats. cluster.<name>.upstream_rq_503 counts both Envoy-generated 503s and 503s forwarded from the upstream. The response flag is the only way to tell them apart.

Per-request, not per-cluster. Flags attach to individual requests. There is no cluster.<name>.response_flags.UO counter.

Comma-separated when multiple apply. A single request can carry UC,URX: the upstream closed the connection and retries are exhausted. Naive string matching on logs misses these.

Alerting on response flags therefore requires a log pipeline, a sidecar counter, or a Lua/Wasm filter that increments a custom stat. Envoy itself does not aggregate them.

Where to find them

Response flags appear wherever %RESPONSE_FLAGS% is in the access log format string. The default Envoy format includes it. In Istio, the default sidecar format also includes it, but custom configurations sometimes strip it; verify before you rely on it.

A typical line:

[2024-01-15T10:23:45.123Z] "GET /api/v1/users HTTP/2" 503 UO 0 91 3 - "..." "..." "abc123"

The response code is 503; the field immediately after it (UO) is the response flag. When no flag is set, Envoy emits -. A dash means Envoy did not flag this request, not that the request was healthy: a 404 forwarded from an upstream also produces - because Envoy did not intervene.

To extract counts without a full log pipeline:

# Default Envoy format: %RESPONSE_FLAGS% is the 6th whitespace-delimited field
# (the quoted request line splits into three). Adjust if your format is custom.
awk '{print $6}' /var/log/envoy/access.log | sort | uniq -c | sort -rn

The flag reference

The flags group by where in the request lifecycle they fire. Read them in that order when debugging: configuration first, then circuit breaking, then upstream connection, then timeouts, then downstream.

Configuration errors: NR and NC

These are the only flags that almost always indicate a bug in the configuration, not a transient network issue. Any non-zero rate of NR or NC in production is a configuration error.

FlagMeaningTypical responseFirst thing to check
NRNo route found503Route table does not match the request. Classic sign of a bad xDS push.
NCNo cluster found503Cluster referenced by the route was removed or never existed.

NR immediately after an xDS push is the textbook signature of a bad config deployment. The control plane reports success, Envoy accepted the config, but the route table no longer matches traffic that worked a minute ago. Check update_rejected (per-resource NACKs) and listener_manager.listener_create_failure; a config that is valid but wrong produces no NACK, so a clean update_success does not rule this out.

Circuit breaker and saturation: UO

UO means Envoy fast-failed a request locally because a per-cluster circuit breaker was open. Envoy generated the 503; the upstream never saw the request.

FlagMeaningTypical responseFirst thing to check
UOUpstream overflow (circuit breaker tripped)503circuit_breakers.<priority>.cx_open, rq_pending_open, rq_open, rq_retry_open

UO should be zero in steady state. Any sustained UO rate means a breaker is open, which means either the upstream is slow (filling the connection pool) or the limits are set too low for the workload. Increasing the limits is rarely the right fix; the breaker is correctly reporting that the upstream cannot absorb more load. See Envoy circuit breaker open: cx_open, rq_pending_open, and fast-failed requests.

Upstream connection failures: UF and UC

Both indicate a problem reaching the upstream, but at different points in the connection lifecycle.

FlagMeaningTypical responseFirst thing to check
UFUpstream connection failure503cluster.<name>.upstream_cx_connect_fail. TCP connect failed: port not open, host down, or firewall blocking.
UCUpstream connection termination502 or 503cluster.<name>.upstream_rq_rx_reset. Upstream sent RST or closed the connection mid-request.

UF means the SYN never completed (or was RST immediately). UC means the connection was established but the upstream tore it down during the request. The distinction maps to two different upstream failure modes: host not listening versus host crashing mid-response. See Envoy upstream_cx_connect_fail: failed TCP connections to upstream hosts.

Upstream timeouts: UT and URX

These two often co-occur. UT says Envoy gave up waiting. URX says Envoy exhausted its retry budget while trying.

FlagMeaningTypical responseFirst thing to check
UTUpstream request timeout504cluster.<name>.upstream_rq_timeout, upstream_rq_per_try_timeout. Upstream is slow or timeout is misconfigured.
URXUpstream retry limit exceededLast response receivedcluster.<name>.upstream_rq_retry, upstream_rq_retry_overflow. Retries did not rescue the request.

URX paired with UC (UC,URX) is a common combination: upstream keeps closing connections, retries keep failing, and the retry budget is exhausted.

Downstream disconnects: DC and DPE

These describe what the client did, not what Envoy did wrong. They are often noise.

FlagMeaningTypical responseFirst thing to check
DCDownstream connection terminationNo response sentOften benign. Investigate only if correlated with latency spikes.
DPEDownstream protocol error4xxClient sent malformed protocol data.

DC is the most over-alerted flag. Clients close connections for many benign reasons: page navigation, mobile app backgrounding, cancellation of slow requests. A baseline DC rate is normal. Investigate only when DC spikes alongside elevated latency, which suggests clients are timing out waiting for Envoy.

Local resets and rate limiting: LR and RL

FlagMeaningTypical responseFirst thing to check
LRLocal reset503Envoy reset the connection locally. Often correlates with circuit breaker or filter behavior.
RLRate limited429Local token bucket or global rate limit service rejected the request.

RL pairs with the rate limit service over_limit result counter. LR is less specific; check circuit breaker and filter stats to find the cause.

Authorization and headers: UAEX and IH

FlagMeaningTypical responseFirst thing to check
UAEXUnauthorized external service403ext_authz.denied. External auth service denied the request.
IHInvalid header400Request had a header Envoy rejects by spec.

UAEX spikes after a credential rotation or policy change. If ext_authz.error is also climbing, the auth service itself is degraded.

Idle and injection: SI and DI

FlagMeaningTypical responseFirst thing to check
SIStream idle timeout408Stream was idle past the configured idle timeout. Common with mobile or long-poll traffic.
DIDelay injectedVariesFault injection filter is active. Verify it is intentional.

DI only appears when a fault injection filter is configured. Seeing it unexpectedly means someone enabled chaos testing in production.

Multiple flags at once

Envoy emits flags as a comma-separated list when more than one applies. The combination carries information that a single flag does not.

CombinationWhat it tells you
UC,URXUpstream closed the connection, retry budget exhausted
UF,URXCould not connect to upstream, retries exhausted
UT,URXUpstream timed out, retries exhausted
UO aloneCircuit breaker open, no upstream attempt made
NR aloneConfiguration error, no upstream attempt made

A retry-related flag (URX) paired with a connection flag (UC, UF) is the signature of a retry storm: the upstream is failing, retries are firing, and the budget is gone. Check the upstream_rq_retry / upstream_rq_total ratio. See Envoy connection pool exhaustion: a slow upstream that fills the pool for the broader pattern.

A diagram for the common 503

When you see a 503, the response flag tells you which subsystem produced it. The flow maps the common cases.

flowchart TD
    REQ[Client request arrives] --> ROUTE{Route exists?}
    ROUTE -- no --> NR[NR: 503]
    ROUTE -- yes --> CLUSTER{Cluster exists?}
    CLUSTER -- no --> NC[NC: 503]
    CLUSTER -- yes --> BREAKER{Circuit breaker open?}
    BREAKER -- yes --> UO[UO: 503, fast-fail]
    BREAKER -- no --> CONNECT{TCP connect ok?}
    CONNECT -- no --> UF[UF: 503]
    CONNECT -- yes --> MIDCONN{Upstream stays connected?}
    MIDCONN -- no --> UC[UC: 502 or 503]
    MIDCONN -- yes --> TIMEOUT{Responds in time?}
    TIMEOUT -- no --> UT[UT: 504]
    TIMEOUT -- yes --> RETRY{Retries needed?}
    RETRY -- exhausted --> URX[URX: last response]
    RETRY -- ok --> OK[2xx, 3xx, or 4xx]

Patterns worth alerting on

Because response flags are access-log only, you need a log pipeline or a custom Lua/Wasm counter to alert on them. Once you have that, common severity guidance is:

  • Any NR or NC in production. Configuration error. Page or ticket immediately.
  • Sustained UO. A breaker is open. Investigate pool and limit configuration; usually the upstream is slow, not Envoy.
  • Elevated UF, UC, or UT. Upstream is failing or slow. Correlate with upstream_cx_connect_fail, upstream_rq_rx_reset, upstream_rq_timeout.
  • Baseline DC. Track it but do not page on it alone. Page only when it correlates with latency spikes.
  • URX climbing with retry / total above 0.3. Retry storm. The retry policy is amplifying the failure.

Gotchas

  • - is not “healthy”. A dash means Envoy did not set a flag. The request may still have failed (for example, a 500 forwarded from the upstream). The flag tells you why Envoy intervened, not whether the response was successful.
  • DC is mostly noise. Mobile clients, page navigations, and request cancellations all produce DC. Investigate only when it correlates with latency.
  • NR after an xDS push is a bad deploy. If NR appears right after a control plane push, the route table is wrong. Check update_rejected and the control plane push logs.
  • Flags are access-log only. There is no Prometheus counter for UO. If you need real-time alerting, add a Lua filter that increments a custom stat, or feed logs to a pipeline that counts them.
  • Multiple flags concatenate. A grep anchored on UO will not match UO,URX. Account for comma-separated values in log parsing.
  • Newer flags exist. Envoy has added flags for DNS resolution failure, overload manager actions, drop overload, and others not covered here. If you see a flag you do not recognize, source/common/stream_info/utility.h in the Envoy repository is the canonical list.

How Netdata helps

  • Correlate response flags with the stats that explain them. A spike in UO in the logs should be read against circuit_breakers.default.cx_open, rq_pending_open, upstream_rq_pending_overflow, and upstream_cx_active. Per-second metrics make the timing of the breaker trip visible against the upstream latency rise that caused it.
  • Distinguish Envoy-generated 5xx from upstream-forwarded 5xx. cluster.<name>.upstream_rq_503 counts both. Overlay log-derived UO and NR counts on the 503 rate to see which fraction Envoy produced versus forwarded.
  • Catch the upstream-side causes early. UF lines up with upstream_cx_connect_fail. UC lines up with upstream_rq_rx_reset. UT lines up with upstream_rq_timeout. Per-second collection shows these leading indicators before the flags appear in logs.
  • Surface retry storms. URX in the logs pairs with upstream_rq_retry, upstream_rq_retry_overflow, and the upstream_rq_total / downstream_rq_total ratio. Watching these together tells you whether retries are helping or amplifying.
  • Catch configuration regressions fast. Cross-check NR and NC against control_plane.connected_state, update_rejected, and listener_manager.listener_create_failure so a bad xDS push is visible within seconds, not at the next log scan.