The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / envoy / envoy-upstream-connection-reuse-churn ▌

Operations Guides

Envoy connection churn: a low reuse ratio and keepalive misconfiguration

Envoy’s upstream connection pool amortizes TCP (and TLS) handshakes across many requests. When it stops doing that, you have connection churn: the pool establishes a fresh connection for nearly every request, and the reuse ratio collapses toward 1.0.

The headline signal is the ratio upstream_rq_total / upstream_cx_total. For HTTP/1.1 with keepalive working, expect a number well above 1, often in the tens. For HTTP/2 with stream multiplexing, expect far higher still. When the ratio sits near 1.0, every request pays the full handshake tax.

Churn rarely pages anyone. The cluster looks healthy: membership_healthy is stable, 503s are not firing, error rates are flat. What you see instead is an elevated upstream_cx_connect_ms rate, extra worker CPU burned on TLS, more file descriptors cycling through the process, and a tail-latency bump on upstream_rq_time for every request that lands on a cold connection. This article covers how to compute the ratio correctly, identify the specific close-cause from Envoy’s per-connection counters, and fix the three common root causes: keepalive disabled, max_connection_duration recycling, and DNS TTL churn on STRICT_DNS clusters.

What this means

The reuse ratio compares two monotonic counters on the cluster stat prefix:

  • cluster.<name>.upstream_rq_total: total requests (including retries) sent upstream.
  • cluster.<name>.upstream_cx_total: total upstream connections ever established.

ratio = upstream_rq_total / upstream_cx_total. For a clean operational read, compute it from per-window rate deltas rather than cumulative counters since process start, which are dominated by cold-start noise.

Envoy also exposes cluster.<name>.upstream_rq_per_cx, a histogram that directly measures the number of requests handled per upstream connection across all HTTP protocols. This is the official, version-stable way to read reuse quality. A healthy distribution is centered well above 1; a distribution pinned at 1 means the pool is one-shotting connections.

A low ratio is expensive in three ways. First, every churned connection adds a TCP handshake (and, when upstream TLS is configured, a TLS handshake) to per-request latency, which shows up as an elevated upstream_cx_connect_ms rate. Second, TLS handshakes are typically the dominant CPU consumer in Envoy, so churn translates directly into worker CPU. Third, each connection consumes two file descriptors (downstream plus upstream) plus kernel socket overhead, so churn inflates FD cycling and can interact with FD exhaustion during spikes.

Envoy tells you exactly why each connection closed. The cluster stats include dedicated close-cause counters that decompose upstream_cx_total into its contributors. The diagnostic flow is to find which counter is consuming your pool.

flowchart TD
  A["Low reuse ratio
rq_total / cx_total near 1"] --> B["Inspect per-connection close counters"] B --> C{"Which counter is rising?"} C -->|upstream_cx_max_requests| D["max_requests_per_connection = 1
keepalive effectively disabled"] C -->|upstream_cx_max_duration_reached| E["max_connection_duration recycling pool"] C -->|upstream_cx_close_notify| F["Idle timeout or upstream GOAWAY"] C -->|none obvious| G["STRICT_DNS churn on short TTL"]

Common causes

CauseWhat it looks likeFirst thing to check
Keepalive disabled (max_requests_per_connection: 1)upstream_cx_max_requests is the dominant close-cause; ratio pinned near 1.0 even on HTTP/1.1Cluster proto max_requests_per_connection
max_connection_duration set too lowupstream_cx_max_duration_reached climbs in step with upstream_cx_totalcommon_http_protocol_options.max_connection_duration
STRICT_DNS with short DNS TTLupdate_success rate is high; upstream_cx_total spikes after each resolution; membership stableCluster type: STRICT_DNS, respect_dns_ttl, upstream DNS TTL
Upstream sending GOAWAY or Connection: closeupstream_cx_close_notify rising; pool otherwise healthyUpstream server close behavior, upstream idle timeout
Idle timeout too aggressiveupstream_cx_idle_timeout is the dominant close-causecommon_http_protocol_options.idle_timeout

Quick checks

The admin port is 9901 in standalone Envoy and 15000 in Istio sidecar mode. Adjust accordingly.

# Reuse ratio inputs for a specific cluster (lifetime counters)
curl -s http://localhost:9901/stats | grep -E 'cluster\.my_cluster\.(upstream_rq_total|upstream_cx_total)'

# Per-connection request histogram (the official reuse signal)
curl -s http://localhost:9901/stats | grep 'upstream_rq_per_cx'

# Close-cause counters - which one is consuming your pool?
curl -s http://localhost:9901/stats | grep -E 'upstream_cx_(max_requests|max_duration_reached|idle_timeout|close_notify)'

# Rate of new connection establishment (should be low at steady state)
curl -s http://localhost:9901/stats | grep 'upstream_cx_total'

# TCP (and TLS) connect time - elevated when churn forces fresh handshakes
curl -s http://localhost:9901/stats | grep 'upstream_cx_connect_ms'

# DNS update cadence for STRICT_DNS / LOGICAL_DNS clusters
curl -s http://localhost:9901/stats | grep -E 'cluster\.my_cluster\.update_(success|failure)'

# Active pool size and pending requests
curl -s http://localhost:9901/stats | grep -E 'cluster\.my_cluster\.(upstream_cx_active|upstream_rq_pending_active)'

For high-traffic proxies, prefer /stats/prometheus with the ?usedonly filter and a 30s+ scrape interval. Scraping the plain /stats endpoint on a very large config can be expensive and can itself contribute to the latency symptoms you are investigating.

How to diagnose it

  1. Compute the ratio over a clean window. Sample upstream_rq_total and upstream_cx_total 60 seconds apart and divide the deltas. Do not rely on cumulative counters, which are dominated by cold-start behavior.
  2. Confirm with upstream_rq_per_cx. If the histogram is pinned at 1, the pool is one-shotting connections. If it is centered at, say, 5, reuse is happening but is weaker than expected.
  3. Find the close-cause. Sum the close-cause counters (upstream_cx_max_requests, upstream_cx_max_duration_reached, upstream_cx_idle_timeout, upstream_cx_close_notify) and find the dominant contributor. The largest one names your failure mode.
  4. Correlate with DNS. If no single close-cause dominates but upstream_cx_total spikes in lockstep with update_success on a STRICT_DNS cluster, DNS churn is resetting the pool. Check dns.cares.resolve_total for the resolution rate.
  5. Check per-worker state. Connection pools are per-worker and per-host. A cluster with 10 hosts across 8 workers has up to 80 independent pools. Aggregate stats can hide a single host or worker that is churning while the rest are healthy. Drill into /clusters?format=json for per-host detail during incidents.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
upstream_rq_total / upstream_cx_totalThe reuse ratio; the headline churn signalHTTP/1.1 near 1.0; HTTP/2 well below ~100
upstream_rq_per_cxOfficial histogram of requests per connectionDistribution pinned at 1
upstream_cx_max_requestsConnections closed due to max_requests_per_connectionDominant close-cause counter
upstream_cx_max_duration_reachedConnections closed due to max_connection_durationClimbing in step with upstream_cx_total
upstream_cx_idle_timeoutConnections closed due to idle timeoutDominant close-cause counter
upstream_cx_close_notifyConnections closed via GOAWAY or Connection: closeRising; indicates upstream-driven close
upstream_cx_connect_msTCP (and TLS) connect latencyElevated rate when churn forces fresh handshakes
cluster.<name>.update_successDNS update cadence (STRICT_DNS / LOGICAL_DNS)High rate correlating with pool resets
upstream_cx_activeCurrent pool sizeSpiky rather than steady at steady-state traffic

Fixes

Keepalive disabled (max_requests_per_connection: 1)

Per Envoy’s cluster proto documentation, setting max_requests_per_connection to 1 “will effectively disable keep alive.” This is the single most common cause of a reuse ratio pinned at 1.0. It is usually a copy-paste from a tuning guide or a leftover from debugging.

Fix: remove the field entirely (the default is unlimited) or set it to a large value. A finite cap can help spread load across connections and bound the blast radius of a single connection, but the cap should be in the hundreds or thousands, not 1.

max_connection_duration recycling

max_connection_duration defaults to 0 (unlimited). The Envoy timeouts FAQ notes that setting it can help with DNS-based clusters where resolved addresses may change even when upstreams stay healthy. But if it is set low (for example, 30 seconds), every connection in the pool cycles on that interval, and the reuse ratio floors at roughly duration / mean_interarrival_time.

Fix: raise it, remove it, or pair it with LOGICAL_DNS so the pool does not need to recycle to pick up DNS changes. If you genuinely need cycling for DNS reasons, measure the resulting reuse ratio and confirm the latency cost is acceptable.

STRICT_DNS churn

Per the service discovery documentation, STRICT_DNS clusters drain and recreate connection pools whenever DNS resolution returns a different IP set. The docs explicitly warn that with changing DNS, STRICT_DNS “would lead to draining connection pools, connection cycling, etc.” If respect_dns_ttl is enabled and upstream DNS records have short TTLs (common with cloud load balancers and service discovery systems), the pool resets on every refresh.

The recommended alternative is LOGICAL_DNS. Per the docs, with LOGICAL_DNS, “connections stay alive until they get cycled.” LOGICAL_DNS treats the resolved address set as a single logical host, so connection pools survive DNS refreshes as long as the resolved IP does not fully disappear.

Tradeoff: LOGICAL_DNS does not load balance across multiple resolved IPs the way STRICT_DNS does. Use it when you are pointing at a single logical upstream (a load balancer, a service VIP) rather than a pool of discrete endpoints. For real endpoint pools, prefer EDS via the control plane, which pushes endpoint updates without churning connections the way DNS does.

Upstream-driven close (GOAWAY / idle timeout)

If upstream_cx_close_notify is the dominant close-cause, the upstream is closing connections. For HTTP/2 and HTTP/3 this is a GOAWAY; for HTTP/1.1 it is a Connection: close header. The upstream server may be applying its own idle timeout or its own max-requests-per-connection policy.

Envoy does not propagate downstream Connection: close headers to upstream. The upstream connection lifecycle is decoupled from the downstream. So if you see upstream_cx_close_notify rising on the upstream side, it is the upstream server doing it, not your downstream clients.

Fix: align Envoy’s idle_timeout (default 1 hour) with the upstream’s idle behavior. If the upstream forces closes earlier than Envoy’s timeout, lower Envoy’s idle timeout to match so you do not hold dead connections. Check the upstream’s own keepalive and max-requests settings as well.

Prevention

  • Track the reuse ratio per cluster. The playbook’s expert monitoring tier lists upstream_rq_total / upstream_cx_total as a Level 4 signal. Treat it as a baseline metric, not an incident-only check.
  • Alert on close-cause balance. In a healthy pool, no single close-cause counter should dominate. Alert when any one counter exceeds the others by a meaningful margin.
  • Default to LOGICAL_DNS or EDS for short-TTL upstreams. Reserve STRICT_DNS for cases where the resolved IP set is genuinely stable.
  • Pair max_connection_duration changes with a reuse-ratio measurement. Any change to connection lifetime policy should be validated against the ratio.
  • Audit cluster templates for max_requests_per_connection. A single misconfigured cluster template can churn across every service that imports it.

How Netdata helps

  • Per-second collection of upstream_rq_total and upstream_cx_total lets you compute the reuse ratio on a tight window without waiting for a slow scrape interval, which matters when churn starts mid-incident.
  • ML anomaly detection flags sudden drops in the reuse ratio even when no static threshold would catch it, for example when a config push silently sets max_requests_per_connection: 1.
  • Correlating upstream_cx_connect_ms with the close-cause counters in a single view shortens the path from “latency is up” to “the pool is churning because of max_connection_duration.”
  • Worker CPU and TLS handshake rate sit alongside the cluster stats, so the CPU cost of churn is visible in real time rather than inferred.
  • DNS update counters for STRICT_DNS and LOGICAL_DNS clusters pair with the reuse ratio to confirm DNS-driven resets without a separate dashboard.