The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / envoy / envoy-membership-healthy-dropping ▌

Operations Guides

Envoy membership_healthy dropping: reading the single most important cluster signal

If you watch only one availability metric per Envoy cluster, watch the ratio of cluster.<name>.membership_healthy to cluster.<name>.membership_total. When the ratio collapses, the remaining healthy hosts take proportionally more load, and the cluster is one or two failures away from panic mode. Upstream 5xx rate, latency, circuit breaker state, and retries are all downstream of host availability.

The hard part is reading the gauge correctly, not collecting it. A drop can mean the upstream is genuinely broken, the control plane removed endpoints, or outlier detection ejected hosts that are still passing active health checks. Each case has a different response.

What the signal actually means

cluster.<name>.membership_healthy is a gauge. It reports the current count of hosts in the cluster that Envoy considers available for load balancing, after both active health checking and outlier detection have applied their logic. It is inclusive of both subsystems, which is the first thing operators get wrong.

The full membership family for a cluster:

StatTypeMeaning
cluster.<name>.membership_healthygaugehosts healthy enough to receive traffic
cluster.<name>.membership_degradedgaugehosts in degraded state
cluster.<name>.membership_excludedgaugehosts excluded from panic-threshold calculations
cluster.<name>.membership_totalgaugefull cluster membership as the control plane sees it
cluster.<name>.membership_changecountertotal membership change events

What makes membership_healthy the right anchor is that it already incorporates the two subsystems you would otherwise have to reason about separately. A host that passes active health checks but has been ejected by outlier detection does not count as healthy. A host that is healthy by both subsystems counts as healthy. You do not have to do that math yourself.

The number to compute and alert on is the ratio membership_healthy / membership_total, not the absolute count. A cluster with 9 of 10 hosts healthy (0.90) is in far better shape than one with 2 of 3 (0.67), even though both have only one unhealthy host. The ratio tells you how concentrated the remaining load is about to become.

How to read it correctly

  • Ratio below 0.5 is critical. Below this point the cluster is one failed host away from panic mode, and the remaining healthy hosts are absorbing at least double their normal share.
  • Any non-zero membership_degraded warrants investigation. Degraded is not healthy. It is Envoy saying “this host is partially broken, route it less traffic.” A cluster with zero unhealthy hosts but several degraded hosts is trending the wrong way.
  • membership_healthy == 0 with membership_total > 0 is total upstream failure for that cluster, but only if traffic is actually flowing to it. Gate the alert: require upstream_rq_total > 0 to rule out idle clusters, and server.live == 1 (gauge; also available as the state field on the /server_info admin endpoint) to rule out draining. Check Envoy uptime via the /server_info admin endpoint to rule out cold start and warming. Sustain the check across at least two health-check intervals to avoid paging on flaps.
  • Clusters with zero routed traffic can sit unhealthy with no user impact. This is the false positive you must gate out. The traffic floor (upstream_rq_total > 0) is not optional.

Why a drop is not always a failure

This is the section that saves you from paging at 3 a.m. on a deployment.

In EDS-based deployments (Istio service mesh, any xDS control plane that pushes endpoints), membership_total is not a constant. It shifts as the control plane adds and removes endpoints. A scale-down event reduces membership_total. A rolling deploy churns it. If membership_healthy drops in lockstep with membership_total and the ratio stays roughly constant, the cluster is not getting less healthy. It is getting smaller, by design.

The failure case is the opposite: membership_total is flat or growing while membership_healthy drops. That is hosts failing in place. The ratio is the truth, and you confirm it by watching whether the two gauges move together or apart.

flowchart TD
    A["membership_healthy drops"] --> B{"membership_total also dropping?"}
    B -- "yes, in step" --> C["Likely control-plane churn
check update_success / update_empty"] B -- "no, flat or growing" --> D["Hosts failing in place
this is the real signal"] C --> E{"Ratio stable?"} E -- "yes" --> F["Scale event or deploy. Not an outage."] E -- "no" --> D D --> G["Correlate: upstream_cx_connect_fail,
health_check.failure,
outlier_detection.ejections_active"]

Two adjacent stats help you tell control-plane churn from real loss:

  • cluster.<name>.update_success and cluster.<name>.update_empty should be incrementing if the membership change came from EDS. update_empty in particular means the control plane pushed an update with zero endpoints, which is expected during scale-to-zero or a deploy lull.
  • cluster.<name>.update_failure means the update failed or Envoy refused the new config. Envoy keeps the old config and the membership numbers do not move. This is a control-plane problem, not a host-health problem. (NACK counts are tracked at the xDS subscription level — cluster_manager.cds.update_rejected for CDS — not as a per-cluster stat.)

What to correlate when the drop is real

Once you have confirmed the ratio is genuinely falling (hosts failing in place, not being removed by the control plane), three signals tell you why.

upstream_cx_connect_fail. This counter fires when the TCP connection attempt to an upstream host fails: process crash, port not listening, FD exhaustion, listen backlog overflow, firewall blocking. A rising connect_fail rate that tracks the falling membership_healthy is a host-availability problem, not a host-performance problem. Connect failures also feed outlier detection, so a host failing to connect will often be ejected on top of failing its health check.

health_check.failure and the health-check family. Active health checks run on the main thread, not worker threads. A rising failure rate here means the host is accepting the connection but not responding correctly to the probe: slow app, wrong response, timeout. The health-check interval bounds your detection latency (see the next section), so the rate of failure is the leading indicator and the count of failed hosts is the lagging one.

outlier_detection.ejections_active. Outlier detection is passive, based on real traffic, and operates independently from active health checks. A host can be passing its health checks and still be ejected by outlier detection because real requests are failing on it. membership_healthy reflects both. If you see the gauge dropping but health_check.failure is flat, check ejections_active and the per-type counters (ejections_enforced_consecutive_5xx, ejections_enforced_success_rate, ejections_enforced_failure_percentage) to find out which outlier rule is firing.

The other direction matters too. A mass-ejection event where ejections_active approaches membership_total is often a correlated infrastructure failure (network, AZ, shared dependency) rather than independent host failures. Hosts get ejected, load concentrates on the survivors, the survivors slow down, they get ejected, and the cluster falls through the panic threshold.

Detection latency is bounded by the health-check interval

Active health checks detect a failed host only as fast as the configured interval allows. A 30-second health-check interval means up to 30 seconds between a host going dark and Envoy marking it unhealthy. During that window, requests are still being routed to the dead host and failing.

The interval is the dominant term. If your SLO for upstream failure detection is faster than your health-check interval, the interval is the bug. Tighten the interval, or add passive outlier detection to catch failures that active checks would otherwise wait out.

Two practical implications:

  • Sustain the alert across at least two health-check intervals. A single failed check is a flap. Two in a row is a host.
  • Outlier detection does not save you for cold hosts. Outlier detection only acts on hosts that are receiving traffic. A host that just failed and has not received a request yet will not be ejected by outlier detection. Active health checks are how you detect a newly-failed host that has no traffic to fail. The two subsystems are complementary, not redundant.

Panic threshold changes the meaning of the ratio

When the healthy percentage drops below the panic threshold (default 50%, configurable per cluster), Envoy stops honoring health status and load-balances across all hosts, including unhealthy ones. This is intentional: it prevents the few remaining healthy hosts from being overloaded to death.

It also changes how you read the gauge. Once panic mode is active, membership_healthy is no longer the count of hosts receiving traffic. Envoy is routing to ejected and unhealthy hosts because the alternative is worse. Increased error rates during panic mode are the designed behavior, not a second failure layered on top of the first.

If you alert on upstream_rq_5xx and the cluster is below panic threshold, the 5xx rate is expected to be bad. The actionable signal is the ratio climbing back above the threshold, not the error rate. Do not treat panic-mode error rates as a new incident.

A stats-matcher gotcha

The HTTP health check filter, when run in “computed from upstream cluster health” mode, does not probe a backend. It reads the cluster’s in-memory membership gauges and returns 200 or 503 based on them. Those gauges are subject to stats_matcher: if you have configured reject_all: true or an exclusion list that drops the membership stats, the gauges are never instantiated and read as zero, causing the filter to return 503 regardless of actual cluster health. Envoy’s own documentation warns that excluding stats may affect behavior in undocumented ways; upstream issue #8771 documents a related case where an excluded metric broke cluster health-check interval selection.

The fix is to allowlist the membership stats explicitly in your stats_matcher:

stats_config:
  stats_matcher:
    inclusion_list:
      patterns:
      - suffix: membership_healthy
      - suffix: membership_degraded
      - suffix: membership_total
      - suffix: live

This is an easy mistake to make when pruning stats for cardinality reasons. The symptom is a health check endpoint that is permanently unhappy even though /clusters?format=json shows healthy hosts.

Signals to watch in production

SignalWhy it mattersWarning sign
membership_healthy / membership_totalThe ratio is the cluster availability signal.Sustained drop below 0.5, or any sustained drop with membership_total flat.
membership_degradedDegraded hosts are partially broken.Any non-zero value warrants investigation.
membership_changeCounts membership churn events.Spikes during steady state suggest control-plane instability.
upstream_cx_connect_failHosts not accepting connections.Rate tracking the membership_healthy drop confirms host-level failure.
health_check.failureActive probes failing.Rising rate is the leading indicator before hosts are marked unhealthy.
outlier_detection.ejections_activePassive ejection of hosts with bad real traffic.Hosts passing health checks but ejected means real traffic is failing.
update_success / update_emptyEDS updates landing.update_empty sustained for a cluster that should have endpoints.
update_failureEnvoy rejecting or failing to apply config.Any non-zero value; the control plane pushed something invalid.
lb_healthy_panicPanic mode active.Cluster is routing to all hosts including unhealthy ones.

How Netdata helps

  • Per-second collection of the membership_healthy, membership_degraded, membership_excluded, and membership_total gauges means you see the ratio move as it happens. Detection latency in your monitoring should not be the bottleneck when Envoy’s own health-check interval already sets a floor.
  • ML anomaly detection on the ratio catches gradual drift (one host every few minutes) that fixed thresholds miss, even when the absolute number is still above a static threshold.
  • Correlating membership_healthy with upstream_cx_connect_fail, health_check.failure, and outlier_detection.ejections_active on a single timeline lets you distinguish host-level failure from control-plane churn from passive ejection in seconds.
  • The same correlation works for ruling out false positives: membership_total moving in lockstep with membership_healthy, plus update_empty incrementing, shows up as a deploy or scale event rather than an outage.
  • Alerting on the ratio with a traffic floor (upstream_rq_total > 0) and a sustained-duration window (at least two health-check intervals) keeps you from paging on cold-start, warming, or idle clusters.