The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / envoy / envoy-outlier-detection-mass-ejection ▌

Operations Guides

Envoy outlier detection mass ejection: when passive health checks empty a cluster

membership_healthy is collapsing. outlier_detection.ejections_active is climbing toward membership_total, and ejections_overflow is ticking up because Envoy wanted to eject more hosts than max_ejection_percent allows. This is the outlier detection mass ejection pattern: one of the few Envoy failure modes that can amplify a partial upstream degradation into a cluster-wide outage.

The mechanism is subtle because outlier detection is doing exactly what it was configured to do. It is a passive health check: it ejects hosts based on the actual traffic they are serving, not on synthetic probes. When one host starts returning 5xx or its success rate drops, ejecting it is correct. The problem is what happens next. Load concentrates on the survivors, they get slower, their success rates drop, and they get ejected too. The cascade feeds itself.

This article covers how to recognize the cascade in real time, distinguish a genuine systemic upstream problem from too-sensitive thresholds, and stop the cascade without making things worse. It assumes you already understand the broad failure pattern catalogue in the Envoy operations hub.

What this means

Outlier detection is independent from active health checks. A host can pass active health checks (synthetic probes succeed) and still be ejected by outlier detection because real traffic is failing. The reverse is also true. When you see membership_healthy falling while membership_total stays constant, the most likely cause is outlier detection, not EDS-driven membership changes.

The cascade has a defined end state. Once the healthy host percentage drops below the panic threshold (default 50%), Envoy stops protecting the survivors and starts load balancing across all hosts, including ejected ones. This is intentional: it prevents the few remaining healthy hosts from being overloaded to death. The tradeoff is that traffic now flows to known-bad endpoints. If panic_threshold is set to 0%, Envoy instead returns 503 “no healthy upstream” when all hosts are ejected. Either way, users see degraded or failed responses, and the operator sees a system that looks broken even though Envoy is behaving as designed.

The critical question during the incident is not “why is Envoy ejecting hosts” but “are these ejections correct, or is outlier detection reacting to a self-inflicted load spike?”

flowchart TD
    A[Some hosts 5xx or
success_rate drops] --> B[Outlier detection
ejects a host] B --> C[Load concentrates
on remaining hosts] C --> D[Survivors slow down,
success_rate falls] D --> E{Threshold trips
again?} E -->|yes| B E -->|max_ejection_percent hit| F[ejections_overflow climbs
ejections_active plateaus] F --> G{healthy below
panic_threshold 50%?} G -->|yes| H[Panic mode: route to ALL hosts
including ejected] G -->|panic_threshold = 0| I[503 no healthy upstream]

Common causes

CauseWhat it looks likeFirst thing to check
Too-sensitive consecutive_5xx threshold (default 5)Ejections fire on a single burst of errors during a deploy or GC pause; surviving hosts recover after ejectionejections_enforced_consecutive_5xx rate vs ejections_detected_consecutive_5xx
success_rate ejection on a small clusterCluster has fewer than 5 hosts (the success_rate_minimum_hosts default), so one slow host triggers a chain reactionejections_enforced_success_rate and cluster size
Genuine systemic upstream degradationAll hosts degrade at once: shared dependency failure, AZ network issue, DB contentionPer-host error pattern from /clusters, correlation across hosts
max_ejection_percent misconfigured (default 10%)Ejections unbounded or capped too high; cluster can be emptied instead of shedding loadOutlier detection config and ejections_overflow behavior
Stats matcher hides ejections_activeIn some mesh deployments, the ejection gauge is not exported, making the cascade invisibleEnvoy stats config, whether ejections_active appears at all
Pre-1.28 max_ejection_percent calculation bugWith small clusters, more hosts ejected than the percentage allowedFixed in 1.28.0 (changelog: “Outlier detection will always respect max_ejection_percent now”, runtime guard envoy.reloadable_features.check_mep_on_first_eject can revert); PR 27624 was part of the enforcement work

Quick checks

Run these against the Envoy admin interface. Default port is 9901; Istio sidecars use 15000 with /healthz/ready on 15021.

# Check current ejection volume per cluster
curl -s http://localhost:9901/stats | grep 'outlier_detection.ejections_active'

# Confirm the cap is being hit (Envoy wanted to eject more than max_ejection_percent allows)
curl -s http://localhost:9901/stats | grep 'ejections_overflow'

# Break down which detector is firing
curl -s http://localhost:9901/stats | grep -E 'ejections_enforced_(consecutive_5xx|success_rate|consecutive_gateway_failure|failure_percentage|local_origin_success_rate)'

# Compare detected vs enforced (detected fires even when capped by max_ejection_percent)
curl -s http://localhost:9901/stats | grep -E 'ejections_detected_|ejections_enforced_total'

# See the membership picture
curl -s http://localhost:9901/stats | grep -E 'membership_(healthy|degraded|excluded|total)'

# Per-host health status to spot correlated vs independent failures
curl -s http://localhost:9901/clusters?format=json | jq '.cluster_statuses[].host_statuses[].health_status'

# Latency on the survivors - rising P99 confirms load concentration
curl -s http://localhost:9901/stats/prometheus | grep 'envoy_cluster_upstream_rq_time'

# Confirm outlier detection config (consecutive_5xx, success_rate settings, max_ejection_percent)
curl -s http://localhost:9901/config_dump | jq '[.. | objects | select(has("outlier_detection")) | .outlier_detection]'

All of these are read-only. None of them touch the data plane.

How to diagnose it

  1. Confirm it is actually a mass ejection, not an EDS membership change. If membership_total is stable while membership_healthy drops and ejections_active rises in lockstep, outlier detection is the cause. If membership_total is also dropping, the control plane is removing endpoints, which is a different incident.

  2. Quantify the ejection rate. Sample ejections_active twice, ten seconds apart. A flat value means the cascade has saturated at the cap. A rising value means hosts are still being ejected. ejections_overflow increasing confirms Envoy wanted to eject more hosts than max_ejection_percent allows, which is a strong signal the cluster is in worse shape than the cap can express.

  3. Identify the detector type. The ejections_enforced_* counters tell you which rule is firing. A dominant ejections_enforced_consecutive_5xx points to error-burst sensitivity. A dominant ejections_enforced_success_rate points to slow-host cascades where survivors degrade under concentrated load. ejections_enforced_failure_percentage fires on aggregate failure-rate thresholds.

  4. Separate detected from enforced. ejections_detected_* counters increment whenever the detector trips, even if the ejection was suppressed by max_ejection_percent. A large gap between ejections_detected_total and ejections_enforced_total means many hosts are failing the detector but only some are being ejected. This is the cap working as intended, but it also means the upstream is systemically degraded.

  5. Distinguish systemic upstream failure from threshold sensitivity. Pull per-host health from /clusters?format=json. If all hosts are failing independently with similar error patterns, the upstream has a real problem: shared database, AZ network, config deploy. If only a few hosts are failing and the cascade is driven by load concentration on survivors, the thresholds are too sensitive for the cluster size.

  6. Verify observability is intact. In mesh environments with a custom stats matcher, the ejections_active counter may be filtered out. Without it, you cannot observe the cascade through stats even if the cap is still enforcing internally. Confirm ejections_active is exposed and non-zero. Stats matcher exclusion affects observability only: the cap is enforced from internal ejection state regardless of whether the stat is created.

  7. Check whether panic mode is already active. If membership_healthy / membership_total is below 50% (the default panic threshold), Envoy is routing to all hosts including ejected ones. The elevated error rate you are seeing is panic-mode behavior, not an additional failure. Do not treat it as a new incident.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
outlier_detection.ejections_activeGauge of currently ejected hosts; the operational view of the cascadeTrending toward membership_total, or sustained above 50% of membership
outlier_detection.ejections_overflowCounter of ejections aborted by max_ejection_percentAny sustained increase: the cluster is in worse shape than the cap can address
ejections_enforced_* per typeTells you which detector rule is driving the cascadeOne type dominating points at the root cause class
ejections_detected_* vs ejections_enforced_totalDetected fires even when capped; gap reveals suppressed demandLarge and growing gap means systemic upstream degradation
membership_healthy / membership_totalRatio crossing panic threshold changes routing behavior entirelyRatio below 50% means Envoy is routing to known-bad hosts
upstream_rq_time P99Load concentration shows up as latency on survivors before it shows up as errorsP99 rising while ejections_active rises is the cascade signature
upstream_rq_5xx and response flagsDistinguishes upstream-originated errors from Envoy-generated 503sUF and UO flags during an ejection cascade need different handling

Fixes

Too-sensitive thresholds

The most common cause and the easiest to fix. consecutive_5xx (default 5) is aggressive for some workloads, and success_rate ejection on small clusters is fragile because the sample size is too small to distinguish noise from real degradation.

  • Raise consecutive_5xx if deploys, GC pauses, or brief upstream restarts are triggering ejections.
  • Raise success_rate_minimum_hosts (default 5) so the detector has enough samples. On a five-host cluster, success-rate ejection is at the statistical boundary.
  • Review split_external_local_origin_errors. If local-origin errors (connect failures, resets) are lumped with upstream 5xx, a network blip can look like an upstream outage.

Changes to outlier detection config are applied via xDS. Verify with update_success and check update_rejected to confirm Envoy accepted the new config.

Genuine systemic upstream degradation

If all hosts are failing independently, disabling outlier detection is the wrong move. The ejections are correct; the upstream is broken. Address the root cause: shared dependency, AZ network issue, bad deploy.

During the incident, consider whether outlier detection is making things worse. If the upstream is slowly recovering and outlier detection keeps ejecting hosts that are nearly healthy again, the cascade extends the outage. Temporarily lowering max_ejection_percent can help: a tighter cap forces Envoy into panic mode sooner, which spreads load across all hosts including the recovering ones. This trades targeted ejection for broad degradation, which is preferable when no host is genuinely healthy.

max_ejection_percent misconfiguration

Default is 10%. Setting it to 100% lets outlier detection empty the cluster. Setting it too low forces panic mode on minor degradation. The right value depends on cluster size: on a three-host cluster, 10% rounds to zero ejections allowed, while on a fifty-host cluster, 10% allows five. Tune to your operational reality.

The always_eject_one_host option (added in 1.31.0) overrides max_ejection_percent to guarantee at least one bad host is ejected even in small clusters. Enable it deliberately; it can defeat the cap on tiny clusters.

Stats matcher hiding the ejection gauge

If a stats matcher filters out ejections_active, you lose visibility into the cascade. Confirm whether the cap is still enforcing internally by comparing ejections_enforced_total against cluster membership and the configured max_ejection_percent. Adjust the stats matcher to include ejections_active so the gauge is exported.

Premature unejection

successful_active_health_check_uneject_host defaults to true: a single successful active health check unejects a host that outlier detection ejected. During a marginal-host cascade, this causes flapping. The host is unejected, receives load, fails again, and is re-ejected. Setting it to false (the config field was added in 1.28.0, replacing the runtime option) forces the host to serve its full ejection time before re-admission, which stabilizes flapping at the cost of slower recovery.

Prevention

  • Do not treat outlier detection as a replacement for active health checks. Outlier detection only acts on hosts that are receiving traffic. Active health checks catch newly-failed hosts that have not yet been routed to. Use both.
  • Tune thresholds to cluster size. success_rate ejection on a cluster below 5 hosts does not fire by default; at exactly 5 it is at the statistical boundary. Either raise success_rate_minimum_hosts or disable success-rate ejection for small clusters.
  • Alert on ejections_overflow, not just ejections_active. ejections_overflow means the cluster is in worse shape than the cap can express. It is the leading indicator that a cascade is being artificially limited.
  • Account for panic threshold behavior in your dashboards. When membership_healthy / membership_total crosses 50%, the meaning of “healthy” changes. Annotate this transition or your incident timeline will be misleading.
  • Verify ejections_active is exposed in mesh environments. A stats matcher that filters it out makes the cascade invisible through stats.
  • Set max_ejection_percent deliberately. The default is reasonable for most clusters, but small clusters and large clusters need different values. Document the reasoning.

How Netdata helps

  • Per-second collection of outlier_detection.ejections_active and ejections_overflow makes the cascade visible in real time, including the moment ejections_overflow starts climbing. Alerting on sustained ejections_overflow gives warning that the cap is being hit before the cluster enters panic mode.
  • Correlating ejections_active with upstream_rq_time P99 and membership_healthy in a single view confirms whether load concentration is driving the cascade. The ejections_enforced_* breakdown by detector type is collected separately, so you can see immediately whether consecutive_5xx or success_rate is the dominant cause.
  • Per-cluster dashboards let you compare ejection behavior across clusters of different sizes, which is essential for tuning thresholds.