The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-httpd / apache-httpd-balancer-member-error ▌

Operations Guides

Apache balancer member in error state: reading balancer-manager and failover

You opened balancer-manager, or someone pasted a screenshot into the incident channel, and one of your BalancerMembers shows Err in the status column. Traffic is still flowing, but the pool is quietly running on fewer backends than you think. If enough members flip to error state, Apache stops proxying entirely and returns 503s even though httpd itself is healthy.

This state is mod_proxy_balancer doing its job: it detected failures against a backend and pulled that member out of rotation so requests stop dying on it. The problem is that the balancer tells you almost nothing about why, and it keeps the member sidelined on its own retry schedule regardless of whether the backend has recovered.

This article covers how to read the error state in balancer-manager, the directives that govern when a member is marked bad and when it comes back, how failover to spares and standbys behaves, and how to find the backend-side root cause.

What this means

When mod_proxy cannot establish a connection to a backend, or when a configured failure condition is met, the balancer marks that worker as being in error state and stops routing requests to it. In the error log you will see a line like:

AH00959: ap_proxy_connect_backend disabling worker for (backend-host:port) for 60s

That 60s is the member’s retry interval. The default retry is 60 seconds: after a worker enters error state, the balancer will not attempt to use it again until the interval elapses, at which point it tries the worker on a new request. If the backend is still broken, the worker goes straight back to error state for another interval.

In balancer-manager, worker status is shown as a short list of flags: the manager UI renders them as words (Err, Dis, Stby, Spar, Drn, Ign, Stop, HcFl), and the same states map to the single-letter flags used in status= parameters. The ones that matter for this incident:

FlagMeaning
(none)Ok, worker is in rotation
E (Err)Error, worker failed and is sidelined until retry
D (Dis)Disabled, administratively removed from rotation
S (Stop)Stopped
N (Drn)Drain, finish existing sticky sessions, take no new ones
R (Spar)Hot spare, used as a drop-in replacement for an unusable worker in the same set
H (Stby)Hot standby, used only when all workers and spares in the set are unavailable
I (Ign)Ignore-errors, always treated as available
C (HcFl)Failed a dynamic health check (mod_proxy_hcheck)

The state machine for a normal member:

stateDiagram-v2
  Ok --> Error: connection failure or failonstatus
  Error --> Ok: retry interval elapsed, backend accepts request
  Error --> Error: retry attempt fails, interval restarts
  Ok --> Drain: set to +N via balancer-manager
  Error --> HotSpareTakesOver: spare (R) in same lbset

Two properties of this design matter during an incident. First, a member in error state with active traffic is a silent capacity reduction: nothing pages, but your pool just lost a fraction of its throughput. Second, if every member of a balancer ends up in error state, the balancer has nowhere to send requests and Apache returns 503 for every proxied request. In proxy troubleshooting, 503 from an otherwise healthy Apache almost always means the pool is exhausted or all members are errored.

Common causes

CauseWhat it looks likeFirst thing to check
Backend process down or crashedErr on the member, AH00959 in the error log, connection refused when you curl the backend directlycurl -sv http://backend:port/health from the Apache host
Network problem between Apache and backendAH00959 with timeout rather than refusal, intermittent error stateDirect curl timing; check for packet loss or conntrack saturation (dmesg for nf_conntrack: table full)
Backend slow enough to trip timeoutsMember flaps in and out of error state; 504s in the access log; scoreboard filling with W statesCompare %D on proxied requests against a direct backend curl
failonstatus configured and backend returns matching codesMember goes to error state even though the backend is up and answering; error log shows status-based failure, not connect failureCheck the BalancerMember config for failonstatus; check the backend’s recent 5xx
Firewall or SELinux blocking proxy connectionsConnection refused or permission denied from Apache, but backend healthy from other hostsaudit.log or journalctl for SELinux denials
Flapping backend (crash loop, GC storms, deploys)Member oscillates between Ok and Err; users see intermittent 502/503Correlate AH00959 timestamps with backend restarts and deploys
bybusyness load-balancing methodRecovered member stays in error state, or returns to Ok but receives no traffic, until Apache restartsCheck lbmethod on the ProxySet

Quick checks

All read-only. Run them from the Apache host.

# 1. See current member status in balancer-manager (if enabled and reachable)
curl -s http://localhost/balancer-manager | grep -E 'Worker|Status|Err'

# 2. Find the moments members were disabled, and the retry interval in effect
grep "AH00959" /var/log/apache2/error.log | tail -20
# RHEL path: /var/log/httpd/error_log

# 3. Test each backend directly, bypassing Apache
curl -sv --max-time 5 http://backend-host:port/health -o /dev/null

# 4. Check proxy error rates in the access log (502 = bad response/refused,
#    503 = pool exhausted or all members errored, 504 = backend timeout)
tail -5000 /var/log/apache2/access.log | awk '$9 ~ /^50[234]$/ {print $9}' | sort | uniq -c

# 5. Check whether Apache workers are piling up waiting on backends
curl -s http://localhost/server-status?auto | grep -E "BusyWorkers|IdleWorkers|Scoreboard"

# 6. Look for proxy connect and read failures
grep -E "AH01114|AH00898" /var/log/apache2/error.log | tail -10

# 7. Rule out SELinux on RHEL-family systems
grep -i "denied.*httpd" /var/log/audit/audit.log 2>/dev/null | tail -5

Note on check 1: balancer-manager only shows balancers defined outside <Location> containers. If your balancer is defined inside a <Location> block, it will not appear or cannot be controlled dynamically, and you will have to read state from the error log instead. If your balancer-manager page looks empty or incomplete, verify that first.

How to diagnose it

  1. Confirm which members are affected and since when. Use balancer-manager or grep AH00959 from the error log. The AH00959 line names the worker and prints the retry interval, so you know exactly when the balancer will try it again.

  2. Determine the failure type: refused, timeout, or status-based. Connection refused means the backend is not listening (process down, wrong port, firewall). Timeout means the path or the backend is slow. Status-based ejection only happens if you configured failonstatus, in which case the backend is answering but returning codes you told Apache to treat as fatal. The remediation is completely different for each.

  3. Test the backend directly from the Apache host. A fast direct curl with a normal response plus a member stuck in Err points at the balancer configuration or a transient that has not yet hit its retry window. A slow or failing direct curl means the backend is the incident; Apache is just the messenger.

  4. Check whether the pool is still serving. Count how many members remain in Ok state versus the pool total, and compare with current request rate. Two members down out of eight during off-peak is a ticket. Two down out of three during peak means the remaining member is about to be the story: watch BusyWorkers, the listen queue (ss -ltn), and 503 counts.

  5. If members flap, correlate with backend events. Pull AH00959 timestamps and line them up against backend restarts, deploys, and GC pauses. A member that re-enters error state within one retry interval of recovering is telling you the backend recovers for a few seconds and then falls over again.

  6. Check for the bybusyness trap. If the balancer uses lbmethod=bybusyness, a worker that recovers may stay in error state, or return to Ok but receive no requests, until Apache is restarted. This does not happen with the default byrequests method. If you are on bybusyness and a recovered member is not taking traffic, that is a known behavior, not a new failure.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Balancer member status (per member)A member in error state with active traffic is silent capacity lossAny member in Err, D, or C during production hours
AH00959 events in the error logExact record of when members are disabled and the retry intervalRepeating for the same member; intervals shorter than your backend’s real recovery time
502/503/504 rate on proxied paths502 = refused or invalid response, 503 = pool exhausted or all members errored, 504 = backend timeoutAny sustained non-zero rate; any 503 at all
Backend response time (direct probe)Distinguishes “backend dead” from “backend slow”P95 above 2x baseline; anything approaching ProxyTimeout
BusyWorkers / IdleWorkersErrored members push load onto fewer backends, which hold workers longerBusyWorkers climbing while the Ok member count drops
Scoreboard W state countWorkers waiting on slow backends show as WW dominating the scoreboard at normal request rate
Listen queue Recv-QIf the remaining members cannot keep up, connections queueSustained non-zero Recv-Q

The key correlation for this incident class: member status changes lead, 5xx rates lag. If you alert only on 503s, you find out after the pool is fully drained. Watching member state directly gives you the interval between “first member errored” and “last member errored”, which is where you can still act.

Fixes

Backend is down or refusing connections

Fix the backend. Apache is behaving correctly. While the backend is down, decide whether the balancer should fail over (spare or standby member) or fail fast (short retry, short timeouts) so clients get a quick error instead of a held worker and a slow 503.

Member flapping because the backend is intermittently slow

A slow backend that occasionally crosses a timeout will ping-pong in and out of error state, and every re-entry costs another retry interval of reduced capacity. Two knobs govern this:

  • retry on the BalancerMember: seconds to wait before retrying an errored worker. Default 60. retry=0 means always retry the worker, which effectively disables the sidelining behavior. That is appropriate when you would rather spread load onto a degraded backend than concentrate it on fewer members, but every failed attempt then costs a connection timeout before failover.
  • failontimeout: when set, an IO read timeout against the backend forces the worker into error state just like a connect failure. It is available in Apache HTTP Server 2.4.5 and later.

If the real problem is backend latency, neither knob fixes it; they only decide how gracefully Apache absorbs it. Reducing ProxyTimeout temporarily makes Apache fail fast instead of holding workers in W state, which protects the rest of the server at the cost of more visible errors.

Backend returns 5xx and you want (or do not want) that to eject the member

failonstatus=500,503 on the BalancerMember tells the balancer to treat those response codes from the backend as failures and put the member into error state. This is useful when a backend is “up” but broken. It is dangerous when the 5xx is request-specific (one bad request, not a sick backend), because one poisoned request type can sideline a healthy member for a full retry interval. If members enter error state while direct health checks pass, look here first.

Pool fully errored: 503s on everything

When all members of a balancer are in error state, forcerecovery (default On) makes Apache try all workers again immediately rather than waiting out the retry intervals, on the theory that trying something beats a guaranteed 503. If you set forcerecovery=Off, a fully errored balancer hard-fails with 503 until retry intervals expire, because the immediate recovery that skips retry timers is what forcerecovery=On enables.

For planned resilience, configure a hot spare (status=+R, drop-in replacement for an unusable worker in the same lbset) or a hot standby (status=+H, activates only when all workers and spares in the set are unavailable). Expect one visible artifact: the first request during the failover transition can return 503 before the standby takes over from the next request onward. That single 503 is normal behavior, not a second failure.

Recovered member not taking traffic on bybusyness

Known issue with lbmethod=bybusyness: a recovered worker can stay in error state, or return to Ok without receiving requests, until httpd is restarted. If you are affected, either restart Apache during a low-traffic window or switch the balancer to the default byrequests method, which does not have this behavior. A restart is disruptive: it drops connections unless you use apachectl graceful, and even graceful restarts briefly overlap old and new children, so plan for the memory overlap.

Prevention

  • Watch member status continuously, not during incidents. A pool that silently loses one of three members has lost a third of its capacity with zero alerts. Alert on any member in Err or C state with active traffic.
  • Size retry to your backend’s real recovery time. If backends typically recover in 5 seconds, a 60-second retry leaves capacity on the floor after every blip. If backends take minutes, a short retry just adds failed connection attempts.
  • Restrict balancer-manager. It must sit behind Require directives (localhost only is the sane default). The interface has an XSS history in older versions, it lets anyone who can reach it disable your backends, and only balancers defined outside <Location> containers are controllable through it anyway.
  • Set ProxyPassInherit Off if you use balancer-manager for dynamic changes. Inheritance of ProxyPass directives into vhosts can cause inconsistent behavior with manager-driven changes.
  • Use mod_proxy_hcheck for active health checks if you want members ejected based on probing rather than live request failures. Failed dynamic health checks show as the C flag and are re-enabled by the checker when the backend recovers.
  • Monitor the backend independently of Apache. “Apache 503” and “backend down” look identical from outside. Direct backend probes separate the two before the incident call starts.
  • Test failover before you need it. Deliberately stop one backend in staging and watch: the AH00959 line, the flag change, the spare takeover, and the single 503 on transition. Any surprise here belongs in staging, not production.

How Netdata helps

  • Per-code error rates on proxied paths: Netdata’s Apache collector and web log parsing split 502, 503, and 504, so you can tell “backend refused” from “pool drained” without grepping access logs mid-incident.
  • Scoreboard state distribution over time: a rising W count alongside a dropping BusyWorkers-to-capacity ratio is the signature of members being sidelined and remaining backends holding workers longer. Point-in-time balancer-manager checks miss the trend.
  • Error log event correlation: AH00959 timestamps aligned against the 5xx timeline show whether member ejections precede or follow user-visible errors, which tells you whether the balancer is protecting you or fighting you.
  • Backend-side metrics on the same dashboard: putting the backend’s own latency and process health next to Apache’s proxy errors collapses the “is it Apache or the backend” question to one glance.
  • Alerts on proxy error rate with per-member context: alerting on any sustained 502/503/504 rate catches the drained pool before it becomes a full 503 outage.

Netdata’s Apache HTTP Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Apache HTTP Server monitoring with Netdata

Netdata monitors Apache HTTP Server with per-second metrics from mod_status, pre-built dashboards, and ML-powered anomaly detection. Watch busy versus idle workers and the scoreboard state mix, requests per second, bytes served per second, and request processing duration alongside the rest of your stack, so you catch the worker-exhaustion, slow-backend, and memory incidents in these runbooks before they page anyone.