The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-httpd / apache-httpd-slow-backend-cascade ▌

Operations Guides

Apache slow backend cascade: how one slow upstream starves the whole worker pool

Users report the site down. The load balancer has pulled the node from rotation. You SSH in expecting to find Apache melting, and instead find a perfectly calm process: normal CPU, normal memory, no crash dumps in the error log. The port is listening, but new connections hang or get refused.

This is the slow backend cascade: the classic “Apache outage” in reverse-proxy deployments. Apache is not broken. Every worker is blocked waiting on a backend that has gone slow, and the frontend has run out of execution slots as a consequence.

The failure is deceptive because every conventional host metric looks healthy. Workers are waiting, not working. CPU is idle. Memory is flat. The incident is only visible in the scoreboard, the backend latency, and eventually the listen backlog.

What this means

When Apache proxies a request, the worker that accepted it stays occupied for the entire round trip: accept the client connection, forward the request, wait for the backend response, relay it back. Worker occupancy time per proxied request is therefore roughly the backend response time plus transfer time.

Now apply Little’s law. If you serve 100 proxied requests per second and the backend answers in 200 ms, you need about 20 workers. If the backend degrades to 5 seconds per response, you need 500 workers for the same traffic. If MaxRequestWorkers is below that, the excess does not slow down gracefully. It queues in the kernel listen backlog, and once the backlog fills, new connections are refused. The degradation curve is cliff-edge: normal service right up to saturation, then instant denial.

The cascade:

flowchart TD
  A[Backend goes slow: DB locks, GC pause, dependency down] --> B[Each proxied request holds a worker longer]
  B --> C[Workers pile into W state, IdleWorkers falls to zero]
  C --> D[New connections queue in listen backlog]
  D --> E[Backlog fills, connections refused]
  E --> F[LB health checks time out, node pulled from rotation]
  C --> G[Apache CPU and memory stay normal: workers waiting, not working]

Two properties make this pattern easy to misread:

  • Direct, non-proxied requests may still work. A static file or a locally served health page can return instantly if a worker is free, while every proxied path hangs. “Apache responds to curl” does not rule this out.
  • The request rate paradoxically drops. Total Accesses counts completed requests. When workers are stuck, completions fall even though client demand is unchanged or rising.

Common causes

The root cause is always on the backend side. Apache is the victim, not the perpetrator.

CauseWhat it looks likeFirst thing to check
Backend database lock contention or slow queriesBackend response time 5x or more above baseline, gradual onsetBackend health endpoint directly, bypassing Apache
Backend memory exhaustion or GC pauseLatency spikes in bursts, backend recovers then degrades againBackend’s own memory and GC metrics
Network partition or packet loss between Apache and backendConnection timeouts rather than slow responses, 502s alongside 504sDirect curl to backend from the Apache host
Backend’s own external dependency failureBackend up and accepting connections but slow on specific pathsWhich proxied URL patterns are slow versus fast
Proxy connection pool exhaustion on top of slow backend503s appearing before MaxRequestWorkers is reachedError log for proxy errors, balancer-manager status

Quick checks

Run these from the Apache host. All are read-only.

# 1. Scoreboard state distribution: the signature of this incident
curl -s http://localhost/server-status?auto | grep "Scoreboard:" | \
  awk '{print $2}' | fold -w1 | sort | uniq -c | sort -nr

# 2. Worker utilization
curl -s http://localhost/server-status?auto | grep -E "BusyWorkers|IdleWorkers"

# 3. Listen backlog depth: Recv-Q is current queue, Send-Q is the max
ss -ltn | grep -E ':80\s|:443\s'

# 4. Listen overflow counter: connections already dropped
nstat -a | grep ListenOverflows 2>/dev/null || netstat -s | grep -i "listen\|overflow"

# 5. Proxy and worker-exhaustion errors
grep -E "AH01114|AH00484" /var/log/apache2/error.log | tail -20
# RHEL path: /var/log/httpd/error_log

# 6. 5xx breakdown from recent traffic
tail -1000 /var/log/apache2/access.log | \
  awk '$9 ~ /^5/ {c[$9]++} END {for (k in c) print k, c[k]}'

# 7. Backend health directly, bypassing Apache entirely
curl -s -o /dev/null -w "TTFB: %{time_starttransfer}s Total: %{time_total}s HTTP: %{http_code}\n" \
  --max-time 10 http://backend-host:backend-port/health

# 8. Does Apache itself still work? Request something local and non-proxied
curl -s -o /dev/null -w "%{http_code} %{time_total}s\n" --max-time 5 http://localhost/server-status?auto

What you are looking for: a scoreboard dominated by W with IdleWorkers at or near zero, a growing Recv-Q, 504s (and then 503s) in the access log, and a backend that is slow or unreachable when queried directly.

How to diagnose it

  1. Read the scoreboard first. Count the state distribution. The cascade signature is W states climbing toward MaxRequestWorkers. Note that W is ambiguous by design: a worker in W could be writing to the client, waiting on the backend, or doing internal processing. Disambiguate with the steps below.

  2. Confirm Apache is not resource-bound. Check CPU and memory. In this pattern both are normal because workers are blocked on I/O, not computing. High CPU or climbing RSS points to a different failure pattern (see the related guides).

  3. Check the listen backlog. Non-zero and growing Recv-Q on the listening socket confirms new connections are arriving faster than workers free up. A rising ListenOverflows counter means connections are already being dropped.

  4. Separate 504 from 503. 504 means a backend connected but did not respond within ProxyTimeout: the backend is slow or hung. 503 means the proxy path had nothing usable: proxy pool exhausted, a balancer member in error state, or all balancer workers busy. Worker saturation itself queues excess connections in the listen backlog rather than answering 503. 504s appearing first, then 503s as the pool drains, is the textbook progression.

  5. Test the backend directly from the Apache host. This splits the problem in half. If the backend is slow when hit directly, Apache is exonerated and the incident moves to the backend. If the backend is fast directly but slow through Apache, suspect the network path, the proxy connection pool, or DNS resolution.

  6. Test a non-proxied path on Apache. If /server-status or a static file answers quickly while proxied paths hang, the diagnosis is confirmed: the frontend is healthy and starved by the upstream.

  7. Check the error log for corroboration. Proxy connection failures (AH01114) and AH00484: server reached MaxRequestWorkers setting tell you how far the cascade has progressed. AH00484 is the definitive confirmation that the worker pool is fully saturated.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Scoreboard W state countDirect view of workers blocked mid-requestW above 50% of workers sustained
IdleWorkersThe last buffer before queuing startsFalling toward zero at stable traffic
Backend response timeThe root cause, and the earliest signalP95 above 2x baseline; 5x is active cascade
504 rateBackend responses exceeding ProxyTimeoutAny sustained occurrence
503 rateProxy pool exhaustion or balancer member errorAny occurrence in production
Listen backlog Recv-QConnections queuing before reaching a workerSustained non-zero; approaching ListenBacklog (default 511) is critical
ListenOverflows counterConnections already refusedAny increase
Request completion rate (Total Accesses delta)Falls paradoxically during the cascadeDrop while client demand is unchanged
AH00484 in error logApache explicitly reporting pool saturationAny occurrence
Apache CPU and RSSThe negative signal that distinguishes this patternNormal resources plus all of the above

The correlation that closes the diagnosis in one glance: high W count, elevated backend response time, zero idle workers, and normal Apache CPU and memory. No other failure pattern produces that combination.

Fixes

Immediate mitigation

Check the backend and act on it, not on Apache. Take the node out of LB rotation if health checks are flapping, then work the backend problem. Restarting Apache buys minutes at best: the new workers immediately re-block on the same slow backend. Do not make Apache restarts your first move.

Fail fast with a lower ProxyTimeout. If the backend is slow but not dead, temporarily reducing ProxyTimeout makes workers give up sooner and return an error instead of holding the slot for the full timeout. This converts “hang until the pool drains” into “fast 504s that clients can retry.” It is a pressure valve, not a cure: some legitimate slow requests will now fail, and the change requires a graceful reload to take effect.

If the backend is dead, say so quickly. A fast 503 is better for clients behind a retrying load balancer than a 60-second hang.

Backend-side resolution

The actual fix lives wherever the backend is: kill the locking query, resolve the GC pressure, restore the failed dependency, fix the network path. Until backend response time returns to baseline, every Apache-side change is symptom management.

Proxy pool sizing

The proxy connection pool is per child process, and its default max equals ThreadsPerChild (1 for prefork). That default is too small for most production workloads, and it produces 503s under moderate load that look like worker exhaustion but are not. Size the pool for expected concurrency, and remember the headroom rule from the playbook: pool_utilization = request_rate x backend_avg_response_time / pool_size. When backend latency doubles, pool utilization doubles. Aim for 2x expected concurrent proxied requests per child at peak.

Prevention

  • Monitor the backend as a first-class signal. Backend response time is the leading indicator for this entire failure class. Alert on backend P95 exceeding 2x baseline before workers start piling up. The playbook’s capacity model applies directly: if peak BusyWorkers trends track backend latency, fixing the backend buys more headroom than raising MaxRequestWorkers.
  • Set a deliberate ProxyTimeout. The default inherits from Timeout. Pick a value tied to your backend’s actual SLA so a hung upstream fails in seconds, not a minute.
  • Health-check the critical path, not a static file. A health check that only exercises local content will keep reporting green while every proxied path is down. Probe through the proxy to the backend.
  • Watch the scoreboard state distribution continuously. A rising W fraction at flat request rate is the earliest frontend-side warning, minutes before the backlog fills.
  • Know your MPM. This cascade hits all MPMs, but interpretation differs: on event MPM, keepalive connections are handled by the listener thread and tracked via ConnsAsyncKeepAlive, so proxied-request blocking shows up cleanly as W states rather than being mixed with K noise.
  • Keep idle headroom. At least 25% of MaxRequestWorkers idle at peak. The degradation curve is cliff-edge; there is no graceful degradation zone to catch you.

How Netdata helps

  • Netdata’s Apache collector polls server-status and charts BusyWorkers, IdleWorkers, and the full scoreboard state breakdown per second, so the W-state pile-up is visible as it builds rather than after the backlog overflows.
  • Correlating worker utilization against request completion rate exposes the signature paradox of this incident: workers maxed while completed requests fall.
  • Backend response time and 502/503/504 rates from log or endpoint monitoring sit on the same dashboard as Apache’s own CPU and memory, making the “healthy Apache, dying upstream” contrast obvious in one view.
  • Listen backlog depth and TCP listen overflow counters from the host are collected alongside Apache metrics, closing the loop from backend latency to worker saturation to refused connections.
  • Anomaly detection on backend latency and scoreboard states catches the slow drift phase (backend degrading over minutes) before it becomes the cliff-edge phase.

Netdata’s Apache HTTP Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Apache HTTP Server monitoring with Netdata

Netdata monitors Apache HTTP Server with per-second metrics from mod_status, pre-built dashboards, and ML-powered anomaly detection. Watch busy versus idle workers and the scoreboard state mix, requests per second, bytes served per second, and request processing duration alongside the rest of your stack, so you catch the worker-exhaustion, slow-backend, and memory incidents in these runbooks before they page anyone.