The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-httpd / apache-httpd-503-service-unavailable ▌

Operations Guides

Apache 503 Service Unavailable: worker exhaustion versus proxy pool exhaustion

Your access log is filling with 503s and users are reporting the site is down. The status code tells you almost nothing: Apache returns 503 for two root causes that look identical in the access log but need completely different fixes.

The first is frontend worker exhaustion. Every worker slot in the scoreboard is busy, Apache has hit MaxRequestWorkers, and new connections queue in the kernel backlog until they time out. The second is a mod_proxy failure: a backend is down, marked errored, or the per-child proxy connection pool is too small, so Apache refuses to forward requests even though its own workers are mostly idle.

The access log cannot distinguish these. The error log can. Triage order: error log first, then the scoreboard, then balancer state. Fix the pool that is actually exhausted.

What this means

A 503 from Apache means “I cannot service this request right now,” but the reason lives in one of two layers:

  1. Frontend layer (MPM workers). All worker slots are occupied. The exact worker-limit code varies by MPM: AH00484 (event), AH00286 (worker), or AH00161 (prefork), followed by server reached MaxRequestWorkers setting; Apache logs the condition once per generation. New connections pile into the listen backlog (default ListenBacklog is 511, capped by net.core.somaxconn) and eventually get refused or time out.

  2. Proxy layer (mod_proxy / mod_proxy_balancer). Apache has free workers but cannot get a usable backend connection. Causes: the backend is down and the proxy worker is in error state (default retry=60 means Apache will not retry that backend for 60 seconds after a failure), all balancer members are errored, or the proxy connection pool (max parameter) is exhausted. The error log shows AH00959: ap_proxy_connect_backend disabling worker for (hostname) for 60s and AH01114: HTTP: failed to make connection to backend.

The dangerous misdiagnosis is conflating them: seeing 503s, assuming MaxRequestWorkers is too low, raising it, and making things worse. If the real bottleneck is the proxy pool or a dead backend, more frontend workers just means more workers blocked waiting on the same dead backend, plus more memory consumed.

flowchart TD
  A[503s in access log] --> B{Error log: MPM worker-limit message present?}
  B -- Yes --> C[Frontend worker exhaustion]
  B -- No --> D{Error log: AH00959 / AH01114 / proxy errors?}
  D -- Yes --> E[Backend down or in error state]
  D -- No --> F{Scoreboard idle workers?}
  C --> G[Check scoreboard state mix: W vs R vs K]
  E --> H[Check balancer-manager and retry state]
  F -- "IdleWorkers = 0" --> C
  F -- "IdleWorkers > 0" --> I[Proxy pool too small: tune max on ProxyPass/BalancerMember]

Common causes

CauseWhat it looks likeFirst thing to check
Traffic exceeds MaxRequestWorkersMPM worker-limit message in error log, scoreboard mostly W, IdleWorkers 0BusyWorkers vs configured MaxRequestWorkers
Slow backend holding workersScoreboard filling with W, backend latency elevated, 504s before 503sCurl the backend directly, bypassing Apache
Backend down, proxy worker in error state503s for up to 60s after failure, AH00959 and AH01114 in error logretry state in balancer-manager
Proxy pool too small (default max)503s under moderate load, IdleWorkers > 0, no worker-limit messagemax on ProxyPass/BalancerMember vs concurrent proxied requests
Stuck graceful restartsMany G states, scoreboard full but below MaxRequestWorkersRestart frequency in error log
Slow clients / SlowlorisMany R states, low throughput relative to connection countScoreboard R count, source IP concentration
Keepalive hoarding (prefork/worker MPM)Many K states consuming slotsMPM in use, KeepAliveTimeout
Log disk full (workers stuck in L)Scoreboard dominated by L, throughput near zerodf -h on log filesystem

Quick checks

All read-only and safe to run during an incident. Paths shown for Debian/Ubuntu; on RHEL use /var/log/httpd/error_log and httpd in place of apache2.

# 1. The single most decisive check: which 503 is this?
grep -E "AH00484|AH00286|AH00161|AH00959|AH01114" /var/log/apache2/error.log | tail -20

# 2. Worker utilization right now
curl -s http://localhost/server-status?auto | grep -E "BusyWorkers|IdleWorkers"

# 3. Scoreboard state mix: where are workers stuck?
curl -s http://localhost/server-status?auto | grep "Scoreboard:" | \
  awk '{print $2}' | fold -w1 | sort | uniq -c | sort -nr

# 4. Is the listen backlog filling? (Recv-Q on LISTEN sockets)
ss -ltn | grep -E ':80\s|:443\s'

# 5. Backend health, bypassing Apache entirely
curl -s -o /dev/null -w "backend: %{http_code} in %{time_total}s\n" \
  --max-time 5 http://backend-host:port/health

# 6. Established connections from Apache to the backend (adjust port)
ss -tn state established dport = :8080 | wc -l

# 7. Balancer member state, if balancer-manager is enabled
curl -s http://localhost/balancer-manager 2>/dev/null | grep -E 'Worker|Status'

# 8. 5xx breakdown from the access log (status is field 9 in combined format)
tail -1000 /var/log/apache2/access.log | \
  awk '$9 ~ /^5/ {c[$9]++} END {for (s in c) print s, c[s]}'

Note on checks 2 and 3: /server-status requires mod_status and should be IP-restricted. ExtendedStatus On has been the default since 2.3.6 when mod_status is loaded, so the scoreboard line is normally present.

How to diagnose it

  1. Read the error log first. AH00484, AH00286, or AH00161 with server reached MaxRequestWorkers setting means the frontend pool is the problem. AH00959 ... disabling worker for (hostname:port) for 60s or AH01114: HTTP: failed to make connection to backend means proxy/backend. If you see neither, check whether the 503s come from an ErrorDocument or application handler instead of Apache itself.

  2. Snapshot the scoreboard. IdleWorkers = 0 with BusyWorkers at or near configured MaxRequestWorkers confirms frontend saturation. IdleWorkers > 0 while 503s continue means Apache has capacity and the failure is downstream.

  3. If frontend exhaustion, look at the state mix. The distribution tells you what is holding workers:

    • Mostly W: workers sending replies or, in proxy mode, blocked waiting for backends. Correlate with backend latency.
    • Mostly R: slow request bodies, slow clients, or Slowloris. Normal traffic rarely exceeds 5% in R.
    • Mostly K on prefork or worker MPM: keepalive connections holding workers hostage; KeepAliveTimeout too long. On event MPM, significant K in the scoreboard is abnormal because keepalive is handled by the listener thread (ConnsAsyncKeepAlive).
    • Many G: graceful restart pile-up; old generations lingering.
    • Many L: log disk full or log pipe stall. Check df -h.
  4. If proxy failure, check the backend directly. Curl the backend from the Apache host. If it is down or slow, that is your incident; Apache is a victim, not the cause. With the default retry=60, the proxy worker stays in error state for 60 seconds after a failure, so 503s persist for up to a minute after the backend recovers. Do not restart Apache over this; it fixes itself when the retry window expires.

  5. If the backend is healthy and IdleWorkers > 0, check pool sizing. The default max for a proxy worker equals ThreadsPerChild for the active MPM, and is 1 for prefork. Pools are per-child-process: total backend connections = max x number of children. Under moderate concurrency the default exhausts quickly and Apache returns 503 with completely idle frontend workers. This is the most misdiagnosed cause of Apache 503s.

  6. Check the listen backlog. Sustained Recv-Q > 0 with high worker utilization means connections are queuing at the kernel level. Brief spikes during bursts are normal; sustained growth is the cliff edge before connection refused.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
MPM worker-limit events (AH00484/AH00286/AH00161) in error logDefinitive confirmation of MaxRequestWorkers saturationAny occurrence
BusyWorkers / MaxRequestWorkersPrimary frontend saturation indicator, cliff-edge behaviorSustained >80%, IdleWorkers = 0
Scoreboard state distributionTells you what is holding workers: W, R, K, G, or L>50% in non-idle states; R >20%
AH00959 / AH01114 proxy errorsBackend connection failures and error-state transitionsAny sustained rate
503 rate split by causeAccess log cannot split; error log canAny 503s in production
Backend response time (direct)Separates “backend down” from “backend slow”P95 >2x baseline
Established Apache-to-backend connectionsApproximates proxy pool utilization per childApproaching max x children
Listen backlog Recv-QLeading indicator before user-visible failuresSustained >0, approaching 511
Balancer member statusWhich backends are errored or disabledAny member errored with live traffic

Fixes

Frontend worker exhaustion

Raise MaxRequestWorkers, but only with the memory math done. MaxRequestWorkers x per-child RSS must stay under roughly 70% of RAM. For prefork with mod_php, measure real per-child RSS first (ps -C apache2 -o rss --no-headers) because 50MB+ per child is common. For threaded MPMs, ServerLimit must be >= MaxRequestWorkers / ThreadsPerChild, and ServerLimit defaults to 16 for worker and event (256 for prefork) if unset, which silently caps how high MaxRequestWorkers can go. Raising ServerLimit requires a full restart, not a graceful reload. MaxRequestWorkers was renamed from MaxClients in 2.3.13/2.4; the old name still works as a deprecated alias.

Fix what is holding workers instead. If the scoreboard is full of W and the backend is slow, raising the limit only buys a bigger queue of blocked workers. Reduce ProxyTimeout temporarily to fail fast, and fix the backend. If it is full of R, tighten mod_reqtimeout: RequestReadTimeout header=20-40,MinRate=500 body=20,MinRate=500. If it is full of K on prefork/worker, lower KeepAliveTimeout or migrate to event MPM. If it is full of G, reduce graceful-restart frequency; GracefulShutdownTimeout applies to graceful-stop, not a graceful reload. Bound request and keepalive time with Timeout, KeepAliveTimeout, and ProxyTimeout, and investigate any stuck request.

Backend down or in error state

Fix the backend, then wait out the retry window. With the default retry=60, Apache will not retry an errored backend for 60 seconds. You can lower retry on the ProxyPass or BalancerMember (for example retry=5) so recovered backends come back into rotation faster, at the cost of probing a flapping backend more often. If all members of a balancer are in error state, forcerecovery=On (the default since Apache 2.4.2) instructs the balancer to recover all workers immediately, without considering each worker’s retry timeout. Apache’s documentation warns that this can deepen an already overloaded backend’s trouble; in that case set it Off.

Stock mod_proxy only learns a backend is dead by failing a real user request. Without active health checks, the first requests after a backend dies always eat the failure. Apache 2.4.21 and later include mod_proxy_hcheck; if it is not already enabled, verify with apachectl -M | grep proxy_hcheck before configuring active checks.

Proxy pool too small

Size max to your real concurrency. Set max on ProxyPass or BalancerMember to roughly 2x the expected concurrent proxied requests per child at peak. Pools are per-child: on prefork, N children x max connections hit the backend, which can overwhelm it; on worker/event, fewer children means max must be larger to reach the same total. Estimate with pool_utilization = request_rate x backend_avg_response_time / pool_size and keep utilization well under 1.

Enable backend keepalive. keepalive=On on ProxyPass reuses backend connections instead of opening a new TCP connection per request, reducing latency and connection churn on both sides.

Prevention

  • Derive MaxRequestWorkers from memory, never from a guess: available_RAM x 0.7 / measured_per_child_RSS. Re-measure after module or application changes.
  • Set MaxConnectionsPerChild to a finite value (5000-10000) so leaky modules cannot grow children without bound. The default of 0 is wrong for mod_php and mod_perl deployments.
  • Size proxy pools explicitly. Never run production reverse proxies on the default max. Document the per-child math next to the directive.
  • Alert on MPM worker-limit messages (AH00484/AH00286/AH00161) and proxy error patterns, not just on 5xx rate. The error log is the only place the two 503 causes are distinguishable automatically.
  • Sample the scoreboard continuously. Point-in-time snapshots during an incident are too late; the state distribution trend tells you whether workers are filling with W, R, or K before the cliff.
  • Keep logs on their own filesystem so a log explosion cannot starve the OS or trigger the log-stall variant of worker exhaustion.
  • Test health checks on the critical path. A health check that only fetches a static file will pass while every proxied request 503s.

How Netdata helps

  • Scoreboard state distribution over time. Netdata collects the Apache scoreboard continuously, so you can see the W/R/K mix building toward exhaustion minutes before IdleWorkers hits zero, instead of discovering it from 503s.
  • BusyWorkers and IdleWorkers trends. Worker utilization graphed against configured MaxRequestWorkers makes the cliff edge visible and gives you the runway estimate for capacity planning.
  • 5xx correlation with request rate. A paradoxical drop in completed requests alongside rising 503s and full workers is the signature of the slow-backend cascade; seeing all three on one dashboard shortens the “is it Apache or the backend” question to seconds.
  • Backend latency alongside worker states. Correlating direct backend response time with W-state accumulation confirms or rules out the proxy cause without log spelunking during an incident.
  • Error log pattern alerting. Alerting on MPM worker-limit and proxy error messages as distinct conditions routes the page to the right fix: frontend capacity versus backend health.

Netdata’s Apache HTTP Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Apache HTTP Server monitoring with Netdata

Netdata monitors Apache HTTP Server with per-second metrics from mod_status, pre-built dashboards, and ML-powered anomaly detection. Watch busy versus idle workers and the scoreboard state mix, requests per second, bytes served per second, and request processing duration alongside the rest of your stack, so you catch the worker-exhaustion, slow-backend, and memory incidents in these runbooks before they page anyone.