The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-httpd / apache-httpd-backend-response-time ▌

Operations Guides

Apache backend response time: telling 'Apache is slow' from 'the backend is slow'

Users report the site is slow. Apache is up, requests eventually complete. Someone says “Apache is slow,” someone else says “the app is slow,” and both are guessing, because from the outside the two are indistinguishable: the client just sees a slow response.

When Apache proxies to a backend, the %D value in your access log bundles three things together: Apache’s own processing, time waiting for the backend, and time transferring the response to the client. A slow backend, a saturated Apache, and a client on a bad connection all produce the same inflated %D. Without separate instrumentation per component, you cannot attribute the latency, and teams routinely burn hours tuning Apache when the database is the problem.

This article is about closing that gap: measuring backend response time separately, reading the scoreboard to see where workers are stuck, and telling “backend down” from “backend slow” from “Apache itself is the bottleneck.”

What this means

A proxied request has three latency contributors, and only one of them is Apache:

  1. Backend time: establishing the backend connection (if no pooled connection exists) and waiting for the response. Usually the dominant term in reverse-proxy deployments.
  2. Apache processing: URI translation, access control, rewrite rules, filter chain (compression, headers). Typically small unless mod_security or complex mod_rewrite is in play.
  3. Client transfer: sending the response body. A 100MB download to a 1Mbps client shows %D of roughly 800 seconds, and none of that is Apache or backend slowness.

The slow-backend cascade is the failure mode to fear. A backend gets slow, each proxied request holds an Apache worker longer, workers accumulate in W state, IdleWorkers drops toward zero, new requests queue in the listen backlog, and you see 504s and then 503s as the pool fully exhausts. Apache is healthy the entire time; it is starved of workers by the backend. This is the most common cause of “Apache outage” in proxy deployments.

The scoreboard cannot resolve the ambiguity on its own: a worker in W state could be writing to the client, waiting on a backend, or doing internal processing. All three show as W. You need backend timing as an independent signal.

Common causes

CauseWhat it looks likeFirst thing to check
Backend application slow (DB locks, GC pause, slow queries)Scoreboard filling with W states, 504s starting, Apache CPU and memory normalcurl the backend directly, bypassing Apache, and time it
Backend down or unreachable502s (connection refused), fast failures rather than slow onesError log for AH01114 and “(111)Connection refused”
Network partition between Apache and backendConnection timeouts to backend, not response timeoutsError log proxy timeout messages; connect vs response time split
Proxy connection pool too small503s under moderate load, workers not exhaustedmax on ProxyPass/BalancerMember vs concurrent proxied requests
Slow clients inflating %DHigh %D only on large responses, backend fast, scoreboard mostly healthyCorrelate %D with response size; check ConnsAsyncWriting on event MPM
Apache-internal processing (mod_security, rewrite, TLS)High %D with fast backend and small responses; elevated Apache CPUCPU per request; TTFB vs total time split

Quick checks

# 1. Scoreboard state distribution: where are workers stuck?
curl -s http://localhost/server-status?auto | grep Scoreboard | \
  sed 's/Scoreboard: //' | fold -w1 | sort | uniq -c | sort -rn

# 2. Busy vs idle workers
curl -s http://localhost/server-status?auto | grep -E "BusyWorkers|IdleWorkers"

# 3. Time the backend directly, bypassing Apache entirely
curl -s -o /dev/null -w "TTFB: %{time_starttransfer}s  Total: %{time_total}s\n" \
  http://backend-host:backend-port/health

# 4. Time the same path through Apache
curl -s -o /dev/null -w "TTFB: %{time_starttransfer}s  Total: %{time_total}s\n" \
  http://localhost/proxied-path

# 5. Proxy errors in the access log (502/503/504 tell different stories)
tail -5000 /var/log/apache2/access.log | awk '$9 ~ /^50[234]$/ {print $9}' | sort | uniq -c

# 6. Proxy error detail in the error log
grep -E "AH01114|AH00898|proxy:|Connection refused" /var/log/apache2/error.log | tail -20

# 7. Current backend connections from this Apache instance
ss -tn state established dport = :8080 | wc -l   # adjust to your backend port

# 8. P95 of %D for proxied requests (assumes %D is last field)
grep "/proxied-path" /var/log/apache2/access.log | tail -1000 | \
  awk '{print $NF}' | sort -n | awk '{a[NR]=$1} END {print "p95 (us):", a[int(NR*0.95)]}'

Checks 3 and 4 are the decisive pair. If the backend is slow when hit directly, the conversation about Apache config is over. If the backend is fast directly but slow through Apache, the problem is on the Apache side: pool exhaustion, queuing, or client transfer.

How to diagnose it

Apache has no built-in per-request backend timing. %D cannot be decomposed after the fact, so diagnosis has two phases: a live comparison you can run immediately, and instrumentation you add so the next incident is answerable from logs alone.

flowchart TD
  A[Slow proxied responses] --> B{Backend slow when curled directly?}
  B -- Yes --> C[Backend is the problem: DB, GC, app, network to backend]
  B -- No --> D{Scoreboard dominated by W states, IdleWorkers near zero?}
  D -- Yes --> E[Workers held by something: check pool exhaustion and 503s]
  D -- No --> F{High %D only on large responses?}
  F -- Yes --> G[Client transfer time, not server slowness]
  F -- No --> H[Apache-internal: CPU, mod_security, rewrite, TLS]
  1. Probe the backend directly. Run check 3 from the Apache host itself, so the network path matches what mod_proxy uses. Probe the actual slow endpoint, not just a health page: a backend can answer /health in 2ms while the real query takes 20 seconds.
  2. Probe through Apache on the same path. Compare TTFB and total time. The delta between direct and proxied measurements is Apache’s overhead plus client transfer. On localhost with a small response, Apache’s own overhead should be milliseconds.
  3. Read the scoreboard. Workers accumulating in W state plus high direct backend timing means workers are waiting on the backend. Healthy scoreboard plus fast backend timing points at response sizes and client behavior instead.
  4. Split connect failure from response slowness. A backend that refuses connections fails fast (502, “Connection refused”). A backend that accepts and never answers fails slow (504 at ProxyTimeout, default 60s). Different incidents, different owners. The error log distinguishes them.
  5. Check the proxy pool. If the backend is fast but you see 503s under load, suspect the pool. The default max for proxy workers equals ThreadsPerChild (1 for prefork), and pools are per-child-process, so total backend capacity is max times the number of children. Pool exhaustion is a cliff: 503 immediately, no queuing.
  6. Instrument for next time. The live comparison works during an incident, but the real fix is making backend timing a permanent log field (see Fixes).

One caveat on ProxyTimeout: it is a per-I/O socket timeout, not a total request deadline. A backend that dribbles data slowly but continuously may never trigger it, even if the full transfer takes minutes, so a “no 504s” log does not prove the backend is healthy.

Instrumenting backend timing

Three practical options, in increasing order of effort.

Backend-injected response header. Have the backend application emit its own processing time as a response header, for example X-Backend-Time, and log it in Apache:

LogFormat "%h %l %u %t \"%r\" %>s %b %D %{X-Backend-Time}o" combined_timing
CustomLog /var/log/apache2/access.log combined_timing

%D is total time in microseconds; %{X-Backend-Time}o is the backend’s self-reported time. The difference, minus client transfer, is Apache’s overhead. This gives you a per-request decomposition in the access log, which is exactly what %D alone cannot provide. It requires the backend to cooperate; before relying on the value, confirm your framework’s exact header name and units.

mod_log_debug. With hook=all, mod_log_debug logs a message at each phase of request processing, and the microsecond timestamps in the error log let you reconstruct where time went inside Apache. mod_log_debug is marked Experimental in 2.4 and per-phase logging is verbose; use it for a bounded investigation on a single vhost, not as permanent fleet-wide config.

ProxyStatus On. This makes mod_status display per-backend proxy worker status alongside the scoreboard, so you can see which balancer members are busy, in error, or disabled. It does not give per-request backend timing, but it tells you whether specific backends are being marked down, which distinguishes “one backend sick” from “all backends slow.”

A separate option worth knowing: mod_proxy_hcheck (2.4.21+) runs out-of-band health checks against backends using TCP, OPTIONS, HEAD, or GET, independent of live traffic. This separates “backend alive” from “backend fast under real load,” which are different questions.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Backend direct-probe latency (TTFB and total)The only clean measurement of backend time, independent of ApacheP95 above 2x baseline sustained
Scoreboard W state countWorkers waiting on backends pile up hereW above 50% of workers sustained, with normal CPU
IdleWorkersHeadroom before queuing startsTrending toward zero during a latency event
502 / 503 / 504 split502 = backend invalid/refused, 503 = pool or worker exhaustion, 504 = backend timeoutAny sustained proxy error rate; 504-then-503 is the cascade signature
Proxy pool utilization per backendPool exhaustion is a cliff with no queuingBusy proxy workers approaching max; AH00934 logs all workers are busy. Unable to serve
%D correlated with response sizeSeparates client-induced latency from server-inducedHigh %D concentrated on large responses only
Listen backlog Recv-QConfirms workers cannot keep up once exhaustedSustained Recv-Q > 0 alongside high W count

Fixes

If the backend is slow. The fix belongs to the backend team, but you can limit the blast radius at Apache. Reduce ProxyTimeout temporarily so workers fail fast instead of being held for the full 60-second default; requests fail with 504 quickly, workers recycle, and Apache keeps serving whatever it can. If the backend is fully unresponsive, take the Apache instance out of LB rotation rather than letting it absorb the queue.

If the proxy pool is too small. Raise max on the ProxyPass or BalancerMember directive, remembering the pool is per-child-process: total backend connections are max times the number of children. Size it at roughly 2x expected concurrent proxied requests per child at peak. In prefork, be careful in the other direction: N processes times M backends times pool size can overwhelm the backend with connections. Confirm connection reuse rather than confusing it with TCP keepalive probes: keep disablereuse=Off (and enablereuse=On for schemes such as FastCGI that opt in). Do not raise MaxRequestWorkers to fix pool exhaustion; that treats the wrong bottleneck and is a common misdiagnosis spiral.

If client transfer is inflating your latency data. Do not “fix” this by changing Apache. Fix your alerting: filter large responses out of latency percentiles, or alert on the backend header time instead of raw %D. On event MPM, slow-client writes are handled asynchronously and tracked in ConnsAsyncWriting, which is the right signal to watch for client-side slowness.

If the latency is genuinely Apache-internal. Check CPU per request. The usual suspects are TLS handshakes without session resumption, mod_security rule cost, complex mod_rewrite, and compression. That is a different article’s worth of diagnosis; you only reach this branch after backend time and client transfer are excluded.

Prevention

  • Log backend time per request. Add the backend timing header to your LogFormat now, not during the next incident. Without it, every latency investigation starts from the same blind spot.
  • Probe backends directly in your monitoring. A health check through Apache tests the whole chain; a health check against the backend tests the backend. You need both, graphed side by side.
  • Alert on the cascade signature, not just the endpoint state. W states rising, IdleWorkers falling, and 504s appearing together is the early phase. Paging on 503s means you find out after the pool is already exhausted.
  • Separate connect failure from response slowness in dashboards. They have different causes, different owners, and different fixes.
  • Size the proxy pool deliberately. The defaults are too small for most production workloads, and pool exhaustion masquerades as an Apache capacity problem.
  • Use health checks that exercise the real service path. A check that only fetches a static file from Apache reports green while every proxied request times out.

How Netdata helps

  • Netdata collects the mod_status scoreboard continuously, so W-state accumulation, BusyWorkers growth, and IdleWorkers drain are visible as time series, not point-in-time snapshots you happened to catch.
  • Per-second collection of request rate, worker utilization, and connection states lets you watch the cascade sequence (workers filling, throughput dropping, backlog growing) as it develops rather than reconstructing it afterward.
  • Correlating Apache worker states with system CPU and memory on the same dashboard makes the “workers waiting, not working” pattern obvious: that combination is the signature of a backend problem, not an Apache problem.
  • 5xx responses broken down by status code let you watch the 504-to-503 progression that marks a slow backend turning into full pool exhaustion.
  • Because Netdata also monitors the backend host and its services, you can put backend latency and resource pressure next to Apache’s worker states in one view, which is the correlation this entire article depends on.

Netdata’s Apache HTTP Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Apache HTTP Server monitoring with Netdata

Netdata monitors Apache HTTP Server with per-second metrics from mod_status, pre-built dashboards, and ML-powered anomaly detection. Watch busy versus idle workers and the scoreboard state mix, requests per second, bytes served per second, and request processing duration alongside the rest of your stack, so you catch the worker-exhaustion, slow-backend, and memory incidents in these runbooks before they page anyone.