The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-httpd / apache-httpd-how-it-works-in-production ▌

Operations Guides

How Apache HTTPD actually works in production: a mental model for operators

Most Apache incidents are misdiagnosed for the same reason: the operator is reading signals without knowing which execution model produced them. A scoreboard full of K states is a crisis on prefork and a rounding error on event. “Apache is down” and “the backend is down” look identical from outside a reverse proxy. MaxRequestWorkers means something different depending on whether your concurrency unit is a process or a thread.

This article is the model layer. It covers the abstractions every Apache runbook, alert, and tuning decision depends on: the three MPMs, the scoreboard, worker pool arithmetic, the kernel listen backlog, the module pipeline, mod_proxy connection pools, and APR memory pools. Almost every Apache outage is one of a small set of resource exhaustion patterns, and each maps directly onto a structure described here.

This is grounded in Apache 2.4 behavior. Where other versions differ in a way that matters operationally, it is called out.

The MPM decides everything

The Multi-Processing Module determines how an incoming connection maps to an execution unit. It is compiled in or loaded once, and it shapes the meaning of nearly every metric you will collect. Check which one you are running before interpreting anything:

# Identify the active MPM
apachectl -V 2>/dev/null | grep MPM
# or
httpd -V 2>/dev/null | grep MPM

prefork. One process per connection. Each child handles exactly one request at a time. No threading, which makes it safe for non-thread-safe modules (classic mod_php is the usual reason it still exists). Per-process RSS is typically 10-50 MB or more depending on loaded modules. The resource you run out of is memory, because every concurrent connection costs a full process.

worker. Hybrid. Multiple child processes, each running multiple threads, each thread handling one connection. Far more memory-efficient than prefork. The failure mode to internalize: a single stuck backend can hold a thread indefinitely, starving that child’s thread pool.

event. An evolution of worker, and the default in Apache 2.4. A dedicated listener thread manages keepalive connections asynchronously (epoll on Linux, kqueue on BSD) and hands a connection to a worker thread only when an actual request arrives. This decouples keepalive connection count from worker thread consumption. On event, keepalive connections show up in ConnsAsyncKeepAlive in server-status, not as K states in the scoreboard. Significant K in the scoreboard on event is abnormal and worth investigating; the same picture on prefork is routine.

Two event-MPM caveats from the 2.4 documentation that matter in production: it falls back to worker-like behavior for connection filters that declare themselves incompatible with async operation, and for output filters that must buffer the whole response body (CGI, FCGI, and some proxied content paths). And the event MPM’s well-known “scoreboard is full, not at MaxRequestWorkers” error means all active-process slots are occupied even though active requests are below MaxRequestWorkers, often because graceful generations linger. Since 2.4.24, event may use allocated slots up to ServerLimit for graceful processes while MaxRequestWorkers continues to bound active processes; the error remains possible at that higher boundary.

How a request moves through the server

The event MPM path, since that is what most production 2.4 installs run:

flowchart LR
  C[Client] -->|SYN| KB[Kernel listen backlog]
  KB -->|accept| L[Listener thread]
  L -->|new request| W[Worker thread pool]
  L -.->|idle keepalive| KA[Async keepalive via epoll]
  KA -->|request arrives| W
  W --> P[Module pipeline phases]
  P --> H[Handler or mod_proxy]
  H --> B[Backend connection pool]
  H --> R[Response to client]

On prefork there is no listener thread or async keepalive: a child process accepts a connection and holds it for its entire lifetime, including keepalive idle time. That single difference explains why prefork melts under connection floods that event handles trivially.

The scoreboard: Apache’s most diagnostic structure

The scoreboard is a shared-memory segment with one slot per possible worker, each slot recording the worker’s current state. Its size is fixed at startup as ServerLimit times ThreadsPerChild and cannot grow dynamically. mod_status is just a reader of this segment.

The states that drive diagnosis:

StateMeaningWhat a lot of them usually means
_Waiting (idle)Healthy headroom
RReading requestSlow clients, slow uploads, or Slowloris
WSending replySlow clients, or workers blocked on a slow backend
KKeepaliveWorker slots held by idle connections (prefork/worker problem)
DDNS lookupHostname-based access control blocking workers
LLoggingLog pipe stall or full log disk
GGracefully finishingOld generation lingering after graceful restart
.Open slot, no processUnused pre-allocated capacity

Two structural facts matter. First, W is ambiguous: writing bytes to the client, waiting on a backend, and doing internal processing all render as W, so a scoreboard full of W tells you workers are held, not why. Second, a slot in . state with load present means the scoreboard has room but children are not spawning, which points at memory or process limits rather than Apache config.

Worker pool arithmetic

Three directives bound the pool, and they interact:

  • ThreadsPerChild: threads per child process. Default 25 on worker/event.
  • ServerLimit: ceiling on child processes. Default 16 on worker/event, 256 on prefork.
  • MaxRequestWorkers: the actual cap on simultaneous requests, renamed from MaxClients in 2.3.13.

On worker/event, the defaults multiply out to 16 x 25 = 400. On prefork, where each process is one worker, the default MaxRequestWorkers is 256. The sizing rule that prevents the most common catastrophic misconfiguration:

MaxRequestWorkers <= available_memory_for_apache / per_worker_memory

Setting MaxRequestWorkers to 1000 on a 4 GB host with mod_php children at 50 MB is a 50 GB commitment against 4 GB of RAM. The result is swap thrash and an OOM kill cascade under load. Derive the limit from measured per-child RSS, and keep Apache’s theoretical maximum (MaxRequestWorkers x worst-case child RSS) under roughly 70% of RAM.

One operational trap from the 2.4 docs: changes to ServerLimit (and ThreadLimit) are ignored during a graceful restart. They only take effect on a full stop and start. If you raised ServerLimit, did a graceful, and saw no effect, that is why.

The kernel listen backlog

When all workers are busy, new connections do not fail immediately. They queue in the kernel’s TCP listen backlog, sized by ListenBacklog (default 511) and silently capped by net.core.somaxconn if that is lower. When the backlog fills, clients get RST or timeouts: the port is open, the process is alive, and the service is effectively down.

This is the leading indicator of worker exhaustion. ss -ltn shows Recv-Q (current backlog depth) against Send-Q (the configured maximum) for listening sockets. Brief non-zero Recv-Q during bursts is normal; sustained Recv-Q means connections arrive faster than Apache accepts them, and it appears before user-visible latency because those connections have not reached a worker yet.

The module and filter pipeline

Every request passes through ordered phases: URI translation, access control, authentication, content generation by a handler, then logging. Every module hook can block, fail, or add latency, and input/output filters (mod_ssl, mod_deflate, mod_headers) sit in a chain around the handler, each allocating from the request’s memory pool.

Operationally, this explains three things. A worker in W state may be stuck inside any module in the chain, not “sending a reply.” mod_deflate buffers before it can send the first byte, which inflates time-to-first-byte without anything being wrong. And complex filter chains inflate per-request memory, which feeds back into the worker pool sizing math above.

mod_proxy backend pools

When Apache reverse-proxies, each child process maintains its own pool of backend connections. Two facts cause most proxy-side incidents:

  • The default pool max equals ThreadsPerChild (so 1 on prefork). Total backend concurrency is max times the number of children, and the defaults are far too small for production load.
  • Pool exhaustion is a cliff. Without acquire, an unavailable worker returns HTTP 503 immediately; with acquire, Apache waits only for the configured milliseconds, then returns SERVER_BUSY. There is no unbounded queue.

This produces the signature misdiagnosis: 503s under moderate load, operator raises MaxRequestWorkers (the wrong bottleneck), nothing improves. Separately, a slow backend holds frontend workers in W state for the full backend latency, so backend slowness consumes your entire worker pool from the inside while Apache’s own CPU and memory look normal. Proxy-specific error codes decode as: 502 (backend refused or answered garbage), 503 (pool exhausted or all balancer members errored), 504 (backend exceeded ProxyTimeout).

APR memory pools

Apache allocates through Apache Portable Runtime pools rather than per-object malloc. Memory for a connection or request comes from a pool that is destroyed wholesale when that connection or request ends. This is fast and prevents most classic leaks, but freed memory returns to the process’s allocator, not necessarily to the OS, so a child that has served many requests can hold a large RSS indefinitely. MaxMemFree (default 2048 KB) caps how much free memory the allocator retains, but it limits the free list, not total process size. The reliable bound is MaxConnectionsPerChild: the default 0 means children never recycle, which is the wrong default for anything running mod_php or mod_perl. A finite value (commonly 5000-10000) forces periodic recycling and turns unbounded leak growth into a sawtooth.

Where the model pays off

Each characteristic Apache failure is one of these structures saturating:

  • Worker exhaustion: all slots in W, R, D, or K; backlog fills; the MPM worker-limit message (for example AH00484) appears once in the error log.
  • Slow backend cascade: W states climb, IdleWorkers falls, 504s then 503s, while request rate paradoxically drops because queued requests never complete.
  • Memory exhaustion: MaxRequestWorkers x per-child RSS exceeds RAM; swap, then OOM kills, then respawn-and-leak-again.
  • Slowloris: many connections trickling bytes hold workers in R indefinitely; mod_reqtimeout is the defense.
  • Log stall: disk full or dead log pipe; workers finish requests but block in L; the server looks alive and serves nothing.
  • Graceful restart pile-up: G states accumulate, multiple child generations coexist, and memory multiplies. Reduce reload frequency; GracefulShutdownTimeout applies to graceful-stop, not a graceful reload.

Signals to watch in production

SignalWhy it mattersWarning sign
BusyWorkers / MaxRequestWorkersPrimary saturation gauge; degradation is a cliff, not a slopeSustained above 80%; IdleWorkers at zero
Scoreboard state distributionTells you where workers are held, not just that they are heldW dominant with normal RPS; R above ~20%; any L spike
Listen backlog Recv-QFills before users see failuresSustained non-zero; rising ListenOverflows counter
MPM worker-limit message in error logApache explicitly reporting pool exhaustionAny occurrence
ConnsAsyncKeepAlive (event only)Confirms keepalive offload is workingLow async count alongside many scoreboard K states on event
Per-child RSS trendSets the real MaxRequestWorkers ceiling; exposes leaksMonotonic growth per PID over hours or days
502/503/504 rate (proxy)Separates backend failure from Apache failureAny sustained rate; 503 at moderate load means pool sizing
(MaxRequestWorkers x avg RSS) / RAMThe OOM tripwireAbove ~70%

How Netdata helps

  • The Apache collector polls server-status?auto and charts BusyWorkers, IdleWorkers, request rate, and the full scoreboard state distribution over time, so a drift toward W or R dominance is visible before the pool empties.
  • Async connection counters (ConnsTotal, ConnsAsyncKeepAlive, ConnsAsyncWriting, ConnsAsyncClosing) are charted on event MPM, which makes the “why are there K states on event” question answerable at a glance.
  • Because Netdata also collects per-process RSS, system memory, and TCP listen socket stats from the same host, you can correlate worker utilization with the backlog filling and with per-child memory growth in one view, instead of sampling ss and ps by hand during an incident.
  • Error rate and latency signals from log-based collection sit on the same dashboard as the scoreboard, which is exactly the join you need to distinguish a slow backend cascade (workers held, Apache CPU idle, 504s rising) from genuine overload.

Apache monitoring in Netdata brings these signals together with per-second metrics and anomaly detection.

The Netdata solution

Apache HTTP Server monitoring with Netdata

Netdata monitors Apache HTTP Server with per-second metrics from mod_status, pre-built dashboards, and ML-powered anomaly detection. Watch busy versus idle workers and the scoreboard state mix, requests per second, bytes served per second, and request processing duration alongside the rest of your stack, so you catch the worker-exhaustion, slow-backend, and memory incidents in these runbooks before they page anyone.