The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-httpd / apache-httpd-keepalive-consuming-workers ▌

Operations Guides

Apache keepalive consuming workers: KeepAliveTimeout, the K state, and MPM choice

The symptom looks like a capacity problem: Apache logs its MPM worker-limit code (for example event-MPM AH00484), new connections start queuing, and users see slow responses or 503s. But the request rate is low, and nothing in CPU, memory, or bandwidth explains it. Then you open the scoreboard and see it: a wall of K states. Most of your workers are not serving requests. They are parked on idle keepalive connections, waiting for a next request that may never come.

On the prefork and worker MPMs, each persistent connection holds a worker slot for the full KeepAliveTimeout, even though no work is happening. With a timeout of 15 to 60 seconds and enough concurrent clients, the pool drains at a fraction of the request rate you would expect. On the event MPM this problem largely disappears, because keepalive handling is offloaded to a listener thread.

What this means

HTTP keepalive lets a client reuse one TCP connection for multiple requests, avoiding the cost of reconnecting (and re-doing the TLS handshake) per request. The tradeoff is that Apache must decide what to do with the connection between requests. The answer depends entirely on the MPM:

  • prefork: one child process per connection. A keepalive connection pins an entire process in the K (keepalive read) state until the next request arrives or KeepAliveTimeout expires.
  • worker: one thread per connection. Same problem, one level cheaper: a keepalive connection pins a thread.
  • event: a dedicated listener thread watches idle keepalive sockets asynchronously (epoll/kqueue) and hands the connection to a worker thread only when a request actually arrives. Idle keepalive connections show up in the ConnsAsyncKeepAlive counter in server-status, not in worker slots.

So on prefork and worker, effective capacity is not MaxRequestWorkers divided by request rate. It is MaxRequestWorkers divided by (request rate x total worker hold time), where hold time includes the keepalive wait. A client that makes one request and then holds the connection for a 30-second KeepAliveTimeout occupies a worker 100 times longer than a request that takes 300 ms.

flowchart LR
  subgraph PF["prefork / worker MPM"]
    C1[client connection] --> W1[worker process/thread]
    W1 -->|request done| K1["K state: waits KeepAliveTimeout"]
    K1 -->|timeout or next request| C1
  end
  subgraph EV["event MPM"]
    C2[client connection] --> L[listener thread]
    L -->|request arrives| W2[worker thread]
    W2 -->|response flushed| L
    L -->|idle| K2["ConnsAsyncKeepAlive (no worker held)"]
  end

Common causes

CauseWhat it looks likeFirst thing to check
KeepAliveTimeout too high for the MPMScoreboard dominated by K, MaxRequestWorkers hit at modest RPSScoreboard state distribution and the configured timeout
Running prefork/worker when event would fitSame as above, chronic rather than incidentalapachectl -V output for the active MPM
Clients holding connections without pipelining requestsMany established connections, low request rate relative to connection countss connection count vs Total Accesses delta
K states on event MPMAbnormal: keepalive offloading is failing (e.g., a connection filter incompatible with event forces fallback to worker-style handling)ConnsAsyncKeepAlive vs scoreboard K count
Load balancer or proxy in front reusing connections slowlyA small number of source IPs hold many keepalive slotsss -tn source IP breakdown

Quick checks

All of these are read-only.

# 1. Confirm which MPM is active
apachectl -V 2>/dev/null | grep -i mpm || httpd -V | grep -i mpm

# 2. Pull the machine-readable status page
curl -s http://localhost/server-status?auto

# 3. Count scoreboard states - look for a dominant K population
curl -s http://localhost/server-status?auto | grep "Scoreboard:" | \
  awk '{print $2}' | fold -w1 | sort | uniq -c | sort -nr

# 4. Busy vs idle workers
curl -s http://localhost/server-status?auto | grep -E "BusyWorkers|IdleWorkers"

# 5. On event MPM: check async connection counters
curl -s http://localhost/server-status?auto | grep -E "^Conns"

# 6. Confirm worker exhaustion events in the error log
grep -hE 'AH(00161|00286|00484)' /var/log/apache2/error.log /var/log/httpd/error_log 2>/dev/null | tail -5

# 7. Compare connection count to request throughput
ss -tn 'state established and ( sport = :80 or sport = :443 )' | wc -l

The diagnostic signature: K count is a large fraction of all slots, BusyWorkers is near MaxRequestWorkers, IdleWorkers is zero, but the Total Accesses delta over a minute is modest. High hold time, low work.

How to diagnose it

  1. Confirm the MPM. Everything downstream depends on this: apachectl -V | grep -i mpm, or check the loaded mpm_* module in your config. If you are on prefork because of mod_php or another non-thread-safe module, keepalive tuning matters more, not less.

  2. Read the scoreboard. Take the state distribution (check 3 above). On prefork/worker, K above roughly 30% of slots is the pattern. On event, keepalive should barely appear in the scoreboard at all; a significant K population on event is itself a finding, because it means connections are not being offloaded to the listener thread. Apache falls back to worker-style handling for connection filters that declare themselves incompatible with event, in which case one worker thread is reserved per connection again.

  3. Separate K from W. If the scoreboard is dominated by W (sending reply) instead of K, you have a different problem: slow clients or a slow backend holding workers during response generation. Keepalive tuning will not fix that. See Apache scoreboard states explained for the full state-by-state reading.

  4. Check the queue and the error log. Run ss -ltn on the listening ports: a growing Recv-Q while workers are stuck in K confirms new connections are queuing behind idle keepalive holds. The MPM worker-limit code in the error log (prefork AH00161, worker AH00286, event AH00484) confirms the pool hit its ceiling.

  5. Quantify the hold time. Pull the current KeepAliveTimeout, KeepAlive, and MaxKeepAliveRequests values from the running config. Defaults in 2.4.x are KeepAlive On, KeepAliveTimeout 5, MaxKeepAliveRequests 100. If someone raised the timeout to 15, 30, or 60 seconds “to reduce handshake overhead,” that is your smoking gun on prefork/worker.

  6. Check who is holding the connections. ss -tn sport = :80 | awk '{print $5}' | cut -d: -f1 | sort | uniq -c | sort -rn | head shows whether a few source IPs (a load balancer, a monitoring system, a single misbehaving client) account for most of the held connections.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Scoreboard K fraction (prefork/worker)Direct measure of workers parked on idle keepaliveSustained above ~30% of slots
Scoreboard K on event MPMShould be near zero; keepalive lives in the listener threadAny significant, sustained K population
ConnsAsyncKeepAlive (event only)Idle keepalive handled cheaply by the listenerNot an alarm by itself; compare against scoreboard K to detect offloading failure
BusyWorkers / MaxRequestWorkersSaturation against the pool ceilingAbove 80% sustained with low RPS
Listen backlog Recv-QConnections queuing before a worker accepts themSustained non-zero alongside high K
MPM worker-limit codeDefinitive pool exhaustionAny occurrence during this pattern
Established connections vs request rateReveals connection hoardingConnection count high, Total Accesses delta low

Fixes

Reduce KeepAliveTimeout (prefork/worker)

This is the direct fix. Apache’s own performance guidance recommends keeping the timeout low and discourages raising it above roughly 60 seconds; the 5-second default exists precisely to bound this effect. For high-traffic prefork or worker deployments, 1 to 2 seconds is a common choice. Most browsers fire their follow-up requests (CSS, JS, images) within a second or two of the first response, so a short timeout preserves the practical benefit of keepalive while freeing workers quickly.

KeepAlive On
KeepAliveTimeout 2

Tradeoff: shorter timeouts mean clients reconnect more often. On HTTPS that means more TLS handshakes, which cost CPU. If you cut the timeout and see CPU rise on TLS-heavy sites, that is the cost of the trade, and it is usually still cheaper than exhausted workers.

Bound MaxKeepAliveRequests

MaxKeepAliveRequests caps how many requests one connection can serve before Apache closes it (default 100). Lowering it forces periodic connection turnover, which recycles workers and also helps bound per-child memory growth. It does not fix the idle-hold problem on its own; combine it with the timeout change.

Move to the event MPM

If nothing in your module set requires prefork (classic mod_php is the usual reason), event removes the problem structurally: keepalive connections stop consuming worker threads at all. Two caveats:

  • Switching MPM means disabling one mpm_* module and enabling another, then a full Apache restart, so plan it as a disruptive change. Validate module compatibility (mod_php in particular) on a staging host first.
  • Some connection filters are incompatible with event and force it back into worker-style one-thread-per-connection behavior for those connections, which quietly reintroduces the problem for the affected traffic.

After switching, verify with the scoreboard: K should nearly vanish and ConnsAsyncKeepAlive should carry the idle load instead.

Raise MaxRequestWorkers only after fixing the hold time

Adding workers to absorb keepalive holds is buying memory to park idle connections. On prefork especially, each worker is a full process, and MaxRequestWorkers x per-child RSS must stay within RAM. Fix KeepAliveTimeout first; then re-evaluate pool size against real concurrency. See Apache MaxRequestWorkers tuning for the sizing math.

Put a keepalive-friendly layer in front

If you cannot leave prefork (mod_php) and cannot tolerate short timeouts, a reverse proxy or load balancer that handles keepalive efficiently in front of Apache absorbs the idle connections and speaks to Apache over a small, well-behaved connection pool. This moves the problem off the scarce resource (Apache workers) onto a component built for cheap connection handling.

Prevention

  • Alert on the K fraction, not just BusyWorkers. BusyWorkers at 95% tells you the pool is full; the K fraction tells you why, and it rises before the pool fills.
  • Alert on K appearing on event MPM. It means offloading has failed and you are silently back to worker-style connection costs.
  • Pin KeepAliveTimeout in config management and treat increases as a reviewed change. A well-meaning bump from 5 to 30 seconds is how this incident usually starts.
  • Watch connections-per-request ratio over time. A rising ratio at constant RPS is an early warning that clients or intermediaries are holding connections longer.
  • Re-validate after any MPM or module change. Switching MPMs, adding a connection filter, or fronting Apache with a new LB all change keepalive economics.

How Netdata helps

  • Netdata collects the Apache scoreboard continuously via server-status, so you see the K population as a time series rather than a point-in-time snapshot during the incident.
  • BusyWorkers and IdleWorkers are charted together, making the “full pool, low throughput” divergence visible at a glance.
  • On event MPM, ConnsAsyncKeepAlive and the other async connection counters are charted alongside worker states, so offloading failures (scoreboard K rising while async keepalive stays flat) stand out.
  • Correlating the scoreboard breakdown with requests per second on the same dashboard separates keepalive hoarding (K-dominant, low RPS) from slow-backend starvation (W-dominant) in seconds.
  • Error log alerting catches AH00484 the moment the pool ceiling is hit, instead of after users report 503s.

Netdata’s Apache HTTP Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Apache HTTP Server monitoring with Netdata

Netdata monitors Apache HTTP Server with per-second metrics from mod_status, pre-built dashboards, and ML-powered anomaly detection. Watch busy versus idle workers and the scoreboard state mix, requests per second, bytes served per second, and request processing duration alongside the rest of your stack, so you catch the worker-exhaustion, slow-backend, and memory incidents in these runbooks before they page anyone.