The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-httpd / apache-httpd-cpu-saturation ▌

Operations Guides

Apache CPU saturation: TLS handshakes, mod_deflate, mod_rewrite, and mod_security

Apache serving static content is almost never CPU-bound. The request path (accept, read, sendfile, log) is cheap. When httpd processes start eating cores, the cause is nearly always something doing per-request or per-connection computation: TLS handshakes, mod_deflate compression, mod_rewrite regex evaluation, or mod_security rule processing.

The symptom is gradual, not a cliff. Latency creeps up as CPU saturates. Unlike worker exhaustion, which fails hard at 100% utilization, CPU saturation degrades service progressively, which makes it easy to ignore until p99 latency is unacceptable.

The other trap is tooling. mod_status reports a CPULoad value averaged over the entire server uptime. On a server that has been running for weeks, a current 100% CPU storm barely moves it. If you are using CPULoad to decide whether you have a CPU problem right now, you are looking at the wrong number.

What this means

CPU saturation means request latency is bounded by compute, not by workers, memory, or backends. Workers are busy doing actual work, not waiting. The scoreboard will show W states, but unlike a slow-backend cascade, CPU is elevated and backend health is fine.

Two distinct shapes:

  • All children hot, proportional to traffic. Something makes every request or every connection expensive. Usual suspects: TLS handshakes without session resumption, mod_deflate at a high compression level, mod_security with a heavy ruleset.
  • One child pegged at 100%, others normal. A single CPU-bound request spinning. The classic cause is a mod_rewrite redirect loop, where a rule rewrites a URL that matches the same rule again. One runaway child barely shows in total system CPU on a multi-core box, but that request never completes and the core is gone.

High system CPU (kernel time) rather than user CPU points at syscall volume: high new-connection rate, many small I/O operations, context switching. On Apache this most often means connection churn, which ties back to TLS handshakes and short keepalive timeouts.

flowchart TD
  A[High Apache CPU] --> B{One child at 100%?}
  B -->|yes| C[mod_rewrite redirect loop]
  B -->|no| D{User or system CPU?}
  D -->|system| G[Connection churn: check new-connection rate and TLS]
  D -->|user| E{CPU per request rising?}
  E -->|yes| F[mod_security or mod_deflate or regex cost]
  G --> H{Session resumption working?}
  H -->|no| I[Full handshake on every connection]
  H -->|yes| J[Connection rate itself too high: keepalive, HTTP/2]

Common causes

CauseWhat it looks likeFirst thing to check
TLS full handshakes on every connectionCPU scaling with new-connection rate to 443, high system CPU, latency elevated on first request per connectionTest resumption with openssl s_client -reconnect; check SSLSessionCache config
mod_deflate compressionUser CPU scales with bytes served; large text responses are expensiveDeflateCompressionLevel; which content types are being compressed
mod_rewrite redirect loopExactly one child at 100% CPU on one core, request never completes, same URL repeating in access logps for the hot child; find the looping URL in the access log
Complex mod_rewrite regexElevated CPU per request, proportional to request rate, all children similarRule count and regex complexity; .htaccess vs main config
mod_security rule processingCPU per request high even at moderate traffic; worse on POSTs and large bodiesRule set size; whether response body inspection is on

Quick checks

All read-only and safe on a live server.

# Per-process CPU, hottest first (Debian name fallback included)
ps -C httpd -o pid,%cpu,cputime --sort=-%cpu 2>/dev/null | head || \
  ps -C apache2 -o pid,%cpu,cputime --sort=-%cpu | head

One child at 100% with everyone else near zero is the rewrite-loop signature. All children equally hot is a per-request or per-connection cost.

# User vs system CPU split for all Apache processes
top -bn1 -p $(pgrep -d',' httpd 2>/dev/null || pgrep -d',' apache2) | tail -n +8

High %sy relative to %us points at connection churn and syscall volume, not rule processing.

# New inbound TCP connection rate: PassiveOpens delta over 10s
awk '/^Tcp:/ {print $7}' /proc/net/snmp
sleep 10
awk '/^Tcp:/ {print $7}' /proc/net/snmp

PassiveOpens is the kernel counter for accepted inbound connections. Divide the delta by 10 for connections per second, and compare against your request rate. If connections per second is close to requests per second, clients are not reusing connections, and every connection costs a TLS handshake. This counts all listeners; on a dedicated web host that is overwhelmingly 80/443 traffic.

# Test TLS session resumption (use your real vhost name for SNI)
openssl s_client -connect localhost:443 -servername example.com -reconnect 2>/dev/null | grep -c "Reused"

Zero (or near zero) “Reused” lines means resumption is broken and every connection pays the full handshake cost.

# Request rate from the Total Accesses delta (not ReqPerSec, which is a lifetime average)
curl -s http://localhost/server-status?auto | grep "Total Accesses:"
sleep 30
curl -s http://localhost/server-status?auto | grep "Total Accesses:"
# Look for the looping URL: same path repeated at high frequency
tail -2000 /var/log/apache2/access.log 2>/dev/null | awk '{print $7}' | sort | uniq -c | sort -rn | head || \
  tail -2000 /var/log/httpd/access_log | awk '{print $7}' | sort | uniq -c | sort -rn | head

How to diagnose it

  1. Establish the shape. Run the ps check above. One hot child means a runaway request (step 4). Uniformly hot children mean a systemic per-request cost (steps 2, 3, 5, 6).

  2. Compute CPU per request. Total CPU time consumed over an interval divided by requests completed in the same interval (from the Total Accesses delta). This number should be stable over time. If it is rising, per-request cost is growing: new mod_security rules, more content being compressed, more rewrite rules, or a traffic-mix shift toward expensive endpoints. If it is flat but total CPU is up, volume is up.

  3. Check the connection side. If system CPU is high and the new-connection rate is high relative to request rate, test session resumption with openssl s_client -reconnect. Zero reuse means every connection is a full handshake. Full handshakes with RSA key exchange are roughly an order of magnitude more CPU than a resumed session; on a busy HTTPS site this dominates everything else.

  4. Hunt the single hot child. If one child is pegged, find what it is serving. The full server-status page (not ?auto) shows the current request per slot. Correlate with the access log: a redirect loop shows the same URL (or its redirect target) repeating at high frequency. The usual mechanism is a mod_rewrite loop in per-directory context (.htaccess): the [L] flag stops the current pass, but Apache reinjects the rewritten URI and runs the ruleset again from the top, so a rule that matches its own output loops forever.

  5. Profile mod_deflate. If user CPU tracks bytes served, check what you are compressing and at what level. DeflateCompressionLevel trades CPU for ratio; higher levels cost more CPU for diminishing size returns. Also check whether you are compressing content that is already compressed (images, video, archives), which burns CPU for zero benefit.

  6. Profile mod_security. If CPU per request is high on a site running mod_security with the OWASP CRS, the ruleset itself is the cost. Check whether response body inspection is enabled; it is a major additional cost. The failure mode here is gradual: every rule you add makes every request slightly more expensive, and there is no single moment it breaks.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
CPU per request (CPU time / request delta)Stable per-request cost is the baseline; growth means a rule, filter, or module got more expensiveUpward trend over days
New-connection rate vs request rateReveals whether handshakes or request processing dominates CPUConnection rate approaching request rate (no connection reuse)
TLS resumption rateFull handshakes are the most CPU-intensive thing Apache doesResumption near zero, or dropping after a config change
User vs system CPU splitSeparates rule/compression cost (user) from syscall and connection churn (system)System CPU >30% of total
Per-child CPU distributionA single hot child is a different incident than uniform loadOne child at 100% on one core
Total Apache CPU vs available coresYou want headroom before latency grows unboundedlySustained >70-80% of cores at peak
Scoreboard W states with high CPUDistinguishes “workers computing” from “workers waiting on backend”Many W plus high CPU (backend problems show many W with normal CPU)

Fixes

TLS handshake cost

Enable and share a session cache. SSLSessionCache defaults to none, which means session-ID resumption does not work at all unless you configure it. Use the shared-memory cache so it works across child processes, for example SSLSessionCache shmcb:/path/to/cache(size). Without a shared cache, a client reconnecting to a different child pays a full handshake. Session tickets (RFC 5077) work across children without a shared cache, but the ticket key is static until restart, so after every restart every client does a full handshake anyway. After any crash or restart the cache is empty: expect a CPU burst of full handshakes during warmup.

Prefer ECDSA certificates. ECDSA handshakes are significantly cheaper than RSA-2048 for the server. If you are on RSA certs and CPU-bound on handshakes, this is one of the largest single wins available.

Reduce connection churn. Longer KeepAliveTimeout and HTTP/2 both cut the number of connections, and therefore handshakes, per unit of traffic. HTTP/2 multiplexing means one connection carries many requests.

Watch your OpenSSL version. Apache links against the system OpenSSL, so OpenSSL’s handshake cost is Apache’s handshake cost. OpenSSL 3.5 changed the default TLS supported groups to prefer the hybrid post-quantum key share (X25519MLKEM768), and the OpenSSL project’s own benchmark showed handshake time rising from about 13 ms to 15.5 ms with the change, so a CPU jump after an OS upgrade that bumped OpenSSL is a real possibility. If CPU jumped after an OS upgrade, check the version before blaming Apache.

mod_rewrite loops and regex cost

Fix the loop with [END], not [L]. In per-directory context (.htaccess), [L] only stops the current pass; the rewritten URI is reinjected and the ruleset runs again. The [END] flag (Apache 2.4+) terminates rewrite processing entirely and is the correct way to stop loops. A rule that rewrites a path that still matches its own pattern needs either [END], a RewriteCond guard that excludes the rewritten form, or a redesign.

Move rules out of .htaccess. Besides enabling the reinjection loop behavior, AllowOverride costs per-request filesystem stat() calls on every directory component of every URL. Put the rules in the main config inside <Directory> blocks and set AllowOverride None. This removes both the loop mechanism and the I/O tax.

Simplify hot-path regex. Rules evaluated per request with expensive backtracking patterns are a steady CPU tax at scale. Anchor patterns, avoid nested quantifiers, and order rules so cheap matches short-circuit first.

mod_deflate cost

Lower DeflateCompressionLevel. Higher levels buy smaller output with more CPU. The compiled default is zlib’s default (Z_DEFAULT_COMPRESSION), which zlib documents as equivalent to level 6. If you are CPU-bound, drop toward the low end and measure the bandwidth increase; for most text content the size difference between level 1 and level 9 is far smaller than the CPU difference.

Compress only what benefits. Restrict compression to text content types (HTML, CSS, JS, JSON, XML). Compressing JPEG, video, or archives costs CPU and yields nothing. Note that mod_deflate buffers output to compress it, which raises time to first byte; that is expected behavior, not a separate bug.

mod_security cost

Audit the ruleset against actual traffic. Every active rule is evaluated per request. Tune out rules that never fire on your traffic profile and disable categories you do not need. Rule cost is a ratchet: it only grows unless you deliberately prune.

Reconsider response body inspection. Inspecting response bodies roughly doubles the inspection work and is a common source of a large latency and CPU jump when a ruleset is first enabled. Disable it unless you have a specific need.

Check for known pathological modes. There are reports of specific request patterns causing outsized CPU when SecRuleEngine is DetectionOnly rather than On, and of the persistent collection storage growing unbounded and progressively slowing rule evaluation. If CPU grows slowly over weeks without a config change, check the age and size of the persistent collection files (the IP collection is typically ip.pag on disk) and test whether cycling them restores performance.

Prevention

  • Track CPU per request as a first-class metric. It is the earliest signal for every cause in this article except the rewrite loop. Trend it and alert on drift, not on absolute CPU.
  • Load-test TLS configuration changes. Turning off session tickets, rotating to a new certificate type, or changing the session cache is a CPU change, not just a security change. Measure handshake cost before and after.
  • Test rewrite rules against their own output. Any rule whose substitution can match its own pattern is a loop candidate in per-directory context. Prefer [END] and explicit RewriteCond guards.
  • Keep 30% CPU headroom at peak. CPU degradation is gradual and invisible until it is not. Keep Apache CPU under 70% of available cores at the highest normal traffic period.
  • Expect the cold-start handshake burst. After any restart, session caches are empty and every connection is a full handshake. Do not size CPU headroom assuming warm-cache behavior right after a deploy.

How Netdata helps

  • Netdata charts per-process CPU for httpd/apache2 children, which makes the “one hot child” rewrite-loop signature visible immediately instead of hidden inside a total-CPU average.
  • It tracks Total Accesses as a rate, so you can derive CPU per request and watch it drift, which is the leading indicator for mod_security, mod_deflate, and regex cost growth.
  • Connection metrics on the listener (and ConnsTotal / async connection counters on the event MPM) let you compare new-connection rate against request rate to see whether handshakes or request processing dominates.
  • The scoreboard state distribution separates CPU-bound workers (many W with high CPU) from backend-bound workers (many W with normal CPU), which decides which guide you need next.
  • Per-second collection catches the post-restart TLS handshake burst and short-lived CPU spikes that interval-based polling averages away.

Netdata’s Apache HTTP Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Apache HTTP Server monitoring with Netdata

Netdata monitors Apache HTTP Server with per-second metrics from mod_status, pre-built dashboards, and ML-powered anomaly detection. Watch busy versus idle workers and the scoreboard state mix, requests per second, bytes served per second, and request processing duration alongside the rest of your stack, so you catch the worker-exhaustion, slow-backend, and memory incidents in these runbooks before they page anyone.