The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-httpd / apache-httpd-tls-handshake-cpu ▌

Operations Guides

Apache TLS handshake CPU: broken session resumption and the HTTPS CPU wall

The symptom looks like a capacity problem: Apache processes pinned near 100% CPU, request latency climbing, and no obvious cause in the access log. Traffic is up, but not absurdly. The workers are not stuck on a backend. The scoreboard is busy but not exhausted. And yet the box is melting.

On HTTPS-heavy servers, the usual suspect is the TLS handshake. A full TLS handshake with an RSA-2048 certificate costs on the order of 15 ms of CPU time per connection. Session resumption cuts that cost by roughly 10x. When resumption works, repeat clients skip the expensive public-key operation entirely. When resumption is broken, every connection pays the full price, and at a few hundred new connections per second the math stops working.

That is the HTTPS CPU wall: the point where the cost of full handshakes per second exceeds the CPU you have, and no amount of worker tuning gets it back. This article covers how to confirm that broken session resumption is the cause, why resumption breaks in Apache, and how to fix it.

What this means

TLS offers two ways to resume a session: session IDs (server-side state) and session tickets (RFC 5077, client-held encrypted state). Apache 2.4’s defaults quietly undermine the first mechanism:

  • SSLSessionCache defaults to none. The inter-process session cache is disabled entirely out of the box.
  • Without a shared cache, a session ID is only valid in the child process that created it. With many children and a load-balanced client reconnecting to a random child, cache hits are rare.
  • Session tickets are on by default (SSLSessionTickets on) and work across children, but they have their own traps: Apache generates one random ticket key at startup and never rotates it, and with TLS 1.3 a graceful restart can silently break ticket-based resumption.

The result is a server where most reconnecting clients do a full handshake anyway. Because a full handshake is the single most CPU-intensive thing Apache does, the failure shows up as a CPU problem, not as a TLS error. Nothing is logged. A silently full or absent session cache means old sessions evicted, resumption rate drops, CPU increases, and no error logged anywhere.

flowchart TD
  A[New connection to :443] --> B{Session resumed?}
  B -- yes --> C[Abbreviated handshake
low CPU] B -- no --> D[Full handshake
~15 ms CPU with RSA-2048] E[Resumption broken:
SSLSessionCache none,
shmcb missing or too small,
TLS 1.3 + graceful restart] --> B D --> F[CPU saturated at
moderate connection rate] F --> G[Latency climbs,
handshake queue grows]

Common causes

CauseWhat it looks likeFirst thing to check
SSLSessionCache none (the default)Low resumption rate since forever; CPU tracks new-connection rate linearlyapachectl -t -D DUMP_RUN_CFG or grep the config for SSLSessionCache
shmcb cache too smallResumption worked at low traffic, degraded as traffic grew; no error loggedCache size vs concurrent session count; test resumption under load
mod_socache_shmcb not loaded (2.2 to 2.4 upgrade)Config references shmcb but startup fails or cache is silently unsupportedError log for SSLSessionCache: 'shmcb' session cache not supported
TLS 1.3 resumption broken after graceful restartResumption works after a full restart, stops after apachectl gracefulTest openssl s_client -reconnect before and after a graceful restart
Session tickets disabled, cache not sharedFull handshakes for every client that lands on a different childCheck SSLSessionTickets and whether any shared cache is configured
RSA-2048 certificate on a high-connection-rate vhostHigh per-handshake cost even when resumption works; CPU per connection highopenssl s_client shows cert type; openssl speed rsa2048 ecdsap256 shows the gap
No HTTP/2, clients opening parallel connectionsMany connections per client, each needing a handshakeConnection count vs request count; check if mod_http2 is loaded

Quick checks

All read-only, safe to run on a production box.

# 1. Confirm CPU is the bottleneck and it is Apache consuming it
ps -C httpd -o pid,%cpu --sort=-%cpu 2>/dev/null || ps -C apache2 -o pid,%cpu --sort=-%cpu

# 2. New-connection rate: PassiveOpens in /proc/net/snmp counts accepted
#    TCP connections (all ports). Sample twice and take the delta.
awk '$1=="Tcp:"{print $7}' /proc/net/snmp
sleep 10
awk '$1=="Tcp:"{print $7}' /proc/net/snmp
# Note: the delta of established counts (ss -tn state established) is NOT a
# new-connection rate, because closed connections vanish from the list.

# 3. Test session resumption directly against the server
openssl s_client -connect localhost:443 -reconnect 2>/dev/null | grep -c "Reused"
# Expect 5 (six connections, five reused). 0 means resumption is fully broken.

# 4. Check the configured session cache
grep -rEi 'SSLSessionCache|SSLSessionTickets' /etc/httpd/ /etc/apache2/ 2>/dev/null

# 5. Verify the shmcb provider module is loaded
apachectl -M 2>/dev/null | grep socache

# 6. Look for the cache-provider error and OCSP stapling errors
grep -E "shmcb|AH01929|AH02217" /var/log/apache2/error.log /var/log/httpd/error_log 2>/dev/null | tail

# 7. See what certificate type clients pay for
echo | openssl s_client -connect localhost:443 -servername $(hostname) 2>/dev/null | \
  openssl x509 -noout -text | grep "Public Key Algorithm"

Check 3 is the decisive one. s_client -reconnect establishes six connections on one session and reports how many were reused. If it prints 0, resumption is broken end to end, regardless of what the config says.

How to diagnose it

  1. Correlate CPU with new connections, not requests. Pull CPU and the new-connection rate to :443 for the same window. If CPU tracks connection rate rather than request rate, handshakes are the load. If CPU tracks request rate instead, look at mod_deflate, mod_rewrite, or mod_security first; see Apache CPU saturation.

  2. Measure resumption from the client’s perspective. Run openssl s_client -connect localhost:443 -reconnect and count “Reused” lines. Run it again against the public hostname through the load balancer, not just localhost, because the LB may terminate TLS itself or distribute connections across Apache children.

  3. Determine which resumption mechanism should be working. If SSLSessionTickets on (the default), tickets should resume across children. If tickets are off, resumption depends entirely on the server-side cache, which must be shmcb shared memory. Per-child in-memory caching does not survive process boundaries.

  4. Check the cache configuration and module. Confirm mod_socache_shmcb is loaded (it moved to a separate module in 2.4) and that SSLSessionCache shmcb:/path(size) is set. A missing module on an upgraded config produces SSLSessionCache: 'shmcb' session cache not supported. Loading a newly added module requires a full restart, not a reload.

  5. Test the graceful-restart edge case. If you run TLS 1.3 and recently did apachectl graceful, re-run the -reconnect test. There is a known behavior where TLS 1.3 resumption silently falls back to full handshakes after a graceful restart with session tickets on; TLS 1.2 is unaffected.

  6. Estimate cache sizing. The commonly copied example shmcb:/path/to/ssl_scache(512000) allocates 512 KiB. On a busy server this turns over fast and effective resumption collapses with no error logged. If resumption degrades at peak but works off-peak, cache size is the suspect. There is no official per-entry size for shmcb; operator estimates range from a few hundred bytes per TLS 1.2 session to roughly 1 KB with larger TLS 1.3 sessions, so size empirically from the mod_status cache counters rather than a formula.

  7. Rule out a post-restart handshake burst. After any crash or hard restart, the session cache is empty and you get a transient burst of full handshakes. That is normal cold-start behavior. It is a problem only if CPU stays high after the cache should have warmed.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Session resumption rate (cache hit/miss)The direct measure of the problemBelow ~80% on a server with repeat visitors
New-connection rate to :443New connections are potential full handshakesRising without a matching request-rate rise
Apache CPU per coreFull handshakes are the top CPU consumerSustained >80% at moderate traffic
CPU per requestDistinguishes per-connection cost from per-request costRising when connection rate rises, flat when request rate rises
Connections vs requests ratioHTTP/2 multiplexing lowers handshake countMany connections per request (HTTP/1.1 parallel connections)
TTFB on new connectionsHandshake time lands in the first byteElevated TTFB concentrated on new connections
OCSP stapling errors (AH01929, AH02217)Silent stapling failure adds client-side latencyAny occurrence in the error log

Apache exposes no native metric for resumption rate. You infer it from the correlation above, or measure it periodically with openssl s_client -reconnect. That gap is why this failure goes undiagnosed for so long.

Fixes

Enable and size the shared session cache

Load mod_socache_shmcb and configure a shared-memory cache, sized for your concurrent session count rather than the 512 KiB example from the docs:

LoadModule socache_shmcb_module modules/mod_socache_shmcb.so
SSLSessionCache shmcb:/var/run/httpd/ssl_scache(10485760)
SSLSessionCacheTimeout 300

The path differs by distribution. A full restart (not graceful) is required after loading the module, so schedule it: a restart drops in-flight connections and briefly empties the session cache. SSLSessionCacheTimeout defaults to 300 seconds; note the timeout is only checked when a session is presented, so stale entries are never purged in the background. Raise the timeout if clients typically return within a longer window, but size the cache so it does not fill first.

Deal with the TLS 1.3 graceful-restart bug

If resumption dies after every graceful restart and you serve TLS 1.3, the documented workaround is SSLSessionTickets off, which switches OpenSSL to stateful (session ID) resumption backed by the shmcb cache. Tradeoff: tickets are off, so resumption now depends entirely on your shared cache being enabled and sized correctly. Do not set tickets off without configuring shmcb, or resumption gets worse.

Prefer ECDSA certificates

ECDSA P-256 signing is roughly an order of magnitude faster than RSA-2048 per operation, which directly cuts the CPU cost of every full handshake. Measure the gap on your own hardware:

# Compare signing performance on this host
openssl speed rsa2048 ecdsap256

Tradeoff: older clients may lack ECDSA support, so dual RSA+ECDSA certificate deployment is common during transition. Even with resumption fixed, ECDSA shrinks the residual cost of the handshakes that cannot be resumed (first visits, cache misses, post-restart bursts).

Reduce handshake count with HTTP/2

HTTP/2 multiplexes many requests over one connection, which directly reduces the number of handshakes your clients need. Requirements: mod_http2 and the event MPM. Enabling it under mpm_prefork logs AH10034: The mpm module (prefork.c) is not supported by mod_http2 as a warning and HTTP/2 stays inactive, which usually means migrating from mod_php to PHP-FPM first. This does not fix broken resumption, but it lowers the connection rate, which lowers the price of any resumption gap.

Rotate ticket keys deliberately

Apache generates one random ticket key at startup and keeps it for the life of the process. There is no automatic rotation, which weakens forward secrecy for resumed sessions. The practical approaches are periodic restarts (daily, via cron) or SSLSessionTicketKeyFile pointing at a 48-byte random key file that you replace on a schedule, with a restart to pick up the new key. Key rotation via restart empties the shmcb cache as a side effect, so expect a brief full-handshake burst afterward.

Prevention

  • Set SSLSessionCache explicitly. The default is none. Any HTTPS server with repeat clients should have a sized shmcb cache in the base config.
  • Alert on resumption rate. A periodic openssl s_client -reconnect probe from an external host catches regressions that no log line ever will.
  • Graph CPU against new-connection rate. The divergence between request-driven and connection-driven CPU is the earliest warning.
  • Test resumption in deploy pipelines. Run the -reconnect probe after every config change or restart that touches mod_ssl, including graceful restarts.
  • Keep Apache current. TLS session handling has had real security fixes, including CVE-2025-23048 (access control bypass via session resumption, fixed in 2.4.64) and CVE-2024-47252 (mod_ssl log injection). Resumption bugs are not only a performance concern.

How Netdata helps

  • Per-second CPU per process, so handshake-driven load is visible as a sharp ramp correlated with traffic events rather than a five-minute average that hides it.
  • New-connection and established-connection counts on :443, letting you overlay connection rate against CPU and see immediately whether load tracks connections or requests.
  • Connections-per-request visibility that shows whether HTTP/2 multiplexing is actually reducing connection churn.
  • Certificate expiry monitoring, so the certificate side of TLS (including renewals that restart Apache and empty the session cache) is watched alongside performance.
  • Restart and uptime tracking, which explains post-restart full-handshake bursts when the cache empties.
  • Correlating these signals on one dashboard shortens the path from “CPU is high” to “resumption rate collapsed at 14:02, right after the graceful restart.”

Netdata’s Apache HTTP Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Apache HTTP Server monitoring with Netdata

Netdata monitors Apache HTTP Server with per-second metrics from mod_status, pre-built dashboards, and ML-powered anomaly detection. Watch busy versus idle workers and the scoreboard state mix, requests per second, bytes served per second, and request processing duration alongside the rest of your stack, so you catch the worker-exhaustion, slow-backend, and memory incidents in these runbooks before they page anyone.