The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-httpd / apache-httpd-ocsp-stapling-failure ▌

Operations Guides

Apache OCSP stapling failure: the silent handshake that adds client latency

OCSP stapling failure is a textbook “silently catastrophic” signal: nothing in the error log at default levels, no 5xx, no worker pile-up, and yet every new TLS handshake is slower than it should be. When Apache cannot fetch or attach a stapled OCSP response, each client that wants revocation information makes its own OCSP request to the CA’s responder, adding hundreds of milliseconds before the handshake completes. The server looks healthy from the inside. The slowness only exists on the client side.

None of the usual Apache signals move: request rate is normal, BusyWorkers are normal, the scoreboard is normal, the error log is quiet. The only way to see the failure is to probe the handshake itself or to know which two error codes to grep for.

This guide covers how stapling is supposed to work, how to confirm it is broken, the small set of causes that account for almost all failures, and how to stop it from regressing silently after the next restart or certificate renewal.

What this means

With OCSP stapling, Apache periodically fetches a signed revocation status for its certificate from the CA’s OCSP responder (the URL comes from the certificate’s AIA extension) and attaches (“staples”) it to the TLS handshake. The client gets proof of freshness without contacting the CA.

When stapling is broken, Apache completes handshakes without a staple. Clients that check revocation (browsers under enterprise policy, many API clients, anything enforcing Must-Staple) then make their own round trip to the OCSP responder before trusting the connection. That round trip is on the critical path of connection establishment.

Two properties make this nasty:

  • No startup prefetch and no persistent cache. Apache does not fetch OCSP responses at startup; it fetches on demand during the first handshake that needs one, and the cache lives in shared memory (shmcb), so it does not survive restarts. The first clients after every restart get no staple, and every restart resets the clock.
  • Default logging hides it. Not all failure modes emit messages at default log levels. You either probe the handshake externally or know to look for AH01929 and AH02217.
flowchart LR
  subgraph Stapled["Stapling working"]
    C1[Client] -->|TLS handshake| A1[Apache]
    A1 -->|periodic refresh| R[CA OCSP responder]
    A1 -->|cert + staple| C1
  end
  subgraph Broken["Stapling broken"]
    C2[Client] -->|TLS handshake| A2[Apache]
    A2 -->|cert, no staple| C2
    C2 -->|own OCSP lookup, +100s of ms| R2[CA OCSP responder]
  end

The right-hand path is what “silent” means: every client pays the latency tax individually, and none of that traffic touches your logs or metrics.

Common causes

CauseWhat it looks likeFirst thing to check
Stapling not fully configuredSSLUseStapling on set but no SSLStaplingCache, or stapling never enabledConfig: both directives present; SSLStaplingCache is mandatory and has no default
Missing issuer chainAH02217 in error log; no staple served for that vhostWhether SSLCertificateFile includes the intermediate certificates
Stapling cache too smallAH01929 in error log; responses evicted or never storedCache size vs. number of certificates (responses can be up to ~10 KB each)
Egress blocked to OCSP responderStaples expire and never refresh; no staple served after cache lifetimeOutbound connectivity from the server to the responder URL in the cert’s AIA extension
Responder errors propagated to clientsClients see OCSP errors instead of a plain handshakeSSLStaplingReturnResponderErrors still at its default on
Restart cleared the cacheStaples missing for a window after every restart; first handshakes slowUptime vs. when clients stopped receiving staples
Must-Staple certificateHandshakes fail (not just slow) for enforcing clients when no staple is availableCertificate extensions for the RFC 7633 Must-Staple flag

Quick checks

These are all read-only.

# 1. Does the server actually staple? Look for the OCSP response section.
echo | openssl s_client -connect example.com:443 -servername example.com -status 2>/dev/null | grep -A2 "OCSP response"
# Working:  "OCSP Response Status: successful"
# Broken:   "OCSP response: no response sent"

# 2. Grep the error log for the two stapling codes.
grep -E "AH01929|AH02217" /var/log/apache2/error.log | tail -20
# RHEL path: /var/log/httpd/error_log

# 3. Confirm both directives are configured (paths vary by distro).
grep -rn "SSLUseStapling\|SSLStaplingCache" /etc/apache2/ /etc/httpd/ 2>/dev/null

# 4. Extract the OCSP responder URL the certificate points to.
openssl x509 -in /path/to/cert.pem -noout -ocsp_uri

# 5. Test egress from the server to that responder (must be run on the Apache host).
curl -s -o /dev/null -w "%{http_code}\n" --max-time 5 "$(openssl x509 -in /path/to/cert.pem -noout -ocsp_uri)"

# 6. Check whether the served chain includes the issuer certificate.
echo | openssl s_client -connect example.com:443 -servername example.com 2>/dev/null | grep -E "^ *[0-9]+ s:|^ *i:"

Check 1 is the definitive test and the one to automate. Run it per vhost, not just the default: each vhost can have a different certificate and a different stapling outcome. Check 6 matters because AH02217 (issuer not configured) is one of the two log-visible causes.

How to diagnose it

  1. Establish the symptom externally. Run check 1 from a host outside the load balancer, against each HTTPS vhost. “No response sent” on any of them is your confirmation. If staples are present everywhere, the latency is elsewhere; see Apache backend response time: telling ‘Apache is slow’ from ’the backend is slow’.

  2. Correlate with restarts. Compare ServerUptimeSeconds from mod_status with when stapling broke. If stapling fails only in the first minutes after a restart and then recovers, you are seeing the empty shmcb cache plus on-demand fetch, not a permanent misconfiguration.

  3. Classify the cause from the error log. AH01929 means the stapling cache cannot hold the response (cache too small). AH02217 means the issuer chain is not configured for a certificate stapling is trying to serve. No message at all, with staples missing, points at responder reachability or responder-side errors.

  4. Test responder reachability from the Apache host. Check 5 above. OCSP responder URLs are typically plain HTTP from the AIA extension, so an egress proxy or firewall rule that only allows 443 outbound will silently block refresh. This is one of the most common causes and produces no Apache-side error at default log levels.

  5. Verify the chain on disk. For AH02217: since 2.4.8, SSLCertificateChainFile is deprecated and SSLCertificateFile is expected to load the intermediates directly. If the file contains only the leaf, stapling cannot build the validation path and refuses to staple for that vhost.

  6. Check for Must-Staple. If clients report hard handshake failures rather than slowness, inspect the certificate for the Must-Staple extension. With Must-Staple, a missing staple is a connection failure for enforcing clients, which converts this from a latency bug into an outage.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Staple presence per vhost (external probe with -status)The only direct measure of the failure“No response sent” on any vhost beyond a short post-restart window
AH01929 / AH02217 counts in error logThe only server-side log evidence at reachable log levelsAny occurrence
Client-observed TTFB for TLS connectionsStapling failure shows up as handshake-phase latency, not request latencyBaseline shift in connection setup time with no change in %D
%D from access logRequest-time latency stays flat during stapling failure; that flatness is the discriminatorIf TTFB rises but %D does not, suspect the handshake, not the app
Certificate days-to-expiry per vhostRenewal events are when chains change and stapling silently breaksNew cert deployed without chain verification
TLS handshake CPUClients doing their own OCSP checks reconnect more aggressively in some stacksCPU up with no request-rate change

Note the deliberate inclusion of %D as a negative signal. Stapling failure inflates connection setup, which happens before the request, so %D never sees it. Flat request latency plus slow connections is exactly the fingerprint.

Fixes

Complete the configuration

Both directives are required. SSLUseStapling defaults to off, and SSLStaplingCache has no default at all; stapling without a cache is silently non-functional. A minimal working block:

SSLUseStapling on
SSLStaplingCache shmcb:/var/run/ocsp(131072)

The cache path and size differ by distro; size it comfortably above the number of stapled certificates multiplied by the worst-case response size (responses can approach 10 KB each). Too small is not a soft failure: it is exactly what AH01929 reports. A graceful reload applies the change; the cache will be empty either way until the first fetches complete.

Fix the issuer chain (AH02217)

Append the intermediate certificates to the file referenced by SSLCertificateFile (leaf first, then intermediates). Since 2.4.8 this is the supported mechanism; do not resurrect SSLCertificateChainFile. Verify with check 6 that the server now sends the full chain, then re-run check 1.

Stop propagating responder failures to clients

SSLStaplingReturnResponderErrors defaults to on, which means a responder-side error is handed to the client instead of Apache simply falling back to an unstapled handshake. The production tuning guidance in the official SSL how-to recommends turning it off and tightening the refresh behavior:

SSLStaplingReturnResponderErrors off
SSLStaplingResponderTimeout 4
SSLStaplingStandardCacheTimeout 172800
SSLStaplingErrorCacheTimeout 60

The tradeoff: with responder errors suppressed, clients receive no revocation status during responder outages and must decide for themselves whether to soft-fail. For most deployments that is preferable to injecting errors into every handshake. Be aware that SSLStaplingFakeTryLater (default on) synthesizes a “tryLater” response when the responder is unreachable. In the shipped 2.4.x source this synthesis only reaches clients when SSLStaplingReturnResponderErrors is also enabled; with it off, mod_ssl answers the handshake without a staple and clients soft-fail, matching the documented interaction. Test against your actual clients after changing these.

Open egress to the responder

Allow outbound connections from the Apache hosts to the responder URL from check 4. This is infrastructure, not Apache config, which is why it gets missed: the TLS config looks perfect and nothing in Apache logs the refused connection. If egress must go through a proxy, note that mod_ssl’s stapling fetcher has limited proxy support; verify behavior on your version before relying on it.

Handle restarts and renewal deliberately

You cannot make the cache persistent, but you can shrink the impact window: probe each vhost with check 1 immediately after restarts and certificate renewals, and alert on a staple that is still missing after the expected on-demand fetch has had time to complete. If you use Let’s Encrypt or another short-lived cert, fold the chain-completeness check into the renewal hook, because that is when AH02217 regressions happen.

If you run mod_md for certificate management, its own stapling implementation (MDStapling on) is reported to behave better than mod_ssl’s. MDStapling requires Apache HTTP Server 2.4.42 or later with mod_md built in; mod_md is part of the 2.4 tree and needs libcurl available at build time, but no extra compile flag.

Prevention

  • Probe stapling, not just the port. Add an external per-vhost check equivalent to check 1. This is the single highest-value step; every other item is a secondary control.
  • Alert on the two AH codes. AH01929 and AH02217 belong in the same error-log alerting set as AH00484. See Apache error log monitoring: severity levels, AH codes, and what to alert on.
  • Gate renewals on chain completeness. A renewal that swaps in a leaf-only SSLCertificateFile passes configtest and breaks stapling silently.
  • Treat Must-Staple as an availability commitment. Do not request Must-Staple certificates unless stapling is monitored as a first-class signal; with it, every stapling failure is a handshake failure.
  • Track it as a maturity item. Stapling success rate is a Level 4 (“expert”) signal in the Apache monitoring maturity model precisely because most teams only discover it after a latency complaint they cannot reproduce server-side.

How Netdata helps

  • Access-log latency percentiles from %D give you the flat-request-latency baseline that makes the handshake-phase fingerprint visible: connection time degrades while request time does not.
  • Certificate expiry monitoring per endpoint covers the adjacent silent-TLS failure and alerts before renewal-time chain regressions become incidents.
  • Error-log monitoring surfaces AH01929 and AH02217 the moment they appear instead of after a client complains.
  • mod_status collection keeps worker, scoreboard, and uptime context alongside TLS symptoms, so you can confirm in one view that stapling failure is not accompanied by worker or backend problems.
  • Restart correlation via uptime tracking lets you tie a stapling gap directly to the restart or reload that emptied the cache.

Netdata’s Apache HTTP Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Apache HTTP Server monitoring with Netdata

Netdata monitors Apache HTTP Server with per-second metrics from mod_status, pre-built dashboards, and ML-powered anomaly detection. Watch busy versus idle workers and the scoreboard state mix, requests per second, bytes served per second, and request processing duration alongside the rest of your stack, so you catch the worker-exhaustion, slow-backend, and memory incidents in these runbooks before they page anyone.