The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-httpd / apache-httpd-error-log-monitoring ▌

Operations Guides

Apache error log monitoring: severity levels, AH codes, and what to alert on

Metrics tell you that something is wrong with Apache. The error log tells you what. When BusyWorkers pegs at MaxRequestWorkers, the scoreboard shows saturation, but the error log hands you the MPM worker-limit code (for example AH00484) with “server reached MaxRequestWorkers setting.” When a child dies, the process table shows a respawn; the error log shows “Segmentation fault.”

This guide covers how to read the log efficiently: the severity hierarchy, the AH code scheme introduced in 2.4, the patterns worth alerting on, and the configuration gotchas that silently suppress the messages you need.

How Apache decides what to log

Every error log line carries a severity. LogLevel sets a floor: messages at the configured level and everything more severe are written, everything below is dropped.

LevelMeaningTypical content
emergApache cannot run at allFatal startup failures, cannot initialize
alertImmediate action requiredSevere runtime failures
critCritical conditionsFailed socket operations, resource failures
errorError conditionsRequest processing failures, proxy errors, permission denied
warnWarningsRecoverable problems, deprecated behavior
noticeNormal but significantStartup, shutdown, “resuming normal operations”
infoInformationalRequest-level detail, 404s in 2.4
debugDebug outputPer-module internals, very high volume

The default is warn. First, startup/shutdown [notice] messages such as “resuming normal operations” and “caught SIGTERM” are below that floor but Apache preserves them in regular-file logs; this is useful for restart forensics and easy to forget. Second, anything a module logs at info is invisible by default, which matters more than most operators realize (see the gotchas below).

In 2.4 there are also trace1 through trace8 levels below debug.

The practical reading rule: [error] is the working level for request and module failures, [crit] and above should be near zero on a healthy server, and any [emerg] or [alert] is a stop-and-look event. As a rough baseline, a sustained rate above 10 [error] lines per minute is abnormal for most deployments; baseline your own.

The AH code scheme

Apache 2.4 tags nearly every log message with a unique code of the form AHxxxxx, for example AH00484. Each code maps to one specific message at one specific location in the source tree. This matters operationally for two reasons:

  • Codes are stable greppable anchors. Message text varies with arguments (paths, PIDs, client addresses); the code does not. Alert rules should match the code, not the full message.
  • Codes disambiguate similar messages. 503 in the access log can mean worker exhaustion or a balancer member in error state. The error log distinguishes them because they carry different codes.

Apache 2.2 logs have no codes; the format there was fixed and message matching is the only option. The error log line format became customizable in 2.4 via ErrorLogFormat, so on a mixed fleet do not assume your parsing regexes port across versions.

What to alert on

Not everything in the error log is signal. Crawler 404s and client-side noise dominate volume; the codes below are the ones that correlate with real production failure modes.

PatternWhat it meansSeverity
AH00484 / AH00286 / AH00161 MaxRequestWorkers reachedWorker pool exhausted; new connections are queuing in the listen backlogPage corroborator; ticket at minimum
Segmentation fault / child exit signalChild process crash, usually a module bugTicket always; multiple per minute, escalate
No space left on deviceLog or work filesystem full; logging and some operations will failPage
Too many open files (EMFILE)Per-process FD limit hit; accepts, file opens, and proxy connects failPage when sustained with traffic
AH01114 proxy connection failureBackend refused or unreachableTicket
AH01136 reverse proxy worker busyPer-backend proxy connection pool exhaustedTicket; page if user-facing
AH00898 bad status line from backendBackend returned a malformed response (502 territory)Ticket
AH01075 error dispatching requestProxy could not hand off the request to the backendTicket
AH02429 response header too longBackend sending headers Apache will not proxyTicket
AH01929, AH02217OCSP stapling problemsTicket; silently degrades TLS latency
“resuming normal operations” clusterRestart events; many close together means restart pile-upTicket

Two nuances on the headliners:

  • AH00484 fires once per exhaustion event, not per rejected request. A single line can represent thousands of queued connections. Never alert on “more than N occurrences” of it; any occurrence during production traffic is worth investigating. Corroborate with BusyWorkers, listen queue depth, and 503s before paging. See Apache AH00484: server reached MaxRequestWorkers setting.
  • Segfaults on threaded MPMs are worse than on prefork. On prefork a segfault kills one child. On worker or event, a crashing thread can take the entire process and all of its threads with it. The log line alone does not tell you which; check the scoreboard and process count around the event.

Some [error] lines are client-side noise: broken pipes when clients disconnect, “File does not exist” for scanner probes. Baseline these and filter them from alert rules, or you will train the team to ignore the log.

LogLevel gotchas that hide real incidents

Setting the level too high. A LogLevel of crit or alert feels like noise reduction. It also hides every proxy failure, every permission problem, and every module error, because those log at [error]. If your error log looks suspiciously clean during an incident, check the level first.

Per-module levels are the right tool for targeted verbosity. Instead of raising the global level to debug and drowning, raise one module: LogLevel warn ssl:debug keeps the floor at warn but turns on debug for mod_ssl. Conversely, LogLevel ssl:warn quiets one noisy module without touching the rest.

2.4 moved 404s to info. In 2.4, “File does not exist” messages were downgraded from error to info. With the default warn floor they vanish entirely. This has silently broken fail2ban rules and security monitoring that scraped for 404s after upgrades from 2.2. To restore them for the core module only: LogLevel warn core:info.

Debug volume is explosive. A single request at LogLevel debug generates many lines per module hook. Use per-module scoping, and never leave global debug on in production. Besides the noise, every one of those lines is a synchronous write (next section).

Virtual hosts can have their own error logs. A per-VHost ErrorLog directive redirects that vhost’s messages away from the main log. If you alert only on the main log, vhost-scoped failures are invisible. Either aggregate all vhost logs into your collection pipeline or consolidate them.

Failure modes of the logging pipeline itself

The error log is on the request hot path, and it has its own failure modes.

flowchart TD
  E[Event in a worker] --> W{Write path}
  W --> F[ErrorLog file - synchronous write]
  W --> P[Piped log program]
  W --> V[Per-VHost error log]
  F -->|disk slow or full| L[Worker blocks in scoreboard L state]
  P -->|program dies| S[AH00106 - reliable logger restarted]
  K[Kernel OOM killer] -->|never reaches Apache| D[dmesg / kernel log only]

Synchronous writes block workers. File-based logging is synchronous: the worker blocks until the write completes. On slow or saturated storage, workers accumulate in the L (Logging) scoreboard state, which is indistinguishable from worker exhaustion from the outside. A sustained L count above a few percent is a storage or log-pipeline problem, not a traffic problem.

Piped logs can fail silently. ErrorLog "|/path/to/program" hands log lines to a child process. If that program dies, upstream Apache restarts reliable piped-loggers and logs AH00106; if writes stall or are redirected elsewhere, workers can block or lose log data. Do not assume a pipe failure appears only as a child SIGPIPE crash. Piped logging with rotatelogs avoids copytruncate races and unbounded file growth, but you trade one failure mode for another: monitor the pipe process itself.

A full log disk is a service outage, not a logging outage. When the log filesystem fills, workers finish requests but cannot log them and hang in L state. The port stays open; the server serves nothing. The signature: scoreboard dominated by L, disk at 100%, error log stops updating. Keep logs on their own filesystem and alert on it before it fills.

Some fatal events never appear here. When the kernel OOM killer terminates httpd children, nothing is written to the Apache error log, because the process was killed externally. The evidence is in dmesg or the journal: dmesg | grep -i oom. If you see unexplained child respawns with a clean error log, check the kernel log before assuming an Apache bug. The same applies to nf_conntrack table saturation and SELinux denials, which surface here as generic “Permission denied” or dropped connections but are explained only in system logs.

Quick reference: reading the log during an incident

# Errors and above, most recent first
grep -E "\[error\]|\[crit\]|\[alert\]|\[emerg\]" /var/log/apache2/error.log | tail -20
# RHEL path: /var/log/httpd/error_log

# Worker exhaustion events
grep -E "AH00484|AH00286|AH00161" /var/log/apache2/error.log | tail -20

# Child crashes (check both Apache and kernel logs)
grep -i "segfault\|segmentation" /var/log/apache2/error.log | tail -20
dmesg | grep -i "segfault.*apache\|segfault.*httpd" | tail -20

# Restart history: frequency and type
grep -E "resuming normal operations|caught SIGTERM|graceful restart" /var/log/apache2/error.log | tail -20

# The things the error log will never tell you
dmesg | grep -i oom | tail -20

Pair each pattern with its metric corroborator: MPM worker-limit codes (AH00484/AH00286/AH00161) with BusyWorkers/IdleWorkers and listen queue depth, segfaults with child process churn, L-state buildup with disk space and I/O latency on the log filesystem, proxy AH codes with 502/503/504 rates in the access log. The error log gives you the cause; the metrics tell you the blast radius.

How Netdata helps

  • Error-level log rate as a first-class signal. Netdata can track parsed log lines by custom fields and pattern matches, including an error-severity field from a custom Apache error-log parser; configure that explicitly instead of assuming the standard access-log job reads the error log.
  • AH00484 correlated with saturation metrics. A MaxRequestWorkers event next to BusyWorkers, listen queue depth, and 503 rate on one dashboard is the difference between “page now” and “watch it.”
  • Scoreboard L state next to disk metrics. Workers blocking on log writes show up as Logging-state growth alongside log filesystem utilization and I/O latency, which pinpoints a log stall without log diving.
  • Restart and crash context. Uptime resets and child churn correlated with system-level events (including OOM kills from the kernel side) close the gap where the error log itself is silent.
  • Per-second granularity. Error log bursts during deploys and graceful-restart pile-ups are visible at the resolution they actually happen at, rather than averaged away.

Netdata’s Apache HTTP Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Apache HTTP Server monitoring with Netdata

Netdata monitors Apache HTTP Server with per-second metrics from mod_status, pre-built dashboards, and ML-powered anomaly detection. Watch busy versus idle workers and the scoreboard state mix, requests per second, bytes served per second, and request processing duration alongside the rest of your stack, so you catch the worker-exhaustion, slow-backend, and memory incidents in these runbooks before they page anyone.