The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-httpd / apache-httpd-proxy-connection-refused ▌

Operations Guides

Apache proxy connection refused: AH01114 / (111) Connection refused to backend

Your error log is filling with lines like these:

AH00957: HTTP: attempt to connect to 127.0.0.1:8080 (localhost) failed
AH01114: HTTP: failed to make connection to backend: localhost

Clients see 502 or 503 responses. Apache itself is running fine. The failure is on the outbound leg: mod_proxy tried to open a TCP connection to the backend and the kernel told it no.

The important thing about (111)Connection refused is what it rules out. ECONNREFUSED means the backend host’s network stack answered immediately with a TCP RST. The host is reachable; nothing is listening on that port, or something actively rejected the connection. That is a different failure from a backend that is slow, hung, or firewalled, which produces a connect timeout after a long wait and a different error path (AH01075: Error dispatching request to, or a 504 after ProxyTimeout). Learning to separate “fast refusal” from “slow timeout” is the core diagnostic move on this page, and it is the difference between “restart the backend” and “go look at the network.”

What this means

When mod_proxy cannot establish a backend connection, you typically see this sequence in the error log:

AH00957: HTTP: attempt to connect to <host>:<port> (<name>) failed
AH00959: ap_proxy_connect_backend disabling worker for (<name>) for 60s
AH01114: HTTP: failed to make connection to backend: <name>

The AH00959 line matters operationally: after a connect failure, Apache marks that proxy worker as being in an error state and will not use it again for retry seconds (default 60). If the backend comes back in 5 seconds, Apache will still refuse to try it for the rest of the cooldown window. With a single backend behind ProxyPass, that means the site stays down for up to a minute after the backend recovers. With a balancer, requests fail over to healthy members, and the errored member rejoins after its retry expires.

If retry=0 is set, there is no cooldown: every incoming request immediately retries the dead backend, and every failure logs the full AH00957/AH01114 sequence. A busy site with retry=0 and a dead backend can produce a serious error-log flood on top of the outage.

Common causes

CauseWhat it looks likeFirst thing to check
Backend process down or crashedAH01114 on every request, instant refusal, 502sss -ltn on the backend host: is anything on that port?
Wrong host or port in ProxyPassRefusal starts right after a config change or deploycurl -v http://<backend>:<port>/ from the Apache host
IPv4/IPv6 mismatchLog shows attempt to connect to [::1]:<port>Backend listens on 127.0.0.1 but localhost resolves to ::1
SELinux (RHEL/Fedora)AH00957 reports (13)Permission denied instead of (111); some backported proxy handlers use AH02454 for UDS targetsgetenforce, then ausearch -m avc -ts recent
AppArmor (Debian/Ubuntu)Permission denied on outbound connect, no config changejournalctl / audit log for AppArmor denials on the apache2 profile
FD exhaustion in the Apache childAH01114 plus “Too many open files” elsewhere in the log/proc/<pid>/limits and FD count per child
Worker stuck in error-state cooldownRefusals continue after backend is confirmed healthyTime since last AH00959 vs configured retry
Backend listen queue fullIntermittent refusals under load, backend “up”Backend’s own ss -ltn Recv-Q and overflow counters

One more distinction: if the connect attempt hangs for the full connection timeout rather than failing instantly, that is not “connection refused.” That is an unreachable or packet-dropping path (firewall DROP, wrong subnet, dead route), and it belongs to the timeout playbook, not this one.

Quick checks

All read-only. Run from the Apache host unless noted.

# 1. See the actual error sequence and which target is failing
grep -E "AH00957|AH00959|AH01114|AH01075|AH02454" \
  /var/log/httpd/error_log /var/log/apache2/error.log 2>/dev/null | tail -30

# 2. Test the backend connect directly, with timing
#    time_connect near zero + "Connection refused" = fast RST (this page)
#    time_connect == timeout = slow/unreachable path (different problem)
curl -sv --max-time 5 -o /dev/null \
  -w "connect: %{time_connect}s total: %{time_total}s\n" \
  http://<backend-host>:<backend-port>/

# 3. Confirm what Apache is actually resolving and dialing
getent hosts <backend-hostname>   # does localhost return ::1 first?

# 4. Check SELinux state and recent denials (RHEL/Fedora)
getenforce
ausearch -m avc -ts recent | grep -i httpd | tail -10

# 5. Check the FD situation on Apache children
for pid in $(pgrep 'httpd|apache2'); do
  echo -n "PID $pid: "; ls /proc/$pid/fd 2>/dev/null | wc -l
done
grep "Max open files" /proc/$(pgrep -o 'httpd|apache2')/limits

# 6. Count live backend connections from Apache's side
ss -tn state established dport = :<backend-port> | wc -l

# 7. If using a balancer, check member status
curl -s http://localhost/balancer-manager 2>/dev/null | grep -iE 'worker|status'

On the backend host, the two questions that close the case most of the time:

# Is anything listening where Apache expects it?
ss -ltn | grep ':<backend-port>'

# Is the backend process even alive?
pgrep -a -f '<backend-process-name>'

How to diagnose it

Work the path from Apache outward. Each step eliminates one layer.

flowchart TD
  A[AH01114 in error log] --> B{curl to backend from Apache host}
  B -->|instant refused| C{something listening on port?}
  B -->|permission denied| D[SELinux or AppArmor denial]
  B -->|connects fine| E[Apache-side problem: cooldown, FDs, config]
  C -->|no| F[backend down or wrong port]
  C -->|yes, but different IP family| G[IPv4/IPv6 mismatch]
  E --> H{AH00959 within retry window?}
  H -->|yes| I[wait or reduce retry]
  H -->|no| J[check FD limits and proxy pool]
  1. Confirm fast refusal vs slow timeout. Run check 2 above. time_connect returning near-zero with “Connection refused” confirms this page. A time_connect equal to your --max-time means the SYN went nowhere; investigate routing and firewalls instead.

  2. Read the target in the log line. AH00957 prints the exact host and port Apache dialed. Verify it matches what you intended in ProxyPass or the BalancerMember line. A surprising port or hostname here means a config or DNS problem, not a backend problem.

  3. Check for IPv6. If the log shows [::1]:<port> and the backend binds only 127.0.0.1, that is the whole incident. localhost resolving to ::1 first is common on default installs.

  4. Check for permission denied, not refused. If the proxy connect log says (13)Permission denied, the backend may be perfectly healthy. AH02454 is a Unix-domain-socket connect signal in some backported proxy handlers; current upstream commonly logs AH00957 for the failed connect, so check the exact package before relying on either ID. On enforcing SELinux systems, httpd is not allowed to make arbitrary outbound TCP connections by default. Confirm with ausearch before changing anything.

  5. Rule out the error-state cooldown. If the backend is now healthy but requests still fail, look at the timestamp of the last AH00959 and compare it to retry. During the cooldown Apache will not even attempt the connection. The balancer-manager page, if enabled, shows the member in error state.

  6. Check Apache’s own resources. Per-child FD exhaustion produces AH01114 because the child literally cannot open a new socket. Look for “Too many open files” in the error log; the playbook’s FD exhaustion pattern lists proxy connection failures (AH01114) as a primary symptom.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
AH01114/AH00957 rate in error logDirect count of backend connect failuresAny sustained non-zero rate
502 vs 503 vs 504 split in access log502/refused means dead or unreachable backend; 503 can mean pool exhaustion or all members errored; 504 means slow, not dead502s without corresponding backend restarts
AH00959 “disabling worker” eventsTracks how often workers enter error-state cooldown and for how longRepeating disable/enable flapping
Backend connect time (curl -w '%{time_connect}' probe)Separates “down” (instant RST) from “slow” (long connect or response)Connect time drifting from ~0 toward timeout
BusyWorkers / scoreboard W statesDead backends fail fast; slow backends hold workers. Refused connections do not exhaust workers, but the resulting retry storms and client refreshes canWorkers climbing while 502s flow
Per-child FD count vs limitFD exhaustion manifests as AH01114Any child above ~70% of its limit
Balancer member statusWhich members are in error state right nowAny member errored with traffic active

The down-versus-slow distinction deserves emphasis because it changes the response. A dead backend fails every connect in microseconds; workers are freed immediately and the damage is “only” the 502s. A slow backend holds workers in W state for seconds each, which is the slow-backend-cascade pattern that takes the whole server down. Connect time plus scoreboard state tells you which one you have within a minute.

Fixes

Backend process down

Restore the listener, then find out why it died. Check the backend’s own logs and dmesg for OOM kills. If this is a recurring crash, restarting it without root cause work just schedules the next AH01114 flood. Note the cooldown: even after the backend is up, Apache may refuse to use it until retry expires, or with a balancer you can reset the member from balancer-manager.

Wrong target in ProxyPass

Fix the host or port in ProxyPass / BalancerMember, run apachectl configtest, then apachectl graceful. If the error started right after a deploy or config-management run, diff the current config against the previous revision before touching anything else.

IPv4/IPv6 mismatch

Use 127.0.0.1 explicitly in the ProxyPass target instead of localhost, or make the backend listen on both families. The config fix is one character of intent and removes the dependency on resolver order entirely.

SELinux denial (RHEL/CentOS/Fedora)

The canonical fix for httpd making outbound proxy connections:

# Allow httpd to make network connections (persistent)
setsebool -P httpd_can_network_connect 1

Verify with getsebool httpd_can_network_connect. Do not disable SELinux or set it permissive to fix this; the boolean exists precisely for the reverse-proxy use case.

AppArmor (Debian/Ubuntu)

Less commonly hit, but AppArmor profiles can deny outbound connects the same way. Check the audit log for denials against the apache2 profile and adjust the profile rather than removing it.

Error-state and retry tuning

retry on ProxyPass or BalancerMember controls the cooldown after a connect failure (default 60 seconds). Tradeoffs:

  • Lower retry (e.g. 5-10s): faster recovery after a brief backend bounce, at the cost of more failed attempts against a genuinely dead backend.
  • retry=0: always retry immediately. Reasonable behind a balancer with multiple members; dangerous with a single backend because every request retries and logs the full failure sequence.
  • Balancers: failonstatus lets you push a member into error state when it returns specific HTTP codes, and forcerecovery (2.4.2+) forces immediate recovery of all members if every member is errored, ignoring retry. Both are useful for backends that fail “up” (accepting connections but returning garbage).

Do not confuse the balancer-level timeout (maximum wait for a free member) with the per-member connection timeout; they are different knobs on different lines. connectiontimeout on the member controls how long Apache waits for the TCP connect to complete, and ProxyTimeout (default: the global Timeout, usually 60s) bounds waits on established backend connections. None of these cause ECONNREFUSED, but mis-set values change how the failure presents.

FD exhaustion

If AH01114 arrives alongside “Too many open files,” raise the per-process limit in the systemd unit (LimitNOFILE=65536) and reload. Sizing rule from the playbook: each client connection, each backend proxy connection, and each log file costs an FD, so the limit should cover roughly MaxRequestWorkers x 2 plus log and static overhead. This requires a restart, not a graceful reload.

Prevention

  • Health-probe the backend path, not just Apache. A localhost check against a static file will pass while every proxied request 502s. Probe through the proxy to the backend’s health endpoint.
  • Alert on the error-log sequence, not just 5xx rate. AH00957/AH01114 rate is a cleaner, earlier signal than access-log 502 percentages, and AH00959 tells you cooldowns are happening.
  • Pin the backend address family. Use explicit IPs in proxy targets so resolver behavior cannot change underneath you.
  • Set the SELinux boolean at provisioning time. httpd_can_network_connect should be part of the base image or config management for any host running httpd as a reverse proxy.
  • Size FD limits for the proxy case. Doubling connections per worker (client plus backend) is the common way FD limits that “were fine” suddenly are not.
  • Choose retry deliberately. Default 60s is safe; know that it extends every backend blip into a minute-long outage window for single-backend setups, and that retry=0 trades that for log volume and hammering a dead service.

How Netdata helps

  • Netdata’s Apache collector polls server-status every second, so BusyWorkers, idle workers, and request rate are visible at the granularity where a retry storm or client-refresh pile-up actually unfolds.
  • The web log collector parses access logs into per-status-code rates, so the 502/503/504 split is a first-class chart rather than an awk pipeline you run during the incident.
  • Correlating 5xx-by-code against worker utilization separates “backend dead” (errors up, workers flat) from “backend slow” (errors up, workers saturated in W) at a glance.
  • System-level charts (per-process FD usage, TCP connection states, network errors) catch the FD-exhaustion and refused-connection variants without per-host log diving.
  • Anomaly detection on error-log rates surfaces the first AH01114 burst rather than the thousandth, which is usually the difference between a ticket and a page.

Netdata’s Apache HTTP Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Apache HTTP Server monitoring with Netdata

Netdata monitors Apache HTTP Server with per-second metrics from mod_status, pre-built dashboards, and ML-powered anomaly detection. Watch busy versus idle workers and the scoreboard state mix, requests per second, bytes served per second, and request processing duration alongside the rest of your stack, so you catch the worker-exhaustion, slow-backend, and memory incidents in these runbooks before they page anyone.