The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-httpd / apache-httpd-ulimit-nofile ▌

Operations Guides

Apache open-file limits: raising ulimit and systemd LimitNOFILE correctly

The default per-process open-file limit on most Linux distributions is 1024, far too low for production Apache. Every client connection, every backend proxy connection, every log file, and every pipe consumes a file descriptor in every child process, and 1024 runs out long before MaxRequestWorkers does. When a child hits the limit, the failure is a cliff edge: “Too many open files” in the error log, failed accepts, failed backend connections, and intermittent 5xx responses, often in only some children at first, which makes the symptom look random.

The confusion comes from the limit not living in one place. Operators edit /etc/security/limits.conf, run ulimit -n in a shell, or tweak Apache configuration, and none of it sticks because the running httpd is managed by systemd and systemd sets the limit itself. This article covers where the limit actually comes from, how to raise it, why it requires a full restart, how to size it, and the adjacent limits that still bite after LimitNOFILE is fixed.

Where the limit actually comes from

For a service-managed httpd, the limit is set by systemd at process start, not by login configuration and not by Apache itself.

  • systemd unit LimitNOFILE=: The authoritative source for httpd started by systemd. Whatever is in the unit file or in drop-in overrides under /etc/systemd/system/httpd.service.d/ (RHEL family) or /etc/systemd/system/apache2.service.d/ (Debian family) wins.
  • /etc/security/limits.conf: Applied by PAM at login time for interactive sessions. It does not apply to services started by systemd. Editing it and restarting httpd changes nothing, which is the most common dead end in this investigation.
  • ulimit -n in a shell or init script: Only affects processes started from that shell. On SysV-era systems this was the mechanism (setting ulimit in /etc/sysconfig/httpd or the init script). On systemd systems it is irrelevant for the service.
  • Apache configuration: There is no Apache directive that raises the FD limit. Apache inherits whatever the process was started with.
flowchart TD
  A[httpd start requested] --> B{Started by systemd?}
  B -- yes --> C[LimitNOFILE from unit file]
  C --> D[Drop-in overrides in .d directory applied last]
  D --> E[Resulting limit visible in /proc/PID/limits]
  B -- no: shell or init script --> F[ulimit -n of the starting shell]
  G[/etc/security/limits.conf] -. ignored for systemd services .-> C

The definitive check is always against the running process, not against any config file:

# Check the effective FD limit of the running Apache parent
cat /proc/$(pgrep -o 'httpd|apache2')/limits | grep "Max open files"

If the value there does not match what you configured, you configured it in the wrong place.

Why raising the limit needs a full restart

systemctl reload httpd and apachectl graceful send SIGUSR1. The existing parent process stays alive and spawns new children under the same inherited limits. systemd applies LimitNOFILE when it starts a process, and a graceful reload never goes back through systemd’s process creation path, so the new value is never picked up.

The only way to apply a changed LimitNOFILE is a full stop and start of the unit. That is disruptive: it drops all in-flight connections. Plan it for a low-traffic window or drain the node from the load balancer first.

# Apply a new FD limit (DISRUPTIVE: drops all connections)
systemctl daemon-reload
systemctl restart httpd    # or apache2 on Debian/Ubuntu

Do not let log rotation or config management trigger hard restarts casually; sending the parent SIGHUP is also a hard restart and drops all connections.

Sizing the limit

The limit is per child process, and each child holds FDs for everything it touches. A sizing formula that covers production reality:

Required FDs per child = (concurrent connections per child x 2 if proxying) + one FD per distinct log file + internal overhead

  • Client connections: one FD each. On worker/event MPM, concurrent connections per child tracks ThreadsPerChild plus async keepalive connections on event.
  • Proxying: each backend connection is another FD. In reverse-proxy deployments, count client connection plus backend connection per in-flight request, hence the x2.
  • Log files: every distinct CustomLog and ErrorLog target is held open by every child. With many VirtualHosts each writing separate access and error logs, this term alone can exceed 1024 before a single client connects. Hundreds of vhosts times two log files each means hundreds of FDs per child as a floor.
  • Overhead: pipes, shared memory, SSL session cache, module internals. Leave room.

Then apply headroom: set the limit to at least 2x the theoretical maximum FD usage. FD exhaustion is a cliff-edge failure with no graceful degradation, and you want headroom for events like log rotation briefly opening new files. For most production deployments LimitNOFILE=65536 is a reasonable floor; vhost-heavy or high-concurrency proxy deployments may justify more.

Also check the system-wide ceiling: /proc/sys/fs/file-nr shows allocated, free, and maximum. fs.file-max must accommodate all httpd children plus everything else on the box. Per-process limits are the usual binding constraint, but on shared hosts the system limit can be hit first.

If the math surprises you, the fix is often not a bigger number. Consolidating vhost logs into a single log with the %v format directive (split later with split-logfile) collapses the per-child log FD count dramatically and is usually the right structural fix.

Procedure

This assumes systemd, which covers RHEL 7+ and derivatives, Debian 8+, and Ubuntu 16.04+. Adjust the service name: httpd on RHEL family, apache2 on Debian family.

  1. Measure first. Record current FD usage per child so you have a baseline and can validate your sizing:
# FD count and limit per Apache process
for pid in $(pgrep 'httpd|apache2'); do
  count=$(ls /proc/$pid/fd 2>/dev/null | wc -l)
  limit=$(awk '/Max open files/{print $4}' /proc/$pid/limits)
  echo "PID $pid: $count / $limit"
done
  1. Create a drop-in override. Do not edit the vendor unit file; package updates will overwrite it.
# Create the drop-in directory and file
mkdir -p /etc/systemd/system/httpd.service.d
cat > /etc/systemd/system/httpd.service.d/limits.conf <<'EOF'
[Service]
LimitNOFILE=65536
EOF

LimitNOFILE also accepts soft:hard syntax (e.g., LimitNOFILE=32768:65536) if you want a lower soft limit with a hard ceiling. A single value sets both.

  1. Reload systemd and fully restart the service. Disruptive: drains connections.
systemctl daemon-reload
systemctl restart httpd
  1. Verify against the running process (see next section). Do not skip this; a typo in the drop-in silently leaves you at the old limit.

Verifying it worked

# Confirm the running parent picked up the new limit
cat /proc/$(pgrep -o 'httpd|apache2')/limits | grep "Max open files"

# Spot-check children too; children inherit the parent's limits
for pid in $(pgrep 'httpd|apache2'); do
  awk -v p=$pid '/Max open files/{print p, $4, $5}' /proc/$pid/limits
done

You should see the new value in the “Max open files” row for parent and children. If it still shows 1024 or 4096, the drop-in was not read: check the path, the service name, that you ran daemon-reload, and that you did a full restart rather than reload.

Common pitfalls

  • Editing /etc/security/limits.conf and expecting it to apply. It does not, for systemd services. This is the single most common wasted hour in this task.
  • Running systemctl reload and assuming the new limit is live. The running parent keeps the old limit; only a full restart applies it. Worse, cat /proc/.../limits is the only honest check, because systemctl show httpd -p LimitNOFILE reports what systemd would apply at next start, not what the running process has.
  • TasksMax capping you from the other direction. systemd also imposes a process-and-thread limit via the cgroup pids controller. If TasksMax is lower than what MaxRequestWorkers implies (processes plus threads), Apache cannot spawn the children you configured, and the kernel logs fork rejections from the pids controller. This looks nothing like an FD problem but shows up in the same investigation. Raise TasksMax= in the same drop-in if your worker math requires it. The default changed across systemd releases: 512 for services on older systemd (the v228-era DefaultTasksMax), then 15% of the system’s PIDs limit on newer releases (systemd 238 and later; check the effective value with systemctl show httpd -p TasksMax).
  • Distro wrapper scripts that set ulimit themselves. Debian’s apache2ctl honors an APACHE_ULIMIT_MAX_FILES environment variable and runs ulimit -n before starting Apache. It is still present in current Debian/Ubuntu packaging (the default is ulimit -n 8192), and Debian’s apache2.service actually starts Apache through apachectl, so on Debian both mechanisms genuinely interact: a ulimit -n lower than the hard limit lowers it, and one higher fails with a warning. Pick one mechanism, preferably the systemd drop-in, and remove the other.
  • Raising the limit to mask an FD leak. If per-child FD count grows monotonically over time, raising LimitNOFILE just delays the cliff. Common leak sources: backend connections in CLOSE_WAIT that never get reaped (often a backend that does not close connections properly), proxy pool connections not returned, or old log handles held across rotation. Track per-child FD counts over days before and after the change.
  • select() and FDs above 1023. The systemd documentation warns that select(2) cannot handle file descriptors above 1023 on Linux. Apache’s worker and event MPMs use poll/epoll internally and are not affected, but a third-party module or CGI that calls select() can misbehave once a child holds more than 1023 FDs. If you run exotic modules, this is worth knowing; it is not a reason to keep the limit at 1024.
  • mod_proxy CLOSE_WAIT accumulation. If backends keep sockets half-closed, FDs accumulate in CLOSE_WAIT inside Apache children. Raising the limit buys time; the real fix is proxy timeout tuning or fixing the backend’s connection handling.

Signals to monitor

SignalWhy it mattersWarning sign
Per-child FD count (/proc/[pid]/fd) vs limitDirect measure of how close each child is to the cliffAny child above 70% of its limit
Effective limit (/proc/[pid]/limits)Confirms configuration actually appliedValue lower than what you configured after a restart
“Too many open files” in error logDefinitive exhaustion signal; failure already happeningAny occurrence; PAGE if repeated with active request failures
Proxy connection failures (AH01114)FD exhaustion in children breaks backend connects firstAppearing alongside “Too many open files”
CLOSE_WAIT count on Apache socketsBackend connections not reaped; FD leak in progressPersistent non-zero CLOSE_WAIT growing over time
System-wide fs.file-nr vs fs.file-maxThe ceiling above the per-process limitAllocated approaching maximum
Number of vhosts x separate log filesThe structural driver of per-child FD floorLog FDs alone approaching the limit

How Netdata helps

  • Netdata charts per-process file descriptor counts against their limits, so you see each Apache child’s FD usage trending toward the ceiling long before “Too many open files” appears.
  • Correlating FD growth with connection-state charts (especially CLOSE_WAIT) separates a leak from legitimate load growth, which decides whether you raise the limit or fix the backend.
  • Error log pattern monitoring catches “Too many open files” and AH01114 as they happen, instead of after users report intermittent 5xx.
  • Restart and uptime context on the same dashboard makes it obvious whether a limit change actually took effect (new limit visible after a full restart) or was silently not applied (graceful reload only).
  • System-wide fs.file-nr alongside per-process counts shows when the box-level ceiling, not the per-process limit, is the binding constraint.

Netdata’s Apache HTTP Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Apache HTTP Server monitoring with Netdata

Netdata monitors Apache HTTP Server with per-second metrics from mod_status, pre-built dashboards, and ML-powered anomaly detection. Watch busy versus idle workers and the scoreboard state mix, requests per second, bytes served per second, and request processing duration alongside the rest of your stack, so you catch the worker-exhaustion, slow-backend, and memory incidents in these runbooks before they page anyone.