The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / apache-httpd / apache-httpd-somaxconn-listenbacklog ▌

Operations Guides

Apache ListenBacklog vs net.core.somaxconn: the silently truncated accept queue

You tuned Apache for burst absorption. You set ListenBacklog 2048 (or left the default 511, reasoning it was generous). Then a traffic spike arrived, workers saturated, and connections were refused far earlier than your capacity model predicted. The scoreboard and MaxRequestWorkers get the blame, but the real culprit is often one layer down: the kernel quietly rewrote your backlog at listen() time and never told anyone.

Apache passes ListenBacklog as the backlog argument to listen(2). Per the listen(2) man page, if the backlog argument is greater than the value in /proc/sys/net/core/somaxconn, the kernel silently caps it to that value. No log line, no warning from Apache, no error return. The effective maximum accept queue on any listening socket is always min(ListenBacklog, net.core.somaxconn).

What follows: the mechanism, how to read the truncated backlog from ss, how to raise both values together, and why the default 511 is too small on busy servers even when Recv-Q reads zero most of the time. For where the accept queue sits in Apache’s wider saturation model, see How Apache HTTPD actually works in production.

What the accept queue is and why it matters

When a client completes the TCP three-way handshake with Apache’s listening socket, the connection lands in the kernel’s accept queue. It stays there until an Apache worker calls accept() and takes it. This queue is the last buffer before service denial: while it has room, a momentarily saturated Apache still absorbs connection bursts; when it fills, the kernel starts dropping or resetting incoming connections and clients see connection refused or timeouts.

The queue exists because accept rate and arrival rate are never perfectly matched. Workers are busy finishing requests, children are spawning, a graceful restart just happened. A deep enough queue smooths those gaps. A shallow one turns every transient stall into refused connections, and a load balancer watching TCP connect behavior may pull the server from rotation before Apache has logged anything at all.

Two failure modes follow:

  1. Queue overflow under saturation. All workers busy, queue fills, SYNs get dropped or RST. The server appears up (the port is open, the process runs) but is unreachable.
  2. Silent truncation at configuration time. You believe the queue holds 511 or 2048 connections. It actually holds 128, because net.core.somaxconn on your kernel defaults lower than your ListenBacklog, and the kernel never mentioned it.

How the two limits interact

Apache httpd 2.4 defaults ListenBacklog to 511; current 2.4 source accepts any positive integer and rejects only values below 1.

The kernel side has its own default, and it changed at a version boundary that still matters:

  • Linux before 5.4: net.core.somaxconn defaults to 128.
  • Linux 5.4 and later: defaults to 4096.

The practical consequences by platform:

PlatformKernelsomaxconn defaultEffective backlog with ListenBacklog 511
RHEL 7 / CentOS 73.10128128 (truncated)
RHEL 8 base 4.18 kernelscommonly 4.18commonly 128128 (truncated)
RHEL 95.144096511
Ubuntu 20.04+ / Debian 11+5.4+4096511

The RHEL 8 row describes the base kernel value; distribution backports or images may change it. Always read /proc/sys/net/core/somaxconn instead of inferring it from the kernel version.

So on a stock kernel below 5.4 with somaxconn=128, Apache’s default 511 is silently truncated to 128. The server you think absorbs a 511-connection burst actually absorbs 128. On kernel 5.4 and later, no truncation happens by default, but the moment you raise ListenBacklog above 4096 the same trap reappears.

One more layer: tuning tooling can change somaxconn out from under you. The throughput-performance profile in RHEL’s tuned service sets net.core.somaxconn to 2048. Set ListenBacklog 4096 on such a system without checking and you get 2048.

flowchart TD
  A[ListenBacklog in Apache config
default 511] --> C[listen backlog argument] B[net.core.somaxconn
128 pre-5.4, 4096 since 5.4] --> C C --> D{Kernel compares at listen time} D -->|backlog greater than somaxconn| E[Silently capped
no warning anywhere] D -->|backlog within somaxconn| F[Used as configured] E --> G[Effective accept queue
= min of the two] F --> G G --> H{Queue full under load?} H -->|yes| I[Connections dropped or RST
clients see refused or timeout]

Reading the effective backlog with ss

You cannot detect truncation from Apache. Apache logs nothing about it, and apachectl configtest validates syntax, not kernel behavior. The only reliable check is to ask the kernel what the listening socket actually got.

# Show current and maximum accept queue for Apache listeners
ss -ltn | grep -E ':80\s|:443\s'

For LISTEN sockets, the column meanings are inverted from what most people expect:

  • Recv-Q: current number of connections sitting in the accept queue, waiting for accept().
  • Send-Q: the maximum backlog the socket actually has. This is the effective, post-truncation value.

If your config says ListenBacklog 511 and ss shows Send-Q of 128, the kernel capped you. Compare Send-Q against your configured value on every Apache host; a mismatch is the truncation, and the fix is on the kernel side.

For a live view of queue depth with more detail:

# Detailed view; current queue depth appears as unacked
ss -lti sport = :80

The current accept queue depth shows up as unacked:N in this output, not in a field named “backlog”. The kernel maps the accept queue counter to the unacked field for listening sockets, which confuses people who go looking for a backlog label.

Finally, check whether connections are already being dropped:

# Listen overflow counters
nstat -a | grep -i listen
netstat -s | grep -i "listen"

Rising ListenOverflows (and ListenDrops) means the queue has already filled and the kernel has already refused connections. That is the trailing indicator; by the time it moves, users felt it.

Raising both values in tandem

The rule: never tune one side without the other, and always verify with ss afterward.

  1. Decide the target queue depth. Base it on burst absorption, not steady state. How many connections can arrive during the worst-case window where all workers are busy: a slow backend stall of a few seconds, a graceful restart, a TLS handshake burst? For most busy servers, 1024 to 4096 is a reasonable range. There is no portable upper-bound guarantee; set a value the kernel accepts and verify the resulting socket Send-Q.

  2. Set the kernel side first.

# Apply immediately
sysctl -w net.core.somaxconn=4096

# Persist across reboot
echo 'net.core.somaxconn = 4096' > /etc/sysctl.d/90-apache-backlog.conf
sysctl --system
  1. Set the Apache side. Add or adjust the directive in the server config:
ListenBacklog 4096

The constraint is simply that the configured value must be greater than zero; the kernel may cap it.

  1. Restart Apache fully. This step is easy to get wrong. A graceful reload is not enough to apply a new backlog to existing listening sockets; httpd must be restarted after the sysctl change for the new backlog to take effect. A hard restart drops active connections, so do it during a low-traffic window or behind a load balancer with connection draining. Run apachectl configtest first, then restart.

  2. Verify.

# Confirm the socket got what you configured
ss -ltn | grep -E ':80\s|:443\s'

Send-Q should now equal your ListenBacklog. If it shows the old value, either the restart did not happen or something re-applied a lower somaxconn (check for competing sysctl files and tuned profiles).

Containers and Kubernetes

Changing somaxconn on the host does not affect running containers. Each container gets its own network namespace, and the namespace initializes somaxconn from the kernel’s build-time constant (128 pre-5.4, 4096 since 5.4), not from the host’s current value. If Apache runs in a container, set the sysctl inside the container’s namespace:

# Docker: set the sysctl for the container's net namespace
docker run --sysctl net.core.somaxconn=4096 ...

On Kubernetes, use securityContext.sysctls on the pod. net.core.somaxconn is classified as an unsafe sysctl, so the kubelet must allow it, for example with --allowed-unsafe-sysctls=net.core.somaxconn. If you skip this, the pod sets ListenBacklog into a namespace capped at its default and you are back to silent truncation.

Why 511 is too low on busy servers even when Recv-Q reads zero

Operators look at ss, see Recv-Q at zero, and conclude the backlog is fine. That confuses the snapshot with the risk. Recv-Q is a point-in-time sample of a queue that fluctuates in milliseconds; a zero reading only means the queue was empty at that instant. It says nothing about the burst you have not had yet.

The queue exists for the bad windows, not the good ones:

  • Worker saturation events. When all workers are busy (a slow backend holding threads, a Slowloris wave, a genuine traffic spike), new connections stack in the accept queue. The queue is the only buffer; after it, connections are refused. On the worker pool’s cliff-edge degradation curve, queue depth is the difference between “degraded for a few seconds” and “load balancer pulled the node”.
  • Graceful restarts. Old children drain while new ones spawn. Accept rate dips. A shallow queue turns a routine reload into refused connections.
  • Cold start bursts. After a crash restart, SSL session caches are empty and full TLS handshakes spike CPU. Accept slows exactly when reconnecting clients hammer the listener.

A server doing thousands of connections per second with a 511-deep queue has sub-second burst absorption. The operational guidance: Recv-Q consistently zero but Send-Q still at the default 511 on a high-traffic server is a planning item, not an all-clear. Raise it before the incident, not after. The tradeoff is modest: a longer queue means saturated servers hold connections longer before refusing them, so clients wait instead of failing fast. Behind a load balancer with aggressive health checks, that is usually the right trade; for fail-fast architectures, keep the queue shorter deliberately, but know you chose it.

Signals to watch in production

SignalWhy it mattersWarning sign
ss Send-Q on Apache listenersThe effective, post-truncation backlog. The only place the truth lives.Send-Q lower than configured ListenBacklog
ss Recv-Q (or unacked in ss -lti)Current queue depth. Leading indicator before user-visible failure.Sustained non-zero; Recv-Q above 10; approaching Send-Q
ListenOverflows / ListenDrops countersProof connections have already been refused.Any sustained increase
BusyWorkers / MaxRequestWorkersThe queue only fills when workers cannot accept fast enough.Sustained above 80%; IdleWorkers at zero
AH00484 in the error logApache explicitly reporting worker pool exhaustion; queue fills next.Any occurrence
503 rate and LB health check failuresThe downstream symptoms once the queue overflows.503s appearing; LB removing the node

Correlate queue depth with worker utilization before touching ListenBacklog. If Recv-Q grows while workers are saturated, more backlog buys seconds, but the real fix is worker capacity or backend latency. If Recv-Q grows while workers are idle, something else is blocking accept(), and a bigger queue will only hide it longer.

How Netdata helps

  • Netdata collects Apache scoreboard and worker metrics from mod_status, so you can watch BusyWorkers and IdleWorkers trending toward saturation, the condition that starts filling the accept queue.
  • TCP listener and connection state metrics from the host let you track queue behavior and connection states alongside Apache’s own view, on the same dashboard and timeline.
  • AH00484 MaxRequestWorkers events and 5xx rates can be correlated against queue growth to distinguish “backlog too small” from “workers exhausted”.
  • Anomaly detection on worker utilization surfaces the slow drift toward saturation that makes a 511-deep queue insufficient long before the first refused connection.
  • Per-second granularity catches the brief saturation windows, graceful restarts and burst arrivals, that minute-resolution polling misses entirely.

Netdata’s Apache HTTP Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Apache HTTP Server monitoring with Netdata

Netdata monitors Apache HTTP Server with per-second metrics from mod_status, pre-built dashboards, and ML-powered anomaly detection. Watch busy versus idle workers and the scoreboard state mix, requests per second, bytes served per second, and request processing duration alongside the rest of your stack, so you catch the worker-exhaustion, slow-backend, and memory incidents in these runbooks before they page anyone.