The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / php-fpm / php-fpm-listen-backlog-overflow ▌

Operations Guides

PHP-FPM listen backlog overflow: the kernel silently dropping connections

The signature: a PHP-FPM pool where every status page field looks healthy yet the web server sporadically returns 502s. Active processes are not pinned at pm.max_children, the listen queue field reads zero, the slow log is quiet, opcache hit rate is normal. The 502s cluster into short bursts that may or may not line up with known load events.

The kernel is dropping incoming FastCGI connections because the listen backlog has filled. PHP-FPM has no visibility into these drops. Once the backlog is at capacity, the kernel refuses to enqueue new connections, and the status page cannot report a queue depth above the configured listen.backlog value because those connections never enter a queue FPM could observe. The failure happens entirely below PHP-FPM’s instrumentation layer.

The fix is not to enlarge the backlog. A deeper queue only delays drops; it does not make the pool drain faster. The real fixes are more workers, faster requests, or fewer arrivals.

What this means

PHP-FPM workers process one request at a time. When all workers are busy, the kernel parks incoming FastCGI connections in the socket backlog, sized by listen.backlog in the pool configuration. The master process eventually calls accept() and hands each connection to an idle worker.

The backlog’s effective size is min(listen.backlog, net.core.somaxconn). PHP’s built-in default is version- and platform-dependent: PHP < 8.2 uses 511 on Linux/macOS; PHP 8.2+ uses -1 on Linux, FreeBSD, OpenBSD, DragonFly BSD, and macOS, which Linux clamps to net.core.somaxconn. Other platforms use 511. net.core.somaxconn itself varies by distribution and configuration (older hosts commonly used 128, many modern hosts use 4096). Some hardened or cloud-optimized images ship net.core.somaxconn = 128, which silently caps even a large listen.backlog. Always confirm with sysctl net.core.somaxconn before reasoning about effective backlog depth.

Once the effective backlog fills, kernel behavior depends on socket type:

  • TCP listener: the kernel increments TcpExtListenOverflows and TcpExtListenDrops in /proc/net/netstat and drops the new connection. Visible via nstat or netstat -s.
  • Unix domain socket (the common case for nginx plus PHP-FPM on the same host): no SYN retry, no TCP counter increments. The kernel refuses the connect() attempt. nginx logs an upstream connect failure and returns 502. Linux does not expose a portable cumulative Unix-socket accept-queue drop counter in procfs; dynamic tracing can inspect connect and backlog state, but it is not a standard drop counter.

The PHP-FPM side sees nothing of this. The listen queue field in the status page is a snapshot of current backlog depth, and by the time you read it the burst may have passed, so it reads 0. The max listen queue high-water mark can show that the pool touched its limit, but not how many connections were rejected.

flowchart TD
    A[nginx opens FastCGI conn] --> B{Worker free?}
    B -- yes --> C[accept and dispatch]
    B -- no --> D{Backlog full?}
    D -- no --> E[enqueue in kernel]
    E --> C
    D -- yes --> F[kernel drops connection]
    F --> G[nginx logs 502]
    F -. invisible to .-> H[FPM status page]
    F --> I[TcpExtListenDrops on TCP; no counter on Unix sockets]

Common causes

CauseWhat it looks likeFirst thing to check
Worker pool fully saturatedactive processes pinned at pm.max_children during burstsFPM full status page, per-worker request URIs and durations
Slow dependency draining workersCPU low, active high, slow log full of DB or curl tracesslow log stack traces and downstream service latency
listen.backlog far above somaxconnPHP config shows 65535 but effective depth is much smallerss -xlnp Send-Q vs sysctl net.core.somaxconn
Transient bursts under slow poll interval502s appear and vanish; status page samples show nothingnstat -az ListenDrops rate vs polling cadence
Reload window with no workers502s line up with SIGUSR2 reloads (cron, logrotate)FPM error log reload events vs nginx error timestamps

The most insidious row is the fourth. With a 10-second polling interval, a 200-millisecond burst that fills and drains the backlog produces zero visible samples in the FPM status page but plenty of kernel counter increments and 502s in the nginx error log.

Quick checks

# TCP listener: kernel drop counters (cumulative since boot)
nstat -az | grep -E 'ListenDrops|ListenOverflows'

# Same counters in a friendlier form
netstat -s | grep -i listen

# Effective backlog vs current queue depth (TCP)
ss -tlnp | grep 9000

# Effective backlog vs current queue depth (Unix socket)
ss -xlnp | grep php

# Confirm sysctl cap
sysctl net.core.somaxconn

# PHP-FPM status page fields
curl -s http://127.0.0.1/fpm-status | grep -E '^(listen queue|max listen queue|listen queue len)'

Interpretation rules:

  • On ss -tlnp and ss -xlnp LISTEN lines, Recv-Q is the current accept queue depth and Send-Q is the effective maximum backlog after somaxconn clamping. If Recv-Q approaches Send-Q, the kernel is about to start dropping. If Recv-Q equals Send-Q for any sustained period, drops are almost certainly happening.
  • nstat values are cumulative since boot. To detect active drops, sample, wait a minute, sample again. Any delta on ListenDrops or ListenOverflows is a problem.
  • The status page field listen queue len reflects what PHP-FPM asked for in listen.backlog, not what the kernel actually granted. The only source of truth for the effective cap is Send-Q on the ss LISTEN line.
  • For Unix sockets there is no portable cumulative drop counter exposed in procfs. Dynamic tracing with bpftrace or perf can inspect connect and backlog state, but it is not a drop counter. Use ss -xlnp Recv-Q vs Send-Q as the real-time signal, and correlate with nginx error log timestamps.

How to diagnose it

  1. Confirm kernel drops are happening. On a TCP listener, watch nstat -az | grep -E 'ListenDrops|ListenOverflows' over a short window. Any non-zero delta is conclusive. On a Unix socket, monitor Recv-Q via ss -xlnp and correlate with nginx error timestamps; there is no equivalent portable cumulative drop counter in procfs.
  2. Cross-reference with nginx error log timestamps. Filter for upstream connect failures during the same window: grep -E "connect.*failed|Connection refused|no live upstreams" /var/log/nginx/error.log. The 502s should line up with periods when drops incremented.
  3. Confirm workers are actually busy at the drop moments. Poll the FPM full status page at 1-second intervals during a burst and count workers in Running state. If running worker count equals pm.max_children at those moments, you have confirmed the precondition.
  4. Check max listen queue historical peak. A non-zero value on a freshly restarted pool tells you the pool has already touched its backlog since the last restart. If max listen queue is close to listen queue len, you have been at the edge.
  5. Inspect per-worker request URIs. A small number of slow endpoints holding workers hostage is the most common pattern. Pull curl -s 'http://127.0.0.1/fpm-status?full&json' and look for repeated URIs with high request duration. If a handful of endpoints account for most of the worker time, fixing those is faster than adding workers.
  6. Verify the effective backlog matches intent. Compare sysctl net.core.somaxconn against the Send-Q column in ss -xlnp for the FPM socket. If they differ, the pool config is not doing what you think.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
TcpExtListenDrops / TcpExtListenOverflows (TCP only)Direct kernel proof of dropped acceptsAny non-zero rate during traffic
Recv-Q on the FPM LISTEN socket via ssCurrent backlog depth, real-time, works for TCP and Unix socketsApproaching Send-Q, the effective cap
Send-Q on the FPM LISTEN socket via ssEffective backlog after somaxconn clampingLower than the configured listen.backlog
FPM listen queue (status page)Inside-pool view of pending connectionsNon-zero during normal traffic means saturation
FPM max listen queue (high-water mark)Historical peak since restartApproaching listen queue len
nginx upstream connect failuresThe user-visible symptomSpikes that correlate with TcpExtListenDrops
FPM active processes / pm.max_children ratioThe precondition for backlog fillSustained near 1.0 during bursts

The first two rows are the only ones that prove the failure. Everything else is a leading indicator or corroborating signal.

Fixes

Add workers

If memory headroom permits, raise pm.max_children. The capacity formula is max_children = (available_memory * 0.7 - OS_overhead) / avg_worker_PSS, using PSS rather than RSS to avoid overstating memory by 30-50% on shared opcache pages. More workers means faster accept() drain, which means the backlog never fills.

This is the right fix when per-worker request duration is healthy (matches baseline p95) and the problem is purely traffic volume exceeding concurrency.

Make requests faster

If workers are busy because a small set of endpoints is slow, fix those endpoints first. Adding workers to compensate for slow database queries or external API timeouts only delays the next saturation event. Inspect the slow log and per-worker request URIs to find the offenders, then optimize the code, add timeouts, or circuit-break against the slow backend.

This is the right fix when slow log entries cluster on specific endpoints, or when active processes are high but CPU is low (workers are I/O-blocked, not computing).

Raise net.core.somaxconn only if it is the real cap

If sysctl net.core.somaxconn is below your configured listen.backlog, the kernel is silently clamping the queue. Raising somaxconn is reasonable as a one-time system tuning step. But raising it without addressing the underlying saturation only buys more queue capacity before drops start. It does not increase request throughput.

Do not just raise listen.backlog

Increasing listen.backlog is the most tempting and least effective fix. A deeper queue means longer waits for queued requests, which means higher latency for the requests that do get served. It also means larger bursts of work arriving at the workers once the queue drains, which can deepen the original saturation. The default listen.backlog (511 on PHP < 8.2, -1 on PHP 8.2+) is sufficient for almost all workloads once the pool is correctly sized.

Prevention

  • Kernel drop counters, continuously. TcpExtListenDrops on TCP listeners is the only signal that catches this failure mode before users report 502s. For Unix sockets, monitor Recv-Q / Send-Q ratio instead. A 1-second polling interval is appropriate; 10 seconds will miss transient bursts.
  • Recv-Q relative to Send-Q on the FPM listening socket. A trendline of Recv-Q / Send-Q gives early warning before drops begin.
  • Status page and worker count in the same view. A non-zero listen queue while active processes is below max_children suggests a worker spawn or socket issue rather than pure saturation.
  • Set request_slowlog_timeout in production. Without slow log traces, you cannot tell the difference between “lots of traffic” and “a few slow endpoints eating the pool.”
  • Set request_terminate_timeout to a finite value. A stuck request permanently removes a worker, silently reducing capacity and pushing the pool toward backlog overflow.
  • Set pm.max_requests to 500-1000. Memory leaks reduce effective worker count over time, narrowing the gap between normal load and saturation.
  • Coordinate timeout hierarchies. If nginx fastcgi_read_timeout is longer than FPM request_terminate_timeout, FPM kills the worker while nginx keeps waiting; if shorter, the worker keeps running for a response nobody will read.
  • Verify reload behavior. SIGUSR2 produces a brief no-worker window during which the backlog can fill quickly. If logrotate fires SIGUSR2 at midnight and 502 bursts cluster around that time, see PHP-FPM graceful reload: the brief no-worker window on SIGUSR2.

How Netdata helps

Netdata collects the signals that prove or rule out backlog overflow at per-second resolution, which matters because the failure mode is often sub-10-seconds and invisible to slower pollers.

  • TCP listener drops and overflows (TcpExtListenDrops, TcpExtListenOverflows) are collected as first-class metrics from /proc/net/netstat, so you can alert on rate-of-change rather than eyeballing cumulative counters.
  • Recv-Q and Send-Q on listening sockets give you the effective backlog vs current queue depth ratio without a custom collector. This works for both TCP and Unix socket listeners.
  • PHP-FPM status page fields (active processes, listen queue, max listen queue, max children reached) appear in the same dashboard as the kernel counters, so you can correlate the inside-pool view with the outside-pool drop view.
  • Composite alerts combining “active processes near max_children” with “kernel ListenDrops incrementing” give you continuous confirmation of the diagnostic steps above, without manual nstat polling.