The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / php-fpm / php-fpm-workers-blocked-on-io ▌

Operations Guides

PHP-FPM workers busy at low CPU: blocked on the database, API, or filesystem

PHP-FPM active processes is climbing toward pm.max_children, the listen queue is building, users are seeing slow responses or intermittent 502s, and yet system CPU is flat at 10 to 20 percent. The pool looks saturated in terms of worker slots, but the CPU headroom suggests the box is barely working.

Each PHP-FPM worker handles exactly one request at a time. When a worker blocks on I/O (a database query, an external API response, a DNS lookup, or a filesystem operation), it holds its worker slot for the entire duration of that wait while consuming essentially zero CPU. It is “active” from FPM’s scoreboard perspective but idle from the kernel scheduler’s perspective.

Raising pm.max_children in response rarely helps and often makes things worse: new workers will also block on the same dependency, multiplying the downstream connection load without improving throughput. The fix is to identify what the workers are waiting on and address that dependency, or add a timeout that lets workers fail fast.

What this means

The diagnostic signature is the divergence between active processes and system CPU utilization. When active processes approaches pm.max_children and CPU stays low, the workers are not computing. They are parked in the kernel’s I/O wait state, holding slots that no amount of CPU can free.

This contrasts with the CPU-bound case, where high active processes correlates with high CPU because workers are genuinely executing PHP. In the I/O-bound case, the PHP runtime has issued a blocking syscall (a recvfrom on a database socket, a curl_easy_perform on an HTTP connection, a read on an NFS mount) and the process is descheduled until data arrives or the call times out.

The downstream consequence is the same as any worker exhaustion: as blocked workers accumulate, fewer slots remain for normal traffic. The listen queue starts filling. Once it overflows, the kernel drops connections and the web server returns 502. The PHP application code is often completely healthy. It is patiently waiting for a dependency that is slow, unreachable, or contended.

The key field on the status page that confirms this is last request cpu, available in full mode (?full). For workers in Idle state, this shows the CPU percentage consumed by their last completed request. A value near zero for a request whose request duration was several seconds is definitive proof the worker spent that time blocked on I/O, not executing PHP. Note that for workers still in Running state, last request cpu is always zero because CPU accounting only completes when the request finishes.

flowchart TD
    A[Active processes climbing] --> B{System CPU low?}
    B -- Yes --> C[I/O-bound worker drain confirmed]
    B -- No --> D[CPU-bound: check opcache, compute paths]
    C --> E[Read slow log for stack traces]
    E --> F{Top of stack shows}
    F -- PDO / mysqli --> G[Database lock or slow query]
    F -- curl / stream --> H[External API or DNS latency]
    F -- file I/O functions --> I[NFS stall or disk saturation]
    F -- session_start --> J[Session lock contention]

Common causes

CauseWhat it looks likeFirst thing to check
Database lock contention or slow querySlow log stack traces show PDO or mysqli functions; database shows lock waits or long-running queriesSlow query log; pg_locks / information_schema.INNODB_TRX
Database connection limit reachedWorkers block acquiring a connection; DB shows connection count at max_connectionsSHOW PROCESSLIST or pg_stat_activity vs DB max connections
External API or DNS latencySlow log shows curl or stream functions at top of stackCall the API endpoint directly and measure latency
NFS or shared filesystem stallWorkers enter D state (uninterruptible sleep); dmesg shows rpc_wait_bit_killablenfsstat retransmits, mountstats, dmesg
Session lock contentionSlow log shows blocking at session_start; affects concurrent requests from same session IDCheck session.save_handler and concurrent same-session requests
cgroup I/O throttlingWorkers block on file I/O despite healthy storage; cgroup IO limits setCheck systemd slice IOWeight or cgroup v2 I/O limits

Quick checks

# Check active processes vs max_children ratio
curl -s http://127.0.0.1/fpm-status | grep -E "active processes|max children|idle processes"

# Check for I/O-bound workers: low last request cpu with long request duration
curl -s 'http://127.0.0.1/fpm-status?full' | grep -E "request duration|last request cpu|state"

# Check slow request counter rate (run twice, 10 seconds apart, compare values)
curl -s http://127.0.0.1/fpm-status | grep "slow requests"

# Read recent slow log entries (path varies by distro and pool config)
tail -100 /var/log/php-fpm/slow.log

# Check if slow log is even configured
php-fpm -tt 2>&1 | grep -E "slowlog|request_slowlog_timeout"

# Check for workers in D state (uninterruptible sleep, typical of NFS stalls)
ps -eo pid,stat,wchan,cmd | grep '[p]hp-fpm' | grep -v master | awk '$2 ~ /^D/'

# Check database connection count from this host
# PostgreSQL:
psql -c "SELECT count(*) FROM pg_stat_activity WHERE client_addr = '<this_host>';"
# MySQL:
mysql -e "SELECT COUNT(*) FROM information_schema.processlist WHERE HOST LIKE '<this_host>%';"

# Check kernel-level listen queue overflow counter
nstat | grep -i ListenOverflows

# Check what outbound connections FPM workers are holding
ss -tnp | grep php-fpm | awk '{print $5}' | cut -d: -f1 | sort | uniq -c | sort -rn | head

How to diagnose it

  1. Confirm the divergence. Pull the status page and compare active processes against your system CPU metric. If active is above 80 percent of pm.max_children and CPU is below 30 percent, you are in I/O-bound territory. Pull ?full and check workers in Running state with high request duration values.

  2. Verify the slow log is configured. request_slowlog_timeout defaults to 0, which disables it entirely. If php-fpm -tt shows it unset or 0, you are flying blind. Set it to 5 seconds (or whatever threshold fits your baseline latency) and reload. The slow log is the single most important diagnostic tool for this pattern. When a worker exceeds the threshold, PHP-FPM attaches to or reads the worker with the platform tracing mechanism (ptrace, /proc/<pid>/mem, or Mach VM APIs), captures a backtrace, then detaches or resumes it. The resulting stack trace shows where the worker was parked.

  3. Read the slow log stack traces. Group entries by the top of the call stack. If you see PDO methods (PDO::query, PDOStatement::execute) or mysqli calls at the top, the database is the bottleneck. If you see curl_exec or stream wrappers, an external API is slow. If you see fopen, fread, or session_start, the filesystem or session locking is the issue. If you see rpc_ functions, NFS is involved.

  4. Check the dependency directly. Once the slow log identifies the category, measure the dependency in isolation. Run the slow query against the database with EXPLAIN ANALYZE. Curl the external API endpoint from the FPM host. Check nfsstat or mountstats for NFS retransmits and timeouts. The dependency’s own metrics will confirm the diagnosis.

  5. Check the connection multiplier. Count how many database connections originate from this FPM host. Each FPM worker that opens a database connection during request processing holds one connection. N active workers means up to N concurrent database connections from this host alone. If you have multiple app servers, multiply accordingly. Compare against the database’s max_connections or connection pool size.

  6. Check for D-state workers. If ps shows workers in state D (uninterruptible sleep), they are stuck in a kernel I/O syscall, typically NFS. These workers cannot be killed with SIGTERM. The only recovery is fixing the NFS server or rebooting. dmesg will often show the WCHAN as rpc_wait_bit_killable.

  7. Check for self-request deadlocks. If a PHP worker makes an HTTP request back to the same application (e.g., calling its own API endpoint), it consumes a second worker slot. Under load, workers end up waiting for each other. Use ss -tnp | grep php-fpm to check whether FPM workers are connecting to the application’s own gateway address.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
active processes / max_children ratioPrimary saturation indicatorSustained above 0.8 with low CPU
System CPU utilization vs active processesDivergence reveals I/O-bound patternActive climbing, CPU flat
slow requests counter (rate of change)Quantifies how many requests exceed slowlog thresholdSustained non-zero rate
Per-worker request duration and last request cpuIdentifies individual blocked workersLong duration with near-zero CPU
listen queue depthLeading indicator before connection dropsAny sustained non-zero value
Database connection count from FPM hostValidates the N-workers-equals-N-connections mathApproaching DB max_connections
Kernel ListenOverflows counterDetects drops invisible to FPMAny increment
Worker state distribution (D vs R vs S)D-state workers indicate kernel I/O stallAny D-state workers

Fixes

If the database is the bottleneck

The slow log will show PDO or mysqli at the top of the stack. Check for lock contention, missing indexes causing full table scans, or queries that regressed after a schema change. If the issue is connection exhaustion rather than query slowness, the fix is a connection pooler. At scale, N FPM workers each opening a connection means N connections from this host. Put PgBouncer (PostgreSQL) or ProxySQL (MySQL) between FPM and the database. This lets hundreds of PHP workers share a smaller pool of database connections.

Persistent connections (PDO::ATTR_PERSISTENT) are a simpler mitigation that caps database connections at pm.max_children, since each worker reuses one persistent connection. This is adequate for single-host deployments but does not multiplex the way an external pooler does.

If an external API is slow

The slow log will show curl_exec or stream functions. PHP’s libcurl has no default timeout, so a worker calling a dead or slow API will wait indefinitely. Set explicit CURLOPT_TIMEOUT and CURLOPT_CONNECTTIMEOUT in application code. A typical safe pattern is 2 seconds for connect, 5 to 10 seconds for total request time. Add circuit breaker logic so repeated failures short-circuit instead of queuing more workers behind the same dead dependency.

If NFS or the filesystem is stalled

Workers accessing files on NFS will enter D state when the NFS server stalls. They cannot be killed or recycled until the syscall completes. The kernel call stack will show rpc_wait_bit_killable. The only fix is to restore the NFS server. To prevent recurrence, consider whether the application truly needs NFS for the code paths that are blocking, or whether those files can be served from local storage. Check cgroup I/O limits if running under systemd slices with IOWeight or IODeviceWeight, as these can throttle filesystem access even when underlying storage is healthy.

If session lock contention is the cause

The slow log will show blocking at session_start. With file-based sessions, concurrent requests from the same session ID serialize via flock(LOCK_EX) on the session file. AJAX-heavy pages firing multiple parallel requests for the same user will all queue behind one lock. The fix is to call session_write_close() as early as possible in the request lifecycle, after the last session write. For a structural fix, switch to Redis or Memcached session handlers, which have different locking semantics.

If the dependency cannot be fixed immediately

Set or lower request_terminate_timeout so stuck workers are killed after a hard limit rather than holding their slot forever. This trades a 502 response for the stuck request (better than a hung connection) and frees the worker slot for subsequent traffic. The timeout should be coherent with the web server’s fastcgi_read_timeout to avoid phantom workers that continue executing after nginx has already returned a 504 to the user.

You can also selectively kill blocked workers to free capacity temporarily:

# Ask a specific blocked worker to start graceful shutdown
kill -SIGQUIT <worker_pid>

Do not kill D-state workers. SIGQUIT will not interrupt a kernel I/O wait.

Prevention

  • Enable the slow log on every production pool. Set request_slowlog_timeout to a value appropriate for your baseline latency (5 seconds is a reasonable starting point). This is the single highest-value configuration change for diagnosing I/O-bound workers.

  • Set explicit timeouts on all outbound calls. Database queries, HTTP API calls, and cache connections should all have fail-fast timeouts. Default PHP cURL timeout is effectively infinite.

  • Set request_terminate_timeout. A hard ceiling prevents stuck workers from permanently consuming slots. 30 to 60 seconds is typical, adjusted for your application’s legitimate long-running endpoints.

  • Install a connection pooler at scale. If your FPM worker count times the number of app servers exceeds the database’s max_connections, you need PgBouncer or ProxySQL. Do not rely on each worker managing its own connection.

  • Monitor the active-vs-CPU divergence as a composite signal. Alerting on active processes alone will fire during legitimate traffic spikes. Alerting on the divergence (high active, low CPU, slow requests incrementing) catches the I/O-bound drain pattern specifically.

  • Avoid NFS for hot code paths. If the application reads templates, configs, or user uploads from NFS, any server stall will pin workers in D state. Prefer local storage or a CDN for read-heavy paths.

How Netdata helps

  • Netdata collects PHP-FPM status page metrics at per-second resolution, which matters because saturation events unfold in seconds. The divergence between active processes and system CPU is visible in real time when both metrics share the same collection cadence.

  • The slow request counter rate is collected alongside active and idle process counts, so you can see whether worker saturation correlates with slow log activity in the same time window.

  • Per-worker request duration data from full status mode lets you distinguish bimodal distributions (some fast, some extremely slow) from uniform slowdown, which narrows the diagnosis between endpoint-specific blocking and systemic backend failure.

  • Database connection metrics from Netdata’s PostgreSQL, MySQL, and generic database collectors can be overlaid against FPM active process counts. When both climb together, the database connection multiplier is the likely cause.

  • ML anomaly detection flags the specific pattern of rising active processes with flat CPU as anomalous, even before it crosses a static threshold. This is useful because the absolute values are workload-dependent, but the divergence shape is a reliable signal.

  • Kernel-level socket metrics, including listen queue depth and overflow counters, provide the layer below FPM’s own visibility. These catch connection drops that the status page cannot report.