The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / php-fpm / php-fpm-active-processes-high ▌

Operations Guides

PHP-FPM active processes near max_children: reading pool utilization

The ratio of active processes to pm.max_children is the most important saturation signal PHP-FPM exposes. Each worker handles one request at a time, so the active worker count is your current concurrent request load. Divided by the configured ceiling, it behaves like a capacity gauge: near 1.0 means no headroom; sustained near 1.0 means one slow query away from queuing.

This is a reading guide for that ratio: what the numerator and denominator actually mean, how to distinguish a normal burst from a sustained problem, how to use the high-water mark (max active processes) for capacity planning, and which correlated signals disambiguate “busy”, “saturated”, and “broken”. It is not a tuning guide; see the related guides at the end for that.

active processes counts workers currently executing a request. idle processes counts workers ready to accept one immediately. total processes is their sum. pm.max_children is the configured ceiling on how many workers the pool may run. The saturation ratio is active processes / pm.max_children. In static mode the denominator equals total processes; in dynamic and ondemand the master spawns workers up to that ceiling on demand.

What the ratio means

Each active worker is occupied for the full duration of a request, including time spent blocked on I/O. A worker waiting 4 seconds on a database query is “active” from PHP-FPM’s perspective even though it is consuming essentially zero CPU. This is why the active count can sit at the ceiling while CPU remains low: the workers are not computing, they are waiting.

Ratio bandInterpretation
0.0 to 0.7 sustained at peakHealthy headroom. Bursts can be absorbed without queuing.
0.7 to 0.85 sustained at peakTight. Capacity planning should be in motion.
Above 0.85 sustainedDanger zone. One slow dependency or traffic burst away from queuing.
1.0 sustainedNo headroom. Every new request waits for a worker to finish.
1.0 plus non-zero listen queueConfirmed exhaustion. Requests are stacking in the socket backlog.

The boundary between “busy” and “saturated” is the listen queue, not the active count alone. Hitting 100% of pm.max_children means the next request will queue, but queuing is only observable once listen queue goes non-zero. Until then you are at the ceiling but the backlog is still absorbing the slack.

Burst versus sustained

A brief spike to 100% is normal. PHP-FPM is designed to absorb short bursts using the listen backlog as a shock absorber. A flash crowd, a cache invalidation, or a deploy-time opcache warmup can push active to the ceiling for a few seconds and then drain. This is not an incident.

Sustained near-ceiling is the problem. The rule of thumb: above 80% of pm.max_children sustained during normal traffic means insufficient headroom. The failure mode is cliff-edged, not gradual. Once the ceiling is reached, latency does not rise linearly. Requests stack in the socket backlog and time out in batches, so the transition from “fine” to “502s for everyone” can happen in seconds.

Sampling cadence matters. The status page is a point-in-time snapshot. A 10- or 30-second poll can miss a transient queue buildup entirely, showing listen queue = 0 while users saw intermittent 502s between polls. For operational alerting on saturation, poll at per-second intervals. For capacity planning, 10 seconds is adequate because you are tracking trends, not catching spikes.

The high-water mark: max active processes

active processes is the current snapshot. max active processes is the high-water mark since the pool (master) started: the peak concurrent active worker count observed over the lifetime of the process. For capacity planning it is usually more useful than the current value, because it tells you how close you have come to the wall even when things look calm right now.

Two readings matter:

  • If max active processes equals pm.max_children, the pool hit the ceiling at least once since start. This counter alone does not tell you whether requests queued, only that every worker was busy simultaneously at some point. Cross-check max listen queue: if it is also non-zero, queuing happened.
  • If max active processes sits well below pm.max_children (for example, 60% at peak), you have headroom you may be able to reclaim. Oversized pools waste memory; the high-water mark tells you whether pm.max_children can be lowered safely.

Both high-water counters reset on master restart. SIGUSR2 graceful reload re-execs the master, so the counters reset on every reload. If you reload frequently during deploys, the historical peak is less useful because you lose context. Track reloads alongside these counters, and note that accepted conn also resets on reload, which can produce artificial spikes or negative deltas in monitoring systems that compute rates across the boundary.

Correlations that disambiguate

The active ratio alone tells you utilization. It does not tell you why workers are busy, whether they are doing useful work, or whether users are already being harmed. Three correlations resolve almost every ambiguity.

flowchart TD
  A[active / max_children] --> B{sustained above 0.8?}
  B -- no --> C[healthy headroom]
  B -- yes --> D{listen queue greater than 0?}
  D -- no --> E[at ceiling, not yet queuing]
  D -- yes --> F[confirmed exhaustion]
  F --> G{CPU high or low?}
  G -- low --> H[I/O-bound: slow log, DB, API, sessions]
  G -- high --> I[compute-bound: more workers or more CPU]
Signal combinationWhat it meansFirst thing to check
High active, zero listen queueAt ceiling but backlog is absorbing it. Headroom is gone but users are not yet blocked.Whether this is sustained or a burst.
High active, growing listen queueConfirmed worker exhaustion. Requests are stacking.Slow log and per-worker request URI to find what is holding workers.
High active, low CPUWorkers are blocked on I/O, not computing. Classic slow-dependency pattern.Database, external API, NFS, DNS, session locks. Raising max_children may help only briefly.
High active, high CPUWorkers are genuinely computing. Compute-bound workload or opcache thrash.Opcache hit rate and memory; whether more CPU or more workers is the right lever.
Active at ceiling, listen queue empty, web server returning 502Backlog already overflowed. Connections are being refused at the kernel level before reaching the queue.Kernel ListenOverflows/ListenDrops counters (`nstat -az
Low active, high latencyA few workers running extremely slow code. Throughput is minimal despite spare-looking capacity.Per-worker request duration from ?full status; slow log.

The most common production pattern is high active plus low CPU. A slow database query, a hung external API call, or file-based session lock contention ties up workers for seconds or minutes while using no CPU. The pool looks saturated, but the root cause is upstream, not in PHP-FPM. Raising max_children in this case buys time but does not fix the problem, because the new workers will block on the same dependency.

Per-pool and per-mode reading

Read each pool independently. If you run multiple pools, aggregate active-process counts across pools are misleading. A memory leak or a slow endpoint in one pool does not affect the others, and one saturated pool behind a round-robin load balancer means every Nth request is slow even when aggregate utilization looks fine.

The process manager mode changes how you interpret the numbers:

  • static: total processes always equals pm.max_children. The integrity check active + idle == max_children should hold. The max children reached counter is always 0 because the master never tries to spawn beyond the fixed pool, so do not alert on it. Watch listen queue and max active processes instead.
  • dynamic: Workers scale between pm.min_spare_servers and pm.max_children. max children reached is meaningful: each increment is a moment the master wanted to spawn a worker but was blocked by the ceiling. Watch its rate of change, not the absolute value.
  • ondemand: Workers spawn on request and die after pm.process_idle_timeout. active and total of 0 at idle is normal, not a failure. The cost is cold-start latency on the first request after idle, which can look like a brief saturation spike as workers fork and opcache warms.

Gotchas when reading the count

Counting bug under FastCGI keepalive. With nginx fastcgi_keep_conn on, the status page can report active processes and total processes far above pm.max_children (upstream reports show values like 3347 active against a 1500 ceiling). The active counter is incremented during the “reading headers” stage and decremented during the “accepting” stage, and with keepalive there is one accepting stage but multiple reading-headers stages. A fix has been proposed (PR #19191, July 2025) but is still an open draft at the time of writing and is not included in any released PHP version. If you use fastcgi_keep_conn on and see active exceeding max_children, you cannot trust the ratio.

Status page polling consumes a worker. Each status request occupies a worker slot for the duration of the FastCGI request. Polling aggressively during a saturation event adds load to an already-loaded pool. Use the /ping endpoint for sub-second liveness (it returns a fixed string and does not require a worker slot in the same way) and reserve the full status page for 15-30 second monitoring polls or per-second polling only during active incidents. If you need the status page reachable even when the main pool is fully saturated, configure a separate status listener (pm.status_listen, available since PHP 8.0).

Status page exposure. The status endpoint reflects the request URI into its output. On unpatched versions, requesting ?html or ?xml output from a browser session with sensitive cookies can execute injected script (CVE-2026-6735, patched in PHP 8.2.31, 8.3.31, 8.4.21, 8.5.6). Restrict the status path to localhost or trusted IPs regardless of version, and avoid loading it in a browser session that has access to the admin interface.

total processes below max_children in static mode. If total processes is less than pm.max_children in static mode, workers are dying faster than the master can replace them. This is a crash loop or a fork failure, not a capacity problem, and the active ratio becomes misleading because the effective ceiling is lower than configured.

Snapshot timing. The status page is a single point-in-time read. Between two polls an entire saturation event can occur and resolve. A current active of 40% does not rule out a spike to 100% two seconds ago. Always pair the current snapshot with the high-water mark (max active processes, max listen queue) to catch events that happened between polls.

Reading it manually

The numbers are available from the status page in text, JSON, or OpenMetrics format (OpenMetrics was added in PHP 8.1.0, enabling native Prometheus scraping without an exporter). The status page does not expose pm.max_children, so pull it from the pool config.

# Adjust path to match your installation. JSON field names use spaces,
# not underscores: "active processes", "max listen queue", etc.
POOL_CONF=/etc/php/8.3/fpm/pool.d/www.conf
MAX_CHILDREN=$(awk -F= '/^[[:space:]]*pm\.max_children[[:space:]]*=/{
    gsub(/[[:space:]]/,"",$2); print $2; exit}' "$POOL_CONF")

curl -s "http://127.0.0.1/fpm-status?json" | python3 -c "
import sys, json
d = json.load(sys.stdin)
active = d['active processes']
ceiling = ${MAX_CHILDREN:-0}
ratio = active / ceiling if ceiling else 0
print(f\"active={active} max_children={ceiling} ratio={ratio:.2f}\")
print(f\"max active (high-water)={d['max active processes']}\")
print(f\"listen queue={d['listen queue']} (max seen={d['max listen queue']})\")"

Adjust the URL to match your pm.status_path and web server config. The defaults above assume the web server proxies /fpm-status to the pool. Using total processes in the denominator instead of pm.max_children is only correct in static mode; in dynamic and ondemand it overstates utilization whenever the pool is not full.

For the per-worker view that reveals which endpoints are holding workers, append &full and read request uri and request duration per process.

How Netdata helps

  • Per-second sampling of active processes, idle processes, total processes, and listen queue catches transient saturation events that 10- or 30-second polls miss entirely.
  • The active and total series sit next to CPU, request rate, and downstream dependency latency on one timeline, which is the disambiguation step that separates I/O-bound saturation from compute-bound saturation.
  • Per-pool charts keep each pool’s saturation independent, so one noisy pool behind a load balancer does not get averaged away.
  • Anomaly detection on the active series flags sustained-near-ceiling behavior as distinct from normal bursty traffic, reducing noise on brief spikes that drain on their own.