The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / php-fpm / php-fpm-ondemand-cold-start ▌

Operations Guides

PHP-FPM ondemand cold start: fork and warmup latency on the first request

PHP-FPM’s ondemand process manager exists to reclaim memory when a pool is doing nothing. The tradeoff is paid in latency on the first request after an idle period: the master must fork a worker, that worker must handle its first request, and if OPcache has no compiled bytecode for the requested scripts, PHP must parse and compile them from disk. On a warm dynamic or static pool, none of that happens. On an ondemand pool that has been idle, all of it happens on the critical path of a single request.

This is not a bug and it is not an outage. It is the designed behavior of the mode. The most common operational mistake is treating the symptoms of ondemand as incidents: paging on total processes = 0, paging on a latency spike that clears in seconds, or chasing the fork cost as if it were a regression. The second most common mistake is deploying ondemand on a workload where the cold-start cost is unavoidable and user-facing, then being surprised that every poll-driven client pays the penalty repeatedly.

What it is and why it matters

In ondemand mode, the pool starts with zero workers. The master process holds the listening socket and waits. When a connection arrives and there is no idle worker to hand it to, the master forks a worker. That worker accepts the connection, initializes PHP, executes the request, returns the response, and then, after pm.process_idle_timeout (default 10s) with no further work, the master kills it.

The steady state of an idle ondemand pool is total processes = 0. This is the single most important thing to internalize. Every monitoring rule, every alert, every dashboard built on the assumption that zero workers means a dead pool will fire constantly against an ondemand pool that is working exactly as designed. See the PHP-FPM idle processes at zero: no burst headroom left guide for the contrast with dynamic and static modes, where zero idle is a genuine warning.

The reason ondemand exists is memory. A static pool of 50 workers holds 50 PHP processes resident whether they are doing work or not. A dynamic pool holds at least pm.min_spare_servers resident. An ondemand pool holds zero when idle. On a host running many small, bursty PHP applications (shared hosting, low-traffic control panels, cron-triggered webhooks), the memory savings are real and large. The cost is that the first request after an idle gap pays a latency penalty that a warm pool does not.

How it works

The cold-start latency on the first request after idle is the sum of three costs that a warm worker avoids entirely.

flowchart TD
    A[Request arrives] --> B{Idle worker?}
    B -- no --> C[Master forks worker]
    C --> D[Worker inits PHP runtime]
    D --> E[OPcache has bytecode?]
    E -- no --> F[Compile scripts from disk]
    F --> G[Execute request]
    E -- yes --> G
    B -- yes --> G
    G --> H[Response + return to idle]
    H --> I{More work within
process_idle_timeout?} I -- yes --> G I -- no --> J[Master kills worker] J --> K[total processes = 0]

1. Fork cost. The master calls fork() to create a new worker process. On modern Linux this is cheap relative to the rest of the stack but not free: page tables are copied, file descriptor tables are duplicated, and the new process must be scheduled. The fork itself is typically the smallest component of cold-start latency, measured in low single-digit milliseconds on an unconstrained host.

2. Request initialization and application bootstrap. The forked worker inherits PHP module state (SAPI, loaded extensions, php.ini settings) from the master via fork(); it does not re-read configuration or reload extensions at fork time. What the first request pays is RINIT plus application bootstrap (autoloader resolution, service container hydration, configuration loading). These per-request costs exist regardless of pool mode, but ondemand pays them on a process that has never served a request. Framework applications with heavy bootstrap paths are most affected.

3. OPcache compile penalty. This is the component that varies the most. OPcache stores compiled bytecode in shared memory managed by the FPM master process, so it persists across worker kills. If the requested scripts are already cached (compiled by a previous worker, with the master still running and the cache not cleared), the cold-start worker skips compilation. If OPcache is empty (master was restarted, cache was cleared on deploy, or scripts were evicted due to opcache.memory_consumption limits), every script in the request path must be parsed and compiled from disk. For a simple script this adds a few milliseconds. For a framework application with hundreds of included files, the compile penalty dominates the cold-start cost and can push the first request into the tens or hundreds of milliseconds.

A critical distinction: after a simple idle period (workers killed, master still running), OPcache is warm and the cold-start penalty is primarily the fork cost plus first-request overhead. After a master restart or deploy, OPcache is cold and the compile penalty applies. The worst case is a burst of traffic arriving immediately after a restart against a zero-worker pool with cold OPcache. The PHP-FPM mental model guide covers this burst-against-cold-pool pattern. Under low traffic the penalty is a single slow request. Under burst traffic where many requests arrive at once against a zero-worker pool, the master can fork workers rapidly as their connections arrive, each paying the compile cost, and CPU spikes across the pool while throughput stays low. This is the opposite of the normal high-traffic signature, where CPU rises because workers are doing useful work.

Where it shows up in production

The places ondemand cold start shows up as a problem are consistent across deployments.

Intermittent user-facing latency on low-traffic endpoints. A health check, an admin panel, an infrequently used webhook: any endpoint that goes minutes or hours between requests is a candidate for a perceptible cold-start delay. Users report “the first click is slow, then it’s fine.” If the pool is ondemand and process_idle_timeout is short, this is the expected pattern, not a defect.

Monitoring that fires on the wrong signal. The two most common false alarms are total processes = 0 paged as a pool outage, and first-request latency spikes paged as a regression. Both are correct observations of the data and incorrect interpretations of the system. The PHP-FPM active processes near max_children guide covers the adjacent failure of reading saturation metrics without accounting for pm mode.

Polling clients that re-trigger cold starts. The classic example is desktop sync clients or mobile apps that poll every 30 seconds. Each poll, if it arrives after the worker has been killed by process_idle_timeout, pays the full cold-start cost. The pool is never genuinely warm. Nextcloud’s server-tuning guidance explicitly does not recommend ondemand because desktop and mobile clients poll roughly every 30 seconds and repeatedly trigger cold starts. Several applications with polling-heavy clients document ondemand as unsuitable for exactly this reason.

Burst arrivals against a zero-worker pool. If traffic arrives in tight bursts separated by idle gaps longer than process_idle_timeout, every burst pays cold start. A cron-driven batch that fires 20 concurrent requests at an ondemand pool that has been idle for a minute will fork 20 workers, each compiling from an empty OPcache if the master was recently restarted, and the burst will be visibly slow. A warm pool would absorb the same burst in milliseconds.

Container and cgroup-constrained hosts. On a host under memory pressure or a container with tight CPU limits, the fork and compile costs inflate. What is 10-50ms on an unconstrained host can become hundreds of milliseconds on a throttled container, because the worker initialization contends with cgroup CPU limits and the compile phase does real CPU work. See PHP-FPM in containers: cgroup limits and the silent OOM kill for the adjacent container-specific failure modes.

Tradeoffs and when to use it

Ondemand is a deliberate trade of memory for latency. The decision to use it should be made explicitly, not by default.

Ondemand is worth it when:

  • The pool is genuinely idle for long stretches. Shared hosting, low-traffic control panels, cron-triggered endpoints, internal tooling used a few times a day. Here the memory savings are real and the cold-start cost is paid rarely and by users who tolerate it.
  • Memory is the binding constraint. If the host cannot afford to keep N workers resident, ondemand is the lever that lets you run the pool at all. The alternative is not a warm dynamic pool, it is not running the application.
  • First-request latency is not user-facing. Background jobs, webhooks where the caller has generous timeouts, internal health checks that report liveness rather than latency. In these cases the cold-start cost is invisible to anything that matters.

Ondemand is not worth it when:

  • Traffic is continuous or bursty-with-short-gaps. If the pool would be warm most of the time anyway under dynamic, the memory savings vanish and the cold-start tax remains. Use dynamic with pm.min_spare_servers sized to your burst pattern.
  • Clients poll on short intervals. Mobile and desktop sync clients, health checks from aggressive load balancers, any client that reconnects faster than process_idle_timeout. These clients pay the cold-start cost on every interaction. Either raise process_idle_timeout above the poll interval or switch to dynamic.
  • The application is latency-sensitive and framework-heavy. A Laravel or Symfony application with hundreds of included files pays a large compile penalty when OPcache is cold. After a master restart or deploy, the first wave of requests to an ondemand pool pays full compile cost. For these workloads, dynamic or static with a properly sized OPcache is almost always the better choice.
  • You need predictable p99. Ondemand introduces a bimodal latency distribution: fast when warm, slow when cold. If your SLO is on tail latency, the cold-start tail is hard to eliminate without keeping the pool warm.

Tuning levers if you keep ondemand.

  • pm.process_idle_timeout (default 10s) controls how long a worker stays alive after its last request. Lower values reclaim memory faster but increase cold-start frequency. Values below 10s are almost always wrong because each spawn/kill cycle is the most expensive thing the pool does. Raise it above your typical inter-request gap if clients poll.
  • pm.max_children sets the ceiling on simultaneous cold starts during a burst. If a burst arrives against a zero-worker pool, the master forks up to this many workers at once, each paying compile cost if OPcache is cold. Size it for memory, not for steady-state concurrency.
  • OPcache sizing matters more in ondemand than in dynamic, because a cold pool after restart means a cold cache. Ensure opcache.memory_consumption and opcache.max_accelerated_files are sized for the codebase so the cache survives long enough to benefit subsequent cold starts within a traffic window.
  • Pre-warming the cache before directing traffic (hitting critical endpoints after a restart or during a deploy window) reduces the compile penalty for the first real users. This is the same technique recommended for any cold-start scenario.

A known upstream issue is that ondemand does not always scale down as expected: the idle-child selection algorithm can leave workers alive longer than configured, so a pool may sit at 20-30 workers when you expect zero. As of PHP 8.5.10, upstream php/php-src issue #12798 remains open. If you observe workers accumulating in an ondemand pool that should be idle, this is a known behavior rather than a misconfiguration of process_idle_timeout.

Signals to watch in production

SignalWhy it mattersWarning sign
total processesCore to understanding ondemand. Zero when idle is normal; non-zero is the live worker count, not the configured ceiling.Alerting on zero will page constantly. Alert on the pool being unresponsive to traffic, not on worker count.
listen queueThe earliest signal of real saturation. In ondemand, brief queue during cold start is normal and clears in seconds.Sustained non-zero queue with traffic present means workers are not keeping up, not that cold start is the problem.
active processes / total processes ratioSaturation percentage in ondemand is misleading because total is the live count, not max_children. The ratio can hit 100% with room to scale.Do not use this ratio as a capacity signal in ondemand. Use max children reached and listen queue instead.
First-request latency (application-level)The direct measurement of cold-start cost. Expect bimodal distribution: fast when warm, slow when cold.A flat latency distribution with no cold-start tail suggests the pool is staying warm (good) or you are not sampling the cold requests.
OPcache hit rateConfirms whether cold starts are paying a compile penalty. OPcache persists across worker kills, so low hit rate points to a restart, deploy clear, or cache eviction, not idle.Hit rate that never recovers between bursts means OPcache is too small or being cleared.
max children reached counterMeaningful in ondemand, unlike static. Increments when the pool wanted to fork but hit the ceiling.Rate of change during normal traffic means max_children is too low for the burst pattern.
Fork frequencyEach fork is a cold-start cost. High fork rate with low throughput means workers are being killed and respawned too aggressively.Correlates with process_idle_timeout being too short for the traffic pattern.

The monitoring posture for ondemand is different from dynamic and static. Do not alert on total processes = 0. Do not alert on a single slow first request. Do alert on sustained listen queue > 0, on max children reached climbing, and on the pool failing to accept traffic at all (ping failure with incoming requests). See PHP-FPM listen queue growing: the earliest signal of saturation for the saturation signal that does warrant alerting regardless of pm mode.

How Netdata helps

  • Per-second polling of total processes, active processes, and idle processes shows the ondemand fork and kill cycle directly, including the brief window where workers exist between first request and idle timeout. At coarser polling intervals this behavior is invisible.
  • Anomaly detection on first-request latency distinguishes the expected cold-start tail from a genuine regression. Rather than alerting on every slow first request, it flags only latency that deviates from the established pattern for this pool.
  • Correlating fork events with OPcache hit rate and CPU shows whether cold-start latency is dominated by compile cost (OPcache miss) or by fork overhead. This separates “OPcache is too small” from “the pool is cycling too aggressively.”
  • The PHP-FPM collector exposes max children reached and listen queue as first-class signals, so alerting can target the real saturation indicators rather than the misleading active/total ratio that ondemand distorts.
  • Dashboards that place FPM pool metrics next to web server 502/504 rates and cgroup CPU throttling make it obvious whether a cold-start spike is user-visible (504s from the web server) or absorbed (slow but successful requests).