The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / php-fpm / php-fpm-graceful-reload-window ▌

Operations Guides

PHP-FPM graceful reload: the brief no-worker window on SIGUSR2

A PHP-FPM graceful reload (kill -USR2 <master_pid>, systemctl reload php8.3-fpm) is not graceful in the nginx sense. There is no overlap window where new workers handle fresh traffic while old workers drain. The old pool is torn down first, the master re-execs itself, and only then does it fork replacement workers. For a measurable interval zero workers are serving requests.

If you page on a 502 spike, a listen-queue jump, or a failed ping probe during a deploy, you have probably already met this window. The pattern is narrow, predictable, and bounded by process_control_timeout. Treating it as an incident is one of the most common PHP-FPM false alarms.

What it is and why it matters

SIGUSR2 is the only signal that triggers a true graceful reload of PHP-FPM. SIGUSR1 reopens log files only and never touches workers. SIGHUP is not handled and defaults to OS termination.

When SIGUSR2 arrives, the master does the following, in order:

  1. Send SIGQUIT to every worker, asking them to finish the current request and exit.
  2. Wait up to process_control_timeout seconds.
  3. Escalate to SIGTERM for any worker still alive.
  4. Escalate to SIGKILL for any worker still alive after that.
  5. Re-exec itself via execvp(), re-reading its configuration.
  6. Fork new workers from the freshly re-exec’d master.

Step 5 is the part most operators miss. nginx’s reload model spawns new workers first, then drains the old pool in parallel. PHP-FPM inverts this: drain first, then re-exec, then fork. There is no overlap, and there is a hard moment in the lifecycle where no workers are bound to the FastCGI socket.

The listen socket file descriptor is preserved across execvp() and inherited by the new master. Connections that arrive during the window are not refused outright; they accumulate in the kernel listen backlog and are accepted once new workers come up. If the backlog fills first, the kernel drops them and the web server sees connection failures (502).

How it works

The diagram below shows the reload sequence. The horizontal gap between “old workers gone” and “new workers accept()ing” is the no-worker window.

sequenceDiagram
    participant Op as Operator / systemd
    participant M as Master
    participant OW as Old workers
    participant NW as New workers
    participant K as Kernel listen backlog

    Op->>M: SIGUSR2
    M->>OW: SIGQUIT
    Note over M,OW: wait process_control_timeout
(default 0) M->>OW: SIGTERM (still alive) M->>OW: SIGKILL (still alive) Note over M: zero workers bound K->>K: new conns pile in backlog M->>M: execvp() re-exec master
(listen fd inherited) M->>NW: fork workers NW->>K: accept() queued conns

The width of the no-worker window is roughly:

window = process_control_timeout
       + time to drain workers that ignore SIGQUIT
       + execvp() cost (config parse, pool init)
       + fork + PHP runtime init for the first workers

With the default process_control_timeout = 0, workers that do not exit on SIGQUIT are killed immediately on the next tick, so the drain component collapses. The remaining components (execvp, fork, runtime init) are usually well under a second on a warm box, but they are non-zero. Under memory pressure, on a slow filesystem, or with many pools, the window stretches.

Two configuration values dominate the window:

  • process_control_timeout: the grace period granted to workers before they are force-killed. Accepts unit suffixes (s, m, h, d); a bare integer is treated as seconds. Match it to max_execution_time and you defeat the purpose; keep it small.
  • listen.backlog: the kernel queue that absorbs connections during the window. The default is version- and platform-dependent: PHP < 8.2 uses 511 on Linux/macOS; PHP 8.2+ uses -1 there (clamped to net.core.somaxconn on Linux), while other platforms retain 511.

Where it shows up in production

The no-worker window is most visible during deploys, logrotate runs, and config pushes. Common surfaces:

  • systemctl reload php*-fpm in a deploy pipeline. The reload is fast but not instantaneous; any in-flight or arriving requests during the window queue or fail.
  • logrotate reloads. A postrotate hook that sends SIGUSR2 (instead of SIGUSR1, which only reopens logs) triggers a full pool recycle on every rotation. This is a classic misconfiguration. SIGUSR1 is the right signal for log rotation.
  • Config-only changes shipped via reload. These cost the same window as a code deploy.
  • Multiple reloads in rapid succession, e.g. a deploy tool that fires SIGUSR2 in a loop on health check failure. On PHP < 7.4 this could crash the master entirely (PHP bug #74083). The fix landed in PHP 7.4.0, and the PHP bug record does not document a backport. On supported PHP the master blocks signals around execvp(), but back-to-back reloads still amplify the no-worker window.

The window is also visible at the web server edge. nginx logs show connect() failed (11: Resource temporarily unavailable), upstream prematurely closed, or plain Connection refused during the window, all stamped within the same second or two as the reload.

When this matters (and how to bound it)

For most traffic patterns a sub-second no-worker window hidden behind a deep backlog is invisible. The window becomes a problem when any of these are true:

  • Your web server has a tight fastcgi_connect_timeout and treats a stalled accept as a 502 before the new workers come up.
  • Your traffic arrival rate is high enough to fill the listen backlog during the window.
  • You have workers that ignore SIGQUIT and continue running without an enforceable time limit, which can hold the master in the drain phase.
  • Your monitoring treats any ping failure or 502 as a page.

Mitigations, roughly in order of cost:

  • Set process_control_timeout deliberately. Default 0 minimises the drain component but kills workers mid-request the moment they ignore SIGQUIT. A small value (1-2 seconds) gives in-flight requests a chance to finish without stretching the window much.
  • Make sure logrotate uses SIGUSR1, not SIGUSR2. SIGUSR1 reopens logs without touching workers. SIGUSR2 is a full reload.
  • Avoid stacking reloads. One reload per deploy. A deploy tool retrying reloads on health-check failure is papering over a different problem.
  • Suppress availability alerts for ~120 seconds after an intentional reload. This is the single most effective noise reduction. Apply it to ping, 502 rate, and listen queue alerts, keyed to a reload or restart event.
  • Do not assume reload clears OPcache. Whether the OPcache shared memory segment survives execvp() depends on the PHP build and shared memory backend. Verify empirically with opcache_get_status() before and after a reload before tuning your deploy around either behaviour.

Signals to watch in production

These are the signals that move during a normal reload. Knowing their expected shape during the window is what separates “expected transient” from “real incident.”

SignalWhy it matters during reloadExpected shape vs. red flag
Active processesDrops to 0 as workers drainBrief zero, then climbs back. Sustained zero after 120s is a red flag.
Total processesHits 0 (or just the master) during the windowBrief dip, then restoration. If it never climbs, new workers are not spawning.
Listen queue depthAbsorbs arrivals while no workers can acceptBrief spike that drains as workers come up. Sustained growth past the window is saturation, not reload.
Pool ping responsePing queues behind the kernel backlog and may fail or stallBrief failure or elevated latency. Persistent failure past 120s means the master did not come back.
Accepted connections rateDrops during the window, then recoversBrief dip is expected. No recovery means workers are not accepting.
Web server 502/504 rateSpikes as connections fail or stallTight spike around the reload timestamp. Persistent elevation is a different incident.
PHP-FPM error logLogs NOTICE: Reloading ... and NOTICE: ready to handle connectionsReload lines bracket the window. Repeated Reloading lines without ready mean reloads are stacking.

The 120-second suppression window applies to the alerting signals above (ping, 502 rate, listen queue). Outside that window, every one of these signals means a real problem.

How Netdata helps

  • Per-second polling of PHP-FPM status fields catches the actual shape of the no-worker window. Ten-second polling will often miss the dip and recovery entirely, leaving you with an alert and no graph.
  • Correlating active processes, listen queue, and accepted connections on a single timeline makes the difference between “expected reload transient” and “real saturation” obvious at a glance.
  • Annotated reload events (from the master’s NOTICE: Reloading log lines) overlaid on the worker count chart let you confirm that the 502 spike and the reload share a timestamp.
  • ML anomaly detection on ping latency and 502 rate distinguishes a bounded reload-shaped spike from an open-ended anomaly.
  • Web server upstream error metrics alongside the PHP-FPM pool view let you confirm the failure is on the FPM socket and not the web server itself.