The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / php-fpm / php-fpm-slow-log-setup ▌

Operations Guides

PHP-FPM slow log: turning on request_slowlog_timeout to see what is slow

When PHP-FPM workers pile up, the status page tells you they are busy. It does not tell you why. Active processes climb toward pm.max_children, the listen queue fills, and you end up correlating timestamps against database slow query logs, APM traces, and application error logs to reconstruct what happened. The slow log closes that gap. When a request exceeds request_slowlog_timeout, PHP-FPM snapshots the worker and writes a backtrace naming the exact script and call that blocked.

It ships disabled. request_slowlog_timeout defaults to 0: no slow log is ever written. This is the most common monitoring gap in PHP-FPM deployments. Teams know workers are busy but not what they are busy on. This guide covers enabling the slow log, tuning it, verifying it end to end, and avoiding the platform-specific traps (Docker capabilities, SELinux, user isolation) that leave the file empty even after you think you have turned it on.

Pick a value below your user-visible latency budget and below request_terminate_timeout. Five seconds is a reasonable starting point for most web workloads. Enable it on every production pool and stop guessing.

What this captures

When a worker exceeds request_slowlog_timeout, the master captures a snapshot of where the worker was executing:

sequenceDiagram
    participant W as Worker
    participant M as FPM master
    participant L as slowlog file
    Note over W: request exceeds request_slowlog_timeout
    M->>W: attach/read (PTRACE_ATTACH, pread, or Mach)
    M->>L: write timestamp, PID, script, line, backtrace
    M->>W: detach/resume

The mechanism is platform-specific. On Linux and BSD it uses PTRACE_ATTACH to stop the worker; alternative builds read /proc/<pid>/mem or use Mach VM APIs and send SIGSTOP explicitly. PHP-FPM reads the worker’s call stack, writes the trace to the slowlog path, then detaches or resumes the worker. This briefly pauses the already-slow request. The cost is negligible next to the diagnostic value: you learn not just that a request was slow but the function and file where it was caught.

Each slow log entry contains:

  • a timestamp and pool name
  • the worker PID
  • the script filename and line number where execution was caught
  • a backtrace of stack frames (depth governed by request_slowlog_trace_depth, default 20, available since PHP 7.2.0)
  • the request URI

The status page also exposes a cumulative slow requests counter. The counter is what you alert on. The file is what you read when the counter moves.

One caveat from the mechanism: writing the trace is a point-in-time snapshot. The worker may have spent most of its time blocked deeper in the call stack than where it was caught. Multiple entries for the same endpoint over time build a more reliable picture than any single trace.

Prerequisites

  • Pool configuration access: write access to the pool config, typically /etc/php/<version>/fpm/pool.d/www.conf.
  • Reload authority: enabling the slow log requires a graceful reload (SIGUSR2) or full restart. During SIGUSR2, PHP-FPM sends SIGQUIT to existing workers and re-execs the master after the last worker exits; it does not overlap old and new workers.
  • A writable slow log path: the master must be able to create and append to the slowlog file. The directory must be owned by the master process user and not world-writable.
  • ptrace permission: the master must be allowed to ptrace the worker. This is the most common failure mode. See “Common pitfalls”.

Procedure

  1. Open the pool configuration for the pool you want to instrument.

    # Locate the pool config for your PHP version
    ls /etc/php/*/fpm/pool.d/
    
  2. Set the timeout. Add or uncomment request_slowlog_timeout. Bare integers are interpreted as seconds; s, m, h, d suffixes are also accepted.

    request_slowlog_timeout = 5
    

    Five seconds is a reasonable starting point for most web workloads. Lower it to 1 or 2 seconds to catch latency regressions early. Raise it to 10 or 30 seconds if your application has legitimately long endpoints and the log is too noisy.

  3. Set the slow log path. The default varies by distribution and install prefix. Set it explicitly so you know where to look.

    slowlog = /var/log/php-fpm/www-slow.log
    
  4. Optionally tune trace depth. The default of 20 frames is usually enough. Raise it if you run a deep framework stack and the bottom of the trace is truncated.

    request_slowlog_trace_depth = 20
    

    This directive is available since PHP 7.2.0. On earlier versions the depth is not configurable.

  5. If the pool runs as a different user than the master, enable dumpable processes. When workers run as www-data but the master runs as root (the common case), ptrace from the master to the worker is blocked unless process.dumpable is set to yes. The directive defaults to no.

    process.dumpable = yes
    

    This directive is available in PHP 7.0.29, 7.1.17, 7.2.5, and later; on older point releases the different-user ptrace path may fail.

  6. Create the log directory and fix ownership.

    # Create the directory owned by the master user (often root)
    mkdir -p /var/log/php-fpm
    chown root:root /var/log/php-fpm
    chmod 0755 /var/log/php-fpm
    
  7. Validate the configuration before reloading.

    # Test config syntax without applying it
    php-fpm -tt 2>&1 | grep -E 'slowlog|request_slowlog_timeout|process.dumpable'
    

    php-fpm -tt tests the configuration and prints the resolved values. Do not skip this. A syntax error in the pool config will prevent the reload from completing and leave the master running the old config, or fail to start after a full restart.

  8. Reload PHP-FPM.

    # Graceful reload (SIGUSR2)
    systemctl reload php-fpm
    

    Schedule the reload during low traffic. SIGUSR2 is the documented graceful reload signal, but treat any FPM reload as potentially disruptive until you have observed behavior under your specific PHP version and traffic pattern.

Verifying it works

Do not assume the slow log works because the config reloaded. The ptrace path fails silently in several common environments. Verify end to end.

  1. Confirm the resolved config.

    # Verify the master actually parsed the directives
    php-fpm -tt 2>&1 | grep -E 'slowlog|request_slowlog_timeout'
    

    You should see your timeout as a non-zero value and your slow log path.

  2. Trigger a deliberately slow request. Drop a temporary test script that sleeps longer than your timeout. Remove it when done.

    <?php sleep(10); ?>
    

    Request it through the web server, or hit it directly via FastCGI.

  3. Check the slow log file.

    tail -50 /var/log/php-fpm/www-slow.log
    

    You should see an entry with the test script’s filename, the sleep() line, and a backtrace. If the file is empty but the status counter incremented, the master believed it logged but the write went to a stale file descriptor (see the log rotation pitfall) or the worker was killed before the trace finished.

  4. Check the slow requests counter.

    curl -s http://127.0.0.1/fpm-status | grep "slow requests"
    

    The counter should increment once per slow request. It is cumulative since pool start, so compare deltas, not absolute values.

  5. If the file is empty and the counter did not move, work through “Common pitfalls”. The most likely cause on containerized or SELinux-enforced hosts is a denied ptrace.

Common pitfalls

Empty slow log with no errors (Docker). Containers drop CAP_SYS_PTRACE by default. Without it, ptrace(ATTACH) fails with “Operation not permitted” and the slow log stays empty even though the master logs a WARNING that it is logging the request. Add the capability to the container:

docker run --cap-add=SYS_PTRACE ...

For Kubernetes, add SYS_PTRACE to the container’s security context capabilities. This is the single most reported slow-log issue in containerized deployments.

Empty slow log with no errors (SELinux). On RHEL and CentOS 7+, SELinux blocks ptrace even when process.dumpable = yes is set. The error in the FPM log is failed to ptrace(ATTACH) child: Operation not permitted (1). Generate and load a local policy module:

# Build a policy from denied audit entries
grep ptrace /var/log/audit/audit.log | audit2allow -M php_ptrace
semodule -i php_ptrace.pp

Empty slow log with no errors (user isolation). If the pool runs as a non-root user and the master runs as root, process.dumpable must be yes (see step 5 above). Without it, ptrace fails silently. This is common on shared hosting and cPanel-style deployments where the master and pool users differ.

request_terminate_timeout fires first. If request_terminate_timeout is set lower than request_slowlog_timeout, the master kills the worker before the slow log can capture the trace. Always keep request_slowlog_timeout below request_terminate_timeout (for example, 5s versus 30s). If you see workers dying with no slow log entry and request_terminate_timeout is configured, check the ordering.

Log rotation breaks the slow log. If the slow log is rotated without sending SIGUSR1 to the master, FPM keeps the old (now deleted) file descriptor open and writes go nowhere. Configure logrotate to signal the master:

postrotate
    kill -USR1 $(cat /run/php-fpm.pid 2>/dev/null) 2>/dev/null || true
endscript

If your slow log suddenly stops receiving entries after a logrotate run, this is why.

Worker stuck after tracing or status reads. Upstream issue php/php-src#7931 reports workers stuck in the Processing state with slow-log tracing enabled. A maintainer linked it to the scoreboard-access problem fixed by php/php-src#8049 (fpm_scoreboard_copy), which landed in PHP 8.0.17 and 8.1.4. If slow-log tracing or full-status reads leave workers in Processing, upgrade beyond those builds or disable the slow log until you can.

Log integrity note (CVE-2024-9026). A low-severity vulnerability (CVSS 3.3) in PHP-FPM allows limited log manipulation when catch_workers_output = yes is set. The advisory describes possible log pollution of up to four characters from worker output. It was patched in PHP 8.1.30, 8.2.24, and 8.3.12. It does not directly affect slow log integrity, but it is relevant when you are enabling or auditing FPM logging features. Run a patched version.

Signals to monitor

SignalWhy it mattersWarning sign
slow requests counter (rate)Direct indicator of requests exceeding the thresholdSustained non-zero rate, or rate above 2x rolling baseline
Slow log script and line concentrationLocalizes the slow code pathOne script or function appearing repeatedly across entries
active processes near pm.max_childrenShows slow requests are consuming worker capacityActive climbing in lockstep with the slow request rate
Listen queue depthConfirms slow requests are causing user-visible queuingNon-zero and growing while slow requests increment
Per-worker request duration (full status)Shows bimodal distribution: fast normal requests plus stuck outliersMultiple workers at 10x the median duration
Slow log entry volume per hourTracks regression or backend degradation over timeStep change after a deploy or dependency incident

Correlate the counter rate with the slow log file. The counter tells you when to look. The file tells you where to look.

How Netdata helps

  • Per-second polling of the slow requests counter catches rate spikes that a 10 or 30 second poll interval misses. PHP-FPM saturation events unfold in seconds, and the slow request rate is the leading indicator before the listen queue fills.
  • Correlation in one view: the slow request rate sits alongside active processes, idle processes, listen queue depth, and per-worker request duration, so you can see whether slow requests are driving worker exhaustion without pivoting between tools.
  • Anomaly detection on the slow request rate flags sudden shifts without hand-tuned thresholds, which matters because the right threshold depends on your application’s normal latency profile.
  • Rate-of-change alerts on max children reached and the listen queue complement the slow log: the slow log explains why workers are stuck, while the queue and max-children counters show when that stuckness is about to spill over into user-visible errors.
  • Per-pool visibility when you run multiple pools, so a slow log spike in one pool is not masked by healthy aggregates in another.