The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / php-fpm / php-fpm-504-gateway-timeout ▌

Operations Guides

PHP-FPM 504 Gateway Timeout: requests accepted but never finishing in time

A 504 from nginx means the FastCGI connection was established, the request was handed to a worker, and the worker never produced a response before nginx gave up. This is a different failure from a 502, where nginx received an invalid, early-close, or no response from the upstream (for example, because the socket was missing, the backlog was full, or the master was dead). The FastCGI channel was healthy. The work was not.

The default fastcgi_read_timeout in nginx is 60 seconds, measured between successive read operations on the socket, not for the response as a whole. Once it fires, nginx closes its side and returns 504. PHP-FPM has no idea this happened. The worker keeps executing, builds the full response, and writes it to a socket nobody is reading. This is the phantom worker problem, and it is the central operational trap of the 504 symptom.

The fix is never just “raise the timeout.” The timeout is doing its job: it is telling you that a subset of requests exceeds your user-facing latency budget. The actual work is to identify which requests are slow, why they are slow, and align the three timeout layers (PHP max_execution_time, FPM request_terminate_timeout, nginx fastcgi_read_timeout) so stuck workers are reclaimed instead of left running on abandoned requests.

What this means

The worker pool may not even be saturated. A single endpoint that hangs for 90 seconds produces 504s for the users hitting it while the rest of the site stays fast. The signature is:

  • nginx error log shows upstream timed out lines referencing the FastCGI socket.
  • PHP-FPM ping answers. The status page renders.
  • active processes is nowhere near max_children.
  • The slow log (if configured) shows the same endpoints repeatedly.

The failure mechanism is a timeout cascade across three independent clocks that almost never start aligned.

flowchart TD
    A[Request enters FPM worker] --> B{Work is CPU-bound?}
    B -->|Yes| C[max_execution_time 30s kills it]
    B -->|No, I/O blocked| D[max_execution_time does NOT fire]
    D --> E{request_terminate_timeout set?}
    E -->|0, disabled| F[Worker runs until script ends]
    E -->|e.g. 30s| G[Master kills worker]
    F --> H{nginx fastcgi_read_timeout fires first?}
    H -->|Yes| I[nginx returns 504 to client]
    I --> J[Worker keeps running: phantom worker]
    H -->|No| K[Response delivered to client]

The three clocks that matter:

  • PHP max_execution_time (default 30s). On typical non-Windows PHP-FPM builds it counts CPU time, not wall-clock. sleep(), network I/O, database queries, and system calls do not count toward it. If PHP was built with --enable-zend-max-execution-timers (ZTS builds default to it since PHP 8.3), elapsed wall time including blocking waits counts instead. This build-dependence is a common source of confusion in 504 diagnosis.
  • PHP-FPM request_terminate_timeout (default 0, disabled). When set, the FPM master kills the worker after the configured wall-clock duration. Enforced at the master level: ini_set() and set_time_limit() cannot extend it. If it is 0, nothing in PHP-FPM will ever kill a stuck worker.
  • nginx fastcgi_read_timeout (default 60s). Measured between successive reads, not for the total response. A script that trickles partial output can keep resetting this timer and run far past the nominal 60 seconds.

The recommended ordering is max_execution_time no later than request_terminate_timeout, with fastcgi_read_timeout highest. If fastcgi_read_timeout is lower than request_terminate_timeout, nginx returns 504 while the FPM worker keeps running: the phantom worker scenario.

Common causes

CauseWhat it looks likeFirst thing to check
Slow upstream dependency (database, external API, Redis, NFS)Slow log stack traces show PDO::*, curl_exec, Redis::*, or flock at the top; per-worker request duration is bimodalThe dependency’s own metrics: DB slow query log, API latency, Redis latency
request_terminate_timeout = 0 with a stuck requestactive processes stays flat, one or two workers show request duration climbing into minutes with no corresponding CPUPool config: the request_terminate_timeout line
fastcgi_read_timeout lower than request_terminate_timeoutnginx returns 504 but FPM status shows the worker still Running the same URI seconds laternginx fastcgi_read_timeout vs pool request_terminate_timeout
max_execution_time treated as wall-clockScript blocks on network I/O, never hits the 30s limit, runs until nginx times outphp.ini max_execution_time and where the script actually spends time (slow log)
Session lock contention (file-based sessions)Slow log shows session_start near the top; affects only same-session concurrent requestslsof on session files; session_write_close() placement in code

Quick checks

Read-only and safe to run during an incident. Adjust paths and status URLs to match your distribution and configuration.

# Confirm 504s are FastCGI timeouts, not proxy timeouts
grep "upstream timed out" /var/log/nginx/error.log | tail -20

# nginx: 504 count in the current hour
grep "$(date '+%d/%b/%Y:%H')" /var/log/nginx/access.log | grep -c " 504 "

# FPM: is the master alive and the socket listening?
pgrep -f "php-fpm: master" && ss -lxn | grep php

# FPM: pool saturation snapshot
curl -s http://127.0.0.1/fpm-status

# FPM: per-worker detail, including request duration and script
curl -s 'http://127.0.0.1/fpm-status?full' | grep -E "pid|state|request duration|request URI|script"

# FPM: slow log tail (only useful if request_slowlog_timeout is set)
tail -100 /var/log/php-fpm/slow.log

# FPM: verify pool-level timeout configuration
php-fpm -tt 2>&1 | grep -E "request_terminate_timeout|request_slowlog_timeout"

# PHP: max_execution_time from the FPM SAPI (CLI defaults to 0/unlimited, do not use php -i)
grep -r max_execution_time /etc/php/*/fpm/ 2>/dev/null

If request_slowlog_timeout is 0 or absent from the php-fpm -tt output, the slow log is disabled. You will not be able to localize the slow path from PHP-FPM alone until you enable it.

How to diagnose it

  1. Confirm the 504 is a FastCGI timeout, not a proxy timeout. The nginx error line should reference the FastCGI socket or the fastcgi_pass upstream, not proxy_pass. If you see upstream timed out against a proxy_pass upstream, this article is not the right diagnostic path.

  2. Capture the slow path before it moves. Pull the full status page and look for workers in Running state with request duration values above your p99. Note the script and request URI fields. The duration field is in microseconds.

  3. Read the slow log. If request_slowlog_timeout is set (for example, 5s), every entry shows the script path and a PHP backtrace captured when the master sent SIGSTOP to the worker. The frame at the top of the trace is where the worker was blocked. Common signatures: PDO::query or PDO::prepare (database), curl_exec or file_get_contents on an HTTP URL (external API), Redis::* (cache), flock or session_start (session lock).

  4. Check for phantom workers. Compare the timestamp of a recent nginx 504 with the FPM full status. If the same worker is still Running the same URI after nginx has already returned 504, you have phantom workers. This confirms a timeout ordering problem: fastcgi_read_timeout is firing before request_terminate_timeout (or before the script finishes, if request_terminate_timeout is 0).

  5. Verify the three timeout layers and their ordering. Check max_execution_time in php.ini, request_terminate_timeout in the pool config, and fastcgi_read_timeout in nginx. Keep max_execution_time no later than request_terminate_timeout, and keep fastcgi_read_timeout highest. Any other ordering produces either premature 504s or phantom workers.

  6. Correlate with the dependency. If the slow log points at the database, check the database slow query log and connection count at the same timestamp. If it points at an external API, check the API’s latency from the app host. The goal is to confirm the slow path is upstream of PHP, not inside PHP itself.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
nginx 504 rateThe user-facing symptom itselfAny sustained non-zero rate on PHP endpoints
Per-worker request duration (full status)Localizes which workers are stuck and on which scriptWorkers with durations more than 10x the p99 baseline
Slow log entry rateTells you what is slow, not just that something is slowSudden increase above baseline; concentration on specific scripts
active processes / max_children ratioRules worker exhaustion in or out as a compounding factorSustained above 80% means slow requests are also consuming pool capacity
request_terminate_timeout configurationDetermines whether stuck workers are ever reclaimedValue of 0 means stuck workers run until the script ends or the pool restarts
Dependency latency (DB, API, cache)The root cause in most 504 incidentsSpikes that line up with the nginx 504 rate

Fixes

Align the three timeout layers

The ordering must keep fastcgi_read_timeout highest; max_execution_time can equal request_terminate_timeout. A common working configuration:

  • max_execution_time = 30 (php.ini). CPU-bound work is killed with a fatal error that produces a trace.
  • request_terminate_timeout = 30 (pool config). Wall-clock safety net. The master kills the worker, so there is no PHP-level trace, but the worker is reclaimed. Cannot be overridden by ini_set() or set_time_limit().
  • fastcgi_read_timeout = 60 (nginx). Gives nginx a window longer than FPM’s terminate timeout, so FPM reclaims the worker before nginx gives up. This avoids phantom workers.

Some operators set request_terminate_timeout slightly higher than max_execution_time to give CPU-bound scripts a chance to hit max_execution_time first, since the fatal error produces a more useful trace than a master kill. The exact values are application-dependent. The ordering is not.

If you have long-running endpoints (webhooks, synchronous report generation, exports), exclude them from the global timeout policy. The right answer is usually a job queue, not a global timeout high enough to accommodate the slowest endpoint.

Enable the slow log

In the pool config:

request_slowlog_timeout = 5
slowlog = /var/log/php-fpm/slow.log

Set the threshold low enough to catch the requests causing 504s but high enough that normal traffic does not flood the log. For an application with a p99 around 200ms, a threshold of 5 seconds catches only pathological requests. For a heavier application, 10 seconds may be more appropriate. The slow log mechanism uses SIGSTOP/SIGCONT to capture the trace, which adds a small delay to the already-slow request. Under extreme conditions (thousands of slow requests per second), the logging I/O itself can compound the problem.

Address the actual slow path

The timeout alignment stops the bleeding. The slow log tells you where to look. Common fixes:

  • Database: add missing indexes, fix lock contention, or move the query behind a connection pool (pgbouncer, ProxySQL) if the bottleneck is connection acquisition rather than query execution.
  • External API: set an explicit, low timeout in the HTTP client. PHP’s cURL default has no timeout (infinite) unless your application sets one. Add a circuit breaker so a failing API does not convert every FPM worker into a waiting thread.
  • Session locks: call session_write_close() as early as possible in the request lifecycle, or switch to Redis or Memcached session handlers, which have different locking semantics.
  • NFS or shared filesystem: check for stall or high latency on the mount. Consider local caching of the files in question.

Reclaim phantom workers during the incident

If you are mid-incident and workers are stuck on abandoned requests, you can reclaim individual workers:

# Identify long-running workers from full status, then send SIGQUIT (graceful)
kill -SIGQUIT <worker_pid>

SIGQUIT tells the worker to finish the current request and exit. For a worker already running a phantom request that nobody will read, this ends the useless work. The master respawns a replacement. Avoid SIGKILL on individual workers unless the worker is not responding to signals. Note: if the worker is blocked in a system call (network I/O, disk I/O), the signal handler will not run until the call returns or is interrupted.

Prevention

  • Enable request_slowlog_timeout on every production pool. It is the single most direct signal for localizing slow work. Without it, 504 diagnosis degrades into correlating nginx error timestamps with APM traces and database logs by hand.
  • Set request_terminate_timeout to a finite value. A value of 0 means any stuck request (deadlock, infinite loop, hung socket) permanently occupies a worker until the next FPM restart.
  • Order the three timeouts correctly and document why. max_execution_time no later than request_terminate_timeout, with fastcgi_read_timeout highest. Record the chosen values and reasoning so the next operator does not undo the alignment.
  • Monitor the slow log entry rate, not just the counter. The counter resets on restart. Alert on rate of change and on concentration against specific scripts.
  • Route long-running endpoints through a job queue or a dedicated nginx location. Do not pollute the global timeout policy with per-endpoint exceptions.
  • On PHP 7.3+, consider request_terminate_timeout_track_finished. Defaults to no, so the terminate timeout stops applying at request end: after fastcgi_finish_request() and while shutdown functions run. Setting it to yes applies the timeout unconditionally through that finishing stage in current PHP 7.3–8.5 sources.

How Netdata helps

  • Per-second collection of active processes, idle processes, listen queue, and max children reached from the FPM status page lets you see the saturation signature of a slow-request cascade as it forms, not minutes later.
  • The PHP-FPM collector surfaces the slow requests counter as a rate, so a spike in slow log entries is visible alongside the nginx 504 rate without manual log correlation.
  • Correlating nginx upstream error rate with FPM pool saturation in a single view confirms whether a 504 is a slow-worker problem (this article) or a connection-refused problem (worker exhaustion or listen backlog overflow).
  • ML-based anomaly detection on active process count and average request duration can surface the bimodal latency distribution that precedes user-visible 504s.