The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / php-fpm / php-fpm-upstream-prematurely-closed ▌

Operations Guides

PHP-FPM "upstream prematurely closed connection while reading response header"

nginx logs this error at [error] level when the FastCGI connection to PHP-FPM closes before response headers arrive. The client gets a 502 Bad Gateway. Unlike connect() failed (111: Connection refused) (no worker available) or upstream timed out (110: Connection timed out) (worker too slow), this error means the connection was established, the worker began executing PHP, and the worker vanished before delivering a response.

The worker process died mid-request. The PHP-FPM master reaps the dead child via SIGCHLD and forks a replacement, but the in-flight request is lost. nginx sees the socket close while still waiting on headers and logs “upstream prematurely closed connection while reading response header from upstream”.

Four root-cause families: worker segfault (SIGSEGV or SIGBUS, usually an extension bug or memory corruption), OOM kill (SIGKILL from the kernel or cgroup), request_terminate_timeout firing (PHP-FPM itself killed the worker for exceeding the per-request hard limit), or a PHP fatal where error output exceeded nginx’s FastCGI buffer or the worker aborted before flushing headers. Diagnosis is timestamp correlation across the nginx error log, the FPM error log, dmesg, and the application error log.

What this means

The failure is per-request. Only the request running in the dead worker is affected. A single 502 in the log is usually a one-off segfault on an unusual input. Sustained 502 rates imply a repeating trigger: a code path that segfaults on every call, memory pressure that kills workers as soon as they grow, or a request_terminate_timeout set lower than your slowest legitimate endpoints.

Distinguish this from capacity failures. Worker exhaustion and listen-queue overflow produce connect() failed or no live upstreams, not “upstream prematurely closed”. If you treat a crash-driven 502 as a capacity problem and raise pm.max_children, you will not fix the crashes; on a memory-constrained host you will accelerate OOM kills.

pm.max_requests recycling is not a cause of this error. Workers that hit pm.max_requests finish their current request and deliver the response before self-terminating, then the master spawns a replacement. There is no mid-request interruption.

flowchart TD
    A[nginx accepts request] --> B[FastCGI conn to FPM worker]
    B --> C[Worker starts executing PHP]
    C --> D{Worker dies mid-request}
    D -- SIGSEGV / SIGBUS --> E[Extension bug or memory corruption]
    D -- SIGKILL --> F[OOM killer or cgroup limit]
    D -- request_terminate_timeout --> G[Per-request hard timeout fired]
    D -- PHP fatal / abort --> H[Fatal error, headers never flushed]
    E --> I[Socket closes before headers]
    F --> I
    G --> I
    H --> I
    I --> J[nginx logs "upstream prematurely closed"]
    J --> K[Client receives 502]

Common causes

CauseWhat it looks likeFirst thing to check
Worker segfault (SIGSEGV/SIGBUS)FPM log: child N exited on signal 11 or signal 7. Same URI repeats across deaths.request uri in full status page; recent extension or PHP upgrade.
OOM kill (SIGKILL)FPM log: child N exited on signal 9. dmesg shows Out of memory: Kill process.Per-worker RSS, system/cgroup memory, pm.max_requests.
request_terminate_timeout firedFPM log: execution timed out ... terminating. Hits the same long-running endpoints.Compare timeout value against slowest endpoint durations and request_slowlog_timeout.
PHP fatal / oversized error outputApplication log shows fatal at same timestamp. Often with XDebug or verbose stack traces in production.nginx fastcgi_buffer_size and PHP log_limit; production error_reporting.
Extension abort (SIGABRT/SIGBUS)FPM log: exited on signal 6 or signal 7. Correlates with a specific PHP version bump.Recent extension upgrade; rlimit_core and process.dumpable for core dumps.

Quick checks

All read-only. Scope to the incident window.

# nginx errors around the incident timestamp (adjust path to your distro)
grep "upstream prematurely closed" /var/log/nginx/error.log | tail -20

# Distinguish from related nginx upstream errors
grep -E "connect\(\) failed|no live upstreams|upstream timed out" /var/log/nginx/error.log | tail -20

# PHP-FPM worker deaths (adjust log path - common: /var/log/php-fpm/error.log or www-error.log)
grep "exited on signal" /var/log/php-fpm/error.log | tail -30

# Segfaults and bus errors specifically
grep -E "signal 11|signal 7|SIGSEGV|SIGBUS" /var/log/php-fpm/error.log | tail -30

# request_terminate_timeout kills
grep "execution timed out" /var/log/php-fpm/error.log | tail -30

# OOM kills from the kernel (signal 9 deaths usually originate here)
dmesg -T | grep -iE "out of memory|oom-kill|killed process" | tail -30
journalctl -k --since "1 hour ago" | grep -i oom

# Current pool state (path depends on pm.status_path config)
curl -s http://127.0.0.1/fpm-status

# Effective per-pool timeout settings
php-fpm -tt 2>&1 | grep -E "request_terminate_timeout|request_slowlog_timeout|pm.max_requests|emergency_restart"

If request_terminate_timeout is 0 in the output, that family is ruled out. The deaths are coming from crashes or OOM.

How to diagnose it

  1. Anchor on a single nginx error timestamp. Pick one 502 from the log and note the second-precise time. Every check below is filtered to that window.

  2. Pull FPM child-exit entries for that window. A signal-11, signal-7, signal-9, or execution timed out entry within a few seconds of the nginx error confirms a worker died on that request. The exit signal tells you the family: 11 or 7 is a crash, 9 is OOM, the timeout message is request_terminate_timeout.

  3. Cross-reference dmesg for OOM kills. A signal-9 death with a matching Out of memory: Killed process line in dmesg (or memory.events.oom_kill in the cgroup on containers) confirms memory pressure, not an extension bug. The absence of an OOM entry with a signal-9 death should prompt a closer look at systemd KillMode/OOMPolicy and orchestrator eviction.

  4. Pull the per-worker request uri from full status during recurrence. If crashes cluster on a single endpoint, the trigger is in that code path or its inputs. Use curl -s http://127.0.0.1/fpm-status?full and watch which URI is in the Running state immediately before each death.

  5. Check the application error log at the same timestamp. A PHP fatal (Fatal error: ..., Allowed memory size of ... exhausted) recorded at the moment of the 502 means the worker aborted on an application condition. If the fatal produced a large stack trace, suspect nginx fastcgi_buffer_size being too small to hold the headers plus the error output.

  6. Look for repeating crash loops. Many exited on signal entries in a tight window, combined with the master logging failed processes threshold ... reached, initiating reload (when emergency_restart_threshold is configured), means a crash loop, not a one-off. Roll back the most recent deployment or extension upgrade.

  7. Verify timeout coherence across the request path. nginx fastcgi_read_timeout, PHP-FPM request_terminate_timeout, and PHP max_execution_time should be coherent. If request_terminate_timeout is set lower than your slowest legitimate endpoint, it will kill workers on healthy slow requests and surface here.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Worker exit rate by signalSeparates crashes (11/7) from OOM (9) from timeout kills from normal recycling.Sustained non-zero rate of signal 11/7/9 outside reload windows.
Per-worker RSSLeading indicator for OOM-driven signal-9 deaths.Monotonic growth with pm.max_requests = 0.
total processes vs configured pmCrash loops leave the pool below max_children while the master plays catch-up.Count dropping below expected mid-traffic in static mode.
request_terminate_timeout kill countEach kill is a 502 to the client that requested that endpoint.Counter climbing on specific slow endpoints.
System/cgroup memory pressureOOM kills happen at the kernel level, invisible to FPM logs except as a signal-9 child exit.Available memory trending toward zero; cgroup memory.events.oom_kill incrementing.
Emergency restart eventsWhen configured, repeated triggers indicate a systemic crash source.Two or more “failed processes threshold” entries within 15 minutes.
Slow log entriesIdentifies which endpoints approach the request_terminate_timeout boundary.Slow-log stack traces matching the killed request URIs.
nginx 502 rateThe user-facing symptom; the earliest external signal.Any sustained non-zero 502 rate on PHP routes.

Fixes

Worker segfaults (signal 11 or 7)

The crash is almost always in a C extension or the PHP runtime itself, not in PHP userland. Identify the trigger URI from full status, then narrow the cause.

  • If a recent PHP or extension upgrade preceded the crashes, roll back and re-test before bisecting.
  • Enable core dumps for definitive stack traces: set rlimit_core = unlimited in the pool config and ensure process.dumpable = yes (needed when the worker runs under a different user/group than the master). Then inspect with coredumpctl or your configured core path.
  • If the crash is reproducible on a single endpoint, block or rate-limit that route at the nginx layer while the extension is fixed.
  • Do not paper over a segfault by raising emergency_restart_threshold. That buys availability at the cost of masking the bug.

OOM kills (signal 9)

The worker was killed by the kernel or cgroup because it (or the pool in aggregate) exceeded the memory budget.

  • Verify the binding constraint. On bare metal, compare avg_worker_RSS * pm.max_children + OS_overhead against total RAM. In containers, check the cgroup memory.max against memory.current and memory.events.oom_kill.
  • Set pm.max_requests to a finite value (500 to 1000 is a common starting range) to bound per-worker growth. With the default of 0, leaks accumulate indefinitely.
  • Use PSS rather than RSS for capacity math. Shared opcache pages inflate RSS by 30 to 50 percent. Pull PSS from /proc/<pid>/smaps_rollup or smem.
  • If a specific request path spikes memory (large result sets, image processing, unbounded caches), cap it in application code or raise memory_limit only for that pool with eyes on the RSS impact.

request_terminate_timeout firing

PHP-FPM killed the worker because the request exceeded the configured hard limit. This is intentional behavior, but it produces this error.

  • Confirm the timeout value is intentional. Default in upstream PHP is 0 (disabled). Some distro or site packages ship a non-zero value. Check the effective value with php-fpm -tt rather than assuming.
  • Compare the timeout against the slowest legitimate endpoint. If request_terminate_timeout is 30 seconds and your report endpoint legitimately takes 45, either raise the limit for that pool or move the workload to a queue.
  • If long-running shutdown handlers or fastcgi_finish_request() work is surviving the timeout, enable request_terminate_timeout_track_finished (available since PHP 7.3.0) so the limit covers the full request lifecycle.
  • Coordinate with nginx fastcgi_read_timeout. If nginx times out first, you get a 504 and a phantom worker still running an unwanted response; if FPM kills first, you get this 502.

PHP fatal or oversized error output

The worker hit a fatal error and either aborted or produced headers plus error output that exceeded nginx’s FastCGI buffer.

  • Check the application error log at the incident timestamp for the actual fatal. A missing fatal entry suggests the worker died before the error handler ran, pointing back to a crash.
  • If fatals only reproduce with XDebug or high error_reporting in production, disable XDebug in production and lower error_reporting to production-appropriate levels. Large stack traces are the common cause of header overflow.
  • Increase nginx fastcgi_buffer_size and fastcgi_buffers if legitimate error output is filling them. The defaults are platform-dependent (commonly 4k or 8k).
  • Raise PHP-FPM log_limit (PHP 7.3+, default 1024) so stack traces in the FPM log are not truncated mid-investigation.

Prevention

  • Set pm.max_requests to a finite value. Default 0 means unbounded growth and is the enabler for most OOM-driven premature closes.
  • Set request_terminate_timeout deliberately. Either disable it (0) and rely on max_execution_time plus application timeouts, or pick a value above your slowest legitimate endpoint. Documenting the choice prevents silent regressions.
  • Configure emergency_restart_threshold and emergency_restart_interval. Defaults are 0 (disabled). A reasonable starting point like emergency_restart_threshold = 10 and emergency_restart_interval = 60 gives the master a circuit breaker for crash loops instead of forking into oblivion.
  • Enable request_slowlog_timeout. The slow log is the only signal that tells you which code path approaches the request_terminate_timeout boundary before users see 502s.
  • Track per-worker RSS over time. Catch the leak before the OOM killer does. PSS is more accurate than RSS for the capacity math.
  • Keep PHP and extensions current on a supported branch. Segfault-causing bugs in older extension versions are a frequent root cause after a PHP upgrade exposes them.
  • Verify timeout coherence on every nginx or FPM config change. fastcgi_read_timeout, request_terminate_timeout, and max_execution_time must tell a consistent story across the request path.
  • Monitor kernel OOM events and cgroup memory.events.oom_kill. FPM only sees these as a signal-9 child exit after the fact.

How Netdata helps

  • The PHP-FPM collector surfaces active processes, idle processes, total processes, listen queue, and max children reached per second, so a crash-loop drop in total processes is visible in the same window as the nginx 502 spike rather than minutes later.
  • Per-process RSS collection (via the apps or cgroup plugins, depending on deployment) lets you see the monotonic per-worker memory climb that precedes signal-9 OOM deaths, alongside system and cgroup memory pressure.
  • The nginx collector surfaces upstream error counts and 4xx/5xx response rates, so the 502 pattern is correlatable with FPM worker exits in a single timeline.
  • ML anomaly detection on total processes, accepted connection rate, and worker exit rate flags the deviation from baseline that precedes a crash-driven 502 wave, which is useful when the underlying segfault is intermittent.
  • Kernel OOM kill counters and cgroup memory.events.oom_kill are collected directly, so a signal-9 child exit can be matched to a kernel-level kill event without switching tools.