The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / php-fpm / php-fpm-child-exited-on-signal-sigsegv ▌

Operations Guides

PHP-FPM "child N exited on signal 11 (SIGSEGV)": worker segfaults

The log line is unambiguous:

WARNING: [pool www] child 12345 exited on signal 11 (SIGSEGV) after 3.4 seconds from start

A worker hit a memory protection fault and the kernel killed it. FPM forks a replacement, but the in-flight request is gone: the client sees a truncated response or a 502, and in a tight loop a whole pool can die to a crash storm. A SIGSEGV is almost never a bug in your PHP code. PHP-level fatal errors exit cleanly with a 500 and a stack trace. Signal 11 means native code dereferenced a bad pointer, so the fault lives in a C extension, OPcache, the JIT, or the PHP runtime itself.

This page covers how to read the log line, what the kernel can tell you that FPM cannot, how to capture a usable stack trace, and how to narrow a noisy crasher to a specific endpoint and cause.

What this means

FPM writes the line when a worker is killed by a signal it cannot recover from. The signal number classifies the failure:

SignalNameWhat it usually means
11SIGSEGVSegmentation fault: native code touched unmapped memory or violated protection. The classic extension or engine bug.
7SIGBUSMemory mapping problem, often a corrupted shared segment or a truncated memory-mapped file.
6SIGABRTAssertion failure inside a native library; the C runtime aborted on purpose.
9SIGKILLOOM kill or external kill. Not a segfault.

Three facts to internalize:

  1. dmesg is more informative than the FPM log. The FPM line says “child died with signal 11”; the kernel tells you the fault address, the instruction pointer, the offending module, and an error code that classifies the access.
  2. The crash is almost always triggered by a specific request. The FPM status page in ?full mode shows the request URI of each active worker, so the dying worker’s URI is the single best lead.
  3. emergency_restart_threshold is disabled by default. Without it, FPM respawns individual crashed workers forever and never does a full restart, which can mask a slow crash loop.

One crash a day is a bug to chase. Several a minute is a crash storm: prioritize mitigation (block the URI, disable a suspect extension, roll back a change) before forensics.

flowchart td
    A[SIGSEGV in FPM log] --> B{Rate?}
    B -- single or occasional --> C[Identify URI from status page]
    B -- storm --> D[Stabilize: block URI, disable suspect]
    C --> E[Check dmesg for fault address and module]
    D --> E
    E --> F{Recent change?}
    F -- yes --> G[Roll back PHP, extension, or opcache config]
    F -- no --> H[Capture core via rlimit_core, gdb backtrace]
    H --> I[Bisect extensions: disable, test, re-enable one by one]
    G --> J[Monitor crash rate after fix]
    I --> J

Common causes

CauseWhat it looks likeFirst thing to check
Native C extension bug (ddtrace, tideways, newrelic, Xdebug, ioncube, APCu or imagick combos)Crashes started after an extension upgrade, or only on certain request shapes. dmesg names the extension .so.php -m and the extension changelog. Disable non-core extensions and re-enable one by one.
Corrupted OPcache shared memory or opcache.file_cacheCrashes cluster after deploys that change file paths, or recur until FPM is restarted. Hit rate looks normal.Disable opcache.file_cache (keep opcache.enable=1) and reload.
JIT tracing bugCrashes only with opcache.jit enabled, often under specific workloads (CLI tools, long-running workers).Set opcache.jit=off (or disable) and reload.
Deep recursion or stack overflowOne endpoint consistently crashes the worker that handles it. Backtrace shows deep frames.Reproduce locally. Raising ulimit -s is a stopgap only.
PHP version regression after upgradeCrashes began within hours of a PHP upgrade; reproduces on a clean app.Roll back to the previous PHP patch release and read the changelog for segfault fixes.
CVE or remote triggerCrashes correlate with specific inbound request patterns (TLS handshake, proxy use).Check the PHP changelog and security advisories. Apply the next patch release.

Quick checks

Run these read-only before changing anything. They tell you the rate, the failing endpoint, and what the kernel saw.

# Crash rate from the FPM error log in the current hour
grep "exited on signal 11" /var/log/php-fpm/error.log | grep "$(date '+%d-%b-%Y %H')" | wc -l

# Kernel-side detail the FPM log omits: fault address, IP, error code, module
dmesg --since "1 hour ago" | grep -A2 -B2 php-fpm
journalctl -k --since "1 hour ago" | grep php-fpm

# Crashing endpoint from the per-worker request URI (full status)
curl -s 'http://127.0.0.1/fpm-status?full' | grep -E "request URI|state"

# Process count to spot a crash loop (compare against pm.max_children)
pgrep -cf "php-fpm: pool"

# Is the emergency-restart circuit breaker even configured?
php-fpm -tt 2>&1 | grep -E "emergency_restart_threshold|emergency_restart_interval"

# Recent upgrades that line up with the start of the crashes
zgrep -h "php" /var/log/dpkg.log* /var/log/yum.log* 2>/dev/null | tail -20

The kernel error field in the dmesg segfault line is worth decoding. A common value is error 4: on x86_64 that is a user-space read of a non-present page, the typical signature of a null or stale pointer dereference inside an extension.

How to diagnose it

  1. Get the rate. A few crashes a day is a different problem from three a minute. If the accepted-connections rate is dropping at the same time, there is user impact. If total processes is bouncing, you are in a respawn loop.

  2. Identify the dying worker’s URI. Poll fpm-status?full at 1-second intervals during the crash window and capture the request URI of any worker in Running state immediately before a death. Crashes that always cluster on one URI point to a code path; crashes spread evenly across URIs point to engine-level state (OPcache, JIT, extension globals).

  3. Pull dmesg for fault detail. For each crashed PID, find the matching kernel segfault line. The pattern php-fpm[PID]: segfault at ADDR ip IP sp SP error N in MODULE[...] names the module where the fault happened. That alone narrows “PHP crashed” to “the tracing extension crashed” or “opcache.so crashed”.

  4. Line up changes. Compare the first crash timestamp against recent deploys, package upgrades, INI edits, and traffic shifts. Segfaults that begin within an hour of a change are usually that change. PHP patch releases, extension upgrades, and INI flips (opcache.jit=on, opcache.file_cache=/path) are the usual suspects.

  5. Capture a core if dmesg is not enough. Core dumps are disabled by default. To enable per pool, set rlimit_core = unlimited in the pool config and ensure the system core_pattern writes somewhere with space:

# Check the current core pattern
cat /proc/sys/kernel/core_pattern

# On systemd distros this typically pipes to systemd-coredump; read with:
coredumpctl list
coredumpctl info <pid>

# To write plain core files instead (until next reboot).
# WARNING: this changes core_pattern globally for the whole host, not just FPM.
echo '/tmp/core-%e.%p' | sudo tee /proc/sys/kernel/core_pattern

Reload FPM with SIGUSR2, reproduce the crash, and inspect the dump:

# Load the core against the PHP binary that produced it
gdb /path/to/php-fpm /tmp/core-php-fpm.<pid>

# Inside gdb
(gdb) bt full
(gdb) thread apply all bt

Core dumps are large (often hundreds of MB) and fill disk fast under a crash loop. Disable collection as soon as you have one or two good stacks.

  1. Bisect extensions. If the backtrace names an extension, version-check it against the upstream changelog. If the backtrace is inconclusive, disable every non-core extension and re-enable them one at a time, exercising the crashing URI between each step. Keep OPcache in the picture but toggle opcache.file_cache and opcache.jit separately.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Worker exit rate (signal 11 only)Direct measure of crash frequency. Sustained non-zero rate is abnormal.More than one SIGSEGV per minute, or any step up from a flat baseline.
Emergency restart eventsFPM crossed emergency_restart_threshold and execvp()’d itself. Users see a brief 502 window.grep "failed processes threshold" /var/log/php-fpm/error.log returns anything. Two or more in 15 minutes is a persistent crasher.
Total process countCrash-loop signature: count oscillates instead of staying flat (static) or tracking demand (dynamic).active + idle consistently below pm.max_children in a static pool.
Per-worker request URIIdentifies the endpoint that triggers the crash; the highest-value field during an incident.One URI dominating the Running set right before each crash.
Accepted connections rateDrops when workers cannot stay alive long enough to accept work.Throughput falling while upstream traffic holds steady.
dmesg segfault linesProvides fault address, instruction pointer, and module name that the FPM log omits.Any new php-fpm[...]: segfault at ... line.

Fixes

Treat these as mitigation first, root cause second. A SIGSEGV is a native bug; the durable fix is almost always a configuration change, an extension upgrade, a PHP patch release, or a code change to avoid the trigger.

OPcache file_cache corruption

If crashes cluster after symlink-style releases or never resolve until FPM is fully restarted, disable opcache.file_cache and keep opcache.enable=1. File-cache corruption after path-changing deploys is a well-repeated production failure mode. Clearing the cache directory and reloading is the short-term fix; leaving opcache.file_cache off is the long-term fix on deployment patterns that move file paths.

JIT tracing bugs

If crashes coincide with opcache.jit being on, set opcache.jit=off (or disable) and reload. JIT tracing faults are typically workload-specific, so a clean test suite will not reproduce them. Disable JIT first, confirm the crash stops, then track the upstream bug.

Extension bugs

If dmesg or the gdb backtrace names an extension, version-check it. The standard sequence: snapshot loaded extensions (php -m), disable every non-core extension, reload, exercise the crashing URI, and re-enable one at a time. Common offenders are tracing APM extensions (ddtrace, tideways, newrelic), ioncube loaders, and certain APCu or imagick version combinations. The extension changelog usually lists segfault fixes by version.

Stack overflow

If the backtrace is dominated by hundreds of repeating frames, raising the stack size limit is only a stopgap. The real fix is in the code: unbounded recursion, recursive serializers, or template engines that recurse on user-supplied input.

PHP version regression

If the first crash lines up with a PHP upgrade, roll back to the previous patch release, then read the upstream changelog for segfault fixes in newer releases. Test upgrades in staging with realistic traffic replay before promoting. Native crashes rarely surface under low-volume smoke tests.

CVEs and remote triggers

If crashes correlate with specific inbound traffic (a TLS handshake, a proxied HTTPS request, a malformed body), treat it as a potential security issue. Check the PHP changelog and security advisories, and stage the next patch release. Network-level rate limiting or WAF rules can buy time while the patch rolls out.

Prevention

  • Configure emergency_restart_threshold and emergency_restart_interval. Both default to 0 (disabled). A reasonable starting point is emergency_restart_threshold = 10 with emergency_restart_interval = 60, giving FPM a circuit breaker against a runaway crasher.
  • Alert on the SIGSEGV rate, not the absolute count. A counter that drifts up over a week is as actionable as a sudden spike; a daily absolute hides both.
  • Track FPM, PHP, and extension versions in your inventory. When a segfault wave starts, the first question is “what changed”. Without version history you cannot answer it.
  • Test PHP upgrades with traffic replay. Segfaults are workload-specific. A smoke test of the top three endpoints will miss a crash that only fires on the fortieth.
  • Keep core dump plumbing ready but off. Know your core_pattern, budget the disk space, and rehearse the gdb commands before you need them at 3 a.m.
  • Pin extensions explicitly. Auto-updated extensions have introduced segfaults in the wild. Treat extension upgrades with the same change-management rigor as PHP upgrades.

How Netdata helps

  • Per-second worker-death tracking shows a crasher starting within seconds instead of after a polling delay, and the rate-of-change view distinguishes a single bad request from a storm.
  • Anomaly detection on the SIGSEGV rate surfaces a slow upward drift that a fixed threshold would miss, and on the accepted-connections rate catches the throughput drop that often precedes user-visible 502s.
  • The full-status poll captures the per-worker request URI, so the dying worker’s endpoint is recorded alongside the crash instead of reconstructed from logs after the fact.
  • Correlating crash rate against deployment markers, package installs, and PHP version changes answers the “what changed” question without manual log archaeology.
  • Emergency restart events are tracked as discrete signals rather than buried in log volume, so a brief self-heal does not pass silently.