The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / uwsgi / uwsgi-worker-respawn-loop ▌

Operations Guides

uWSGI worker respawn loop: 'DAMN ! worker died :( trying respawn'

The uWSGI master logs DAMN ! worker N (pid: XXX) died, killed by signal S :( trying respawn ... when it receives SIGCHLD for a worker that exited unexpectedly. The next line is typically Respawned uWSGI worker N (new pid: XXXX). In a respawn loop, these pairs repeat rapidly: the master forks a replacement, the replacement dies, and the cycle continues.

The signal number S is your primary diagnostic clue. Signal 9 (SIGKILL) means OOM kill or harakiri. Signal 11 (SIGSEGV) means a C extension crashed. Signal 6 (SIGABRT) means an assert failed. A worker that dies less than one second after spawn is hitting a startup failure, not a request-time problem.

Distinguish a crash loop from normal worker recycling first. With max-requests configured, workers deliberately exit after serving N requests and respawn. That is healthy. A crash loop burns CPU on repeated fork and startup cycles while serving zero traffic.

What this means

uWSGI does not throttle crash-induced respawns aggressively. If the new worker hits the same fatal condition (a bad import, a C extension segfault, memory pressure), it dies again within milliseconds. The result is a tight loop: fork, import, crash, log, fork again.

Each iteration consumes CPU for the fork syscall and application startup. In lazy-apps mode, every respawn re-imports the entire application. A 4-worker loop with a 2-second startup time burns worker-seconds on startup while serving zero requests. Accepting worker count fluctuates or stays at zero, throughput collapses, and the kernel listen queue fills with connections nobody will serve.

One non-crash case produces the same log pattern: harakiri. When harakiri fires, the master kills the worker with SIGKILL and respawns it. The log shows killed by signal 9, but you will also see HARAKIRI ON WORKER N (pid: XXXX, try: 1) !!!. If you see signal 9 without a HARAKIRI line, suspect the OOM killer.

flowchart TD
    A["DAMN ! worker died"] --> B{"Signal in log?"}
    B -->|"9 SIGKILL"| C{"HARAKIRI log present?"}
    B -->|"11 SIGSEGV"| D["Segfault in C extension or libc"]
    B -->|"6 SIGABRT"| E["Assert failure or abort"]
    C -->|Yes| F["Request exceeded harakiri timeout"]
    C -->|No| G{"Worker alive less than 1s?"}
    G -->|Yes| H["Startup crash or import error"]
    G -->|No| I["OOM killer terminated worker"]
    D --> J{"Alpine with threads > 1?"}
    J -->|Yes| K["musl libc incompatibility"]
    J -->|No| L["Check C extension versions"]

Common causes

CauseWhat it looks likeFirst thing to check
OOM kill (signal 9, no HARAKIRI)Worker killed mid-request, RSS near system or container limit, no startup error in app logsdmesg or kernel logs for OOM-killer entries
Harakiri kill (signal 9, with HARAKIRI log)Worker killed at exactly the harakiri timeout, harakiri_count risingDownstream dependency health, harakiri-verbose traceback
Startup crash (any signal, worker alive less than 1s)Worker dies immediately after spawn, never serves a request, tight respawn loopApplication logs for ImportError, SyntaxError, missing env vars
C extension segfault (signal 11)Worker crashes mid-request or during init, no Python tracebackCore dump, C extension versions, application logs
Alpine musl + threads (signal 11)Constant SIGSEGV on Alpine with threads > 1uWSGI version, base image

Quick checks

# Read the kill signal from uWSGI logs
grep -E "DAMN ! worker .* died" /var/log/uwsgi/app.log | tail -20

# Check for HARAKIRI lines alongside the deaths
grep -E "HARAKIRI ON WORKER" /var/log/uwsgi/app.log | tail -20

# Check respawn_count and harakiri_count from the stats server
uwsgi --connect-and-read 127.0.0.1:9191 | jq '[.workers[] | {id: .id, respawn_count: .respawn_count, harakiri_count: .harakiri_count, status: .status, pid: .pid}]'

# Count accepting workers (pid > 0, not cheap, accepting connections)
uwsgi --connect-and-read 127.0.0.1:9191 | jq '[.workers[] | select(.pid > 0 and .status != "cheap" and .accepting == 1)] | length'

# Check for OOM killer activity in kernel logs
dmesg | grep -i "oom\|killed process" | tail -20

# Check per-worker RSS from stats server
uwsgi --connect-and-read 127.0.0.1:9191 | jq '.workers[] | select(.pid > 0) | {id: .id, rss_mb: (.rss / 1048576)}'

# Validate uWSGI options/config (loads plugins and, in non-lazy mode, the app; then exits)
uwsgi --ini /etc/uwsgi/app.ini --no-server

# To catch application import errors regardless of mode
python -c "from your_module import application"

# Check worker age from last_spawn timestamps
uwsgi --connect-and-read 127.0.0.1:9191 | jq --argjson now "$(date +%s)" '[.workers[] | select(.pid > 0) | {id: .id, pid: .pid, last_spawn_age: ($now - .last_spawn)}]'

How to diagnose it

  1. Read the signal number from the log. The DAMN ! worker N (pid: XXX) died, killed by signal S line includes the signal. Signal 9 is SIGKILL (OOM or harakiri). Signal 11 is SIGSEGV (segfault). Signal 6 is SIGABRT (assert). This single number narrows the cause significantly.

  2. Distinguish harakiri from OOM. If the signal is 9, check for HARAKIRI ON WORKER log lines at the same timestamp. If present, the worker was killed for exceeding the request timeout. If absent, suspect the OOM killer.

  3. Check worker lifetime. If workers are dying within seconds of spawn, you have a startup crash loop. Look at last_spawn timestamps in the stats server. Workers that are repeatedly very young are crashing during initialization, not during request processing.

  4. Check the OOM killer. Run dmesg | grep -i oom to see if the kernel or cgroup OOM killer terminated the worker. In containers, check the container runtime logs for OOMKilled events. If RSS was climbing before the kills, this is memory exhaustion, not a code bug.

  5. Check respawn_count vs harakiri_count. Both are per-worker monotonic counters in the stats server. If respawn_count is increasing but harakiri_count is not, the respawns are not harakiri-driven. Factor in max-requests recycling (step 6) to isolate crash-driven respawns: crash respawns = total respawns minus harakiri kills minus max-requests recycles.

  6. Distinguish from max-requests recycling. If max-requests is configured, some respawns are expected. Calculate the expected rate: total_requests_per_second / max_requests_per_worker * num_workers. If the actual respawn rate matches this, the respawns are normal recycling. If it exceeds this, workers are crashing.

  7. Check application logs for startup errors. With lazy-apps enabled, each worker imports the application independently. An ImportError, SyntaxError, missing environment variable, or failed database migration will crash every worker on startup. Run uwsgi --ini /etc/uwsgi/app.ini --no-server to validate the configuration (uWSGI has no --check-config option). To catch import errors, load the callable directly in Python: python -c "from your_module import application".

  8. Check for C extension segfaults. Signal 11 with no Python traceback usually means a C extension crashed. Check for core dumps. Alpine/musl builds of uWSGI with threads enabled have a documented history of segfaults (unbit/uwsgi issues #1514 and #2508), with no single upstream fix version covering them; test your exact uWSGI, Python, and musl versions, or switch to a glibc-based image.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
respawn_count (per-worker, delta)Direct measure of worker lifecycle churnRate exceeding expected max-requests cadence
harakiri_count (per-worker, delta)Distinguishes harakiri kills from crashesNon-zero delta means requests are timing out
Accepting worker countIf zero, service cannot serve requestsDrops to zero with master alive means total outage
Worker RSSMemory growth leading to OOM killsRSS approaching system or container memory limit
Worker lifetime (last_spawn age)Short lifetime signals startup crash loopWorkers consistently alive less than a few seconds
Exception count (delta)Application errors that may precede crashesSpike in exceptions before worker deaths
System OOM events (dmesg)Kernel-level evidence of memory pressureAny OOM-kill entries referencing uWSGI worker PIDs

Fixes

OOM kill loop

If the OOM killer is targeting workers (signal 9, no HARAKIRI log, OOM entries in dmesg), the worker pool is consuming more memory than available. The respawned worker immediately grows back to the same RSS and gets killed again, creating an infinite loop.

Reduce per-worker memory by lowering reload-on-rss so workers recycle before approaching the system limit. Reduce the worker count if total RSS (workers times per-worker RSS) exceeds available memory. If running in a container, increase the memory limit. Consider max-requests as a complementary recycling mechanism.

Tradeoff: Fewer workers means less concurrency. Measure the actual per-worker RSS under load before reducing the count.

Startup crash loop (import error, missing dependency)

If workers die within seconds of spawn, the application fails to initialize. Check application logs for the specific error: ImportError, SyntaxError, missing environment variable, database migration not applied.

Run uwsgi --ini /etc/uwsgi/app.ini --no-server to validate the configuration and, in non-lazy mode, import the application. To catch import errors directly, run python -c "from your_module import application".

Roll back the deployment if the error appeared after a code change.

C extension segfault (signal 11)

If workers crash with signal 11 and no Python traceback, a C extension is faulting. Common culprits include database drivers, XML parsers, and ML libraries.

Check the C extension versions against known compatibility issues. On Alpine Linux with threads greater than 1, verify the exact uWSGI and Python versions against the reported segfault issues (unbit/uwsgi #1514 and #2508) or switch to a glibc-based image.

Enable core dumps if not already configured. A core file from the crashed worker gives you a backtrace into the C extension.

Harakiri-driven respawns

If respawns correlate 1:1 with harakiri_count increases, workers are being killed for exceeding the request timeout, not crashing. The root cause is downstream: a slow database query, an unresponsive external API, or a network partition to a dependency.

Check harakiri-verbose output for the blocked syscall. See uWSGI harakiri-verbose: finding the blocked syscall behind a timeout. Do not disable harakiri to silence the respawns. Fix the downstream dependency or tune the timeout.

Prevention

  • Validate before deploy. Run uwsgi --ini app.ini --no-server in your CI pipeline to catch config and (in non-lazy mode) application import errors before they reach production.
  • Set reload-on-rss below the OOM threshold. Workers should recycle gracefully before the kernel kills them. This converts a violent OOM kill into a graceful recycle.
  • Monitor respawn_count delta against expected max-requests cadence. If you know your traffic rate and max-requests value, you can compute the expected recycling rate. Alert when actual respawns exceed it.
  • Enable harakiri-verbose. When a harakiri fires, this logs the blocked syscall and wchan, which is the difference between a fast diagnosis and a slow one.
  • Pin and test uWSGI on Alpine. If using Alpine Linux, test the exact uWSGI and Python versions with threads enabled before promotion; there is no single upstream fix version for the reported Alpine segfaults (unbit/uwsgi #1514 and #2508).

How Netdata helps

  • Per-second respawn tracking lets you see crash loops as they form, not minutes later. Respawn rate spikes are immediately visible alongside CPU usage spikes from repeated fork cycles.
  • Correlating respawn_count with harakiri_count separates crash-driven respawns from harakiri-driven ones without manual log parsing. If both rise together, it is harakiri. If only respawn rises, it is a crash.
  • Worker RSS trending shows memory growth leading to OOM kills. You can see the sawtooth pattern of reload-on-rss recycling versus the linear climb of a worker approaching the OOM threshold.
  • Accepting worker count drops to zero during a startup crash loop, even while the master process appears healthy. Alerting on this catches the outage that a master-PID health check misses.
  • Anomaly detection on respawn rate flags deviations from the normal max-requests cadence without requiring a static threshold.