The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / uwsgi / uwsgi-nginx-502-bad-gateway ▌

Operations Guides

uWSGI behind nginx: 502 Bad Gateway and 'upstream prematurely closed connection'

The error string upstream prematurely closed connection while reading response header from upstream in the nginx error log means nginx had an established connection to a uWSGI worker, the worker accepted the request, and then the connection closed before nginx received a complete response header. The worker vanished mid-response.

This is distinct from two other errors that also produce 5xx responses but have different root causes:

nginx log stringHTTP statusWhat happenedWhere to look
upstream prematurely closed connection while reading response header from upstream502Worker accepted the request, then died mid-responseuWSGI stats: harakiri_count, respawn_count; dmesg
connect() ... failed (111: Connection refused) while connecting to upstream502No connection established: uWSGI not listening, socket permissions, or backlog fullMaster alive? Socket perms? ss -ltn
upstream timed out (110: Connection timed out) while reading response header from upstream504Worker accepted but did not finish within nginx’s timeoutuWSGI stats: worker busy ratio, avg_rt

This article covers the first row only. The root cause is always on the uWSGI side. The worker was killed by harakiri, crashed with a segfault, was killed by the kernel OOM killer, or was SIGKILL’d by the master enforcing an evil-reload-on-rss threshold. Diagnose from uWSGI, not from nginx.

The exit is always involuntary. Normal worker recycling via max-requests or reload-on-rss waits for the current request to finish before exiting. A worker that vanishes mid-response was killed by SIGKILL (from harakiri, from evil-reload-on-rss, or from the kernel OOM killer) or crashed with a segfault.

flowchart TD
    A["nginx returns 502"] --> B{"Error string in log?"}
    B -- "prematurely closed" --> C["Worker died mid-response"]
    B -- "Connection refused 111" --> D["Worker never accepted"]
    B -- "timed out 110" --> E["Worker too slow: 504"]
    C --> F{"harakiri_count rising?"}
    F -- "Yes" --> G["Request exceeded timeout"]
    F -- "No" --> H{"Non-harakiri respawns?"}
    H -- "Yes" --> I{"dmesg evidence?"}
    I -- "oom-kill" --> J["OOM killer"]
    I -- "segfault" --> K["C extension crash"]
    I -- "nothing" --> L["evil-reload-on-rss"]
    H -- "No" --> M["buffer-size overflow"]

Common causes

CauseWhat it looks likeFirst thing to check
Harakiri killRequest exceeded configured timeout. Master sent SIGKILL. harakiri_count rising.harakiri_count delta in stats server
evil-reload-on-rssWorker RSS exceeded threshold. Master SIGKILL’d mid-request. respawn_count rising, no harakiri.Config for evil-reload-on-rss; write_errors
Worker segfaultC extension crashed. Worker killed by signal 11 (SIGSEGV). Respawn without harakiri.uWSGI log for signal 11; dmesg
OOM killerKernel killed worker for memory pressure. Respawn without harakiri or signal.dmesg for oom-killer
buffer-size overflowRequest headers exceed the 4096 byte default buffer. Connection closed without response.uWSGI log for invalid request block size

Quick checks

These commands are read-only and safe to run during an incident. Adjust the stats socket address (127.0.0.1:9191 in these examples) to match your deployment. If the stats server uses a UNIX socket, substitute uwsgi --connect-and-read /path/to/stats.sock. If --stats-http is enabled, curl http://127.0.0.1:9191 also works.

# Check harakiri count (per-worker, monotonic, never reset even on respawn)
uwsgi --connect-and-read 127.0.0.1:9191 | jq '[.workers[].harakiri_count] | add'

# Check total respawn count (includes harakiri, max-requests, and crash respawns)
uwsgi --connect-and-read 127.0.0.1:9191 | jq '[.workers[].respawn_count] | add'

# Check write errors (broken pipes from mid-request kills)
uwsgi --connect-and-read 127.0.0.1:9191 | jq '[.workers[].cores[].write_errors] | add'

# Check per-worker RSS in MB (identify OOM or evil-reload-on-rss candidates)
uwsgi --connect-and-read 127.0.0.1:9191 | jq '.workers[] | select(.pid > 0) | {id: .id, rss_mb: (.rss / 1048576)}'

# Check currently busy workers and what they are serving
uwsgi --connect-and-read 127.0.0.1:9191 | jq '[.workers[] | select(.status == "busy") | .cores[] | select(.in_request == 1) | .vars[] | select(startswith("REQUEST_URI=")) | .[12:]]'

# Check for OOM kills in kernel log
dmesg -T | grep -i "out of memory\|oom-kill"

# Check for segfaults in kernel log
dmesg -T | grep -i "segfault"

# Confirm the exact nginx error string
grep "prematurely closed" /var/log/nginx/error.log | tail -20

How to diagnose it

  1. Confirm the error type. Grep the nginx error log for the exact string. If it says Connection refused, the worker never accepted the connection: check socket permissions, whether uWSGI is listening, and whether the backlog is full. If it says timed out, the worker was too slow and nginx gave up first: that is a 504 latency problem, not a mid-response death.

  2. Check harakiri_count delta. If it is rising, workers are being killed by the harakiri timer. The request exceeded the configured timeout. The root cause is a slow or hung request, typically a downstream dependency (database, external API, DNS resolution). Check the REQUEST_URI of busy workers in the stats server (inside each busy core’s vars array; there is no top-level uri field) to identify which endpoint is timing out. If harakiri-verbose is enabled, the uWSGI log will include the blocked syscall and wchan from /proc/<pid>/syscall and /proc/<pid>/wchan.

  3. Isolate non-harakiri respawns. Subtract the harakiri_count delta from the respawn_count delta. Both are per-worker, monotonic counters that persist through respawns. The remainder includes max-requests recycling, segfaults, and OOM kills. If the non-harakiri respawn rate exceeds what max-requests recycling alone would produce, workers are crashing or being OOM-killed. Check dmesg for evidence.

  4. Check for evil-reload-on-rss. If your config includes evil-reload-on-rss, the master sends SIGKILL to workers mid-request when RSS exceeds the threshold. This produces 502s on whatever request was in flight at that moment, with no harakiri increment. Check whether write_errors are elevated (clients disconnected before the response was sent). The fix is to switch to reload-on-rss, which is graceful: the worker finishes the current request, then exits and respawns.

  5. Check buffer-size. If the 502s correlate with specific requests that have large headers (long URIs, large cookies), check the uWSGI log for invalid request block size. The default buffer-size is 4096 bytes. Requests with headers exceeding this are discarded, and nginx sees a closed connection. Note: this does not kill the worker. The worker stays alive, but the connection is closed without a response.

  6. Align timeouts. Verify the relationship between nginx uwsgi_read_timeout and uWSGI’s harakiri. If uwsgi_read_timeout is shorter than harakiri, nginx returns a 504 to the client while the worker keeps processing (wasted capacity, and the worker eventually dies to harakiri anyway). If uwsgi_read_timeout is equal to or longer than harakiri, harakiri fires first, the worker dies, and nginx sees the disconnect as a 502. Set uwsgi_read_timeout >= harakiri so harakiri is the effective limiter.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
harakiri_count (delta)Each increment is one worker killed mid-request, one client got 502Any sustained non-zero rate
respawn_count (delta)Workers dying for any reason: harakiri, crash, OOM, max-requests recyclingRate exceeding expected max-requests cadence
respawn_count minus harakiri_count (delta)Non-harakiri respawns: includes max-requests recycling, crashes, and OOMDelta exceeding expected max-requests cadence
write_errors (delta)Client connection closed before response was sent (broken pipe from mid-request kill)Sustained increase correlates with evil-reload-on-rss or harakiri
Worker RSS (bytes)Memory pressure leading to OOM kill or evil-reload-on-rss killRSS approaching system limits or configured threshold
avg_rt per worker (microseconds)Latency approaching harakiri threshold means kills are imminentavg_rt approaching the configured harakiri value
Worker busy ratioAll workers busy means saturation, which leads to slow requests and harakiriSustained 100% busy across all active workers

Fixes

Harakiri kills: fix the downstream dependency

If harakiri_count is rising, the problem is not the timeout value. The problem is that requests are taking too long. Raising harakiri only delays the 502; it does not fix the root cause.

Identify the slow endpoint from the REQUEST_URI of busy workers in the stats server (per-core vars array). Check downstream dependencies: database query latency, external API response times, DNS resolution times. Add application-level timeouts on all downstream calls so a hung dependency returns a fast error instead of blocking until harakiri fires. Enable harakiri-verbose for diagnostic backtraces showing where the worker was blocked.

evil-reload-on-rss: switch to reload-on-rss

evil-reload-on-rss sends SIGKILL to workers mid-request when RSS exceeds the threshold. Every request in flight at that moment gets a 502. Replace it with reload-on-rss, which triggers a graceful exit: the worker finishes the current request, then recycles.

The tradeoff: reload-on-rss allows RSS to exceed the threshold during the grace period. If memory pressure is severe enough that the extra headroom matters, the real fix is addressing the memory leak or increasing system memory.

Segfaults: identify the C extension

Signal 11 (SIGSEGV) in the uWSGI log means a C extension crashed. Common culprits include database drivers, XML and HTML parsers, and image processing libraries. If segfaults correlate with specific request types, identify the code path and the library involved.

Persistent, unexplained segfaults can also indicate memory corruption from a previous OOM event. Treat them as suspicious until you identify the faulting library.

OOM kills: add memory headroom or reduce workers

The Linux OOM killer targets the highest-RSS process. With multiple uWSGI workers, one gets killed first, is respawned by the master, grows back, and the cycle repeats. Check dmesg for confirmation.

Options: reduce worker count to lower aggregate memory demand, configure reload-on-rss (graceful) at a threshold below the OOM danger zone so workers recycle before the kernel intervenes, or add system memory. Verify that max-requests is set so workers recycle proactively before RSS grows dangerously.

buffer-size overflow: increase the buffer

If the uWSGI log shows invalid request block size: X (max 4096), increase buffer-size to accommodate the largest expected request headers. Values of 16384 or 32768 are common for applications with large cookies or long URIs.

Prevention

  • Configure harakiri and harakiri-verbose. Without harakiri, stuck workers have no timeout and permanently consume a worker slot. Set it to 2-3x your expected maximum legitimate request duration. The default is disabled. Enable harakiri-verbose for the blocked syscall and wchan when harakiri fires.
  • Prefer reload-on-rss over evil-reload-on-rss. The graceful variant finishes the current request before recycling. The evil variant kills mid-request and produces 502s.
  • Align uwsgi_read_timeout with harakiri. Set uwsgi_read_timeout >= harakiri so harakiri is the effective limiter and nginx sees the 502 (diagnostic) rather than timing out first and returning a 504 (less diagnostic, and the worker keeps processing uselessly).
  • Monitor harakiri_count and respawn_count as rate-of-change deltas, not absolute values. Both are per-worker, monotonic counters that never reset, even on respawn.
  • Set buffer-size appropriately for your workload if requests carry large headers.
  • Ensure all downstream calls have their own timeouts. A missing downstream timeout on a database query, HTTP client, or DNS resolver is the most common root cause of harakiri storms and the 502s they produce.

How Netdata helps

  • Per-second collection of harakiri_count, respawn_count, worker status, avg_rt, and RSS from the uWSGI stats server. You see the kill the moment it happens, not when nginx 502 alerts fire minutes later.
  • Correlate harakiri_count spikes with downstream metrics (database query latency, external API response time) to identify the slow dependency causing the timeout.
  • Anomaly detection on respawn_count flags crash-induced respawns separately from routine max-requests recycling.
  • RSS growth-rate trends surface memory pressure before OOM kills or evil-reload-on-rss triggers.
  • Write error rate monitoring catches mid-request kills (broken pipes from SIGKILL) that otherwise appear only as intermittent 502s in nginx logs.