The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / uwsgi / uwsgi-how-it-works-in-production ▌

Operations Guides

How uWSGI actually works in production: a mental model for operators

uWSGI is a pre-fork application server. A single master process spawns a pool of worker processes, each running a full copy of your application. The master never serves requests. It manages worker lifecycle, enforces timeouts, and coordinates graceful reloads. Every request flows through the same path: a client connects to a socket, the kernel queues that connection in a listen backlog, and a worker calls accept() to pull it out and process it synchronously.

The simplicity hides several cliff-edge failure modes. When all workers are busy, there is no graceful degradation. Requests pile up in the kernel backlog until it fills, and then the kernel silently drops new connections. When a worker hangs, there is no recovery unless you configured harakiri. When a reload fails, every worker can disappear at once.

This article covers the abstractions you need before the runbooks: the master-worker relationship, the kernel backlog as the primary buffer, the worker state machine, the harakiri watchdog, the cheaper subsystem, and the difference between threading and async modes.

What it is and why it matters

The core design has three components:

  1. Master process: Loads the application, opens the listening socket, then forks workers. Does not handle HTTP requests. Its job is worker management: spawning, monitoring, enforcing per-request timeouts (harakiri), and coordinating reloads.

  2. Worker processes: Each worker is an OS-level process with its own memory space, containing a full copy of the application. Workers compete for incoming connections by calling accept() on the shared listening socket. When a worker accepts a connection, it processes that request to completion before accepting the next one.

  3. Kernel listen backlog: The buffer between incoming connections and available workers. When a client connects, the kernel completes the TCP handshake and places the connection in this queue. When a worker calls accept(), it pulls the next connection. If all workers are busy and the queue fills, the kernel drops new connections.

Every production failure in uWSGI traces back to one of these components. Worker exhaustion is worker pool saturation. Harakiri storms are the watchdog firing repeatedly. Memory creep is per-worker process copies growing over time. Reload failures are the fork-replace lifecycle going wrong.

How it works

The master-worker fork model

By default, uWSGI loads the application once in the master process, then forks workers. This uses copy-on-write (COW) semantics: workers share the master’s memory pages until they write to them, at which point the kernel duplicates the modified page.

The lazy-apps option changes this: each worker loads the application independently after fork. This costs more memory and increases startup time, but avoids problems with libraries that are not fork-safe: database connection pools, thread locks, file handles initialized at import time. The older lazy option is deprecated and discouraged by the uWSGI project; it changes many internal defaults and exists only for backward compatibility.

Request flow through the kernel backlog

flowchart TD
    LB[nginx / LB] -->|incoming connections| SK[Kernel listen backlog]
    SK -->|accept| W1[Worker 1: app copy]
    SK -->|accept| W2[Worker 2: app copy]
    SK -->|accept| Wn[Worker N: app copy]
    Master[Master process] -->|fork + respawn| W1
    Master -->|fork + respawn| W2
    Master -->|fork + respawn| Wn
    Master -.->|harakiri: SIGKILL| W2
    W1 -->|response| LB

The listen backlog is the primary saturation point. It has a hard limit, set by --listen (default 100) and capped by the kernel’s net.core.somaxconn. When all workers are busy, incoming connections accumulate here. Once the backlog fills, the kernel drops new connections with no application-level log entry, no error counter, and no uWSGI-level signal.

One instrumentation gap: uWSGI’s listen_queue stats field is unreliable on standard Linux. The TCP measurement via TCP_INFO is inconsistent across kernel versions, and the UNIX socket measurement requires a non-standard kernel ioctl. The load field in the stats JSON is identical to listen_queue, not average latency as its name implies: the uWSGI master sets both from the same measured backlog value, with a source comment noting that load “should be something more advanced based on different values”. Both fields almost always read 0. Measure queue depth externally:

ss -ltn  # TCP: check Recv-Q against the socket
ss -lxn  # UNIX sockets: check Recv-Q

Worker state machine

Each worker cycles through a simple state machine:

StateMeaning
idleWorker is alive and waiting to accept a new connection
busyWorker is processing a request
cheapWorker has been scaled down by the cheaper subsystem (pid=0)
pauseWorker is intentionally suspended, for example during reload
sig<N>Worker is handling a signal

A worker stuck in busy indefinitely means a request has hung. This is the most common silent failure: the worker appears alive in the process table but is doing nothing useful. Without harakiri, that worker stays stuck forever, permanently reducing capacity.

Harakiri: the per-request watchdog

Harakiri is uWSGI’s safety net for stuck workers. When configured via --harakiri, the master starts a timer each time a worker accepts a request. If the worker exceeds the timeout, the master sends SIGKILL and respawns it.

Key behaviors:

  • Harakiri is disabled by default. Without it, a stuck worker stays stuck forever. The absence of harakiri configuration is itself a risk.
  • Each harakiri kill increments the per-worker harakiri_count (never reset on respawn) and respawn_count.
  • --harakiri-verbose logs the blocked syscall and wchan when harakiri fires (Linux only, reads /proc/<pid>/syscall and /proc/<pid>/wchan).
  • If --harakiri-graceful-timeout is set, the master sends SIGTERM first, giving the worker a chance to clean up before SIGKILL.

Harakiri is not a failure indicator. It is a safety mechanism. The problem is whatever is causing requests to hang, not the kill itself.

Reloads: re-forking the pool

A graceful reload (triggered by SIGHUP or by touching a reload-trigger file in Emperor mode) tells all workers to finish their current requests and exit. The master then forks a fresh pool from the potentially updated application code.

This creates a capacity window: old workers drain while new workers start. If the application has slow startup (heavy imports, ML model loading, cache warming), this window can stretch to seconds or minutes. --chain-reload mitigates this by cycling workers one at a time instead of all at once.

A reload with broken application code (import error, syntax error, missing dependency) kills the old workers, and the new workers fail to start. The master repeatedly tries to spawn workers that immediately die. The result is zero accepting workers.

Version note: on uWSGI 2.0.x (the current stable line), SIGTERM means “brutally reload the stack,” not “shut down.” Set die-on-term = true if you want SIGTERM to behave as convention expects. This is documented to change in the unreleased 2.1 branch.

The cheaper subsystem: dynamic worker scaling

The cheaper subsystem adjusts the worker count based on demand. Algorithms:

  • spare (default): Maintains a minimum number of idle workers. Spawns new workers when idle count drops below the threshold.
  • spare2: A variant of spare with separate scaling thresholds.
  • backlog (Linux TCP only): Scales based on the kernel listen queue depth. Does not work with UNIX domain sockets.
  • busyness (requires cheaper_busyness plugin): Scales based on actual worker utilization over time.

When workers are scaled down by cheaper, they appear in the stats JSON with status: "cheap" and pid: 0. Monitoring must account for this: alerting on a fixed expected worker count generates false positives when cheaper legitimately scales down.

Threading vs async

Threaded mode: Each worker runs N threads that share the worker’s memory space. In Python, threads share the GIL, so CPU-bound work is serialized within a worker despite multiple threads. Threads help for I/O-bound workloads where the thread is blocked on a downstream response, but they do not provide CPU parallelism. Multiple worker processes are still needed for that. enable-threads is ON by default as of uWSGI 2.0.27 (the 2.0.27 release notes say “set enable threads by default”). On older versions, set enable-threads = true explicitly for application-generated threads to function.

Async mode (gevent, asyncio): Each worker multiplexes many concurrent I/O-bound requests via an event loop. A single worker can handle hundreds or thousands of concurrent connections, but only if the application code yields to the event loop consistently. The uWSGI documentation warns: “If you are in doubt, do not use async mode.”

In async mode, the meaning of “worker busy” changes. A worker is “busy” when the event loop is running, which is almost always. Per-core in_request counts become the true concurrency indicator, not the binary busy/idle status. Worker busy ratio, the primary utilization signal in pre-fork mode, is nearly useless in async mode.

Where it shows up in production

The mental model maps to characteristic failure patterns:

Worker exhaustion. All workers are busy, requests queue in the kernel backlog, then get dropped when the backlog fills. This manifests as rising latency followed by connection failures. The cliff is sharp: there is no gradual degradation between “keeping up” and “dropping connections.”

Harakiri death spiral. A downstream dependency (database, external API) becomes unresponsive. Every request blocks on the dependency, exceeds the harakiri timeout, and the worker is killed and respawned. The respawned worker immediately accepts a new request that also blocks. Throughput collapses to near zero while respawn rate spikes. The system burns CPU on fork cycles while serving zero useful traffic.

Memory creep. Each worker holds a full copy of the application. RSS grows over time due to memory leaks, Python memory fragmentation, or cache bloat. Workers grow until OOM-killed or recycled by max-requests or reload-on-rss. This produces a sawtooth RSS pattern.

Stuck workers without harakiri. If harakiri is not configured, a worker that hangs on blocking I/O stays stuck forever. Capacity silently degrades as workers accumulate in stuck state. By the time you notice, the service may already be unresponsive, and harakiri_count reads zero, giving a false sense of safety.

Reload blackout. A graceful reload with broken application code leaves zero running workers. Old workers are killed, new workers fail to start, and the master churns through respawn attempts while serving nothing.

Tradeoffs and when to use them

Pre-fork vs threads vs async: Pre-fork (separate processes) is the safest default. It provides true isolation and CPU parallelism. Threads add concurrency within a process but are GIL-constrained in Python and introduce shared-state bugs. Async modes offer the highest concurrency for I/O-bound workloads but require careful application design and change the meaning of monitoring signals.

lazy-apps vs default fork: Default fork saves memory via COW but can break with non-fork-safe libraries. lazy-apps is safer but costs more memory and startup time. Each respawned worker re-imports the entire application.

Cheaper vs fixed workers: Cheaper reduces resource usage during low-traffic periods but introduces monitoring variability. Fixed workers provide predictable capacity but waste resources during quiet periods. If you use cheaper, alert on utilization ratios and accepting worker count, not absolute worker counts.

reload-on-rss vs evil-reload-on-rss: Both recycle workers when RSS exceeds a threshold. reload-on-rss is graceful: the worker finishes its current request, then exits. evil-reload-on-rss sends SIGKILL mid-request, which means the client sees a broken response. Monitor write errors alongside respawn rates to detect when mid-request kills are happening.

Signals to watch in production

SignalWhy it mattersWarning sign
Accepting worker countPrimary availability metric. Zero accepting workers with a running master is critical.Count drops to zero or fluctuates rapidly
Worker busy ratioCurrent concurrency utilization. At 100%, additional requests queue in the kernel backlog with no uWSGI-level visibility.Sustained at or near 100%
Harakiri rateRequests are hanging or running too long. Each kill means a dropped request and a respawn cycle.Any sustained non-zero rate
Average response time (avg_rt)Application responsiveness trend (avg_rt is recomputed as (old + new) / 2 for every request). Approaches harakiri timeout before workers start dying.Sustained increase or approaching configured harakiri
Worker RSSMemory health per worker. Consistent growth across all workers indicates a leak.Sustained positive slope over hours
Respawn rateWorker lifecycle churn. Normal with max-requests recycling. Abnormal when driven by crashes or harakiri.Rate significantly exceeds expected max-requests cadence
Socket backlog depth (external)Primary saturation signal. uWSGI’s internal listen_queue is unreliable on Linux. Use ss to measure externally.Sustained non-zero Recv-Q
Exception rateUnhandled application errors reaching the WSGI layer.Baseline-relative increase

How Netdata helps

  • Per-second worker state visibility: Netdata’s uWSGI collector pulls the stats server JSON at per-second resolution, capturing worker busy ratios, accepting worker counts, and state transitions that coarser polling intervals miss.

  • Correlating harakiri with downstream latency: When harakiri rate spikes, overlay uWSGI harakiri counts against database query latency, external API response times, and system-level CPU and memory metrics in the same dashboard. The root cause of a harakiri storm is almost always downstream.

  • Detecting memory creep before OOM: Per-worker RSS collected over time reveals the sawtooth pattern of leaks and recycling. Anomaly detection can flag RSS growth trends before they hit the reload-on-rss threshold or trigger an OOM kill.

  • Backlog pressure without internal metrics: Since uWSGI’s listen_queue is unreliable on standard Linux, correlate host-level TCP and socket metrics with worker busy ratios for a more accurate picture of backlog pressure.

  • Distinguishing recycling from crashes: Overlay respawn rate with harakiri count to immediately see whether respawns are driven by healthy max-requests recycling or harakiri-induced crashes.