The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / uwsgi / uwsgi-monitoring-maturity-model ▌

Operations Guides

uWSGI monitoring maturity model: from survival to expert

A four-level progression for uWSGI monitoring, from bare liveness checks to deep signal correlation. The levels are cumulative: you cannot skip to Level 3 by tracking RSS growth trends while ignoring harakiri rate. Each tier closes a specific class of blind spot that the previous tier could not see.

The model assumes the uWSGI stats server is enabled with --stats <address>. Without it, every level above survival is unreachable. HTTP access to the stats server requires the additional --stats-http flag; otherwise use uwsgi --connect-and-read <addr> for TCP sockets or socat - UNIX-CONNECT:<path> for UNIX sockets.

The four levels

flowchart TD
    L4["Level 4: Expert
per-endpoint harakiri, worker age skew,
connection-pool correlation, PSS divergence"] L3["Level 3: Mature
per-worker distribution, RSS trends,
serving URI, spooler/emperor, fd count"] L2["Level 2: Operational
busy ratio, harakiri/respawn/exception deltas,
per-worker RSS, avg_rt"] L1["Level 1: Survival
master alive, worker accepting,
non-zero throughput"] L1 --> L2 --> L3 --> L4

Level 1: survival

Goal: know whether uWSGI is alive and processing requests.

Three signals form the survival floor:

  • Master process alive. Check via PID file or process monitoring. If the master is dead, the entire application is down: no workers exist, no requests are processed.
  • At least one worker accepting. From the stats server, count workers where pid > 0 and accepting == 1 and status != "cheap". Zero accepting workers with a running master means the service cannot accept new requests. Connections queue in the kernel backlog until they time out.
  • Non-zero throughput during expected traffic hours. Sum workers[].requests across all workers and compute the delta between polling intervals. A sudden drop with no corresponding decrease in incoming traffic indicates workers are stuck or the application is rejecting requests.

What Level 1 catches: total outages, master process death, complete worker absence.

What Level 1 misses: worker starvation, memory leaks, harakiri storms, listen queue overflow, degraded latency. The service can pass all three checks while dropping 100% of real traffic because all workers are stuck on blocking calls. A health check that hits the stats endpoint will return a response even during complete worker starvation because the stats server is served by the master process, not by workers.

Level 2: operational

Goal: detect degradation before it becomes an outage.

Level 2 adds utilization, error rates, and resource pressure signals. This is the minimum a production deployment should have.

  • Worker busy ratio. Count workers where status == "busy" divided by alive workers (pid > 0 and status != "cheap"). Sustained 100% means every additional request queues in the kernel backlog with no uWSGI-level visibility. Brief spikes to 100% during traffic bursts are normal if they self-resolve within seconds.
  • Harakiri rate (delta). Sum workers[].harakiri_count across all workers and compute the delta between polling intervals. This counter is per-worker and monotonic, never reset even on respawn. Track the rate, not the absolute value. If harakiri is not configured, this counter is always 0, which is not a sign of health but a monitoring blind spot.
  • Respawn rate (delta). Sum workers[].respawn_count and compute the delta. Normal causes include max-requests recycling and reload-on-rss. To isolate crash-induced respawns from harakiri respawns, subtract the harakiri rate from the respawn rate.
  • Exception rate (delta). Sum workers[].exceptions and compute the delta. These are unhandled exceptions that reach the WSGI layer. Application-level error handlers that catch and return 500s do not increment this counter.
  • Average response time (avg_rt) per worker. The workers[].avg_rt field is an exponential moving average with factor 0.5, computed internally as (old_avg_rt + current_request_time) / 2. Each new request contributes 50% of the new value. After approximately 7 requests, older contributions fall below 1%. This makes avg_rt responsive to recent changes but volatile. A single slow request can move it significantly. Do not use it as a latency SLI; use access logs or application-level metrics for that.
  • Per-worker RSS. The workers[].rss field reports resident set size in bytes. The --memory-report option must be enabled for RSS and VSZ values to appear in the stats output. Consistent growth across all workers indicates a memory leak. Divergence, where one worker is much larger than others, indicates a request-specific memory issue.
  • Listen queue awareness. The listen_queue field in uWSGI stats is unreliable on standard Linux: TCP measurement via TCP_INFO varies across kernel versions, and UNIX socket measurement requires a non-standard kernel ioctl. In current 2.0.x, load is identical to listen_queue (the master writes the same backlog value to both), not average latency despite its name. The listen_queue_errors field exists in the JSON output but is never incremented in the 2.0.x source, so it reads 0. At Level 2, the minimum requirement is awareness that these fields are unreliable and must be measured externally.

What Level 2 catches: worker pool exhaustion, coarse memory leak detection, harakiri storms, application error spikes, capacity saturation.

What Level 2 misses: which endpoint is causing the problem, whether respawns are healthy recycling or crashes, per-worker asymmetry, file descriptor exhaustion, spooler or emperor subsystem failures.

Level 3: mature

Goal: diagnose root causes, not just detect symptoms. Level 3 adds per-worker granularity, subsystem health, and external signals that compensate for uWSGI’s internal limitations.

  • Per-worker request distribution. Compare workers[].requests across workers. Uneven distribution suggests load imbalance, lock contention, or thundering herd effects. Enable --thunder-lock for uniform accept() distribution across workers.
  • RSS growth trend. Track the slope of RSS over hours, not the point-in-time value. Linear growth across all workers is a leak. A sawtooth pattern with periodic respawns is healthy recycling via max-requests or reload-on-rs.
  • Currently serving URI. The per-core request variables (cores[].vars, populated while in_request == 1) identify which endpoint a busy worker is serving. During starvation, if all busy workers show the same URI, that endpoint is the root cause.
  • Respawn classification. Expected respawn rate equals (total_requests_per_second / max_requests_per_worker) * num_workers. Respawn rate significantly above this, especially when tracking harakiri count 1:1, indicates crashes rather than recycling.
  • Write and read error rates per core. The workers[].cores[].write_errors and workers[].cores[].read_errors fields track socket errors during request and response handling. Sustained high write error rates indicate slow responses causing clients to disconnect. Suppressed from stats output if --stats-no-cores is enabled.
  • Spooler health (if configured). The spoolers[] array exposes tasks (pending count), running (0/1), and respawns. Growing tasks count means the spooler cannot keep up. Any non-zero respawn rate indicates spooler instability.
  • Emperor and Vassal health (if Emperor mode). Each vassal needs independent monitoring. Emperor health does not imply vassal health. A vassal can die and fail to restart due to a broken config while the Emperor continues running. The count of active vassals versus expected vassals is the primary health indicator.
  • Kernel overflow counters. Use nstat -az TcpExtListenOverflows TcpExtListenDrops to detect silent connection drops. These counters are system-wide, not per-socket. On multi-service hosts, correlate with per-socket queue depth via ss -ltn or ss -lxn to attribute drops to uWSGI specifically.
  • File descriptor count per worker. Check with ls -1 /proc/<worker_pid>/fd | wc -l against ulimit -n. File descriptor exhaustion causes silent connection failures with no uWSGI-level signal. Usage should stay below 80% of the soft limit.
  • Signal queue depth. Top-level signal_queue (master) and per-worker workers[].signal_queue. Should be 0 in steady state. Any sustained non-zero value means internal uWSGI signals (timers, file monitors, custom signals) are backing up.
  • Accepting worker count (cheaper-aware). When the cheaper subsystem is active, worker count fluctuates by design. Alert when accepting workers drop below the cheaper minimum, not against a fixed expected count. Workers with status: "cheap" have pid: 0 and are intentionally scaled down.

What Level 3 catches: root cause identification during incidents, which endpoint is slow, whether memory growth is systemic or per-worker, spooler or emperor failures, silent connection drops at the kernel level, file descriptor leaks.

What Level 3 misses: deep correlations between uWSGI metrics and downstream dependencies, precise per-endpoint failure attribution, copy-on-write memory accounting, predictive indicators.

Level 4: expert

Goal: predict failures and correlate uWSGI signals with system-level and downstream signals. Level 4 signals are typically added after the second or third major incident reveals a gap.

  • Per-endpoint harakiri attribution. Identify which endpoints are timing out. The stats server does not expose per-endpoint harakiri counts natively, and the metrics subsystem exposes only per-worker and per-core counters, not per-endpoint breakdowns. Operators typically derive this by correlating harakiri log entries with access logs.
  • Worker age skew. Track workers[].last_spawn timestamps to detect asymmetric recycling. If one worker respawns far more frequently than others, it may be hitting a specific code path that causes crashes or memory bloat.
  • Connection-pool correlation. Correlate worker busy ratio and response time with downstream connection pool utilization (database, Redis, external APIs). uWSGI does not natively expose database connection pool metrics, and no bundled plugin correlates them with worker stats. This requires external instrumentation on the downstream systems, displayed alongside uWSGI metrics on the same timeline.
  • PSS divergence. RSS over-reports per-worker usage because shared pages (shared libraries, copy-on-write pages) are counted fully for each process. PSS (Proportional Set Size) distributes shared pages proportionally, giving better accounting. uWSGI’s stats server provides RSS and VSZ but not PSS. PSS requires reading /proc/<pid>/smaps_rollup externally. Track how quickly workers diverge from the master post-fork to measure copy-on-write efficiency over the worker’s lifetime.
  • Stuck request age detection. When a core has in_request == 1, compute current_time - workers[].cores[].req_info.request_start to get the elapsed time of the in-flight request (req_info.request_start is a UNIX timestamp, present in current 2.0.x while the core is in request; suppressed by --stats-no-cores). This detects stuck requests before they trigger harakiri, especially important when harakiri is not configured. The vars field in cores can expose request headers and URI of the in-flight request, useful for diagnosing which endpoint is stuck.
  • External socket queue monitoring. Use ss -ltn 'sport = :PORT' for TCP or ss -lxn for UNIX sockets. Check Recv-Q for current queue depth and Send-Q for the configured backlog limit. This compensates for the unreliable listen_queue stats field and provides the real saturation picture.
  • Running time per request. Compute workers[].running_time / workers[].requests for the true cumulative average processing time per request. This is more stable than avg_rt for trend analysis, though it does not reflect recent changes as quickly. Note that running_time is not reset on respawn in the 2.0.x source (it keeps accumulating across respawns), while delta_requests is reset to 0 each time a worker respawns.

What Level 4 catches: predictive failure indicators, downstream dependency correlations, copy-on-write memory accounting, early stuck-request detection before harakiri fires, true per-request latency trends.

Signal reference by level

LevelSignals addedKey question answered
1: SurvivalMaster alive, accepting worker count, throughputIs uWSGI alive?
2: OperationalBusy ratio, harakiri/respawn/exception deltas, avg_rt, per-worker RSSIs it degrading?
3: MaturePer-worker distribution, RSS trends, serving URI, respawn classification, spooler/emperor health, kernel overflows, fd count, signal queueWhat is broken and where?
4: ExpertPer-endpoint harakiri, worker age skew, connection-pool correlation, PSS divergence, stuck request age, external socket queueWhy is it breaking and what fails next?

Configuration prerequisites

Several signals across all levels require explicit configuration beyond --stats:

  • --memory-report: must be enabled for RSS and VSZ values to appear in stats output. Without it, all memory-related signals at Level 2 and above are blind.
  • --harakiri <seconds>: without it, stuck workers have no timeout. harakiri_count is always 0, which looks healthy but is a monitoring blind spot. A single hung request can reduce capacity by one worker indefinitely.
  • --harakiri-verbose: enables logging of the blocked syscall and wchan when harakiri fires (Linux only, reads /proc/<pid>/syscall and /proc/<pid>/wchan). Essential for diagnosing what workers are stuck on.
  • --stats-http: required if you want to access the stats server via HTTP with curl. Without it, the stats server serves raw JSON on a socket and you must use uwsgi --connect-and-read or socat.
  • --thunder-lock: required for uniform request distribution across workers. Without it, the kernel’s accept() thundering herd behavior can cause uneven load that masquerades as a worker-level problem.

How Netdata helps

Netdata collects uWSGI stats server output at per-second resolution, which matters for signals that change quickly during incidents. The correlations that shorten diagnosis time:

  • Worker busy ratio + avg_rt at per-second granularity catches capacity exhaustion before the listen queue fills. At 10-second polling intervals, a complete starvation event can run for up to 10 seconds undetected.
  • Harakiri rate + respawn rate correlation distinguishes a harakiri death spiral (respawns track harakiri 1:1, throughput collapses) from healthy max-requests recycling (steady respawns, zero harakiri).
  • Per-worker RSS trends reveal the sawtooth pattern of memory recycling versus the linear growth of a genuine leak, visible over hours rather than at a single point in time.
  • Kernel TCP overflow counters alongside uWSGI worker metrics surface the silent connection drops that uWSGI’s own listen_queue and listen_queue_errors fields cannot report.
  • Cross-layer correlation places uWSGI worker signals next to system CPU, memory, and network metrics on the same timeline. When response time increases, the correlation with downstream dependency health is immediately visible rather than requiring a separate dashboard.