The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / uwsgi / uwsgi-all-workers-busy ▌

Operations Guides

uWSGI all workers busy: reading the busy ratio before the queue fills

When every uWSGI worker shows status: "busy", the next incoming request does not wait in a place you can see. It lands in the kernel socket backlog, which uWSGI cannot reliably measure on standard Linux. If that backlog fills, the kernel drops connections silently. No log entry, no error counter, no uWSGI-level signal. The application looks alive but stops serving real traffic.

The worker busy ratio is the earliest internal indicator that you are approaching that cliff. It tells you what proportion of your alive, non-cheaped workers are currently processing requests. Reading it correctly requires understanding what “busy” actually means, what it does not mean, and where the visibility gap starts.

What this means

The busy ratio is straightforward to define: count workers where status == "busy" and divide by workers where pid > 0 and status != "cheap". Cheaped workers (scaled down by the cheaper subsystem, pid: 0) are excluded from both numerator and denominator.

At 100%, every alive worker is occupied. The next connection goes to the kernel listen queue (the socket backlog). uWSGI has no visibility into that queue on standard Linux: the listen_queue and load stats fields are unreliable. The listen_queue field relies on TCP_INFO behavior that varies across kernel versions, and the UNIX socket measurement requires a non-standard kernel ioctl. The load field mirrors listen_queue, not an average latency metric as its name suggests. The listen_queue_errors field exists in the JSON but is never incremented anywhere in the uWSGI 2.0.x source (it is only reported from a field nothing writes), so it reads 0 on current releases. Both fields are almost always 0 regardless of actual backlog depth.

This is the core problem: the moment you need visibility the most (all workers busy, queue filling), uWSGI goes dark. You must measure the kernel backlog externally with ss.

flowchart LR
    A[Request arrives] --> B{Worker available?}
    B -- yes --> C[Worker accepts
status = busy] C --> D[Request processed] D --> E[Worker returns to idle] B -- no --> F[KERNEL backlog
no uWSGI visibility] F --> G{Backlog full?} G -- no --> H[Request queues
latency increases] G -- yes --> I[Connection dropped
TCP RST or refused] H --> B

What “busy” actually means

“Busy” means the worker is not accepting new connections. It does not mean the worker is doing CPU work. A worker blocked on a database query, waiting on an external HTTP call, or sleeping in a blocking I/O call counts as busy. This is the single most common misinterpretation.

In threaded mode, “busy” means at least one thread is active. It does not mean all threads are occupied. You need per-core in_request data for thread-level concurrency visibility.

In async mode (gevent or asyncio), “busy” means the event loop is running. A worker is rarely idle because it multiplexes many concurrent requests. The busy ratio is nearly meaningless in this mode. If you are running gevent, track per-core request counts and greenlet-level concurrency instead.

Common causes

CauseWhat it looks likeFirst thing to check
Downstream dependency slowdownAll workers busy, avg_rt rising, throughput fallingIn-flight request URI from busy cores (cores[].vars), then check the downstream service (DB, cache, API)
Insufficient worker countBusy ratio sustained above 80% during normal traffic, listen queue nonzero at peaksCompare worker count against traffic volume and avg_rt
Traffic spikeBusy ratio jumps to 100% suddenly, throughput still high or increasingCompare incoming request rate against recent baseline
Stuck workers without harakiriWorkers stuck in busy indefinitely, busy ratio climbs and never recoversCheck if harakiri is configured. Check per-core in_request and request age
Single slow endpointAll busy workers show the same URIIn-flight request URI from each busy worker’s cores[].vars

Quick checks

These commands read from the uWSGI stats server. Adjust the socket address to match your deployment. The stats server serves raw JSON on a socket by default. Use uwsgi --connect-and-read for TCP sockets and socat - UNIX-CONNECT:<path> for UNIX sockets. If --stats-http is enabled, curl http://<addr> also works.

# Compute the busy ratio (alive, non-cheaped workers only)
uwsgi --connect-and-read 127.0.0.1:9191 | jq \
  '([.workers[] | select(.status == "busy")] | length) as $busy
   | ([.workers[] | select(.pid > 0 and .status != "cheap")] | length) as $alive
   | if $alive > 0 then ($busy / $alive * 100) else 0 end'

# List each worker's status and PID (in-flight request URI lives in cores[].vars)
uwsgi --connect-and-read 127.0.0.1:9191 | jq \
  '.workers[] | select(.pid > 0) | {id, status, pid, avg_rt}'

# Count accepting workers (alive, non-cheaped, accepting == 1)
uwsgi --connect-and-read 127.0.0.1:9191 | jq \
  '[.workers[] | select(.pid > 0 and .status != "cheap" and .accepting == 1)] | length'

# Check avg_rt across alive workers
# <!-- TODO: verify whether avg_rt is reported in microseconds or milliseconds -->
uwsgi --connect-and-read 127.0.0.1:9191 | jq \
  '[.workers[] | select(.pid > 0 and .status != "cheap")]
   | if length > 0 then (map(.avg_rt) | add / length) else 0 end'

# Check harakiri count (monotonic, never reset even on respawn)
uwsgi --connect-and-read 127.0.0.1:9191 | jq \
  '[.workers[].harakiri_count] | add'

# Check request age for in-flight requests (requires cores, suppressed by --stats-no-cores)
# <!-- TODO: verify req_info.request_start field name and timestamp format across uWSGI versions -->
uwsgi --connect-and-read 127.0.0.1:9191 | jq \
  --argjson now "$(date +%s)" \
  '[.workers[] | select(.pid > 0) | .id as $wid
   | .cores[] | select(.in_request == 1)
   | {worker: $wid, core: .id, age_seconds: ($now - .req_info.request_start)}]'

# Measure kernel socket backlog externally (TCP)
ss -ltn 'sport = :8000'

# Measure kernel socket backlog externally (UNIX socket)
ss -lxn | grep uwsgi

The ss output shows Recv-Q (current queue depth) and Send-Q (backlog limit). Recv-Q should be 0 in steady state. Any sustained non-zero value means workers cannot accept fast enough.

How to diagnose it

  1. Confirm the busy ratio is sustained, not a brief spike. Brief spikes to 100% during traffic bursts are normal if they self-resolve within seconds. Poll at 1-second intervals if possible. At 10-second polling, a complete starvation event can run for up to 10 seconds before detection.

  2. Check what workers are busy on. Read the in-flight request URI from each busy worker’s cores[].vars (the request variables, populated only while in_request == 1). If all busy workers show the same URI, a single endpoint is the bottleneck. If URIs are diverse, the problem is systemic (downstream dependency, resource contention).

  3. Check the kernel listen queue externally. Do not trust listen_queue or load in the stats JSON. Use ss -ltn for TCP or ss -lxn for UNIX sockets. If Recv-Q is non-zero and growing, the kernel backlog is filling.

  4. Correlate with avg_rt. Rising avg_rt with stable or falling throughput means each request takes longer. The avg_rt field is an exponential moving average computed as (old_avg_rt + current_request_time) / 2, giving roughly 50% weight to the most recent request. This makes it responsive to recent changes but volatile. Compare it against your configured harakiri timeout. If avg_rt approaches harakiri, workers will start dying.

  5. Check harakiri rate. Sum harakiri_count across all workers and track the delta between polling intervals. If harakiri is rising and respawn rate tracks it closely, you may be in a death spiral where every respawned worker immediately gets stuck again. If harakiri is not configured at all, harakiri_count is always 0 and stuck workers have no timeout. The absence of harakiri configuration is itself a risk.

  6. Check accepting worker count. A worker can have status: "idle" but accepting: 0, meaning it will not take new requests. This can happen during graceful reloads. Count workers where pid > 0 and status != "cheap" and accepting == 1. If this count is zero while the master is alive, the service cannot accept any new requests.

  7. Check kernel-level overflow counters. Use nstat -az TcpExtListenOverflows TcpExtListenDrops to see if the kernel has been dropping connections. These are cumulative system-wide counters since boot, not per-socket. Run the command twice with an interval between to measure the rate of change. Correlate with per-socket listen queue depth to attribute drops to uWSGI.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Worker busy ratioPrimary concurrency utilization. At 100%, next request queues in kernel backlog with no uWSGI visibilitySustained above 80% is limited headroom. 100% sustained is starvation
Accepting worker countHow many workers can serve requests right now. A worker can be idle but not acceptingZero accepting workers with master alive is critical
avg_rt (per worker)Application responsiveness. EMA gives 50% weight to last requestApproaching harakiri timeout means workers will start dying
Harakiri rate (delta)Requests exceeding timeout and being killed. Each harakiri is a dropped request and a respawnAny sustained non-zero rate in a normally-zero deployment
Kernel listen queue (Recv-Q from ss)External measurement of backlog depth, since uWSGI’s internal measurement is brokenAny sustained non-zero value means workers cannot keep up
TcpExtListenOverflowsKernel-level confirmation that connections are being droppedAny non-zero rate of change means users see connection failures
Request throughput (delta requests)Throughput trend. Drop with no traffic decrease means workers are stuckSudden drop from baseline
Stuck request ageElapsed time of in-flight requests, from per-core req_info.request_startRequest age approaching harakiri, or exceeding 60s if harakiri is not configured

Fixes

Downstream dependency slowdown

If all workers are busy and avg_rt is rising, the root cause is almost always downstream. Check database connection counts, query latency, cache hit rates, and external API response times. Correlate uWSGI worker metrics with downstream dependency health on the same dashboard.

If a specific endpoint is identified, consider blocking or rate-limiting it at the load balancer to free workers for other traffic. At the application level, fail fast: return a 503 immediately instead of waiting for a downstream timeout.

Insufficient worker count

If the busy ratio is consistently above 80% during normal traffic with no downstream issue, you need more workers or faster request processing. Adding workers increases memory consumption (each worker is a full process copy). Verify that total worker RSS stays below 70-80% of system RAM, accounting for copy-on-write sharing.

If you use the cheaper subsystem, ensure cheaper minimum is at least 20% of maximum workers to maintain headroom.

Stuck workers without harakiri

If harakiri is not configured, a single hung request permanently consumes a worker slot. Workers accumulate in stuck state with no recovery. Configure --harakiri at 2-3x your expected maximum legitimate request duration, and enable --harakiri-verbose for diagnostic backtraces of the blocked syscall.

Listen backlog too small

The default --listen is 100 and the Linux kernel default net.core.somaxconn is 128 on older kernels (4096 since kernel 5.4). For high-traffic services, a brief downstream hiccup can fill this in under a second. Increase both --listen and net.core.somaxconn to absorb brief spikes without dropping connections.

Prevention

  • Track busy ratio over time, not just point-in-time. A ratio that creeps from 60% to 80% over weeks is a capacity planning signal.
  • Maintain at least 20% idle workers under normal traffic. The degradation curve is cliff-edge: there is no graceful degradation between “keeping up” and “dropping connections.”
  • Always configure harakiri. Without it, stuck workers are permanent. The absence of harakiri configuration is a monitoring blind spot.
  • Monitor the kernel listen queue externally with ss. Do not rely on uWSGI’s listen_queue or load stats fields.
  • Monitor TcpExtListenOverflows at the kernel level for dropped-connection confirmation.
  • Use chain reload (--chain-reload) for deployments. Standard graceful reload can briefly reduce capacity to near-zero during worker replacement, since all old workers drain before new workers finish loading; chain reload cycles workers one at a time to keep partial capacity.
  • Do not use the stats endpoint as a health check. The stats server is served by the master process and remains responsive during complete worker starvation. Health checks must go through the worker pool.

How Netdata helps

  • Per-second busy ratio collection from the uWSGI stats server, computed as busy workers over alive non-cheaped workers. Saturation events shorter than 10 seconds are invisible at typical polling intervals.
  • Correlated timelines of busy ratio, avg_rt, harakiri rate, and request throughput on a single dashboard. When all four move together, the composite pattern points at capacity exhaustion, death spiral, or downstream failure.
  • Cheaper-aware accepting worker count, distinguishing expected cheaped-down workers from workers that stopped accepting.
  • Anomaly detection on busy ratio and related signals, useful where static thresholds produce false positives on small instances or during batch processing.
  • External socket queue depth via ss-based collection or eBPF, providing visibility into the kernel backlog where uWSGI’s own stats go dark.